Image processing methods, apparatuses, electronic devices, storage media, and software products
By applying multiple processing strategies and adjusting image feature parameters, a target image with superior performance to the original image was generated, solving the problem of low image resolution in zoom processing and improving image performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2024-11-25
- Publication Date
- 2026-05-26
AI Technical Summary
When zooming an image based on a large zoom ratio, the processed image may lose detail information, resulting in low image resolution and poor performance.
The first image is processed using at least two image processing strategies to generate at least two second images. A target image is generated by adjusting the baseline image and the reference image. The index parameters of the target image are greater than those of the second images.
It improves the resolution and similarity of images, generates target images with better performance than the original images, and solves the problem of low resolution in zoom processing.
Smart Images

Figure CN122093671A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing, and more particularly to an image processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] In related technologies, fixed-focus cameras can be used to shoot in different shooting scenarios, and the zoom ratio of the captured image can be adjusted using digital zoom. However, when zooming the image based on a large zoom ratio, the processed image may lose detail information, resulting in low image resolution. Thus, the performance of the processed image is poor. Summary of the Invention
[0003] To overcome the problems in related technologies, this disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product to improve the performance of the image obtained after processing a first image.
[0004] According to a first aspect of the present disclosure, an image processing method is provided, the method comprising:
[0005] The first image is processed using at least two image processing strategies to obtain at least two second images. The second images obtained using different image processing strategies have different index parameters, and the index parameters of each second image are greater than those of the first image. The index parameters are positively correlated with the performance of the images.
[0006] The target image is obtained based on at least two second images and the index parameters of at least two second images;
[0007] The index parameters of the target image are greater than those of the second image.
[0008] In one embodiment, obtaining the target image based on at least two second images and index parameters of at least two second images includes:
[0009] The second image whose index parameters meet the first index condition is determined as the reference image; wherein, the index condition is used to indicate that: the resolution of the second image is greater than a preset resolution threshold, or the similarity between the first image and the second image is greater than a preset similarity threshold.
[0010] The second image whose index parameters meet the second index condition is determined as the reference image; wherein the first index condition and the second index condition are different.
[0011] The target image is obtained by adjusting the reference image based on the reference image.
[0012] In one embodiment, the first image is processed based on at least two image processing strategies to obtain at least two second images, including:
[0013] At least two sets of image feature parameters are extracted from the first image, and / or the image feature parameters extracted from the first image are compressed based on at least two compression ratios to obtain at least two sets of image feature parameters; wherein the image processing strategies corresponding to different sets of image feature parameters are different.
[0014] Based on the image processing strategies corresponding to the feature parameters of each group of images, the first image is processed to obtain at least two second images.
[0015] In one embodiment, the indicator parameters include the image resolution and the similarity between the first image and the second image; the resolution of each second image is greater than that of the first image; based on the image processing strategy corresponding to each set of image feature parameters, the first image is processed to obtain at least two second images, including:
[0016] Based on the feature parameters of the first image and the descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a preset resolution threshold; wherein, the descriptive information is determined based on the type of each object in the first image and / or the relationship between each object.
[0017] Based on the second image feature parameters, the first image is processed to obtain a second image with a similarity greater than a preset similarity threshold;
[0018] Wherein, the first image feature parameter includes global features of the first image, the second image feature parameter includes local features of the second image; and / or, the data volume of the first image feature parameter is less than the data volume of the second image feature parameter.
[0019] In one embodiment, the method further includes:
[0020] Based on the type of each object in the first image, determine the processing parameters of the first image region where each object is located;
[0021] Based on the feature parameters and descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a resolution threshold, including:
[0022] Based on the processing parameters of each first image region, the feature parameters of the first image, and the descriptive information of the first image, each first image region of the first image is processed to obtain a second image with a resolution greater than the resolution threshold.
[0023] In one embodiment, determining the processing parameters of the first image region where each object is located, based on the type of each object in the first image, includes:
[0024] If the first image includes an object of a first target type, determine the pose of the object of the first target type;
[0025] Based on the pose of the first target type, determine the processing parameters of the first image region where the object of the first target type is located;
[0026] There is a correspondence between the pose and processing parameters of the first target type.
[0027] In one embodiment, the method further includes:
[0028] Based on the types of each object in the first image and the relationships between each object, the first descriptive information is determined;
[0029] In response to the detection of a preset input operation, the first description information is adjusted based on the preset input operation to obtain the second description information;
[0030] Based on the feature parameters and descriptive information of the first image, the first image is processed, including:
[0031] The first image is processed based on the first image feature parameters and the second descriptive information.
[0032] In one embodiment, the method further includes:
[0033] The associations in the first description information that do not meet the preset description conditions are deleted to obtain the processed first description information;
[0034] The first description information is adjusted based on a preset input operation to obtain the second description information, including:
[0035] The processed first description information is adjusted based on the preset input operation to obtain the second description information.
[0036] In one embodiment, adjusting a reference image based on a reference image to obtain a target image includes:
[0037] Based on the type of object in each first image region of the first image, select the region to be adjusted from each second image region in the reference image that matches the first image region;
[0038] Based on a third image region in the reference image that matches the region to be adjusted, the region to be adjusted in the baseline image is adjusted to obtain the target image.
[0039] In one embodiment, a first indicator condition is used to indicate that the resolution of the reference image is greater than a resolution threshold. Based on the type of objects in each first image region of the first image, selecting regions to be adjusted from each second image region in the reference image that matches the first image region includes:
[0040] For each region of the first image, a second target type is determined from a preset type based on the type of the object in the first image region; wherein there is a correspondence between the object type and the preset type;
[0041] If a first image region of a first image matches a second image region of a reference image, determine the image content of a second target type in the first image region and the similarity between the image content of the second target type in the second image region.
[0042] The second image region in the baseline image with a similarity less than the similarity threshold is identified as the region to be adjusted.
[0043] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising:
[0044] The processing module is configured to process the first image based on at least two image processing strategies to obtain at least two second images; wherein the index parameters of the second images obtained based on different image processing strategies are different, and the index parameters of each second image are greater than the index parameters of the first image, and the index parameters are positively correlated with the performance of the image.
[0045] The acquisition module is configured to obtain the target image based on at least two second images and the index parameters of at least two second images;
[0046] The index parameters of the target image are greater than those of the second image.
[0047] In one embodiment, the acquisition module includes:
[0048] The determination module is configured to determine the second image that meets the first indicator condition as the reference image; wherein the indicator condition is used to indicate that: the resolution of the second image is greater than a preset resolution threshold, or the similarity between the first image and the second image is greater than a preset similarity threshold.
[0049] The second image whose index parameters meet the second index condition is determined as the reference image; wherein the first index condition and the second index condition are different.
[0050] The adjustment module is configured to adjust the reference image based on the reference image to obtain the target image.
[0051] In one embodiment, the processing module is further configured as follows:
[0052] At least two sets of image feature parameters are extracted from the first image, and / or the image feature parameters extracted from the first image are compressed based on at least two compression ratios to obtain at least two sets of image feature parameters; wherein the image processing strategies corresponding to different sets of image feature parameters are different.
[0053] Based on the image processing strategies corresponding to the feature parameters of each group of images, the first image is processed to obtain at least two second images.
[0054] In one embodiment, the indicator parameters include the image resolution and the similarity between the first image and the second image; the resolution of each second image is greater than the resolution of the first image; the processing module is further configured to:
[0055] Based on the feature parameters of the first image and the descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a preset resolution threshold; wherein, the descriptive information is determined based on the type of each object in the first image and / or the relationship between each object.
[0056] Based on the second image feature parameters, the first image is processed to obtain a second image with a similarity greater than a preset similarity threshold;
[0057] Wherein, the first image feature parameter includes global features of the first image, the second image feature parameter includes local features of the second image; and / or, the data volume of the first image feature parameter is less than the data volume of the second image feature parameter.
[0058] In one embodiment, the determining module is further configured as follows:
[0059] Based on the type of each object in the first image, determine the processing parameters of the first image region where each object is located;
[0060] The processing module is also configured as follows:
[0061] Based on the processing parameters of each first image region, the feature parameters of the first image, and the descriptive information of the first image, each first image region of the first image is processed to obtain a second image with a resolution greater than the resolution threshold.
[0062] In one embodiment, the determining module is further configured as follows:
[0063] If the first image includes an object of a first target type, determine the pose of the object of the first target type;
[0064] Based on the pose of the first target type, determine the processing parameters of the first image region where the object of the first target type is located;
[0065] There is a correspondence between the pose and processing parameters of the first target type.
[0066] In one embodiment, the determining module is further configured as follows:
[0067] Based on the types of each object in the first image and the relationships between each object, the first descriptive information is determined;
[0068] In response to the detection of a preset input operation, the first description information is adjusted based on the preset input operation to obtain the second description information;
[0069] The processing module is also configured as follows:
[0070] The first image is processed based on the first image feature parameters and the second descriptive information.
[0071] In one embodiment, the processing module is further configured as follows:
[0072] The associations in the first description information that do not meet the preset description conditions are deleted to obtain the processed first description information;
[0073] The adjustment module is also configured as follows:
[0074] The processed first description information is adjusted based on the preset input operation to obtain the second description information.
[0075] In one embodiment, the determining module is further configured as follows:
[0076] Based on the type of object in each first image region of the first image, select the region to be adjusted from each second image region in the reference image that matches the first image region;
[0077] The adjustment module is also configured to adjust the region to be adjusted in the reference image based on a third image region in the reference image that matches the region to be adjusted, so as to obtain the target image.
[0078] In one embodiment, the first indicator condition is used to indicate that the resolution of the reference image is greater than a resolution threshold, and the determining module is further configured to:
[0079] For each region of the first image, a second target type is determined from a preset type based on the type of the object in the first image region; wherein there is a correspondence between the object type and the preset type;
[0080] If a first image region of a first image matches a second image region of a reference image, determine the image content of a second target type in the first image region and the similarity between the image content of the second target type in the second image region.
[0081] The second image region in the baseline image with a similarity less than the similarity threshold is identified as the region to be adjusted.
[0082] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0083] processor;
[0084] Memory used to store computer programs or instructions;
[0085] The processor executes computer programs or instructions to implement the steps in any of the image processing methods in the first aspect described above.
[0086] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, comprising:
[0087] When a computer program or instruction in a storage medium is executed by a processor, the steps in any of the image processing methods in the first aspect described above are implemented.
[0088] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program or instructions, which, when executed by a processor, implement the steps of any of the image processing methods in the first aspect described above.
[0089] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0090] In this embodiment, the first image can be processed using different image processing strategies to obtain second images with different index parameters and superior performance compared to the first image. Furthermore, a target image with superior performance compared to each of the first images can be generated jointly based on at least two second images with superior performance and their index parameters. This effectively improves the performance of the first image and yields a second image with better performance.
[0091] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0092] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0093] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0094] Figure 2 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0095] Figure 3 This is a schematic diagram illustrating a process for extracting descriptive information according to an exemplary embodiment.
[0096] Figure 4 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0097] Figure 5 This is a schematic diagram illustrating a process for acquiring a reference image according to an exemplary embodiment.
[0098] Figure 6 This is a timing diagram illustrating an image processing method according to an exemplary embodiment.
[0099] Figure 7 This is a structural block of an image processing apparatus according to an exemplary embodiment.
[0100] Figure 8 This is a structural block diagram of an electronic device according to an exemplary embodiment.
[0101] Figure 9 This is a block diagram of an apparatus according to an exemplary embodiment. Detailed Implementation
[0102] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0103] The image processing method shown in this embodiment can be applied to electronic devices with image processing capabilities. Here, the electronic device can include a mobile electronic device or a fixed electronic device. The mobile electronic device can include devices such as mobile phones, tablets, laptops, and in-vehicle electronic devices. The fixed electronic device can include desktop computers, smart TVs, etc. In some embodiments, the operating system of the electronic device can include an Input Output System (IOS) operating system, an Android operating system, etc.
[0104] It should be noted that electronic devices may include, but are not limited to, mobile communication electronic devices, portable entertainment devices, wearable devices, home appliances, augmented reality devices, virtual reality devices, and special-purpose devices. Among these, mobile communication electronic devices may include, but are not limited to, mobile phones, tablets, and smartwatches; portable entertainment devices may include, but are not limited to, digital cameras; wearable devices may include, but are not limited to, smart bracelets and smart glasses; home appliances may include, but are not limited to, televisions and video recorders; augmented reality (AR) devices may include, but are not limited to, AR glasses; virtual reality (VR) devices may include, but are not limited to, VR glasses; and special-purpose devices may include, but are not limited to, professional cameras (such as SLR cameras and point-and-shoot cameras).
[0105] It should be noted that the execution entity of the embodiments of this disclosure can be the central processing unit (CPU) in an electronic device in terms of hardware, and can be, for example, a related background service or application in an electronic device in terms of software, without limitation.
[0106] To better understand the technical solutions in the embodiments of this disclosure, the image processing methods in related technologies are described by way of example below:
[0107] Due to limitations in camera module size and cost, electronic devices generally do not use lenses with exceptionally high zoom ratios. However, some high-end electronic devices can use super-resolution algorithms to process images with high zoom ratios, improving image clarity. In this case, some high-end electronic devices can support 120x digital zoom.
[0108] After zooming the original image based on the zoom ratio, traditional convolutional neural networks can be used to perform super-resolution processing on the zoomed image to improve image sharpness. However, the ability of general convolutional neural networks to generate ultra-high-resolution images using super-resolution methods is relatively weak; they can only supplement some texture on the existing texture of the original image. Furthermore, for ultra-high zoom scenes, the super-resolution images obtained by general super-resolution algorithms exhibit obvious smearing and distortion, severely impacting the user experience. For example, telephoto photography is one of the main scenarios where mobile phone photography cannot completely replace camera photography, and users have high demands for this scenario, requiring high zoom ratios for image processing. When using traditional convolutional neural networks to generate ultra-high-resolution images in telephoto photography scenarios, the generated images often suffer from poor realism, leading to a poor user experience.
[0109] For example, ultra-high-definition algorithms based on convolutional neural networks can be implemented using Real Enhanced Super-Resolution Generative Adversarial Networks (GANs). This approach introduces a discriminator, where the generator and discriminator compete and cooperate against each other during training. The generator learns the distribution characteristics of the training data, aiming to generate images that can fool the discriminator and achieve a high degree of realism. The discriminator, on the other hand, aims to accurately distinguish between real and generated samples, maximizing the difference between them and feeding back a signal to the generator. RealESRGAN constructs data pairs by simulating a series of degradation processes, training a generator with super-resolution capabilities, and can achieve good generated ultra-high-definition effects on real images.
[0110] While RealSRGAN can remove noise and mitigate some image degradation, it struggles to improve texture loss and blurring in severely degraded telephoto scenes. RealSRGAN's generative capabilities are relatively poor; it can only enhance existing textures and cannot generate new, sharp textures.
[0111] Furthermore, current popular ultra-high-definition (UHD) image generation algorithms also employ Artificial Intelligence Generated Content (AIGC) technology, which can generate detailed high-definition images. AIGC methods use large models trained on massive amounts of data, enabling the model to generate corresponding image content based on prompts. Because the prompt information is highly compressed, one descriptive word typically corresponds to one target, thus AIGC models possess extremely strong generation capabilities. When users see low-quality images, they can often associate them with what a clear image would look like; the most intuitive example is that users can understand some incomplete or blurry text. If large models can fully integrate the image information of the input image into the generated image, they have the potential to produce high-definition images that meet user needs.
[0112] While AIGC technology can produce high-definition and detailed images, its fidelity is poor in some scenes, easily resulting in noticeable texture errors. For photography scenarios, it is necessary to ensure consistency between the generated image and the real scene. At the same time, mobile devices have high limitations on power consumption and latency; power consumption cannot be too high, and the shooting time cannot be too long.
[0113] Based on this, this disclosure proposes a lightweight ultra-high-resolution large-scale model that can be applied to telephoto photography scenarios. It can improve image clarity and enrich image texture while maintaining the realism of the input image. It should be noted that this invention can also be used in general photography scenarios or for enhancing album images.
[0114] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 1 As shown, the method includes:
[0115] Step 11: Process the first image based on at least two image processing strategies to obtain at least two second images; wherein the index parameters of the second images obtained based on different image processing strategies are different, and the index parameters of each second image are greater than the index parameters of the first image, and the index parameters are positively correlated with the performance of the image.
[0116] In one embodiment, the image processing strategy can be a strategy for processing the first image based on an image generation model. Here, processing the first image based on at least two image processing strategies can be understood as processing the first image based on at least two image generation models. Exemplarily, the image generation model is a trained model for image enhancement. For example, the image generation model may include, but is not limited to, at least one of the following: an AIGC model, a GAN model, and a diffusion model.
[0117] In one embodiment, there is a correspondence between an image processing strategy and the index parameters of a second image obtained based on that image processing strategy. For example, different image processing strategies may correspond to different index parameters for the second image.
[0118] In one embodiment, an image may correspond to at least one type of index parameter. Each type of index parameter indicates the image's performance at different latitudes.
[0119] In one embodiment, the image's metrics parameters include at least one of the following: image resolution, color depth, and signal-to-noise ratio (SNR). It should be noted that image resolution indicates the richness of detail in the image, color depth indicates the richness of colors, and SNR indicates the proportion of noise in the image. Here, image resolution, color depth, and SNR can accurately indicate the quality of an image.
[0120] In one embodiment, the indicator parameters of the image may further include: the similarity between the image and the first image. For example, the indicator parameters of the second image may include: the similarity between the second image and the first image.
[0121] It should be noted that the index parameters of each second image are greater than those of the first image. This can be understood as at least one of the resolution, color depth, and signal-to-noise ratio of each second image being greater than at least one of the resolution, color depth, and signal-to-noise ratio of the first image.
[0122] In one embodiment, the index parameters of each second image are all greater than the index parameters of the first image. This can be understood as at least one type of index parameter of each second image being greater than the index parameters of the first image. For example, the second image may include a second image A and a second image B. The resolution of the second image A may be greater than the resolution of the first image, and the color depth of the second image B may be greater than the color depth of the first image.
[0123] In one embodiment, the index parameters of each second image are all greater than the index parameters of the first image. This can be understood as, for the same type of index parameters, the index parameters of each second image are all greater than the index parameters of the first image. For example, the resolution of each second image may be greater than the resolution of the first image. Alternatively, the signal-to-noise ratio (SNR) of each second image may be greater than the SNR of the first image.
[0124] In one embodiment, image index parameters can be determined based on image feature parameters. These image feature parameters may include at least one of the following: texture feature parameters, semantic feature parameters, color feature parameters, contour feature parameters, contrast, brightness, and saturation. For example, the similarity between a first image and a second image can be obtained by comparing at least one type of image feature parameters of the first image and the second image.
[0125] In one embodiment, the resolution of each second image can be negatively correlated with the similarity between the second image and the first image. It should be noted that during the process of processing the first image based on image processing parameters to improve its resolution, detailed information may be added to the first image. However, this added detailed information may contain errors, resulting in a significant difference between the second image with added detailed information and the initial first image. In this case, the higher the resolution of the second image, the greater the difference between the second image and the first image may be, leading to a potentially lower similarity between the first and second images.
[0126] Step 12: Obtain the target image based on at least two second images and the index parameters of at least two second images;
[0127] The index parameters of the target image are greater than those of the second image.
[0128] In one embodiment, obtaining a target image based on at least two second images and their index parameters includes: fusing the at least two second images based on their index parameters to obtain the target image.
[0129] In one embodiment, the index parameters of the second image can be positively correlated with the fusion ratio of the second image. The fusion ratio of the second image can be understood as the proportion of each second image in the process of fusion processing at least two second images.
[0130] In one embodiment, for each second image, weighted index parameters of the second image can be weighted based on weight parameters to obtain weighted index parameters of the second image; wherein, the weight parameters include: weight values corresponding to each index parameter respectively; based on the weighted index parameters corresponding to each second image, at least two second images are fused. The weighted index parameters of the second image can be positively correlated with the fusion ratio corresponding to the second image. It should be noted that the weighted processing of each index parameter of the second image based on the weight parameters can be based on the weight values of the weight parameters, and weighting the index parameters corresponding to the weight values separately. Here, the fusion ratio corresponding to each second image can be accurately determined based on the weighted index parameters corresponding to each second image during the fusion processing of at least two second images. For example, when the number of second images is greater than two, there may be a second image whose fusion ratio is zero. That is, there may be at least one second image that is not fused into the target image.
[0131] In one embodiment, fusing at least two second images based on index parameters to obtain a target image includes: determining at least two images to be fused from the at least two second images based on index parameters; and fusing the at least two images to be fused to obtain the target image. The index parameters of the images to be fused may be greater than a preset first parameter threshold. Here, the second image with better performance can be selected from the at least two second images for fusion processing, and not all second images can be fused. This ensures that the performance of the target image obtained after fusion processing is better than that of the second images while reducing the speed of fusion processing of the second images.
[0132] In one embodiment, at least two images to be fused can be determined from at least two second images based on at least two types of index parameters of each second image; wherein at least one type of index parameter of the images to be fused can be greater than a parameter threshold. The types of index parameters greater than the parameter threshold corresponding to different images to be fused can be different. In this way, the second images with better performance in at least one dimension can be combined to generate a target image with better performance in all dimensions.
[0133] In one embodiment, a base image and a reference image can be determined from at least two second images based on index parameters of at least two second images; the base image is then adjusted based on the reference image to obtain the target image.
[0134] It should be noted that any reference image and / or reference image in this disclosure can be an image determined from the images to be fused. The process of adjusting the reference image based on the reference image can be understood as the process of fusing the reference image and the reference image.
[0135] In one embodiment, for each second image, weighted index parameters of the second image can be weighted based on weight parameters to obtain weighted index parameters of the second image; wherein the weight parameters include weight values corresponding to each index parameter; a reference image is determined based on the weighted index parameters corresponding to each second image; wherein the weighted index parameters of the reference image are greater than a preset second parameter threshold. For example, the weighted index parameters of the reference image are greater than the weighted index parameters of at least two second images other than the reference image.
[0136] It should be noted that any second parameter threshold of this disclosure may be greater than any first parameter threshold of this disclosure.
[0137] Here, the second image with the best overall performance can be designated as the reference image. This simplifies the process of adjusting the reference image to improve its performance. Thus, the speed of adjusting the reference image is increased, as is the performance of the target head image obtained after adjusting the reference image.
[0138] In one embodiment, a reference image is determined based on various index parameters of at least two images. The reference image has at least one index parameter that is greater than the index parameter of the baseline image.
[0139] In one embodiment, a reference image is determined from at least two second images based on a first type of index parameter; wherein the first type of index parameter of the reference image is greater than a preset third parameter threshold. A reference image is determined from at least two second images based on a second type of index parameter; wherein the second type of index parameter of the reference image is greater than a preset fourth parameter threshold.
[0140] In one embodiment, the first type of index parameter of the reference image is greater than the first type of index parameter of the reference image. The second type of index parameter of the reference image is less than the second type of index parameter of the reference image.
[0141] In one embodiment, a second image whose index parameters meet a first index condition is determined as a baseline image; wherein, the index condition can be used to indicate that at least one type of index parameter is greater than a preset parameter threshold; a second image whose index parameters meet a second index condition is determined as a reference image. Different index conditions may correspond to different index parameters. The baseline image is adjusted based on the reference image to obtain the target image.
[0142] For example, a second image whose first type of indicator parameters meet a first indicator condition can be determined as a baseline image; wherein, the first indicator condition is used to indicate that the first type of indicator parameters are greater than a preset parameter threshold; a second image whose second type of indicator parameters meet a second indicator condition can be determined as a reference image. The second indicator condition is used to indicate that the second type of indicator parameters are greater than a preset parameter threshold; the baseline image is adjusted based on the reference image to obtain the target image.
[0143] In one embodiment, different indicator conditions correspond to different numbers of indicator parameters. For example, a first indicator condition may be used to indicate that a first number of indicator parameters are greater than a preset parameter threshold, and a second indicator condition may be used to indicate that a second number of indicator parameters are greater than the preset parameter threshold. The first number and the second number may be different.
[0144] It should be noted that different indicator parameters may have different threshold values.
[0145] For example, the first indicator condition can be used to indicate that the resolution of the second image is greater than a preset resolution threshold, the second indicator condition can be used to indicate that the color depth of the second image is greater than a preset depth threshold, and the similarity of the second image reaches a preset similarity threshold. In this case, the first number of indicator parameters corresponding to the first indicator condition can be 1, and the second number of indicator parameters corresponding to the first indicator condition can be 2.
[0146] Here, the second image may include a second image C and a second image D. The resolution of the second image C may be greater than a preset resolution threshold, and the similarity of the second image C may be less than a preset similarity threshold. The color depth of the second image may be less than a preset depth threshold. The resolution of the second image D may be less than a preset resolution threshold, and the similarity of the second image C may be greater than a preset similarity threshold. The color depth of the second image may be greater than a preset depth threshold. At this time, the index parameters of the second image C satisfy the first index condition, and the index parameters of the second image D satisfy the second index condition. The second image C can be determined as the base image, and the second image D can be determined as the reference image. Therefore, the second image C can be adjusted based on the second image D to obtain a target image that satisfies the first and second index conditions.
[0147] In one embodiment, the index parameters of the target image include at least one of the following: image resolution, color depth, signal-to-noise ratio, and similarity between the target image and the first image. In one embodiment, the index parameters of the target image are greater than the index parameters of the second image, which can be understood as the weighted index parameters of the target image being greater than the weighted index parameters corresponding to the second image.
[0148] In this embodiment, the first image can be processed using different image processing strategies to obtain second images with different index parameters and superior performance compared to the first image. Furthermore, a target image with superior performance compared to each of the first images can be generated jointly based on at least two second images with superior performance and their index parameters. This effectively improves the performance of the first image and yields a second image with better performance.
[0149] In one embodiment, obtaining the target image based on at least two second images and index parameters of at least two second images includes:
[0150] The second image whose index parameters meet the first index condition is determined as the reference image; the index condition is used to indicate that the resolution of the second image is greater than a preset resolution threshold, or the similarity between the first image and the second image is greater than a preset similarity threshold.
[0151] The second image whose index parameters meet the second index condition is determined as the reference image; wherein the first index condition and the second index condition are different.
[0152] The target image is obtained by adjusting the reference image based on the reference image.
[0153] In one embodiment, a first indicator condition is used to indicate that the resolution of the second image is greater than a preset resolution threshold, and a second indicator condition is used to indicate that the similarity between the first image and the second image is greater than a preset similarity threshold. Here, any first type of indicator parameter in this disclosure can be the resolution of an image, and any second type of indicator parameter in this disclosure can be the similarity between the first image and the second image.
[0154] It should be noted that designating the second image that meets the first criterion as the baseline image can be understood as designating the second image with a resolution greater than the resolution threshold as the baseline image. Alternatively, designating the second image with a resolution greater than the resolution threshold as the baseline image can also be understood as designating the image with the highest resolution among all the second images as the baseline image.
[0155] It should be noted that designating a second image that meets the second criterion condition as a reference image can be understood as designating the second image with a similarity greater than the similarity threshold as the reference image. Alternatively, designating the second image with a similarity greater than the similarity threshold as the reference image can also be understood as designating the image with the highest similarity among all second images as the reference image.
[0156] In another embodiment, a first indicator condition is used to indicate that the similarity between the first image and the second image is greater than a preset similarity threshold. A second indicator condition is used to indicate that the resolution of the second image is greater than a preset resolution threshold. Here, any first type of indicator parameter in this disclosure can be the similarity between the first image and the second image, and any second type of indicator parameter in this disclosure can be the resolution of the image.
[0157] It should be noted that designating the second image that meets the first criterion as the baseline image can be understood as designating the second image with a similarity greater than the similarity threshold as the baseline image. Alternatively, designating the second image with a similarity greater than the similarity threshold as the baseline image can also be understood as designating the image with the highest similarity among all the second images as the baseline image.
[0158] It should be noted that designating a second image that meets the second criterion condition as a reference image can be understood as designating a second image with a resolution greater than a resolution threshold as a reference image. Alternatively, designating a second image with a resolution greater than the resolution threshold as a reference image can also be understood as designating the image with the highest resolution among all second images as the reference image.
[0159] In one embodiment, the reference image can be adjusted based on the reference image to obtain a target image that satisfies the first index condition and the second index condition.
[0160] In one embodiment, a reference image can be adjusted based on a reference image to obtain a target image with a first weighted index parameter. The first weighted index parameter is greater than a second weighted index parameter corresponding to the reference image, and also greater than a third weighted index parameter corresponding to the reference image.
[0161] In this embodiment, a reference image whose index parameters meet the first index condition can be adjusted based on a reference image whose index parameters meet the second index condition. This allows for the joint generation of a target image with superior performance in multiple dimensions, based on a reference image and a reference image that exhibits superior performance in at least one dimension. For example, a target image can be generated with a resolution greater than a resolution threshold and a similarity greater than a similarity threshold. Thus, the performance of the first image is effectively improved through the generated target image.
[0162] In one embodiment, the first image is processed based on at least two image processing strategies to obtain at least two second images, including:
[0163] At least two sets of image feature parameters are extracted from the first image, and / or the image feature parameters extracted from the first image are compressed based on at least two compression ratios to obtain at least two sets of image feature parameters; wherein the image processing strategies corresponding to different sets of image feature parameters are different.
[0164] Based on the image processing strategies corresponding to the feature parameters of each group of images, the first image is processed to obtain at least two second images.
[0165] In one embodiment, at least two sets of image feature parameters are extracted from the first image, wherein the proportion of global features and / or the proportion of local features in the different sets of image feature parameters are different. Global features can be image features extracted from the entire image region of the first image, and local features are image features extracted from a local image region of the first image.
[0166] In one embodiment, a first proportion of global features in the image feature parameters is positively correlated with the resolution of the second image generated based on the image feature parameters. And / or, a second proportion of local features in the image feature parameters is negatively correlated with the resolution of the second image generated based on the image feature parameters. It should be noted that during the generation of the second image based on the image feature parameters, the content of the first image represented by global features is more blurred compared to local features. In this case, richer and more detailed image content can be generated based on global features, and the generation of the second image is not overly limited by the image content of the first image, thereby effectively improving the resolution of the generated second image.
[0167] In one embodiment, the first proportion of global features in the image feature parameters is negatively correlated with the similarity to the second image generated based on the image feature parameters. And / or, the second proportion of local features in the image feature parameters is positively correlated with the similarity to the second image generated based on the image feature parameters. It should be noted that in the process of generating the second image based on the image feature parameters, the content of the first image is more refined than that of the global features. In this case, by enhancing the detailed information of the first image based on local features, the image content of the first image can be better restored based on the local features, thereby increasing the similarity between the first image and the second image generated based on the first image.
[0168] In one embodiment, the compression ratio is negatively correlated with the amount of data in each set of image feature parameters, and the amount of data in the image feature parameters is negatively correlated with the resolution of the second image generated based on the image feature parameters. It should be noted that the higher the compression ratio, the smaller the amount of data in the image feature parameters obtained after compression processing of the image feature parameters extracted from the first image, and the more blurred the content of the first image represented by the compressed image feature parameters. In this case, richer and more detailed image content can be generated based on the compressed image feature parameters, and the generation of the second image is not overly limited by the image content in the first image, thereby effectively improving the resolution of the generated second image.
[0169] In one embodiment, the compression ratio is negatively correlated with the amount of data in each set of image feature parameters, while the amount of data in the image feature parameters is positively correlated with the similarity to the second image generated based on the image feature parameters. It should be noted that the smaller the compression ratio, the larger the amount of data in the image feature parameters obtained after compression processing of the image feature parameters extracted from the first image. The compressed image feature parameters represent the content of the first image in a more refined manner. In this case, by enhancing the detail information of the first image based on the image feature parameters, the image content of the first image can be better restored based on the larger amount of image feature data, thereby increasing the similarity between the first image and the second image generated based on the first image.
[0170] In one embodiment, based on the image processing strategies corresponding to each set of image feature parameters, the first image is processed to obtain at least two second images, including: generating the second images based on the image feature parameters using an image generation model corresponding to each set of image feature parameters. The image generation model can be a pre-trained algorithm model. The image generation model can be used to extract image feature parameters from the first image generated from the input image model, and process these image feature parameters to obtain a second image with index parameters higher than those of the first image.
[0171] In one embodiment, the index parameters of the second image obtained based on different groups of image feature parameters are different.
[0172] In one embodiment, the indicator parameters include: image resolution and the similarity between the first image and the second image. The resolution and / or similarity of the second image obtained based on different image feature parameters will differ. Furthermore, the resolution and similarity of the second image are negatively correlated.
[0173] In this embodiment, at least two sets of image feature parameters can be extracted from the first image, and / or the extracted image feature parameters can be compressed at different compression ratios to obtain image feature parameters with different data volumes. Then, different image processing strategies can be used to process the first image for each set of image feature parameters, thereby obtaining at least two second images with different index parameters. This ensures the richness of the obtained second images, and allows for the joint generation of a high-performance target image based on the various second images with different index parameters.
[0174] In one embodiment, the indicator parameters include the image resolution and the similarity between the first image and the second image; the resolution of each second image is greater than that of the first image; based on the image processing strategy corresponding to each set of image feature parameters, the first image is processed to obtain at least two second images, including:
[0175] Based on the feature parameters of the first image and the descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a preset resolution threshold; wherein, the descriptive information is determined based on the type of each object in the first image and / or the relationship between each object.
[0176] Based on the second image feature parameters, the first image is processed to obtain a second image with a similarity greater than a preset similarity threshold;
[0177] Wherein, the first image feature parameter includes global features of the first image, the second image feature parameter includes local features of the second image; and / or, the data volume of the first image feature parameter is less than the data volume of the second image feature parameter.
[0178] In one embodiment, the relationships between objects can be relative positional relationships and / or containment relationships. Containment relationships indicate that one object is located within another object. Relative positional relationships indicate the distance and / or relative orientation between objects. It should be noted that the descriptive information can be high-level semantic information from the first image. During the process of generating a second image based on the first image, this descriptive information can be used to supplement missing details in the first image, thereby effectively improving the resolution of the generated second image.
[0179] In one embodiment, object types can be selected from a preset set of object types, and / or, relationships between objects can be selected from a preset set of relationships. It should be noted that since the number of object types and the number of prepositions describing the relationships between objects are finite, object types and preset relationships can be pre-set. That is, a limited set of object types and preset relationships can be stored offline. This facilitates the rapid retrieval of necessary information from the offline storage during the identification of object types and / or relationships in the first image, thereby accelerating the process of identifying relevant information about objects in the first image. This reduces the computational load of extracting descriptive information from the first image.
[0180] In one embodiment, the number of object types may be less than a preset threshold.
[0181] In one embodiment, when the number of object types exceeds a certain threshold, a first similarity can be determined between the object types. Object types with a first similarity greater than the first similarity threshold are then merged into the same type. For example, the object types include a first type and a second type. The first type could be a peony, and the second type could be a hydrangea. If the first similarity between the first type and the second type reaches the first similarity threshold, then the first type and the second type can be merged into the same type: flower. This effectively reduces the amount of information stored for description.
[0182] It should be noted that, in order to distinguish it from the first similarity and the first similarity threshold, the similarity between any first image and the second image of this disclosure can be referred to as the second similarity, and the similarity threshold corresponding to any second similarity of this disclosure can be referred to as the second similarity threshold.
[0183] In one embodiment, the number of preset object types can be less than a preset threshold. Thus, during the process of identifying the object types in the first image, the number of object types can be controlled to be less than the preset threshold, eliminating the need to reduce the number of object types after the object types in the first image have been identified. For example, object types may include: buildings, greenery, flowers, animals, birds, human bodies, and man-made objects.
[0184] In one embodiment, when the computing power margin of the electronic device exceeds a preset computing power threshold, the descriptive information of the first image can be determined based on a preset first algorithm model. The first algorithm model can be a graph-to-text model. For example, the graph-to-text model can be a Contrastive Language-Image Pretraining (CLIP) or a Large Language and Vision Assistant (LLaVa) model. When the computing power margin of the electronic device is less than the preset computing power threshold, the descriptive information of the first image can be determined based on a second algorithm model. The preset algorithm model can be a graph-to-text model. The second algorithm model can be a general object detector.
[0185] In one embodiment, the amount of data for the model parameters of the second algorithm model is less than the amount of data for the model parameters of the first algorithm model. And / or, the second algorithm model extracts descriptive information from the image faster than the first algorithm model. And / or, the amount of descriptive information extracted from the image by the second algorithm model is less than the amount of descriptive information extracted from the image by the first algorithm model.
[0186] In one embodiment, the image feature parameters may include global features and local features. A first proportion of global features in the first image feature parameters is greater than a first proportion of global features in the second image feature parameters. A second proportion of local features in the first image feature parameters is less than a second proportion of local features in the second image feature parameters.
[0187] In one embodiment, image feature parameters extracted from the first image can be compressed based on a first compression ratio to obtain first image feature parameters with a first data volume. Image feature parameters extracted from the first image can be compressed based on a second compression ratio to obtain second image feature parameters with a second data volume. The first data volume can be less than the second data volume. The first compression ratio can be greater than the second compression ratio.
[0188] In one embodiment, the same set of image feature parameters extracted from the first image can be compressed based on a first compression ratio and a second compression ratio to obtain the first image feature parameters and the second image feature parameters.
[0189] In another embodiment, a first set of image feature parameters extracted from the first image can be compressed based on a first compression ratio to obtain first image feature parameters with a first data volume. A second set of image feature parameters extracted from the first image can be compressed based on a second compression ratio to obtain second image feature parameters with a second data volume. The first data volume can be less than the second data volume. The first compression ratio can be greater than the second compression ratio. Here, different compression ratios are used to compress different sets of image feature parameters. The first proportion of global features and / or the second proportion of local features in different sets of image feature parameters are different.
[0190] In one embodiment, a first set of image feature parameters and a second set of image feature parameters can be extracted from a first image. Specifically, the first proportion of global features in the first set of image feature parameters is greater than the first proportion of global features in the second set of image feature parameters. Conversely, the second proportion of local features in the first set of image feature parameters is less than the second proportion of local features in the second set of image feature parameters.
[0191] In one embodiment, the first image can be processed based on the first image feature parameters to obtain a second image with a resolution greater than a preset resolution threshold; the first image can be processed based on the second image feature parameters to obtain a second image with a similarity greater than a preset similarity threshold; wherein the first image feature parameters include global features of the first image, the second image feature parameters include local features of the second image; and / or, the data volume of the first image feature parameters is less than the data volume of the second image feature parameters.
[0192] In this embodiment, on one hand, detailed information in the first image can be supplemented based on descriptive information describing each object in the first image, thereby effectively improving the resolution of the first image and obtaining a second image with a resolution greater than a resolution threshold. Alternatively, the first image can be fuzzily described based on first image feature parameters including global features and / or features with a small data volume, thereby supplementing detailed information in the first image with a stronger generation intensity, thus generating a second image with a resolution greater than a resolution threshold. On the other hand, the first image can be precisely described based on second image feature parameters including local features and / or features with a large data volume, thereby supplementing detailed information in the first image with a weaker generation intensity, generating a second image with a similarity greater than a similarity threshold.
[0193] In one embodiment, the method further includes:
[0194] Based on the type of each object in the first image, determine the processing parameters of the first image region where each object is located;
[0195] Based on the feature parameters and descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a resolution threshold, including:
[0196] Based on the processing parameters of each first image region, the feature parameters of the first image, and the descriptive information of the first image, each first image region of the first image is processed to obtain a second image with a resolution greater than the resolution threshold.
[0197] In one embodiment, the image regions where each object in the first image is located can be segmented to obtain various first image regions. For example, semantic segmentation can be used to segment the detected objects, obtaining masks corresponding to each object. It should be noted that the area covered by this mask is the first image region where the object is located. Here, based on processing parameters corresponding to the object type, the first image regions covered by different masks can be precisely processed using the masks corresponding to each object.
[0198] In one embodiment, the type of object in each first image region can be determined. And / or, the association between each first image region can be determined to determine the association between each object corresponding to each first image region. Here, the association between each object can be accurately determined for each segmented first image region.
[0199] In one embodiment, the processing parameters can be positively correlated with the amount of image information data added to the first image region. It should be noted that the more image information data added to the first image region, the more detailed information is added to the first image region, and the higher the clarity of the first image region.
[0200] In one embodiment, the index parameters of the processed first image region differ depending on the processing parameters. The processing parameters may be positively correlated with resolution, and / or negatively correlated with similarity.
[0201] In one embodiment, the processing parameters may include noise parameters. The first image region can be processed based on the noise parameters corresponding to the first image region.
[0202] In one embodiment, processing a first image region based on noise parameters corresponding to the first image region includes: processing the first image region based on the noise parameters corresponding to the first image region using a preset diffusion model. The noise parameters can be used to indicate the proportion of noise added to the first image region during the processing of each first image region based on the diffusion model.
[0203] It should be noted that the image region processing based on the diffusion model includes both forward diffusion and reverse diffusion processes. During forward diffusion, noise is gradually added to the image region, eventually making the data distribution in the image region approximate a standard Gaussian distribution. In the reverse generation process, noise is gradually removed starting from the standard Gaussian noise to generate ultra-high-resolution image data. The noise parameter here can be used to indicate the proportion of noise added to the image region during the forward diffusion process.
[0204] For example, the type of object may include: animal fur, green plants and flowers, buildings, and man-made objects. For animal fur and green plants and flowers, the noise parameter may be a first noise parameter, and for buildings and man-made objects, the noise parameter may be a second noise parameter. The first noise parameter may be less than the second noise parameter.
[0205] Here, to suppress the randomness of the diffusion model, image feature parameters of the low-quality first image can be embedded into the initial sampling noise. For the first image region containing objects such as animal fur and green plants, the noise ratio in that region can be increased during the process of embedding image feature parameters into the initial sampling noise. This can improve the resolution of the generated first image region, meaning that more detailed information in the first image region can be generated with a stronger generation intensity. Conversely, for buildings and man-made objects, the proportion of image feature parameters can be increased during the process of embedding image feature parameters into the initial sampling noise, reducing the randomness of the detailed information in the image region generated by the diffusion model.
[0206] It should be noted that, for the diffusion model, noise is gradually added to the image feature parameters until the noise-added image feature parameters conform to a Gaussian distribution. That is, there are multiple processes of adding noise to the image feature parameters. Here, increasing the proportion of the image feature parameters is equivalent to adding noise from an intermediate step, thus spreading the noise.
[0207] In one embodiment, an encoder can be used to encode the first image to obtain image feature parameters in the first image. Here, the image feature parameters can be understood as the encoding result of encoding the first image using the encoder.
[0208] In one embodiment, a diffusion model can be used to fuse the image feature parameters and random noise matrix of the first image based on noise parameters to obtain image feature parameters with added noise; a diffusion model can be used to generate target feature parameters based on descriptive information and image feature parameters with added noise; and a decoder can be used to decode the target feature parameters to obtain a second image with a resolution greater than a resolution threshold.
[0209] In one embodiment, the process of constructing a diffusion model can be as follows:
[0210] Step 1: Based on the first sample image group, train the encoder until the encoder meets the training conditions; wherein, the encoder is used to extract image feature parameters from the images input to the encoder.
[0211] Here, the images in the first sample image group can be images captured in telephoto shooting scenarios. The trained encoder can be an encoder adapted to the degradation of telephoto images. In this case, the encoder is an autoencoder; after encoding the input image, it outputs an encoded vector (i.e., image feature parameters). This encoded vector, after decoding, can generate an image identical to the input image. This approach allows the encoder to have a certain degree of robustness in processing blurred images captured in telephoto shooting scenarios.
[0212] Step 2: Based on the second set of sample images, train the decoder until the decoder meets the training conditions.
[0213] Here, the resolution of the images in the second sample image group can be greater than a resolution threshold. The images in the second sample image group can be high-resolution images. A decoder can be trained using high-resolution images to produce clear images after decoding the encoded vectors based on the decoder.
[0214] Step 3: Establish an algorithm model between the trained encoder and decoder. The input of the algorithm model is the output of the encoder, and the output of the algorithm model is the input of the decoder. The algorithm model can be trained based on a third set of sample images until the training conditions are met, resulting in a diffusion model. The training conditions can be the convergence of the algorithm model's parameters.
[0215] Here, after training is complete, the learning rate can be reduced, and the model parameters in the encoder and decoder can be fine-tuned.
[0216] Step 4: The inference steps of the diffusion model are compressed using a progressive distillation method until the inference steps of the diffusion model are less than a preset threshold.
[0217] For example, if the original inference steps of the trained diffusion model are 1000, the incremental distillation method can compress the inference steps of the diffusion model to 3 to 5, which is 166 times faster than the original 1000 inference steps.
[0218] It should be noted that when training the diffusion model, an incremental distillation method can be used to reduce the number of inference steps. Additionally, some model parameters can be pruned and some operators replaced to improve the running speed of the diffusion model on electronic devices.
[0219] Compared to related technologies that process different regions of a first image based on uniform processing parameters, leading to mismatches between the first image regions and the processing parameters, this embodiment adapts to the type of each object, applying processing parameters tailored to the type of object in each region of the first image. This ensures good processing results for each region of the first image.
[0220] In one embodiment, determining the processing parameters of the first image region where each object is located, based on the type of each object in the first image, includes:
[0221] If the first image includes an object of a first target type, determine the pose of the object of the first target type;
[0222] Based on the pose of the first target type, determine the processing parameters of the first image region where the object of the first target type is located;
[0223] There is a correspondence between the pose and processing parameters of the first target type.
[0224] In one embodiment, the type of the object may include: a first target type and types other than the first target type.
[0225] In one embodiment, a first processing parameter can be used to process the first image region containing the first target type; a second processing parameter can be used to process the first image region containing objects of other types besides the first target type. That is, the first image can be divided into two first image regions: one containing objects of the first target type, and the other containing objects of the first target type. Different processing parameters can be applied to each of the two first image regions. This simplifies the type of processing parameters determined and simplifies the process of processing each first image region separately, thereby improving image processing speed.
[0226] In one embodiment, the resolution of the image region obtained based on the first processing parameters is greater than the resolution of the image region obtained based on the second processing parameters.
[0227] In one embodiment, the first processing parameters include: a first processing sub-parameter and a second processing sub-parameter; when the pose of an object of the first target type is a preset first pose, the object of the first target type in the first image region with the first pose is processed based on the first processing sub-parameter. When the object of the first target type has a second pose different from the first pose, a pose adjustment operation can be performed on the object of the first target type in the first image with the second pose based on the second processing sub-parameter; the second processing sub-parameter is used to adjust the pose of the first target type to the first pose; the object of the first target type in the first image region with the first pose is processed based on the first processing sub-parameter, and the object of the first target type in the first image region with the first pose is adjusted to the second pose.
[0228] It should be noted that the first processing sub-parameter can be a pre-determined parameter for processing objects of the first target type with a first pose. Here, it is not necessary to pre-determine different processing parameters for objects of the first target type with different poses. Instead, objects of the first target type with different poses can be uniformly corrected to the first pose, and then processed based on the pre-determined first processing sub-parameter. This reduces the amount of data for processing parameters and ensures good processing results for objects of the first target type based on the processing parameters.
[0229] For example, the object of the first target type may include: text. If text is detected in the first image and the text is in a second pose, the pose of the text can be adjusted to a first pose based on the second processing parameters, and the text with the second pose can be subjected to ultra-high definition processing based on the first processing parameters; after obtaining the ultra-high definition text region, the pose of the text can be restored to the first pose.
[0230] It should be noted that the first orientation here indicates that the text is parallel to the horizontal direction. The second orientation indicates that the text is tilted relative to the horizontal direction.
[0231] It should be noted that the first image region of an object with the first target type can be segmented from the first image, and after processing the first image region, the processed first image region can be pasted back into the original image.
[0232] Here, ultra-high-definition images specifically for text can be generated through text detection, text orientation correction, text recognition, generation of ultra-high-definition images of text, text orientation rotation, and pasting the rotated text back onto the original image.
[0233] For example, the object of the first target type may further include: a face. When a face is detected in the first image and the face is in a second pose, the pose of the face can be adjusted to a first pose based on the second processing parameters, and the face with the second pose can be subjected to ultra-high definition processing based on the first processing parameters; after obtaining the ultra-high definition face region, the pose of the face can be restored to the first pose.
[0234] It should be noted that the first pose can be a face facing outwards. The second pose can be a face turned to the side facing outwards.
[0235] In one embodiment, when the object of the first target type is a face, key feature points of the face can be determined; the pose of the face can be determined based on the relative positional relationship between the key feature points. Exemplarily, the key feature points can be used to indicate at least one of the following features of the face: eyebrows, eyes, nose, ears, and mouth.
[0236] Here, ultra-high-definition image generation specifically for faces can be achieved through face detection, facial landmark localization, face orientation correction, generation of ultra-high-definition face images, face orientation rotation, and pasting the rotated face back onto the original image.
[0237] In one embodiment, the first image can be processed using a convolutional neural network to obtain a second image. The resolution of the second image is higher than that of the first image. It should be noted that when image degradation is severe, the semantic perception capability of a typical convolutional neural network will decrease.
[0238] Based on this, a semantic awareness method can be provided in this disclosure, the method comprising:
[0239] Obtain descriptive information for the first image. This information includes descriptive statements and object categories. The descriptive statements describe the relationships between the objects. Based on the descriptive statements and object types of the first image, guide the text-based graph model to generate a high-quality second image.
[0240] Here, a degradation-aware extractor can be trained first, aiming to align the image feature parameters extracted from the low-quality first image with those from the high-quality image. The first half of the degradation-aware extractor is then used as the encoder. The encoder maps the image feature parameters of the first image to the latent space of the diffusion model. During image generation, the encoder encodes the low-quality first image into image feature parameters (which can be understood as feature vectors) and embeds them into initial random Gaussian noise. The descriptive information of the first image is used as prompts, and the cross-attention mechanism proposed in Pixel-Aware Stable Diffusion (PASD) is employed to learn semantic guidance, ultimately ensuring the generation of ultra-high-definition images with semantic fidelity.
[0241] It should be noted that while this method can generate high-resolution images, the computational burden of extracting input information is significant, making it difficult to deploy on mobile devices. Furthermore, it exhibits some loss of control in areas rich in information, such as text and fine textures, easily generating noticeably erroneous textures.
[0242] Based on this, in one embodiment, a second image with a resolution greater than a resolution threshold can be generated based on the first image using a first image generation model; a second image with a similarity greater than a similarity threshold can be generated based on the first image using a second image generation model. A target image is generated based on the second image with a resolution greater than the resolution threshold, the second image with a similarity greater than the similarity threshold, and the index parameters of each second image. Here, the advantages of combining the second image with a resolution greater than the resolution threshold and the second image with a similarity greater than the similarity threshold can be combined to improve the clarity and realism of the generated target image.
[0243] In one embodiment, the first image generation model may include at least one of the following: a diffusion model, a face ultra-high definition model, and a text ultra-high definition model.
[0244] In one embodiment, if the first image includes an object of a first target type, and the object of the first target type is a face, the first image is processed based on a face ultra-high resolution model to obtain a second image with a resolution greater than a resolution threshold.
[0245] The ultra-high-resolution face model described here can perform any of the image processing operations for faces disclosed herein. Based on the type of each object in the first image, the ultra-high-resolution face model determines the processing parameters for the first image region where each object is located; based on the processing parameters of each first image region and the first image feature parameters, it processes each first image region of the first image separately to obtain a second image with a resolution greater than a resolution threshold. For example, the ultra-high-resolution face model can use first processing parameters to process the first image region where a face is located; and it can use second processing parameters to process the first image region where non-face objects are located.
[0246] In one embodiment, if the first image includes an object of a first target type, and the object of the first target type is text, the first image is processed based on a text ultra-high definition model to obtain a second image with a resolution greater than a resolution threshold.
[0247] The super-resolution text model described here can perform any of the image processing operations for text disclosed herein. The super-resolution text model can determine the processing parameters of the first image region where each object is located based on its type in the first image; based on the processing parameters of each first image region and the first image feature parameters, it processes each first image region of the first image separately to obtain a second image with a resolution greater than a resolution threshold. For example, the super-resolution text model can use first processing parameters to process the first image region where text is located; the super-resolution text model can use second processing parameters to process the first image region where non-text objects are located.
[0248] Here, considering the higher demands of users for face and text scenes, separate models can be designed for these two scenes to process the first image, in order to utilize stronger prior information and generate clear and accurate textured images.
[0249] In one embodiment, the second image generation model can be an image generation model obtained by training a convolutional neural network. The second image generation model can process the first image based on any second image feature parameters of this disclosure to obtain a second image with a similarity greater than a similarity threshold.
[0250] Compared to processing all objects of the first target type with the same parameters, which may lead to a mismatch between the objects and the processing parameters, this embodiment of the present disclosure, when the first image includes objects of the first target type, can adapt to the pose of the objects of the first target type and accurately determine the processing method for the first image region where the objects of the first target type are located. This improves the accuracy and flexibility of processing objects of the first target type.
[0251] In one embodiment, the method further includes:
[0252] Based on the types of each object in the first image and the relationships between each object, the first descriptive information is determined;
[0253] In response to the detection of a preset input operation, the first description information is adjusted based on the preset input operation to obtain the second description information;
[0254] Based on the feature parameters and descriptive information of the first image, the first image is processed, including:
[0255] The first image is processed based on the first image feature parameters and the second descriptive information.
[0256] In one embodiment, processing the first image based on first image feature parameters and first image description information includes: processing the first image based on first image feature parameters and first description information.
[0257] In one embodiment, the first descriptive information can be compressed based on a third compression ratio to obtain compressed first descriptive information. The first image is then processed based on the first image feature parameters and the compressed first descriptive information.
[0258] In one embodiment, image feature parameters extracted from the first image can be compressed based on a first compression ratio to obtain first image feature parameters. First descriptive information can be compressed based on a third compression ratio to obtain compressed first descriptive information. The first image can be processed based on the first image feature parameters and the compressed first descriptive information. The first compression ratio can be greater than the third compression ratio. Alternatively, the first compression ratio and the third compression ratio can be the same.
[0259] Here, the first image can be processed based on the compressed first description information and the compressed first image feature parameters. In this way, while reducing the amount of data used to process the first image and improving the processing speed, the blurred first image feature parameters and blurred first description information can be used to better supplement the missing details in the first image, thereby improving the resolution of the obtained second image.
[0260] In one embodiment, adjusting the first description information based on a preset input operation to obtain second description information includes: determining input information corresponding to the preset input operation. The input information indicates at least one of the following: the type of the second object, the association between the first object and the second object, and the association between the various second objects. The second description information is obtained by adjusting the first description information based on the input information.
[0261] For example, the first descriptive information can be used to indicate that the type of the first object includes birds and grass, and the relationship between the birds and grass can be that the birds are located on the grass. The input information can be used to indicate that the type of the second object is a snake, and the relationship between the second object and the first object can be that the snake is located on the grass. Then, the second descriptive information obtained by adjusting the first descriptive information based on the input information can be used to indicate that the first image includes birds, snakes, and grass, and both the birds and snakes are located on the grass.
[0262] In one embodiment, the input information can also be used to instruct adjustments to the type of the first object and / or the association between the first objects. For example, the first descriptive information can be used to indicate that the type of the first object includes a bird and grass, and the association between the bird and grass is that the bird is located on the grass. The input information can be used to adjust the bird to a snake. Then, the second descriptive information obtained after adjusting the first descriptive information based on the input information can be used to indicate that the first image includes a snake and grass, and the snake is located on the grass.
[0263] In one embodiment, the second description information can be compressed based on a third compression ratio to obtain compressed second description information; the first image is then processed based on the first image feature parameters and the compressed second description information.
[0264] In this embodiment, the first descriptive information can be adjusted based on the content input by the user's preset input operation to obtain second descriptive information that meets the user's needs. Furthermore, a second image that meets the user's needs can be intelligently generated based on the second descriptive information that meets the user's needs.
[0265] In one embodiment, the method further includes:
[0266] The associations in the first description information that do not meet the preset description conditions are deleted to obtain the processed first description information;
[0267] The first description information is adjusted based on a preset input operation to obtain the second description information, including:
[0268] The processed first description information is adjusted based on the preset input operation to obtain the second description information.
[0269] It should be noted that the preset description conditions can be that the relationships in the first description information conform to the description rules of natural language.
[0270] In one embodiment, a pre-trained natural language model can be used to determine whether the relationships in the first descriptive information meet preset descriptive conditions. It should be noted that determining whether the relationships in the first descriptive information meet the preset descriptive conditions can be understood as making a reasonableness judgment on the relationships in the first descriptive information. If the relationship is unreasonable, the relationship is deleted, while the type of the object corresponding to the relationship is retained.
[0271] For example, the first description information can be used to indicate that the types of objects include: people and crows, and the association between people and crows indicates that a person is on a crow. Here, the association between people and crows does not conform to the preset description conditions, that is, the association between people and crows is illogical. In this case, the association between people and crows can be deleted from the first description information, while retaining the types of objects indicated by the first description information.
[0272] In one embodiment, the image regions where each object in the first image is located can be segmented to obtain each first image region and its mask. The relative positional relationships between objects in the first image can be determined based on the relative positional relationships between the masks. The inclusion relationships between objects in the first image can be determined based on the inclusion relationships between the masks.
[0273] In one embodiment, a target object in a first image can be determined. The target object is an object of a first target type, and / or, the target object is located at a preset position. The preset position can be the location of the center region of the first image. The order in which preset operations are performed on each object can be determined based on the distance between each object and the target object. Distance and order can be positively correlated. The preset operations can be used to determine the type of each object and / or the association relationship between each object. The association relationship includes: the relative positional relationship between each object and / or the containment relationship between each object.
[0274] For example, the target object of the image can be used as a reference. In a scene containing a person, the person is taken as the target object; in a scene not containing a person, objects near the center of the image are taken as the target object; and if the center of the image does not contain any objects, only the distribution of objects around the perimeter of the image is described. When describing the relative positional relationship between objects, one can extend outward from the main object. For objects without a containment relationship, their relative positional relationship can be described as "adjacent," "not far," and "far away," depending on their distance. For example, a person is adjacent to a tree, and not far from the tree is grass. For objects with a containment relationship, this is represented by the mask of one object being completely within the mask of another object. To describe the containment relationship, two objects with a containment relationship can be directly described as one object being inside another object. It should be noted that the containment relationship obtained solely from the mask is not entirely accurate and may lead to an object appearing to be inside another object when it is not actually so. For example, a bird flying in front of a building might be associated with the bird being inside the building based on the mask relationship, which is clearly not realistic. Therefore, we perform a rationality assessment on the descriptive statements generated in the aforementioned process. We use a natural language model to determine whether the statement is reasonable. If it is unreasonable, we delete the descriptive statement while retaining the object type in the description. We can also support user input of specific descriptive statements to control image generation.
[0275] In one embodiment, associations in the second description information that do not meet preset description conditions are deleted to obtain processed second description information. The first image is processed based on the first image feature parameters and the second description information, including: processing the first image based on the first image feature parameters and the processed second description information.
[0276] In this embodiment of the disclosure, associations in the first description information that do not conform to preset description conditions can be deleted. That is, associations in the first description information that do not conform to common sense can be deleted to obtain first description information that conforms to preset description conditions. In this way, further image processing operations can be performed based on the first description information that conforms to preset description conditions, thereby ensuring the accuracy of the process of processing the first image.
[0277] In one embodiment, adjusting a reference image based on a reference image to obtain a target image includes:
[0278] Based on the type of object in each first image region of the first image, select the region to be adjusted from each second image region in the reference image that matches the first image region;
[0279] Based on a third image region in the reference image that matches the region to be adjusted, the region to be adjusted in the baseline image is adjusted to obtain the target image.
[0280] In one embodiment, a second image region with a resolution greater than a resolution threshold and / or a similarity less than a similarity threshold can be identified as the region to be adjusted.
[0281] In one embodiment, there is a correspondence between the object type and the processing parameters for processing the first image region. The processing parameters are positively correlated with the resolution of the processed first image region, and negatively correlated with the similarity of the processed first image region. A second image region in the reference image can be determined based on the object type, where the resolution is greater than a resolution threshold and / or the similarity is less than a similarity threshold.
[0282] In one embodiment, if a first image region of a first image includes an object of a first target type, a second image region that matches the first image region including the object of the first target type can be selected as the region to be adjusted.
[0283] In one embodiment, adjusting the region to be adjusted in the reference image based on a third image region in the reference image that matches the region to be adjusted, to obtain a target image, includes: adjusting the image content of the region to be adjusted in the reference image based on the image content of the third image region that matches the region to be adjusted, to obtain the target image. It should be noted that adjusting the second image content of the region to be adjusted in the reference image based on the first image content of the third image region that matches the region to be adjusted, to obtain the target image, can also refer to fusing the first image content and the second image content to obtain the target image.
[0284] In this embodiment, on one hand, the region to be adjusted can be selected from each of the second image regions in the reference image that match the first image regions, based on the type of objects in each of the first image regions of the first image. That is, the region to be adjusted that needs to ensure good performance can be accurately selected from the reference image according to the type of the object. On the other hand, the region to be adjusted in the reference image can be adjusted based on a third image region in the reference image that matches the region to be adjusted, thereby ensuring that the advantages of the second and third image regions can be combined to adjust the region to be adjusted, thus ensuring that the adjusted region has good performance.
[0285] In one embodiment, a first indicator condition is used to indicate that the resolution of the reference image is greater than a resolution threshold. Based on the type of objects in each first image region of the first image, selecting regions to be adjusted from each second image region in the reference image that matches the first image region includes:
[0286] For each region of the first image, a second target type is determined from a preset type based on the type of the object in the first image region; wherein there is a correspondence between the object type and the preset type;
[0287] If a first image region of a first image matches a second image region of a reference image, determine the image content of a second target type in the first image region and the similarity between the image content of the second target type in the second image region.
[0288] The second image region in the baseline image with a similarity less than the similarity threshold is identified as the region to be adjusted.
[0289] In one embodiment, the image content of the preset type may include at least one of the following: texture content, color content, object type, and text content. In one embodiment, determining a second target type from the preset types based on the object type in the first image region includes: determining the type of the first image region based on the object type in the first image region; wherein the type of the first image region includes: a first type and a second type; when processing the first image, for the first image region of the first type, processing parameters greater than a first threshold are applied to the first image region; when processing the first image, for the first image region of the second type, processing parameters less than a first threshold are applied to the first image region. If the type of the first image region is the first type, the second target type is determined to be the object type; if the type of the first image region is the second type, the second target type is determined to be texture content.
[0290] In one embodiment, when the second target type is an object, if the type of the object in the first image region is different from the type of the object in the second image region that matches the first image region, then the second image region is determined as the region to be adjusted. If the type of the object in the first image region is the same as the type of the object in the second image that matches the first image region, then the second image region is not determined as the region to be adjusted.
[0291] It should be noted that for the first image region using larger processing parameters—that is, for a strongly generated region where the supplementary detail information can be significantly enhanced—strong texture generation is acceptable. However, it is necessary to ensure the semantic consistency between the second image region matching the first image region and the semantic result of the first image region. Therefore, it is sufficient to ensure that the object types in the first image region and the second image region matching the first image region are consistent.
[0292] In one embodiment, when the image content of the second target type is texture content, if the similarity between the texture content in the first image region and the texture content of the second image region matching the first image region is less than a similarity threshold, then the second image region is determined as a region to be adjusted. If the similarity between the texture content of the first image region and the texture content of the second image region matching the first image region is greater than the similarity threshold, then the second image region is not determined as a region to be adjusted.
[0293] It should be noted that for the first image region using smaller processing parameters—that is, regions requiring weaker generation to supplement detailed information, such as the first image region of objects like buildings and man-made objects—generating a strong texture on top of the existing texture is acceptable. However, if a strong texture is generated at a location in the first image region that has no texture at all, meaning the similarity between the texture content in the first image region and the texture content in the matching second image region is less than a similarity threshold, it needs to be determined as abnormal texture generation. In this case, a third image region with a similarity greater than the similarity threshold will be used to enhance the second image region. The generation intensity in the second image region is reduced by fusing the third and second image regions.
[0294] In one embodiment, the similarity of the text content in all matching first and second image regions can be determined; the second image region with a similarity less than a similarity threshold is determined as the region to be adjusted.
[0295] It should be noted that after generating the second image, the recognizability of the text content in the second image must be no worse than that of the text content in the first image. If the recognizability of the text content in the second image region of the reference image is worse than that of the text content in the first image region of the first image, then a third image region is needed to enhance the text content in the second image region that does not meet the requirements, in order to improve the recognizability of the text content.
[0296] In one embodiment, the similarity of text content in a matching first image region and a second image region can be determined based on a text ultra-high resolution model.
[0297] In one embodiment, text content can be extracted from the first image and the second image respectively using Optical Character Recognition (OCR) technology. During the generation of the second image, the text content in the second image can be generated based on the text content extracted from the first image using OCR technology.
[0298] In one embodiment, the color content of the reference image and the base image can be fused to obtain the target image. It should be noted that a color correction method can be used to color-align the reference image generated by any of the first image generation models of this disclosure and the reference image generated by the second image generation model, and a fusion method can be used to recreate the boundary effect produced by stitching together different images, thereby obtaining a high-definition, non-degradable, high-quality target image.
[0299] In one embodiment, the image content of each preset type in the first image can be compared with the image content of each preset type in the reference image to obtain the similarity between each first image region of the first image and the second image region that matches the first image region of the first image. If the similarity corresponding to any preset type of image content in the second image region is less than the similarity threshold, the second image region is determined as the region to be adjusted.
[0300] It should be noted that although every effort is made to avoid abnormal generation during the generation of the reference image, the quality of the first image may be poor, and the textures of some low-quality areas may be ambiguous, leading to some abnormal results in the generated second image. To address this issue, this disclosure includes a consistency correction module. This module corrects errors in the generated reference image through color consistency correction, texture consistency correction, object consistency correction, and text consistency correction, thereby obtaining a target image with good image quality.
[0301] It should be noted that in the consistency correction scheme adopted in this approach, traditional methods can be used for texture detection and texture consistency correction, or a deep learning network can be used to extract texture content from the first image and the reference image, and determine the similarity of the extracted texture content between the first image and the reference image. In the object consistency correction and text consistency correction schemes, the similarity between the image feature parameters of the first image extracted based on the deep learning network and the image feature parameters of the reference image extracted based on the deep learning network can be determined. The image feature parameters are intermediate features generated during the image generation process. The image feature parameters can characterize the object type and / or text content. Alternatively, the similarity of the image content of the first image and the finally generated reference image can be directly determined. This image content can be the final output result of the deep learning network. This image content includes: the object type and / or text content.
[0302] In this embodiment, based on the type of object in the first image region, the second target type of the image content to be verified in the second image region can be accurately determined. This allows for the accurate identification of the region to be adjusted from the reference image based on the similarity between the matching image content of the second target type in the first and second image regions. This ensures the accuracy of the region to be adjusted, thereby guaranteeing a good adjustment effect on the region to be adjusted in the reference image based on the reference object.
[0303] Figure 2 This is an image processing method illustrated according to an exemplary embodiment, such as... Figure 2 As shown, the method includes:
[0304] Step 201: Determine the first image.
[0305] Step 202: Encode the first image using an encoder to obtain the feature parameters of the first image.
[0306] Step 203: Determine the random noise.
[0307] Step 204: Determine the description information of the first image.
[0308] Step 204: Using the first image generation model, process the first image feature parameters output by the encoder based on random noise and descriptive information.
[0309] The first image generation model may include a diffusion model, which strongly supplements detailed information in the first image.
[0310] Step 205: The decoder decodes the first target feature parameters output by the diffusion model.
[0311] Step 206: Obtain the reference image.
[0312] Step 207: Process the first image using the second image generation model.
[0313] The second image generation model can be an image generation model obtained by training a regular convolutional neural network. The second image generation model is weaker in supplementing detailed information in the first image.
[0314] Step 208, obtain the reference image.
[0315] Step 209: If there are anomalies in the reference image, the reference image is corrected based on the reference image.
[0316] It should be noted that an anomaly in the reference image can be determined when there is poor consistency between the reference image and the first image.
[0317] Step 210: Obtain the target image.
[0318] Figure 3 This is a schematic diagram illustrating a process for extracting descriptive information according to an exemplary embodiment, such as... Figure 3 As shown, the method includes:
[0319] Step 301: Determine the first image.
[0320] Step 302: Detect each object in the first image.
[0321] Step 303: Determine the type of each object in the first image.
[0322] Step 304: Perform semantic segmentation on each object detected in the first image.
[0323] Step 305: Determine the relationships between the segmented objects.
[0324] Step 306: Determine the input information entered by the user.
[0325] The input information here can be any of the input information disclosed herein.
[0326] Step 307: Determine the description information based on the type of each object, the relationship between each object, and the input information.
[0327] Step 308: Determine the reasonableness of the description information.
[0328] The reasonableness judgment of the description information can be understood as determining whether the relationship between the various objects in the description information meets the preset description conditions.
[0329] Step 309: Delete unreasonable descriptive statements in the description information to obtain reasonable description information.
[0330] Description statements are used to describe the relationships between various objects. Deleting unreasonable description statements here can be understood as deleting relationships in the description information that do not meet the preset description conditions.
[0331] Step 310: Generate a description vector based on reasonable description information.
[0332] Step 311: Determine the noise parameters corresponding to each segmented object based on the type of each object.
[0333] It should be noted that after the random noise matrix is initialized, the encoded result obtained by the encoder after processing the first image can be fused with random noise to varying degrees based on the diffusion model to control the generation intensity. The noise parameter can be understood as the proportion of random noise added to the encoded result by the diffusion model. For regions where high generation intensity is desired, a higher proportion of noise is fused; for regions where low intensity is desired, a lower proportion of noise is fused. Higher generation intensity allows for the addition of more detailed information in the image region.
[0334] It's important to note that the key to improving the fidelity of ultra-high-definition image generation methods lies in maintaining consistency between the generated image and the high-level semantic information in the low-quality input image. Therefore, accurately extracting information from the low-quality input image is crucial. Currently, mainstream generative ultra-high-definition methods employ graph-to-text models to obtain image descriptions. For example, PASD uses CLIP to obtain the description of the input image, and SUPIR uses the LLaVa large model to extract the image description. However, graph-to-text models are computationally intensive and difficult to deploy on mobile devices. The semantic descriptions in common ultra-high-definition image generation algorithms typically include: the types of objects in the scene, the relative positional relationships between objects in the scene, and detailed descriptions of the objects in the scene.
[0335] In this disclosure, a lightweight, general-purpose object detector is employed to acquire the types and location information of common objects in the image, simplifying the extraction of descriptive information. To better control the generation intensity of different objects, semantic segmentation can be used to obtain their precise masks. Since the encoder performs high compression when encoding the first image into a feature vector and feeding it into the diffusion model, general algorithms for generating ultra-high-definition images do not require precise target location descriptions for each object; only the relative positions of each object need to be described to achieve accurate generation of ultra-high-definition images.
[0336] Figure 4 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 4 As shown, the method includes:
[0337] Step 401: Determine the first image.
[0338] Step 402: If the first image does not contain text or faces, the first image is encoded using an encoder to obtain the feature parameters of the first image.
[0339] Step 403: Determine the random noise.
[0340] Step 404: Determine the description vector of the first image;
[0341] Step 405: Using a diffusion model, the first image feature parameters output by the encoder are processed based on random noise and descriptive information.
[0342] Step 406: The decoder decodes the first target feature parameters output by the diffusion model.
[0343] Step 407: Obtain the reference image.
[0344] Step 408: If the first image contains text, detect the text in the first image based on the text super-resolution model.
[0345] Step 409: Correct the posture of the text based on the text ultra-high definition model.
[0346] Step 410: Generate a reference image based on the corrected text.
[0347] Step 411: If the first image includes a face, detect the face in the first image based on the face super-resolution model.
[0348] Step 412: Correct the pose of the face based on the ultra-high-definition face model.
[0349] Step 413: Generate a reference image based on the corrected face.
[0350] Step 414: Process the first image using the second image generation model.
[0351] The second image generation model can be an image generation model obtained by training a regular convolutional neural network. The second image generation model is weaker in supplementing detailed information in the first image.
[0352] Step 415, obtain the reference image.
[0353] Step 416: If there are anomalies in the reference image, the reference image is corrected based on the reference image.
[0354] It should be noted that an anomaly in the reference image can be determined when there is poor consistency between the reference image and the first image.
[0355] Step 417: Obtain the target image.
[0356] Figure 5 This is a method for obtaining a reference image according to an exemplary embodiment, such as... Figure 5 As shown, the method includes:
[0357] Step 501: Encode the first image using an encoder to obtain the feature parameters of the first image.
[0358] Step 502: Determine the random noise.
[0359] After performing step 502, as follows Figure 5 As shown, the first target feature parameters can be determined using a diffusion model based on the description vector, random noise, and first image feature parameters. When the diffusion model determines the first target feature parameters with a step size of n, the inference step size of the diffusion model can be shortened using a progressive distillation method.
[0360] Step 503: The decoder decodes the first target feature parameters output by the diffusion model.
[0361] Step 504: Obtain the reference image.
[0362] Figure 6 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 5 As shown, the method includes:
[0363] Step 601: Determine the first image.
[0364] Step 602: Determine the reference image.
[0365] For the first image and the reference image, perform the following steps 503 to 506 respectively:
[0366] Step 603: Determine the color content of the image.
[0367] Step 604: Determine the texture content of the image.
[0368] Step 605: Determine the type of objects in the image.
[0369] Step 606: Determine the text content in the image.
[0370] Step 607: Determine the similarity of image content in the first image and the second image by at least one of the following: color content, texture content, object type, and text content.
[0371] Step 608: Determine the reference image.
[0372] Step 609: If the similarity of the image content in the first image and the second image is less than the similarity threshold, the reference image and the baseline image are fused based on the fusion matrix.
[0373] Step 610: Obtain the target image.
[0374] In related technologies, images obtained in telephoto shooting scenarios exhibit significant issues such as smearing, distortion, and texture loss. Furthermore, the text-based image models in the ultra-high-resolution methods of these technologies have a large number of parameters and long iteration steps, making them difficult to deploy on mobile devices. These technologies also contain numerous obvious errors in images containing faces and / or text.
[0375] Based on this, such as Figures 2 to 6 As shown, this disclosure presents a generative ultra-resolution method with strong generation capabilities and high fidelity that can be applied to telephoto photography scenarios. This disclosure can generate high-fidelity images with low computational load. It should be noted that this solution is designed for situations where electronic devices have insufficient computing power, and can be selected for image processing in telephoto photography scenarios, while also being applicable to other image processing scenarios.
[0376] The disclosed solution has the following advantages:
[0377] 1. High Clarity: In telephoto photography scenarios, the high zoom ratio required for digital zoom can lead to severe distortion and blurring in the initial image being processed. In such cases, conventional ultra-high-resolution image processing methods struggle to generate clear, undegraded images. However, the image processing method disclosed herein introduces a large model with generative capabilities, overcoming the adverse effects of telephoto degradation and producing high-definition images with clear and detailed textures, thus improving the telephoto photography capabilities of mobile devices. The generative method proposed in this solution can significantly improve image clarity.
[0378] 2. High Fidelity: Generative models are highly random, easily producing obviously erroneous textures that negatively impact user experience. This solution employs multiple strategies to improve the accuracy of the generated results. For example, a general object detection model is used to extract semantic information from the image. The input image encoding is embedded during network initialization, and the image generation model is guided based on the semantic information of the first image, thus ensuring consistency between the generated target image and the content of the first image. For the generated results, salient texture, text, and object type verification are also performed to further ensure semantic consistency between the generated results and the input content. Furthermore, face and text scenes are processed separately to ensure the accuracy of the generated images in face and text scenes.
[0379] 3. Lightweight and easy to deploy on electronic devices: Current mainstream image generation models require around 1000 inference steps, which is time-consuming. Improving fidelity introduces large image-to-text models, making them difficult to deploy on electronic devices. This solution uses a lightweight object detection method to acquire descriptive information, significantly reducing the computational cost of information acquisition. Simultaneously, a progressive distillation method is used to extensively prune and compress the large model, reducing the inference step size. This achieves a speed increase of tens of times while retaining sufficient generation capabilities, ultimately enabling the inference process to be completed within seconds on electronic devices.
[0380] Figure 7 This is an image processing apparatus exemplarily shown according to an embodiment of the present disclosure, the apparatus comprising:
[0381] The processing module 71 is configured to process the first image based on at least two image processing strategies to obtain at least two second images; wherein the index parameters of the second images obtained based on different image processing strategies are different, and the index parameters of each second image are greater than the index parameters of the first image, and the index parameters are positively correlated with the performance of the image.
[0382] The acquisition module 72 is configured to obtain the target image based on at least two second images and the index parameters of at least two second images;
[0383] The index parameters of the target image are greater than those of the second image.
[0384] In one embodiment, the acquisition module 72 includes:
[0385] The determination module is configured to determine the second image that meets the first indicator condition as the reference image; wherein the indicator condition is used to indicate that: the resolution of the second image is greater than a preset resolution threshold, or the similarity between the first image and the second image is greater than a preset similarity threshold.
[0386] The second image whose index parameters meet the second index condition is determined as the reference image; wherein the first index condition and the second index condition are different.
[0387] The adjustment module is configured to adjust the reference image based on the reference image to obtain the target image.
[0388] In one embodiment, the processing module 71 is further configured to:
[0389] At least two sets of image feature parameters are extracted from the first image, and / or the image feature parameters extracted from the first image are compressed based on at least two compression ratios to obtain at least two sets of image feature parameters; wherein the image processing strategies corresponding to different sets of image feature parameters are different.
[0390] Based on the image processing strategies corresponding to the feature parameters of each group of images, the first image is processed to obtain at least two second images.
[0391] In one embodiment, the indicator parameters include the resolution of the image and the similarity between the first image and the second image; the resolution of each second image is greater than the resolution of the first image; the processing module 71 is further configured to:
[0392] Based on the feature parameters of the first image and the descriptive information of the first image, the first image is processed to obtain a second image with a resolution greater than a preset resolution threshold; wherein, the descriptive information is determined based on the type of each object in the first image and / or the relationship between each object.
[0393] Based on the second image feature parameters, the first image is processed to obtain a second image with a similarity greater than a preset similarity threshold;
[0394] Wherein, the first image feature parameter includes global features of the first image, the second image feature parameter includes local features of the second image; and / or, the data volume of the first image feature parameter is less than the data volume of the second image feature parameter.
[0395] In one embodiment, the determining module is further configured as follows:
[0396] Based on the type of each object in the first image, determine the processing parameters of the first image region where each object is located;
[0397] Processing module 71 is also configured as follows:
[0398] Based on the processing parameters of each first image region, the feature parameters of the first image, and the descriptive information of the first image, each first image region of the first image is processed to obtain a second image with a resolution greater than the resolution threshold.
[0399] In one embodiment, the determining module is further configured as follows:
[0400] If the first image includes an object of a first target type, determine the pose of the object of the first target type;
[0401] Based on the pose of the first target type, determine the processing parameters of the first image region where the object of the first target type is located;
[0402] There is a correspondence between the pose and processing parameters of the first target type.
[0403] In one embodiment, the determining module is further configured as follows:
[0404] Based on the types of each object in the first image and the relationships between each object, the first descriptive information is determined;
[0405] In response to the detection of a preset input operation, the first description information is adjusted based on the preset input operation to obtain the second description information;
[0406] Processing module 71 is also configured as follows:
[0407] The first image is processed based on the first image feature parameters and the second descriptive information.
[0408] In one embodiment, the processing module 71 is further configured to:
[0409] The associations in the first description information that do not meet the preset description conditions are deleted to obtain the processed first description information;
[0410] The adjustment module is also configured as follows:
[0411] The processed first description information is adjusted based on the preset input operation to obtain the second description information.
[0412] In one embodiment, the determining module is further configured as follows:
[0413] Based on the type of object in each first image region of the first image, select the region to be adjusted from each second image region in the reference image that matches the first image region;
[0414] The adjustment module is also configured to adjust the region to be adjusted in the reference image based on a third image region in the reference image that matches the region to be adjusted, so as to obtain the target image.
[0415] In one embodiment, the first indicator condition is used to indicate that the resolution of the reference image is greater than a resolution threshold, and the determining module is further configured to:
[0416] For each region of the first image, a second target type is determined from a preset type based on the type of the object in the first image region; wherein there is a correspondence between the object type and the preset type;
[0417] If a first image region of a first image matches a second image region of a reference image, determine the image content of a second target type in the first image region and the similarity between the image content of the second target type in the second image region.
[0418] The second image region in the baseline image with a similarity less than the similarity threshold is identified as the region to be adjusted.
[0419] Figure 8This is a structural block diagram illustrating an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0420] Reference Figure 8 The electronic device 800 may include one or more of the following components: processing component 802, memory 804, power supply component 806, multimedia component 808, audio component 810, input / output (I / O) interface 812, sensor component 814, and communication component 816.
[0421] Processing component 802 typically controls the overall operation of electronic device 800, such as operations associated with at least one of display, telephone call, data communication, camera operation, and recording operation. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.
[0422] Memory 804 is configured to store various types of data to support the operation of electronic device 800. Examples of such data include at least one of the following: instructions for any application or method operating on electronic device 800, contact data, phonebook data, messages, pictures, and videos. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0423] Power supply component 806 provides power to various components of electronic device 800. Power supply component 806 may include at least one of the following: a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 800.
[0424] Multimedia component 808 includes a screen that provides an output interface between electronic device 800 and user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen may be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When device 800 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0425] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when electronic device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.
[0426] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, and buttons. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0427] Sensor assembly 814 includes one or more sensors for providing state assessment of various aspects of electronic device 800. For example, sensor assembly 814 can detect the on / off state of electronic device 800, the relative positioning of components such as the display and keypad of electronic device 800, changes in position of electronic device 800 or one of its components, the presence or absence of user contact with electronic device 800, orientation or acceleration / deceleration of electronic device 800, and temperature changes of electronic device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include an optical sensor, such as a Complementary Metal Oxide Semiconductor (CMOS) or Charge Coupled Device (CCD) image sensor, for use in imaging applications.
[0428] In some embodiments, the sensor assembly 814 may also include, but is not limited to, at least one of the following: an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, and a temperature sensor.
[0429] Communication component 816 is configured to facilitate wired or wireless communication between electronic device 800 and other devices. Electronic device 800 can access wireless networks based on communication standards, such as Wi-Fi, 4G, 5G, or combinations thereof. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth (BT), and other technologies.
[0430] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components.
[0431] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including executable instructions or a computer program, which can be executed by a processor 820 of an electronic device 800 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0432] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a mobile electronic device, enables the mobile electronic device to perform any of the image processing methods described above in the embodiments of this disclosure. For example, the image processing method includes:
[0433] The first image is processed using at least two image processing strategies to obtain at least two second images. The second images obtained using different image processing strategies have different index parameters, and the index parameters of each second image are greater than those of the first image. The index parameters are positively correlated with the performance of the images.
[0434] The target image is obtained based on at least two second images and the index parameters of at least two second images;
[0435] The index parameters of the target image are greater than those of the second image.
[0436] This disclosure provides a computer program product comprising a computer program or executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program or executable instructions from the computer-readable storage medium and executes the computer program or executable instructions, causing the computer device to perform any of the image processing methods described above in this disclosure.
[0437] Figure 9 This is a block diagram illustrating a display device 900 according to an exemplary embodiment. For example, device 900 may be provided as a server. (Refer to...) Figure 9 The apparatus 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by memory 932 for storing instructions, such as application programs, that can be executed by the processing component 922. The application programs stored in memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 922 is configured to execute instructions to perform the aforementioned image processing method:
[0438] The first image is processed using at least two image processing strategies to obtain at least two second images. The second images obtained using different image processing strategies have different index parameters, and the index parameters of each second image are greater than those of the first image. The index parameters are positively correlated with the performance of the images.
[0439] The target image is obtained based on at least two second images and the index parameters of at least two second images;
[0440] The index parameters of the target image are greater than those of the second image.
[0441] Device 900 may also include a power supply component 926 configured to perform power management of device 900, a wired or wireless network interface 950 configured to connect device 900 to a network, and an input / output (I / O) interface 958. Device 900 can operate an operating system stored in memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.
[0442] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the foregoing claims.
[0443] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: The first image is processed using at least two image processing strategies to obtain at least two second images; wherein the index parameters of the second images obtained using different image processing strategies are different, and the index parameters of each second image are greater than the index parameters of the first image, and the index parameters are positively correlated with the performance of the image. The target image is obtained based on the at least two second images and the index parameters of the at least two second images; The index parameters of the target image are greater than those of the second image.
2. The image processing method according to claim 1, characterized in that, Obtaining the target image based on the at least two second images and the index parameters of the at least two second images includes: The second image whose index parameters satisfy the first index condition is determined as the reference image; wherein, the index condition is used to indicate that: the resolution of the second image is greater than a preset resolution threshold, or the similarity between the first image and the second image is greater than a preset similarity threshold. The second image whose index parameters satisfy the second index condition is determined as the reference image; wherein the first index condition and the second index condition are different; The target image is obtained by adjusting the reference image based on the reference image.
3. The image processing method according to claim 1 or 2, characterized in that, The process of processing the first image based on at least two image processing strategies to obtain at least two second images includes: At least two sets of image feature parameters are extracted from the first image, and / or the image feature parameters extracted from the first image are compressed based on at least two compression ratios to obtain at least two sets of image feature parameters; wherein the image processing strategies corresponding to different sets of image feature parameters are different. Based on the image processing strategies corresponding to the feature parameters of each group of images, the first image is processed to obtain at least two second images.
4. The image processing method according to claim 3, characterized in that, The index parameters include the resolution of the image and the similarity between the first image and the second image; the resolution of each of the second images is greater than the resolution of the first image; The image processing strategy based on the image feature parameters of each group processes the first image to obtain at least two second images, including: Based on the first image feature parameters and the first image description information, the first image is processed to obtain a second image with a resolution greater than a preset resolution threshold; wherein, the description information is determined based on the type of each object in the first image and / or the relationship between each object. Based on the second image feature parameters, the first image is processed to obtain a second image with a similarity greater than a preset similarity threshold; Wherein, the first image feature parameter includes global features of the first image, and the second image feature parameter includes local features of the second image; and / or, the data volume of the first image feature parameter is less than the data volume of the second image feature parameter.
5. The image processing method according to claim 4, characterized in that, The method further includes: Based on the type of each object in the first image, determine the processing parameters of the first image region where each object is located. The step of processing the first image based on the first image feature parameters and the description information of the first image to obtain a second image with a resolution greater than a resolution threshold includes: Based on the processing parameters of each of the first image regions, the first image feature parameters, and the description information of the first image, each of the first image regions of the first image is processed to obtain a second image with a resolution greater than the resolution threshold.
6. The image processing method according to claim 5, characterized in that, The process of determining the processing parameters of the first image region where each object is located based on the type of each object in the first image includes: If the first image includes an object of a first target type, determine the pose of the object of the first target type; Based on the pose of the first target type, determine the processing parameters of the first image region where the object of the first target type is located; There is a correspondence between the pose of the first target type and the processing parameters.
7. The image processing method according to claim 4, characterized in that, The method further includes: Based on the types of each object in the first image and the relationships between each object, first descriptive information is determined; In response to the detection of a preset input operation, the first description information is adjusted based on the preset input operation to obtain the second description information; The processing of the first image based on the first image feature parameters and the description information of the first image includes: The first image is processed based on the first image feature parameters and the second description information.
8. The image processing method according to claim 7, characterized in that, The method further includes: The associations in the first description information that do not meet the preset description conditions are deleted to obtain the processed first description information; The step of adjusting the first description information based on the preset input operation to obtain the second description information includes: The processed first description information is adjusted based on the preset input operation to obtain the second description information.
9. The image processing method according to claim 2, characterized in that, The step of adjusting the reference image based on the reference image to obtain the target image includes: Based on the type of object in each first image region of the first image, select the region to be adjusted from each second image region in the reference image that matches the first image region; Based on a third image region in the reference image that matches the region to be adjusted, the region to be adjusted in the reference image is adjusted to obtain the target image.
10. The image processing method according to claim 9, characterized in that, The first indicator condition is used to indicate that the resolution of the reference image is greater than the resolution threshold. The step of selecting the region to be adjusted from each of the second image regions in the reference image that match the first image regions, based on the type of objects in each first image region of the first image, includes: For each first image region of the first image, a second target type is determined from a preset type based on the type of the object in the first image region; wherein there is a correspondence between the object type and the preset type; If a first image region of the first image matches a second image region of the reference image, the image content of the second target type in the first image region and the similarity between the image content of the second target type in the second image region are determined. The second image region in the reference image whose similarity is less than the similarity threshold is determined as the region to be adjusted.
11. An image processing apparatus, characterized in that, The device includes: The processing module is configured to process a first image based on at least two image processing strategies to obtain at least two second images; wherein the index parameters of the second images obtained based on different image processing strategies are different, and the index parameters of each second image are greater than the index parameters of the first image, and the index parameters are positively correlated with the performance of the image; The acquisition module is configured to obtain a target image based on the at least two second images and the index parameters of the at least two second images; The index parameters of the target image are greater than those of the second image.
12. An electronic device, characterized in that, include: processor; Memory used to store computer programs or instructions; The processor executes the computer program or instructions to implement the steps of the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing a computer program or instructions, characterized in that, When the computer program or instructions in the storage medium are executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 10.