Image optimization method and device, equipment, storage medium and vehicle
By acquiring low-resolution images and using super-segment models for super-segment processing and detail repair, the problem of high-quality image generation cost in the prior art is solved, and resource saving and time efficiency are improved.
Patent Information
- Application Number
- CN202311757393.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, high-quality image generation requires high computer hardware resources, is time-consuming and has the possibility of memory overload, resulting in excessive image generation cost.
By acquiring low-resolution images, super-segment processing is performed using a preset super-segment model to obtain the image to be repaired, and the subject area of the image to be repaired is repaired to obtain the target image.
While ensuring image quality, it effectively reduces the cost of image generation and avoids excessive consumption of hardware resources and waste of time.
Smart Images

Figure CN120182137A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an image optimization method, device, equipment, storage medium and vehicle. Background Art
[0002] Text-to-image is a computer generation task aimed at converting text descriptions or natural language texts into corresponding images. In this task, the computer model needs to understand the prompt words input by the user and generate images that match the prompt words.
[0003] In the related art, if the user has a high requirement for the detail richness of the generated image, the computer model will directly convert the input prompt words into high-definition and high-richness images. For example, 1024*1024 high-definition images. However, directly converting the prompt words into high-definition and high-richness images by the computer model requires high hardware resources of the computer, and is extremely time-consuming. At the same time, there is a possibility of memory overload, so the cost of image generation is too high. Summary of the Invention
[0004] Embodiments of this application provide an image optimization method, device, equipment, storage medium and vehicle, which can solve the problem of high cost of existing high-quality image generation.
[0005] In a first aspect, embodiments of this application provide an image optimization method, and the method includes:
[0006] Obtain a low-resolution image, where the low-resolution image is converted from prompt words input by the user;
[0007] Perform super-resolution processing on the low-resolution image by using a pre-set super-resolution model to obtain an image to be repaired;
[0008] Perform detail repair on the main area of the image to be repaired to obtain a target image.
[0009] In some embodiments, before performing super-resolution processing on the low-resolution image by using the pre-set super-resolution model, the method further includes:
[0010] Obtain a super-resolution model to be trained and a real image set, where the real image set includes multiple real images with the same image resolution;
[0011] Train the super-resolution model by using the real image set until the loss function of the super-resolution model converges.
[0012] In some embodiments, training the super-resolution model by using the real image set until the loss function of the super-resolution model converges includes:
[0013] Downsample the real images in the real image set to obtain downsampled images;
[0014] Perform super-resolution processing on the downsampled images through the super-resolution model to obtain super-resolution images, where the image resolution of the super-resolution images is the same as that of the real images;
[0015] Determine the super-resolution images as the predicted real probabilities of the real images;
[0016] Adjust the model parameters of the super-resolution model based on the predicted real probabilities until the predicted real probabilities are greater than or equal to a preset target probability threshold, and determine that the loss function of the super-resolution model converges.
[0017] In some embodiments, the detailed repair of the main region of the image to be repaired to obtain a target image includes:
[0018] Perform semantic segmentation on the image to be repaired to determine the main region to be detailedly repaired in the image to be repaired;
[0019] Add noise to the main region of the image to be repaired to obtain an image to be processed;
[0020] According to the description information of the image to be repaired, perform noise reduction processing on the main region with added noise in the image to be processed to obtain a target image after detailed repair.
[0021] In some embodiments, the performing semantic segmentation on the image to be repaired to determine the main region to be detailedly repaired in the image to be repaired includes:
[0022] Perform semantic segmentation on the image to be repaired to identify at least one target object in the image to be repaired;
[0023] Determine the target object with the largest number of pixel points among the at least one target object as the main object;
[0024] Determine the region where the main object is located as the main region to be detailedly repaired in the image to be repaired.
[0025] In some embodiments, the adding noise to the main region of the image to be repaired to obtain an image to be processed includes:
[0026] Determine the non-main region in the image to be repaired, where the non-main region is the region outside the main region in the image to be repaired;
[0027] Adjust the pixel values of the pixel points in the non-main region of the image to be repaired to 0;
[0028] Obtain Gaussian noise with random initialization, and superimpose the Gaussian noise on the main area of the image to be restored after adjusting the pixel values to obtain an image to be processed.
[0029] In some embodiments, the denoising process for the main area of the image to be processed with added noise according to the description information of the image to be restored to obtain the target image after detailed restoration includes:
[0030] Encode the image to be processed to obtain an initial latent feature image;
[0031] Perform multiple iterative denoising processes on the initial latent feature image according to the description information of the image to be restored to obtain a restored latent feature image;
[0032] Decode the restored latent feature image to obtain the target image.
[0033] In some embodiments,
[0034] In some embodiments, the performing multiple iterative denoising processes on the initial latent feature image according to the description information of the image to be restored to obtain a restored latent feature image includes:
[0035] Encode the description information of the image to be restored to obtain a text embedding vector;
[0036] Embed the text embedding vector into an image generation model, and use the image generation model embedded with the text embedding vector to perform multiple iterative denoising processes on the initial latent feature image to obtain the restored latent feature image.
[0037] In a second aspect, an image optimization device provided by an embodiment of the present application includes:
[0038] A first acquisition module, configured to acquire a low-resolution image, where the low-resolution image is converted from a prompt word input by a user;
[0039] A super-resolution module, configured to perform super-resolution processing on the low-resolution image using a preset super-resolution model to obtain an image to be restored;
[0040] A repair module, configured to perform detailed repair on the main area of the image to be restored to obtain a target image.
[0041] In a third aspect, an image optimization device provided by an embodiment of the present application includes: a processor and a memory storing computer program instructions;
[0042] When the processor executes the computer program instructions, the above image optimization method is implemented.
[0043] Fourthly, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above image optimization method is implemented.
[0044] Fifthly, an embodiment of the present application provides a vehicle, which includes computer program instructions. When the computer program instructions are executed by a processor, the above image optimization method is implemented.
[0045] In an embodiment of the present application, a low-resolution image is obtained, where the low-resolution image is converted from a prompt word input by a user; the low-resolution image is super-resolved by using a preset super-resolution model to obtain an image to be repaired; the main area of the image to be repaired is repaired in detail to obtain a target image. That is to say, the prompt word input by the user can be first converted into a low-resolution image with relatively rough details, and then the low-resolution image is super-resolved to obtain an image to be repaired, and the main area of the image to be repaired is repaired in detail to improve the quality of the image. In this way, compared with the prior art that directly converts the prompt word into a picture with rich overall details by consuming a large amount of computing power, the prompt word can be first converted into a low-quality image with relatively low computing power, and then after simple super-resolution of the image, only the main part of the image is repaired in detail, which can effectively reduce the cost of image generation while ensuring the image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. Obviously, the following described drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0047] Figure 1 is a schematic flowchart of an image optimization method provided by an embodiment of the present application;
[0048] Figure 2 is a schematic structural diagram of an image optimization device provided by an embodiment of the present application;
[0049] Figure 3 is a schematic hardware structure diagram of an image optimization device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0051] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article or device comprising the elements.
[0052] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.
[0053] Specifically, to solve the problems of the prior art, the embodiments of the present application provide an image optimization method, device, equipment, storage medium, and vehicle. First, the image optimization method provided by the embodiments of the present application will be introduced below.
[0054] Figure 1 The flowchart of an image optimization method provided by an embodiment of the present application is shown. This method can be applied to the in-vehicle computer of a vehicle or a cloud server communicatively connected to the vehicle. The method includes the following steps:
[0055] S110, obtain a low-resolution image, where the low-resolution image is converted from a prompt word input by a user;
[0056] In this embodiment, a text-to-image model can be used to convert the prompt word input by the user into a low-resolution image, and the prompt word is used to describe the content of the low-resolution image that the user wants to generate. For example, the prompt word can be "generate a little dog" or "generate a big tree".
[0057] In addition, the input prompt words can include positive prompt words and negative prompt words. The meaning of positive prompt words is the content that is expected to be included in the target image, while negative prompt words indicate the content that is not expected to be included in the target image.
[0058] Exemplarily, the text-to-image model can be the base model in the latent diffusion model (Stable Diffusion model). A sampler with a relatively fast iteration speed can be loaded into the base model. The base model can convert the text content input into the model into a low-resolution image with insufficient details and a low resolution at a relatively fast speed.
[0059] S120, perform super-resolution processing on the low-resolution image using a pre-set super-resolution model to obtain an image to be repaired;
[0060] In this embodiment, after obtaining, the trained super-resolution model is used to perform super-resolution processing on the entire low-resolution image to initially improve the overall resolution of the low-resolution image and obtain an image to be repaired.
[0061] Exemplarily, the super-resolution model can be the Enhanced Super-Resolution Generative Adversarial Network (ESRGAN), and ESRGAN is used to perform 2X super-resolution on the low-resolution image. Specifically, ESRGAN can use deep learning techniques and generative adversarial networks to improve the spatial resolution of low-resolution images. 2X super-resolution can generate an image with a resolution twice that of the low-resolution image in both the horizontal and vertical directions by processing the low-resolution image, thereby improving the visual quality and details of the image. For example, the resolution of the low-resolution image is 512*512, and the resolution of the image to be repaired after super-resolution processing is 1024*1024.
[0062] As an optional embodiment, the super-resolution model can be a Residual in Residual Dense Block (RRDB) module without Batch Normalization (BN). Exemplarily, the RRDB module can include three interconnected Dense Blocks, and each Dense Block consists of five convolutional layers. These convolutional layers may have different filters and feature map depths for learning different levels of representations of the image. In this way, the super-resolution model has better generalization.
[0063] S130, perform detail repair on the main region of the image to be repaired to obtain the target image.
[0064] In this embodiment, the details of the super-resolved image to be repaired are still relatively rough. The main area of the image can be determined in the image to be repaired, and then the missing details in the main area can be restored or repaired, so as to improve the detail richness of the main area. Among them, the main area is the main part of the image.
[0065] In the embodiment of the present application, a low-resolution image is obtained, where the low-resolution image is converted from a prompt word input by the user; the low-resolution image is super-resolved by using a pre-set super-resolution model to obtain an image to be repaired; the details of the main area of the image to be repaired are repaired to obtain a target image. That is to say, the prompt word input by the user can be first converted into a low-resolution image with relatively rough details, and then the low-resolution image is super-resolved to obtain an image to be repaired, and the details of the main area in the image to be repaired are repaired to improve the quality of the image. In this way, compared with directly converting the prompt word into a picture with rich overall details at the cost of a large amount of computing power in the prior art, the prompt word can be first converted into an image with low quality through relatively low computing power, and then after simple super-resolution of the image, only the main part of the image is repaired in detail, which can effectively reduce the cost of image generation while ensuring the image quality.
[0066] As an optional embodiment, before using the pre-set super-resolution model to super-resolve the low-resolution image, the method further includes:
[0067] Obtain a super-resolution model to be trained and a set of real images, where the set of real images includes multiple real images with the same image resolution;
[0068] Use the set of real images to train the super-resolution model until the loss function of the super-resolution model converges.
[0069] In this embodiment, before applying the super-resolution model to the super-resolution of the low-resolution image, the model parameters of the super-resolution model need to be initialized first, and then the super-resolution model can be trained sequentially with multiple real images with the same resolution until the loss function of the super-resolution model converges to a certain value or stabilizes, and it can be considered that the training of the super-resolution model is completed.
[0070] As an optional embodiment, using the set of real images to train the super-resolution model until the loss function of the super-resolution model converges includes:
[0071] Perform downsampling on the real images in the set of real images to obtain downsampled images;
[0072] Super-resolve the downsampled images through the super-resolution model to obtain super-resolved images, and the image resolution of the super-resolved images is the same as the image resolution of the real images;
[0073] Determine the predicted true probability that the super-resolution image is the true image;
[0074] Based on the predicted true probability, adjust the model parameters of the super-resolution model until the predicted true probability is greater than or equal to a preset target probability threshold, and determine that the loss function of the super-resolution model converges.
[0075] In this embodiment, the super-resolution model can be a residual in-residual dense block without batch normalization. During the training of the super-resolution model, a set of true images composed of multiple true images with the same resolution can be obtained, and then the fake image corresponding to each true image, that is, the super-resolution image, can be obtained, and then the super-resolution model can be trained using the super-resolution image and the true image.
[0076] Specifically, taking the first true image as an example, the first downsampled image corresponding to the first true image can be obtained by performing downsampling processing on the first true image. The resolution of the first downsampled image is lower than that of the first true image, but the first true image and the first downsampled image match in content.
[0077] Then, the super-resolution model can be used to perform super-resolution processing on the first downsampled image to obtain the first super-resolution image. The first super-resolution image is a fake image predicted from the first downsampled image to the first true image, and the resolution of the first super-resolution image is the same as that of the first true image. Subsequently, the first super-resolution image and the corresponding first true image can be compared to calculate the predicted true probability that the first super-resolution image is the true image, and the discriminator loss and generator loss in the super-resolution model can be calculated from this predicted true probability. And through iterative training, the probability that the super-resolution image is the true image can be increased, thereby reducing the discriminator loss and generator loss and improving the performance of the model until the discriminator loss and generator loss converge, then it is considered that the super-resolution model has completed training.
[0078] Specifically, the predicted true probability of the super-resolution model can be calculated by the following formula:
[0079]
[0080] where x r is the first true image, x f is the predicted first super-resolution image, C(·) is the raw output of the discriminator before activation, E x (·) is the expectation of the predicted distribution of the super-resolution image, and σ is a scaling factor.
[0081] As an alternative embodiment, the detailed repair of the main region of the image to be repaired to obtain the target image includes:
[0082] Perform semantic segmentation on the image to be repaired to determine the main area in the image to be repaired that requires detailed repair;
[0083] Add noise to the main area of the image to be repaired to obtain an image to be processed;
[0084] According to the description information of the image to be repaired, perform noise reduction on the main area with added noise in the image to be processed to obtain the target image after detailed repair.
[0085] In this embodiment, semantic segmentation can label the pixels in the image as belonging to specific categories to identify and distinguish different objects, structures, or regions in the image to be repaired. Specifically, the area that is centered and has the largest coverage area in the image to be repaired can be determined as the main area of the image to be repaired.
[0086] Exemplarily, the image to be repaired can be semantically segmented into multiple regions by the SAM (Segment Anything Model) segmentation model, and the region that is centered and has the largest coverage area among the multiple regions can be determined as the main area.
[0087] After determining the main area of the image to be repaired, a pre-set random number seed can be obtained, and the random number seed can be converted into an initialized Gaussian noise, and then the Gaussian noise can be added to each pixel in the main area, that is, an image to be processed with blurred noise on the main area can be obtained.
[0088] After adding noise to the main area in the image to be processed, the main area with added noise can be denoised to complete the repair of the image. Specifically, the image to be processed with added noise can be input into an image-to-image model. The image-to-image model can predict the noise on the image to be processed and use the loaded sampler to denoise the predicted noise, so as to obtain the target image after detailed repair.
[0089] Exemplarily, the image-to-image model can be the base model in the latent diffusion model (Stable Diffusion model).
[0090] In addition, the description information of the image to be repaired can be used to guide the image-to-image model to predict and remove the noise on the image to be processed. The description information of the image to be repaired can characterize the image content of the image to be repaired. Specifically, computer vision extraction techniques can be used to extract the features of the image to be repaired, and then these features can be converted into natural language descriptions, that is, the description information of the image to be repaired can be obtained. The prompt words input by the user can also be directly used as the description information of the image to be repaired.
[0091] In an embodiment of the present application, a to-be-repaired image is obtained. By performing semantic segmentation on the to-be-repaired image, the main region of the to-be-repaired image is determined based on the result of the semantic segmentation; noise is added to the main region of the to-be-repaired image; and noise reduction processing is performed on the main region of the to-be-repaired image with added noise to obtain a target image. That is to say, a low-resolution image can be first converted into a to-be-repaired image with relatively rough details, and then by performing semantic segmentation on the to-be-repaired image, the main region of the to-be-repaired image is determined, and only the main region is subjected to noise addition and noise reduction to enrich the details of the main region in the image and improve the quality of the image. In this way, compared with directly converting a prompt into an overall image with rich details at the cost of a large amount of computing power in the prior art, the prompt can be first converted into a to-be-repaired image with low quality by relatively low computing power, and then only the main part of the to-be-repaired image is subjected to detail repair, effectively reducing the cost of image generation while ensuring the image quality.
[0092] As an optional embodiment, the performing semantic segmentation on the to-be-repaired image to determine the main region to be detail-repaired in the to-be-repaired image includes:
[0093] Performing semantic segmentation on the to-be-repaired image to identify at least one target object in the to-be-repaired image;
[0094] Determining the target object with the largest number of pixel points among the at least one target object as the main object;
[0095] Determining the region where the main object is located as the main region to be detail-repaired in the to-be-repaired image.
[0096] In this embodiment, by performing semantic segmentation on the to-be-repaired image, a semantic label can be assigned to each pixel point in the to-be-repaired image, such as "person", "vehicle", "road". Each semantic label is used to represent an object category. All pixel points with the same semantic label and a connection relationship can be determined as the same target object.
[0097] If there is only one target object in the to-be-repaired image, the target object can be directly determined as the main object in the to-be-repaired image; if there are multiple target objects in the to-be-repaired image, the target object with the largest number of pixel points among the multiple target objects can be determined as the main object.
[0098] Then, the region where the main object is located can be determined as the main region in the to-be-repaired image. In this way, the main region can be accurately and quickly determined in the to-be-repaired image.
[0099] As an optional embodiment, the adding noise to the main region of the to-be-repaired image to obtain an image to be processed includes:
[0100] Determine the non-subject area in the image to be repaired, where the non-subject area is the area outside the subject area in the image to be repaired;
[0101] Adjust the pixel values of the pixel points in the non-subject area of the image to be repaired to 0;
[0102] Obtain randomly initialized Gaussian noise, and superimpose the Gaussian noise on the subject area of the image to be repaired after adjusting the pixel values to obtain an image to be processed.
[0103] In this embodiment, after determining the subject area in the image to be repaired, the area outside the subject area in the image to be repaired can be determined as the non-subject area. Since in the image to be repaired, the subject area is the area where the subject object is located in the image, and the subject object is the object that needs to be prominently shown in the image, the subject area has relatively high requirements for the picture quality and image details, while the non-subject area has relatively low requirements for the picture quality and image details.
[0104] In order to save the computing power of the model, only the subject area in the image to be repaired needs to be repaired in detail, and the non-subject area does not need to be repaired in detail. Therefore, by masking the non-subject area, the pixel values of the pixel points in the subject area are not adjusted, and the values of all pixel points in the non-subject area are set to 0.
[0105] After adjusting the pixel values of the non-subject area, when adding noise to the image to be repaired, the masked non-subject area will not be affected by the noise. That is to say, when adding noise to the masked image to be repaired, the noise will only be added to the subject part of the image to be repaired.
[0106] In the above way, it is possible to quickly and accurately add noise only to the subject area in the image to be repaired.
[0107] As an optional embodiment, the step of performing noise reduction processing on the subject area where noise is added to the image to be processed according to the description information of the image to be repaired to obtain a target image with detailed repair includes:
[0108] Encode the image to be processed to obtain an initial latent feature image;
[0109] According to the description information of the image to be repaired, perform iterative noise reduction processing on the initial latent feature image multiple times to obtain a repaired latent feature image;
[0110] Decode the repaired latent feature image to obtain the target image.
[0111] In this embodiment, after adding noise to the image to be repaired after mask processing, an image with blurred noise, that is, the image to be processed, can be obtained. The blurred area in the image to be processed is the main area of the image. Subsequently, the image to be processed with added noise can be encoded by an autoencoder, that is, the image to be processed with added noise is mapped into the latent space to generate an initial latent feature image.
[0112] Subsequently, the initial latent feature image can be input into the base model of the latent diffusion model (Stable Diffusion model). Since the non-main areas have been masked, after loading a suitable sampler in the base model, only the noise of the main area can be predicted, and based on the predicted result, noise reduction processing is performed on the main area in the initial latent feature image to obtain a repaired latent feature image.
[0113] After obtaining the repaired latent feature image, a variational autoencoder can be used to decode the repaired latent feature image to obtain a visible target image.
[0114] In this embodiment, by setting the pixel values of the areas outside the main area that require detail repair and super-resolution to zero, the non-main areas do not participate in the noise addition and noise reduction processes, and only a small amount of computing power is required to complete the detail enrichment of the specified area.
[0115] As an alternative embodiment, the repeatedly performing noise reduction processing on the initial latent feature image according to the description information of the image to be repaired to obtain a repaired latent feature image includes:
[0116] Encoding the description information of the image to be repaired to obtain a text embedding vector;
[0117] Embedding the text embedding vector into an image generation model, and using the image generation model embedding the text embedding vector to perform multiple iterations of noise reduction processing on the initial latent feature image to obtain the repaired latent feature image.
[0118] In this embodiment, before performing detail repair on the image to be repaired, a pre-trained image description generation model can be obtained. The image description generation model has been trained with a large-scale image corpus and can generate relatively accurate description information of the image.
[0119] After inputting the image to be repaired into the image description generation model, the image description generation model will output a text content that describes the main content contained in the image to be repaired. This description can include information such as objects, scenes, actions, etc. Therefore, the output text content can be determined as the description information of the image to be repaired.
[0120] The text content describing information can be mapped to a high-dimensional vector space to obtain a text embedding vector. Then, the initial latent feature image and the text embedding vector can be input into an image generation model, and the text embedding vector is used to guide the image generation model to perform multiple iterations of noise reduction processing on the initial latent feature image to obtain a repaired latent feature image.
[0121] In this embodiment, the description information obtained by describing the image to be repaired can be used to guide the accurate detail repair of the main region in the image to be repaired.
[0122] Based on the image optimization method provided in the above embodiment, correspondingly, the present application also provides a specific implementation manner of an image optimization device. Please refer to the following embodiments.
[0123] First, refer to Figure 2 , the image optimization device 200 provided in the embodiment of the present application includes the following modules:
[0124] The first acquisition module 201 is configured to acquire a low-resolution image, where the low-resolution image is converted from a prompt word input by a user;
[0125] The super-resolution module 202 is configured to perform super-resolution processing on the low-resolution image by using a preset super-resolution model to obtain an image to be repaired;
[0126] The repair module 203 is configured to perform detail repair on the main region of the image to be repaired to obtain a target image.
[0127] The device can acquire a low-resolution image, where the low-resolution image is converted from a prompt word input by a user; perform super-resolution processing on the low-resolution image by using a preset super-resolution model to obtain an image to be repaired; perform detail repair on the main region of the image to be repaired to obtain a target image. That is to say, the prompt word input by the user can be first converted into a low-resolution image with relatively rough details, and then the low-resolution image is subjected to super-resolution processing to obtain an image to be repaired, and detail repair is performed on the main region of the image to be repaired to improve the quality of the image. In this way, compared with directly converting the prompt word into a picture with rich overall details by consuming a large amount of computing power in the prior art, the prompt word can be first converted into an image with low quality by using relatively low computing power, and then the image is simply super-resolved, and only the main part of the image is subjected to detail repair, which can effectively reduce the cost of image generation while ensuring the image quality.
[0128] As an implementation manner of the present application, the above image optimization device 200 may further include:
[0129] A second acquisition module, configured to acquire a super-resolution model to be trained and a set of real images, where the set of real images includes multiple real images with the same image resolution;
[0130] A training module, configured to train the super-resolution model using the set of real images until the loss function of the super-resolution model converges.
[0131] As an implementation manner of this application, the above training module may further include:
[0132] A downsampling unit, configured to perform downsampling processing on the real images in the set of real images to obtain downsampled images;
[0133] A super-resolution unit, configured to perform super-resolution processing on the downsampled images through the super-resolution model to obtain super-resolution images, where the image resolution of the super-resolution images is the same as the image resolution of the real images;
[0134] A determination unit, configured to determine the predicted real probability of the super-resolution images as the real images;
[0135] A parameter adjustment unit, configured to adjust the model parameters of the super-resolution model based on the predicted real probability until the predicted real probability is greater than or equal to a preset target probability threshold, and determine that the loss function of the super-resolution model converges.
[0136] As an implementation manner of this application, the above repair module 203 may further include:
[0137] A segmentation unit, configured to perform semantic segmentation on the image to be repaired to determine the main body area to be repaired in detail in the image to be repaired;
[0138] A noise addition unit, configured to add noise to the main body area of the image to be repaired to obtain an image to be processed;
[0139] A noise reduction unit, configured to perform noise reduction processing on the main body area with added noise in the image to be processed according to the description information of the image to be repaired to obtain a target image after detailed repair.
[0140] As an implementation manner of this application, the above segmentation unit may further include:
[0141] A segmentation subunit, configured to perform semantic segmentation on the image to be repaired to identify at least one target object in the image to be repaired;
[0142] A first determination subunit, configured to determine the target object with the largest number of pixel points in the at least one target object as the main body object;
[0143] A second determination subunit, configured to determine the area where the main object is located as the main area to be detailed-repaired in the image to be repaired.
[0144] As an implementation manner of the present application, the above-mentioned noise addition unit may further include:
[0145] A third determination subunit, configured to determine the non-main area in the image to be repaired, where the non-main area is the area outside the main area in the image to be repaired;
[0146] An adjustment subunit, configured to adjust the pixel values of the pixel points in the non-main area of the image to be repaired to 0;
[0147] A noise addition subunit, configured to obtain randomly initialized Gaussian noise and superimpose the Gaussian noise on the main area of the image to be repaired after adjusting the pixel values to obtain an image to be processed.
[0148] As an implementation manner of the present application, the above-mentioned noise reduction unit may further include:
[0149] An encoding subunit, configured to encode the image to be processed to obtain an initial latent feature image;
[0150] A noise reduction subunit, configured to perform multiple iterative noise reduction processes on the initial latent feature image according to the description information of the image to be repaired to obtain a repaired latent feature image;
[0151] A decoding subunit, configured to decode the repaired latent feature image to obtain the target image.
[0152] As an implementation manner of the present application, the above-mentioned noise reduction subunit may also be used to:
[0153] Encode the description information of the image to be repaired to obtain a text embedding vector;
[0154] Embed the text embedding vector into an image generation model, and use the image generation model embedded with the text embedding vector to perform multiple iterative noise reduction processes on the initial latent feature image to obtain the repaired latent feature image.
[0155] The image optimization device provided by the embodiments of the present invention can implement each step in the above method embodiments. To avoid repetition, it will not be elaborated here.
[0156] Figure 3 FIG. shows a schematic hardware structure diagram of an image optimization device provided by an embodiment of the present application.
[0157] The image optimization device may include a processor 301 and a memory 302 storing computer program instructions.
[0158] Specifically, the above-mentioned processor 301 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.
[0159] The memory 302 may include a mass storage for data or instructions. By way of example and not limitation, the memory 302 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 302 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 302 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 302 is a non-volatile solid-state memory.
[0160] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0161] The processor 301 reads and executes the computer program instructions stored in the memory 302 to implement any one of the image optimization methods in the above embodiments.
[0162] In one example, the image optimization device may further include a communication interface 303 and a bus 310. Among them, as Figure 3 shown, the processor 301, the memory 302, and the communication interface 303 are connected through the bus 310 and complete communication with each other.
[0163] The communication interface 303 is mainly used to implement communication between the various modules, devices, units, and / or devices in the embodiments of the present application.
[0164] The bus 310 includes hardware, software, or both, and couples the components of the image optimization device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable bus or a combination of two or more of these. Where appropriate, the bus 310 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0165] The image optimization device may be based on the above embodiments, so as to implement the image optimization method and device in combination with the above.
[0166] In addition, in combination with the image optimization method in the above embodiments, the embodiments of the present application may provide a computer storage medium for implementation. Computer program instructions are stored on the computer storage medium; when the computer program instructions are executed by a processor, any one of the image optimization methods in the above embodiments is implemented, and the same technical effects can be achieved. To avoid repetition, details are not described here again. Among them, the above computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc., which are not limited herein.
[0167] In addition, the embodiments of the present application further provide a vehicle, including computer program instructions, and when the computer program instructions are executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0168] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0169] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.
[0170] It should also be noted that in the exemplary embodiments mentioned in the present application, some methods or systems are described based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, can be different from the order in the embodiments, or several steps can be executed simultaneously.
[0171] As described above with reference to the flowcharts and / or block diagrams of methods, apparatuses, and vehicles according to embodiments of the present disclosure. It should be understood that each block in the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to generate a machine such that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more blocks of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can also be implemented by dedicated hardware that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0172] The above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.
Claims
1. An image optimization method, characterized in that, The method includes: Obtain a low-resolution image, where the low-resolution image is converted from a prompt word input by a user; Perform super-resolution processing on the low-resolution image using a pre-set super-resolution model to obtain an image to be repaired; Perform detail repair on the main area of the image to be repaired to obtain a target image.
2. The image optimization method according to claim 1, characterized in that, Before performing super-resolution processing on the low-resolution image using the pre-set super-resolution model, the method further includes: Obtain a super-resolution model to be trained and a set of real images, where the set of real images includes multiple real images with the same image resolution; Train the super-resolution model using the set of real images until the loss function of the super-resolution model converges.
3. The image optimization method according to claim 2, characterized in that, The training of the super-resolution model using the set of real images until the loss function of the super-resolution model converges includes: Perform downsampling processing on the real images in the set of real images to obtain downsampled images; Perform super-resolution processing on the downsampled images through the super-resolution model to obtain super-resolution images, where the image resolution of the super-resolution images is the same as the image resolution of the real images; Determine the predicted real probability of the super-resolution images as the real images; Adjust the model parameters of the super-resolution model based on the predicted real probability until the predicted real probability is greater than or equal to a pre-set target probability threshold, and determine that the loss function of the super-resolution model converges.
4. The image optimization method according to claim 1, characterized in that, The performing detail repair on the main area of the image to be repaired to obtain a target image includes: Perform semantic segmentation on the image to be repaired to determine the main area to be detail-repaired in the image to be repaired; Add noise to the main area of the image to be repaired to obtain an image to be processed; Perform noise reduction processing on the main area with added noise in the image to be processed according to the description information of the image to be repaired to obtain a target image after detail repair.
5. The image optimization method according to claim 4, characterized in that, The performing semantic segmentation on the image to be repaired to determine the main area to be detail-repaired in the image to be repaired includes: Perform semantic segmentation on the image to be repaired to identify at least one target object in the image to be repaired; Determine the target object with the largest number of pixel points among the at least one target object as the main object; Determine the area where the main object is located as the main area to be detail-repaired in the image to be repaired.
6. The image optimization method according to claim 4, characterized in that, The adding noise to the main area of the image to be repaired to obtain an image to be processed includes: Determine the non-main area in the image to be repaired, where the non-main area is the area outside the main area in the image to be repaired; Adjust the pixel values of the pixel points in the non-main area of the image to be repaired to 0; Obtain randomly initialized Gaussian noise and superimpose the Gaussian noise on the main area of the image to be repaired after adjusting the pixel values to obtain an image to be processed.
7. The image optimization method according to claim 4, characterized in that, The performing noise reduction processing on the main area with added noise in the image to be processed according to the description information of the image to be repaired to obtain a target image after detail repair includes: Encode the image to be processed to obtain an initial latent feature image; Performing multiple iterations of noise reduction processing on the initial latent feature image according to the description information of the image to be repaired to obtain a repaired latent feature image; Decoding the repaired latent feature image to obtain the target image.
8. The image optimization method according to claim 7, characterized in that, The performing multiple iterations of noise reduction processing on the initial latent feature image according to the description information of the image to be repaired to obtain a repaired latent feature image includes: Encoding the description information of the image to be repaired to obtain a text embedding vector; Embedding the text embedding vector into an image generation model, and using the image generation model embedded with the text embedding vector to perform multiple iterations of noise reduction processing on the initial latent feature image to obtain the repaired latent feature image.
9. An image optimization device, characterized in that, The apparatus includes: A first acquisition module, configured to acquire a low-resolution image, where the low-resolution image is converted from a prompt word input by a user; A super-resolution module, configured to perform super-resolution processing on the low-resolution image by using a preset super-resolution model to obtain an image to be repaired; A repair module, configured to perform detail repair on the main region of the image to be repaired to obtain a target image.
10. An image optimization device, characterized in that, The image optimization device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image optimization method according to any one of claims 1-8 is implemented.
11. A computer storage medium, characterized in that, Computer program instructions are stored on the computer storage medium, and when the computer program instructions are executed by the processor, the image optimization method according to any one of claims 1-8 is implemented.
12. A vehicle, characterized in that, Including at least one of the following: The image optimization apparatus according to claim 9; The image optimization device according to claim 10; The computer storage medium according to claim 11.