Image generation method and device
By integrating the first image texture features with high definition and the second image profile features with low definition in the mobile terminal, the problem of image details loss and low definition in ultra-telephoto shooting of the mobile terminal is solved, and the clarity and reality of the image are improved.
Patent Information
- Application Number
- CN202510500127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing mobile terminal ultra-telephoto shooting technology results in serious loss of image details and low clarity, making it impossible to effectively reconstruct complex textures, and it is easy to produce shadows when shooting moving targets, and has limited adaptability.
By acquiring a first image with high definition and a second image with low definition, the image generation model is used to fuse the texture features of the first image with the contour features of the second image to generate a target image with higher definition.
On the basis of maintaining the image outline, the clarity, detail and reality of the image are significantly improved, the imaging quality is improved, and the needs of different shooting scenes and users.
Smart Images

Figure CN120339090A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to an image generation method and apparatus. Background Art
[0002] In the field of mobile terminal photography, ultra-long focal length shooting usually relies on the method of cropping and zooming (cropping and magnifying a local area of the image). Due to the limitation of the mobile terminal size, it is impossible to be equipped with a large ultra-long focal length lens like a single-lens reflex camera. Therefore, when shooting scenes that require high magnification zoom, such as birds, the imaging quality is often poor. Although the existing ultra-long focal length shooting technology of mobile terminals can meet the user's demand for long-distance shooting to a certain extent, due to the small sensor size, the resolution, details, and clarity of the image after cropping and magnifying will be greatly affected. The existing long focal length shooting technology has the following defects: poor ability to reconstruct complex textures, easy to generate false details, resulting in serious loss of image details and low clarity. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide an image generation method and apparatus, which can solve the problem that the related shooting technology is prone to cause serious loss of image details and low clarity.
[0004] In a first aspect, the embodiments of this application provide an image generation method, including:
[0005] Obtain a first image and a second image, where the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than the clarity of the second image;
[0006] Input the first image and the second image into an image generation model to obtain a first fused feature image, where the image features of the first fused feature image include the texture features of the first image and the contour features of the second image;
[0007] Based on the first fused feature image, obtain a target image corresponding to the second image, where the clarity of the target image is greater than the clarity of the second image.
[0008] In a second aspect, the embodiments of this application provide an image generation apparatus, including:
[0009] A first acquisition module, configured to obtain a first image and a second image, where the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than the clarity of the second image;
[0010] The first processing module is configured to input the first image and the second image into an image generation model to obtain a first fused feature image, where the image features of the first fused feature image include the texture features of the first image and the contour features of the second image;
[0011] The second obtaining module is configured to obtain a target image corresponding to the second image based on the first fused feature image, and the clarity of the target image is greater than that of the second image.
[0012] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can be run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0014] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.
[0015] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.
[0016] In the embodiment of the present application, two first images and second images with different clarity and matching attribute information are input into an optimized image generation model. The image generation model fuses the texture features of the first image with high clarity and the contour features of the second image with low clarity. Therefore, the first fused feature image obtained enhances the texture details while ensuring the contour of the second image. Therefore, the target image obtained based on the first fused feature image improves the clarity, details, and realism while maintaining the contour of the second image. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flowchart of an image generation method provided by an embodiment of the present application;
[0018] Figure 2 is a schematic structural diagram of a YOLO model provided by an embodiment of the present application;
[0019] Figure 3 is a schematic flowchart of the extraction process of image features provided by an embodiment of the present application;
[0020] Figure 4 It is a schematic flowchart of the specific process for generating a predicted image provided by an embodiment of the present application;
[0021] Figure 5 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;
[0022] Figure 6 It is a structural block diagram of an electronic device provided by an embodiment of the present application;
[0023] Figure 7 It is a structural block diagram of another electronic device provided by an embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope protected by the present application.
[0025] The terms "first", "second", etc. in the specification of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.
[0026] The existing long - focal - length shooting technologies have the following defects:
[0027] Poor imaging quality: The cropping and magnification method of ultra - long - focal - length shooting results in serious loss of image details and low clarity;
[0028] Poor interaction when shooting moving targets: When shooting with ultra - long - focal - length, the field of view (FOV) of the camera is small. When the shooting target is moving, it is very difficult to capture the picture, and the motion blur of the target is very serious when the picture is captured;
[0029] Lack of realism: The generated images have differences from the actual high - definition images in terms of color, light and shadow, and texture;
[0030] Limited adaptability: It cannot be flexibly adjusted according to different shooting scenarios and user requirements, and has insufficient processing ability for different types of targets or complex backgrounds.
[0031] Therefore, the embodiments of the present application provide an image generation method and apparatus. The low-quality image and high-definition image of the target object are input into the optimized image generation model. The image generation model fuses the texture features of the first image with high clarity and the contour features of the second image with low clarity, so that the generated target image improves clarity, details, and realism while maintaining the contour of the second image.
[0032] The following will describe the image generation method provided by the embodiments of the present application in conjunction with the accompanying drawings through specific embodiments and their application scenarios.
[0033] As Figure 1 shown, the embodiments of the present application provide an image generation method, which may specifically include the following steps:
[0034] Step 101, obtain a first image and a second image, where the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than the clarity of the second image.
[0035] Obtain a first image and a second image with matching attribute information. The first image and the second image are two images with different clarities, and both images include the target object.
[0036] The difference between the first image and the second image is as follows: The first image has higher clarity, higher pixel density, and more pixels per unit area, and can clearly present minute details and textures; the second image has lower clarity, fewer pixels and lower density, and is prone to blurring, jaggedness, or mosaic phenomena when magnified. Moreover, the first image can accurately restore complex textures and layering, with sharp edges and no blurring; the second image has blurred details, low distinguishability between the target and the background, and is prone to information loss due to noise or compression.
[0037] Step 102, input the first image and the second image into the image generation model to obtain a first fusion feature image, where the image features of the first fusion feature image include the texture features of the first image and the contour features of the second image.
[0038] The image generation model is a multi-modal hybrid guided generation model that adopts the large framework of Stable Diffusion and ControlNet. Stable Diffusion is a text-to-image model, and ControlNet guides the generation process by adding additional control conditions, such as edge detection, depth maps, etc., so as to more precisely control the output results.
[0039] Input the first image with higher clarity and the second image with lower clarity into an image generation model. Through the image generation model, fuse the texture features of the first image and the contour features of the second image to obtain a first fused feature image. Thus, the first fused feature image retains the contour features of the second image, enhancing the texture of the second image and ensuring a significant improvement in the clarity, details, and realism of the generated image.
[0040] Step 103: Obtain a target image corresponding to the second image based on the first fused feature image, where the clarity of the target image is greater than that of the second image.
[0041] Output the first fused feature image as the target image corresponding to the second image. The clarity of this target image is greater than that of the second image. Thus, the target image improves in clarity, details, and realism while maintaining the contour of the second image.
[0042] In the above embodiments of the present application, input the first image and the second image with different clarities and matching attribute information into the optimized image generation model. Through the image generation model, fuse the texture features of the first image with high clarity and the contour features of the second image with low clarity. Thus, the obtained first fused feature image enhances the texture features while ensuring the contour of the second image. Therefore, the target image obtained based on the first fused feature image improves in clarity, details, and realism while maintaining the contour of the second image.
[0043] In an optional specific embodiment, step 101 of obtaining the first image and the second image includes:
[0044] Obtain the captured second image, where the second image includes a target object;
[0045] According to the attribute information of the second image, screen out the first image that matches the attribute information of the second image from the image database. The attribute information includes: geographical location information of the shooting location and category information of the target object.
[0046] First, capture the second image, which includes a target object. Based on the second image, obtain attribute information such as the geographical location information of the shooting location of the second image and the category information of the target object in the second image. Traverse the geographical location information of the shooting location and the category information of the target object of each image in the image database, and then screen out the first image with matching geographical location information and matching type information of the target object from the image database.
[0047] It should be noted that when the user takes the second image, there is no need to frequently move the terminal manually to find the target object. When the target detection model locates the target object within the navigation frame, it will automatically follow the target object and adjust the Region of Interest (ROI) of the cropping frame according to digital zoom to achieve dynamic cropping. In addition, when the target object is detected, if the movement amplitude of the target object is large, if a conventional exposure meter is used, the main body of the photographed target object usually has obvious motion blur. Therefore, adjust the exposure value or switch to the motion capture mode according to the motion amplitude of the target object main body (i.e., the quantization index shaking value). Because the motion of the target object under the ultra-long focal length FOV will be magnified and is more obvious than at low magnification, the shaking value is multiplied by the value of digital zoom to equal the motion index of the motion amplitude at this focal length. The exposure meter follows the exposure settings of the long focal point, and adjusts the exposure value according to the calculated motion index. The larger the motion index, the shorter the exposure time and the smaller the motion blur.
[0048] In one embodiment, taking the ultra-long focal length preview photo-taking screen as an example, design a target detection algorithm exclusive to the target object to locate the target object to be photographed in real time. The detection model selects the You Only Look Once (YOLO) series of models, which are fast and accurate. The model structure is as Figure 2 shown. The model receives the preview screen image in real time and outputs the detection frame and category of the target object in the preview frame.
[0049] After the image to be detected (such as 512*512) enters the YOLO model, it will first enter the Backbone part of the main network, which is responsible for extracting useful features from the input image. It is usually a convolutional neural network trained in large-scale image classification tasks. The backbone network captures hierarchical features at different scales, extracts low-level features (such as edges and textures) in earlier layers, and extracts high-level features (such as object parts and semantic information) in deeper layers. Corresponding to the P1 to P5 layers of the Backbone in the figure, the resolutions of the feature maps are 256*256, 128*128, 64*64, 32*32, and 16*16 respectively. The resolution of the feature map is getting lower and lower, and the semantic information it represents is more advanced and global. The Neck layer is an intermediate component connecting the Backbone and the Head. It aggregates and refines the features extracted by the backbone network, usually focusing on strengthening the spatial and semantic information at different scales. It includes additional convolutional layers, feature pyramid networks, etc. to improve the representativeness of the features. Corresponding Figure 2P3, P4, and P5 in the [Chinese context] respectively enter the Neck layer after convolution to obtain N3, N4, and N5. The features of N5 and N4 will be upsampled and fused with the features of N4 and N3 respectively to obtain multi-scale feature representations. Head is the last component of the object detection model, which is responsible for making predictions based on the features provided by the Backbone and Neck. It usually consists of one or more sub-networks for specific tasks, performing tasks such as classification and localization. As Figure 2 shown, the features of N5, N4, and N3 in the Neck layer then enter three detection heads (Head1, Head2, and Head3) respectively, corresponding to output detection results of different scales and sizes. The output results are absolute coordinates (i.e., coordinates in the image) and categories.
[0050] The training dataset required for training the object detection model uses the images in the image database and a part of the synthetic dataset to augment the training dataset. For example: paste the classification dataset images of CUB-200 (an existing open-source dataset) of the target object onto the natural dataset (i.e., existing images of landscapes, buildings, etc.) to obtain the synthetic dataset. The model training adopts two-stage training. In the first stage, the Backbone is pre-trained on the training dataset. In the second stage, the Backbone is frozen, and the classification head and detection head are trained on the organized training dataset.
[0051] In one embodiment, the image database includes multiple images, and each image includes the target object, that is, the image database is the image database of the target object. Among them, the target object can be objects that require morphological fidelity such as birds, insects, and flowers.
[0052] Among them, the attribute information includes but is not limited to: geographical location information of the shooting location, category information of the target object, time information, weather information, etc. The geographical location information includes: landmarks and coordinates.
[0053] In one embodiment, the attribute information includes the geographical location information of the shooting location and the category information of the target object. Obtain the image with the same geographical location information as the second image and the highest category similarity of the target object in the image database as the first image.
[0054] Whether the geographical location information is the same geographical location information can be judged by the following method:
[0055] Judge whether the landmarks of the two shooting locations are the same landmark. If they are the same landmark, further judge whether the distance between the coordinates of the two shooting locations is within the preset distance. If it is within the preset distance, it means that the geographical location information of the two images is the same geographical location information.
[0056] Taking the target object as a bird as an example, the construction process of the image database is described as follows:
[0057] The first step is to collect data:
[0058] Professional databases: Obtain bird images and their distribution information from professional bird databases and research institutions. For example, the Bird Atlas dataset, which covers bird observation records in multiple countries and regions, including information such as the species, quantity, distribution area, and observation time of birds;
[0059] Citizen science platforms: Integrate data from citizen science platforms, such as the global bird observation data platform (eBird), etc. There are a large number of observation records and images uploaded by bird enthusiasts on these platforms;
[0060] Web crawlers: Use web crawler technology to crawl bird images and corresponding geographical information from relevant natural history websites, ecological research websites, etc.
[0061] Purchase of image data: Purchase image data from professional image rights companies.
[0062] The second step is to organize and annotate the data:
[0063] Perform quality screening on the images collected in the first step, remove blurred, severely occluded, or unclear images to ensure that the images have high clarity and usability. Classify and annotate the collected images according to the species of birds, geographical regions, etc. Among them, obtain GPS information from the core metadata standard (Exchangeable Image File Format, EXIF) information of the digital image file of the jpg image.
[0064] The third step is to construct and store the image database:
[0065] Build an image database of birds, with fields including the category of birds (i.e., bird subtypes), GPS coordinates, location of significant landmarks (set to null if the information is unknown), image path (i.e., the storage location of the image for indexing), tokens after feature extraction of the image, etc., to facilitate subsequent real-time retrieval. Represented by pseudocode as:
[0066] Bird_Database = {
[0067] "Region GPS": "xxx xxx",
[0068] "Significant landmark": "XXXX Lake",
[0069] "Bird subtype": "Mandarin duck class",
[0070] "Image path": "xxx xxx",
[0071] "Feature token path": "xxx xxx"
[0072] }
[0073] In an optionally specific embodiment, the step of screening out the first image that matches the attribute information of the second image from the image database according to the attribute information of the second image includes:
[0074] Screen out a set of candidate images from the image database whose similarity to the attribute information of the second image meets a preset condition according to the attribute information of the second image;
[0075] Extract the first image feature of the second image and the second image features of each candidate image in the set of candidate images;
[0076] Determine the candidate image corresponding to the second image feature with the highest similarity to the first image feature in the set of candidate images as the first image.
[0077] When the user actually takes a picture of the target object, a set of candidate images that meet the preset conditions will be screened out from the image database through the GPS geographical location, typical landmarks, and the category of the target object output by the target detection model, forming a set of candidate images. The specific screening process is as follows:
[0078] If the landmark of the shooting location of the image in the image database belongs to the same landmark as the landmark of the shooting location of the captured image, then it is screened out to obtain a preliminarily screened set of candidate images. If there is no landmark in the captured image, the area formed with the shooting position in the captured image as the center and a radius of XX kilometers is used as the candidate area, and the images in the image database whose shooting locations are in the candidate area are screened out to obtain a preliminarily screened set of candidate images.
[0079] Then, images with the same category information as the target object in the captured image are further screened out from the preliminarily screened set of candidate images, and the images with the same landmark as the captured image and the same category information of the target object are used as candidate images to form a set of candidate images.
[0080] In an embodiment, as Figure 3 shown, a pre-trained deep residual network ResNet50 is used as the extractor of image features. Each candidate image in the set of candidate images will be offline inferred to obtain a 1024-dimensional feature (i.e., the second image feature). The 1024-dimensional feature map can fully express the low-level semantic information and high-level semantic information of the image. By extracting image features in an offline inference manner, it does not occupy terminal resources.
[0081] The second image is passed through the same ResNet50 model to obtain 1024-dimensional features (i.e., the first image features). Calculate the cosine similarity between the first image features and the second image features of each candidate image in the candidate image set, and determine the candidate image with the highest similarity as the first image that matches the second image.
[0082] In an optional specific embodiment, the step 102 inputs the first image and the second image into an image generation model to obtain a first fused feature image, including:
[0083] Input the first image and the second image into the image generation model to obtain the first hidden layer features corresponding to the first image and the second hidden layer features corresponding to the second image;
[0084] Fuse the first hidden layer features and the second hidden layer features to obtain third image features, where the third image features include the texture features in the first hidden layer features and the contour features in the second hidden layer features;
[0085] Decode the third image features to obtain the first fused feature image.
[0086] The second image is encoded into second hidden layer features after passing through the encoder Encoder. The first image also enters the ControlNet model after passing through the Encoder to obtain the corresponding first hidden layer features. Fuse the texture features in the first hidden layer features and the contour features in the second hidden layer features. The obtained third image features not only retain the contour features of the second image but also add the texture features of the second image. The third image features can be decoded through the decoder Decoder to obtain the first fused feature image, which not only retains the contour features of the second image but also enhances the texture of the second image.
[0087] In an optional specific embodiment, the step of fusing the first hidden layer features and the second hidden layer features to obtain third image features includes:
[0088] Input the text description information into the image generation model, and based on the second image and the text description information, obtain fourth image features, where the text description information includes geographical location information and semantic information;
[0089] Perform cross-attention calculation on the first hidden layer features and the second hidden layer features to obtain fused features, and generate the third image features based on the fused features and the fourth image features.
[0090] Input the text description information about geographical location information and semantic information, as well as the second image, into the semantic information extraction module Redux of the image generation model to obtain the fourth image feature including semantic information and geographical location information. Calculate the cross-attention between the texture feature of the first hidden layer feature and the contour feature of the second hidden layer feature to obtain the fusion feature. Then, based on the fusion feature and the fourth image feature, the third image feature related to the text description information and the second image can be generated.
[0091] Among them, the cross-attention calculation between the first hidden layer feature and the second hidden layer feature is specifically as follows:
[0092] The first hidden layer feature and the second hidden layer feature are respectively represented as the key-value pairs Q (query), K (key), and V (value) calculated by the transformer layer. Q lq Performs an attention mechanism operation with Kref and Vref, and Qref performs an attention mechanism operation with Klq and Vlq. Among them, Qlq, Klq, and Vlq are the key-value pairs of the second image, and Qref, Kref, and Vref are the keys of the first image. In addition, the prompt participates in the supervision of the network throughout the process to ensure that the output image is not distorted compared to the input image, that is, the CLIP Loss constrains the consistency between the text description of the output image and the training corresponding prompt.
[0093] In an optional specific embodiment, after obtaining the target image, the target image can be post-processed and optimized. Specifically: Since there may be a certain color error between the target image and the input second image, color correction is required. A color correction method based on wavelet transform can be used. The low-frequency component separated by wavelet transform usually contains the overall color information of the image. By adjusting the low-frequency component, the overall color deviation can be corrected. After extracting and adjusting the wavelet components, they are recombined through inverse wavelet transform to generate the color-corrected image. In addition, contrast enhancement, sharpening adjustment, and artificial intelligence blurring can be performed according to needs to further improve the final imaging effect.
[0094] The training process of the image generation model is described below:
[0095] Obtain the first sample image and the second sample image in the image database. Each image in the image database includes a target object, and the attribute information of the first sample image matches the attribute information of the second sample image;
[0096] Perform blind degradation processing on the second sample image to obtain the third sample image. The attribute information of the second sample image is the same as the attribute information of the third sample image;
[0097] Input the first sample image and the third sample image into the initial image generation model to obtain a predicted image;
[0098] Compare the predicted image and the first sample image through a target loss function to obtain a target loss value, and iteratively train the initial image generation model based on the target loss value to obtain an optimized image generation model.
[0099] Specifically, obtain the first sample image and the attribute information of the first sample image, traverse each sample image and its attribute information in the image database, determine whether the attribute information of each sample image matches the attribute information of the first sample image, and determine the sample image whose attribute information matches the attribute information of the first sample image as the second sample image.
[0100] Use the first sample image as the high-definition reference image, and the second sample image as the ground truth (GT) image of the initial image generation model. Use this GT image to perform blind degradation processing to obtain a third sample image with lower clarity to simulate the picture taken by the terminal's ultra-long focal length. Among them, the attribute information of the third sample image before and after the blind degradation processing remains unchanged. Input the first sample image and the third sample image into the initial image generation model before training to obtain a predicted image.
[0101] Compare the predicted image and the third sample image through a target loss function, calculate the target loss value, and iteratively train the initial image generation model based on the target loss value to obtain an optimized image generation model.
[0102] In one embodiment, the target loss function is represented by the following formula:
[0103] L = L2 Loss + LPIPS Loss + CLIP Loss
[0104] Where L represents the target loss function;
[0105] L2 Loss is also called Mean Squared Error (MSE). L2 Loss is one of the most commonly used regression loss functions, which calculates the mean of the sum of the squares of the differences between the predicted value and the true value;
[0106] The LPIPS Loss is a deep learning loss function used to evaluate the perceptual distance between images. Its core idea is to extract features from different layers of a pre-trained convolutional neural network (usually VGG or AlexNet), and then calculate the Euclidean distance between these features to measure image similarity. The Learned Perceptual Image Patch Similarity Loss (LPIPS Loss) is considered to perform better perceptually than traditional pixel-level losses (such as L2 Loss) because it is closer to human visual perception and can capture image quality differences that cannot be reflected by low-level pixel changes.
[0107] The Contrastive Language-Image Pretraining Loss (CLIP Loss) uses contrastive loss to adjust the embeddings of images and text, making paired images and text more similar and unpaired ones more different.
[0108] In an optionally specific embodiment, the blind degradation process includes:
[0109] Blur, resize, and compress the second sample image to obtain a fourth sample image;
[0110] Degrade the fourth sample image to obtain a degraded image;
[0111] Decode the degraded image to obtain a third sample image;
[0112] Wherein, the degraded image is in the format of an unprocessed original image, and the fourth sample image and the third sample image are in the same processed image format.
[0113] The blind degradation process consists of four stages:
[0114] First-order blind degradation: Blur the second sample image with various isotropic kernels and anisotropic kernels, that is, blur with the same weight in all directions and blur with obvious directionality.
[0115] Second-order blind degradation: Randomly upsample or downsample the second sample image (resize), with the sampling factor set to 0.25 - 4.0, that is, randomly enlarge or reduce the size of the third image, and the sampling method is random, including nearest neighbor interpolation (nearest) and bicubic interpolation (bicubic), etc. This degradation simulates the pixel loss introduced by ultra-long focal digital cropping.
[0116] Third-order blind degradation: Compress the second sample image with a random compression rate to simulate the information loss introduced after image compression.
[0117] Fourth-order blind degradation: Existing terminal image reconstruction algorithms decode raw-format images (raw format), which are multiple frames of unprocessed original images, into RGB images through an artificial intelligence network. Although this process greatly restores the original information of the images, it inevitably loses some details, resulting in smeared decoded images. Therefore, in this application, the fourth sample image in RGB format after the first three steps of blind degradation is degraded into multiple frames of raw images (i.e., degraded images), and then decoded into the third sample image in RGB format through an artificial intelligence network to obtain a slightly smeared image. The low-resolution image obtained after this step can be closer to the effect of real ultra-long focal length shooting.
[0118] In an optional specific embodiment, the step of inputting the first sample image and the third sample image into the initial image generation model to obtain a predicted image includes:
[0119] Input the first sample image and the third sample image into the initial image generation model to obtain the third hidden layer feature corresponding to the first sample image and the fourth hidden layer feature corresponding to the third sample image.
[0120] Input the text description information into the initial image generation model, and based on the third sample image and the text description information, obtain the first sample image feature. The text description information includes geographical location information and semantic information;
[0121] Perform cross-attention calculation on the third hidden layer feature and the fourth hidden layer feature to obtain a sample fusion feature, and based on the sample fusion feature and the first sample image feature, generate a second sample image feature, and perform decoding processing on the second sample image feature to obtain a predicted image.
[0122] Mount LoRA to fine-tune the Stable Diffusion unet base model (i.e., the initial image generation model). The traditional unet base model takes random noise as input, and the text features encoded by the CLIP model are used to guide the model to perform denoising processing, and finally obtain the hidden layer features of a high-definition image. Finally, it is mapped into an image highly related to the text prompt through a decoder Decoder. When training the above text-to-image mapping model, only semantic information or image quality information is considered, resulting in low accuracy of the generated image features. The Redux model is a semantic information extraction module. Similar to the CLIP model, Redux receives semantic text and generates features of a fixed dimension. It directly takes an image as input and outputs semantic features of the same dimension as the CLIP model, and the description features are more stable and closer to the image. Therefore, in the framework of this application, Redux replaces CLIP, which plays a role in ensuring fidelity. In addition, geographical location information is directly introduced into the Redux model, and the features finally output by the model include the semantic information and geographical location information of the image.
[0123] In the embodiment of the present application, LoRA is mounted on the basis of the base model to fine-tune the base model, the geographical location information is spliced with the original semantic prompt, and the mapping between the geographical location information and the high-definition image of the specific region is established.
[0124] As Figure 4 shown, first, after the fine-tuned base model is mounted with LoRA, it has the mapping ability between geographical location information and the first sample image. The base model has three parts of input. The first part is the third sample image, which is encoded into the fourth hidden layer feature after passing through the encoder. The first sample image also enters the ControlNet model through the encoder to obtain the corresponding third hidden layer feature. Inputting the third sample image and geographical location information into Redux can output the first sample image feature including semantic information and geographical location information, where the geographical location information includes GPS geographical location and landmarks. Sending this first sample image feature, the third hidden layer feature, and the fourth hidden layer feature into the base model for calculation, the second sample image feature is obtained. Passing the second sample image feature through the decoder can obtain the predicted image.
[0125] In an optional specific embodiment, the predicted image and the target image can both be added to the image database to form a dynamic image database.
[0126] In the above embodiment of the present application, the low-definition image and the high-definition image of the target object are input into the optimized image generation model. The optimized image generation model can combine the attribute information of the low-definition image and the attribute information of the high-definition image to enhance the texture of the low-definition image, ensuring that the generated image has a significant improvement in clarity, details, and realism. Moreover, through the generation ability of the Stable Diffusion model, combined with the guidance of the high-definition image database and text description information, the clarity, details, and realism of the low-definition image can be effectively enhanced, making the generated image close to or even exceed the effect of single-lens reflex shooting, significantly improving the imaging quality. Also, the construction and update of the image database are based on a large amount of actual data, which can continuously adapt to the types of new target objects and shooting environments, ensuring that the generation effect of the model is continuously improved as the data becomes richer. And, by using the target detection model to capture the target object in the navigation frame in real time, the user does not need to frequently move the mobile terminal for shooting, improving the capture experience. And, without increasing the hardware cost, the high-quality ultra-long focal length target object shooting function is realized through software algorithms.
[0127] The execution subject of the image generation method provided by the embodiment of the present application can be an image generation device. In the embodiment of the present application, taking the image generation device executing the image generation method as an example, the image generation device provided by the embodiment of the present application is described.
[0128] As Figure 5 shown, an embodiment of the present application further provides an image generation device 500, which specifically includes:
[0129] A first acquisition module 501, configured to acquire a first image and a second image, where the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than the clarity of the second image;
[0130] A first processing module 502, configured to input the first image and the second image into an image generation model to obtain a first fused feature image, where the image features of the first fused feature image include the texture features of the first image and the contour features of the second image;
[0131] A second acquisition module 503, configured to obtain a target image corresponding to the second image based on the first fused feature image, where the clarity of the target image is greater than the clarity of the second image.
[0132] Optionally, the first acquisition module 501 is specifically configured to:
[0133] Acquire a captured second image, where the second image includes a target object;
[0134] According to the attribute information of the second image, screen out a first image from the image database that matches the attribute information of the second image, where the attribute information includes: geographical location information of the shooting location and category information of the target object.
[0135] Optionally, when the first acquisition module 501 screens out a first image from the image database that matches the attribute information of the second image according to the attribute information of the second image, it is specifically configured to:
[0136] According to the attribute information of the second image, screen out a candidate image set from the image database whose similarity to the attribute information of the second image meets a preset condition;
[0137] Extract a first image feature of the second image and second image features of each candidate image in the candidate image set;
[0138] Determine the candidate image corresponding to the second image feature with the highest similarity to the first image feature in the candidate image set as the first image.
[0139] Optionally, the first processing module 502 is specifically configured to:
[0140] Input the first image and the second image into the image generation model to obtain the first hidden layer features corresponding to the first image and the second hidden layer features corresponding to the second image;
[0141] Fuse the first hidden layer features and the second hidden layer features to obtain third image features, where the third image features include the texture features in the first hidden layer features and the contour features in the second hidden layer features;
[0142] Decode the third image features to obtain the first fused feature image.
[0143] Optionally, when the first processing module 502 fuses the first hidden layer features and the second hidden layer features to obtain third image features, it is specifically configured to:
[0144] Input the text description information into the image generation model, and based on the second image and the text description information, obtain fourth image features, where the text description information includes geographical location information and semantic information;
[0145] Perform cross-attention calculation on the first hidden layer features and the second hidden layer features to obtain a fused feature, and based on the fused feature and the fourth image features, generate the third image features.
[0146] Optionally, the training process of the image generation model includes:
[0147] A third acquisition module for acquiring a first sample image and a second sample image in an image database, where each image in the image database includes a target object, and the attribute information of the first sample image matches the attribute information of the second sample image;
[0148] A second processing module for performing blind degradation processing on the second sample image to obtain a third sample image, where the attribute information of the second sample image is the same as the attribute information of the third sample image;
[0149] A third processing module for inputting the first sample image and the third sample image into an initial image generation model to obtain a predicted image;
[0150] A training module for comparing the predicted image and the first sample image through a target loss function to obtain a target loss value, and iteratively training the initial image generation model based on the target loss value to obtain an optimized image generation model.
[0151] Optionally, the blind degradation processing includes:
[0152] Blur process, size adjustment process, and compression process are performed on the second sample image to obtain a fourth sample image;
[0153] A degradation process is performed on the fourth sample image to obtain a degraded image;
[0154] The degraded image is decoded to obtain a third sample image;
[0155] Wherein, the degraded image is in the format of an original image that has not been processed, and the fourth sample image and the third sample image are in the same processed image format.
[0156] The image generation device provided in the embodiments of the present application can implement Figures 1 to 4 each process implemented by the method embodiments. To avoid repetition, details are not described herein again.
[0157] The image generation device in the embodiments of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than terminals. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0158] The image generation device in the embodiments of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0159] Optionally, as Figure 6As shown in the figure, an embodiment of the present application further provides an electronic device 600, including a processor 601 and a memory 602. A program or instruction that can run on the processor 601 is stored on the memory 602. When the program or instruction is executed by the processor 601, each step of the above image generation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be elaborated here.
[0160] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.
[0161] Figure 7 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.
[0162] The electronic device 1000 includes, but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.
[0163] Those skilled in the art can understand that the electronic device 1000 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 1010 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 7 The electronic device structure shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0164] In one embodiment, the processor 1010 is configured to obtain a first image and a second image, the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than the clarity of the second image;
[0165] Input the first image and the second image into an image generation model to obtain a first fused feature image, and the image features of the first fused feature image include the texture features of the first image and the contour features of the second image;
[0166] Based on the first fused feature image, obtain a target image corresponding to the second image, and the clarity of the target image is greater than the clarity of the second image.
[0167] Optionally, when the processor 1010 obtains the first image and the second image, it is specifically configured to:
[0168] Obtain the captured second image, and the second image includes a target object;
[0169] Based on the attribute information of the second image, a first image that matches the attribute information of the second image is filtered out from the image database, where the attribute information includes: geographical location information of the shooting location and category information of the target object.
[0170] Optionally, when the processor 1010 filters out a first image that matches the attribute information of the second image from the image database according to the attribute information of the second image, it is specifically configured to:
[0171] According to the attribute information of the second image, a set of candidate images whose similarity to the attribute information of the second image meets a preset condition is filtered out from the image database;
[0172] Extract the first image feature of the second image and the second image features of each candidate image in the set of candidate images;
[0173] The candidate image corresponding to the second image feature with the highest similarity to the first image feature in the set of candidate images is determined as the first image.
[0174] Optionally, when the processor 1010 inputs the first image and the second image into the image generation model to obtain a first fused feature image, it is specifically configured to:
[0175] Input the first image and the second image into the image generation model to obtain a first hidden layer feature corresponding to the first image and a second hidden layer feature corresponding to the second image;
[0176] Perform feature fusion on the first hidden layer feature and the second hidden layer feature to obtain a third image feature, where the third image feature includes the texture feature in the first hidden layer feature and the contour feature in the second hidden layer feature;
[0177] Perform decoding processing on the third image feature to obtain the first fused feature image.
[0178] Optionally, when the processor 1010 performs feature fusion on the first hidden layer feature and the second hidden layer feature to obtain a third image feature, it is specifically configured to:
[0179] Input the text description information into the image generation model, and based on the second image and the text description information, obtain a fourth image feature, where the text description information includes geographical location information and semantic information;
[0180] Perform cross-attention calculation on the first hidden layer features and the second hidden layer features to obtain fused features, and generate the third image feature based on the fused features and the fourth image feature.
[0181] Optionally, during the training process of the image generation model, the processor 1010 is specifically configured to:
[0182] Obtain a first sample image and a second sample image from the image database, where each image in the image database includes a target object, and the attribute information of the first sample image matches the attribute information of the second sample image;
[0183] Perform blind degradation processing on the second sample image to obtain a third sample image, where the attribute information of the second sample image is the same as the attribute information of the third sample image;
[0184] Input the first sample image and the third sample image into the initial image generation model to obtain a predicted image;
[0185] Compare the predicted image and the first sample image through a target loss function to obtain a target loss value, and iteratively train the initial image generation model based on the target loss value to obtain an optimized image generation model.
[0186] Optionally, the blind degradation processing includes:
[0187] Perform blurring processing, size adjustment processing, and compression processing on the second sample image to obtain a fourth sample image;
[0188] Perform degradation processing on the fourth sample image to obtain a degraded image;
[0189] Decode the degraded image to obtain a third sample image;
[0190] Wherein, the degraded image is in the original image format without processing, and the fourth sample image and the third sample image are in the same processed image format.
[0191] It should be understood that in the embodiments of the present application, the input unit 1004 may include a Graphics Processing Unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes the image data of static pictures or videos obtained by an image capturing device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. The other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.
[0192] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area may store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchlink dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.
[0193] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1010 either.
[0194] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0195] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs, etc.
[0196] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0197] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0198] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0199] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0201] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
Claims
1. An image generation method, characterized in that, Including: Obtain a first image and a second image, where the attribute information of the first image matches the attribute information of the second image, and the clarity of the first image is greater than that of the second image; Input the first image and the second image into an image generation model to obtain a first fused feature image, where the image features of the first fused feature image include the texture features of the first image and the contour features of the second image; Based on the first fused feature image, obtain a target image corresponding to the second image, where the clarity of the target image is greater than that of the second image.
2. The method according to claim 1, wherein The obtaining of the first image and the second image includes: Obtain the captured second image, where the second image includes a target object; According to the attribute information of the second image, screen out a first image from the image database that matches the attribute information of the second image, where the attribute information includes geographical location information of the shooting location and category information of the target object.
3. The method according to claim 2, wherein The screening out of the first image from the image database that matches the attribute information of the second image according to the attribute information of the second image includes: According to the attribute information of the second image, screen out a set of candidate images from the image database whose similarity to the attribute information of the second image meets a preset condition; Extract the first image features of the second image and the second image features of each candidate image in the set of candidate images; Determine the candidate image corresponding to the second image feature with the highest similarity to the first image feature in the set of candidate images as the first image.
4. The method according to claim 1, wherein The inputting of the first image and the second image into the image generation model to obtain a first fused feature image includes: Input the first image and the second image into the image generation model to obtain a first hidden layer feature corresponding to the first image and a second hidden layer feature corresponding to the second image; Perform feature fusion on the first hidden layer feature and the second hidden layer feature to obtain a third image feature, where the third image feature includes the texture features in the first hidden layer feature and the contour features in the second hidden layer feature; Perform decoding processing on the third image feature to obtain the first fused feature image.
5. The method according to claim 4, wherein The performing of feature fusion on the first hidden layer feature and the second hidden layer feature to obtain a third image feature includes: Input text description information into the image generation model, and based on the second image and the text description information, obtain a fourth image feature, where the text description information includes geographical location information and semantic information; Perform cross-attention calculation on the first hidden layer feature and the second hidden layer feature to obtain a fused feature, and based on the fused feature and the fourth image feature, generate the third image feature.
6. The method according to claim 1, wherein The training process of the image generation model includes: Obtain a first sample image and a second sample image from the image database, where each image in the image database includes a target object, and the attribute information of the first sample image matches the attribute information of the second sample image; Perform blind degradation processing on the second sample image to obtain a third sample image, where the attribute information of the second sample image is the same as that of the third sample image; Input the first sample image and the third sample image into the initial image generation model to obtain a predicted image; Compare the predicted image and the first sample image through the target loss function to obtain a target loss value, and iteratively train the initial image generation model based on the target loss value to obtain an optimized image generation model.
7. The method according to claim 6, characterized in that The blind degradation processing includes: Perform blurring processing, size adjustment processing, and compression processing on the second sample image to obtain a fourth sample image; Perform degradation processing on the fourth sample image to obtain a degraded image; Decode the degraded image to obtain a third sample image; Wherein, the degraded image is in the original image format without processing, and the fourth sample image and the third sample image are in the same processed image format.
8. An image generation device, characterized in that, It includes: A first acquisition module for acquiring a first image and a second image, where the attribute information of the first image matches that of the second image, and the clarity of the first image is greater than that of the second image; A first processing module for inputting the first image and the second image into an image generation model to obtain a first fused feature image, where the image features of the first fused feature image include the texture features of the first image and the contour features of the second image; A second acquisition module for obtaining a target image corresponding to the second image based on the first fused feature image, where the clarity of the target image is greater than that of the second image.
9. An electronic device, characterized in that, It includes a processor and a memory, where the memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.
10. A readable storage medium, characterized in that, The program or instructions are stored on the readable storage medium, and when the program or instructions are executed by the processor, the steps of the method according to any one of claims 1-7 are implemented.