Image processing method and apparatus, electronic device, and readable storage medium

By acquiring short-focal-length and long-focal-length images and using edge and semantic information to augment the long-focal-length images, the problem of low image generation accuracy in existing technologies is solved, enabling the generation of images with a wider field of view and improving image quality and consistency.

WO2026092388A1PCT designated stage Publication Date: 2026-05-07VIVO MOBILE COMM CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2025-10-27
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In the existing technology, when electronic devices generate extended images from telephoto images, the accuracy is low, and the physical and semantic consistency between the extended image and the original image cannot be strictly guaranteed.

Method used

By acquiring short-focal-length and long-focal-length images, the long-focal-length image is expanded using image edge information and semantic information to generate a target image with a wider field of view. Image processing is performed using a control condition model and a diffusion model.

Benefits of technology

It improves the accuracy of generated images, ensures semantic consistency between the extended image and the original image, and enhances image quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025130249_07052026_PF_FP_ABST
    Figure CN2025130249_07052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of images, and discloses an image processing method and apparatus, an electronic device, and a readable storage medium. The image processing method comprises: on the basis of a first image and a second image, generating a third image; on the basis of image edge information of the third image and image semantic information of the third image, expanding image content of the second image to generate a target image. The first image is an image acquired by a first camera, and the second image is an image acquired by a second camera. The first image and the second image are images of a same scene. The first camera has a focal length less than that of the second camera. The third image is an image generated by replacing first image content in the first image with second image content, and the second image content is image content in the second image that is the same as the first image content.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, apparatus, electronic devices and readable storage media

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411552597.0, filed in China on November 1, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application belongs to the field of image technology, specifically relating to an image processing method, apparatus, electronic device, and readable storage medium. Background Technology

[0004] With the development of electronic device technology, users frequently use the camera applications of electronic devices to take pictures. Currently, users can usually only take and acquire images at a specific focal length at a time, such as telephoto or short focal length images.

[0005] In related technologies, in order to break through the limitations of traditional electronic device photography and achieve high-quality full-focal-length portrait effects in a single shot, electronic devices typically use telephoto lenses to capture clearer images due to the limited number of cameras. Then, based on a stable diffusion model, an extended image corresponding to the telephoto image is generated to obtain an image with a wider field of view.

[0006] Because the expanded images generated are based solely on images taken at long focal lengths, they are random and therefore have low accuracy. Summary of the Invention

[0007] The purpose of this application is to provide an image processing method, apparatus, electronic device, and readable storage medium that can ensure that the extended image corresponding to the image captured by the telephoto lens maintains strict physical and semantic consistency with the image captured by the telephoto lens, thereby improving the accuracy of the generated extended image.

[0008] In a first aspect, embodiments of this application provide an image processing method, which includes: generating a third image based on a first image and a second image; and expanding the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image; wherein the first image is an image acquired by a first camera, the second image is an image acquired by a second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the first image content in the first image with the second image content, wherein the second image content is the same as the first image content in the second image.

[0009] Secondly, embodiments of this application provide an image processing apparatus, which includes a processing module and a generation module. The processing module is used to generate a third image based on a first image and a second image. The generation module is used to expand the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, the second image is an image acquired by a second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the first image content in the first image with the second image content, wherein the second image content is the same as the first image content in the second image.

[0010] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0011] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0012] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0013] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0014] In this embodiment, the electronic device expands the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, and the second image is an image acquired by a second camera. Both the first and second images depict the same scene. The focal length of the first camera is shorter than that of the second camera. The third image is generated by replacing the content of the first image in the first image with the content of the second image. The content of the second image is the same as the content of the first image. Because the focal length of the first camera is shorter than that of the second camera, the first image acquired by the first camera is typically a short-focal-length image with a wider field of view but insufficient clarity, while the second image acquired by the second camera is typically a long-focal-length image with higher clarity but a limited field of view. In this solution, the electronic device generates a third image with a wider field of view by acquiring both short-focal-length and long-focal-length images. It then applies conditional constraints to the expanded content of the second image based on the image edge information and image semantic information of the third image, thereby ensuring high accuracy of the generated target image and guaranteeing semantic consistency between the second and target images. Attached Figure Description

[0015] Figure 1 is a schematic diagram of an image processing method provided in an embodiment of this application;

[0016] Figure 2 is a schematic diagram of the structure of an adapter model provided in an embodiment of this application;

[0017] Figure 3 is a second schematic diagram of an image processing method provided in an embodiment of this application;

[0018] Figure 4 is a schematic diagram of a diffusion model structure provided in an embodiment of this application;

[0019] Figure 5 is a flowchart illustrating a specific example of an image processing method provided in an embodiment of this application.

[0020] Figure 6 is a schematic diagram of the structure of an image processing device provided in an embodiment of this application;

[0021] Figure 7 is one of the hardware structure diagrams of an electronic device provided in an embodiment of this application;

[0022] Figure 8 is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The following explains some terms and nouns used in the embodiments of this application.

[0025] Stable Diffusion is a deep learning-based text-to-image diffusion model that combines diffusion and autoencoders. The model uses an autoencoder to transform the image into a representation in the latent space, a text encoder to process the text, a cross-attention mechanism to learn the correlation between the text and image content, and a convolutional neural network (U-Net) to generate the predicted image. It is trained on a massive dataset containing images of various sizes and shapes, the diversity of which allows Stable Diffusion to generate a wide range of realistic images. Furthermore, it enables the generation of new pixels that match the style and content of the original image.

[0026] ControlNet is a neural network primarily designed to control a pre-trained Stable Diffusion model. It locks the parameters of the Stable Diffusion model and clones them into a trainable copy of ControlNet. Control conditions, such as edge maps and segmentation maps, are input to control the final output of the Stable Diffusion model. Training the ControlNet-Outpainting model on a large number of masked images allows it to infer and generate masked regions, improving the inpainting capabilities of Stable Diffusion. This results in the generation of new pixels on the image, achieving seamless image expansion.

[0027] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0028] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0029] The identifiers in this application are text, symbols, images, etc. used to indicate information, and may be used as carriers for displaying information in the form of identifiers or other containers, including but not limited to text identifiers, image identifiers, symbol identifiers, etc.

[0030] It should be noted that the image processing method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the image processing method provided in this application.

[0031] Currently, most mainstream electronic devices on the market are equipped with multi-focal length cameras to meet different user shooting needs. Telephoto cameras have a narrower field of view and a smaller subject area, making them suitable for shooting distant objects and capturing clearer details. Short-focal length cameras, on the other hand, offer a wider field of view and subject area, presenting a broader image while emphasizing perspective and enhancing the image's impact, thus better meeting the diverse shooting requirements of users.

[0032] In related technologies, in order to break through the limitations of traditional electronic device photography and achieve high-quality full-focal-length portrait effects in a single shot, electronic devices typically use telephoto lenses to capture clearer images due to the limited number of cameras. Then, they employ a diffusion-based image expansion algorithm to generate an expanded image corresponding to the telephoto lens image, thereby obtaining an image with a wider field of view.

[0033] However, on the one hand, given the limited number of cameras on electronic devices, current devices need to install multiple high-precision cameras with different focal lengths to directly implement the technique of cutting from long focal lengths to short focal lengths, placing extremely high demands on the hardware. On the other hand, although electronic devices can directly extend images based on the generative algorithms in Stable Diffusion, these algorithms only infer short focal length images from existing long focal length images, predicting images with a wider field of view. Due to the randomness of generative algorithms, it is impossible to strictly guarantee the physical and semantic alignment of the expanded region with the original image, resulting in insufficient accuracy of the extended image.

[0034] In the image processing method provided in this application embodiment, the image content of the second image is expanded based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, and the second image is an image acquired by a second camera. The first and second images are images of the same scene, and the focal length of the first camera is shorter than that of the second camera. The third image is generated by replacing the content of the first image in the first image with the content of the second image, where the content of the second image is the same as the content of the first image. Because the focal length of the first camera is shorter than that of the second camera, the first image acquired by the first camera is typically a short-focal-length image with a wider field of view but insufficient clarity, while the second image acquired by the second camera is typically a long-focal-length image with higher clarity but a limited field of view. In this solution, the electronic device generates a third image with a wider field of view by acquiring both short-focal-length and long-focal-length images, and applies conditional constraints to the expanded content of the second image based on the image edge information and image semantic information of the third image. This results in a high accuracy of the generated target image and ensures semantic consistency between the second image and the target image.

[0035] The image processing method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0036] The image processing method provided in this application can be executed by an image processing device. Exemplarily, the image processing device can be an electronic device, or a component within that electronic device, such as an integrated circuit or a chip. The image processing method provided in this application will be described executively below using an electronic device as an example.

[0037] This application provides an image processing method. Figure 1 shows a flowchart of an image processing method provided by this application, which can be applied to an electronic device. As shown in Figure 1, the image processing method provided by this application may include the following steps 201 and 202.

[0038] Step 201: The electronic device generates a third image based on the first and second images.

[0039] In some embodiments of this application, the first image described above is an image acquired by a first camera.

[0040] In some embodiments of this application, the second image described above is an image acquired by a second camera.

[0041] In some embodiments of this application, the first image and the second image described above are images of the same scene captured by different cameras.

[0042] In some embodiments of this application, the focal length of the first camera is smaller than that of the second camera.

[0043] In some embodiments of this application, the electronic device is an electronic device with a dual-camera system, that is, the electronic device has at least two cameras.

[0044] For example, if the first camera is a short focal length camera, then the first image is a short focal length image.

[0045] For example, if the second camera is a telephoto camera, then the second image is a telephoto image.

[0046] For example, an electronic device acquires a relatively blurry short-focal-length image and a high-resolution long-focal-length image respectively using the first and second cameras in a dual-camera system. The long-focal-length image can be a 4x image, with a clear picture, but a relatively small overall field of view. The short-focal-length image, while generally blurry, has a wider field of view and includes an extended image corresponding to the short-focal-length image generated from the long-focal-length image—that is, the extended image content required to generate the target image described below—thus providing the possibility for generating the target image.

[0047] In some embodiments of this application, the third image is an image generated by replacing the content of the first image in the first image with the content of the second image.

[0048] In some embodiments of this application, the content of the second image is the same as the content of the first image in the second image.

[0049] In some embodiments of this application, the electronic device acquires key points in a first image and a second image, then matches and aligns the key points of the first image and the second image one by one to obtain the content of the first image and the content of the second image, and then generates a third image by replacing the content of the first image with the content of the second image.

[0050] Step 202: The electronic device expands the image content of the second image based on the image edge information and image semantic information of the third image to generate the target image.

[0051] In some embodiments of this application, the above-mentioned image edge information is used to characterize the edge information of each image region in the third image.

[0052] In some embodiments of this application, the above-mentioned image semantic information is used to characterize image information in a third image.

[0053] In some embodiments of this application, the electronic device inputs the image edge information and image semantic information of the third image into a control condition model to extract constraint condition information, and then inputs the second image and constraint condition information into a diffusion model. The diffusion model expands the image content of the second image according to the constraint condition information to generate a target image.

[0054] In the image processing method provided in this application embodiment, the electronic device expands the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, the second image is an image acquired by a second camera, the first and second images are images of the same scene, the focal length of the first camera is shorter than that of the second camera, and the third image is an image generated by replacing the content of the first image in the first image with the content of the second image. The content of the second image is the same as the content of the first image in the second image. Because the focal length of the first camera is shorter than that of the second camera, the first image acquired by the first camera is usually a short-focal-length image with a wider field of view but insufficient clarity, while the second image acquired by the second camera is usually a long-focal-length image with higher clarity but a limited field of view. In this solution, the electronic device generates a third image with a wider field of view by acquiring both short-focal-length and long-focal-length images, and imposes conditional constraints on the expanded content of the second image based on the image edge information and image semantic information of the third image, thereby ensuring high accuracy of the generated target image and guaranteeing semantic consistency between the second image and the target image.

[0055] Optionally, in some embodiments of this application, step 201 above can be specifically implemented by steps 201a to 201c below.

[0056] Step 201a: The electronic device acquires the first key point of the first image and the second key point of the second image.

[0057] In some embodiments of this application, the electronic device employs a key point detection algorithm to obtain a first key point in a first image and a second key point in a second image.

[0058] For example, the aforementioned key points can be key points corresponding to the main person in a portrait image, or key points corresponding to the main scenery in a landscape image.

[0059] Step 201b: The electronic device matches the first key point and the second key point to determine the content of the first image and the content of the second image.

[0060] In some embodiments of this application, the electronic device matches the same first key point and second key point.

[0061] In other words, the electronic device performs an overlay match between the first image and the second image, and identifies the parts of the first image and the second image that have the same content as the content of the first image and the second image, respectively.

[0062] For example, suppose the first image and the second image are half-body images of Xiaoming. The key points of the first image are the key points of Xiaoming's face, and the key points of the second image are also the key points of Xiaoming's face. The electronic device matches the key points of the left eye, the key points of the right eye, and the key points of the mouth in the two images one by one to determine the content of the first image and the content of the second image.

[0063] Step 201c: The electronic device replaces the content of the first image with the content of the second image to generate a third image.

[0064] For example, the electronic device performs image matching based on the corresponding key points of the first image and the second image, and calculates the transformation matrix to determine the position of the first image in the coordinate system of the second image, such as M(match_x, match_y, match_width, match_height). Then, the electronic device replaces the image content in the position region M of the first image with the image content in the same coordinate position region of the second image to obtain the aforementioned third image.

[0065] It is understandable that the third image obtained by the electronic device based on the first image to complete the second image has roughly the same structural information as the first and second images. Therefore, in the aligned third image, other fields of view of the second image can be completed based on the first image.

[0066] In this way, the electronic device can obtain the magnification required to expand the second image based on the third image, and calculate the size of the final generated target image.

[0067] Optionally, in some embodiments of this application, before the above-mentioned step 202 "the electronic device expands the image content of the second image based on the image edge information and image semantic information of the third image to generate the target image", the image processing method provided in some embodiments of this application further includes the following steps 301 and 302.

[0068] Step 301: The electronic device uses an edge detection operator to detect the third image and obtain the image edge information.

[0069] For example, the edge detection operator described above can be the Canny operator.

[0070] It is understandable that when electronic devices use AI generative algorithms to predict the expanded region of the second image, in order to ensure that the second image after the expansion and filling has the same physical structure information and semantic information as the first image, this embodiment uses the Canny edge detection operator to extract the edge information of the third image.

[0071] Step 302: The electronic device uses an adapter model to extract the semantic information of the third image.

[0072] In some embodiments of this application, the above adapter model includes the IP-Adapter algorithm.

[0073] For example, the IP-Adapter algorithm is an efficient and lightweight adapter for implementing image cues in a pre-trained Stable Diffusion model. It leverages high-level semantic information in images to control image generation within the Stable Diffusion model. Its decoupled cross-attention mechanism processes textual and image features separately in the cross-attention layer, achieving multimodal image generation. This allows image semantic information and textual cues to work well together, giving the generated image the semantic information of a third-party image.

[0074] It should be noted that the execution order of the above steps 301 and 302 can be either step 301 first and then step 302, or step 302 first and then step 301, or they can be executed simultaneously. This application embodiment does not impose any restrictions.

[0075] In this way, the electronic device acquires image edge information and image semantic information to perform image expansion constraints on the image generation process, thereby ensuring the consistency of the physical structure information of the generated target image with the first and second images, and thus improving the accuracy of target image generation.

[0076] Optionally, in some embodiments of this application, step 302 above can be specifically implemented by steps 302a to 302d.

[0077] Step 302a: The electronic device inputs the third image and text prompt information into the adapter model.

[0078] In some embodiments of this application, the above-mentioned text prompt information is used to instruct the adapter model to extract the image semantic information of the third image.

[0079] Step 302b: The electronic device extracts the first image feature information of the third image and the first text feature information of the text prompt information through the cross-attention layer in the adapter model.

[0080] In some embodiments of this application, the adapter model described above is used to extract the image semantic information of an image.

[0081] In some embodiments of this application, the adapter model described above includes at least two cross-attention layers.

[0082] For example, the electronic device extracts the first image feature information of the third image through the first cross-attention layer, and extracts the first text feature information of the text prompt information through the second cross-attention layer.

[0083] For example, after the electronic device inputs the above-mentioned text prompt information and the third image into the adapter model, the first image feature information of the third image and the first text feature information of the text prompt information are extracted through different cross-attention layers in the adapter model.

[0084] Step 302c: The electronic device generates image semantic information based on the first image feature information and the first text feature information through the adapter model.

[0085] Step 302d: The electronic device outputs image semantic information through the adapter model.

[0086] In some embodiments of this application, the electronic device processes the first image feature information and the second text feature information through the hierarchy in the adapter model to generate the image semantic information of the third image, and finally the electronic device outputs the image semantic information through the adapter model.

[0087] For example, as shown in Figure 2, the adapter model comprises an image encoder, a text encoder, layer normalization, a fully connected layer, and a decoupled cross-attention layer. The third image and text prompt information are input into the adapter model. Image feature information is extracted through the image encoder, layer normalization, and fully connected layer, and the first image feature information is further extracted through the decoupled cross-attention layer. Simultaneously, text feature information is extracted through the text encoder, and the first text feature information is further extracted through the decoupled cross-attention layer. The decoupled cross-attention layer performs denoising processing on the first image feature information and the first text feature information at different levels to obtain the final image semantic information.

[0088] In this way, the electronic device separates the cross-attention layers of text features and image features through the decoupled cross-attention mechanism in the adapter model, thus completing multimodal image generation. This allows image semantic information and text prompts to be used in good coordination, thereby extracting the image semantic information of the third image.

[0089] Optionally, in some embodiments of this application, in conjunction with FIG1 and FIG3, the above-mentioned step 202 "the electronic device performs image expansion processing on the second image based on the image edge information and image semantic information of the third image to generate the target image" can be specifically implemented through the following steps 202a to 202d.

[0090] Step 202a: The electronic device inputs the image edge information and image semantic information into the control condition model for processing and outputs the first image condition constraint information.

[0091] In some embodiments of this application, the control condition model described above can be a ControlNet model, which extracts feature information of image edge information and image semantic information through coding blocks as the first image condition constraint information.

[0092] In some embodiments of this application, the image edge information extracted above and the image semantic information extracted by the IP-Adapter are used as input to the ControlNet model, and the second image is extended based on a Stable Diffusion model with strong generation capabilities, such as SDXL, to generate the target image, and the target image maintains physical and semantic consistency with the first and second images.

[0093] Step 202b: The electronic device inputs the first image condition constraint information and the second image into the diffusion model.

[0094] In some embodiments of this application, the aforementioned first image condition constraint information is used to constrain the diffusion model to expand the second image, so that the target image generated by the expanded image maintains semantic and physical consistency with the first and second images.

[0095] Step 202c: The electronic device performs image expansion processing on the second image based on the first image condition constraint information using a diffusion model to generate the target image.

[0096] Step 202d: The electronic device outputs the target image through the diffusion model.

[0097] For example, as shown in Figure 4, this embodiment uses Stable Diffusion(a) as an example to demonstrate how ControlNet(b) adds conditional control to a large pre-trained diffusion model. Stable Diffusion is essentially a neural network model with an encoder, intermediate blocks, and a decoder with skip connections. Both the encoder and decoder contain 12 blocks, and the complete model contains 25 blocks, including the intermediate blocks. Of the 25 blocks, 8 are downsampling or upsampling convolutional layers, while the other 17 are main blocks, each containing 4 ResNet layers and 2 Vision Transformers (ViTs). Each ViT contains multiple cross-attention and self-attention mechanisms.

[0098] For example, referring to Figure 3, "SD Encoder Block A" contains 4 ResNet layers and 2 ViT layers, and "×3" indicates that the block is repeated 3 times. Text prompts are encoded using a CLIP text encoder, and the diffusion time step is encoded using a position-encoded time encoder. The ControlNet architecture is applied to each encoder level of the U-Net.

[0099] Specifically, in this embodiment, ControlNet is used to create trainable copies of 12 encoding blocks and 1 Stable Diffusion intermediate block. The 12 encoding blocks have four resolutions (64×64, 32×32, 16×16, 8×8), each copied three times. The output is added to the 12 skip connections and 1 intermediate block of the U-Net. Since stable diffusion is a typical U-Net architecture, this ControlNet architecture may be applicable to other models. Our method of connecting ControlNet is computationally efficient—because the parameters of the locked copies are frozen, gradient calculations are not required in the initially locked encoder for fine-tuning. This approach speeds up training and saves GPU memory. According to tests on a single NVIDIA A100 PCIe 40GB, optimizing stable diffusion using ControlNet requires only about 23% of GPU memory and 34% of system memory, compared to the time required per training iteration when optimizing stable diffusion without ControlNet. The image diffusion model learns to progressively denoise images and generate samples from the training domain. The denoising process can occur in the pixel space or in the latent space encoded from the training data. Stable Diffusion uses the latent image as the training domain because working in this space has been shown to stabilize the training process. Specifically, Stable Diffusion uses a preprocessing method similar to VQ-GAN to transform the 512×512 pixel spatial image into a smaller 64×64 latent image. To add ControlNet to Stable Diffusion, we first transform each input conditioning image (e.g., edges, pose, depth, etc.) from the 512×512 input size into a 64×64 feature space vector that matches the size of Stable Diffusion, i.e., the first image constraint information mentioned above. This feature space vector is then input into Stable Diffusion so that Stable Diffusion can generate the target image based on the feature space vector.

[0100] In this way, the electronic device introduces the image edge information and image semantic information of the third image to constrain the image expansion process of generating the target image. This allows the expanded image to better maintain the consistency of the original image's edge structure and semantics, thereby giving the target image better structure and realism while also better maintaining the overall structural information of the original image. This avoids the generation of people or clutter that are unrelated to the original image's environment, thus improving the accuracy of the generated target image.

[0101] The following specific examples illustrate the image processing method provided in the embodiments of this application.

[0102] Example 1: As shown in Figure 5, the image processing method may include the following steps A1 to A3.

[0103] Step A1: First, the dual-camera system of the electronic device acquires a high-definition telephoto main camera image, i.e., the second image mentioned above, and a relatively blurry short-focal-length secondary camera image, i.e., the first image mentioned above. Then, the main camera image and the secondary camera image are matched and aligned to generate a main-secondary camera aligned image, i.e., the third image mentioned above.

[0104] Step A2: The electronic device extracts the image edge information and image semantic information of the aligned images of the main and secondary cameras using the edge acquisition method and the adapter model, respectively.

[0105] Step A3: The electronic device expands the image content of the main camera image by inputting the image edge information, image semantic information and the main camera image of the aligned main and secondary cameras into the diffusion model to generate short and long focal length images, i.e. the target image mentioned above.

[0106] Thus, through the embodiments of this application, focal length switching can be performed on already captured photos without the need for additional hardware support, ensuring the physical and semantic consistency between the extended image predicted by the Stable Diffusion generative algorithm and the original image, and achieving the effect of cropping a long-focus scene to a short-focus scene. This improves upon the limitations of traditional mobile phone photography in post-processing of already captured scenes, while enhancing the quality of the generated images and the satisfaction of the final product, and reducing post-processing costs for users.

[0107] Each of the above-described method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application does not impose any restrictions on this.

[0108] The window blurring method provided in this application can be executed by an electronic device or a window blurring device. This application uses a window blurring device executing the window blurring method as an example to illustrate the window blurring device provided in this application.

[0109] Figure 6 shows a possible structural schematic diagram of the image processing apparatus involved in an embodiment of this application. As shown in Figure 6, the image processing apparatus 700 may include a processing module 701 and a generation module 702.

[0110] The processing module 701 is used to generate a third image based on the first image and the second image; the generation module 702 is used to expand the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image; wherein the first image is an image acquired by the first camera, the second image is an image acquired by the second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the content of the first image in the first image with the content of the second image, wherein the content of the second image is the same as the content of the first image in the second image.

[0111] Optionally, in some embodiments of this application, the above-mentioned processing module 701 is specifically used for:

[0112] Obtain the first key point of the first image, and obtain the second key point of the second image;

[0113] Match the first key point and the second key point to determine the content of the first image and the content of the second image.

[0114] Replace the content of the first image with the content of the second image to generate the third image.

[0115] Optionally, in some embodiments of this application, the processing module 701 is further configured to expand the image content of the second image based on the image edge information and image semantic information of the third image, and before generating the target image, use an edge detection operator to detect the third image to obtain image edge information; the processing module 701 is further configured to extract the image semantic information of the third image using an adapter model.

[0116] Optionally, in some embodiments of this application, the above-mentioned processing module 701 is specifically used for:

[0117] The third image and text prompts are input into the adapter model. The text prompts are used to instruct the adapter model to extract the semantic information of the third image.

[0118] The first image feature information of the third image and the first text feature information of the text prompt information are extracted through the cross attention layer in the adapter model.

[0119] Based on the first image feature information and the first text feature information, the adapter model generates image semantic information.

[0120] The adapter model outputs semantic information of the image.

[0121] Optionally, in some embodiments of this application, the above-mentioned generation module 702 is specifically used for:

[0122] Image edge information and image semantic information are input into the control condition model for processing, and the first image condition constraint information is output.

[0123] The first image constraint information and the second image are input into the diffusion model;

[0124] The second image is expanded using a diffusion model based on the conditional constraints of the first image to generate the target image.

[0125] The target image is output using a diffusion model.

[0126] In the image processing apparatus provided in this application embodiment, the image processing apparatus expands the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, the second image is an image acquired by a second camera, the first image and the second image are images of the same scene, the focal length of the first camera is shorter than that of the second camera, and the third image is an image generated by replacing the content of the first image in the first image with the content of the second image. The content of the second image is the same as the content of the first image in the second image. Because the focal length of the first camera is shorter than that of the second camera, the first image acquired by the first camera is usually a short-focal-length image with a wider field of view but insufficient clarity, while the second image acquired by the second camera is usually a long-focal-length image with higher clarity but a limited field of view. In this solution, the image processing apparatus generates a third image with a wider field of view by acquiring both short-focal-length and long-focal-length images, and imposes conditional constraints on the expanded content of the second image based on the image edge information and image semantic information of the third image, thereby ensuring high accuracy of the generated target image and guaranteeing semantic consistency between the second image and the target image.

[0127] The image processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0128] The image processing device in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0129] The image processing apparatus provided in this application embodiment can implement the various processes implemented in the image processing method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0130] Optionally, as shown in FIG7, this application embodiment also provides an electronic device 800, including a processor 801 and a memory 802. The memory 802 stores a program or instructions that can run on the processor 801. When the program or instructions are executed by the processor 801, they implement the various steps of the above-described image processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0131] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0132] Figure 8 is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.

[0133] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0134] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The electronic device structure shown in Figure 8 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0135] The processor 110 is used to generate a third image based on the first image and the second image; the processor 110 is used to expand the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image; wherein the first image is an image acquired by the first camera, the second image is an image acquired by the second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the first image content in the first image with the second image content, wherein the second image content is the same as the first image content in the second image.

[0136] Optionally, in some embodiments of this application, the processor 110 is specifically used for:

[0137] Obtain the first key point of the first image, and obtain the second key point of the second image;

[0138] Match the first key point and the second key point to determine the content of the first image and the content of the second image.

[0139] Replace the content of the first image with the content of the second image to generate the third image.

[0140] Optionally, in some embodiments of this application, the processor 110 is further configured to expand the image content of the second image based on the image edge information and image semantic information of the third image, and before generating the target image, use an edge detection operator to detect the third image to obtain image edge information; the processor 110 is further configured to extract the image semantic information of the third image using an adapter model.

[0141] Optionally, in some embodiments of this application, the processor 110 is specifically used for:

[0142] The third image and text prompts are input into the adapter model. The text prompts are used to instruct the adapter model to extract the semantic information of the third image.

[0143] The first image feature information of the third image and the first text feature information of the text prompt information are extracted through the cross attention layer in the adapter model.

[0144] Based on the first image feature information and the first text feature information, the adapter model generates image semantic information.

[0145] The adapter model outputs semantic information of the image.

[0146] Optionally, in some embodiments of this application, the processor 110 is specifically used for:

[0147] Image edge information and image semantic information are input into the control condition model for processing, and the first image condition constraint information is output.

[0148] The first image constraint information and the second image are input into the diffusion model;

[0149] The second image is expanded using a diffusion model based on the conditional constraints of the first image to generate the target image.

[0150] The target image is output using a diffusion model.

[0151] In the electronic device provided in this application embodiment, the electronic device expands the image content of the second image based on the image edge information and image semantic information of the third image to generate a target image. The first image is an image acquired by a first camera, and the second image is an image acquired by a second camera. The first and second images are images of the same scene, and the focal length of the first camera is shorter than that of the second camera. The third image is an image generated by replacing the content of the first image in the first image with the content of the second image, where the content of the second image is the same as the content of the first image. Because the focal length of the first camera is shorter than that of the second camera, the first image acquired by the first camera is usually a short-focal-length image with a wider field of view but insufficient clarity, while the second image acquired by the second camera is usually a long-focal-length image with higher clarity but a limited field of view. In this solution, the electronic device generates a third image with a wider field of view by acquiring both short-focal-length and long-focal-length images, and imposes conditional constraints on the expanded content of the second image based on the image edge information and image semantic information of the third image. This results in a high accuracy of the generated target image and ensures semantic consistency between the second image and the target image.

[0152] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0153] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0154] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0155] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0156] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0157] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0158] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0159] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0160] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0162] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image processing method, the method comprising: Generate a third image based on the first and second images; Based on the image edge information and image semantic information of the third image, the image content of the second image is expanded to generate the target image; Wherein, the first image is an image acquired by the first camera, the second image is an image acquired by the second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the first image content in the first image with the second image content, wherein the second image content is the same as the first image content in the second image.

2. The method according to claim 1, wherein, The process of generating a third image based on the first and second images includes: Obtain the first key point of the first image, and obtain the second key point of the second image; Match the first key point and the second key point to determine the content of the first image and the content of the second image. The content of the first image is replaced with the content of the second image to generate a third image.

3. The method according to claim 1, wherein, Before expanding the image content of the second image based on the image edge information and image semantic information of the third image to generate the target image, the method further includes: An edge detection operator is used to detect the edge information of the third image. The image semantic information of the third image is obtained using an adapter model.

4. The method according to claim 3, wherein, The step of using an adapter model to obtain the image semantic information of the third image includes: The adapter model generates the image semantic information based on the first image feature information and the first text feature information; The image semantic information is output through the adapter model.

5. The method according to claim 3, wherein, Based on the image edge information and image semantic information of the third image, the image content of the second image is expanded to generate the target image, including: The image edge information and the image semantic information are input into the control condition model for processing, and the first image condition constraint information is output. The first image condition constraint information and the second image are input into the diffusion model; The target image is generated by expanding the image content of the second image based on the first image conditional constraint information using the diffusion model. The target image is output through the diffusion model.

6. An image processing apparatus, the image processing apparatus comprising: Processing module and generation module; The processing module is used to generate a third image based on the first image and the second image; The generation module is used to expand the image content of the second image based on the image edge information and image semantic information of the third image processed by the processing module, and generate a target image. Wherein, the first image is an image acquired by the first camera, the second image is an image acquired by the second camera, the first image and the second image are images of the same scene, the focal length of the first camera is smaller than that of the second camera, and the third image is an image generated by replacing the first image content in the first image with the second image content, wherein the second image content is the same as the first image content in the second image.

7. The apparatus according to claim 6, wherein, The processing module is specifically used for: Obtain the first key point of the first image, and obtain the second key point of the second image; Match the first key point and the second key point to determine the content of the first image and the content of the second image. The first image content is replaced with the second image content to generate the third image.

8. The apparatus according to claim 6, wherein, The processing module is further configured to expand the image content of the second image based on the image edge information and image semantic information of the third image, and before generating the target image, use an edge detection operator to detect the third image to obtain the image edge information of the third image; The processing module is also used to extract the image semantic information of the third image using an adapter model.

9. The apparatus according to claim 8, wherein, The processing module is specifically used for: The third image and text prompt information are input into the adapter model, and the text prompt information is used to instruct the adapter model to extract the image semantic information of the third image; The first image feature information of the third image and the first text feature information of the text prompt information are extracted through the cross attention layer in the adapter model. The adapter model generates the image semantic information based on the first image feature information and the first text feature information; The image semantic information is output through the adapter model.

10. The apparatus according to claim 8, wherein, The generation module is specifically used for: The image edge information and the image semantic information are input into the control condition model for processing, and the first image condition constraint information is output. The first image condition constraint information and the second image are input into the diffusion model; The target image is generated by expanding the image content of the second image based on the first image conditional constraint information using the diffusion model. The target image is output through the diffusion model.

11. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the image processing method as claimed in any one of claims 1 to 5.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the image processing method as described in any one of claims 1 to 5.

13. A computer program product stored in a storage medium, the computer program product being executed by at least one processor to implement the steps of the image processing method as claimed in any one of claims 1 to 5.

14. A chip comprising a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions to implement the steps of the image processing method as claimed in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic equipment

    CN111161176A

  • Model generation method and device, electronic equipment and storage medium

    CN117689996A

  • Image processing method, device and equipment, computer readable storage medium and product

    CN118365747A

  • Image processing method and device, electronic equipment and readable storage medium

    CN119521019A

  • Image fusion method and device

    US20230325994A1