Image processing method and related device

By taking multiple high-definition, low-field-of-view first images within a specific time period and stitching them together, the problem of poor object clarity in group photos is solved, and the generation of high-definition, large-field-of-view target images is achieved.

CN119277192BActive Publication Date: 2025-09-30HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410074980.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-09-30
Estimated Expiration
2044-01-18

AI Technical Summary

Technical Problem

The clarity of people in group photos is poor, especially when using the rear main camera or ultra-wide mode, where the field of view is larger but clarity is affected.

Method used

By taking multiple first images with higher clarity but smaller field of view within a specific time period and stitching them together, the target object in the second image is enhanced, and the image clarity is improved using super-resolution technology and feature alignment technology.

Benefits of technology

It improves the clarity of objects in group photos, combines the advantages of high definition and wide field of view, and solves the problem of poor object clarity in group photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119277192B_ABST
    Figure CN119277192B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide an image processing method and related apparatus, relating to the field of terminal technology. The method includes: shooting multiple first images in a first mode, shooting a second image in a second mode, and then obtaining a target image based on the first and second images. The clarity of the image obtained by shooting in the first mode is higher than the clarity of the image obtained by shooting in the second mode, and the field of view of the image obtained by shooting in the first mode is lower than the field of view of the image obtained by shooting in the second mode. In this way, the target image can be obtained by combining the advantages of a high field of view and high clarity, solving the technical problem of poor clarity of objects in group photos in the related art, and achieving the technical effect of improving the clarity of objects in group photos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to image processing methods and related devices. Background Art

[0002] In daily life, people often use electronic devices to capture multiple subjects, such as group photos. Typically, for large groups, the rear main camera or ultra-wide mode is used to capture photos to ensure everyone is within the frame.

[0003] However, the clarity of the characters in the group photo images needs to be further improved. Summary of the Invention

[0004] The present invention provides an image processing method and related apparatus for use in the field of terminal technology. This embodiment enhances a target image within a second image by using multiple first images with higher clarity and lower field of view than the second image. This solves the technical problem of poor object clarity in group photos in related technologies, achieving the technical effect of improving the clarity of objects in group photos.

[0005] In a first aspect, an embodiment of the present application provides an image processing method, applied to an electronic device, comprising:

[0006] A first interface is displayed, the first interface including a first preview image and a first button; in response to a first operation on the first button, a prompt message and a second button are displayed, the prompt message is used to prompt the capture of images of multiple target objects; within a target time period, multiple first images are captured using the first mode, and the multiple first images are displayed respectively; in response to a second operation on the second button, a second image is captured using the second mode, wherein the second button is continuously displayed during the target time period, the target time period includes a time period from receiving the first operation to receiving the second operation, the clarity of the image captured using the first mode is higher than the clarity of the image captured using the second mode, and the field of view angle of the image captured using the first mode is lower than the field of view angle of the image captured using the second mode; a target object in the second image is enhanced using a stitched image to obtain a target image, the stitched image including an image obtained by stitching multiple first images.

[0007] The first interface may be an interface for taking a group photo. Alternatively, the first interface may be an interface for entering a group photo mode. The first button may be a button for triggering the electronic device to capture an image in the first mode.

[0008] For example, the first interface may be as follows: Figure 5A or Figure 5B In the interface shown, the first button can be as follows Figure 5Aor Figure 5B the shutter button.

[0009] The second button may be a button for triggering the end of capturing the first image.

[0010] The target time period may refer to the time period during which the first image is captured. It should be noted that the first image captured within the target time period only means that the first image is captured between the time period between the receipt of the first button operation and the receipt of the second button operation, and does not mean that the electronic device is actually timing.

[0011] For example, the multiple first images of this embodiment may be as shown in FIG6A- Figure 6B The preview window 201 shown in FIG. 201 may be an image, or an image Figure 6C-6D The preview window 201 shown displays an image.

[0012] Optionally, each first image may include part of the target objects among all objects, the second image may include all target objects, the target image includes all target objects, and the clarity of each target object in the target image is higher than the clarity of the target object in the second image.

[0013] In an embodiment of the present application, the target image in the second image is enhanced by using multiple first images with higher clarity and lower field of view than the second image, thereby solving the technical problem of poor clarity of objects in group photos in the related art and achieving the technical effect of improving the clarity of objects in group photos.

[0014] In conjunction with the first aspect, in one possible implementation, the prompt information includes information for prompting the user to maintain a constant posture of the electronic device, and the plurality of first images are captured by the electronic device using a movable camera;

[0015] Alternatively, the prompt information includes information for guiding the movement trajectory, and the multiple first images are taken by the mobile electronic device.

[0016] In this embodiment, the specific prompt information may be related to the manner in which the electronic device captures the first image. For example, the prompt information when the electronic device controls the movement of the camera and captures the first image during the camera movement may be different from the prompt information when the user moves the electronic device and captures the first image during the camera movement.

[0017] For example, the second button can be Figures 6A-6E The "Finish button" shown in the figure can have the following prompt information: Figure 6A-6B The first prompt message shown, or Figure 6C-6D The second prompt information shown is not limited here.

[0018] It can be understood that by prompting the user to keep the posture of the electronic device unchanged and capturing multiple first images during the process of controlling the movement of the camera by the electronic device, the user does not need to move the electronic device, which can reduce the user's operating steps and improve the intelligence and convenience of the electronic device's shooting.

[0019] In addition, the movement trajectory of the electronic device is guided by prompt information, and multiple first images are captured by the user moving the electronic device. In this way, there is no need to set a device for controlling the movement of the camera in the electronic device, which can simplify the structure of the electronic device.

[0020] In conjunction with the first aspect, in one possible implementation, the image processing method further includes:

[0021] Displaying a second interface, the second interface including a second preview image, a third button, and a fourth button, the third button being in an unselected state, the fourth button being in a selected state, and the second preview image being a preview image captured in the mode of the fourth button;

[0022] The first interface is displayed, including:

[0023] In response to the third operation on the third button, a first interface is displayed, in which the third button is in a selected state and the fourth button is in an unselected state.

[0024] The second interface may refer to the interface displayed when the camera application is opened. The third button may refer to the button that triggers the display of the first interface.

[0025] For example, the second interface may be as follows Figure 4A In the interface shown, the third button can be as follows Figure 4A The photo mode option 2021 shown, the fourth button can be as follows Figure 4A Photo mode options shown.

[0026] In an embodiment of the present application, the third button can be directly displayed on the second interface, so that when the third operation on the third button is detected, the first interface is entered, and then multiple first images can be taken on the first interface. In this way, the efficiency of taking multiple first images can be improved, thereby improving the efficiency of image enhancement.

[0027] In conjunction with the first aspect, in one possible implementation, the image processing method further includes:

[0028] Displaying a third interface, the third interface includes a second preview image, a fourth button, and a fifth button, the fourth button is selected, the second preview image is a preview image captured in the mode of the fourth button, and the fifth button is unselected;

[0029] In response to a fourth operation on the fifth button, a fourth interface is displayed, where the fourth interface includes the third button, and the third button is in an unselected state;

[0030] The first interface is displayed, including:

[0031] In response to the third operation on the third button, a first interface is displayed, in which the third button is in a selected state and the fourth button is in an unselected state.

[0032] Among them, the third interface may refer to the interface displayed when the camera application is opened, and the fifth button may refer to the button that triggers the display of the fourth interface.

[0033] For example, the third interface may be as follows: Figure 4B The fourth button can be as shown in the interface Figure 4B The fourth interface can be as follows: Figure 4C As shown, the third button may refer to Figure 4C Photo mode options 2021 shown.

[0034] For example, the display of the second interface or the third interface may be that the electronic device detects the Figure 3 When the camera application icon 101 is shown, the interface displayed on the display screen of the electronic device.

[0035] In conjunction with the first aspect, in one possible implementation, using the stitched image to enhance the target object in the second image to obtain the target image includes:

[0036] The second image is segmented to obtain an object region mask; the spliced ​​image is corrected based on the object region mask to obtain a reference image, wherein the size of the object region in the reference image matches the size of the object region mask; the second image is super-resolved based on the reference image and the object region mask to obtain a target image.

[0037] The object region may refer to an area requiring image enhancement. For example, if the second image is a group photo of people, the object region may be a portrait region. For another example, if the second image is a photo of multiple plants, the object region may be plants. The size of the object region in the reference image matches the size of the object region mask, which may mean that the size of the object region in the reference image is consistent with the size of the object region mask.

[0038] The target image can be understood as the second image with higher definition. Specifically, the definition of the object in the target image is higher than the definition of the object in the second image.

[0039] In an embodiment of the present application, the second image is segmented to obtain an object area mask; the spliced ​​image is corrected based on the object area mask to obtain a reference image, and the size of the object area in the reference image matches the size of the object area mask; the second image is super-resolved based on the reference image and the object area mask to obtain a target image, that is, the spliced ​​image is first corrected to obtain the reference image, and then super-resolved based on the reference image. In this way, the reference image is more closely matched with the object area in the second image than the spliced ​​image. In this way, the accuracy and efficiency of super-resolution based on the reference image can also be higher than that of super-resolution based on the spliced ​​image.

[0040] In conjunction with the first aspect, in one possible implementation, super-resolving the second image based on the reference image and the object region mask to obtain the target image includes:

[0041] Feature extraction is performed on the reference image and the second image respectively to obtain reference image features of the reference image and second image features of the second image; the reference image features and the second image features are feature aligned based on the object area mask to obtain a pixel-level offset map, and the pixel-level offset map indicates the offset information of each pixel in the second image in the reference image; the reference image is encoded (encoder) to obtain a reference feature map; a deformable convolution operation or an attention-based convolution operation is performed on the reference feature map based on the pixel-level offset map to obtain texture features to be transferred; the texture features to be transferred are transferred to the second image to obtain a target image.

[0042] Among them, texture features can refer to regularly arranged patterns within a certain range in an image. Texture is one of the important features of image processing.

[0043] In an embodiment of the present application, the offset direction and offset amount of each pixel of the second image in the reference image can be determined based on the pixel-level offset map. In this way, the variable convolution operation or the attention-based convolution operation of the encoded feature map takes into account the pixel offset between the reference image and the second image. In this way, the correlation between the obtained migrated texture features and the object area is higher, and the enhanced area of ​​the target image obtained after migration to the second image is also more accurate.

[0044] In conjunction with the first aspect, in one possible implementation, the method further includes:

[0045] Performing single-frame super-resolution on the second image to obtain a second image after single-frame super-resolution;

[0046] Migrating the texture features to be migrated to the second image to obtain a target image includes:

[0047] The texture features to be migrated are migrated to the second image after single-frame super-resolution to obtain the target image.

[0048] The single-frame super-resolution may be performed on the entire second image, including but not limited to super-resolution of the object area and super-resolution of the non-object area. Taking a group portrait as an example, super-resolution may be performed on the portrait area and the background area.

[0049] In the embodiment of the present application, the second image is subjected to single-frame super-resolution to obtain a single-frame super-resolved second image. The texture features to be transferred are then transferred to the single-frame super-resolved second image to obtain the target image. This represents at least two levels of image super-resolution. This further improves the super-resolution effect of the object area, thereby enhancing the clarity of the object area. Furthermore, super-resolution can also be performed on non-object areas, thereby further improving the overall clarity of the image.

[0050] In combination with the first aspect, in a possible implementation method, the target image is obtained by super-resolving a pre-trained super-resolution model based on a reference image, an object area mask and a second image, and the pre-trained super-resolution model includes a feature extraction network, a feature alignment network, an encoder, a convolutional network and a texture transfer network; the feature extraction network is used to extract features from the reference image and the second image respectively to obtain reference image features of the reference image and second image features of the second image; the feature alignment network is used to align the reference image features and the second image features based on the object area mask to obtain a pixel-level offset map; the encoder is used to encode the reference image to obtain a reference feature map; the convolutional network is used to perform a deformable convolution operation or an attention-based convolution operation on the reference feature map based on the pixel-level offset map to obtain texture features to be transferred; the texture transfer network is used to migrate the texture features to be transferred to the second image to obtain a target image.

[0051] In conjunction with the first aspect, in one possible implementation, the target image is obtained by migrating the texture features to be migrated to a second image after single-frame super-resolution, and the pre-trained super-resolution model further includes a single-frame super-resolution network;

[0052] The single-frame super-resolution network is used to perform single-frame super-resolution on the second image to obtain a second image after single-frame super-resolution;

[0053] The texture transfer network is used to transfer the texture features to be transferred to the second image after single-frame super-resolution to obtain the target image.

[0054] In conjunction with the first aspect, in one possible implementation, correcting the stitched image based on the object region mask to obtain a reference image includes:

[0055] The stitched image is stitched with the object region mask, and features are extracted from the stitched image stitched with the object region mask to obtain a stitched feature map; correction processing is performed N times based on the stitched feature map to obtain a reference image, where N is an integer not less than 2;

[0056] Each calibration process includes:

[0057] Perform network regression on the input, as well as mesh motion and bending processing, to obtain the output;

[0058] Among them, the input of the first correction process is the spliced ​​image with the object area mask spliced, the input of the Mth correction process is the output obtained by the M-1th correction process, and the output obtained by the Nth correction process is used as the reference image, 1<M≤N, and M is an integer.

[0059] In the embodiment of the present application, after N correction processes, that is, multi-level correction processes, rough correction processes and refined correction processes can be performed in sequence. In this way, the effect of the correction process can be improved, and then the effect of image enhancement can be improved.

[0060] In a second aspect, the present application also provides an image processing method, including:

[0061] A second image and multiple first images are acquired, where the multiple first images are acquired by shooting using the first mode, and the second images are acquired by shooting using the second mode, wherein the clarity of the images acquired by shooting using the first mode is higher than the clarity of the images acquired by shooting using the second mode, and the field of view of the images acquired by shooting using the first mode is lower than the field of view of the images acquired by shooting using the second mode; the multiple first images are stitched together to obtain a stitched image; and a target object in the second image is enhanced using the stitched image to obtain a target image.

[0062] In an embodiment of the present application, the image processing method may also be executed by a server. For example, after the electronic device captures the first image and the second image, the first image and the second image are sent to the server. The server then returns a target image based on the first image and the second image, and the electronic device stores the target image.

[0063] In a third aspect, the present application also proposes a method for training a super-resolution model, the method comprising:

[0064] A super-resolution model to be trained and training samples are obtained. The super-resolution model to be trained includes a feature extraction network, a feature alignment network, an encoder, a convolutional network and a decoder. The feature extraction network is used to extract features from the reference image and the second image respectively to obtain reference image features of the reference image and second image features of the second image. The feature alignment network is used to align the reference image features and the second image features based on the object area mask to obtain a pixel-level offset map. The pixel-level offset map indicates the offset information of each pixel in the second image in the reference image. The object area mask is obtained by segmenting the second image. The encoder is used to encode the reference image to obtain a reference feature map. The convolutional network is used to perform a deformable convolution operation or an attention-based convolution operation on the reference feature map based on the pixel-level offset map to obtain texture features to be transferred. The decoder is used to decode the texture features (decoder) to obtain texture information. The texture information is used to supervise the training process of the super-resolution model to be trained. The super-resolution model to be trained is trained based on the training samples until the training of the super-resolution model to be trained is completed, thereby obtaining a pre-trained super-resolution model.

[0065] In an embodiment of the present application, after the super-resolution model is trained, image enhancement processing can be performed based on the trained super-resolution model, so that the clarity of the object in the image can be improved.

[0066] In a fourth aspect, an embodiment of the present application provides an image processing device, which may be an electronic device or a chip or chip system within an electronic device. The image processing device may include a display unit and a processing unit. When the image processing device is an electronic device, the display unit may be a display screen. The display unit is used to perform the display step so that the electronic device implements an image processing method described in the first aspect or any possible implementation of the first aspect. When the image processing device is an electronic device, the processing unit may be a processor. The image processing device may also include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the electronic device implements an image processing method described in the first aspect or any possible implementation of the first aspect. When the image processing device is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit so that the electronic device implements an image processing method described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (eg, a register, a cache, etc.), or a storage unit within the electronic device that is located outside the chip (eg, a read-only memory, a random access memory, etc.).

[0067] In a fifth aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, the memory being used to store code instructions, and the processor being used to run the code instructions to execute the method described in the first aspect, the second aspect, the third aspect, any possible implementation of the first aspect, any possible implementation of the second aspect, or any possible implementation of the third aspect.

[0068] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run on a computer, the computer executes the method described in the first aspect, the second aspect, the third aspect, any possible implementation of the first aspect, any possible implementation of the second aspect, or any possible implementation of the third aspect.

[0069] In the seventh aspect, an embodiment of the present application provides a computer program product including a computer program. When the computer program is run on a computer, the computer executes the method described in the first aspect, the second aspect, the third aspect, any possible implementation of the first aspect, any possible implementation of the second aspect, or any possible implementation of the third aspect.

[0070] In an eighth aspect, the present application provides a chip or chip system, which includes at least one processor and a communication interface, the communication interface and the at least one processor being interconnected by a line, and the at least one processor being used to run a computer program or instruction to execute the method described in the first aspect, the second aspect, the third aspect, any possible implementation of the first aspect, any possible implementation of the second aspect, or any possible implementation of the third aspect. The communication interface in the chip can be an input / output interface, a pin, or a circuit, etc.

[0071] In one possible implementation, the chip or chip system described above in this application further includes at least one memory, in which instructions are stored. The memory may be a storage unit within the chip, such as a register, a cache, etc., or a storage unit of the chip (e.g., a read-only memory, a random access memory, etc.).

[0072] It should be understood that the second to sixth aspects of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0074] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;

[0075] Figure 3 A schematic diagram of an interface for opening a camera application provided in an embodiment of the present application;

[0076] Figures 4A-4C A schematic diagram showing a default shooting preview interface provided in an embodiment of the present application;

[0077] Figure 5A-5B A schematic diagram of an interface for entering a group photo mode provided in an embodiment of the present application;

[0078] Figures 6A-6E A schematic diagram of an interface for capturing multiple first images provided in an embodiment of the present application;

[0079] Figure 7 A schematic diagram of an interface for capturing a second image provided in an embodiment of the present application;

[0080] Figure 8 A schematic diagram of a specific implementation of an image processing method provided in an embodiment of the present application;

[0081] Figure 9 A schematic diagram of a specific implementation of image enhancement provided in an embodiment of the present application;

[0082] Figure 10 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0083] To facilitate a clear description of the technical solutions of the embodiments of the present application, some of the terms and technologies involved in the embodiments of the present application are briefly introduced below:

[0084] 1. Super-resolution (SR), also known as super-resolution for short. The counterpart of super-resolution is low-resolution images (LR). Super-resolution is a low-level image processing task that maps a low-resolution image to a high-resolution image in order to enhance image details. There are many reasons for image blurriness, such as various types of noise, lossy compression, downsampling, and so on. Super-resolution is a classic application of computer vision. SR refers to the reconstruction of a corresponding high-resolution image from an observed low-resolution image (in other words, improving the resolution) through software or hardware methods. It has important application value in monitoring equipment, satellite image remote sensing, digital high-definition, microscopic imaging, video coding communications, video restoration, and medical imaging.

[0085] 2. The field of view angle is also called the field of view in optical engineering. The size of the field of view angle determines the field of view of the optical instrument. The field of view angle can also be expressed as FOV, and its relationship with the focal length is as follows: image height = EFL*tan(half FOV); EFL is the focal length; FOV is the field of view angle. The size of the field of view angle determines the field of view of the optical instrument. The larger the field of view angle, the larger the field of view and the smaller the optical magnification. In layman's terms, the target object will not be included in the lens if it exceeds this angle. According to the field of view angle classification, it includes but is not limited to the following types of lenses:

[0086] Standard lens: The viewing angle is about 45 degrees and has a wide range of uses.

[0087] Telephoto lens: The viewing angle is within 40 degrees and can be used for shooting at long distances.

[0088] Wide-angle lens: The viewing angle is more than 60 degrees, the observation range is larger, and the image at close range is distorted.

[0089] 3. Image segmentation is the technique and process of dividing an image into several specific regions with unique properties and identifying objects of interest. It is a key step in the transition from image processing to image analysis. Existing image segmentation methods are mainly categorized as follows: threshold-based, region-based, edge-based, and those based on specific theories. From a mathematical perspective, image segmentation is the process of dividing a digital image into mutually disjoint regions. The image segmentation process is also a labeling process, where pixels belonging to the same region are assigned the same number.

[0090] In an embodiment of the present application, the second image can be segmented to obtain an object area mask. A mask can be understood as data similar to an image. In an embodiment of the present application, the image and the mask can be fused to draw more attention to part of the content in the image. Generally, a mask can be used to extract a region of interest, such as fusing a pre-made region of interest mask with the image to be processed to obtain an image of the region of interest, where the image values ​​in the region of interest remain unchanged, while the image values ​​outside the region are all 0. It can also play a shielding role, using a mask to shield certain areas on the image so that they do not participate in the processing or the calculation of the processing parameters, or only processing or statistics are performed on the shielded area. In an embodiment of the present application, it can be a mask of an object area, such as a mask of a portrait area or a mask of a body area.

[0091] 4. Feature maps refer to image features extracted from the input image through convolution operations in convolutional neural networks. In a convolutional neural network, the output of each layer is a three-dimensional tensor, where the third dimension represents the number of feature maps. Each feature map is obtained by convolving the feature map of the previous layer with several convolution kernels, each corresponding to a specific feature. Therefore, feature maps can be viewed as responses to specific features in the input image and can be used to understand the working principles of convolutional neural networks and visualize their feature extraction process.

[0092] 5. Image features Image features mainly include: color features, shape features, texture features and spatial relationship features.

[0093] 6. The purpose of image stitching rectangularization is to address the problem of irregular boundaries after stitching. Existing image stitching rectangularization methods typically involve two stages: the first stage is to search for an initial grid, which involves placing a regular grid on the stitched image to describe the position of each point in the image; the second stage is to optimize a target grid, which involves deforming the initial grid so that the grid edges align as closely as possible with the rectangular boundaries. Then, by transforming the stitched image from the initial grid to the target grid, a rectangular image is obtained. This process is called grid deformation or grid warping.

[0094] 7. Texture transfer super-resolution processing can be implemented in a variety of ways. For example, texture transfer super-resolution processing can be implemented using a neural network model (texture transfer super-resolution neural network model) to output a super-resolution video. The neural network model can be any neural network model, such as a deep neural network (DNN), a convolutional neural network (CNN), or a combination thereof.

[0095] 8. Other terms

[0096] In the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the terms "first chip" and "second chip" are used solely to distinguish between different chips and do not define their order. Those skilled in the art will understand that terms such as "first" and "second" do not define the quantity or execution order, and do not necessarily define differences.

[0097] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0098] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, c can be single or multiple.

[0099] 9. Electronic devices

[0100] The electronic devices of the embodiments of the present application may include handheld devices, vehicle-mounted devices, etc. with wireless communication functions. For example, some electronic devices include: mobile phones, tablet computers, PDAs, laptop computers, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving cars, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, wearable devices, terminal devices in 5G networks or future evolved public land mobile communication networks. The terminal equipment in the network (PLMN), etc., is not limited to this in the embodiments of the present application.

[0101] As an example and not a limitation, in the embodiments of the present application, the electronic device may also be a wearable device. Wearable devices may also be referred to as wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0102] In addition, in the embodiments of the present application, the electronic device can also be a terminal device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0103] The electronic devices in the embodiments of the present application may also be referred to as: terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, etc.

[0104] In the embodiments of the present application, the electronic device or each network device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also known as main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a Windows operating system. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0105] The following examples illustrate the scenarios of the embodiments of the present application.

[0106] In some example situations, such as Figure 1 As shown, a group photo needs to be taken. Generally, in order to capture an image that can accommodate all the people, it is necessary to use the rear main camera or ultra-wide-angle mode to shoot, so that the field of view is wide enough to accommodate everyone. However, a wide field of view generally means a short focal length, and a short focal length means sacrificing a certain degree of image clarity, which will result in poor clarity of the people in the group photo.

[0107] In other exemplary situations, the focal length of the telephoto lens is longer than that of the main camera or ultra-wide-angle camera, and the image taken with the telephoto lens is clearer than that taken with the main camera or ultra-wide-angle camera, but the field of view of the telephoto lens is also smaller than that of the main camera or ultra-wide-angle camera, and the image taken with the telephoto lens may not be able to accommodate everyone.

[0108] In view of this, the embodiments of the present application propose an image processing method and a related device, which can enhance the second image through multiple first images. Since the first image has higher clarity but a smaller field of view than the second image, the advantages of the field of view of the first image and the clarity of the second image can be combined to obtain a target image with a large field of view and high clarity. In this way, compared with the second image, the clarity of the object in the target image is higher, thereby solving the technical problem of poor clarity of objects in group photos in the related technology and achieving the technical effect of improving the clarity of objects in group photos.

[0109] In order to better understand the embodiments of the present application, the structure of the electronic device according to the embodiments of the present application is introduced below:

[0110] Figure 1 A schematic diagram of the hardware structure of the electronic device 10 is shown.

[0111] The electronic device 10 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0112] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0113] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0114] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0115] In an embodiment of the present application, camera 193 can capture preview image data, first image data, and second image data, and display the image on display screen 194. Camera 193 may include a main camera for capturing preview image data, a telephoto camera for capturing first image data, and a main camera or an ultra-wide-angle camera for capturing second image data. Processor 110 may perform image enhancement processing, for example, by using a neural network processor based on a neural network model.

[0116] Figure 2 A schematic diagram of the software structure of the electronic device 10 is shown.

[0117] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present invention, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.

[0118] Figure 2 This is a block diagram of the software structure of electronic device 100 according to an embodiment of the present invention. A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other via software interfaces. In some embodiments, the Android system is divided into multiple layers: from top to bottom, the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0119] The application layer can include a series of application packages.

[0120] like Figure 2As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0121] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0122] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0123] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.

[0124] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0125] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0126] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).

[0127] The Hardware Abstraction Layer (HAL) is an interface layer between the operating system kernel and upper-level software. Its purpose is to abstract the hardware. It serves as an abstract interface for device kernel drivers, providing application programming interfaces (APIs) that access the underlying devices to higher-level Java API frameworks. The HAL includes multiple library modules, such as the Camera HAL and the Display HAL.

[0128] Each library module implements an interface for a specific type of hardware component. For example, the Camera HAL provides the Camera FWK with an interface to access hardware components like the camera. The Display HAL provides the Display FWK with an interface to access hardware components like the display. When the system framework API requires access to the device's hardware, the Android operating system loads the library module for that hardware component.

[0129] The kernel layer is the layer between hardware and software. It includes at least the display driver, camera driver, audio driver, and sensor driver. For example, the camera driver controls the camera to capture image data. Optionally, the camera driver can include the main camera driver, the ultra-wide-angle camera driver, and the telephoto camera driver.

[0130] The following is an example of the workflow of electronic device software.

[0131] For example, when a touch sensor in a terminal device receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including information such as touch coordinates, touch force, and the timestamp of the touch operation). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the button corresponding to the input event. For example, if the touch operation is a click operation and the virtual button corresponding to the click operation is the "image capture button," the camera application calls the application framework layer's interface, which in turn calls the kernel layer to activate the display driver and the telephoto camera driver, causing the display screen to display the first image captured by the telephoto camera. Then, when the "Done" virtual button is clicked, the kernel layer calls the main camera driver or the ultra-wide-angle camera driver, causing the display screen to display the second image captured by the main camera or the ultra-wide-angle camera. The camera application then generates a target image based on the first and second images and saves the target image in an image storage file, such as an "album."

[0132] The following describes in detail the shooting scenarios provided by this application with reference to some exemplary user interface diagrams.

[0133] It is understood that the terms "interface" and "user interface" in the embodiments of the present application refer to the media interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. A common form of user interface is a graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be an interface element such as an icon, window, button, etc. displayed on the display screen of an electronic device, where a button can include a visual interface element such as an icon, button, menu, tab, text box, dialog box, status bar, navigation bar, widget, etc.

[0134] 1. Open the Camera app ( Figure 3 ).

[0135] like Figure 3As shown, the user interface 100 (also known as the main interface) displays a page with application icons placed on it. The page may include multiple application icons (for example, a camera application, a weather application icon, a calendar application icon, an album application icon, a note application icon, an email application icon, an application store application icon, a settings application icon, etc.). A page indicator may also be displayed below the multiple application icons to indicate the positional relationship between the currently displayed page and other pages. Below the page indicator are multiple application icons (for example, a camera application icon 101, a browser application icon, a message application icon, and a dial application icon). These application icons remain displayed when the page is switched.

[0136] It is understandable that the camera application icon 101 is an icon of a camera application (ie, a camera application). The camera application icon 101 can be used to trigger the launch of the camera application.

[0137] In this embodiment, the electronic device may detect a user operation on the camera application icon 101 , and in response to the operation, the electronic device may display an initial shooting preview interface.

[0138] It is understandable that the user operations mentioned in this application may include but are not limited to touch (for example, click, etc.), voice control, gestures and other operations, and this application does not limit this.

[0139] 2. Display the default shooting preview interface ( Figures 4A-4C ).

[0140] (1) The default shooting preview interface includes the group photo mode option ( Figure 4A ).

[0141] like Figure 4A As shown, the user interface 200 is a shooting interface of the default shooting mode of the camera application, and the user can preview the image and complete the shooting on this interface.

[0142] The user interface 200 may include a preview window 201 , a camera mode option 202 , an album shortcut button, a shutter button (also known as a capture button), and a camera flip button.

[0143] The preview window 201 of the user interface 200 can be used to display a preview image. The preview image displayed in the preview window 201 of the user interface 200 can be an image captured by the camera of the electronic device based on the framing range when the photo mode option is selected.

[0144] One or more shooting mode options may be displayed in the camera mode option 202. The one or more shooting mode options may include: an aperture mode option, a night scene mode option, a smart portrait mode option, a photo mode option, a group photo mode option 2021, a movie mode option, and more options.

[0145] It is understandable that the camera mode option 202 may further include more or fewer shooting mode options.

[0146] (2) The default shooting preview interface does not include the group photo mode option ( Figure 4B-4C ).

[0147] like Figure 4B As shown, the user interface 200 does not include a group photo mode option 2021 .

[0148] In such Figure 4B In the user interface 200 shown, the electronic device can detect the Figure 4B In response to a user operation of the more option button of Figure 4C User interface 300 is shown. In user interface 300, photo mode options other than those displayed in user interface 200 may be displayed. For example, user interface 300 may display a professional mode option, a panoramic mode option, a high-dynamic range (HDR) mode option, a time-lapse photography mode option, a watermark mode option, a document correction mode option, a high-pixel mode option, and a group photo mode option 2021.

[0149] It is understandable that, in the user interface 300 , in addition to the group photo mode option 2021 , other mode options may be displayed according to actual circumstances.

[0150] It should be noted that, in the user interface 200, when the electronic device detects a sliding operation (such as a left or right swipe) on the camera mode option 202, the camera mode option displayed by the camera mode option 202 is updated, thereby displaying the group photo mode option 2021.

[0151] When the electronic device detects a user operation on the group photo mode option 2021 , the electronic device enters the group photo mode.

[0152] 3. Enter group photo mode ( Figure 5A-5B )

[0153] (1) After entering the group photo mode, a low-angle and high-definition preview image is displayed. Figure 5A )

[0154] like Figure 5A As shown, the preview image displayed in the preview window of the user interface 400 may be an image captured by the camera of the electronic device based on the viewing range when the group photo mode option 2021 is selected. Multiple people may be displayed in the preview image.

[0155] For example, Figure 5AAs shown, after entering the group photo mode, the preview image displayed in the preview window can be of low field angle and high definition.

[0156] (2) After entering the group photo mode, a preview image with a high field of view and low resolution is displayed ( Figure 5B ).

[0157] like Figure 5B As shown, the preview image displayed in the preview window of the user interface 400 may be an image with a high field of view and low definition.

[0158] When the electronic device detects a user operation on a shutter button, the electronic device starts capturing multiple images with a low field of view and high definition.

[0159] 4. Take multiple images with low field of view and high resolution ( Figures 6A-6E ).

[0160] (1) Electronic equipment automatically captures multiple images with low field of view and high resolution ( Figure 6A 、 Figure 6B and Figure 6E ).

[0161] like Figure 6A 、 Figure 6B and Figure 6E As shown, when the electronic device detects a user operation on the shutter button, the camera of the electronic device begins to capture and save multiple high-definition images with a low field of view angle. During the process of the camera capturing multiple high-definition images with a low field of view angle, the preview image displayed in the preview window 201 can also change in real time, so that the user can determine whether the multiple high-definition images with a low field of view angle include all people. During the shooting process, a text prompt such as "All people in the group photo are captured separately. Please keep the phone fixed during shooting" can be continuously displayed.

[0162] It is understandable that the prompt information may require that the content of the portrait group photo area in the image with high field angle and low definition is already fully contained in the multiple images with low field angle and high definition in the previous step.

[0163] (2) The user operates an electronic device to capture multiple images with low field of view and high resolution (6C- Figure 6E )

[0164] like Figure 6C-6E As shown, optionally, it can continuously display information such as "All people in the photo are framed separately, please move your phone as estimated below", and provide guidance on the corresponding movement trajectory, and each time the electronic device detects a user operation on the shutter button, it takes a low-field-angle and high-definition multiple image and saves it.

[0165] It should be noted that the guidance of the movement trajectory is not limited to moving in a direction parallel to the horizontal line, but can also be moving in a direction perpendicular to the horizontal line, as long as all the people in the photo can be included in the frame. There is no limitation here.

[0166] When the electronic device detects a user operation on the completion button, it stops capturing images with a low field angle and high definition, and then it can start capturing images with a high field angle and low definition.

[0167] 5. Take images with high field of view and low resolution ( Figure 7 )

[0168] For example, the user interface 400 may display a prompt message to prompt that all people need to be in the frame at the same time. For example, the prompt message may include text prompts such as "All people in the group photo need to be in the frame at the same time. Please keep the phone fixed during the shooting process."

[0169] When the electronic device detects a user operation on the shutter button, an image with a high field angle and low definition is captured.

[0170] It should be noted that it can also be that the electronic device detects Figure 6E When the user operates the Done button shown, an image with a high field angle and low resolution is automatically captured.

[0171] In this way, multiple low-field-angle and high-definition images can be obtained, each low-field-angle and high-definition image includes some of the characters, and the set of characters in all low-field-angle and high-definition images includes all the characters. At the same time, the high-field-angle and low-definition image includes all the characters. Furthermore, image enhancement can be performed based on the high-field-angle and low-definition image and the multiple low-field-angle and high-definition images to obtain the target image, so that the clarity of the characters in the obtained target image is higher than the clarity of the characters in the high-field-angle and low-definition image.

[0172] Then, the target image is saved in the album. When the electronic device detects a user operation on the album shortcut button, the target image can be displayed.

[0173] It is understandable that the shooting scenes in the above examples, in addition to shooting of group photos of people, can also be scenes of shooting multiple animals, multiple plants or multiple target objects obtained by arrangement and combination, which will not be described in detail here.

[0174] It should be noted that the above user interfaces are merely some examples provided for this application and should not be regarded as limitations of this application.

[0175] The following describes the specific implementation of the above embodiment in conjunction with a group photo scene. Figure 8The image processing method shown may include:

[0176] S800: The camera application receives a request to start the camera application.

[0177] In this embodiment, the request to start the camera application can be used to request to start the camera application. Optionally, the electronic device can send a request to start the camera application to the camera application when detecting an operation that triggers the start of the camera application. For example, Figure 3 As shown, when the electronic device detects a user operation on the camera application icon 101 of the user interface 100 , it sends a request to the camera application to start the camera application.

[0178] It is understandable that the above examples are examples of some camera applications receiving requests to start the camera application, and this embodiment is not limited thereto.

[0179] S802: The camera application sends a request to the main camera to start the main camera.

[0180] S804: The main camera starts and collects second preview image data.

[0181] The second preview image data may refer to preview image data collected after the camera application is started.

[0182] S806: The main camera sends the second preview image data to the camera HAL.

[0183] S808: The camera HAL processes the second preview image data and sends the processed image to the display screen for display.

[0184] Since the second preview image data is the image data collected by the main camera when the camera application is started, the second preview image is displayed when the camera application is started. Figure 4A As shown, the second preview image may be a preview image displayed in the preview window 201 of the user interface 200 .

[0185] It is understandable that the above provides some examples of the second preview images, and the second preview image actually displayed is related to the direction of the main camera when it is activated, and this embodiment is not limited to this.

[0186] S810: The camera application receives a request to start a group photo mode.

[0187] The request to open the group photo mode can be used to request to enter the group photo mode. Optionally, the electronic device can send a request to open the group photo mode to the camera application when detecting a trigger operation to enter the group photo mode. For example, Figure 4AAs shown, when the electronic device detects a user operation acting on the group photo mode option 2021, it sends a request to the camera application to turn on the group photo mode. Figure 4B and Figure 4C As shown, when the electronic device detects a user operation on more buttons, it enters Figure 4C In the user interface 300 shown, when the electronic device detects a user operation on the group photo mode option 2021, it sends a request to the camera application to enable the group photo mode.

[0188] It is understandable that the above examples are examples of some camera applications receiving requests to enable group photo mode, and this embodiment is not limited thereto.

[0189] S812: The camera application sends a request to the telephoto camera to start the telephoto camera.

[0190] In this embodiment, the telephoto camera has a longer focal length than the main camera, and the corresponding field of view angle is lower, but the corresponding image clarity is also higher.

[0191] S814: The telephoto camera is started and first preview image data is collected.

[0192] The first preview data may be understood as preview image data collected after entering the group photo mode.

[0193] S816: The telephoto camera sends first preview image data to the camera HAL.

[0194] S818. The camera HAL processes the first preview image data and sends the processed image to the display screen for display.

[0195] In this embodiment, since the first preview image data is the image data collected by the telephoto camera when the telephoto camera is activated, when entering the group photo mode, the first preview image is displayed. Figure 5A As shown, the first preview image may be the preview image displayed in the preview window 201 of the user interface 400 .

[0196] It is understandable that the above provides some examples of the first preview images. The first preview image actually displayed is related to the direction of the telephoto camera when it is activated, and this embodiment is not limited to this.

[0197] In another possible implementation, when entering the group photo mode, the first preview image data may be collected by the main camera. Figure 5B As shown, the first preview image may also be the preview image displayed in the preview window 201 of the user interface 400 .

[0198] It is understandable that by starting the telephoto camera when entering the group photo mode, the telephoto camera can be called in time to capture multiple first images, thereby improving the efficiency of capturing multiple first images and further improving the efficiency of image enhancement.

[0199] S820: The camera application receives a click operation on the capture button.

[0200] The click operation of this embodiment may be one or more times, which is not limited here. In other possible implementations, the click operation may also be a user operation such as a touch operation or a voice control operation.

[0201] S822: The camera application sends a request to the telephoto camera to capture a first image.

[0202] S824: The telephoto camera captures a first image.

[0203] Wherein, the telephoto camera may capture a first image each time it receives a request to capture a first image.

[0204] For example, the telephoto camera is movable, such as Figure 6A and Figure 6B As shown, after receiving a click operation on the shooting button, the telephoto camera is controlled to move and a request to shoot the first image is sent to the telephoto camera at a set interval. Each time the telephoto camera receives a request to shoot the first image, it shoots a first image. Since the telephoto camera is in the process of moving, the position of the telephoto camera is not fixed when the telephoto camera receives a request to shoot the first image at different time points. In this way, the telephoto camera can capture multiple first images from different perspectives, that is, the image content of the multiple first images is different. In this way, the multiple first images can capture all the people. Optionally, if the first image is shot by the telephoto camera, the first image can also be a telephoto image.

[0205] Optionally, the telephoto camera can be moved in ways that include, but are not limited to, translation or rotation about a fixed axis. Exemplarily, the telephoto camera can move in a first direction and / or a second direction, with the first direction and the second direction being perpendicular. In this way, the electronic device can capture multiple first images in different postures, for example, capturing multiple first images in a landscape or portrait orientation. Exemplarily, the telephoto camera can rotate about a first fixed axis and / or about a second fixed axis, with the first fixed axis and the second fixed axis being perpendicular. In this way, the electronic device can also capture multiple first images in different postures.

[0206] Another example is Figure 6C and Figure 6DAs shown, each time a capture button click operation is received, a request to capture the first image is sent. In this way, the user can move the electronic device each time after capturing a first image. In this way, the telephoto camera is located at a different position each time a request to capture the first image is received. In this way, the telephoto camera can capture multiple first images from different perspectives.

[0207] Optionally, the multiple first images of this embodiment can be as follows Figure 6A and Figure 6B The preview window 201 shown displays a low-viewing-angle and high-definition image.

[0208] S826: The telephoto camera sends the multiple first images to the camera application.

[0209] In this embodiment, the camera application may save multiple first images so as to perform image enhancement based on the multiple first images.

[0210] S828: The camera application receives a click operation on the Done button.

[0211] In this embodiment, when the camera application receives the click operation of the completion button, it means that the user believes that the multiple first images can include all the people, and the shooting of the first image can be ended at this time.

[0212] S830: The camera application sends a request to the main camera to capture a second image.

[0213] In this embodiment, the camera application may send a request to the main camera to capture the second image when receiving a click operation of the completion button; or when receiving a click operation of the completion button; Figure 7 The illustrated example acts as a request to the main camera to capture a second image when the user operates the capture button.

[0214] S832: The main camera captures a second image.

[0215] For example, the second image may be Figure 7 Since the second image has a lower definition than the first image, the first image may also be a low-resolution image (LR image).

[0216] S834: The main camera sends the second image to the camera application.

[0217] In another possible implementation, the second image may be captured by an ultra-wide-angle camera.

[0218] S836: The camera application stitches the multiple first images to obtain a stitched image, and uses the stitched image to enhance the person in the second image to obtain a target image.

[0219] In this embodiment, the main purpose of enhancing the person is to make the person clearer, that is, the person's clarity in the target image is higher than that in the second image.

[0220] It should be noted that in some other possible implementations, the second image can be captured with the main camera first, and then multiple first images can be captured with the telephoto camera. In this way, when switching from normal photo mode to group photo mode, preview image data can continue to be collected with the main camera, thereby reducing the frequent switching of cameras, and thus reducing the frequent switching of preview images on the preview interface, which can improve the user experience. In addition, it can also reduce the time wasted due to switching back and forth between cameras, thereby improving the efficiency of image processing.

[0221] It should be noted that in other possible implementations, the first image may be captured using a combination of one or more cameras. This embodiment does not limit how the first image is captured, as long as the first image has higher clarity and a lower field of view than the second image. Similarly, the second image may be captured using a combination of one or more cameras. This embodiment does not limit how the second image is captured, as long as the second image has a wider field of view and lower clarity than the first image.

[0222] Based on the above embodiments, Figure 9 The specific implementation of how to perform image enhancement is explained.

[0223] 1. The first images are stitched together to obtain a stitched image.

[0224] In this embodiment, a plurality of first images may be stitched together to obtain a stitched image containing the complete group portrait area. It is understood that if the first image is a telephoto image, the stitched image may also be understood as a telephoto stitched image.

[0225] Optionally, the image stitching method of this embodiment can be image stitching based on key point matching. For example, it can include but is not limited to the following steps: image preprocessing: including basic operations of digital image processing (such as denoising, edge extraction, histogram processing, etc.), establishing a matching template for the image, and performing certain transformations on the image (such as Fourier transform, wavelet transform, etc.). Then, image registration is performed: using a certain matching strategy, the corresponding position of the template or feature point in the image to be stitched in the reference image is found, and then the transformation relationship between the two images is determined. Then, a transformation model is established: based on the corresponding relationship between the template or image features, the parameter values ​​in the mathematical model are calculated, thereby establishing a mathematical transformation model of the two images. Then, a unified coordinate transformation is performed: based on the established mathematical transformation model, the image to be stitched is transformed into the coordinate system of the reference image to complete the unified coordinate transformation. Then, a fusion reconstruction is performed: the overlapping areas of the images to be stitched are fused to obtain a smooth and seamless panoramic image of the stitched reconstruction.

[0226] 2. Perform portrait segmentation on the second image to obtain a portrait mask.

[0227] The portrait mask may be a mask corresponding to the entire portrait area, that is, the portrait mask is a mask of the entire portrait area. In addition, the portrait mask may refer to a mask corresponding to each portrait, which is not limited here.

[0228] Optionally, the portrait segmentation method of this embodiment can be to perform segmentation through a specific portrait segmentation algorithm. The portrait segmentation algorithm mainly includes two categories: semantic segmentation and instance segmentation. Semantic segmentation refers to assigning each pixel in the image to a specific category, such as human body, background, etc.; while instance segmentation is to group pixels belonging to the same entity together on the basis of semantic segmentation to form independent instances. Optionally, the portrait segmentation of the second image can be performed through a trained portrait segmentation model. Specifically, it can be a portrait segmentation model based on deep learning training, and any neural network with sufficient accuracy can be used.

[0229] It should be noted that the portrait area in this embodiment can include both the face area and the body area. If the portrait area includes both the face area and the body area, then the object area to be super-resolved is also the face area and the body area. This way, the problem of obvious difference in clarity between the face and the body after enhancement can be avoided.

[0230] 3. Rectangularize the image.

[0231] The portrait mask and the stitched image are input into the image rectangularization module to rectangularize the reference image. That is, the stitched image can be rectangularized based on the portrait mask through the image rectangularization module to obtain the reference image.

[0232] In this embodiment, the purpose of image matrixing is mainly to perform correction processing, such as removing black edges around the mosaic image, reshaping the mosaic image into a rectangular size, etc. Optionally, after the portrait area in the mosaic image is reshaped, the size of the object area in the reference image matches the size of the object area mask.

[0233] Optionally, the splicing is combined with the portrait mask for feature extraction. The obtained feature map is passed through the grid regression network and the grid motion and bending (warp, also known as distortion) process to obtain a first-stage output, which represents the coarse-scale network result. The first-stage result is input into the grid regression network again for grid motion prediction. The obtained result grid motion and warping give the final rectangular result.

[0234] Features can be extracted using a feature extractor. Optionally, convolutional pooling blocks are stacked to extract high-level semantic features from the input. For example, eight convolutional layers are used, with the number of filters set to 64, 64, 64, 64, 128, 128, 128, and 128, respectively. Max pooling layers are used after the second, fourth, and sixth convolutional layers.

[0235] The network can be regressed using a mesh motion regressor. Optionally, after feature extraction, an adaptive pooling layer is used to fix the resolution of the feature map. Subsequently, a fully convolutional structure is used as a mesh motion regressor to predict the horizontal and vertical motion of each vertex based on a regular mesh. Assuming a mesh resolution of U×V, the output volume is of size (U+1)×(V+1)×2.

[0236] The mesh motion can be estimated using a progressive regression of the residual. Alternatively, the exact mesh motion can be estimated using a progressive approach. First, the warped image is not used directly as input to the new network, as this would double the computational complexity. Instead, the intermediate feature map is warped. Then, two regressors with the same structure are used to predict the primary mesh motion and the residual mesh motion, respectively. Finally, the feature map can also be warped.

[0237] In this embodiment, the rectangularization process can be performed on the stitched image based on the portrait mask, that is, the portrait mask is used as a reference benchmark. In this way, the rectangularization process has a better effect and is more efficient.

[0238] In another possible implementation, the mesh regression network, mesh movement, and wrapping process may be performed once, or at least three times. This may be set as needed and is not limited here.

[0239] 4. Reference-based super-resolution (RefSR).

[0240] RefSR refers to the fact that, in addition to the low-resolution image, there is also a high-resolution reference image that has similar texture or content to the image to be restored. Utilizing the information in the Ref image can help the SR process.

[0241] In this embodiment, semantic guidance of portrait information can be added, super-resolution and texture migration can be separated, a separate branch can be used to extract high-frequency texture information related to the portrait, and the second image can be super-resolved. Then, the high-frequency texture information is migrated to the super-resolved second image, and finally the portrait super-resolution result (target image) can be obtained.

[0242] For example, a general SR network can be used to process the second image to obtain a single-frame super-resolution result. Then, feature extraction is performed on the LR image and the reference image respectively to obtain the reference image features of the reference image and the second image features of the second image. Then, feature alignment is performed guided by the portrait mask information to obtain a corresponding pixel-level offset map. The pixel-level offset map can indicate the offset information of each pixel in the second image in the reference image, such as the offset direction and offset amount of each pixel in the second image in the reference image.

[0243] Next, the reference image is encoded to obtain a reference feature map (also known as the encoded feature map). The pixel-level offset map is then applied to the encoded feature map through a variable convolution operation or an attention-based convolution operation to obtain the texture features to be transferred. The obtained texture features to be transferred and the single-frame super-resolution result (the second image after the single-frame super-resolution) are then input into the texture transfer module to obtain the final portrait super-resolution result.

[0244] In this embodiment, the offset direction and offset amount of each pixel of the second image in the reference image can be determined based on the pixel-level offset map. In this way, the variable convolution operation or the attention-based convolution operation on the encoded feature map takes into account the pixel offset between the reference image and the second image. In this way, the correlation between the obtained migrated texture features and the object area is also higher.

[0245] Optionally, the aforementioned image super-resolution process can be performed based on a pre-trained super-resolution model. In this embodiment, the pre-trained super-resolution model may include a feature extraction network for respectively extracting features from the reference image and the second image, a feature alignment network for aligning the reference image features and the second image features based on the object region mask, an encoder for encoding the reference image, a convolution network for performing a deformable convolution operation or an attention-based convolution operation on the reference feature map based on a pixel-level offset map, a texture migration network for migrating the texture features to be migrated to the second image, and a single-frame super-resolution network for performing single-frame super-resolution on the second image.

[0246] Optionally, the super-resolution model to be trained can be trained based on the training samples until the training of the super-resolution model to be trained is completed to obtain a pre-trained super-resolution model. In this embodiment, the super-resolution model to be trained may also include a decoder, which is used to decode texture features to obtain texture information, and the texture information can be used to supervise the training process of the super-resolution model to be trained. Specifically, the training samples may include training image samples, and the baseline texture information can be determined based on the training image samples. Then, the training loss can be calculated based on the texture information obtained by decoding and the baseline texture information, and then it is determined whether the training of the super-resolution model to be trained is completed. If the training is not completed, the model parameters of the super-resolution model to be trained are updated and the training is continued.

[0247] It should be noted that the super-resolution model to be trained may include modules or networks related to texture information extraction, while other modules unrelated to texture information extraction can directly use trained modules or networks. Optionally, in this embodiment, the modules related to texture information extraction may include, but are not limited to, at least one of a feature extraction network, a feature alignment network, an encoder, and a convolutional network.

[0248] In another possible implementation, the shape of the reference image may be processed into a shape other than a rectangle, such as a circle, etc., as long as it can correct the portrait area in the stitched image. There is no limitation on the specific shape of the reference image.

[0249] In another possible implementation, the spliced ​​image may not be rectangularized, that is, the spliced ​​image is used as a reference image to perform the super-resolution process.

[0250] It can be understood that by rectangularizing the stitched image to obtain a reference image that better matches the portrait mask, and by rectangularizing the image based on the stitched image and the portrait area mask, prioritizing minimizing the loss of the portrait effect after rectangularization, the super-resolution results are more accurate and efficient. In addition, if the reference image is a stitched image generated based on the telephoto image, and super-resolution is performed based on the stitched telephoto image, the texture of the portrait area in the telephoto image can be applied to the super-resolution process to improve the authenticity of the super-resolution result.

[0251] In another possible implementation, in addition to super-resolution of the portrait area, super-resolution can also be performed on areas of interest other than the portrait area, such as scenery. Furthermore, if the captured image is a non-portrait image, such as an animal or plant image, super-resolution can also be performed on areas of interest such as animals or plants, without limitation.

[0252] It should be noted that image enhancement can refer to the purposeful emphasis on overall or local characteristics of an image, sharpening previously unclear images or emphasizing certain interesting features. This can amplify the differences between features of different objects in the image while suppressing uninteresting features. This improves image quality, enriches information, enhances image interpretation and recognition, and meets the needs of certain specialized analyses. In this embodiment, in addition to super-resolution, certain interesting features can also be emphasized to enhance the image's prominent content and improve image display quality.

[0253] The following specific embodiments are used to describe the technical solutions of the present application in detail. The following specific embodiments can be implemented independently or in combination with each other. The same or similar concepts or processes may not be described in detail in some embodiments.

[0254] The image processing method provided in this embodiment may include:

[0255] A first interface is displayed, where the first interface includes a first preview image and a first button.

[0256] In response to a first operation on the first button, prompt information and a second button are displayed, where the prompt information is used to prompt the user to capture images of multiple target objects.

[0257] During the target time period, a plurality of first images are captured using the first mode and the plurality of first images are displayed respectively.

[0258] In this embodiment, how to trigger the first mode to capture the first image can be referred to the relevant description of FIG5 , which will not be elaborated here.

[0259] In response to a second operation on a second button, a second image is captured using a second mode, wherein the second button is continuously displayed within a target time period, and the target time period includes a time period from receiving the first operation to receiving the second operation. The clarity of the image captured using the first mode is higher than the clarity of the image captured using the second mode, and the field of view angle of the image captured using the first mode is lower than the field of view angle of the image captured using the second mode.

[0260] In this embodiment, how the first mode captures the first image and how the second mode captures the second image can be referred to the relevant description of FIG5 , which will not be elaborated here.

[0261] A target object in the second image is enhanced by using a stitched image to obtain a target image, wherein the stitched image includes an image obtained by stitching a plurality of first images.

[0262] In this embodiment, how to stitch multiple first images to obtain a stitched image and how to enhance the stitched image to obtain a target image can be described with reference to the relevant description of FIG. 5 , and is not limited here.

[0263] In another possible implementation, the image processing method further includes:

[0264] Displaying a second interface, the second interface including a second preview image, a third button, and a fourth button, the third button being in an unselected state, the fourth button being in a selected state, and the second preview image being a preview image captured in the mode of the fourth button;

[0265] Accordingly, the first interface is displayed, including:

[0266] In response to the third operation on the third button, a first interface is displayed, in which the third button is in a selected state and the fourth button is in an unselected state.

[0267] In another possible implementation, the image processing method further includes:

[0268] Displaying a third interface, the third interface includes a second preview image, a fourth button, and a fifth button, the fourth button is selected, the second preview image is a preview image captured in the mode of the fourth button, and the fifth button is unselected;

[0269] In response to a fourth operation on the fifth button, a fourth interface is displayed, where the fourth interface includes the third button, and the third button is in an unselected state;

[0270] The first interface is displayed, including:

[0271] In response to the third operation on the third button, a first interface is displayed, in which the third button is in a selected state and the fourth button is in an unselected state.

[0272] In another possible implementation, the first interface may be directly displayed when the camera application is opened, that is, the default mode when the camera application is opened is the group photo mode, which can be set as needed and is not limited here.

[0273] It should be noted that the module names involved in the embodiments of the present application can be defined as other names as long as the functions of each module can be achieved, and there is no specific restriction on the names of the modules.

[0274] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0275] The image processing method of the embodiment of the present application has been described above. The device for performing the above method provided by the embodiment of the present application is described below. Those skilled in the art will understand that the method and device can be combined and referenced with each other, and the relevant device provided by the embodiment of the present application can perform the steps in the above list sorting method.

[0276] The image processing method provided in the embodiment of the present application can be applied to electronic devices with communication functions. The electronic devices include terminal devices. The specific device form of the terminal device can refer to the above related descriptions and will not be repeated here.

[0277] An embodiment of the present application provides a terminal device, which includes: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the terminal device executes the above method.

[0278] like Figure 10 A schematic diagram of the structure of a chip provided in an embodiment of the present application: Chip 1000 includes one or more (including two) processors 1001 , a communication circuit 1002 , a communication interface 1003 , and a memory 1004 .

[0279] In some embodiments, the memory 1004 stores the following elements: executable modules or data structures, or a subset thereof, or an extended set thereof.

[0280] The method described in the above embodiment of the present application can be applied to the processor 1001, or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1001 or an instruction in the form of software. The above-mentioned processor 1001 can be a general-purpose processor (for example, a microprocessor or a conventional processor), a digital signal processor (digital signal processing, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates, transistor logic devices or discrete hardware components. The processor 1001 can implement or execute the methods, steps and logic block diagrams related to each processing disclosed in the embodiment of the present application.

[0281] The steps of the method described in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. Among them, the software module can be located in a mature storage medium in the field such as a random access memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable read only memory (EEPROM). The storage medium is located in the memory 1004, and the processor 1001 reads the information in the memory 1004 and completes the steps of the above method in combination with its hardware.

[0282] The processor 1001 , the memory 1004 , and the communication interface 1003 can communicate with each other via the communication line 1002 .

[0283] In the above embodiment, the instructions stored in the memory for execution by the processor may be implemented in the form of a computer program product, wherein the computer program product may be pre-written in the memory or downloaded and installed in the memory in the form of software.

[0284] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.

[0285] In one possible implementation, a computer-readable medium may include RAM, ROM, compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage or other magnetic storage devices, or any other medium intended to carry or store the desired program code in the form of instructions or data structures and accessible by a computer. Moreover, any connection is appropriately referred to as a computer-readable medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technology (such as infrared, radio and microwave) is used to transmit software from a website, server or other remote source, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology such as infrared, radio and microwave are included in the definition of medium. Disk and optical disc as used herein include optical disc, laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks generally reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0286] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes the above method.

[0287] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0288] The above specific implementation methods further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included in the scope of protection of the present invention.

Claims

1. An image processing method, characterized in that: Applied to electronic equipment, the method includes: Displaying a first interface, wherein the first interface includes a first preview image and a first button; In response to a first operation on the first button, displaying prompt information and a second button, wherein the prompt information is used to prompt the user to capture images of multiple target objects; Within a target time period, a plurality of first images are captured using a first mode, and the plurality of first images are displayed respectively; In response to a second operation on the second button, a second image is captured using a second mode, wherein the second image includes the entire target object, the second button is continuously displayed during a target time period, the target time period including a period from when the first operation is received to when the second operation is received, the clarity of the image captured using the first mode is higher than the clarity of the image captured using the second mode, and the field of view of the image captured using the first mode is lower than the field of view of the image captured using the second mode; Enhance the target object in the second image using a stitched image to obtain a target image, wherein the stitched image includes an image obtained by stitching the multiple first images; The step of enhancing the target object in the second image by using the stitched image to obtain the target image includes: Segmenting the second image to obtain an object region mask; correcting the spliced ​​image based on the object region mask to obtain a reference image, wherein the size of the object region in the reference image matches the size of the object region mask; wherein the object region is the region requiring image enhancement; The second image is super-resolved based on the reference image and the object region mask to obtain a target image.

2. The method according to claim 1, characterized in that The prompt information includes information for prompting the user to keep the posture of the electronic device unchanged, and the plurality of first images are captured by the electronic device through a movable camera; Alternatively, the prompt information includes information for guiding a movement trajectory, and the plurality of first images are obtained by photographing the electronic device while moving.

3. The method according to claim 1, characterized in that The method further comprises: Displaying a second interface, the second interface including a second preview image, a third button, and a fourth button, the third button being in an unselected state, the fourth button being in a selected state, and the second preview image being a preview image captured in the mode of the fourth button; The displaying of the first interface includes: In response to a third operation on the third button, the first interface is displayed, in which the third button is selected and the fourth button is unselected.

4. The method according to claim 1, wherein The method further comprises: Displaying a third interface, the third interface including a second preview image, a fourth button, and a fifth button, the fourth button being selected, the second preview image being a preview image captured in the mode of the fourth button, and the fifth button being unselected; In response to a fourth operation on the fifth button, displaying a fourth interface, the fourth interface including a third button, the third button being in an unselected state; The displaying of the first interface includes: In response to a third operation on the third button, the first interface is displayed, in which the third button is selected and the fourth button is unselected.

5. The method according to any one of claims 1 to 4, characterized in that The super-resolving the second image based on the reference image and the object region mask to obtain a target image includes: performing feature extraction on the reference image and the second image respectively to obtain reference image features of the reference image and second image features of the second image; aligning features of the reference image and features of the second image based on the object region mask to obtain a pixel-level offset map, wherein the pixel-level offset map indicates offset information of each pixel in the second image in the reference image; Encoding the reference image to obtain a reference feature map; Performing a deformable convolution operation or an attention-based convolution operation on the reference feature map based on the pixel-level offset map to obtain texture features to be migrated; The texture features to be migrated are migrated to the second image to obtain a target image.

6. The method according to claim 5, characterized in that The method further comprises: Performing single-frame super-resolution on the second image to obtain a second image after single-frame super-resolution; Migrating the texture features to be migrated to the second image to obtain a target image includes: The texture features to be migrated are migrated to the second image after single-frame super-resolution to obtain a target image.

7. The method according to claim 5, characterized in that The target image is obtained by super-resolving the reference image, the object region mask, and the second image using a pre-trained super-resolving model, wherein the pre-trained super-resolving model includes a feature extraction network, a feature alignment network, an encoder, a convolutional network, and a texture transfer network; The feature extraction network is used to perform feature extraction on the reference image and the second image respectively to obtain reference image features of the reference image and second image features of the second image; The feature alignment network is used to align the reference image features and the second image features based on the object region mask to obtain a pixel-level offset map; The encoder is used to encode the reference image to obtain a reference feature map; The convolutional network is used to perform a deformable convolution operation or an attention-based convolution operation on the reference feature map based on the pixel-level offset map to obtain texture features to be migrated; The texture migration network is used to migrate the texture features to be migrated to the second image to obtain a target image.

8. The method according to claim 7, characterized in that The target image is obtained by migrating the texture features to be migrated to a second image after single-frame super-resolution, and the pre-trained super-resolution model also includes a single-frame super-resolution network; The single-frame super-resolution network is used to perform single-frame super-resolution on the second image to obtain a second image after single-frame super-resolution; The texture migration network is used to migrate the texture features to be migrated to the second image after single-frame super-resolution to obtain a target image.

9. The method according to any one of claims 1, 6-8, characterized in that Correcting the stitched image based on the object area mask to obtain a reference image includes: splicing the spliced ​​image with the object area mask, and performing feature extraction on the spliced ​​image with the object area mask to obtain a spliced ​​feature map; Performing N correction processes based on the spliced ​​feature map to obtain a reference image, where N is an integer not less than 2; Each calibration process includes: Perform network regression on the input, as well as mesh motion and bending processing, to obtain the output; Among them, the input of the first correction process is the spliced ​​image with the object area mask spliced, the input of the Mth correction process is the output obtained by the M-1th correction process, and the output obtained by the Nth correction process is used as the reference image, 1<M≤N, and M is an integer.

10. An image processing method, characterized in that: include: Acquire a second image and a plurality of first images, wherein the plurality of first images are captured using the first mode, and the second images are captured using the second mode, wherein the clarity of the images captured using the first mode is higher than the clarity of the images captured using the second mode, and the field of view of the images captured using the first mode is lower than the field of view of the images captured using the second mode; splicing the plurality of first images to obtain a spliced ​​image; enhancing the target object in the second image using the stitched image to obtain a target image; The step of enhancing the target object in the second image by using the stitched image to obtain the target image includes: Segmenting the second image to obtain an object region mask; correcting the spliced ​​image based on the object region mask to obtain a reference image, wherein the size of the object region in the reference image matches the size of the object region mask; wherein the object region is the region requiring image enhancement; The second image is super-resolved based on the reference image and the object region mask to obtain a target image.

11. An electronic device, characterized in that: The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method as described in any one of claims 1 to 10.

12. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises computer instructions, and when the computer instructions are executed on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The computer program product comprises a computer program code, and when the computer program code is run on an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Photographing method and device

    CN113810598A

  • Panoramic depth image generation method and device

    CN115022526A

  • Photographing method and device

    WO2022022715A1