Image generation method and device, electronic equipment and computer program product

By identifying and fusing target objects, especially small-scale faces, and using target weights to process images, the problem of oversharpening of face areas is solved and the transition effect of the image is improved.

CN120355618APending Publication Date: 2025-07-22GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446599.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In image processing of existing electronic devices, the face area is prone to oversharpening, resulting in poor image processing effect, especially when the transition between the face and the surrounding area is not smooth enough on a small scale.

Method used

By identifying the target object, especially a small-scale face, the target weight is used to fuse the input image and the output image to generate the target image, so that the clarity of the target area is between the input and the output image, avoiding transition sharpening.

Benefits of technology

The transition effect between the target area and the surrounding area in the image generated by the electronic device is improved, the abruptness within and outside the target area is reduced, and the natural transition effect of the image is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355618A_ABST
    Figure CN120355618A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image generation method and device, electronic equipment and a computer program product. Relates to the technical field of image processing. The method is applied to electronic equipment and comprises the steps that a first input image is processed to obtain a first output image, and the definition of the first output image is higher than that of the first input image; under the condition that the human face exists in the first input image, the first input image and the first output image are fused based on the target weight, a target image is obtained, and the definition of a target area in the target image is larger than that of the first input image and smaller than that of the first output image; the target area is an image area where the target object is located, the target weight comprises an input weight and an output weight, the input weight is the weight of the first input image, and the output weight is the weight of the first output image. According to the scheme, the transition effect of the target area and the surrounding area in the image generated by the electronic equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of image processing technology, and particularly to an image generation method, apparatus, electronic device, and computer program product. Background Art

[0002] With the development of science and technology, various electronic devices have emerged in people's daily lives, and people can use electronic devices for entertainment, learning, etc.

[0003] Currently, in various electronic devices, most of them have image processing functions. For example, a user edits a certain image stored in an electronic device to generate an image they want. Among them, adjusting the clarity of an image is a common processing function. Most electronic devices perform a one-time processing on the image input by the user based on a machine learning model pre-installed inside themselves to generate a clearer image. Although this solution can generate an image with one key, the face area in the generated image is prone to oversharpening, and there are problems such as the transition between the face area and other areas in the image not being smooth enough and the image processing effect being poor. Summary of the Invention

[0004] In order to solve the problems of related technologies, avoid the oversharpening of the target object area, and improve the transition effect between the target object area and the surrounding area in the image generated by the electronic device. Embodiments of the present application provide an image generation method, apparatus, electronic device, and computer program product. The technical solutions are as follows:

[0005] On the one hand, an embodiment of the present application provides an image generation method, which is applied to an electronic device. The method includes:

[0006] Process a first input image to obtain a first output image, and the clarity of the first output image is higher than that of the first input image;

[0007] In the case where there is a target object in the first input image, perform a fusion process on the first input image and the first output image based on a target weight to obtain a target image. The clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image. The target area is the image area where the target object is located. The target weight includes an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image.

[0008] On the other hand, an embodiment of the present application provides an image generation apparatus, which is applied to an electronic device. The apparatus includes:

[0009] A first acquisition module, configured to process a first input image to obtain a first output image, where the clarity of the first output image is higher than that of the first input image;

[0010] A second acquisition module, configured to, when there is a target object in the first input image, perform fusion processing on the first input image and the first output image based on a target weight to obtain a target image, where the clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image, the target area is the image area where the target object is located, the target weight includes an input weight and an output weight, the input weight is the weight of the first input image, and the output weight is the weight of the first output image.

[0011] On the other hand, the present application provides an electronic device, where the electronic device includes a processor and a memory, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements the image generation method described in the above aspect.

[0012] On the other hand, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image generation method described in the above aspect.

[0013] On the other hand, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the image generation method described in the above aspect.

[0014] On the other hand, an embodiment of the present application provides an application publishing platform, where the application publishing platform is used to publish a computer program product, and when the computer program product runs on a computer, it causes the computer to execute to implement the image generation method described in the above aspect.

[0015] The beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0016] In an electronic device, the first input image is processed to obtain a first output image, and the clarity of the first output image is higher than that of the first input image. When there is a target object in the first input image, the first input image and the first output image are fused based on the target weights to obtain a target image. The clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image. The target area is the image area where the target object is located. The target weights include an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image will be fused based on the target weights, so that the clarity of the target area in the finally generated target image is greater than that of the first input image and less than that of the first output image, thereby reducing the abruptness between the target area and the area outside the target area, avoiding the over-sharpening of the target object, and improving the transition effect between the target area and the surrounding area in the image generated by the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0018] Figure 1 Schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application;

[0019] Figure 2 Method flowchart of an image generation method provided by an exemplary embodiment of the present application;

[0020] Figure 3 Method flowchart of an image generation method provided by an exemplary embodiment of the present application;

[0021] Figure 4 Schematic diagram of an image of a face recognized by a first recognition model according to an exemplary embodiment of the present application;

[0022] Figure 5 Method flowchart of an image generation method provided by an exemplary embodiment of the present application;

[0023] Figure 6 Schematic diagram of an image of a face recognized by a second recognition model according to an exemplary embodiment of the present application;

[0024] Figure 7The flowchart of a method for generating an image provided by an exemplary embodiment of the present application;

[0025] Figure 8 The structural block diagram of an image generation device provided by an exemplary embodiment of the present application;

[0026] Figure 9 The structural schematic diagram of another example of the image generation device provided by an embodiment of the present application. Detailed implementation manners

[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0028] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0029] It should be noted that the terms "first", "second", and "third" involved in the embodiments of the present application are used to distinguish similar or different objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0030] The solution provided by the present application can be used in the real scenario where people use electronic devices to optimize images in their daily lives. For the sake of easy understanding, the application scenarios involved in the embodiments of the present application will be briefly introduced below.

[0031] With the development of science and technology, various electronic devices have emerged in people's daily lives, and people can use electronic devices for entertainment, learning, etc. Among them, the image processing function has been applied in most electronic devices, and users can start the corresponding application program in the electronic device to enable the electronic device to process images.

[0032] Please refer to Figure 1 , which shows the structural schematic diagram of an electronic device provided by an exemplary embodiment of the present application. As Figure 1The electronic device shown includes components such as a processor 110, a memory 120, a transceiver 130, a display unit 140, an input unit 150, a sensor 160, an audio circuit 170, and a power module 180.

[0033] The processor 110 is the control center of the electronic device. It uses various interfaces and lines to connect all parts of the electronic device. By running or executing software programs and / or modules stored in the memory 120, and by calling data stored in the memory 120, it executes various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 110 may include one or more processing units. Optionally, the processor 110 may integrate an application processor, which mainly processes the operating system, user interface, and application programs, etc. Of course, it may also include other processors, which are not listed one by one here.

[0034] The memory 120 can be used to store software programs and modules. The processor 110 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 120. The memory 120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store data created according to the use of the electronic device (such as audio data, phone book, etc.). In addition, the memory 120 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices.

[0035] The transceiver 130 can provide wireless communication solutions applied to the electronic device, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The transceiver 130 can be one or more devices integrating at least one communication processing module. For example, integrating an antenna with a baseband processor as the transceiver 130, or integrating an antenna and a modulation / demodulation processor as the transceiver 130, etc., which is not limited here.

[0036] The display unit 140 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device. The display unit 140 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc., without limitation here.

[0037] The input unit 150 can be used to receive input numerical or character information, and generate key signal inputs related to the user settings and function controls of the electronic device. Specifically, the input unit 150 can collect operations of the user on or near it, and drive corresponding connection devices according to a preset program. In addition, the input unit 150 may include a touch panel, which can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel, the input unit 150 may further include other input devices. Specifically, the other input devices may include, but are not limited to, one or more of function keys (such as volume control buttons, switch buttons, etc.), trackballs, joysticks, etc.

[0038] The electronic device may further include at least one sensor 160, such as a gyroscope sensor, a motion sensor, and other sensors. The motion sensor may include an acceleration sensor, which is used to detect the magnitude of acceleration in all directions. When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the electronic device, such as horizontal and vertical screen switching, related games, magnetometer posture calibration, etc.; as for other sensors such as a pressure gauge, a barometer, a hygrometer, a thermometer, an infrared sensor, etc. that the electronic device may also be configured with, they will not be elaborated here.

[0039] The audio circuit 170 may include a speaker and a microphone, and can provide an audio interface between the user and the electronic device. The audio circuit 170 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 170 and then converted into audio data. After the audio data is output to the processor 110 for processing, it is sent to another electronic device, for example, through the video circuit, or the audio data is output to the memory 120 for further processing.

[0040] The electronic device further includes a power module 180 that supplies power to each component. Optionally, the power module 180 can be logically connected to the processor 110 through a power management device, so as to implement functions such as management of charging, discharging, and power consumption management through the power management device.

[0041] Although not shown, the electronic device may further include a camera. Optionally, the camera may be located at the front or rear of the electronic device, which is not limited in the present embodiment.

[0042] It is to be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0043] Optionally, the above-mentioned electronic devices may include but are not limited to wearable devices (such as smart bracelets, smart watches, smart glasses, etc.), mobile phones, tablet computers, laptops, smart glasses, smart watches, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Compression Standard Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) players, desktop computers, laptops, etc.

[0044] Usually, the above Figure 1 The electronic device shown needs to be equipped with a corresponding operating system and run based on the operating system. For example, the operating system of the electronic device can be an Android system, an iOS system, a Linux system, etc.

[0045] Optional, for the above Figure 1 In the electronic device shown, if a user needs to perform image enhancement processing on an image in the electronic device to improve the clarity of the image, the user can input the image into a pre-set image processing model in the electronic device, and use the image processing model to add texture details, noise particles, etc. to the input image to improve clarity.

[0046] In image processing models, a global unified enhancement strategy is usually adopted to enhance all image elements in the input image. This may easily lead to over-sharpening when processing certain specific image elements. For example, for smaller faces in the image, this image processing model may easily introduce excessive noise artifacts in the face area, resulting in unnatural granular texture, causing the transition between the smaller faces and other surrounding areas to be not smooth enough, resulting in poor image processing effects.

[0047] To solve the above problems existing in the related art, avoid the oversharpening of the face region, and improve the transition effect between the face region and the surrounding regions in the image generated by the electronic device, an image generation method is provided in an embodiment of the present application. By identifying the target object, when the input image contains the target object, the input image and the output image are fused according to the target weight, so that the target region in the obtained target image will not be oversharpened.

[0048] Please refer to Figure 2 , which shows a flowchart of an image generation method provided by an exemplary embodiment of the present application. This image generation method can be applied to an electronic device. As Figure 2 shown, this image generation method may include the following steps:

[0049] Step 201, process the first input image to obtain a first output image, and the clarity of the first output image is higher than that of the first input image.

[0050] Optionally, in the present application, the electronic device has the function of processing the first input image to obtain a first output image with higher clarity. For example, an image processing model is configured in the electronic device, and this image processing model can be a generative adversarial network (GAN) model. When any image is used as the first input image, the electronic device can process the first input image through this GAN model to obtain a first output image with higher clarity.

[0051] Step 202, when there is a target object in the first input image, fuse the first input image and the first output image based on the target weight to obtain a target image. The clarity of the target region in the target image is greater than that of the first input image and less than that of the first output image. The target region is the image region where the target object is located. The target weight includes an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image.

[0052] Optionally, during the process of obtaining the first output image, the electronic device can execute the step of determining whether the first input image contains a target object. When it is found that the first input image contains a target object, the first input image and the first output image are fused based on the target weight. Among them, the target weight can be pre-set in the electronic device by the developer.

[0053] For example, in the target weights, the input weight is weight one and the output weight is weight two. The process by which the electronic device performs fusion processing on the first input image and the first output image based on the target weights to obtain the target image can be as follows: Multiply the first input image by weight one + multiply the first output image by weight two, and the resulting fused image is the target image.

[0054] Optionally, the target object can be any object that the developer wants to recognize in advance. For example, the target object is a face, an animal, or a limb. Among them, these objects will be relatively small in the image when the imaging distance is far, and it is easy to have the situation of oversharpening during the processing of the above image processing model. Therefore, in this solution, by combining the recognition of whether there is a target object in the first input image, and when there is a target object in the first input image, the first input image and the first output image of the image processing model are fused based on the target weights, so that the clarity of the target area is greater than the clarity of the first input image and less than the clarity of the first output image, and the transition between the target area and the area outside the target area is smoother.

[0055] In summary, in the electronic device, the first input image is processed to obtain the first output image, and the clarity of the first output image is higher than that of the first input image; when there is a target object in the first input image, the first input image and the first output image are fused based on the target weights to obtain the target image. The clarity of the target area in the target image is greater than the clarity of the first input image and less than the clarity of the first output image. The target area is the image area where the target object is located. The target weights include the input weight and the output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image will be fused based on the target weights, so that the clarity of the target area in the finally generated target image is greater than the clarity of the first input image and less than the clarity of the first output image, thereby reducing the abruptness between the target area and the area outside the target area, avoiding oversharpening of the target object, and improving the transition effect between the target area and the surrounding area in the image generated by the electronic device.

[0056] In a possible implementation manner, a first recognition model is configured in the electronic device. The first recognition model is used to recognize target objects greater than or equal to the first size. The electronic device uses the first recognition model to recognize the first input image, so as to determine whether there is a target object in the first input image.

[0057] Please refer to Figure 3 , which shows a flowchart of a method for generating an image provided by an exemplary embodiment of the present application. This method for generating an image can be applied to an electronic device. AsFigure 3 As shown in Figure 3 , the image generation method may include the following steps:

[0058] Step 301: Process the first input image to obtain a first output image, where the clarity of the first output image is higher than that of the first input image.

[0059] Optionally, the process of processing the first input image in the electronic device may be implemented based on an image processing model, and the image processing model may be a generative adversarial network (GAN) model as exemplified above.

[0060] In a possible implementation manner, the image processing model may be pre-trained by a developer and set in the electronic device. For example, the image processing model is trained based on second target sample data. By training, after an image is input into the image processing model, the image processing model outputs an image with higher clarity. That is, the clarity of the output image of the image processing model is higher than that of the input image of the image processing model. The second target sample data includes original images of various target objects and reference images generated according to the original images. In the reference image, the clarity of the area where the target object is located is lower than the clarity of the area outside the area where the target object is located in the reference image, and the clarity of the area outside the area where the target object is located in the reference image is higher than the clarity of the area outside the area where the target object is located in the original image.

[0061] Optionally, taking the target object as a human face as an example, the above electronic device may obtain scene images including the models and corresponding focal lengths using the GAN model, taken at a distance of 20 - 200 meters in a natural state and containing pedestrians (sparse pedestrians or dense crowds), and images containing various human face situations (near / far, front / side, with expressions or movements, wearing hats / partially blocked, etc.) taken in indoor / outdoor, day / night, front light / back light and other environments. These images of various human faces may also cover different human faces corresponding to various ethnic groups. For these images containing various human faces, the electronic device crops out all small-scale human face areas (with a size of about 200 pixels or less), fills them in a tiled manner into an image patch of a unified size, and performs gamma correction with coefficients of 0.5 and 1.5 respectively to simulate the brightness of the original image under different illuminations, as the original image of the low-resolution image (LR). A copy is made for each LR, and slight bilateral filtering denoising and sharpening processing are performed (additional upsampling processing may also be added in the image super-resolution task), so as to obtain a reference image (GroundTruth, GT). The clarity of the reference image is higher than that of the original image. Data pairs are constructed according to each LR and the corresponding GT, and these data pairs are the second target sample data.

[0062] That is to say, for various face-containing images that can be obtained on the network or by an electronic device, the electronic device can crop all small-scale face regions (with a size of about 200 pixels or less) from various face-containing images, and perform a series of processes to obtain LR. Based on LR, GT is further generated, and LR-GT is used as the training data for the image processing model. That is, LR is the input of the image processing model, and GT is the output of the image processing model. Then, the image processing model is trained, and finally, an output image with higher clarity can be generated, but there is no excessive increase in details in the face part of the output image. That is, for the trained image processing model, if an image is input, the face part in the output image is similar to the GT content in the training process, enabling the image processing model to master the semantic features of small-scale faces and learn not to generate excessive grains and details in the face region.

[0063] It should be noted that if the electronic device does not have the conditions to naturally capture various images during the above training process, it can also capture the public datasets (such as CelebA-HQ, FFHQ, WIDER FACE, etc.) in a dark room to obtain the required face-containing images. The present application does not limit the method of obtaining various face-containing images.

[0064] Step 302, identify the first input image through the first recognition model to obtain the first recognition result.

[0065] Optionally, in the process of obtaining the first output image, a step of determining whether the first input image contains a target object can be performed. For example, the electronic device is further configured with a first recognition model. While inputting the above first input image into the image processing model for processing, the first input image can also be input into the first recognition model to identify the first input image.

[0066] Among them, the first recognition model is used to identify target objects with a size greater than or equal to a first size, and the first size is greater than a first threshold. For example, the first recognition model can be a recognition model in an Image Signal Processor (ISP). Currently, the recognition ability of the recognition model in the ISP is limited, and it can usually only identify target objects with a size greater than or equal to the first size in the image, and cannot identify target objects with a size smaller than the first size in the image. In this solution, based on this recognition model, the first input image is identified to obtain the corresponding first recognition result. Optionally, the first recognition result includes the number of target objects contained in the first input image identified by the first recognition model and the coordinate information of the target objects, thereby indicating whether there is a target object in the input image.

[0067] Taking the target object as a human face as an example, the training of the recognition model in the ISP can be obtained by training based on common training methods. After the recognition model in the ISP performs face recognition on the first input image, the output metadata can be regarded as the corresponding first recognition result. The first recognition result may include the number of human faces in the first input image and the corresponding face coordinates.

[0068] Step 303, when the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, determine that there is a target object in the first input image.

[0069] Optionally, the electronic device can detect the number of target objects included in the first input image in the above first recognition result, compare the number of the target objects with a preset number threshold to know the indication situation of the first recognition result, and then determine whether there is a target object in the first input image.

[0070] For example, still taking the target object as a human face as an example, the electronic device can detect the number of human faces in the first input image in the first recognition result. If the number of human faces in the first input image is greater than 0, it means that the first recognition result indicates that the first input image contains a human face greater than or equal to the first size. In this case, it is determined that there is a target object (human face) in the first input image. On the contrary, if the number of human faces in the first input image is not greater than 0, it means that the first recognition result indicates that the first input image does not contain a human face greater than or equal to the first size.

[0071] In a possible implementation manner, in order to improve the accuracy of the target object recognized by the first recognition model, when the electronic device executes the step of determining that there is a target object in the first input image when the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, it can be as follows: when the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, detect whether the target object greater than or equal to the first size included in the first input image is valid; when at least one target object greater than or equal to the first size is valid, determine that there is a target object in the first input image.

[0072] That is, the electronic device can also detect the validity of the target object recognized by the first recognition model when the first recognition result indicates that the first input image contains a target object greater than or equal to the first size. When at least one of the target objects greater than or equal to the first size is valid, it can be determined that the target object exists in the first input image. For example, for the first input image, after recognition by the above-mentioned first recognition model, if the first recognition result indicates that there are three target objects, the electronic device can detect the validity of the three target objects. When one of the target objects is valid, it can be determined that the target object exists in the first input image. Optionally, the detection of each target object can be performed once or sequentially.

[0073] In one possible implementation, an electronic device detects whether a target object greater than or equal to a first size contained in a first input image is valid as follows: for the first target object, detect whether the first target object is within a preset field of view (FOV); the first target object is any one of target objects greater than or equal to the first size; if the first target object is within the preset FOV, determine that the first target object is valid; if the first target object is not within the preset FOV, determine that the first target object is invalid.

[0074] The preset FOV is determined based on the medium and long focal lengths identified by the first recognition model. For example, if the medium and long focal length of the first recognition model is 15, and its corresponding FOV is FOV 1, then the preset FOV is FOV 1, and the validity of each target object is detected by detecting whether each target object is within the preset FOV.

[0075] Optionally, the electronic device may detect whether the first target object is within a preset field of view FOV as follows: calculate the overlapping area between the first target object and the preset FOV based on the pixel coordinates of the first target object and the preset FOV; when the overlapping area is greater than or equal to a first preset area, determine that the first target object is within the preset FOV, and the first preset area is determined based on the area of the first target object itself; when the overlapping area is less than the first preset area, determine that the first target object is not within the preset FOV.

[0076] For example, for the first target object, the electronic device can obtain the coordinate information (i.e., pixel coordinates) of the first target object through the above first recognition result, calculate the self-area of the first target object according to the pixel coordinates of the first target object, and use half of the self-area of the first target object as the first preset area. After calculating the overlapping area between the first target object and the preset FOV based on the pixel coordinates of the first target object and the preset FOV, if the overlapping area is greater than or equal to half of the self-area of the first target object, it indicates that the first target object is within the preset FOV; if the overlapping area is less than half of the self-area of the first target object, it indicates that the first target object is not within the preset FOV.

[0077] For example, in the case where the target object is a human face, the above first recognition result will include the pixel coordinates (including the upper left corner coordinates and the lower right corner coordinates) of each recognized human face. The electronic device can obtain the area of the first human face based on the pixel coordinates of the first human face. When detecting whether the first human face is valid, calculate the overlapping area between the two based on the pixel coordinates of the first human face and the crop_region parameter of the FOV used by the first recognition model. If the overlapping area is greater than or equal to half of the area of the first human face itself, it is considered that the first human face is within the FOV, and the first human face recognized by the first recognition model based on the ISP is valid. If the overlapping area is less than half of the area of the first human face itself, it is considered that the first human face is not within the FOV, and the first human face recognized by the first recognition model based on the ISP is invalid.

[0078] Still taking the target object being a human face as an example, please refer to Figure 4 , which shows a schematic diagram of an image of a human face recognized by a first recognition model according to an exemplary embodiment of the present application. As Figure 4 shown, in the first input image 401, through the recognition of the first recognition model, each human face 402 included in the first recognition result, through the detection of the FOV here, as long as there is one human face among the human faces 402 that is within the FOV, it can be determined that there is a target object in the first input image. If none of the human faces 402 are within the FOV, it can be determined that there is no target object in the first input image.

[0079] Optionally, during the process of detecting the validity of the target object recognized by the first recognition model, if the detection is performed in sequence, as long as any one of the target objects is detected to be valid, the detection process can be ended, and it can be determined that there is a target object in the first input image.

[0080] Step 304, when there is a target object in the first input image, the first input image and the first output image are fused based on the target weight to obtain a target image. The clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image. The target area is the image area where the target object is located. The target weight includes an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image.

[0081] Optionally, the target weight is preset by the developer. Among them, the input weight is greater than the output weight, and the sum of the input weight and the output weight is one. For example, the input weight is 0.8 and the output weight is (1 - 0.8) = 0.2.

[0082] In this embodiment, when it is recognized by the first recognition model that there is a target object in the first input image and it is determined that there is a target object in the first input image, the first input image and the first output image can be fused based on the preset target weight to obtain a target image. Let input represent the first input image and GANout represent the first output image. When fusing here, it can be done in the following way: 0.8 * input + 0.2 * GANout. That is, the pixel points of the target image are the same as those of the first input image and the first output image. The pixel value of each pixel point in the target image is equivalent to multiplying the corresponding pixel point in the first input image by 0.8 and adding the corresponding pixel point in the first output image multiplied by 0.2, thereby generating the target image.

[0083] In summary, in the electronic device, the first input image is processed to obtain the first output image, and the clarity of the first output image is higher than that of the first input image; when there is a target object in the first input image, the first input image and the first output image are fused based on the target weight to obtain a target image. The clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image. The target area is the image area where the target object is located. The target weight includes an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image will be fused based on the target weight, so that the clarity of the target area in the finally generated target image is greater than that of the first input image and less than that of the first output image, thereby reducing the abruptness between the target area and the area outside the target area, avoiding the over-sharpening of the target object, and improving the transition effect between the target area and the surrounding area in the image generated by the electronic device.

[0084] In a possible implementation manner, a second recognition model is configured in the electronic device. The second recognition model is used to recognize target objects smaller than or equal to a second size, and the second size is smaller than a first threshold. The second recognition model in the electronic device is used to recognize a first input image, so as to determine whether there is a target object in the first input image.

[0085] Please refer to Figure 5 , which shows a flowchart of a method for generating an image provided by an exemplary embodiment of the present application. The method for generating an image can be applied to an electronic device. As Figure 5 shown, the method for generating an image may include the following steps:

[0086] Step 501, process the first input image to obtain a first output image, and the clarity of the first output image is higher than that of the first input image.

[0087] Optionally, the implementation details of the electronic device executing step 501 may refer to the relevant description in step 301 above, and will not be elaborated here.

[0088] Step 502, recognize the first input image through the second recognition model to obtain a second recognition result.

[0089] Optionally, in this embodiment, the electronic device is configured with a second recognition model. The second recognition model is used to recognize target objects smaller than or equal to a second size, and the second size is smaller than a first threshold. In the process of obtaining the first output image, the step of determining whether the first input image contains a target object can be executed. For example, while inputting the first input image into an image processing model for processing, the first input image can also be input into the second recognition model to recognize the first input image.

[0090] Optionally, the second recognition model is trained according to first target sample data. The first target sample data includes synthetic images obtained by fusing various background images with images containing target objects using a preset fusion algorithm. The preset fusion algorithm is used to fuse the target objects in the images containing target objects with various background images after processing the target objects according to a preset processing logic. The preset processing logic includes one or more of scaling according to a preset scaling ratio, rotating according to a preset rotation angle, and adjusting illumination according to preset illumination parameters.

[0091] In this solution, the first target sample data is obtained by fusing various images using a dynamic synthesis strategy, which not only ensures the sample data volume but also includes the position information of various target objects, thereby improving the accuracy of the second recognition model obtained through training. Optionally, the second recognition model can be trained based on the RFBNet (Receptive Field Block Network) model. This network model can introduce a multi-scale receptive field module on the basis of the SSD (Single Shot MultiBox Detector) basic framework, and construct a spatial attention mechanism by paralleling convolutional layers with different dilation rates to achieve enhanced perception of small-scale face features in the distance.

[0092] Optionally, the preset fusion algorithm can be the Poisson fusion algorithm. Taking the target object as a face as an example, the electronic device can obtain each image containing a face in publicly available population datasets such as WIDER FACE and CelebA, and also obtain each image containing a face from face datasets such as FDDB and FFHQ. The face region in the real-shot face dataset is accurately segmented, and using the Poisson fusion algorithm, it is adaptively embedded into high-resolution background images such as DIV8K and HQ50K. During the process, the face scaling ratio (0.1 times - 0.5 times), rotation angle (±30°), and illumination matching parameters (i.e., preset illumination parameters) are randomly adjusted to effectively simulate the multi-level spatial relationship between the face and the background caused by the depth of field change in mobile phone real-shot. The obtained images containing faces, as well as the face features of each face obtained by annotating the images, are used as the first target sample data for training the second recognition model, so that the trained second recognition model can output the face features of each face contained in any input image.

[0093] During the training process, a dual-loss joint optimization strategy is adopted: the bounding box regression loss uses an improved mIoU (Intersection over Union metric) function; the classification loss uses a cross-entropy function based on Softmax. After forward propagation, the samples are sorted in descending order of the loss value, and only the samples with the top 30% of the loss values are used for backpropagation gradient calculation, enabling the network to focus on learning the features of challenging samples such as blurred faces and occluded faces that are prone to confusion. After the training is completed, the second recognition model can recognize target objects with a smaller size that are difficult to recognize by the above-mentioned first recognition model.

[0094] For the trained second recognition model, after the electronic device inputs the first input image into the second recognition model, the second recognition model can recognize the first input image and obtain a second recognition result. Optionally, the second recognition result includes the object features of the target objects recognized as being less than or equal to the second size.

[0095] In a possible implementation manner, the first recognition model and the second recognition model are both configured in the electronic device. The electronic device first recognizes the first input image through the first recognition model to obtain a first recognition result. When the first recognition result indicates that the first input image does not contain a target object with a size greater than or equal to the first size, the electronic device then recognizes the first input image through the second recognition model to obtain a second recognition result. The first recognition result is the recognition result obtained by the first recognition model recognizing the first input image. The implementation manner of the first recognition model may refer to the description in the above Figure 3 embodiment and will not be elaborated here.

[0096] Step 503: When the second recognition result indicates that the first input image contains a target object with a size less than or equal to the second size, it is determined that there is a target object in the first input image.

[0097] Optionally, if the second recognition result contains the object features of the recognized target object with a size less than or equal to the second size, it indicates that the second recognition result indicates that the first input image contains a target object with a size less than or equal to the second size. If the second recognition result does not contain the object features of the recognized target object with a size less than or equal to the second size, it indicates that the second recognition result indicates that the first input image does not contain a target object with a size less than or equal to the second size.

[0098] Optionally, similar to the first recognition result of the first recognition model above, the second recognition result may also include the quantity of each target object. Here, the electronic device can also detect the quantity of the target objects included in the first input image in the second recognition result, and compare the quantity of the target objects with a preset quantity threshold to know the indication situation of the second recognition result, and then determine whether there is a target object in the first input image. The comparison process between the quantity of the target objects and the preset quantity threshold may refer to the description in the above step 303 and will not be elaborated here.

[0099] In a possible implementation manner, in order to improve the accuracy of the target object recognized by the second recognition model, when the electronic device executes the step of determining that there is a target object in the first input image when the second recognition result indicates that the first input image contains a target object with a size less than or equal to the second size, it may be as follows: when the second recognition result indicates that the first input image contains a target object with a size less than or equal to the second size, according to the object features of each target object with a size less than or equal to the second size and the classification label weights corresponding to the object features of each target object, obtain the confidence index of each target object; when there is a confidence index greater than the confidence threshold among the confidence indexes of each target object, it is determined that there is a target object in the first input image.

[0100] Among them, the classification label weight can be an adaptive weight based on Support Vector Machine (SVM). In this solution, an adaptive weight decision model based on SVM can be added. Traditional face detectors usually use a fixed confidence threshold (such as 0.5) for binary classification decision-making. This threshold strategy faces significant challenges in the real-shot mobile phone scenarios: due to differences in shooting conditions (sudden changes in lighting, motion blur, large dynamic range of face scales), the variance of feature distribution is large, and a single threshold is difficult to balance the contradiction between low-quality small faces (the threshold needs to be reduced to improve recall rate) and excluding complex background interference (the threshold needs to be increased to ensure precision rate). To avoid using a single threshold, this solution can train an adaptive weight decision model based on SVM, so that the classification label weights (i.e., adaptive weights) corresponding to the object features of each target object can be obtained according to the object features of each target object in the second recognition result of the second recognition model.

[0101] Among them, training the adaptive weight decision model based on SVM can be as follows. The second recognition model is used to recognize a large number of images, and the object features of each target object in these images are obtained, and the classification label weights corresponding to the object features of each target object are calibrated. Then, the adaptive weight decision model is trained so that the adaptive weight decision model can obtain the corresponding classification label weights for the object features of different target objects, that is, the adaptive weight decision model can obtain the classification label weights corresponding to the object features of each target object based on the object features of each target object in the second recognition result of the second recognition model. This makes the weights of the object features of the target object not fixed, and can more accurately determine whether the target object recognized by the second recognition model is a target object less than or equal to the second size.

[0102] For example, in the traditional solution, for a certain eigenvalue one of a target object, if eigenvalue one is less than the above-set single threshold, it is considered that the target object is a target object less than or equal to the second size; if eigenvalue one is greater than the above-set single threshold, it is considered that the target object is not a target object less than or equal to the second size. However, through the adaptive weight decision model of this solution, after the object features of the target object are recognized by the second recognition model, the corresponding classification label weights can be obtained for the object features of the target object. According to the object features of each target object less than or equal to the second size and the classification label weights corresponding to the object features of each target object, the confidence index of each target object is obtained, and it is determined whether each target object is a target object less than or equal to the second size based on the confidence index of each target object.

[0103] Optionally, the classification method of the SVM-based adaptive weight decision model is exemplified as follows: if w1 * eigenvalue 1 + w2 * eigenvalue 2 + w3 * eigenvalue 3 + constant b > 0, it is classified as 1, and if w1 * eigenvalue 1 + w2 * eigenvalue 2 + w3 * eigenvalue 3 + constant b > 0, it is classified as 0. Here, w1, w2, and w3 are the classification label weights respectively, eigenvalue 1, eigenvalue 2, and eigenvalue 3 are the object features of a target object, the constant b can be obtained through the above training process, 1 represents the classification of a target object less than or equal to the second size, and 0 represents the classification of a target object not less than the second size.

[0104] Of course, here three eigenvalues and the corresponding three classification label weights are used for exemplification. In practical applications, a target object can also contain more eigenvalues. The method for the electronic device to determine whether a target object is a target object less than or equal to the second size is similar, and will not be elaborated here.

[0105] Optionally, after calculating the confidence indicators of each target object, when there is a confidence indicator greater than the confidence threshold among the confidence indicators of each target object, the electronic device determines that there is a target object in the first input image. That is, the electronic device needs to screen the confidence indicators of each target object, and when there is a confidence indicator greater than the confidence threshold among the confidence indicators of each target object, it determines that there is a target object in the first input image.

[0106] Step 504, when there is a target object in the first input image, a target mask image is generated according to the first input image and each first object. The target mask image is the same size as the first input image, and the positions of each first object in the target mask image are the same as the positions of each first object in the first input image. The pixel values within the pixel regions of each first object in the target mask image are the first weight, and the pixel values outside the pixel regions of each first object are the second weight. Each first object is each target object corresponding to each confidence indicator greater than the confidence threshold.

[0107] Among them, the target mask image being the same size as the first input image means that the target mask image has the same dimensions as the first input image and contains the same number of pixels.

[0108] Optionally, the process of the electronic device generating the target mask image can be as follows: When there is a target object in the first input image, an original mask image is generated. The size of the original mask image is the same as that of the first input image, and the number of pixels it contains is also the same. Then, the pixel regions of each first object are determined, each pixel within the pixel region of each first object is set to 1, and each pixel outside the pixel region of each first object is set to 0. The original mask image is subjected to two Gaussian blur processes, all pixels of the original mask image are multiplied by a first value, and the result is limited to a minimum value of a second value to obtain the target mask image. Here, the first weight is equal to the first value, and the second weight is equal to the second value.

[0109] Still taking the target object as a human face as an example, please refer to Figure 6 , which shows a schematic diagram of an image of a human face recognized by a second recognition model according to an exemplary embodiment of the present application. As Figure 6 shown, in the first input image 601, through the recognition of the second recognition model, each human face 602 included in the second recognition result, and each human face 603 corresponding to a confidence index greater than the confidence threshold. In the process of generating the target mask image, the above-mentioned original mask image is first generated, and the pixel regions of each human face 603 corresponding to a confidence index greater than the confidence threshold are determined. Each pixel within the pixel region of each first object in the original mask image is set to 1, and each pixel outside the pixel region of each first object is set to 0. The original mask image is subjected to two Gaussian blur processes, all pixels of the original mask image are multiplied by 0.8, and the result is limited to a minimum value of 0.3 to obtain the target mask image. That is, the pixel value within the pixel region of each first object in the target mask image is 0.8, and the pixel value outside the pixel region of each first object is 0.3. Then, the first weight is equal to 0.8, and the second weight is equal to 0.3. Here, the settings of 0.8 and 0.3 are set by developers based on experience, and in practical applications, they can also be other values less than 1. For example, the pixel value within the pixel region of each first object in the target mask image is set to 0.7, and the pixel value outside the pixel region of each first object is 0.1.

[0110] Step 505, perform a fusion process on the first input image and the first output image according to the target mask image to obtain a target image, where the target image is obtained based on the first product result and the second product result. The first product result is the product result obtained by multiplying each pixel of the first input image by the pixel value of each pixel of the target mask image as the input weight. The second product result is the product result obtained by multiplying each pixel of the first output image by (1 minus the pixel value of each pixel of the target mask image) as the output weight.

[0111] Optionally, in this embodiment, the input weights include a first weight and a second weight, and the output weights include a third weight and a fourth weight; the first weight is the weight corresponding to each pixel point within the target area in the first input image, and the second weight is the weight corresponding to each pixel point outside the target area in the first input image; the third weight is the weight corresponding to each pixel point within the target area in the first output image, and the fourth weight is the weight corresponding to each pixel point outside the target area in the first output image; the first weight is greater than the third weight, the second weight is less than the fourth weight, the sum of the first weight and the third weight is one, and the sum of the second weight and the fourth weight is one.

[0112] Through the generation of the above target mask image, the input weights are equivalent to the pixel values of each pixel point of the target mask image, and the output weights are equivalent to 1 minus the pixel values of each pixel point of the target mask image. Optionally, the electronic device can calculate the pixel values of each pixel point in the target image in the following manner: mask * input + (1 - mask) * GANout. Where mask represents the target mask image, input represents the first input image, and GANout represents the first output image.

[0113] Taking the target object as a human face, and the pixel value within the pixel area of each first object in the target mask image is 0.8, and the pixel value outside the pixel area of each first object is 0.3 as an example, fusing the two images (the first input image and the first output image) through the above mask * input + (1 - mask) * GANout formula is equivalent to fusing the area within the human face into the pixel value obtained by 0.2 * GANout + 0.8 * input, and fusing the area outside the human face into the pixel value obtained by 0.7 * GANout + 0.3 * input, finally achieving the effect of no abnormal particles generated in the human face part and certain protection of the background clarity.

[0114] It should be noted that the above first weight, second weight, third weight, and fourth weight can also be preset by developers in addition to being obtained based on the pixel values in the above target mask image. That is, when performing fusion, the target mask image is not generated, and directly multiply by the set weights and sum them as in step 304 above. Moreover, in this solution, after the electronic device executes the above steps 501 to 503, when it is determined that there is a target object in the first input image, the above step 304 can also be directly applied to perform fusion based on the preset target weights, that is, as in step 304 above, the input weight is greater than the output weight, and the sum of the input weight and the output weight is one. For example, the first input image and the first output image are fused in such a way that the input weight is 0.8 and the output weight is (1 - 0.8) = 0.2.

[0115] In this embodiment, when the input weights include the first weight and the second weight, and the output weights include the third weight and the fourth weight, it is equivalent to realizing the fusion of the first input image and the first output image with different weights inside and outside the face region of the small-sized face, making the pixel transition between the inside and outside of the face region smoother, avoiding over-sharpening of the small-sized face, and improving the transition effect between the face region and the surrounding region in the image generated by the electronic device.

[0116] In summary, in an electronic device, the first input image is processed to obtain a first output image, and the clarity of the first output image is higher than that of the first input image; when there is a target object in the first input image, the first input image and the first output image are fusion-processed based on the target weights to obtain a target image, and the clarity of the target region in the target image is greater than that of the first input image and less than that of the first output image. The target region is the image region where the target object is located, and the target weights include the input weight and the output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image will be fused based on the target weights, so that the clarity of the target region in the finally generated target image is greater than that of the first input image and less than that of the first output image, thereby reducing the abruptness between the inside and outside of the target region, avoiding over-sharpening of the target object, and improving the transition effect between the target region and the surrounding region in the image generated by the electronic device.

[0117] Next, taking the electronic device as a mobile phone, and the above-mentioned first recognition model and second recognition model are set in the mobile phone, the image generation process executed in the mobile phone can be as follows. Please refer to Figure 7 , which shows a method flowchart of an image generation method provided by an exemplary embodiment of the present application. The image generation method can be applied to an electronic device. As Figure 7 shown, the image generation method may include the following steps:

[0118] Step 701, the GAN model processes the first input image to obtain a first output image.

[0119] Among them, the GAN model in the mobile phone processes an input image to obtain a first output image with higher clarity / quality.

[0120] Step 702, the first recognition model of the ISP platform recognizes the first input image.

[0121] Optionally, the mobile phone first performs recognition through the first recognition model, and the recognition result includes the number of each recognized face.

[0122] Step 703, whether a face is detected.

[0123] Optionally, if the number of each face in the recognition result is greater than 0, it is considered that a face is detected; if it is less than 0, no face is detected. If a face is detected, perform Step 704; if no face is detected, perform Step 705.

[0124] Step 704, detect whether the face is within the preset FOV.

[0125] Optionally, the relevant content in the above Figure 3 shown embodiment can be referred to here. If there is a face within the preset FOV, perform Step 707; if there is no face within the preset FOV, perform Step 708.

[0126] Step 705, perform recognition on the first input image through the second recognition model.

[0127] Optionally, when the mobile phone does not detect a face after performing recognition through the first recognition model, it then performs recognition on the first input communication through the second recognition model, and the recognition result includes the face features of each recognized face, such as features like the side length, shape, color, etc. corresponding to the area of the face.

[0128] Step 706, perform classification through the adaptive weight decision model based on SVM.

[0129] Optionally, this is equivalent to the above Figure 5 process of calculating and detecting the confidence index of the recognition result of the second recognition model by the adaptive weight decision model based on SVM in the embodiment. After classification by the adaptive weight decision model, if the number of classifications of target objects less than or equal to the second size is greater than 0, that is, there are small-sized faces, enter Step 709; if the number of classifications of target objects less than or equal to the second size is not greater than 0, that is, there are no small-sized faces, enter Step 708.

[0130] Step 707, fuse the first input image and the first output image in the way that the input weight is 0.8 and the output weight is 0.2.

[0131] Step 708, fuse the first input image and the first output image in the way that the input weight is 0 and the output weight is 1.

[0132] Step 709, fuse the first input image and the first output image with different weights inside and outside the face area.

[0133] Among them, the fusion process is to multiply the pixel value of each pixel point of the first input image by the corresponding weight, and add the pixel value of the corresponding pixel point in the first output image multiplied by the corresponding weight, so as to obtain the pixel value of each pixel point in the target image.

[0134] Optionally, the implementation process of step 709 can be implemented according to the processes of step 504 and step 505 in the above Figure 5 and will not be elaborated here.

[0135] In summary, in an electronic device, the first input image is processed to obtain a first output image, and the clarity of the first output image is higher than that of the first input image; when there is a target object in the first input image, the first input image and the first output image are fused based on the target weight to obtain a target image, and the clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image. The target area is the image area where the target object is located, and the target weight includes an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image will be fused based on the target weight, so that the clarity of the target area in the finally generated target image is greater than that of the first input image and less than that of the first output image, thereby reducing the abruptness between the inside and outside of the target area, avoiding the oversharpening of the target object, and improving the transition effect between the target area and the surrounding area in the image generated by the electronic device.

[0136] The following is an apparatus embodiment of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the present application.

[0137] Please refer to Figure 8 , which shows a structural block diagram of an image generation apparatus provided by an exemplary embodiment of the present application. The image generation apparatus 800 can be used in an electronic device to execute all or part of the steps executed by the electronic device in the methods provided by the above various illustrated embodiments. The image generation apparatus 800 includes:

[0138] A first acquisition module 801, configured to process a first input image to obtain a first output image, where the clarity of the first output image is higher than that of the first input image;

[0139] A second acquisition module 802, configured to, when there is a target object in the first input image, perform a fusion process on the first input image and the first output image based on a target weight to obtain a target image, where the clarity of a target area in the target image is greater than that of the first input image and less than that of the first output image, the target area is an image area where the target object is located, the target weight includes an input weight and an output weight, the input weight is the weight of the first input image, and the output weight is the weight of the first output image.

[0140] In summary, in an electronic device, the first input image is processed to obtain a first output image, and the clarity of the first output image is higher than that of the first input image; when there is a target object in the first input image, a fusion process is performed on the first input image and the first output image based on a target weight to obtain a target image, where the clarity of a target area in the target image is greater than that of the first input image and less than that of the first output image, the target area is an image area where the target object is located, the target weight includes an input weight and an output weight, the input weight is the weight of the first input image, and the output weight is the weight of the first output image. In this solution, after generating a clearer first output image, if there is a target object in the first input image, the first input image and the second input image are fused based on the target weight, so that the clarity of the target area in the finally generated target image is greater than that of the first input image and less than that of the first output image, thereby reducing the abruptness between the target area and the area outside the target area, avoiding the oversharpening of the target object, and improving the transition effect between the target area and the surrounding area in the image generated by the electronic device.

[0141] Optionally, the input weight is greater than the output weight, and the sum of the input weight and the output weight is 1.

[0142] Optionally, the electronic device is configured with a first recognition model, and the first recognition model is used to recognize a target object greater than or equal to a first size, where the first size is greater than a first threshold;

[0143] The apparatus further includes:

[0144] A second acquisition module, configured to, before performing the fusion process on the first input image and the first output image based on the target weight to obtain a target image, recognize the first input image through the first recognition model to obtain a first recognition result;

[0145] A first determination module, configured to determine that there is a target object in the first input image when the first recognition result indicates that the first input image includes a target object greater than or equal to the first size.

[0146] Optionally, the first determination module includes: a first detection unit and a first determination unit;

[0147] The first detection unit is configured to detect whether a target object with a size greater than or equal to the first size included in the first input image is valid when the first recognition result indicates that the first input image includes a target object with a size greater than or equal to the first size;

[0148] The first determination unit is configured to determine that there is a target object in the first input image when at least one target object with a size greater than or equal to the first size is valid.

[0149] Optionally, the first detection unit is further configured to:

[0150] For a first target object, detect whether the first target object is within a preset field of view (FOV); the first target object is any one of the target objects with a size greater than or equal to the first size;

[0151] Determine that the first target object is valid when the first target object is within the preset FOV;

[0152] Determine that the first target object is invalid when the first target object is not within the preset FOV.

[0153] Optionally, the first detection unit is further configured to:

[0154] Calculate an overlapping area between the first target object and the preset FOV according to pixel point coordinates of the first target object and the preset FOV;

[0155] Determine that the first target object is within the preset FOV when the overlapping area is greater than or equal to a first preset area, where the first preset area is determined based on an area of the first target object itself;

[0156] Determine that the first target object is not within the preset FOV when the overlapping area is less than the first preset area.

[0157] Optionally, the electronic device is configured with a second recognition model for recognizing a target object with a size less than or equal to a second size, where the second size is less than a first threshold;

[0158] The apparatus further includes:

[0159] A third acquisition module, configured to, before fusing the first input image and the first output image based on the target weight to obtain a target image, identify the first input image through the second identification model to obtain a second identification result;

[0160] A second determination module, configured to determine that there is a target object in the first input image when the second identification result indicates that the first input image contains a target object smaller than or equal to the second size.

[0161] Optionally, the second determination module includes:

[0162] A first acquisition unit, configured to, when the second identification result indicates that the first input image contains a target object smaller than or equal to the second size, obtain a confidence index of each target object according to the object features of each target object smaller than or equal to the second size and the classification label weights corresponding to the object features of each target object;

[0163] A second determination unit, configured to determine that there is a target object in the first input image when there is a confidence index greater than the confidence threshold among the confidence indexes of each target object.

[0164] Optionally, the apparatus further includes:

[0165] A first generation module, configured to, before fusing the first input image and the first output image based on the target weight to obtain a target image, generate a target mask image according to the first input image and each first object, where the target mask image has the same size as the first input image, and the positions of each first object in the target mask image are the same as the positions of each first object in the first input image, the pixel values within the pixel regions of each first object in the target mask image are the first weight, and the pixel values outside the pixel regions of each first object are the second weight, and each first object is each target object corresponding to each confidence index greater than the confidence threshold;

[0166] The second acquisition module is further configured to

[0167] Fuse the first input image and the first output image according to the target mask image to obtain a target image, where the target image is obtained based on a first product result and a second product result. The first product result is the product result obtained by multiplying each pixel value of the target mask image as the input weight by each pixel of the first input image. The second product result is the product result obtained by multiplying each pixel value of the first output image with (1 minus the pixel value of each pixel of the target mask image) as the output weight.

[0168] Optionally, the input weight includes a first weight and a second weight, and the output weight includes a third weight and a fourth weight;

[0169] The first weight is the weight corresponding to each pixel within the target area in the first input image, and the second weight is the weight corresponding to each pixel outside the target area in the first input image;

[0170] The third weight is the weight corresponding to each pixel within the target area in the first output image, and the fourth weight is the weight corresponding to each pixel outside the target area in the first output image;

[0171] The first weight is greater than the third weight, the second weight is less than the fourth weight, the sum of the first weight and the third weight is 1, and the sum of the second weight and the fourth weight is 1.

[0172] Optionally, the second recognition model is trained according to first target sample data. The first target sample data includes synthetic images obtained by fusing various background images with images containing target objects using a preset fusion algorithm. The preset fusion algorithm is used to fuse the target objects in the images containing target objects with various background images after processing the target objects according to a preset processing logic. The preset processing logic includes one or more of scaling according to a preset scaling ratio, rotating according to a preset rotation angle, and adjusting illumination with preset illumination parameters.

[0173] Optionally, the electronic device is further configured with a first recognition model for recognizing target objects greater than or equal to a first size, where the first size is greater than the second size;

[0174] The third acquisition module is further configured to:

[0175] In a case where the first recognition result indicates that the first input image does not contain a target object greater than or equal to the first size, the second recognition model is used to recognize the first input image to obtain a second recognition result, where the first recognition result is a recognition result obtained by recognizing the first input image through the first recognition model.

[0176] Optionally, the electronic device includes an image processing model, which is trained based on second target sample data. The clarity of the output image of the image processing model is higher than that of the input image of the image processing model. The second target sample data includes original images of various target objects and reference images generated according to the original images. The clarity of the area where the target object is located in the reference image is lower than the clarity of the area outside the area where the target object is located in the reference image, and the clarity of the area outside the area where the target object is located in the reference image is higher than the clarity of the area outside the area where the target object is located in the original image.

[0177] The first obtaining module is further configured to:

[0178] Process the first input image through the image processing model to obtain the first output image.

[0179] Please refer to Figure 9 , which is a schematic structural diagram of another example of the image generation device provided in the embodiments of the present application. Among them, the image generation device 900 may be an electronic device, and the electronic device can implement the functions in the method provided in the embodiments of the present application. Among them, the image generation device 900 may be a chip system. In the embodiments of the present application, the chip system may be composed of chips or may include chips and other discrete devices.

[0180] In terms of hardware implementation, the above communication module may be a transceiver, and the transceiver is integrated in the image generation device 900 to form a communication interface 903.

[0181] The image generation device 900 includes at least one processor 901, which is used to implement or support the image generation device 900 to implement the functions of the electronic device in the method provided in the embodiments of the present application. Exemplarily, the processor 901 may execute steps such as processing the first input image to obtain a first output image, and in a case where a target object exists in the first input image, performing a fusion process on the first input image and the first output image based on a target weight to obtain a target image. For specific details, please refer to the detailed description in the method example, and details are not described here.

[0182] The image generation device 900 may further include at least one memory 902 for storing program instructions and / or data. The memory 902 is coupled to the processor 901. The coupling in the embodiments of the present application is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms for information interaction between devices, units or modules. The processor 901 may cooperate with the memory 902. The processor 901 may execute the program instructions stored in the memory 902. At least one of the at least one memory may be included in the processor.

[0183] The image generation device 900 may further include a communication interface 903 for communicating with other devices through a transmission medium, so that the devices in the image generation device 900 can communicate with other devices. Exemplarily, the other device may be a network-side device. The processor 901 may use the communication interface 903 to send and receive data. The communication interface 903 may specifically be a transceiver.

[0184] In the embodiments of the present application, the specific connection medium between the communication interface 903, the processor 901 and the memory 902 is not limited. In the embodiments of the present application Figure 9 it is shown that the memory 902, the processor 901 and the communication interface 903 are connected through a bus 904. The bus is represented by a thick line in Figure 9 The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only one thick line is used to represent it in

[0185] In the embodiments of the present application, the processor 901 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0186] In the embodiments of the present application, the memory 902 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or may also be a volatile memory, such as a random-access memory (RAM). A memory is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0187] Optionally, the embodiments of the present application further provide an electronic device, which includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, all or part of the steps performed by the electronic device in the image generation method of the above various embodiments are implemented.

[0188] Optionally, the embodiments of the present application further provide a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, all or part of the steps performed by the electronic device in the image generation method of the above various embodiments are implemented.

[0189] Optionally, the embodiments of the present application further provide a chip, which includes a runnable computer program. When the chip executes the computer program, all or part of the steps performed by the electronic device in the image generation method of the above various embodiments are implemented.

[0190] Optionally, the embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the image generation method of the above various embodiments is implemented.

[0191] Optionally, the embodiments of the present application further provide an application publishing platform, which is used to publish a computer program product. When the computer program product runs on a computer, the computer is caused to execute all or part of the steps performed by the electronic device in the image generation method of the above various embodiments.

[0192] It should be noted that: when the device provided in the above embodiments executes the control of the electronic device, only the above division of each functional module is used for illustration. In practical applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0193] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0194] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc.

[0195] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image generation method, characterized in that, Applied to an electronic device, the method includes: Processing a first input image to obtain a first output image, where the clarity of the first output image is higher than that of the first input image; When there is a target object in the first input image, performing a fusion process on the first input image and the first output image based on a target weight to obtain a target image, where the clarity of the target area in the target image is greater than that of the first input image and less than that of the first output image, the target area is the image area where the target object is located, the target weight includes an input weight and an output weight, the input weight is the weight of the first input image, and the output weight is the weight of the first output image.

2. The method according to claim 1, wherein The input weight is greater than the output weight, and the sum of the input weight and the output weight is one.

3. The method according to claim 2, wherein The electronic device is configured with a first recognition model for recognizing target objects greater than or equal to a first size, where the first size is greater than a first threshold; Before performing the fusion process on the first input image and the first output image based on the target weight to obtain a target image, the method further includes: Recognizing the first input image through the first recognition model to obtain a first recognition result; When the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, determining that there is a target object in the first input image.

4. The method according to claim 3, wherein When the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, determining that there is a target object in the first input image includes: When the first recognition result indicates that the first input image contains a target object greater than or equal to the first size, detecting whether the target object greater than or equal to the first size contained in the first input image is valid; When at least one target object greater than or equal to the first size is valid, determining that there is a target object in the first input image.

5. The method according to claim 4, wherein Detecting whether the target object greater than or equal to the first size contained in the first input image is valid includes: For a first target object, detecting whether the first target object is within a preset field of view FOV; the first target object is any one of the target objects greater than or equal to the first size; When the first target object is within the preset FOV, determining that the first target object is valid; When the first target object is not within the preset FOV, determining that the first target object is invalid.

6. The method according to claim 4, wherein The detecting whether the first target object is within the preset field of view FOV includes: Calculating the overlapping area between the first target object and the preset FOV according to the pixel point coordinates of the first target object and the preset FOV; When the overlapping area is greater than or equal to the first preset area, it is determined that the first target object is within the preset FOV, and the first preset area is determined based on the area of the first target object itself; When the overlapping area is less than the first preset area, it is determined that the first target object is not within the preset FOV.

7. The method according to claim 1, wherein The electronic device is configured with a second recognition model, and the second recognition model is used to recognize target objects less than or equal to a second size, and the second size is less than a first threshold; Before performing the fusion processing on the first input image and the first output image based on the target weights to obtain a target image, the method further includes: Recognizing the first input image through the second recognition model to obtain a second recognition result; When the second recognition result indicates that the first input image contains a target object less than or equal to the second size, it is determined that there is a target object in the first input image.

8. The method according to claim 7, wherein The determining that there is a target object in the second input image when the second recognition result indicates that the first input image contains a target object less than or equal to the second size includes: When the second recognition result indicates that the first input image contains a target object less than or equal to the second size, obtaining a confidence index of each target object according to the object characteristics of each target object less than or equal to the second size and the classification label weights corresponding to the object characteristics of each target object; When there is a confidence index greater than the confidence threshold among the confidence indexes of each target object, it is determined that there is a target object in the first input image.

9. The method according to claim 8, wherein Before performing the fusion processing on the first input image and the first output image based on the target weights to obtain a target image, the method further includes: Generating a target mask image according to the first input image and each first object, where the target mask image is the same size as the first input image, and the positions of each first object in the target mask image are the same as the positions of each first object in the first input image. The pixel values within the pixel regions of each first object in the target mask image are the first weights, and the pixel values outside the pixel regions of each first object are the second weights. Each first object is each target object corresponding to each confidence index greater than the confidence threshold; The performing the fusion processing on the first input image and the first output image based on the target weights to obtain a target image includes: Fuse the first input image and the first output image according to the target mask image to obtain a target image, where the target image is obtained based on a first product result and a second product result. The first product result is a product result obtained by multiplying each pixel value of the target mask image as the input weight by each pixel of the first input image. The second product result is a product result obtained by multiplying each pixel value of the first output image by (1 minus each pixel value of the target mask image) as the output weight.

10. The method according to any one of claims 7 to 9, characterized in that The input weight includes a first weight and a second weight, and the output weight includes a third weight and a fourth weight; The first weight is the weight corresponding to each pixel within the target area in the first input image, and the second weight is the weight corresponding to each pixel outside the target area in the first input image; The third weight is the weight corresponding to each pixel within the target area in the first output image, and the fourth weight is the weight corresponding to each pixel outside the target area in the first output image; The first weight is greater than the third weight, the second weight is less than the fourth weight, the sum of the first weight and the third weight is 1, and the sum of the second weight and the fourth weight is 1.

11. The method according to claim 7, wherein The second recognition model is trained according to first target sample data, which includes synthetic images obtained by fusing various background images with an image containing a target object using a preset fusion algorithm. The preset fusion algorithm is used to fuse the target object in the image containing the target object with various background images after processing the target object according to a preset processing logic. The preset processing logic includes one or more of scaling according to a preset scaling ratio, rotating according to a preset rotation angle, and adjusting illumination with preset illumination parameters.

12. The method according to claim 7, wherein The electronic device is further configured with a first recognition model for recognizing a target object greater than or equal to a first size, where the first size is greater than the second size; Obtaining a second recognition result by recognizing the first input image through the second recognition model includes: In the case where the first recognition result indicates that the first input image does not contain a target object greater than or equal to the first size, recognize the first input image through the second recognition model to obtain a second recognition result, where the first recognition result is a recognition result obtained by recognizing the first input image through the first recognition model.

13. The method according to claim 1, characterized in that, The electronic device includes an image processing model, which is trained based on second target sample data. The clarity of the output image of the image processing model is higher than that of the input image of the image processing model. The second target sample data includes the original images of various target objects and the reference images generated according to the original images. The clarity of the area where the target object is located in the reference image is lower than the clarity of the area outside the area where the target object is located in the reference image. The clarity of the area outside the area where the target object is located in the reference image is higher than the clarity of the area outside the area where the target object is located in the original image; Processing the first input image to obtain a first output image includes: Processing the first input image through the image processing model to obtain the first output image.

14. An image generation device, characterized in that, Applied to an electronic device, the device includes: A first acquisition module, configured to process a first input image to obtain a first output image, and the clarity of the first output image is higher than that of the first input image; A second acquisition module, configured to, when there is a target object in the first input image, perform a fusion process on the first input image and the first output image based on target weights to obtain a target image, and the clarity of the target area in the target image is greater than the clarity of the first input image and less than the clarity of the first output image. The target area is the image area where the target object is located, and the target weights include an input weight and an output weight. The input weight is the weight of the first input image, and the output weight is the weight of the first output image.

15. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, the image generation method according to any one of claims 1 to 13 is implemented.

16. A computer program product, characterized in that, Including a computer program, which implements the image generation method according to any one of claims 1 to 13 when executed by a processor.