Generation method and device of stylized image, electronic equipment, medium and product
By recognizing the area ratio of the face region and setting stylization parameters, adjusting the mask occlusion ratio and stylization intensity, the problems of face collapse and insufficient stylization in stylization transfer are solved, and high-quality stylized image generation is achieved.
Patent Information
- Application Number
- CN202511044748.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, excessive stylization intensity during stylization transfer can lead to facial distortion, while insufficient stylization intensity results in inadequate background stylization, making it difficult to improve the overall stylization of the image while maintaining facial quality.
By identifying the face region in the original image to be stylized, a first stylization parameter and a second stylization parameter are set based on the area ratio of the face region, and the occlusion ratio and stylization intensity of the mask are adjusted to generate a stylized image.
While ensuring that the faces do not become distorted, the overall stylization of the images is improved, thus enhancing the quality of the stylized images.
Smart Images

Figure CN120953409A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image stylization transfer technology, and particularly relates to a method for generating stylized images, an apparatus for generating stylized images, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] Stable diffusion-based training is particularly popular in the field of image-to-image processing. The basic steps are: by collecting a certain number of images and performing stylistic fine-tuning, images with a style similar to the training images can be obtained. After training, the stylization intensity is set based on the input images. Generally, the lower the stylization intensity, the smaller the change to the original image; the higher the stylization intensity, the greater the change to the original image.
[0003] However, excessive stylization intensity can cause the face to distort after stylization transfer; while insufficient stylization intensity can result in inadequate background stylization. Summary of the Invention
[0004] This application aims to at least address one of the technical problems existing in the prior art. To this end, this application proposes a method for generating stylized images, an apparatus for generating stylized images, an electronic device, a computer-readable storage medium, and a computer program product, which can maximize the overall stylization of the image and improve the image quality of the stylized image while ensuring the quality of the face in the stylized image.
[0005] In a first aspect, this application provides a method for generating stylized images, comprising:
[0006] Identify face regions in the original image to be styled and transferred;
[0007] Based on the area ratio of the face region in the original image, a first stylization parameter and a second stylization parameter are set. The first stylization parameter is used to characterize the occlusion ratio of the face region by the mask, and the second stylization parameter is used to characterize the stylization intensity of the original image.
[0008] Based on the first stylization parameter and the second stylization parameter, the original image is subjected to stylization transfer to generate a stylized image.
[0009] Secondly, this application provides an apparatus for generating stylized images, the apparatus comprising:
[0010] The recognition module is used to identify face regions in the original image to be styled and transferred;
[0011] The setting module is used to set a first stylization parameter and a second stylization parameter based on the area ratio of the face region in the original image. The first stylization parameter is used to characterize the degree of occlusion of the face region by the mask, and the second stylization parameter is used to characterize the degree of difference between the original image and the stylized image.
[0012] The stylization transfer module is used to perform stylization transfer on the original image based on the first stylization parameter and the second stylization parameter to generate a stylized image.
[0013] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for generating stylized images.
[0014] Fourthly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method for generating stylized images.
[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for generating stylized images.
[0016] The stylized image generation method, apparatus, electronic device, computer-readable storage medium, and computer program provided in this application determine the area proportion of the face region in the original image by identifying the face region in the original image to be stylized. Different proportions of the face region in the original image result in different effects of stylization intensity on the face.
[0017] It's understandable that a smaller proportion of the face region means that even a lower stylization intensity can lead to a poorly rendered face (e.g., completely different from the original face or even unrecognizable); conversely, a larger proportion of the face region means that even a higher stylization intensity may not cause the face to be ruined. Therefore, the proportion of the face region affects the degree to which stylization intensity influences the image quality of the face region.
[0018] Therefore, this application takes into account the proportion of the face region to adaptively set the second stylization parameter of the original image, so as to ensure that the face region does not collapse and to maximize the stylization intensity of the original image.
[0019] Furthermore, by setting the first stylization parameter of the face region, the occlusion ratio of the face region mask is set, thereby adjusting the influence of the second stylization parameter on the face region through the mask. For example, the higher the occlusion ratio, the smaller the area of the face to be stylized, and the smaller the influence of the stylization intensity on the face.
[0020] By combining the first and second stylization parameters, the stylization intensity of the original image can be further improved without compromising the face region. This avoids the impact of excessively strong or weak stylization intensity on the stylized image, thereby improving the image quality of the stylized image.
[0021] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description
[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0023] Figure 1 This is a schematic diagram illustrating an application scenario of the stylized image generation method provided in the embodiments of this application;
[0024] Figure 2 This is a first flowchart illustrating the method for generating stylized images provided in this application embodiment;
[0025] Figure 3 This is a schematic diagram of a first scene of the method for generating stylized images provided in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of the second process of the method for generating stylized images provided in the embodiments of this application;
[0027] Figure 5 This is a schematic diagram of a second scene of the method for generating stylized images provided in the embodiments of this application;
[0028] Figure 6 This is a schematic diagram of a third scene of the method for generating stylized images provided in the embodiments of this application;
[0029] Figure 7 This is a schematic diagram illustrating the principle of the stylized image generation method provided in the embodiments of this application;
[0030] Figure 8 This is a schematic diagram of the module of the stylized image generation apparatus provided in the embodiments of this application;
[0031] Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; and
[0032] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0033] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0034] To facilitate understanding, the application scenarios of this application will be introduced below:
[0035] Please see Figure 1 , Figure 1 This is an application scenario diagram of a method for generating stylized images provided in an embodiment of this application. The application scenario provided in this application includes a terminal device 101 and a server 102, and the method for generating stylized images provided in this application can be executed by at least one of the terminal device 101 and the server 102.
[0036] The terminal device may integrate a client, which can be a client with the function of displaying data information such as text, images, audio, and video, including but not limited to a device control client associated with the processing equipment, so as to realize communication and control with the processing equipment. This client can be a standalone client or an embedded sub-client integrated into a client (e.g., a social client), and there is no limitation on this.
[0037] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This application does not limit this.
[0038] It should be noted that, Figure 1 The number of terminal devices and servers is for illustrative purposes only; the number of terminal devices and servers can be more or less, and there is no limitation herein. Terminal devices and servers can be connected directly or indirectly via wired or wireless communication, and this application does not impose any limitations on this.
[0039] The method for generating stylized images involved in this application can be implemented using cloud technology.
[0040] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing.
[0041] Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to the cloud computing business model. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will all require robust system support, which can only be achieved through cloud computing.
[0042] The stylized image generation method of this application can be implemented based on cloud computing. Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computers, enabling various application systems to obtain computing power, storage space, and information services as needed. The network providing these resources is called the "cloud." From the user's perspective, resources in the "cloud" are infinitely scalable, readily available, on-demand, expandable, and pay-as-you-go.
[0043] As a provider of fundamental cloud computing capabilities, a cloud resource pool (referred to as a cloud platform, generally called an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose from. The cloud resource pool mainly includes: computing devices (virtualized machines containing operating systems), storage devices, and network devices.
[0044] The method for generating stylized images in this application embodiment can be executed by an electronic device, which can be at least one of a server and a terminal device. That is, the method can be executed by the server or the terminal device alone, or by both the server and the terminal device. Therefore, the executing entity of each step will not be described again below.
[0045] Based on the above description, this application provides a method for generating stylized images. The method for generating stylized images will be described in detail below:
[0046] Please see Figure 2 and Figure 3 The stylized image generation method provided in this application embodiment is implemented by steps 011 to 013, which are described in detail below.
[0047] Step 011: Identify the face region S2 in the original image S1 to be styled and transferred.
[0048] Style transfer is a technique that blends the content of one image with the style of another to generate a new image that retains the original content while also having the target style.
[0049] Among them, the face region S2 is the region where the face is located in the original image S1.
[0050] In one alternative embodiment, the face region S2 can be determined by extracting facial features from the original image S1 to identify the face; or, the face region S2 can be identified by a pre-trained face recognition model (such as the YOLO-Face model).
[0051] Please see Figure 4 and 5 In one optional embodiment, step 011 includes:
[0052] Step 0111: Identify faces in the original image and determine the bounding box region of the faces;
[0053] Step 0112: Determine the face region based on the bounding box region.
[0054] Specifically, after recognizing the face in the original image S1, in order to facilitate the setting of the mask and stylization processing, the bounding box region S3 of the face can be determined, such as a rectangular region or a circular region containing the face, and thus the bounding box region S3 is determined as the face region S2.
[0055] In one optional embodiment, there are multiple faces. Based on the area of the bounding box region S3 corresponding to each face, a target face S4 is determined among the multiple faces. Based on the bounding box region S3 corresponding to the target face S4, a face region S2 is determined.
[0056] For example, the target face S4 is a user-defined face that needs to be stylized; or, the target face S4 is the face whose face area is the median of the face areas of the multiple faces; or, the target face S4 is the face with the largest face area among the multiple faces, etc.
[0057] Thus, in scenarios with multiple faces, the image quality of the target face S4 in the stylized image is guaranteed by determining the appropriate target face S4.
[0058] Step 012: Based on the area ratio of the face region in the original image, set the first stylization parameter and the second stylization parameter.
[0059] The area ratio of the face region in the original image refers to the ratio of the area of the face region to the area of the original image.
[0060] The first stylization parameter is used to characterize the proportion of the face region occluded by the mask. For example, the first stylization parameter and the proportion of the face region occluded by the mask are negatively correlated, or the first stylization parameter and the proportion of the face region occluded by the mask are positively correlated.
[0061] Taking the negative correlation between the first stylization parameter and the mask on the occlusion ratio of the face region as an example, the value range of the first stylization parameter is [0, 1]. When the first stylization parameter is 0, the mask occludes the entire face region. When the first stylization parameter is 1, the mask does not occlude any area of the face region. When the first stylization parameter is 0.3, the mask occludes 70% of the face region.
[0062] In one optional embodiment, a mask S5 covers a face region S2. The mask S5 includes multiple evenly distributed occlusion regions S6, and the occlusion ratio of the face region S2 is positively correlated with the area of the occlusion region S6.
[0063] Thus, the occluded regions S6 are multiple and evenly distributed, which ensures a relatively uniform distribution of the unoccluded areas of the face region S2. It can be understood that the more occluded regions S6 there are, the more evenly the face region S2 can be stylized during stylization processing, thus guaranteeing the stylization transfer effect of the face region S2.
[0064] The second stylization parameter characterizes the stylization intensity of the original image. Stylization intensity represents the degree of change to the original image; the greater the stylization intensity, the greater the change to the original image.
[0065] In one alternative embodiment, the stylization intensity and the degree of image difference are positively correlated, with the degree of image difference used to characterize the degree of difference between the original image and the stylized image.
[0066] In other words, the greater the stylization intensity, the greater the degree of change to the original image, which will result in a greater difference between the original image and the stylized image.
[0067] In one alternative embodiment, both the first stylization parameter and the second stylization parameter are positively correlated with the area ratio.
[0068] In other words, the smaller the area ratio, the smaller the first and second stylization parameters; the larger the area ratio, the larger the first and second stylization parameters.
[0069] It's understandable that the smaller the area proportion, the more likely a smaller stylization intensity will cause facial distortion. Therefore, the second stylization parameter needs to be set relatively small. Conversely, the smaller the area proportion, the more necessary it is to reduce the impact of stylization intensity on the facial region to ensure the face doesn't distort. In this case, a smaller first stylization parameter can be set to increase the mask's occlusion ratio of the facial region.
[0070] It is understandable that the area covered by the mask will not be affected by stylization transfer. That is, the pixels of the masked area do not change and remain consistent with the pixels of the corresponding area in the face in the original image, thereby reducing the impact of stylization transfer on the face area.
[0071] After setting a smaller first stylization parameter to reduce the impact of stylization transfer, a second stylization parameter can be appropriately increased to maintain the overall stylization intensity of the original image, avoid an insufficient stylization intensity that would result in inadequate stylization of the original image, and ensure the image quality of the stylized image.
[0072] When the face region occupies a large area, even setting a high stylization intensity will not cause face distortion. Therefore, a larger second stylization parameter can be set. Furthermore, a larger first stylization parameter can be set to minimize occlusion of the face region, thereby improving the stylization transfer effect of the face region (e.g., the style is basically consistent with the style of the target image to be stylized) while ensuring that the face does not distort.
[0073] Please refer to it again. Figure 4 In an optional embodiment, step 012:
[0074] Step 0121: Obtain multiple preset percentage ranges, and set corresponding first preset stylization parameters and second preset stylization parameters for each percentage range;
[0075] Step 0122: Determine the first preset stylization parameter corresponding to the target proportion interval where the area proportion is located as the first stylization parameter, and the corresponding second preset stylization parameter as the second stylization parameter.
[0076] Specifically, the area proportion range can be determined based on the maximum and minimum area proportions, and this range can be divided into multiple proportion intervals. For each proportion interval, a corresponding first preset stylization parameter and a second preset stylization parameter are set. The first and second preset stylization parameters for each proportion interval are empirical values. Experiments can be conducted beforehand to obtain the most suitable first and second preset stylization parameters for each proportion interval, thereby obtaining the best-quality stylized image for the current proportion interval.
[0077] Next, based on the area proportion of the face region in the original image to be stylized, the target proportion interval (such as any preset proportion interval) is determined. Then, the first preset stylization parameter corresponding to the target proportion interval is determined as the first stylization parameter, and the second preset stylization parameter corresponding to the target proportion interval is determined as the second stylization parameter.
[0078] In this way, the corresponding first and second stylization parameters can be accurately determined based on the current area proportion of the face region, thereby improving the image quality of the stylized image.
[0079] In one optional embodiment, the multiple proportion intervals include a first proportion interval, a second proportion interval, a third proportion interval, a fourth proportion interval, and a fifth proportion interval; the first preset stylization parameter corresponding to each of the first proportion interval, the second proportion interval, the third proportion interval, the fourth proportion interval, and the fifth proportion interval increases, and the preset second stylization parameter corresponding to each of the first proportion interval, the second proportion interval, the third proportion interval, the fourth proportion interval, and the fifth proportion interval increases.
[0080] It is understandable that the smaller the area of the face region, the smaller the stylization intensity can be set (i.e., the smaller the corresponding first preset stylization parameter can be set), while the occlusion ratio of the face region can be set larger (i.e., the smaller the corresponding second preset stylization parameter can be set); the larger the area of the face region, the greater the stylization intensity can be set (i.e., the larger the corresponding first preset stylization parameter can be set), while the occlusion ratio of the face region can be set smaller (i.e., the larger the corresponding second preset stylization parameter can be set).
[0081] Therefore, the first and second preset stylization parameters corresponding to the first, second, third, fourth, and fifth proportion intervals all increase sequentially.
[0082] In one optional embodiment, the first proportion interval, the second proportion interval, the third proportion interval, the fourth proportion interval, and the fifth proportion interval are [0, 0.01), [0.01, 0.05), [0.05, 0.1), [0.1, 0.3), and [0.3, 1], respectively.
[0083] It is understandable that for most images, the proportion of a human face will not be too large, generally less than 0.3. Therefore, when dividing the proportion intervals, we can focus on subdividing [0, 0.3] to obtain the first proportion interval, the second proportion interval, the third proportion interval, and the fourth proportion interval, and then divide [0.3, 1] into the fifth proportion interval.
[0084] Thus, based on the distribution of face area proportions in various images, by reasonably dividing the face area proportion range, the area proportion intervals with a large number of distributions are divided into more proportion intervals, and the area proportion intervals with a small number of distributions are divided into fewer proportion intervals. This improves the accuracy of proportion interval division while avoiding too many proportion intervals, thereby improving the accuracy of the first and second preset stylization parameters corresponding to each proportion interval.
[0085] In one alternative embodiment, the input area ratio is fed into a preset neural network model to output a first stylization parameter and a second stylization parameter.
[0086] It is understandable that, in addition to pre-dividing the proportion range based on empirical values and setting the first and second preset stylization parameters that are most suitable for the proportion range, a neural network model can also be trained based on a preset training set. The neural network model can then output the first and second stylization parameters corresponding to different area proportions.
[0087] For example, the training set includes multiple training samples, and different training samples include first and second preset stylization parameters corresponding to different proportion intervals. By training the model to convergence through the training set, the corresponding first and second stylization parameters can be output based on the input area proportion.
[0088] Step 013: Based on the first stylization parameter and the second stylization parameter, perform stylization transfer on the original image to generate a stylized image.
[0089] Specifically, after setting the first stylization parameter and the second stylization parameter, the image region to be stylized in the original image (such as the image region outside the masked area in the original image) and the stylization intensity during stylization transfer can be determined. Thus, by using a reasonable stylization intensity, the image region to be stylized is stylized, thereby generating a stylized image.
[0090] For example, based on the first stylization parameter and the second stylization parameter, a pre-trained stylization transfer model can be used to perform stylization transfer on the original image, thereby generating a stylized image.
[0091] Among them, the stylization transfer model is a model that can apply the style of one image to another image to generate an image with a new style.
[0092] For example, stylistic transfer models include diffusion models, convolutional neural networks (CNNs), generative adversarial networks (GANs), and autoencoders.
[0093] Please refer to it again. Figure 4 In an optional embodiment, step 013:
[0094] Step 0131: Generate a mask image based on the first stylization parameter and the original image;
[0095] Step 0132: Using the original image as guiding information, perform stylization transfer on the mask image based on the second stylization parameter to generate a stylized image.
[0096] Specifically, during stylization transfer, the occlusion ratio of the mask is first determined based on the first stylization parameter. Then, based on the determined mask occlusion ratio, a mask is set on the original image to generate a mask image. In the mask image, the face portions occluded by the mask will not undergo stylization transfer, while the face portions not occluded by the mask will undergo stylization transfer.
[0097] Then, the stylization intensity can be determined based on the second stylization parameter, and then the original image can be stylized and transferred based on the determined stylization intensity to generate a stylized image.
[0098] Specifically, based on a determined stylization intensity, stylization transfer is performed on the image region outside the masked area in the original image to generate a locally stylized image. Then, the locally stylized image and the masked face portion (i.e., the masked area) are fused to generate a complete stylized image.
[0099] The local stylized image includes the image region outside the masked area after stylization transfer, as well as the mask. During fusion, the pixels of the masked area are used to replace the pixels at the corresponding mask positions in the local stylized image, thereby generating the final stylized image.
[0100] In the process of stylization transfer, the original image is used as guiding information to ensure that the stylized image after stylization transfer retains as much content as possible from the original image while possessing the target style.
[0101] When the original image serves as guiding information, its features and attention map can be used to guide the diffusion model to preserve the original spatial layout. Alternatively, during denoising, the diffusion model can guide the denoising process through conditional generation to maintain the structure of the content image. By adjusting the weights of the style loss, the model can achieve different degrees of balance between content and style, thereby generating stylized images that fully preserve the structure of the content image but have different style intensities.
[0102] In one optional embodiment, the original image, the mask image, and the second stylization parameters are input to a preset stylization transfer model for stylization transfer to generate a stylized image.
[0103] It is understandable that a preset stylization transfer model can be used for stylization transfer. Simply input the original image, the mask image, and the second stylization parameters into the preset stylization transfer model to quickly output a stylized image.
[0104] The stylized image generation method of this application determines the area proportion of the face region in the original image by identifying the face region in the original image to be stylized. Different proportions of the face region in the original image result in different effects of stylization intensity on the face.
[0105] It's understandable that a smaller proportion of the face region means that even a lower stylization intensity can lead to a poorly rendered face (e.g., completely different from the original face or even unrecognizable); conversely, a larger proportion of the face region means that even a higher stylization intensity may not cause the face to be ruined. Therefore, the proportion of the face region affects the degree to which stylization intensity influences the image quality of the face region.
[0106] Therefore, this application takes into account the proportion of the face region to adaptively set the second stylization parameter of the original image, so as to ensure that the face region does not collapse and to maximize the stylization intensity of the original image.
[0107] Furthermore, by setting the first stylization parameter of the face region, the occlusion ratio of the face region mask is set, thereby adjusting the influence of the second stylization parameter on the face region through the mask. For example, the higher the occlusion ratio, the smaller the area of the face to be stylized, and the smaller the influence of the stylization intensity on the face.
[0108] By combining the first and second stylization parameters, the stylization intensity of the original image can be further improved without compromising the face region. This avoids the impact of excessively strong or weak stylization intensity on the stylized image, thereby improving the image quality of the stylized image.
[0109] Please see Figure 7The following example illustrates the method for generating stylized images according to this application.
[0110] First, the image is input into the Yolo-face model, which identifies the bounding box region (Bbox) of the face (a rectangle with horizontal coordinates ranging from [x1, x2] and vertical coordinates ranging from [y1, y2]), thereby determining the face region.
[0111] Then, based on the proportion of the face region to the area of the input image (i.e. Figure 7 The stylization intensity (i.e., the proportion of the face in the image) is determined using a preset Denoise mapping table (such as a mapping table of different proportion ranges and corresponding second preset stylization parameters). Figure 7 The process involves determining the denoise intensity (corresponding to the second stylization parameter) and then setting a face mask (i.e., a face mask). Based on the face proportion and a preset mask noise mapping table (such as a mapping table of different proportion ranges and corresponding first preset stylization parameters), the mask occlusion ratio (i.e., the mask noise level, corresponding to the first stylization parameter) is determined.
[0112] Finally, based on the first stylization parameters, the second stylization parameters, and the trained diffusion model (i.e. Figure 7 The diffusion model (in the diffusion model) utilizes the inpainting mode of the graph-generated image to generate stylized images.
[0113] Specifically, the input can be an input image, a mask image (i.e., an image generated after processing the input image based on the first stylization parameter), a second stylization parameter, and other default parameters (such as the type of sampling algorithm) into the diffusion model. After inference by the diffusion model, the generated stylized image is obtained.
[0114] Based on the method described in the above embodiments, this application also provides a stylized image generation apparatus 300 for performing the steps in the above-described stylized image generation method. Please refer to... Figure 8 , Figure 8 This is a schematic diagram of a stylized image generation apparatus 300 provided in an embodiment of this application. The stylized image generation apparatus 300 includes:
[0115] The recognition module 301 is used to identify the face region in the original image to be styled and transferred;
[0116] The setting module 302 is used to set a first stylization parameter and a second stylization parameter based on the area ratio of the face region in the original image. The first stylization parameter is used to characterize the degree of occlusion of the face region by the mask, and the second stylization parameter is used to characterize the degree of difference between the original image and the stylized image.
[0117] The stylization transfer module 303 is used to perform stylization transfer on the original image based on the first stylization parameter and the second stylization parameter to generate a stylized image.
[0118] It should be noted that the specific details of each module unit in the above-mentioned stylized image generation device have been described in detail in the embodiments of the above-mentioned stylized image generation method, and will not be repeated here.
[0119] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0120] In some embodiments, the stylized image generation apparatus in this application can be implemented in hardware, such as an electronic device or a component in an electronic device, such as an integrated circuit or a chip; the stylized image generation apparatus can also be implemented in software, such as as an application installed in an electronic device.
[0121] In some embodiments, please refer to Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 includes a processor 501 and a memory 502. The memory 502 stores a computer program 503 that can run on the processor 501. When the processor 501 executes the program 503, it implements the various processes of the embodiments of the above-described stylized image generation method and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0122] Please see Figure 10 , Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device can be a terminal or a server. Exemplarily, the electronic device 700 includes a central processing unit (CPU) 701, a system memory 704 including random access memory (RAM) 702 and read-only memory (ROM) 703, and a system bus 705 connecting the system memory 704 and the central processing unit 701.
[0123] In some embodiments, the electronic device 700 may also include a basic input / output system 706 that helps transmit information between various devices within the computer, and a mass storage device 707 for storing the operating system 713, the client 714, and other program modules 715.
[0124] In some embodiments, the basic input / output system 706 includes a display 708 for displaying information and an input device 709 for user input, such as a touch panel and other input devices. A touch panel is also called a touchscreen. A touch panel may include both a touch device and a touch controller. Other input devices may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described further here.
[0125] Both the display 708 and the input device 709 are connected to the central processing unit 701 via an input / output controller 710 connected to the system bus 705. The basic input / output system 706 may also include the input / output controller 710 for receiving and processing input from touch panels, other input devices, etc. Similarly, the input / output system 706 also includes output devices such as displays, printers, or other types of output devices.
[0126] Mass storage device 707 is connected to central processing unit 701 via a mass storage controller (not shown) connected to system bus 705. Mass storage device 707 and its associated computer-readable media provide non-volatile storage for electronic device 700. That is, mass storage device 707 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.
[0127] According to various embodiments of this application, the electronic device 700 can also be connected to a remote computer on a network, such as the Internet. That is, the electronic device 700 can be connected to a network 717 via a network interface unit 716 connected to the system bus 705, or the network interface unit 716 can be used to connect to other types of networks or remote computer systems (not shown).
[0128] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described stylized image generation method and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0129] The processor can be the processor in the electronic device described in the above embodiments. The computer-readable storage medium can be a computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc.
[0130] Computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types.
[0131] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method for generating stylized images. The processor may be a processor in the electronic device described in the above embodiments. When executed by the processor, the computer program implements various processes of the embodiments of the above-described method for generating stylized images and achieves the same technical effects; therefore, to avoid repetition, further details are omitted here.
[0132] It is understood that in the specific implementation of this application, data related to user identity or characteristics is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
Claims
1. A method for generating stylized images, characterized in that, include: Identify face regions in the original image to be styled and transferred; Based on the area ratio of the face region in the original image, a first stylization parameter and a second stylization parameter are set. The first stylization parameter is used to characterize the occlusion ratio of the face region by the mask, and the second stylization parameter is used to characterize the stylization intensity of the original image. Based on the first stylization parameter and the second stylization parameter, the original image is subjected to stylization transfer to generate a stylized image.
2. The generation method according to claim 1, characterized in that, The identification of face regions in the original image to be styled for transfer includes: Identify faces in the original image and determine the bounding box region of the faces; The face region is determined based on the bounding box region.
3. The generation method according to claim 1, characterized in that, The number of faces is multiple, and the method further includes: Based on the area of the bounding box region corresponding to each face, the target face among the multiple faces is determined; The face region is determined based on the bounding box region corresponding to the target face.
4. The generation method according to claim 1, characterized in that, The first stylization parameter and the mask have a negative correlation with the proportion of occlusion of the face region, the second stylization parameter and the stylization intensity have a positive correlation, and both the first stylization parameter and the second stylization parameter have a positive correlation with the area ratio.
5. The generation method according to claim 4, characterized in that, The mask covers the face area, and the mask includes multiple evenly distributed occlusion areas. The occlusion ratio of the face area and the area of the occlusion area are positively correlated.
6. The generation method according to claim 4, characterized in that, The stylization intensity and the degree of image difference are positively correlated, and the degree of image difference is used to characterize the degree of difference between the original image and the stylized image.
7. The generation method according to claim 1, characterized in that, The step of setting a first stylization parameter and a second stylization parameter based on the area ratio of the face region in the original image includes: Multiple preset percentage ranges are obtained, and each percentage range is set with a corresponding first preset stylization parameter and a preset second stylization parameter; The first preset stylization parameter corresponding to the target proportion interval where the area proportion is located is determined as the first stylization parameter, and the corresponding preset second stylization parameter is determined as the second stylization parameter.
8. The generation method according to claim 7, characterized in that, The plurality of proportion intervals include a first proportion interval, a second proportion interval, a third proportion interval, a fourth proportion interval, and a fifth proportion interval; the first preset stylization parameter corresponding to each of the first proportion interval, the second proportion interval, the third proportion interval, the fourth proportion interval, and the fifth proportion interval increases, and the preset second stylization parameter corresponding to each of the first proportion interval, the second proportion interval, the third proportion interval, the fourth proportion interval, and the fifth proportion interval increases.
9. The generation method according to claim 1, characterized in that, The step of setting a first stylization parameter and a second stylization parameter based on the area ratio of the face region in the original image includes: The area ratio is input into a preset neural network model to output the first stylization parameter and the second stylization parameter.
10. The generation method according to claim 1, characterized in that, The step of performing stylization transfer on the original image based on the first stylization parameter and the second stylization parameter to generate a stylized image includes: Generate a mask image based on the first stylization parameters and the original image; Using the original image as guiding information, the mask image is stylized and transferred based on the second stylization parameter to generate the stylized image.
11. The generation method according to claim 10, characterized in that, The step of using the original image as guiding information and performing stylistic transfer on the mask image based on the second stylization parameters to generate the stylized image includes: Using the original image as guiding information, stylization transfer is performed on the image region outside the masked area in the masked image to generate a locally stylized image; The stylized image is generated by fusing the local stylized image and the masked occlusion area.
12. The generation method according to claim 10 or 11, characterized in that, The step of performing stylization transfer on the mask image based on the original image, the mask image, and the second stylization parameters to generate the stylized image includes: The original image, the mask image, and the second stylization parameters are input into a preset stylization transfer model to perform stylization transfer, thereby generating the stylized image.
13. An apparatus for generating stylized images, characterized in that, include: The recognition module is used to identify face regions in the original image to be styled and transferred; The setting module is used to set a first stylization parameter and a second stylization parameter based on the area ratio of the face region in the original image. The first stylization parameter is used to characterize the degree of occlusion of the face region by the mask, and the second stylization parameter is used to characterize the degree of difference between the original image and the stylized image. The stylization transfer module is used to perform stylization transfer on the original image based on the first stylization parameter and the second stylization parameter to generate a stylized image.
14. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1-12.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-12.
16. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in claims 1-12.