Image generation method, device and storage medium
By aligning the basic object and reference object images and fusion of style vectors to generate fusion images, the problem of inefficient drawing and drawing of design details in the new design is solved, and efficient local design details migration and image generation are achieved.
Patent Information
- Application Number
- CN202111162116.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In the fields of clothing design, designers need to manually draw and draw local design details of existing styles when designing new styles, resulting in low editing efficiency and inability to achieve large-scale batch design.
By aligning the basic object image and the reference object image, the style vectors of the design key points area and the non-design key points area are extracted, and the fusion image is fused to generate a fusion image, so as to automatically migrate the structural features of the reference object design key points area to the basic object.
This improves image generation efficiency and enables designers to quickly and efficiently migrate local design details of reference objects to basic objects, which not only retains the overall design style of the basic objects, but also integrates local design details of reference objects.
Smart Images

Figure CN114119348B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image generation method, device, and storage medium. Background Art
[0002] In fields such as the fashion design field and the home decoration design field, designers often have a need to draw on local designs from existing designs in new designs.
[0003] For example, in the process of fashion design, designers usually need to draw on some local designs of existing styles to create new styles. However, currently, in the process of image editing for new fashion design, designers refer to some local design details of existing styles and manually draw them into the new design, so that some local detail features of the existing style are incorporated into the new design being drawn. In this way, large-scale batch design cannot be achieved, and the editing efficiency is very low. Summary of the Invention
[0004] Embodiments of the present invention provide an image generation method, device, and storage medium to improve the efficiency of image generation.
[0005] In a first aspect, an embodiment of the present invention provides an image generation method, the method including:
[0006] Align a base object image and a reference object image, and obtain a design key point area of the reference object image and a non-design key point area of the base object image;
[0007] Extract a first set of style vectors corresponding to the base object image;
[0008] Extract a second set of style vectors according to the design key point area of the reference object image and the non-design key point area of the base object image;
[0009] Generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the design key point area of the reference object image are combined in the fused image.
[0010] In a second aspect, an embodiment of the present invention provides an image generation device, the device including:
[0011] An acquisition module, configured to align a base object image and a reference object image, and obtain a design key point area of the reference object image and a non-design key point area of the base object image;
[0012] An extraction module, configured to extract a first set of style vectors corresponding to the base object image; and extract a second set of style vectors according to the key design area of the reference object image and the non-key design area of the base object image.
[0013] A generation module, configured to generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, where the fused image combines the structural features of the key design area of the reference object image.
[0014] In a third aspect, an embodiment of the present invention provides an electronic device, including: an input device, a processor, and a display screen;
[0015] The input device, coupled to the processor and the display screen, is configured to input a base object image and a reference object image;
[0016] The processor is configured to align the base object image and the reference object image, and obtain the key design area of the reference object image and the non-key design area of the base object image; extract a first set of style vectors corresponding to the base object image; extract a second set of style vectors according to the key design area of the reference object image and the non-key design area of the base object image; generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, where the fused image combines the structural features of the key design area of the reference object image;
[0017] The display screen is configured to display the base object image, the reference object image, and the fused image.
[0018] In a fourth aspect, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the image generation method as described in the first aspect.
[0019] In a fifth aspect, an embodiment of the present invention provides an image generation method, the method including:
[0020] In response to a request from a user device to call an image generation service interface, use the processing resources corresponding to the image generation service interface to perform the following steps:
[0021] Align the base object image and the reference object image, and obtain the key design area of the reference object image and the non-key design area of the base object image; extract the first set of style vectors corresponding to the base object image; according to the key design area of the reference object image and the non-key design area of the base object image, extract the second set of style vectors; generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the key design area of the reference object image are combined in the fused image.
[0022] In a sixth aspect, an embodiment of the present invention provides an image generation method, and the method includes:
[0023] Align the base object image and the reference object image, and respectively obtain the key design area and the non-key design area of the base object image and the reference object image;
[0024] Extract the first set of style vectors corresponding to the non-key design area in the base object image;
[0025] Extract the second set of style vectors corresponding to the key design area in the reference object image;
[0026] Generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the key design area of the reference object image are combined in the fused image.
[0027] In the solution provided by the embodiment of the present invention, it is assumed that the user wants to perform image editing on a certain base object, and during the editing process, wants to transfer the structural features (such as shape, contour, pattern, fold, etc.) of the key design area of the reference object to the base object. For this purpose, first, obtain the base object image and the reference object image and perform alignment processing on the two, and moreover, obtain the key design area of the reference object image and the non-key design area of the base object image. Then, extract the first set of style vectors corresponding to the base object image, and combine the key design area of the reference object image and the non-key design area of the base object image to extract the second set of style vectors. In this way, the second set of style vectors includes the style vectors of the key design area in the reference object image and the style vectors of the non-key design area in the base object image. Finally, fuse the first set of style vectors and the second set of style vectors to generate a fused image as the editing result according to the fused third set of style vectors, so that the purpose of automatically and efficiently transferring the structural features of the key design area of the reference object to the base object can be achieved, and the final editing result not only fuses the local design details of the reference object but also retains the original overall design style of the base object. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0029] Figure 1 Flowchart of an image generation method provided by an embodiment of the present invention;
[0030] Figure 2 Schematic diagram of a mask image;
[0031] Figure 3a Schematic diagram of hiding non-design key areas in a reference object image;
[0032] Figure 3b Schematic diagram of hiding corresponding areas in a base object image;
[0033] Figure 4 Schematic diagram of a style vector optimization process provided by an embodiment of the present invention;
[0034] Figure 5 Schematic diagram of another style vector optimization process provided by an embodiment of the present invention;
[0035] Figure 6 Schematic diagram of a mixed image;
[0036] Figure 7 Schematic diagram of an edited image;
[0037] Figure 8 Schematic diagram of the execution process of an image generation method provided by an embodiment of the present invention;
[0038] Figure 9 Application schematic diagram of an image generation method provided by an embodiment of the present invention;
[0039] Figure 10 Flowchart of an image generation method provided by an embodiment of the present invention;
[0040] Figure 11 Schematic diagram of the structure of an image generation device provided by an embodiment of the present invention;
[0041] Figure 12 Schematic diagram of the structure of an electronic device provided by this embodiment;
[0042] Figure 13 Schematic diagram of the structure of another electronic device provided by this embodiment. Detailed implementation manners
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] In addition, the sequence of steps in the following method embodiments is only an example and is not strictly limited.
[0045] The image generation method provided by the embodiments of the present invention can be executed by an electronic device, which can be a terminal device such as a PC or a laptop, or a server. The server can be a physical server or a virtual server. The server can be a physical or virtual server on the user side or a cloud server.
[0046] The solution provided by the embodiments of the present invention can be used to generate an image of a basic object with a specific editing effect. Briefly, a user (the user in this article can be a designer with image design requirements) can, after obtaining an image containing a basic object (referred to as a basic object image) and an image containing a reference object (referred to as a reference object image) and marking the design key point area in the reference object image whose design style is desired to be referred to, implement the solution provided by the embodiments of the present invention to achieve the automatic migration of the structural features presented in the above-mentioned design key point area to the corresponding area in the basic object image. In this way, in the finally generated image, the overall original design style of the basic object is retained, and the local design details of the design key point area in the reference object are integrated.
[0047] The visual features of an image can be divided into three types: low-level, middle-level, and high-level. Among them, the low-level and middle-level can be considered as visual features reflecting the macroscopic level of the image, while the high-level is a visual feature reflecting the microscopic level of the image. The above-mentioned structural features can correspond to the low-level and middle-level features. Taking the clothing scenario as an example, the low-level and middle-level visual features of clothing can include, for example: shape, contour, specific structure, large patterns, etc., and the high-level visual features can include features such as color, texture, and material. Therefore, in the image obtained after automatically editing the basic image, the overall original design style of the basic object is retained, and the local design details of the design key point area in the reference object are integrated. Briefly, it can be understood that the visual features such as color, texture, and material reflected globally in the edited image are the same as those of the basic object, and the structural features of the part corresponding to the design key point area in the edited image are the same as those of the design key point area.
[0048] In different application scenarios, the above-mentioned basic object and reference object will be different. For example, in the clothing design scenario, the basic object and the reference object are two clothing items with different styles, such as two pairs of jeans, two T-shirts, etc. In the home decoration design scenario, the basic object and the reference object can be the decoration design drawings of two living rooms respectively.
[0049] The following will combine the following embodiments to elaborate in detail on the execution process of the image generation method provided by this application.
[0050] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present invention. As Figure 1 shown, it may include the following steps:
[0051] 101. Align the basic object image and the reference object image, and obtain the design key point area of the reference object image and the non-design key point area of the basic object image.
[0052] 102. Extract the first set of style vectors corresponding to the basic object image.
[0053] 103. Extract the second set of style vectors according to the design key point area of the reference object image and the non-design key point area of the basic object image.
[0054] 104. Generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors. The fused image combines the structural features of the design key point area of the reference object image.
[0055] First, obtain the basic object image and the reference object image, and mark the design key point area to be referred to in the reference object image. Among them, the basic object included in the basic object image is the object that needs to be further edited to make it present a specific effect. In practical applications, the basic object image can be an image directly taken of the basic object, or an image designed by the user independently; the reference object image is an image that has been edited previously and contains the reference object, or it can also be an image directly taken of the reference object. For the convenience of subsequent image processing by the model, the sizes of the basic object image and the reference object image need to be adjusted to be the same.
[0056] In practical applications, the basic object and the reference object may have different shapes. To ensure the accuracy of subsequent processing, in the image generation solution provided by the embodiment of the present invention, first align the basic object image and the reference object image, that is, align the reference object image to the basic object image.
[0057] Generally speaking, the image alignment process is as follows: Detect the feature points of the basic object image and the reference object image respectively, and align the basic object image and the reference object image according to the detected feature points.
[0058] Optionally, the process of feature point detection can be specifically implemented as follows:
[0059] Perform sparse feature point detection on the base object image and the reference object image respectively to obtain multiple feature points;
[0060] Perform edge detection on the base object image and the reference object image respectively to obtain the edge of the base object and the edge of the reference object;
[0061] Sample multiple feature points on the edge of the base object and the edge of the reference object respectively.
[0062] That is to say, first, a certain feature point detection model (such as a model trained on datasets such as DeepFashionv1 that can perform sparse feature point detection) can be used to perform sparse feature point detection on the base object image and the reference object image respectively to obtain multiple feature points. At this time, the number of feature points obtained is small. To ensure the image alignment effect, densification processing of the feature points can also be performed. The densification of feature points can be achieved through edge extraction of the image. Therefore, perform edge detection on the base object image and the reference object image respectively to obtain the edge of the base object and the edge of the reference object, and then sample multiple feature points along the edge of the base object and the edge of the reference object to complete the densification processing of the feature points.
[0063] It should be noted that in a certain application scenario, the base object and the reference object are of the same type, such as both being jeans. The numbers of feature points corresponding to different positions on a certain object can be preset. For example, the numbers of feature points from the left vertex of the waistband to the right vertex of the waistband are 1 to 8 in sequence, the number of the feature point at the crotch position is 9, and the numbers of feature points from top to bottom on the outside of the left trouser leg are 10 to 40 in sequence, and so on. Based on this numbering rule, after determining the multiple feature points respectively included in the above-mentioned base object image and reference object image, the numbers corresponding to each feature point can be determined respectively. The feature points with the same number on the two images are a pair of feature points with a corresponding relationship.
[0064] After that, according to the corresponding relationship between the feature points on the base object image and the reference object image, the transformation parameters can be determined, so as to transform the reference object image according to the transformation parameters to achieve alignment with the base object image.
[0065] Through the above-mentioned feature point detection and image alignment processing, even if the shapes of the base object and the reference object are different, the style vectors of the corresponding two images can be fused.
[0066] To generate an image that retains the original design style of the base object while incorporating the local structural features of the design key points area of the reference object, on the one hand, it is necessary to extract the style vector (which can also be called the latent vector) of the base object image to obtain the first set of style vectors; on the other hand, it is also necessary to extract the second set of style vectors that include the style vectors of the design key points area of the reference object and the style vectors of the non-design key points area of the base object image. Among them, since the base object image and the reference object image are images of the same size, after marking the design key points area in the reference object image, it can be understood that the image area in the base object image corresponding to this design key points area can be regarded as the design key points area of the base object image, and the other image areas are called the non-design key points areas in the base object image.
[0067] In practical applications, optionally, a style-based generative network model (style GAN) or its improved neural network model can be used to extract the style vectors. This network model can extract multiple (such as 18) levels of style vectors of the input image.
[0068] Input the base object image into a style GAN model, and the multiple levels of style vectors output by it are called the first set of style vectors (i.e., multiple latent vectors).
[0069] In fact, the input and output of the style GAN model can be interchanged, that is, after inputting an image into the style GAN model, the style GAN model can extract multiple levels of style vectors of the input image, and after inputting multiple levels of style vectors into the style GAN model, the style GAN model can output the corresponding generated image. Therefore, the style vectors can affect the visual features of the generated image. By extracting the above two sets of style vectors and fusing these two sets of style vectors, a generated image that realizes the above local design style transfer effect can be generated based on the fused style vectors - which is called the fused image in the embodiments of the present invention.
[0070] To achieve the extraction of the above second set of style vectors, it is necessary to first obtain the design key points area of the reference object image and the non-design key points area of the base object image, and the obtaining step can be realized by means of the mask image corresponding to the reference object image. Based on the base object image and the reference object image, a mask image corresponding to the reference object image can be generated based on the marked design key points area in the reference object image.
[0071] For ease of understanding, in combination with Figure 2 Exemplarily illustrate the meaning of the mask image. Such as Figure 2As shown in the figure, first, the mask image is a binary black and white image, and the size of the mask image is equal to the size of the reference object image. Suppose the reference object image is an image containing a pair of pants as shown schematically in the figure. The user selects the area of the trouser leg in this image as the design key area. Then, in the mask image, the white pixel area (RGB=(1,1,1)) corresponds to this design key area, while the remaining black pixel area (RGB=(0,0,0)) corresponds to the non-design key area in the reference object image, that is, the area other than the design key area. The generation method of the mask image can be implemented with reference to existing related technologies and will not be elaborated here.
[0072] Obtain the design key area of the reference object image and the non-design key area of the base object image. Specifically, it can be based on the mask image to obtain a first image containing the design key area of the reference object image and a second image containing the non-design key area of the base object image. Then, based on the first image and the second image, a second set of style vectors is extracted.
[0073] It should be noted that the sizes of the first image and the second image here are equal to the sizes of the reference object image and the base object image. Simply put, the first image is an image formed by hiding the non-design key area in the reference object image and only exposing the design key area. The second image is an image formed by hiding the design key area in the base object image and only exposing the non-design key area.
[0074] Specifically, the mask image can be multiplied by the reference object image to hide the non-design key area of the reference object image, so as to obtain the first image containing the design key area of the reference object image. By multiplying the XOR result of the mask image by the base object image to hide the design key area of the base object image, the second image containing the non-design key area of the base object image is obtained. Among them, the mask image has the same size as the reference object image and the base object image.
[0075] Multiplying the mask image by the reference object image means multiplying the RGB values of the corresponding pixels in the two images. Since the mask image only includes a white image area with RGB=(1,1,1) and a black image area with RGB=(0,0,0), the result of multiplying the reference object image by the mask image is that the image area corresponding to the white image area in the reference object image (i.e., the design key area) remains unchanged, but the image area corresponding to the black image area in the reference object area (i.e., the non-design key area) is set to black, thus achieving the purpose of hiding the non-design key area in the reference object image and only exposing the design key area, as Figure 3a shown.
[0076] If the masked image is represented as M, the XOR result of the masked image can be expressed as: 1 - M. Simply put, the XOR result of the masked image is equivalent to subtracting the RGB values of each pixel in it from 1 respectively, and the result is that the pixels originally white in the masked image are set to black, and the pixels originally black are set to white.
[0077] The process of multiplying the XOR result of the masked image by the base object image refers to the above-mentioned process of multiplying the masked image by the reference object image, which will not be elaborated here. The multiplication result is that the pixels in the design key point area of the base object image are all set to black, and the pixels in the non-design key point area are all set to white, so as to achieve the purpose of hiding the design key point area and only exposing the non-design key point area, as Figure 3b shown.
[0078] The process of extracting the second set of style vectors based on the above-mentioned first image and second image can be implemented as:
[0079] Iteratively execute the following optimization process of the second set of style vectors until the second set of style vectors that meet the requirements are obtained:
[0080] Input the first image and the mixed image into the classification model to determine the first loss function value through the classification model; wherein, the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity of the two input images;
[0081] Input the second image and the mixed image into the classification model to determine the second loss function value through the classification model;
[0082] Combine the first loss function value and the second loss function value to optimize the second set of style vectors.
[0083] As can be seen from the above description, in the embodiments of the present invention, the second set of style vectors need to be continuously optimized to obtain the final second set of style vectors that meet the conditions. Initially, the second set of style vectors can be initialized to multiple random vector values.
[0084] For ease of understanding, in combination with Figure 4 to exemplarily illustrate the above-mentioned optimization process of the second set of style vectors.
[0085] The first round of iteration process: In the initial state, a set of style vectors (such as 18 style vectors) can be randomly initialized and input into the generation network model shown in the figure, such as the styleGAN model mentioned above, so as to output an image through this generation network model, which is called a mixed image. Then, the first image and this mixed image are used as a pair of inputs and input into the classification model. The role of this classification model is to determine whether the two input images are of the same class. In the case where the current inputs are the mixed image and the first image, it is to determine whether the mixed image is of the same class as the first image. The way to measure whether two input images are of the same class is: calculate the perceptual similarity between the two input images. If the perceptual similarity between them is very high, it is considered to be of the same class. Therefore, the loss function corresponding to the classification model is used to measure the perceptual similarity between the two input images. In Figure 4 it, assume that the loss function value between the mixed image output by the classification model and the first image is loss1. Similarly, the second image and this mixed image are also used as a pair of inputs and input into the classification model. In Figure 4 it, assume that the loss function value between the mixed image output by the classification model and the second image is loss2. The sum of the two loss function values is denoted as the total loss. Based on this total loss, the second set of style vectors is optimized through the backpropagation process.
[0086] The process of the second round of iteration: The second set of style vectors optimized through the first round of iteration process is input into the generation network model to obtain a second mixed image. Similarly, the above-mentioned first image and this second mixed image are used as a pair of inputs and input into the classification model. The classification model outputs the corresponding loss function value, which is assumed to be denoted as loss1'. The second image and this second mixed image are used as a pair of inputs and input into the classification model. The classification model outputs the corresponding loss function value, which is assumed to be denoted as loss2'. The sum of the two loss function values is denoted as the total loss'. Based on this total loss', the second set of style vectors is further optimized through the backpropagation process.
[0087] And so on. Assume that when the nth round of iteration is executed, if the total loss function value output in this round is less than the set threshold, it is considered that the iteration ends, and the second set of style vectors used in the nth round is used as the finally optimized second set of style vectors for the subsequent style vector fusion process.
[0088] In practical applications, optionally, the above classification model can be, for example, a VGG model, but not limited to this. The loss function of the classification model can be, for example, the LPIPS loss that measures high-level perceptual consistency.
[0089] In an alternative embodiment, in order to further improve the quality of the second set of style vectors, two different types of loss functions can be used for constraint, including the loss for measuring the high-level perceptual consistency and the per-pixel mean squared error loss for low-level texture consistency.
[0090] Based on this, the process of extracting the second set of style vectors based on the first image and the second image can also be implemented as follows:
[0091] Iteratively execute the following optimization process of the second set of style vectors until a satisfactory second set of style vectors is obtained:
[0092] Input the first image and the mixed image into the classification model to determine the first loss function value through the classification model; wherein, the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity of the two input images;
[0093] Input the second image and the mixed image into the classification model to determine the second loss function value through the classification model;
[0094] Perform pixel comparison between the first image and the mixed image to determine the third loss function value; and perform pixel comparison between the second image and the mixed image to determine the fourth loss function value;
[0095] Optimize the second set of style vectors by combining the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value.
[0096] For ease of understanding, in combination with Figure 5 to exemplarily illustrate the above optimization process of the second set of style vectors.
[0097] The first round of iteration process: In the initial state, a set of style vectors (such as 18 style vectors) can be randomly initialized and input into the generation network model shown in the figure. An image is output through this generation network model, which is called a mixed image. Then, the first image and this mixed image are used as a pair of inputs and input into the classification model. Suppose the loss function value between the mixed image output by the classification model and the first image is loss1. Similarly, the second image and this mixed image are also used as a pair of inputs and input into the classification model. Suppose the loss function value between the mixed image output by the classification model and the second image is loss2. In addition, the first image and this mixed image are used as a pair of inputs and input into the pixel contrast module. The pixel contrast module compares the two input images pixel by pixel and determines the corresponding loss function value based on the set loss function and the pixel contrast result, denoted as loss3. Similarly, the second image and this mixed image are used as a pair of inputs and input into the pixel contrast module. Suppose the loss function value output by the pixel contrast module is loss4. The sum of the four loss function values is denoted as the total loss. Based on this total loss, the second set of style vectors is optimized through the backpropagation process.
[0098] The process of the second round of iteration: The second set of style vectors optimized through the first round of iteration process is input into the generation network model to obtain the second mixed image. Similarly, the first image and this second mixed image are used as a pair of inputs and input into the classification model. The classification model outputs the corresponding loss function value, supposed to be denoted as loss1'. The second image and this second mixed image are used as a pair of inputs and input into the classification model. The classification model outputs the corresponding loss function value, supposed to be denoted as loss2'. In addition, the first image and this second mixed image are used as a pair of inputs and input into the pixel contrast module. Suppose the loss function value output by the pixel contrast module is loss3'. Similarly, the second image and this second mixed image are used as a pair of inputs and input into the pixel contrast module. Suppose the loss function value output by the pixel contrast module is loss4'. The sum of the four loss function values in this round is denoted as the total loss'. Based on this total loss', the second set of style vectors is further optimized through the backpropagation process.
[0099] And so on. Suppose when the nth round of iteration is executed, the total loss function value output in this round is less than the set threshold, then it is considered that the iteration ends, and the second set of style vectors used in the nth round is used as the finally optimized second set of style vectors for the subsequent style vector fusion process.
[0100] It can be understood that the above process of continuously optimizing the second set of style vectors is to ensure that the mixed image generated based on the optimized second set of style vectors can be more similar to the first image and the second image. The design key point area in the reference object image is exposed in the first image, and the part of the non-design key point area in the basic object image is exposed in the second image. Therefore, to achieve the above-mentioned similarity purpose, by continuously optimizing the second set of style vectors used to generate the mixed image, finally, the optimized second set of style vectors will contain the style vectors of the design key point area in the reference object image and the style vectors of the non-design key point area in the basic object image. The mixed image generated based on the optimized second set of style vectors will have the following characteristics: some image areas present style features very similar to the non-design key point areas in the basic object image, while other areas present style features very similar to the design key point areas in the reference object image. For ease of understanding, combined with Figure 6 is used for exemplary illustration.
[0101] As Figure 6 shown in, assuming that the basic object is a pair of blue vertical-striped pants and the reference object is a pair of black plaid pants, then in the mixed image generated based on the optimized second set of style vectors, in the image area corresponding to the design key point area in the reference object image, the style features of the design key point area of the reference object are still presented: black plaid; while in the remaining image area, the style features of the non-design key point area in the basic object image are presented: blue vertical stripes.
[0102] After obtaining the optimized second set of style vectors that contain the style vectors of the design key point area in the reference object image and the style vectors of the non-design key point area in the basic object image, and obtaining the first set of style vectors obtained by performing style feature extraction on the basic object image, the first set of style vectors and the second set of style vectors can be fused to generate a fused image as the editing result of the basic object image according to the fused third set of style vectors. In this way, the visual effect of migrating the structural features of the design key point area of the reference object to the corresponding position on the basic object will be presented in the fused image.
[0103] Specifically, as described above, a set of style vectors consists of multiple style vectors. For example, both the first set of style vectors and the second set of style vectors include 18 levels of style vectors. During the fusion process, the fusion levels and fusion ratios of the multiple levels of style vectors contained in the first set of style vectors and the second set of style vectors can be determined first, and then the style vectors at the corresponding fusion levels can be fused according to the fusion ratio.
[0104] In practical applications, a large number of experiments can be conducted in advance to determine the fusion level and fusion ratio that suit the current application scenario. Among them, the fusion level refers to which levels of style vectors among all multiple levels of style vectors are to be fused, which may be all or part. Simply put, the fusion ratio is the fusion weight, that is, the weight value used when fusing a style vector at a certain level in the first group of style vectors with the corresponding style vector at the same level in the second group of style vectors. During the experimental stage or the training stage, the fusion level and fusion ratio can be continuously adjusted until the determined fusion level and fusion ratio enable the finally generated image to present the expected visual effect. Among them, the finally generated image refers to the image generated based on the third group of style vectors obtained by fusing two groups of style vectors using the determined fusion level and fusion ratio.
[0105] It can be understood that for a certain application scenario, the experimentally obtained fusion level and fusion ratio can be applied to different input situations in that application scenario. For example, in the design scenario of jeans, the two pairs of jeans used as the base object and the reference object in this input can adopt the same fusion ratio and fusion level as the other two pairs of jeans used as the base object and the reference object in the next input. Therefore, the fusion level and fusion ratio can be set in advance for different application scenarios, and for the current application scenario, the selected fusion level and fusion ratio can be adopted.
[0106] For ease of understanding, combined with Figure 7 an exemplary illustration of the fusion process of style vectors and the image generated based on the fused style vectors is given.
[0107] In Figure 7 it is still assumed that the base object and the reference object are the Figure 6 situation shown in Figure 7 As shown in Figure 7 the first group of style vectors and the second group of style vectors are fused based on the determined fusion ratio and fusion level, and the fusion result is called the third group of style vectors. The third group of style vectors can be input into a generation network model shown in the figure to obtain the output image, which is the fused image, that is, the image obtained after automatically editing the original base object image. As shown in Figure 7 in the fused image generated based on a group of fused style vectors, only the relatively coarse-grained structural features in the design key point area of the reference object are migrated to the corresponding positions of the base object, while fine-grained features such as color and texture are not migrated, so that the original design style of the base object is still retained in the fused image, and at the same time, the local design details of the reference object are incorporated.
[0108] In summary, based on the solution provided by the embodiments of the present invention, when a user needs to refer to the partial design of a reference object to perform editing processing on an image containing a basic object for a new design, the user only needs to input the basic object image and the reference object image with the design key area marked, and then can call the solution provided by the embodiments of the present invention to implement automated image editing processing, obtaining an editing result that not only retains the original design style of the basic object but also incorporates the partial design details of the reference object, thereby helping to improve the design processing efficiency.
[0109] In the above embodiments, the execution processes of each step are respectively introduced exemplarily. For the convenience of comprehensively understanding the processing logic of the solution provided by the embodiments of the present invention, the overall execution process is exemplarily described in combination with Figure 8 to exemplarily illustrate the overall execution process.
[0110] As Figure 8 shown, in practical applications, a user can load the basic object image and the reference object image through a terminal device such as a PC or a laptop, mark the design key area in the reference object image, and then call the image generation method provided by the embodiments of the present invention. The image generation method provided by the embodiments of the present invention can be provided by a certain application program. Then the user can start the application program, complete the loading of the above two images on the image loading interface, and then start the execution of the subsequent image generation process.
[0111] In combination with Figure 8 , as described above, a corresponding mask image can be first generated based on the reference object image marked with the design key area, and then the extraction and fusion of the above first group of style vectors and the second group of style vectors are respectively performed, and a fused image as the editing result is generated based on the fused style vectors.
[0112] As described above, the image generation method provided by the present invention can be executed in the cloud. Several computing nodes can be deployed in the cloud, and each computing node has processing resources such as computing and storage. In the cloud, a certain service can be organized to be provided by multiple computing nodes. Of course, a single computing node can also provide one or more services. The way for the cloud to provide the service can be to provide a service interface externally, and the user calls the service interface to use the corresponding service. The service interface includes forms such as a Software Development Kit (SDK) and an Application Programming Interface (API).
[0113] For the solution provided in the embodiments of the present invention, the cloud can provide a service interface for an image generation service. The user calls the image generation service interface through a user device to trigger a request to call the image generation service interface to the cloud. The cloud determines a computing node that responds to the request, and uses the processing resources in the computing node to perform the following steps:
[0114] Align the basic object image and the reference object image, and obtain the design key point area of the reference object image and the non-design key point area of the basic object image;
[0115] Extract the first set of style vectors corresponding to the basic object image;
[0116] According to the design key point area of the reference object image and the non-design key point area of the basic object image, extract the second set of style vectors;
[0117] Generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the design key point area of the reference object image are combined in the fused image.
[0118] The detailed process of the image generation service interface using the processing resources to perform image generation processing can refer to the relevant descriptions in the foregoing other embodiments, and will not be elaborated here.
[0119] In practical applications, the above request may directly carry a basic object image and a reference object image marked with a design key point area, and the cloud parses the corresponding images from the request. Furthermore, the cloud generates a mask image and performs subsequent steps. In addition, the cloud can pre-store the models required for the execution of the above steps to call these models to complete the corresponding processing.
[0120] For ease of understanding, in combination with Figure 9 to illustrate by way of example. The user can call the image generation service interface through the user device E1 shown in Figure 9 and upload a service request including a basic object image and a reference object image marked with a design key point area through the interface. In the cloud, as shown in the figure, in addition to deploying a number of computing nodes, a management node E2 running a control service is also deployed. After receiving the service request sent by the user device E1, the management node E2 determines a computing node E3 that responds to the service request. After receiving these two images, the computing node E3 performs steps such as mask image generation, style vector extraction, and fusion, and finally outputs a fused image. The detailed execution process refers to the introduction in the foregoing embodiments and will not be elaborated here. After that, the computing node E3 sends the fused image to the user device E1, and the user device E1 displays the fused image, and the user can perform further editing and other operations based on this.
[0121] Figure 10 The flowchart of an image generation method provided by an embodiment of the present invention is shown as Figure 10 follows, and may include the following steps:
[0122] 1001. Align the base object image and the reference object image, and respectively obtain the design key area and the non-design key area of the base object image and the reference object image.
[0123] 1002. Extract the first set of style vectors corresponding to the non-design key area in the base object image.
[0124] 1003. Extract the second set of style vectors corresponding to the design key area in the reference object image.
[0125] 1004. Generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the design key area of the reference object image are combined in the fused image.
[0126] In this embodiment, for the process of aligning the base object image and the reference object image, reference can be made to the relevant descriptions in the foregoing other embodiments, which will not be elaborated herein.
[0127] In this embodiment, for the process of respectively obtaining the design key area and the non-design key area of the base object image and the reference object image, it can be implemented by means of the mask image introduced in the foregoing embodiments. The specific implementation process can refer to the relevant descriptions in the foregoing embodiments, which will not be elaborated herein. Here, it is assumed that the image obtained only including the non-design key area of the base object image is called image A, and the image obtained only including the design key area of the reference object image is called image B.
[0128] Then, extracting style vectors for image A can obtain the above-mentioned first set of style vectors, and extracting style vectors for image B can obtain the above-mentioned second set of style vectors.
[0129] In this embodiment, a set of style vectors corresponding to the non-design key area of the base object image is used as a set of style vectors corresponding to the base object image, and a set of style vectors corresponding to the design key area of the reference object image is used as a set of style vectors corresponding to the reference object image.
[0130] Fusing the above two sets of style vectors will make the fused set of style vectors contain both the style of the non-design key area of the base object image and the style of the design key area of the reference object image. Then, the fused image generated based on the fused style vectors will contain both the structural features of the base object and the structural features of the design key area of the reference object image.
[0131] The image generation device of one or more embodiments of the present invention will be described in detail below. Those skilled in the art can understand that these devices can all be configured by using commercially available hardware components through the steps taught by this solution.
[0132] Figure 11 The following is a schematic structural diagram of an image generation device provided by an embodiment of the present invention, as Figure 12 shown, the device includes: an acquisition module 11, an extraction module 12, and a generation module 13.
[0133] The acquisition module 11 is used to align the basic object image and the reference object image, and acquire the design key point area of the reference object image and the non-design key point area of the basic object image.
[0134] The extraction module 12 is used to extract the first set of style vectors corresponding to the basic object image; according to the design key point area of the reference object image and the non-design key point area of the basic object image, extract the second set of style vectors.
[0135] The generation module 13 is used to generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, and the structural features of the design key point area of the reference object image are combined in the fused image.
[0136] Optionally, in the process of acquiring the design key point area of the reference object image and the non-design key point area of the basic object image, the acquisition module 11 is specifically used to: acquire the mask image of the reference object image, and the mask image is generated based on the marked design key point area of the reference object image; based on the mask image, acquire a first image including the design key point area of the reference object image and a second image including the non-design key point area of the basic object image.
[0137] Optionally, the obtaining module 11 is specifically configured to: multiply the mask image by the reference object image to obtain a first image including the key design point area of the reference object image; multiply the XOR result of the mask image by the base object image to obtain a second image including the non-key design point area of the base object image. Optionally, in the process of extracting the second set of style vectors, the extraction module 12 is specifically configured to: iteratively execute the following optimization process of the second set of style vectors until a second set of style vectors that meet the requirements is obtained: input the first image and the mixed image into a classification model to determine a first loss function value through the classification model; wherein, the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity between two input images; input the second image and the mixed image into the classification model to determine a second loss function value through the classification model; combine the first loss function value and the second loss function value to optimize the second set of style vectors.
[0138] Optionally, in the optimization process of the second set of style vectors, the extraction module 12 is further configured to: perform pixel comparison on the first image and the mixed image to determine a third loss function value; and perform pixel comparison on the second image and the mixed image to determine a fourth loss function value; combine the first loss function value, the second loss function value, the third loss function value and the fourth loss function value to optimize the second set of style vectors.
[0139] Optionally, in the process of fusing the first set of style vectors and the second set of style vectors, the generating module 13 is specifically configured to: determine the fusion level and fusion ratio of the multi-layer style vectors included in the first set of style vectors and the second set of style vectors respectively; fuse the style vectors at the corresponding fusion levels according to the fusion ratio.
[0140] Optionally, in the process of aligning the base object image and the reference object image, the obtaining module 11 is configured to: respectively perform feature point detection on the base object image and the reference object image; align the base object image and the reference object image according to the detected feature points.
[0141] Wherein, in the process of feature point detection, the obtaining module 11 is specifically configured to: respectively perform sparse feature point detection on the base object image and the reference object image to obtain a plurality of feature points; respectively perform edge detection on the base object image and the reference object image to obtain the edge of the base object and the edge of the reference object; respectively sample a plurality of feature points on the edge of the base object and the edge of the reference object.
[0142] Among them, during the image alignment process, the obtaining module 11 is specifically configured to: determine transformation parameters according to the corresponding relationship between the feature points on the base object image and the reference object image; and transform the reference object image according to the transformation parameters to align it with the base object image.
[0143] Figure 11 The device shown can execute the steps described in the foregoing embodiments. For the detailed execution process and technical effects, refer to the descriptions in the foregoing embodiments, which will not be elaborated herein.
[0144] In a possible design, the above Figure 11 structure of the image generating device shown can be implemented as an electronic device, such as Figure 12 shown, the electronic device may include: an input device 21, a processor 22, and a display screen 23.
[0145] The input device 21 is coupled to the processor 22 and the display screen 23, and is used to input a base object image and a reference object image.
[0146] The processor 22 is configured to align the base object image and the reference object image, and obtain the key design area of the reference object image and the non-key design area of the base object image; extract a first set of style vectors corresponding to the base object image; extract a second set of style vectors according to the key design area of the reference object image and the non-key design area of the base object image; and generate a fused image based on the fusion of the first set of style vectors and the second set of style vectors, where the fused image combines the structural features of the key design area of the reference object image.
[0147] The display screen 23 is used to display the base object image, the reference object image, and the fused image.
[0148] Figure 13 This is a schematic structural diagram of another electronic device provided in this embodiment. As Figure 13 shown, the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0149] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the method steps 101 - step 105 described above. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0150] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read - only memory (EEPROM), erasable programmable read - only memory (EPROM), programmable read - only memory (PROM), read - only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0151] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0152] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front - facing camera and / or a rear - facing camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front - facing camera and / or the rear - facing camera can receive external multimedia data. Each front - facing camera and rear - facing camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0153] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0154] The input / output interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0155] The sensor component 814 includes one or more sensors for providing an assessment of various aspects of the state of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0156] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0157] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0158] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0159] In addition, an embodiment of the present invention provides a non-transitory machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the image generation method provided in the foregoing embodiment.
[0160] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0161] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of adding a necessary general hardware platform, and of course, can also be implemented by a combination of hardware and software. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention may be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An image generation method, characterized in that, Including: Align the base object image and the reference object image, and obtain the design key point area of the reference object image and the non-design key point area of the base object image, where the base object in the base object image has design key points; Extract the first set of style vectors corresponding to the base object image, and the first set of style vectors reflects the design style of the base object image; According to the design key point area of the reference object image and the non-design key point area of the base object image, extract the second set of style vectors, and the second set of style vectors reflects the design style of the non-design key point area and the design details of the design key point area; Iteratively execute the following optimization process of the second set of style vectors until the second set of style vectors that meet the requirements are obtained: Input the first image including the design key point area of the reference object image and the mixed image into the classification model to determine the first loss function value through the classification model, where the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity of two input images; Input the second image including the non-design key point area of the base object image and the mixed image into the classification model to determine the second loss function value through the classification model; Perform pixel comparison on the first image and the mixed image to determine the third loss function value; Perform pixel comparison on the second image and the mixed image to determine the fourth loss function value; Optimize the second set of style vectors by combining the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value; Generate a fused image based on the fusion of the first set of style vectors and the optimized second set of style vectors, and the structural features of the design key point area of the reference object image are combined in the fused image.
2. The method according to claim 1, characterized in that, The obtaining of the design key point area of the reference object image and the non-design key point area of the base object image includes: Obtain the mask image of the reference object image, and the mask image is generated based on the marked design key point area of the reference object image; Based on the mask image, obtain the first image including the design key point area of the reference object image and the second image including the non-design key point area of the base object image.
3. The method according to claim 2, characterized in that, The obtaining of the first image including the design key point area of the reference object image and the second image including the non-design key point area of the base object image based on the mask image includes: Multiply the mask image by the reference object image to obtain the first image including the design key point area of the reference object image; Multiply the exclusive OR result of the mask image by the base object image to obtain the second image including the non-design key point area of the base object image.
4. The method according to claim 1, characterized in that, The aligning of the base object image and the reference object image includes: Perform feature point detection on the base object image and the reference object image respectively; Align the base object image and the reference object image according to the detected feature points.
5. The method according to claim 4, characterized in that, The performing feature point detection on the base object image and the reference object image respectively includes: Performing sparse feature point detection on the base object image and the reference object image respectively to obtain a plurality of feature points; Performing edge detection on the base object image and the reference object image respectively to obtain the edge of the base object and the edge of the reference object; Sampling a plurality of feature points on the edge of the base object and the edge of the reference object respectively.
6. The method according to claim 5, characterized in that, The aligning the base object image and the reference object image according to the detected feature points includes: Determining transformation parameters according to the correspondence between the feature points on the base object image and the reference object image; Transforming the reference object image according to the transformation parameters to align it with the base object image.
7. The method according to claim 1, characterized in that, Fusing the first set of style vectors and the second set of style vectors includes: Determining the fusion level and fusion ratio of the multi-layer style vectors included in the first set of style vectors and the second set of style vectors respectively; Fusing the style vectors at the corresponding fusion level according to the fusion ratio.
8. An image generation method, characterized in that, including: Align the base object image and the reference object image, and respectively obtain the non-design key point area of the base object image and the design key point area of the reference object image. The base object in the base object image has design key points; Extracting a first set of style vectors corresponding to the non-design key point area in the base object image; Extracting a second set of style vectors corresponding to the design key point area in the reference object image, and the second set of style vectors reflects the design details of the design key point area; Iteratively execute the following optimization process of the second set of style vectors until a second set of style vectors that meet the requirements is obtained: Inputting a first image including the design key point area of the reference object image and a mixed image into a classification model to determine a first loss function value through the classification model, where the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity of two input images; Performing pixel comparison on the first image and the mixed image to determine a third loss function value; Optimizing the second set of style vectors by combining the first loss function value and the third loss function value; Iteratively execute the following optimization process of the first set of style vectors until a second set of style vectors that meet the requirements is obtained: Inputting a second image including the non-design key point area of the base object image and the mixed image into the classification model to determine a second loss function value; Performing pixel comparison on the second image and the mixed image to determine a fourth loss function value; Optimizing the first set of style vectors by combining the second loss function value and the fourth loss function value; Generating a fused image based on the fusion of the optimized first set of style vectors and the optimized second set of style vectors, and the structural features of the design key point area of the reference object image are combined in the fused image.
9. An electronic device, characterized in that, including: An input device, a processor, and a display screen; The input device, coupled to the processor and the display screen, is configured to input a base object image and a reference object image, wherein the base object in the base object image has design key points; The processor is configured to align the base object image and the reference object image, and obtain the design key point area of the reference object image and the non-design key point area of the base object image; Extract a first set of style vectors corresponding to the base object image, where the first set of style vectors reflects the design style of the base object image; Extract a second set of style vectors according to the design key point area of the reference object image and the non-design key point area of the base object image, where the second set of style vectors reflects the design style of the non-design key point area and the design details of the design key point area; Generate a fused image based on the fusion of the first set of style vectors and the optimized second set of style vectors, where the fused image combines the structural features of the design key point area of the reference object image; Wherein, the following optimization process of the second set of style vectors is iteratively executed until a second set of style vectors that meet the requirements is obtained: Input a first image including the design key point area of the reference object image and a mixed image into a classification model to determine a first loss function value through the classification model, wherein the mixed image is an image generated based on the second set of style vectors to be optimized, and the loss function corresponding to the classification model is used to measure the perceptual similarity between two input images; input a second image including the non-design key point area of the base object image and the mixed image into the classification model to determine a second loss function value through the classification model; Perform pixel comparison between the first image and the mixed image to determine a third loss function value; perform pixel comparison between the second image and the mixed image to determine a fourth loss function value; Optimize the second set of style vectors by combining the first loss function value, the second loss function value, the third loss function value, and the fourth loss function value; The display screen is configured to display the base object image, the reference object image, and the fused image.
10. A non - transitory machine - readable storage medium, characterized in that, The non-transitory machine-readable storage medium stores executable code, which, when executed by a processor of an electronic device, causes the processor to execute the image generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Method, apparatus, and system for face exchange between images, and computer program product
CN110660037A