Image generation method and device and electronic equipment
By generating masked images and performing feature fusion processing, the problem of low image editing efficiency in the prior art is solved, a simplified image local editing process is realized, and the editing efficiency is improved.
Patent Information
- Application Number
- CN202411944097.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, users need to produce accurate masked images for local image editing, resulting in inefficient image editing.
By responding to a position selection instruction of the initial image, a target position is determined and a mask image is generated based on the position. Then, using multi-time step feature fusion and feature denoising processing, a target image is generated, wherein the second area of the target image contains the image content corresponding to the target text.
On the premise of ensuring the editing effect, the process of user-generating masked images is simplified, image editing efficiency is improved, and the complexity of user input is reduced.
Smart Images

Figure CN120070658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image generation method, apparatus, and electronic device. Background Art
[0002] The latent diffusion model can perform local editing on an image; that is, modify the image content within a local area of the image. In related technologies, in order to enable the latent diffusion model to know the local area to be modified, a user needs to provide a mask image indicating the local area; the user needs to create an accurate mask image to make the local editing result output by the latent diffusion model meet the expectations; however, the process of creating an accurate mask image is cumbersome and time-consuming, reducing the image editing efficiency. Summary of the Invention
[0003] In view of this, an object of the present invention is to provide an image generation method, apparatus, and electronic device to improve the image editing efficiency on the premise of ensuring a good editing effect.
[0004] In a first aspect, an embodiment of the present invention provides an image generation method, the method including: responding to a position selection instruction for an initial image, determining a target position on the initial image; generating a mask image corresponding to the initial image based on the target position; where the mask image is used to: indicate a first area to be edited in the initial image; based on the mask image, performing multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until a target image is generated; where a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0005] In a second aspect, an embodiment of the present invention provides an image generation apparatus, the apparatus including: a position determination module, configured to respond to a position selection instruction for an initial image and determine a target position on the initial image; a mask generation module, configured to generate a mask image corresponding to the initial image based on the target position; where the mask image is used to: indicate a first area in the initial image; an image generation module, configured to perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text based on the mask image until a target image is generated; where a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0006] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, and the processor executing the computer-executable instructions to implement the above image generation method.
[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the above image generation method.
[0008] The embodiments of the present invention bring the following beneficial effects:
[0009] For the above image generation method, apparatus and electronic device, in response to a position selection instruction for an initial image, a target position is determined on the initial image; based on the target position, a mask image corresponding to the initial image is generated; wherein the mask image is used to: indicate a first area to be edited in the initial image; based on the mask image, perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until a target image is generated; wherein a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0010] In this way, the user only needs to provide an image position, and a mask image can be generated based on this image position. The mask image indicates the area to be edited, so as to edit the image content indicated by the target text in this area; this way improves the image editing efficiency on the premise of ensuring a good editing effect.
[0011] Other features and advantages of the present invention will be described in the following description, and in part, will be obvious from the description, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the description, claims and drawings.
[0012] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0013] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present invention;
[0015] Figure 2 It is a schematic diagram of an image generation process provided by an embodiment of the present invention;
[0016] Figure 3 An example diagram for generating a target image provided by an embodiment of the present invention;
[0017] Figure 4 Another example diagram for generating a target image provided by an embodiment of the present invention;
[0018] Figure 5 Another example diagram for generating a target image provided by an embodiment of the present invention;
[0019] Figure 6 A schematic diagram of an image generation device provided by an embodiment of the present invention;
[0020] Figure 7 A schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] For ease of understanding, first, the terms related to the embodiments of the present invention are explained.
[0023] 1. Latent Diffusion Models: Abbreviated as LDMs, namely latent diffusion models, which are deep learning models for generating images and other data types. The latent diffusion model changes the input data by gradually adding noise and then learns how to recover the data from the noise. The latent diffusion model can be used to generate new images or modify existing images.
[0024] 2. Mask: That is, a mask image. In image processing, a mask image is usually a black-and-white image used to indicate which parts need to be processed or edited. The white parts usually indicate the parts to be edited, while the black parts indicate the parts to remain unchanged.
[0025] 3. Semantic loss: Semantic loss is a method for measuring the semantic difference between images, usually used to supervise the learning process of a generation model and can also be used to guide the model to generate appropriate image content according to a given text description.
[0026] 4. Gradient Update: Gradient update refers to the process of adjusting model parameters according to the gradient of the loss function with respect to the model parameters during training. It can also be used to adjust the masked image Mask to better match the desired content changes.
[0027] 5. CLIP: Contrastive Language-Image Pre-training, a multi-modal pre-training model; it is a deep learning model used to solve cross-modal tasks, such as associating text and images. The main function of CLIP is to understand the relationship between text and images and be able to generate or retrieve relevant image content given a text description.
[0028] 6. Alpha-CLIP: An enhanced model based on the aforementioned CLIP model. By introducing an additional alpha channel, the model can focus on the user-specified area without changing the image content. Alpha-CLIP can be used to evaluate the results of image editing. The aforementioned CLIP model is mainly used to evaluate the similarity between the generated image and the text description, but the Alpha-CLIP model mainly focuses on the evaluation of the edited area.
[0029] The working process of Alpha-CLIP is as follows: First, extract the edited area. Alpha-CLIP can extract the edited area from the generated image, so that the edited area can be evaluated separately; then, calculate the similarity, specifically calculate the similarity between the edited area and the editing instruction, or rather, calculate the similarity between the masked and marked edited area and the description text. Alpha-CLIP can quantify whether the edited area meets the editing instructions given by the user, thus providing an objective measure of the editing quality.
[0030] In related technologies, when performing local editing on an image, the user needs to provide an accurate masked image to restrict the local editing area.
[0031] For example, in the Blended Diffusion model, the masked image provided by the user is used to be mixed with text-guided noise during the denoising process to generate an image. However, the masked images provided by users in these methods have great limitations because the success of editing highly depends on the precise shape of the masked image, and it is often cumbersome and time-consuming for users to create an accurate Mask.
[0032] There is also a way to describe the target area through text. However, describing the target area through text may be difficult for users to accurately describe the location, and the model may also be unable to accurately understand the area described by the text, resulting in a poor final editing effect.
[0033] In addition, in the related art, it is a relatively heavy burden for users to provide a mask image or describe the area through text, and it limits the flexibility of image editing.
[0034] Based on this, embodiments of the present invention provide an image generation method, apparatus, and electronic device, which can be applied to generate various types of images.
[0035] See Figure 1 An image generation method shown in the figure. The method includes the following steps:
[0036] Step S102, in response to a position selection instruction for the initial image, determine a target position on the initial image;
[0037] Step S104, based on the target position, generate a mask image corresponding to the initial image; wherein, the mask image is used to: indicate a first area to be edited in the initial image;
[0038] Specifically, the position selection instruction can be generated by the user after performing a position selection operation through a terminal device. For example, by means of a mouse or touch to perform a position selection operation to generate the position selection instruction; the position selection operation can be a click or slide operation on the initial image.
[0039] Specifically, when the position selection operation is a click operation, the click position of the click operation can be used as the target position, and the target position can be understood as a position point; when the position selection operation is a slide operation, the slide path of the slide operation can be used as the target position, and the target position includes multiple position points.
[0040] After determining the target position, a new image can be generated on the initial image. This new image has the same size as the initial image; this new image can also be a new layer generated based on the initial image; the target position is recorded through the new image, and specifically, the position coordinates of the target position can be recorded.
[0041] The target position is used to indicate the regional position of the first area in the initial image; in actual implementation, the first area can be formed by expanding around the target position to the surrounding. The first area is recorded through the mask image; in the mask image, the pixel values in the first area are different from the pixel values outside the first area. For example, the pixel value in the first area is 1, while the pixel value outside the first area is 0.
[0042] Step S106, based on the mask image, perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until the target image is generated; wherein, the second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area;
[0043] The regional position of the second region in the target image corresponds to the regional position of the first region indicated in the mask image in the initial image.
[0044] By encoding the initial image through an image encoder, the image features of the initial image can be obtained. By encoding the target text through a text encoder, the text features of the target text can be obtained. The target text is edited and input by the user and is used to indicate the edited content of the first region. For example, the target text is "a skateboard", "a little cow", etc.
[0045] Feature fusion and feature denoising processing can be achieved through a latent diffusion model. For example, Stable Diffusion model, Denoising Diffusion Implicit Models model, etc. The latent diffusion model needs to perform cyclic iterative processing on the features for multiple time steps. After generating the final features, the final features are decoded to obtain the target image. There can be multiple target images, and the user can select the image that meets the requirements from multiple target images.
[0046] It can be understood that the target image is edited based on the initial image. The user determines the target position in the initial image and inputs the target text. Therefore, the image content indicated by the target text is edited in the first region corresponding to the target position, and the edited image content is obtained in the second region of the target image. For example, if the target text is "a little cow", then based on the initial image, the image content in the first region corresponding to the target position selected by the user is modified, and finally, "a little cow" is generated in the second region of the target image. At the same time, the image content outside the second region retains the image content of the initial image outside the first region, thus realizing local editing of the image.
[0047] The above image generation method responds to a position selection instruction for the initial image, determines a target position on the initial image, generates a mask image corresponding to the initial image based on the target position, where the mask image is used to indicate the first region to be edited in the initial image, and based on the mask image, performs multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until the target image is generated. The second region in the target image contains the image content corresponding to the target text, and the region outside the second region contains the image content corresponding to the initial image. The second region corresponds to the first region.
[0048] In this method, the user only needs to provide an image position, and a mask image can be generated based on this image position. The mask image indicates the region to be edited, so that the image content indicated by the target text can be edited in this region. This method improves the image editing efficiency on the premise of ensuring a good editing effect.
[0049] In a specific implementation, in a specified time step among multiple time steps, the first region in the mask image is adjusted based on the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text; wherein, the intermediate image is generated based on the features after feature fusion and feature denoising in the previous time step.
[0050] In this embodiment, the user only provides a target position. Based on this, in the feature processing of multiple time steps in this embodiment, the mask image needs to be continuously adjusted so that the region shape of the first region indicated by the mask image has a high degree of matching with the contour of the image content indicated by the target text.
[0051] The aforementioned specified time step can be all the time steps of multiple time steps or part of the time steps. The multiple time steps are the total steps, and the specified time step can be the first 50% of the time steps of multiple time steps. That is, in the multiple time steps of feature fusion and feature denoising, in the first 50% of the time steps, the first region in the mask image will change continuously, and in the last 50% of the time steps, the first region will no longer change, and mainly further process the features of the image content within the first region.
[0052] In the specified time step, before performing the feature fusion and feature denoising processing of this time step, the mask image is adjusted first. Specifically, if the specified time step is the first time step, the mask image can be adjusted based on the similarity between the initial image and the target text. If the specified time step is a time step after the first time step, the features obtained after feature fusion and feature denoising in the previous time step of this specified time step are decoded to obtain an intermediate image; the similarity between the intermediate image of the previous time step and the target text is calculated through a preset algorithm, and then the first region is adjusted based on this similarity; for example, the region edge of a region with a higher similarity may be expanded, and the region edge of a region with a lower similarity may be shrunk, etc.; after adjusting the region edge, the region area, region shape, etc. of the first region may all change.
[0053] In the specified time step, after the mask image is adjusted, the feature fusion and feature denoising processing of this time step is performed based on the adjusted mask image.
[0054] In the above method, a mask image is generated based on the position provided by the user, and the editing region indicated by the mask image is continuously adjusted in multiple time steps, so that the shape of the editing region is more matched with the image content indicated by the target text, further improving the image editing effect.
[0055] In a specific implementation, a potential energy map corresponding to the initial image is generated with the target position as the center; wherein, in the potential energy map, the closer the pixel position is to the target position, the greater the corresponding potential energy value; based on a preset potential energy threshold, the potential energy map is binarized to obtain a mask image; wherein, in the potential energy map, the potential energy value greater than the preset potential energy threshold is set to a first value, and the potential energy value less than or equal to the preset potential energy threshold is set to a second value, and the pixel positions corresponding to the first value form a first region.
[0056] The potential energy value in the potential energy map is used to indicate the distance between the pixel position where the potential energy value is located and the target position. For example, the potential energy value of the target position is 1, the potential energy value of the pixel position adjacent to the target position is 0.9, and the potential energy value of a pixel position further away is 0.8, and so on until the potential energy value becomes 0.
[0057] If the potential energy value decreases rapidly with distance, the pixel positions with potential energy values greater than 0 form a local region in the potential energy map, and this local region may be a circular, elliptical or other shaped region; if the potential energy value decreases slowly with distance, the potential energy values of all pixel positions in the potential energy map may be greater than 0.
[0058] The purpose of generating the potential energy map in this embodiment is to indicate the distance between each pixel position and the target position through the potential energy value of each pixel position. The potential energy value of a pixel position closer to the target position is larger, and the probability that this pixel position is included in the first region is greater. The potential energy value of a pixel position farther away is smaller, and the probability that this pixel position is included in the first region is smaller. Thus, the regional position of the first region in the image includes the target position selected by the user to meet the user's position requirements.
[0059] The aforementioned preset potential energy threshold can be set according to requirements. When the preset potential energy threshold is larger, the number of pixel positions greater than the preset potential energy threshold is smaller, and thus the area of the first region is smaller; in one example, when the range of the potential energy value is from 1 to 0, if the preset potential energy threshold is 0.9, the pixel positions corresponding to the potential energy value greater than 0.9 form the first region A; if the preset potential energy threshold is 0.8, the upward positions corresponding to the potential energy value greater than 0.8 form the first region B, and the area of the first region B is larger than the area of the first region A.
[0060] The aforementioned first value and second value are different numerical values. For example, the first value is 1 and the second value is 0, or the first value is 2 and the second value is 1, etc. There are various potential energy values in the potential energy map. After binarization processing, there are only two pixel values in the mask image. Among them, the potential energy value greater than the preset potential energy threshold is the first value, and the potential energy value less than or equal to the preset potential energy threshold is the second value.
[0061] In one example, the pixel values in the mask image include 1 and 0, where 1 is a first value, and the pixel positions where the first value is located constitute the aforementioned first area, and 0 is a second value, and the pixel positions where the second value is located constitute the image area outside the first area. It should be noted that the target position usually has a large potential energy value, so the target position is always located in the first area, thereby ensuring that the first area indicated by the mask image does not deviate from the target position selected by the user.
[0062] Furthermore, a potential energy map in the form of a Gaussian distribution corresponding to the initial image is generated with the target position as the center. The potential energy value in the potential energy map can be understood as the height potential energy of the corresponding pixel position. A height potential energy field is generated around the target position. The height potential energy field is centered on the target position and is Gaussian distributed; the potential energy value at the target position is the largest.
[0063] It can be seen from the above embodiments that, in a specified time step, the first region in the mask image is continuously adjusted so that the region contour of the first region matches the image content corresponding to the target text.
[0064] In a specific adjustment method, in a specified time step among multiple time steps, the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text is determined; the gradient of the similarity relative to the mask image is determined, and the gradient is superimposed on the potential energy map corresponding to the mask image to obtain an updated potential energy map; based on a preset potential energy threshold, the updated potential energy map is binarized to obtain an updated mask image; wherein, in the updated mask image, the first area is updated.
[0065] In actual implementation, the latent space vector output at a specified time step is decoded to obtain an intermediate image. The intermediate image and the target text can be mapped into a feature space to obtain the image features corresponding to the intermediate image and the text features corresponding to the target text, and then the similarity between the image features and the text features is calculated. The similarity can be represented by cosine distance, Euclidean distance or other distance functions.
[0066] In a specific implementation method, the regional image content of the first area indicated by the mask image is extracted from the intermediate image corresponding to the previous time step of the specified time step; the cosine distance between the regional image content and the target text is calculated; wherein the cosine distance indicates the similarity between the regional image content and the target text.
[0067] Since the present embodiment performs local editing on the initial image, it is only necessary to focus on the similarity between the first region and the target text; the first region is segmented from the intermediate image, the regional image features corresponding to the first region are obtained, and the cosine distance between the regional image features and the text features of the target text is calculated, and the cosine distance is used to indicate the similarity between the current regional image content of the first region and the target text.
[0068] In one example, the intermediate image, the current mask image, and the target text are input into the aforementioned Alpha-CLIP model, and the cosine distance between the feature vector of the intermediate image and the feature vector of the target text is calculated through this model.
[0069] Furthermore, the text feature of the target text can specifically be the latent space vector of the target file; by means of backpropagation, the gradient of the similarity with respect to the mask image is calculated; specifically, the gradient of the similarity with respect to the downsampled mask image can be calculated, that is, the aforementioned cosine distance with respect to the downsampled mask image; the larger the gradient, the more important the pixel position represents for generating the image content corresponding to the target text, and the greater the probability that the pixel position belongs to the first region.
[0070] After calculating the gradient, the absolute value of the gradient is calculated, and the absolute value is superimposed on the potential energy map corresponding to the mask image to obtain an updated potential energy map. In actual implementation, for each pixel position in the mask image, a corresponding gradient value is calculated, and this gradient value is superimposed on the potential energy value at the pixel position; when the absolute value of the gradient value is greater than 0, the potential energy value at this pixel position will increase, and thus, the probability that this pixel position is updated to the first region will increase.
[0071] It should be noted that at the specified time step, what is actually updated is the potential energy map corresponding to the mask image, that is, the potential energy value in the potential energy map will be updated after the gradient value is superimposed, and then the updated potential energy map is binarized based on a preset potential energy threshold to obtain an updated mask image.
[0072] In one way, at the specified time step, in addition to updating the potential energy map, a preset potential energy threshold is also increased; this preset potential energy threshold can be gradually increased at each specified time step, that is, as the specified time step changes, the preset potential energy threshold is gradually increasing, and the potential energy values in the regions in the potential energy map that are relatively important for generating the image content corresponding to the target text are also gradually increasing, so as to ensure that the potential energy values in the important regions are always higher than the preset potential energy threshold, and these regions are the first regions for editing the image content corresponding to the target text.
[0073] For the pixel positions where the potential energy value increases slowly, when the potential energy value is lower than the preset potential energy threshold, this pixel position will be updated to a pixel position outside the first region. It can be understood that the first region shrinks at this pixel position. For the pixel positions where the potential energy value increases relatively quickly, which may be outside the first region, as the potential energy value increases, when the potential energy value is higher than the preset potential energy threshold, this pixel position will be updated to the first region. It can be understood that the first region expands at this pixel position.
[0074] Therefore, the final area and shape of the first region are determined by the image content indicated by the target text.
[0075] In a specific implementation, at the first time step, calculate the first Hadamard product result of the text feature of the target text and the downsampled mask image, and calculate the second Hadamard product result of the image feature of the initial image and the inverted image of the downsampled mask image; wherein, the inverted image is used to: indicate the region outside the first region in the initial image; fuse the first Hadamard product result and the second Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature.
[0076] At subsequent time steps after the first time step, calculate the first Hadamard product result of the text feature of the target text and the downsampled mask image, and calculate the third Hadamard product result of the denoised feature corresponding to the previous time step and the inverted image of the downsampled mask image; fuse the first Hadamard product result and the third Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature; decode the denoised feature at the last time step to obtain the target image.
[0077] The process of feature fusion for one time step can be expressed by the following formula:
[0078] z t = z fg ⊙ m latent + z bg ⊙ (1 - m latent )
[0079] Wherein, z t is the fused feature, z fg is the text feature of the target text, which can specifically be the latent space vector of the target text. Since the initial image is locally edited based on the target text, the text feature of the target text can also be understood as the foreground vector; ⊙ is the Hadamard product operator;
[0080] m latent is the downsampled mask image, indicating the region position of the first region. This first region can be understood as the region of the foreground image generated based on the target text; the region outside the first region indicated by the inverted image can be understood as the region of the background image provided by the initial image. Since the pixel values in the mask image consist of 0 and 1, (1 - m latent ) is to invert the downsampled mask image to obtain the inverted image. The inversion process can be understood as inverting the pixel value 0 to pixel value 1 and pixel value 1 to pixel value 0; therefore, in the inverted image, the pixel positions with pixel value 1 form the region outside the first region, and the pixel positions with pixel value 0 form the first region.
[0081] z bg is the denoised feature for the previous time step. When the time step is the first time step, z bg is the image feature of the initial image, that is, the latent space vector of the initial image.
[0082] At each time step, z fg and z bg are fused with the downsampled masked image m latent as the weight; for a pixel position, if the pixel value at this pixel position in m latent is 1, then the feature corresponding to this pixel position in z fg is used at this pixel position. If the pixel value at this pixel position in m latent is 0, then the feature corresponding to this pixel position in z bg is used at this pixel position.
[0083] Based on this, m latent determines which regions in the initial image are edited and which regions are retained; at the last time step, after outputting the denoised feature, the denoised feature is decoded to obtain the target image.
[0084] In the above way, it is possible to continue editing within the specified region and keep the image details outside this region unchanged.
[0085] Combined with the above formula, see Figure 2 the schematic diagram of the image generation process shown.
[0086] To achieve the purpose of locally editing the initial image, the input data includes the initial image, the target position, and the target position; the initial image can generate the image feature of the initial image, that is, the latent space vector z int , in the first time step, z bg is z int , in subsequent time steps, z bg is the denoised feature Z0’ output in the previous time step.
[0087] z fg is the text feature of the target text. Specifically, the target text can be input into the text encoder, and the output feature p is denoised to generate z fg . z fg and z bg After being fused based on the masked image Mt, are input into the denoising network, and the feature p is also input into the denoising network to output the denoised feature Z0’ of the current time step. The denoised feature Z0’ is decoded by the image decoder to generate an intermediate image.
[0088] Generate a potential energy map based on the target position. The feature p, intermediate image, and mask image Mt are input into the Alpha-CLIP model to output a gradient. This gradient is superimposed on the potential energy map to generate the mask image Mt-1, and the mask image Mt for the next time step t is updated based on this mask image Mt-1.
[0089] At the last time step, the denoising network outputs the denoised feature Z0, which is input into the image decoder to output the target image.
[0090] In Figure 3 Three groups of examples are shown. In the first example, the target text is "a green bowl", the green dot represents the target position, a first region is generated based on this target position, the magenta region represents the first region, and the second region corresponding to the first region in the target image finally generates a green bowl. In the second example, the target text is "a small cow", the green dot represents the target position, a first region is generated based on this target position, the magenta region represents the first region, and the second region corresponding to the first region in the target image finally generates a small cow. In the third example, the target text is "people swimming", the green dot represents the target position, a first region is generated based on this target position, the magenta region represents the first region, and the second region corresponding to the first region in the target image finally generates four people swimming.
[0091] In Figure 4 After the user clicks on a target position in the initial image, a first region is generated based on this target position; the red region represents the first region. When the target text is "tombstone", the target image 1 containing the tombstone is generated. When the target text is "cartoon car", the target image 2 containing the cartoon car is generated. When the target text is "snake", the target image 3 containing the snake is generated.
[0092] Figure 5 The adjustment process of the first region indicated by the mask image is shown. In the first example, the target text is "a giraffe". First, the first region is an approximately circular region. After continuous adjustment, the shape of the first region approaches the shape of a giraffe, and finally the target image containing the giraffe is generated. In the second example, the target text is "snow mountain". First, the first region is an approximately circular region. After continuous adjustment, the shape of the first region approaches the shape of a mountain, and finally the target image containing the snow mountain is generated. It should be noted that in Figure 5 In the example, the percentage represents the position of the current time step in the total number of time steps. For example, when the current time step is 48%, the final mask image is determined, and the shape of the first region no longer changes.
[0093] The image generation method provided in this embodiment only requires the user to click with the mouse or touch to provide a reference point, i.e., the aforementioned target position, and provide a text description, and then precise local image editing can be achieved. During the editing process, a mask image Mask will be dynamically generated based on this reference point, and this mask image will be dynamically evolved and adjusted according to the semantic loss to guide the mask image, and finally a mask image that fits the image content indicated by the text description will be generated, so as to achieve local editing of the image.
[0094] This embodiment only uses the single-point position clicked by the user and the text description to dynamically generate the mask image, without the need for the user to manually create the mask image or provide the text description of the editing area, thereby reducing the complexity of user input, simplifying the user input requirements, and reducing the dependence on input data; this embodiment can achieve real-time local image editing, improve the interactivity and real-time performance of the editing process, is easy to operate, user-friendly, and can achieve efficient and accurate local image editing.
[0095] Corresponding to the above method embodiment, refer to Figure 6 An image generation device as shown, the device includes:
[0096] A position determination module 60, configured to respond to a position selection instruction for an initial image and determine a target position on the initial image;
[0097] A mask generation module 62, configured to generate a mask image corresponding to the initial image based on the target position; wherein, the mask image is used to: indicate a first area to be edited in the initial image;
[0098] An image generation module 64, configured to perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text based on the mask image until a target image is generated; wherein, a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0099] The above image generation device responds to a position selection instruction for an initial image, determines a target position on the initial image; generates a mask image corresponding to the initial image based on the target position; wherein, the mask image is used to: indicate a first area to be edited in the initial image; perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text based on the mask image until a target image is generated; wherein, a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0100] In this method, the user only needs to provide an image location, and a mask image can be generated based on this image location. The mask image indicates the area to be edited, so as to edit the image content indicated by the target text in this area; this method improves the image editing efficiency on the premise of ensuring good editing effects.
[0101] The above device further includes a mask adjustment module, which is used to adjust the first area in the mask image based on the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text in the specified time step of multiple time steps; wherein, the intermediate image is generated based on the features after feature fusion and feature denoising processing of the previous time step.
[0102] The above mask generation module is used to: generate a potential energy map corresponding to the initial image with the target position as the center; wherein, in the potential energy map, the pixel positions closer to the target position have larger corresponding potential energy values; perform binarization processing on the potential energy map based on a preset potential energy threshold to obtain a mask image; wherein, in the potential energy map, the potential energy values greater than the preset potential energy threshold are set to a first value, and the potential energy values less than or equal to the preset potential energy threshold are set to a second value, and the pixel positions corresponding to the first value form a first area.
[0103] The above mask generation module is used to: generate a potential energy map in the form of a Gaussian distribution corresponding to the initial image with the target position as the center.
[0104] The above mask adjustment module is used to: determine the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text in the specified time step of multiple time steps; determine the gradient of the similarity with respect to the mask image, and superimpose the gradient on the potential energy map corresponding to the mask image to obtain an updated potential energy map; perform binarization processing on the updated potential energy map based on a preset potential energy threshold to obtain an updated mask image; wherein, in the updated mask image, the first area is updated.
[0105] The above mask adjustment module is used to: extract the regional image content of the first area indicated by the mask image from the intermediate image corresponding to the previous time step of the specified time step; calculate the cosine distance between the regional image content and the target text; wherein, the cosine distance indicates the similarity between the regional image content and the target text.
[0106] The above mask adjustment module is used to: calculate the gradient of the similarity with respect to the downsampled mask image.
[0107] The above mask adjustment module is used to: calculate the absolute value of the gradient, and superimpose the absolute value on the potential energy map corresponding to the mask image to obtain an updated potential energy map.
[0108] The above device further includes a potential energy value increase module, which is used to: increase the preset potential energy threshold.
[0109] The above-mentioned image generation module is used for: at the first time step, calculating the first Hadamard product result of the text features of the target text and the downsampled mask image, and calculating the second Hadamard product result of the image features of the initial image and the inverted image of the downsampled mask image; wherein, the inverted image is used to indicate the area outside the first area in the initial image; fusing the first Hadamard product result and the second Hadamard product result to obtain a fused feature; inputting the fused feature into a denoising network for processing to obtain a denoised feature; at subsequent time steps after the first time step, calculating the first Hadamard product result of the text features of the target text and the downsampled mask image, and calculating the third Hadamard product result of the denoised feature corresponding to the previous time step of the specified time step and the inverted image of the downsampled mask image; fusing the first Hadamard product result and the third Hadamard product result to obtain a fused feature; inputting the fused feature into a denoising network for processing to obtain a denoised feature; decoding the denoised feature of the last time step to obtain the target image.
[0110] The image generation method and device provided in this embodiment only use a single-point reference position clicked by the user and a content description to dynamically generate the mask image Mask, reducing the operation complexity of the user during the image editing process. This simplified operation method enables non-professional users to easily perform local image editing without the need to possess advanced image processing skills or in-depth understanding of the details of Mask creation.
[0111] The image generation method of this embodiment can complete the image editing task in about one second, which greatly speeds up the editing speed and improves the real-time performance. This fast response ability makes this embodiment suitable for application scenarios that require instant feedback, such as online image editing tools or mobile applications.
[0112] Guided by semantic loss in this embodiment, it ensures the semantic consistency between the edited content and the text description provided by the user, thereby improving the accuracy and naturalness of the editing result. This method can more accurately understand and execute the user's editing intention, and generate high-quality results that seamlessly blend with the original image.
[0113] Since this embodiment adopts an editing process that does not require training, it reduces the demand for computing resources. It can run on resource-constrained devices, such as smartphones or tablets, without relying on high-performance computing devices.
[0114] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned image generation method. This electronic device can be a server or a terminal device.
[0115] See Figure 7As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100, and the processor 100 executes the computer-executable instructions to implement the above image generation method.
[0116] Furthermore, Figure 7 The electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected via the bus 102.
[0117] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is implemented through at least one communication interface 103 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0118] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0119] The processor in the above electronic device can implement the following operations in the above image generation method by executing computer-executable instructions:
[0120] An image generation method, in response to a position selection instruction for an initial image, determines a target position on the initial image; based on the target position, generates a mask image corresponding to the initial image; wherein the mask image is used to: indicate a first region to be edited in the initial image; based on the mask image, perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until a target image is generated; wherein the second region in the target image contains the image content corresponding to the target text, and the region outside the second region contains the image content corresponding to the initial image; the second region corresponds to the first region.
[0121] In a specified time step among the multi-time steps, based on the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text, adjust the first region in the mask image; wherein the intermediate image is generated based on the features after the feature fusion and feature denoising processing of the previous time step.
[0122] Generate a potential energy map corresponding to the initial image centered on the target position; wherein, in the potential energy map, the closer the pixel position is to the target position, the greater the corresponding potential energy value; perform binarization processing on the potential energy map based on a preset potential energy threshold to obtain a mask image; wherein, in the potential energy map, the potential energy value greater than the preset potential energy threshold is set to a first value, and the potential energy value less than or equal to the preset potential energy threshold is set to a second value, and the pixel positions corresponding to the first value form a first region.
[0123] Generate a potential energy map in the form of a Gaussian distribution corresponding to the initial image centered on the target position.
[0124] In a specified time step among multiple time steps, determine the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text; determine the gradient of the similarity with respect to the mask image, and superimpose the gradient onto the potential energy map corresponding to the mask image to obtain an updated potential energy map; perform binarization processing on the updated potential energy map based on a preset potential energy threshold to obtain an updated mask image; wherein, in the updated mask image, the first region is updated.
[0125] Extract the regional image content of the first region indicated by the mask image from the intermediate image corresponding to the previous time step of the specified time step; calculate the cosine distance between the regional image content and the target text; wherein, the cosine distance indicates the similarity between the regional image content and the target text.
[0126] Calculate the gradient of the similarity with respect to the downsampled mask image.
[0127] Calculate the absolute value of the gradient, and superimpose the absolute value onto the potential energy map corresponding to the mask image to obtain an updated potential energy map.
[0128] Increase the preset potential energy threshold.
[0129] In the first time step, calculate the first Hadamard product result of the text features of the target text and the downsampled mask image, and calculate the second Hadamard product result of the image features of the initial image and the inverted image of the downsampled mask image; wherein, the inverted image is used to: indicate the region outside the first region in the initial image; fuse the first Hadamard product result and the second Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature; in the subsequent time steps of the first time step, calculate the first Hadamard product result of the text features of the target text and the downsampled mask image, and calculate the third Hadamard product result of the denoised feature corresponding to the previous time step and the inverted image of the downsampled mask image; fuse the first Hadamard product result and the third Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature; decode the denoised feature of the last time step to obtain the target image.
[0130] In this method, the user only needs to provide an image position, and a mask image can be generated based on this image position. The mask image indicates the area to be edited, so as to edit the image content indicated by the target text in this area. This method improves the image editing efficiency on the premise of ensuring a good editing effect.
[0131] This embodiment also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above image generation method.
[0132] The computer-executable instructions stored in the above computer-readable storage medium can implement the following operations in the above image generation method by executing the computer-executable instructions:
[0133] An image generation method, in response to a position selection instruction for an initial image, determines a target position on the initial image; based on the target position, generates a mask image corresponding to the initial image; wherein, the mask image is used to: indicate a first area to be edited in the initial image; based on the mask image, perform multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until a target image is generated; wherein, a second area in the target image contains the image content corresponding to the target text, and the area outside the second area contains the image content corresponding to the initial image; the second area corresponds to the first area.
[0134] In a specified time step among the multi-time steps, based on the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text, adjust the first area in the mask image; wherein, the intermediate image is generated based on the features after the feature fusion and feature denoising processing of the previous time step.
[0135] Centered on the target position, generate a potential energy map corresponding to the initial image; wherein, in the potential energy map, the closer the pixel position is to the target position, the greater the corresponding potential energy value; based on a preset potential energy threshold, perform binary processing on the potential energy map to obtain a mask image; wherein, in the potential energy map, the potential energy value greater than the preset potential energy threshold is set to a first value, and the potential energy value less than or equal to the preset potential energy threshold is set to a second value, and the pixel positions corresponding to the first value form the first area.
[0136] Centered on the target position, generate a potential energy map in the form of a Gaussian distribution corresponding to the initial image.
[0137] At a specified time step among multiple time steps, determine the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text; determine the gradient of the similarity with respect to the masked image, and superimpose the gradient onto the potential energy map corresponding to the masked image to obtain an updated potential energy map; based on a preset potential energy threshold, perform binarization processing on the updated potential energy map to obtain an updated masked image; wherein, in the updated masked image, the first region is updated.
[0138] Extract the regional image content of the first region indicated by the masked image from the intermediate image corresponding to the previous time step of the specified time step; calculate the cosine distance between the regional image content and the target text; wherein, the cosine distance indicates the similarity between the regional image content and the target text.
[0139] Calculate the gradient of the similarity with respect to the downsampled masked image.
[0140] Calculate the absolute value of the gradient, and superimpose the absolute value onto the potential energy map corresponding to the masked image to obtain an updated potential energy map.
[0141] Increase the preset potential energy threshold.
[0142] At the first time step, calculate the first Hadamard product result of the text features of the target text and the downsampled masked image, and calculate the second Hadamard product result of the image features of the initial image and the inverted image of the downsampled masked image; wherein, the inverted image is used to: indicate the region outside the first region in the initial image; fuse the first Hadamard product result and the second Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature; at subsequent time steps after the first time step, calculate the first Hadamard product result of the text features of the target text and the downsampled masked image, and calculate the third Hadamard product result of the denoised feature corresponding to the previous time step and the inverted image of the downsampled masked image; fuse the first Hadamard product result and the third Hadamard product result to obtain a fused feature; input the fused feature into a denoising network for processing to obtain a denoised feature; decode the denoised feature at the last time step to obtain the target image.
[0143] In this method, the user only needs to provide an image position, and then a masked image can be generated based on this image position. The masked image indicates the region to be edited, so as to edit the image content indicated by the target text in this region; this method improves the image editing efficiency while ensuring a good editing effect.
[0144] The computer program product of the image generation method, device and electronic device provided by the embodiments of the present invention includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.
[0145] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0146] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0147] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0148] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0149] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An image generation method, characterized in that: The method comprises: In response to a position selection instruction for an initial image, determining a target position on the initial image; Based on the target position, a mask image corresponding to the initial image is generated; wherein the mask image is used to: indicate a first area to be edited in the initial image; Based on the mask image, performing multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until a target image is generated; The second area in the target image contains image content corresponding to the target text, and the area outside the second area contains image content corresponding to the initial image; the second area corresponds to the first area.
2. The method according to claim 1, characterized in that After the step of generating a mask image corresponding to the initial image based on the target position, the method further includes: In a specified time step among the multiple time steps, the first area in the mask image is adjusted based on the similarity between an intermediate image corresponding to a previous time step of the specified time step and the target text; wherein the intermediate image is generated based on features after feature fusion and feature denoising processing of the previous time step.
3. The method according to claim 1, characterized in that The step of generating a mask image corresponding to the initial image based on the target position includes: Taking the target position as the center, generating a potential energy map corresponding to the initial image; wherein, in the potential energy map, the closer the pixel position is to the target position, the greater the corresponding potential energy value; Based on a preset potential energy threshold, the potential energy map is binarized to obtain a mask image; wherein, in the potential energy map, potential energy values greater than the preset potential energy threshold are set to a first value, and potential energy values less than or equal to the preset potential energy threshold are set to a second value, and the pixel positions corresponding to the first value constitute the first area.
4. The method according to claim 3, characterized in that The step of generating a potential energy map corresponding to the initial image with the target position as the center comprises: generating a potential energy map in the form of a Gaussian distribution corresponding to the initial image with the target position as the center.
5. The method according to claim 2, characterized in that: In a specified time step among the multiple time steps, based on the similarity between the intermediate image corresponding to a previous time step of the specified time step and the target text, the step of adjusting the first area in the mask image comprises: In a specified time step among the multiple time steps, determining a similarity between an intermediate image corresponding to a previous time step of the specified time step and the target text; Determining a gradient of the similarity relative to the mask image, and superimposing the gradient onto a potential energy map corresponding to the mask image to obtain an updated potential energy map; Based on a preset potential energy threshold, the updated potential energy map is binarized to obtain an updated mask image; wherein, in the updated mask image, the first area is updated.
6. The method according to claim 5, characterized in that The step of determining the similarity between the intermediate image corresponding to the previous time step of the specified time step and the target text comprises: Extracting the regional image content of the first region indicated by the mask image from the intermediate image corresponding to the previous time step of the specified time step; The cosine distance between the regional image content and the target text is calculated; wherein the cosine distance indicates the similarity between the regional image content and the target text.
7. The method according to claim 5, characterized in that The step of determining the gradient of the similarity relative to the mask image comprises: calculating the gradient of the similarity relative to the down-sampled mask image.
8. The method according to claim 5, characterized in that The step of superimposing the gradient onto the potential energy map corresponding to the mask image to obtain an updated potential energy map comprises: The absolute value of the gradient is calculated, and the absolute value is superimposed on the potential energy map corresponding to the mask image to obtain an updated potential energy map.
9. The method according to claim 5, characterized in that Before the step of binarizing the updated potential energy map based on a preset potential energy threshold to obtain an updated mask image, the method further includes: increasing the preset potential energy threshold.
10. The method according to claim 1, characterized in that Based on the mask image, the step of performing multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text until the target image is generated includes: At the first time step, a first Hadamard product result of the text feature of the target text and the downsampled mask image is calculated, and a second Hadamard product result of the image feature of the initial image and the inverse image of the downsampled mask image is calculated; wherein the inverse image is used to indicate an area outside the first area in the initial image; The first Hadamard product result and the second Hadamard product result are fused to obtain a fused feature; the fused feature is input into a denoising network for processing to obtain a denoising feature; At a subsequent time step of the first time step, a first Hadamard product result of the text feature of the target text and the downsampled mask image is calculated, and a third Hadamard product result of the denoising feature corresponding to the previous time step and the inverse image of the downsampled mask image is calculated; The first Hadamard product result and the third Hadamard product result are fused to obtain a fused feature; the fused feature is input into a denoising network for processing to obtain a denoising feature; Decode the denoised features of the last time step to obtain the target image.
11. An image generating device, characterized in that: The device comprises: a position determination module, configured to determine a target position on the initial image in response to a position selection instruction for the initial image; A mask generation module, used to generate a mask image corresponding to the initial image based on the target position; wherein the mask image is used to: indicate a first area to be edited in the initial image; An image generation module, used for performing multi-time-step feature fusion and feature denoising processing on the image features of the initial image and the text features of the target text based on the mask image, until a target image is generated; The second area in the target image contains image content corresponding to the target text, and the area outside the second area contains image content corresponding to the initial image; the second area corresponds to the first area.
12. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the image generating method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the image generation method according to any one of claims 1 to 10.
Citation Information
Cited By
Deep neural network pruning method based on dynamic contrast mask and knowledge distillation
CN120952086A