Image Completion Method, Device, Storage Medium and Electronic Device
Through the image completion model, the image encoder and fusion module are used to process images to generate a natural complete image, which solves the problem of poor image completion effect in the prior art and achieves a more natural image completion effect.
Patent Information
- Application Number
- CN202011611826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-12-30
AI Technical Summary
The existing image completion technology is not effective when dealing with complex scenes, and there are obvious color differences and artifacts at the edges of the generated images, which lacks naturalness.
An image completion model is adopted, which includes an image encoder and a fusion module. Image encoding is generated by an image encoder, and a complete image is generated based on channel features and spatial features through a fusion module. The fusion module includes a decoder and an attention submodule through which image encoding is processed to generate a natural complete image.
On the basis of maintaining the image intent, image completion is achieved, artifacts and chromatic aberration are reduced, and the generated complete image is more natural.
Smart Images

Figure CN114764748B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing, and in particular, to an image completion method, an apparatus, a storage medium, and an electronic device. Background Art
[0002] Image completion is a technology for generating alternative content in missing regions of an image to make the image complete, and is widely applied in various scenarios in life. The current image completion technology has unsatisfactory effects when dealing with complex scenarios, and there are obvious color differences and artifacts at the edges of the repaired missing parts of the image, and the generated image is not natural enough. Summary of the Invention
[0003] The purpose of the present disclosure is to provide an image completion method, an apparatus, a storage medium, and an electronic device to solve the above technical problems.
[0004] To achieve the above purpose, the first aspect of the present disclosure provides an image completion method, including: determining the position information to be completed on the image to be completed; inputting the image to be completed and the position information to be completed into an image completion model; obtaining the complete image generated by the image completion model; wherein, the image completion model includes an image encoder and a fusion module, and the step of generating the complete image by the image completion model includes: generating an image code based on the image to be completed and the position information to be completed through the image encoder; generating channel features based on the image code through the fusion module, generating spatial features based on the image to be completed and the position information to be completed, and generating the complete image based on the channel features and the spatial features.
[0005] Optionally, the fusion module includes a decoder and an attention sub-module, and the generating channel features based on the image code through the fusion module includes: generating a plurality of hierarchical features based on the image code through the decoder, and processing the hierarchical features of a plurality of preset hierarchies in the plurality of hierarchical features through the attention sub-module to obtain the channel features.
[0006] Optionally, the fusion module includes an image convolution sub-module, and generating spatial features based on the image to be completed and the position information to be completed includes: performing convolution based on the image to be completed and the position information to be completed through the image convolution sub-module to obtain the spatial features.
[0007] Optionally, the image encoder is trained through the following steps:
[0008] Input multiple first sample images into the training encoder to obtain multiple key values output by the training encoder; input second sample images corresponding to the first sample images into the image encoder to obtain multiple query values output by the image encoder, where the corresponding first sample images and second sample images are images obtained by removing different parts from the same original image; based on the query values and the key values, adjust the parameters of the image encoder and the parameters of the training encoder through a preset encoder training loss function.
[0009] Optionally, the adjusting the parameters of the image encoder and the training encoder based on a preset encoder training loss function includes: adjusting the parameters of the training encoder and the image encoder in a form of momentum update.
[0010] Optionally, the image completion model is trained by one or more of a reconstruction loss function, a perceptual loss function, a style loss function, a total variation loss function, and an adversarial loss function.
[0011] A second aspect of the present disclosure provides an image completion device, the device includes: a position determination module, configured to determine the position information to be completed on the image to be completed; an input module, configured to input the image to be completed and the position information to be completed into the image completion model; an acquisition module, configured to acquire the complete image generated by the image completion model; where the image completion model includes an image encoder and a fusion module, the image encoder is configured to generate an image encoding based on the image to be completed and the position information to be completed; the fusion module is configured to generate channel features based on the image encoding, generate spatial features based on the image to be completed and the position information to be completed, and generate the complete image based on the channel features and the spatial features.
[0012] Optionally, the fusion module includes a decoder and an attention sub-module, the decoder is configured to generate multiple hierarchical features based on the image encoding, and the attention sub-module is configured to process the hierarchical features of multiple preset hierarchies among the multiple hierarchical features to obtain the channel features.
[0013] Optionally, the fusion module includes an image convolution sub-module, configured to perform convolution based on the image to be completed and the position information to be completed to obtain the spatial features.
[0014] Optionally, the fusion module further includes a synthesis sub-module, configured to adjust the formats of the channel features and the spatial features to be consistent; the fusion module is further configured to generate a fusion feature based on the adjusted channel features and the channel features, and generate the complete image based on the fusion feature through a preset image generation algorithm.
[0015] Optionally, the apparatus further includes a training module configured to input multiple first sample images into a training encoder to obtain multiple key values output by the training encoder; input second sample images corresponding to the first sample images into the image encoder to obtain multiple query values output by the image encoder, where the corresponding first sample images and second sample images are images obtained by removing different parts from the same original image; and adjust the parameters of the image encoder and the parameters of the training encoder based on the query values and the key values through a preset encoder training loss function.
[0016] Optionally, the training module is further configured to adjust the parameters of the training encoder and the image encoder in a form of momentum update.
[0017] Optionally, the image completion model is trained by one or more of a reconstruction loss function, a perceptual loss function, a style loss function, a total variation loss function, and an adversarial loss function.
[0018] A third aspect of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method described in the first aspect of the present disclosure are implemented.
[0019] A fourth aspect of the present disclosure provides an electronic device, including a memory and a processor, where a computer program is stored on the memory, and the processor is configured to execute the computer program in the memory to implement the steps of the method described in the first aspect of the present disclosure.
[0020] Through the above technical solutions, at least the following technical effects can be achieved:
[0021] By processing the image to be completed through the image completion model to obtain channel features and spatial features, and generating a complete image based on the channel features and spatial features, image completion can be achieved while maintaining the image meaning and considering the missing position of the image, reducing artifacts and color differences, and making the generated completed image more natural.
[0022] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings are used to provide a further understanding of the present disclosure, and constitute a part of the specification, and are used to explain the present disclosure together with the following specific implementation manners, but do not constitute a limitation to the present disclosure. In the drawings:
[0024] Figure 1 is a flowchart of an image completion method shown according to an exemplary embodiment of the present disclosure.
[0025] Figure 2It is a schematic diagram of an image encoder training process shown according to an exemplary disclosed embodiment.
[0026] Figure 3 It is a schematic structural diagram of an image completion model shown according to an exemplary disclosed embodiment.
[0027] Figure 4 It is a block diagram of an image completion device shown according to an exemplary disclosed embodiment.
[0028] Figure 5 It is a block diagram of an electronic device shown according to an exemplary disclosed embodiment. Detailed implementation manners
[0029] The following will describe the detailed implementation manners of the present disclosure with reference to the accompanying drawings. It should be understood that the detailed implementation manners described herein are only used to illustrate and explain the present disclosure, and are not used to limit the present disclosure.
[0030] Figure 1 It is a flowchart of an image completion method shown according to an exemplary disclosed embodiment. This method can be applied to any electronic device such as a server or a terminal. The image to be completed in the present disclosure can be an image stored in the electronic device, an image obtained from the Internet, or an image obtained through any other means.
[0031] As Figure 1 shown, the method includes the following steps:
[0032] S11. Determine the position information to be completed on the image to be completed.
[0033] The position of the missing area on the picture can be extracted from the image to be completed in the form of a position mask, a mask image, etc., and used as the input value of the image completion model in the form of completion position information. As Figure 1 shown, it is a schematic diagram of an image to be completed and the corresponding completion position information of the image to be completed. Figure 1 The completion position information in
[0034] is in the form of a mask image.
[0035] S12. Input the image to be completed and the position information to be completed into the image completion model.
[0036] S13. Obtain the complete image generated by the image completion model.
[0037] Among them, the image completion model includes an image encoder and a fusion module. The steps for the image completion model to generate the complete image include:
[0038] Generating an image encoding based on the image to be completed and the position information to be completed through the image encoder;
[0038] The fusion module generates channel features based on the image encoding, generates spatial features based on the image to be completed and the position information to be completed, and generates the complete image based on the channel features and the spatial features.
[0039] Among them, the image encoding output by the image encoder can generate multiple hierarchical features after decoding. The channel features can be obtained by processing these hierarchical features. The spatial features can be obtained by convolving the image to be completed and the position information to be completed. The channel features focus on characterizing the semantic content of the image, such as the object type, scene, etc. in the image. The spatial features focus on characterizing the spatial content such as the missing positions in the image. By weighted fusion of the channel features and the spatial features, a fusion feature that can both reflect the image content and emphasize the missing positions can be obtained. Based on this fusion feature, a natural complete image can be generated.
[0040] In a possible implementation manner, the fusion module includes a decoder and an attention sub-module. The decoder can generate multiple hierarchical features based on the image encoding, and the attention sub-module processes the hierarchical features of multiple preset hierarchies among the multiple hierarchical features to obtain the channel features.
[0041] The attention sub-module can be embedded after a preset hierarchy of the decoder to process the output value of the preset hierarchy of the decoder to obtain the channel features for feature fusion.
[0042] It should be noted that due to the different structures of the preset number of layers of the decoder, the output hierarchical features may be different, and there may be multiple intermediate results obtained by the attention module based on the processing of the hierarchical features. The following form can be used to process the intermediate results to obtain the channel features.
[0043] Define the hierarchical features (i.e., intermediate results) processed by the attention module as F = [f1, …, f c , …, f C , then the statistical information z c of the channel features can be obtained through the following calculation method:
[0044]
[0045] Among them, z c is the statistical value of the channel features in the c-th dimension, H GP represents the global pooling layer, H GP (f c ) is the hierarchical feature f c after global pooling, h is the number of rows of the f c matrix, and w is the number of columns of the f c matrix.
[0046] Further, the Sigmoid function can be used as the gating function to obtain the channel weight ω:
[0047] ω = f(W U δ(W D z))
[0048] where f is the sigmoid function and δ is the ReLU activation function. W U and W D are both preset parameters of the convolutional layer. Among them, W U is used to set the number of channels to C / r, and W D is used to set the number of channels to C, where C is the channel number variable and r is the scaling factor.
[0049] The hierarchical features can be weighted and adjusted through the channel weight ω to obtain the weighted hierarchical features where ω c is the channel weight of f c .
[0050]
[0051] Channel features That is to say, the channel features are the set of processed hierarchical features.
[0052] In a possible implementation manner, the fusion module includes an image convolution sub-module, and the spatial features can be obtained by performing convolution on the to-be-complemented image and the to-be-complemented position information through the image convolution sub-module.
[0053] The spatial features can also be processed to make the formats of the spatial features and the channel features consistent.
[0054] Specifically, given a to-be-complemented image x′ q , the to-be-complemented image is transformed in channels and size to match the channel features
[0055] x′ q =(W C x q )↓
[0056] where W C is the parameter of the 1×1 convolutional layer, and ↓ are the downsampling models respectively.
[0057] After obtaining the processed spatial features and the processed channel features, the fusion feature α can be calculated through the following formula:
[0058]
[0059] where, W Dare the parameters of the 1×1 convolutional layer, used to set the number of channels to be the same as that of x′ q and f is the sigmoid function. A is a learnable transformation function composed of multiple convolutional functions, and its parameters can be adjusted during the training phase.
[0060] The final completion result can be obtained by the following formula:
[0061]
[0062] where ⊙ represents the Hadamard product, and B represents the combination function.
[0063] The image completion model in this disclosure is divided into two training processes. The first training process is the training of the image encoder. Multiple first sample images can be input into the training encoder to obtain multiple key values output by the training encoder; the second sample images corresponding to the first sample images are input into the image encoder to obtain multiple query values output by the image encoder. Among them, the corresponding first sample images and the second sample images are images obtained by removing different parts from the same original image; the parameters of the image encoder and the parameters of the training encoder are adjusted based on a preset encoder training loss function to make the query values close to the key values.
[0064] The preset encoder training loss function L can be set as where z q is the key value output by the training encoder, is the query value corresponding to z q and τ is the temperature control hyperparameter.
[0065] Figure 2 The figure shows a schematic diagram of an image encoder training process. The first sample image a (including a1, a2, a3,...) is input into the training encoder to obtain key values c1, c2, c3...; the second sample image b (including b1, b2, b3,... corresponding to a1, a2, a3 respectively) is input into the image encoder to obtain query values d1, d2, d3... The parameters of the image encoder and the parameters of the training encoder are adjusted through the encoder training loss function to make the query values close to the key values. Among them, the sample images include the missing images and the corresponding missing position information to be completed.
[0066] Since the update of the encoder is performed by querying key values close to the query value, and the number of samples required for training is large, considering the memory size of the GPU and the difficulty of feature learning, the update of the continuous dictionary can be implemented in the form of a queue. During training, the latest key values are enqueued, while the old key values are dequeued. In this way, the training efficiency can be improved and the training difficulty can be reduced.
[0067] In a possible implementation manner, the parameters of the training encoder and the image encoder can be adjusted based on the form of momentum update:
[0068] θ k ←mθ k +(1 - m)θ q
[0069] where θ q is the parameter of the training encoder, and θ k represents the parameter of the image encoder. m is the momentum update parameter, and the specific value of this momentum update parameter can be adjusted according to the training needs.
[0070] The second training of the image completion model in this disclosure is the training of the fusion module. Specifically, according to the training requirements, one or more of the reconstruction loss function, perceptual loss function, style loss function, total variation loss function, and adversarial loss function can be set to train the fusion module.
[0071] Among them, the reconstruction loss function L rec can be set as:
[0072]
[0073] Among them, Y is the complete image of the sample image input to the image completion model during training, is the completed image generated by the image completion model based on the sample image during training. The reconstruction loss function is used to reduce the absolute difference between the complete image Y and the completed image obtained during training, so that the image generated by the model is close to the sample image.
[0074] The perceptual loss function L per can be set as:
[0075]
[0076] Among them, φ is the pre-trained VGG-16 network. φ i is the feature map of the i-th pooling layer output by the VGG-16 network. In this solution, the pool-1, pool-2, and pool-3 layers in VGG-16 are used. The perceptual loss function is used to improve the visual quality of the generated image.
[0077] Style loss function L style can be set as:
[0078]
[0079] where C i represents the number of channels of the feature map output by the i-th layer of the pre-trained VGG-16 network.
[0080] Total variational loss function L tv can be set as:
[0081]
[0082] where represents the pixel at position (i + 1, j) on the completed image, and Ω represents the position to be completed in the image.
[0083] Adversarial loss function L adv can be set as:
[0084]
[0085] where G is the image encoder, D is the decoder, and P Y is the distribution function of the complete image Y, and is the distribution function of the completed image
[0086] It should be noted that the above loss functions are optional loss functions in the present disclosure. In practical applications, one or more of them can be selected, or one or more of them can be simply deformed or combined to achieve the training purpose corresponding to the above loss functions.
[0087] In a possible implementation manner, all the above loss functions can be combined to set the objective function L of the training total as:
[0088]
[0089] where is the structural loss function, which is a deformed form of the reconstruction loss function. Specifically, the structural loss function can be set as is the texture loss function, which is obtained by combining the perceptual loss function, style loss function, total variational loss function, and adversarial loss function. Specifically, where k represents that the loss function is calculated at the k-th layer of the decoder, and λ recRepresents the weight factor, which can be adjusted according to training requirements. P and Q are the preset number of layers of the decoder. After experimental verification, choosing P as {1,2,3,4,5,6} and Q as {1,2,3} can achieve better training results.
[0090] Figure 3 is a structural schematic diagram of an image completion model according to an exemplary disclosed embodiment. Figure 3 As shown, the image completion model includes an image encoder and a fusion module, and the fusion module also includes a decoder and an attention submodule. After the image to be completed and the position information to be completed are input into the image completion model, they are encoded by the image encoder, and the image encoding is decoded by the decoder, and the output value of the preset layer of the decoder is processed by the attention submodule embedded in the preset layer of the decoder to obtain channel features. The fusion module convolves the channel features to obtain channel features in the format of C×h×w; on the other hand, the fusion module convolves the image to be completed and the position information to be completed to obtain spatial features, and adjusts the spatial features to obtain spatial features in the format of C×h×w; based on the fusion function A, the spatial features and channel features are fused to obtain the fusion feature α, and the fusion feature α is processed based on the completion function B to obtain the completed image.
[0091] After the image completion model is trained, the image with any missing area is put into the image completion model, and the content and style of the original image can be obtained. Figure 1 The completed image has consistent and natural missing edges.
[0092] Through the above technical solution, at least the following technical effects can be achieved:
[0093] The image to be completed is processed by the image completion model to obtain channel features and spatial features, and a complete image is generated based on the channel features and spatial features. This can achieve image completion while maintaining the image meaning and considering the missing position of the image, reduce artifacts and chromatic aberration, and make the generated completed image more natural.
[0094] Figure 4 is a block diagram of an image completion device according to an exemplary disclosed embodiment. Figure 4 As shown, the device 400 includes:
[0095] The position determination module 410 is used to determine the position information to be completed on the image to be completed.
[0096] The input module 420 is used to input the image to be completed and the position information to be completed into the image completion model.
[0097] An acquisition module 430, configured to acquire the complete image generated by the image completion model.
[0098] Wherein, the image completion model includes an image encoder and a fusion module. The image encoder is configured to generate an image encoding based on the image to be completed and the to-be-completed position information; the fusion module is configured to generate channel features based on the image encoding, generate spatial features based on the image to be completed and the to-be-completed position information, and generate the complete image based on the channel features and the spatial features.
[0099] Optionally, the fusion module includes a decoder and an attention sub-module. The decoder is configured to generate a plurality of hierarchical features based on the image encoding, and the attention sub-module is configured to process the hierarchical features of a plurality of preset hierarchies among the plurality of hierarchical features to obtain the channel features.
[0100] Optionally, the fusion module includes an image convolution sub-module, configured to perform convolution based on the image to be completed and the to-be-completed position information to obtain the spatial features.
[0101] Optionally, the fusion module further includes a synthesis sub-module, configured to adjust the formats of the channel features and the spatial features to be consistent; the fusion module is further configured to generate a fusion feature based on the adjusted channel features and the channel features, and generate the complete image based on the fusion feature through a preset image generation algorithm.
[0102] Optionally, the device further includes a training module, configured to input multiple first sample images into a training encoder to obtain multiple key values output by the training encoder; input a second sample image corresponding to the first sample image into the image encoder to obtain multiple query values output by the image encoder, wherein the corresponding first sample image and the second sample image are images obtained by removing different parts from the same original image; based on the query values and the key values, adjust the parameters of the image encoder and the parameters of the training encoder through a preset encoder training loss function.
[0103] Optionally, the training module is further configured to adjust the parameters of the training encoder and the image encoder in a form of momentum update.
[0104] Optionally, the image completion model is trained by one or more of a reconstruction loss function, a perceptual loss function, a style loss function, a total variation loss function, and an adversarial loss function.
[0105] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0106] Through the above technical solutions, at least the following technical effects can be achieved:
[0107] By processing the image to be completed with an image completion model, channel features and spatial features are obtained, and a complete image is generated based on the channel features and spatial features. Image completion can be achieved while maintaining the meaning of the image and considering the missing positions of the image, reducing artifacts and color differences, and making the generated completed image more natural.
[0108] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment. As Figure 5 shown, the electronic device 500 may include: a processor 501, a memory 502. The electronic device 500 may further include one or more of a multimedia component 503, an input / output (I / O) interface 504, and a communication component 505.
[0109] Among them, the processor 501 is used to control the overall operation of the electronic device 500 to complete all or part of the steps in the above image completion method. The memory 502 is used to store various types of data to support the operation of the electronic device 500. These data may include, for example, instructions for any application or method operating on the electronic device 500, as well as application-related data, such as contact data, received and sent messages, pictures, audio, video, and so on. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. The multimedia component 503 may include a screen and an audio component. Among them, the screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 502 or sent through the communication component 505. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 504 provides an interface between the processor 501 and other interface modules. The above other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 505 is used for wired or wireless communication between the electronic device 500 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, near field communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited here. Therefore, the corresponding communication component 505 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.
[0110] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-described image completion method.
[0111] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-described image completion method are implemented. For example, the computer-readable storage medium may be the above-described memory 502 including program instructions, and the above program instructions may be executed by the processor 501 of the electronic device 500 to complete the above-described image completion method.
[0112] The preferred embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure. For example, each loss function can be deformed, or the outputs of different layers of the decoder can be selected to achieve the above effects.
[0113] In addition, it should be noted that, in the above specific embodiments, the various specific technical features described can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present disclosure will not separately describe various possible combination manners.
[0114] Furthermore, any combination can be made between various different embodiments of the present disclosure as long as it does not violate the idea of the present disclosure, and it should also be regarded as the content disclosed by the present disclosure.
Claims
1. An image completion method, characterized in that, The method includes: determining the to-be-completed position information on the to-be-completed image; inputting the to-be-completed image and the to-be-completed position information into an image completion model; obtaining the complete image generated by the image completion model; wherein, the image completion model includes an image encoder and a fusion module, and the step of the image completion model generating the complete image includes: generating an image encoding based on the to-be-completed image and the to-be-completed position information by the image encoder; generating channel features based on the image encoding by the fusion module, generating spatial features based on the to-be-completed image and the to-be-completed position information, and generating the complete image based on the channel features and the spatial features; The image encoder is trained through the following steps: inputting multiple first sample images into a training encoder to obtain multiple key values output by the training encoder; inputting second sample images corresponding to the first sample images into the image encoder to obtain multiple query values output by the image encoder, wherein the corresponding first sample images and second sample images are images obtained by removing different parts from the same original image; adjusting the parameters of the image encoder and the parameters of the training encoder based on the query values and the key values through a preset encoder training loss function.
2. The method according to claim 1, wherein The fusion module includes a decoder and an attention sub-module. The step of generating channel features based on the image encoding by the fusion module includes: generating multiple hierarchical features based on the image encoding by the decoder, and processing the hierarchical features of multiple preset hierarchies among the multiple hierarchical features through the attention sub-module to obtain the channel features.
3. The method according to claim 1, wherein The fusion module includes an image convolution sub-module. Generating spatial features based on the to-be-completed image and the to-be-completed position information includes: performing convolution on the to-be-completed image and the to-be-completed position information through the image convolution sub-module to obtain the spatial features.
4. The method according to claim 1, wherein The fusion module further includes a synthesis sub-module. Generating the complete image based on the channel features and the spatial features includes: adjusting the formats of the channel features and the spatial features to be consistent through the synthesis sub-module; and generating a fusion feature based on the adjusted channel features and spatial features; and generating the complete image based on the fusion feature through a preset image generation algorithm.
5. The method according to claim 1, characterized in that, Adjusting the parameters of the image encoder and the training encoder by the preset encoder training loss function includes: adjusting the parameters of the training encoder and the image encoder in the form of momentum update.
6. The method according to claim 1, characterized in that, The image completion model is trained through one or more of a reconstruction loss function, a perceptual loss function, a style loss function, a total variation loss function, and an adversarial loss function.
7. An image completion device, characterized in that, The device includes: a position determination module for determining the position information to be completed on the image to be completed; an input module for inputting the image to be completed and the position information to be completed into an image completion model; an acquisition module for acquiring the complete image generated by the image completion model; wherein, the image completion model includes an image encoder and a fusion module, the image encoder is used to generate an image encoding based on the image to be completed and the position information to be completed; the fusion module is used to generate channel features based on the image encoding, generate spatial features based on the image to be completed and the position information to be completed, and generate the complete image based on the channel features and the spatial features, and the image encoder is trained through the following steps: inputting multiple first sample images into a training encoder to obtain multiple key values output by the training encoder; inputting a second sample image corresponding to the first sample image into the image encoder to obtain multiple query values output by the image encoder, wherein the corresponding first sample image and the second sample image are images obtained by removing different parts from the same original image; adjusting the parameters of the image encoder and the parameters of the training encoder based on the query values and the key values through a preset encoder training loss function.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method according to any one of claims 1-6.
9. An electronic device, characterized in that, Comprising: a memory having a computer program stored thereon; a processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image completion method and device and electronic equipment
CN111640076A
Image editing method and device, electronic equipment and storage medium
CN111814566A