Image generation method, device, equipment and storage medium
By identifying and repairing the foreground area in the image, combining the pixel filling and image repair model of the reference background area, the problem of low image authenticity in the prior art is solved, and more realistic image generation is achieved.
Patent Information
- Application Number
- CN202210429410.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-04-22
AI Technical Summary
In the prior art, when deleting the image foreground by cutting the image, the generated image is not very authentic.
By identifying the initial image, the foreground area to be processed is determined, and pixel filling is performed based on the reference background area, and the foreground area is repaired in combination with the image repair model to generate a target image.
While ensuring image generation efficiency, the authenticity of the generated target image is improved.
Smart Images

Figure CN115131464B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to an image generation method, apparatus, device, and storage medium. Background Art
[0002] With the development of artificial intelligence-based image processing technology, some users want to use image processing technology to delete the foreground part in an image and retain the background part in the image for combination with other foreground parts to obtain a new image. For example, some users want to use the background in a PPT (PowerPoint) image obtained by screenshot. In this case, it is necessary to eliminate the foreground image added to the background of the PPT image.
[0003] In the related art, the foreground image on the background is often deleted by means of matting, and then pixel points are filled in the deleted area according to a preset rule, so as to achieve the purpose of eliminating the foreground image on the background.
[0004] However, when filling pixel points according to a preset rule, the authenticity of the generated image after eliminating the foreground image is not high. Summary of the Invention
[0005] Embodiments of this application provide an image generation method, apparatus, device, and storage medium, which can improve the authenticity of the generated image. The technical solution is as follows:
[0006] On the one hand, an image generation method is provided. The method includes:
[0007] Identifying an initial image to obtain at least one first foreground area to be processed in the initial image;
[0008] Based on at least one reference background area surrounding the first foreground area in the initial image, performing pixel filling on a target foreground area in the initial image to obtain a first processed image, where the target foreground area is a first foreground area of a first type among the at least one first foreground area;
[0009] Based on the first processed image and an image restoration model, performing image restoration on a second foreground area in the first processed image to obtain a second processed image, where the second foreground area is a first foreground area other than the target foreground area among the at least one first foreground area;
[0010] Based on the first processed image and the second processed image, generating a target image, where the target image is a background image after removing the at least one first foreground area from the initial image.
[0011] On the one hand, an image generation apparatus is provided. The apparatus includes:
[0012] An image recognition module, configured to recognize an initial image to obtain at least one first foreground region to be processed in the initial image;
[0013] A pixel filling module, configured to perform pixel filling on a target foreground region in the initial image based on at least one reference background region surrounding the first foreground region in the initial image, to obtain a first processed image, where the target foreground region is a first foreground region of a first type among the at least one first foreground region;
[0014] An image restoration module, configured to perform image restoration on a second foreground region in the first processed image based on the first processed image and an image restoration model, to obtain a second processed image, where the second foreground region is a first foreground region other than the target foreground region among the at least one first foreground region;
[0015] An image generation module, configured to generate a target image based on the first processed image and the second processed image, where the target image is a background image after removing the at least one first foreground region from the initial image.
[0016] In a possible implementation manner, the pixel filling module is configured to determine the target foreground region from the at least one first foreground region based on pixel values of pixel points in the at least one reference background region; fill the target foreground region with a target pixel value in the initial image to obtain a first processed image, where the target pixel value is determined based on pixel values of pixel points in the reference background region corresponding to the target foreground region.
[0017] In a possible implementation manner, the pixel filling module is configured to perform histogram statistics on pixel values of pixel points in the at least one reference background region to obtain a pixel histogram of the at least one reference background region; determine the type of the at least one first foreground region based on the pixel histogram of the at least one reference background region; and determine the first foreground region of the first type among the at least one first foreground region as the target foreground region.
[0018] In a possible implementation, the pixel filling module is configured to slide a sliding window on the pixel histogram of any one of the at least one reference background region; in response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram meeting a quantity condition, determining the type of the first foreground region surrounded by the reference background region as the first type; in response to the number of pixel points in the reference background region covered by the sliding window on the pixel histogram not meeting the quantity condition, determining the type of the first foreground region surrounded by the reference background region as the second type, where the second type is different from the first type.
[0019] In a possible implementation, the apparatus further includes:
[0020] A pixel value determination module, configured to divide the pixel points in the reference background region corresponding to the target foreground region into multiple pixel point sets, where the multiple pixel point sets correspond to multiple pixel value intervals; determine a target pixel point set from the multiple pixel point sets, where the target pixel point set is the pixel point set with the largest number of pixel points among the multiple pixel point sets; and determine the median of the pixel value interval corresponding to the target pixel point set as the target pixel value.
[0021] In a possible implementation, the image inpainting module is configured to generate a mask image based on the second foreground region, where the mask image is used to represent the position of the second foreground region in the first processed image; superimpose the mask image on the first processed image to obtain a third processed image; splice the third processed image and the mask image to obtain a spliced image; and input the spliced image into the image inpainting model, and perform downsampling, feature extraction, and upsampling on the spliced image through the image inpainting model to output the second processed image.
[0022] In a possible implementation, the image inpainting module is configured to perform Fourier convolution on the spliced image through the image inpainting model to obtain a downsampled image of the spliced image; perform Fourier convolution and residual connection on the downsampled image through the image inpainting model to obtain image features of the downsampled image; and perform deconvolution on the image features of the downsampled image through the image inpainting model to output the second processed image.
[0023] In a possible implementation manner, the image inpainting module is configured to process the spliced image by using a Fourier convolutional unit through the image inpainting model to obtain a local feature map and a global feature map of the spliced image; splice the local feature map and the global feature map of the spliced image to obtain a downsampled image of the spliced image.
[0024] In a possible implementation manner, the image generation module is configured to fuse the first processed image and the second processed image based on a mask image to obtain the target image, where the mask image is generated based on the position of the second foreground region in the first processed image.
[0025] In a possible implementation manner, the image recognition module is configured to input the initial image into an image recognition model, extract features of the initial image through the image recognition model to obtain a feature map of the initial image; slide a sliding window on the feature map of the initial image through the image recognition model to obtain a plurality of sub-feature maps of the initial image, where the plurality of sub-feature maps correspond to a plurality of regions of the initial image; classify the plurality of regions based on the plurality of sub-feature maps through the image recognition model, and output at least one first foreground region among the plurality of regions, where the first foreground region is a foreground region among the plurality of regions.
[0026] In a possible implementation manner, the device further includes:
[0027] A model training module, configured to obtain a first sample image and a second sample image, where the second sample image is a sample image with a blank foreground region randomly generated in the first sample image; input the second sample image into the image inpainting model, repair the blank foreground region in the second sample image through the image inpainting model, and output a sample repaired image; train the image inpainting model based on the difference information between the first sample image and the sample repaired image.
[0028] In a possible implementation manner, the model training module is configured to perform at least one of the following:
[0029] Input the first sample image and the sample repaired image into a discriminant model, and output a authenticity parameter of the sample repaired image through the discriminant model based on first difference information between the first sample image and the sample repaired image, where the authenticity parameter is used to represent the authenticity of the sample repaired image; train the image inpainting model based on the authenticity parameter;
[0030] Input the first sample image and the sample restored image into a feature extraction model. Through the feature extraction model, perform feature extraction on the first sample image and the sample restored image, and output the first sample image features of the first sample image and the sample restored image features of the sample restored image; based on the second difference information between the first sample image features and the sample restored image features, train the image restoration model.
[0031] On the one hand, a computer device is provided. The computer device includes one or more processors and one or more memories. At least one computer program is stored in the one or more memories. The computer program is loaded and executed by the one or more processors to implement the image generation method.
[0032] On the one hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by a processor to implement the image generation method.
[0033] On the one hand, a computer program product is provided. The computer program product includes program code. The program code is stored in a computer-readable storage medium. A processor of a computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above image generation method.
[0034] Through the technical solution provided by the embodiments of the present application, when generating a target image, the initial image is recognized to obtain a first foreground area to be processed, which is also the foreground area to be processed. Pixel filling is performed on the first foreground area based on a reference background area adjacent to the first foreground area to obtain a first processed image of the target foreground area. This processing method is also a policy-based processing method. Image restoration is performed on a second foreground area in the first restored image through an image restoration model to obtain a second processed image. This processing method is also a model-based processing method. Finally, based on the first processed image and the second processed image, a target image eliminating the first foreground area can be obtained. Combining the policy-plus-model processing method can improve the authenticity of the generated target image while ensuring the image generation efficiency. Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a schematic diagram of the implementation environment of an image generation method provided by an embodiment of this application;
[0037] Figure 2 It is a flowchart of an image generation method provided by an embodiment of this application;
[0038] Figure 3 It is a flowchart of another image generation method provided by an embodiment of this application;
[0039] Figure 4 It is a schematic diagram of a PPT image provided by an embodiment of this application;
[0040] Figure 5 It is a schematic diagram of an initial image provided by an embodiment of this application;
[0041] Figure 6 It is a schematic diagram of another PPT image provided by an embodiment of this application;
[0042] Figure 7 It is a schematic diagram of the structure of an image repair model provided by an embodiment of this application;
[0043] Figure 8 It is a schematic diagram of the structure of a residual unit provided by an embodiment of this application;
[0044] Figure 9 It is a schematic diagram of the structure of a Fourier convolution unit provided by an embodiment of this application;
[0045] Figure 10 It is a schematic diagram of the structure of a frequency domain transformation sub-unit provided by an embodiment of this application;
[0046] Figure 11 It is a flowchart of yet another image generation method provided by an embodiment of this application;
[0047] Figure 12 It is a schematic diagram of the effect of an image repair provided by an embodiment of this application;
[0048] Figure 13 It is a flowchart of a training method of an image repair model provided by an embodiment of this application;
[0049] Figure 14 It is a schematic diagram of the effect of another image repair provided by an embodiment of this application;
[0050] Figure 15 It is a schematic diagram of the structure of an image generation device provided by an embodiment of this application;
[0051] Figure 16 It is a schematic diagram of the structure of a terminal provided by an embodiment of this application;
[0052] Figure 17 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific implementation manners
[0053] To make the purpose, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the drawings.
[0054] In the present application, terms such as "first" and "second" are used to distinguish identical items or similar items with basically the same functions. It should be understood that there is no logical or temporal dependence between "first", "second", and "nth", nor are the quantity and execution order limited.
[0055] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0056] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0057] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge sub-models to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0058] Semantic feature: A feature used to represent the semantics expressed by text. Different texts can correspond to the same semantic feature. For example, the texts "What's the weather like today" and "How's the weather today" can correspond to the same semantic feature. A computer device can map the characters in a text to character vectors, and based on the relationships between the characters, combine and operate on the character vectors to obtain the semantic feature of the text. For example, the computer device can adopt the Bidirectional Encoder Representations from Transformers (BERT) of the codec.
[0059] Mask: A mask is a string of binary codes that performs a product operation on the target field to mask or display a certain character in the target field. For example, if the target field is (1, 1, 0, 1) and the mask is (1, 0, 1, 0), after performing the product operation on the target field and the mask, we get (1, 0, 0, 0). That is to say, the first and third characters in the target field are retained, and the second and fourth characters are "masked" and become 0. Through the mask, we can know the characters retained and "masked" in the target field.
[0060] Normalization processing: Maps a sequence of numbers with different value ranges to the interval (0, 1) to facilitate data processing. In some cases, the normalized values can be directly implemented as probabilities.
[0061] Image inpainting: A technology that repairs and fills specified positions in an image to make the image complete and reasonable.
[0062] Generative Adversarial Networks (GAN): A deep learning model that trains the generative model through the mutual game learning of the generative model and the discriminative model.
[0063] Fourier convolution: A network structure that applies the Fourier transform to the convolutional network and is designed to replace the ordinary convolution.
[0064] It can be understood that in the specific implementation of this application, when it comes to data related to images, etc., when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0065] Figure 1 It is a schematic diagram of the implementation environment of an image generation method provided by the embodiments of this application. See Figure 1 In this implementation environment, it may include a terminal 110 and a server 140.
[0066] The terminal 110 is connected to the server 140 via a wireless network or a wired network. Optionally, the terminal 110 is a vehicle-mounted terminal, a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc., but is not limited thereto. The terminal 110 installs and runs an application program that supports image generation.
[0067] The server 140 is an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The server 140 provides background services for the application program running on the terminal 110.
[0068] Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminal is only one, or the above terminals are dozens or hundreds, or more, and other terminals are also included in the above implementation environment at this time. The embodiments of the present application do not limit the number and device type of the terminals.
[0069] After introducing the implementation environment of the embodiments of the present application, the application scenarios of the embodiments of the present application will be described below in combination with the above implementation environment. In the following description process, the terminal is also the terminal 110 in the above implementation environment, and the server is also the server 140 in the above implementation environment.
[0070] The technical solution provided by the embodiments of the present application can be applied to the scenario of removing the foreground image superimposed on the background image. For example, it can be applied to the scenario of removing the chart in the PPT image. Here, the PPT image refers to the screenshot or the photographed PPT image, the background image is the background of the PPT image, the foreground image is the chart, and the chart includes pictures, graphics, and tables; or it can be applied to the scenario of removing the portrait in the landscape image; or it can be applied to the scenario of removing the stain in the image.
[0071] When the technical solution provided by the embodiments of the present application is applied to the scenario of removing the chart in the PPT image, the user selects the PPT image with the chart to be removed through the terminal, and the terminal sends the selected PPT image to the server. The server uses the technical solution provided by the embodiments of the present application to process the PPT image, and obtains the processed PPT image. There is no chart in the processed PPT image, and the processed PPT image is also the background image of the PPT image selected by the user. The server sends the processed PPT image to the terminal, and the terminal displays the processed PPT image to the user. In some embodiments, the user can add other charts to the processed PPT image to synthesize a new PPT image.
[0072] When the technical solution provided in the embodiment of the present application is applied to the scenario of removing human figures from a landscape image, the user selects a landscape image with a human figure to be removed through the terminal, and the terminal sends the selected landscape image to the server. The server processes the landscape image by using the technical solution provided in the embodiment of the present application to obtain a processed landscape image. There is no human figure in the processed landscape image, and the processed landscape image is also the background image of the landscape image selected by the user. The server sends the processed landscape image to the terminal, and the terminal displays the processed landscape image to the user. In some embodiments, the user can add other human figures to the processed landscape image to synthesize a new landscape image.
[0073] It should be noted that the above is described by taking the technical solution provided in the embodiment of the present application being applied to the scenarios of removing charts from PPT images and removing human figures from landscape images as examples. In other possible implementation manners, the technical solution provided in the embodiment of the present application can also be applied to other image processing scenarios, and the embodiment of the present application does not limit this.
[0074] After introducing the implementation environment and application scenarios of the embodiment of the present application, the image generation method provided in the embodiment of the present application will be described below. Refer to Figure 2 , the technical solution provided in the embodiment of the present application can be executed by the terminal or the server, or jointly executed by the terminal and the server. In the embodiment of the present application, taking the execution subject as the server as an example for description, the method includes:
[0075] 201. The server identifies the initial image to obtain at least one first foreground area to be processed in the initial image.
[0076] Wherein, the initial image is an image to be processed. In different scenarios, the type of the initial scenario is different. For example, the initial scenario image is a PPT image containing charts, or a landscape image containing human figures, etc. At least one first foreground area in the initial image is the foreground area to be processed in the initial image. When the initial image is a PPT image, the at least one first foreground area is the chart in the PPT image.
[0077] 202. The server performs pixel filling on the target foreground area in the initial image based on at least one reference background area surrounding the first foreground area in the initial image to obtain a first processed image, where the target foreground area is the first foreground area of the first type in the at least one first foreground area.
[0078] Among them, the reference background area surrounds the first foreground area, that is, the reference background area is the background area around the first foreground area. Pixel filling of the target foreground area means filling pixel points in the target foreground area to obtain a target foreground area composed of the refilled pixel points. For example, if the initial image is a PPT image and the target foreground area is an icon in the PPT image, after pixel filling the target foreground area with pixel points of the target color, the target foreground area becomes an area of the target color and the original icon is completely replaced. The types of the first foreground area include a first type and a second type. Among them, the first foreground area of the first type can be processed based on the reference background area, and the first foreground area of the second type is processed by the image restoration model in the subsequent steps.
[0079] 203. The server performs image restoration on the second foreground area in the first processed image based on the first processed image and the image restoration model to obtain a second processed image, where the second foreground area is the first foreground area in the at least one first foreground area other than the target foreground area.
[0080] Among them, the second foreground area is the first foreground area in the at least one first foreground area other than the target foreground area, that is, the remaining first foreground area after the above step 202 is processed. The image restoration model is an image restoration model trained based on a first sample image and a second sample image and has the ability to restore images. The second sample image is a sample image with a blank foreground area randomly generated in the first sample image. During the training process, with the first sample image as the supervision, the purpose of training the image restoration model is to eliminate the blank foreground area in the second sample image.
[0081] 204. The server generates a target image based on the first processed image and the second processed image, where the target image is the background image after removing the at least one first foreground area from the initial image.
[0082] Among them, the first processed image is the processed image obtained after processing the target foreground area, and the second processed image is the first processed image after processing the second foreground area. In the target image generated based on the first processed image and the second processed image, both the target foreground area and the second foreground area are eliminated, that is, all the first foreground areas in the initial image are eliminated.
[0083] Through the technical solution provided by the embodiments of the present application, when generating a target image, the initial image is recognized to obtain a first foreground area to be processed, which is also the foreground area to be processed. The first foreground area is pixel-filled based on a reference background area adjacent to it to obtain a first processed image of the target foreground area. This processing method is also a policy-based processing method. The second foreground area in the first repaired image is image-repaired through an image repair model to obtain a second processed image. This processing method is also a model-based processing method. Finally, based on the first processed image and the second processed image, a target image with the first foreground area eliminated can be obtained. Combining the policy-plus-model processing method can improve the authenticity of the generated target image while ensuring the image generation efficiency.
[0084] The above steps 201-204 are a brief introduction to the technical solution provided by the embodiments of the present application. Below, some examples will be combined to provide a more detailed description of the technical solution provided by the embodiments of the present application. See Figure 4 The technical solution provided by the embodiments of the present application can be executed by a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the technical solution being executed by the server as an example for illustration, the method includes:
[0085] 301. The server obtains an initial image.
[0086] Among them, the initial scene image is a PPT image containing a chart, or a landscape image containing a portrait, etc. The embodiments of the present application do not limit this.
[0087] In a possible implementation manner, in response to an operation on the initial image, the terminal sends the initial image to the server. In this implementation manner, the user can control the terminal to send the initial image to the server through the operation on the initial image, and the user can select the initial image by himself, and the efficiency of human-computer interaction is relatively high.
[0088] For example, the terminal displays an image selection page, and the image selection page includes multiple candidate images. In response to a click operation on the initial image among the multiple candidate images, the terminal sends the initial image to the server, and the server obtains the initial image. In this case, the multiple candidate images are images stored on the terminal. In some embodiments, when the multiple candidate images are images stored on the server, in response to a click operation on the initial image on the image selection page, the terminal sends an image selection instruction to the server, and the image selection instruction carries the identifier of the initial image. After receiving the image selection instruction, the server obtains the identifier of the initial image from the image selection instruction. The server queries based on the identifier of the initial image to obtain the initial image.
[0089] For example, taking this initial image as a PPT image, correspondingly, the multiple candidate images are multiple candidate PPT images, and the multiple candidate PPT images are all PPT images obtained by screenshot or photograph. For example, referring to Figure 4 , a PPT image 400 is provided. In this PPT image 400, there are three foreground images 401-403. In some embodiments, the multiple candidate PPT images belong to the same PPT file. The terminal displays an image selection page, and the image selection page includes multiple candidate PPT images. In response to a click operation on any one of the multiple candidate PPT images, the terminal sends the clicked candidate PPT image to the server, and the selected candidate PPT image is also the initial image. The server receives the candidate PPT image.
[0090] In some embodiments, the image selection page is a page of an online document. In this case, the online document provides a function for processing PPT images, and the user can select a PPT image for processing through the image selection page of the online document.
[0091] 302. The server recognizes the initial image to obtain at least one first foreground region to be processed in the initial image.
[0092] Among them, at least one first foreground region in the initial image is the foreground region to be processed in the initial image. When the initial image is a PPT image, the at least one first foreground region is a chart in the PPT image. In some embodiments, processing the at least one first foreground region means deleting the at least one first foreground region.
[0093] In a possible implementation manner, the server inputs the initial image into an image recognition model, extracts features of the initial image through the image recognition model to obtain a feature map of the initial image. The server uses a sliding window to slide on the feature map of the initial image through the image recognition model to obtain multiple sub-feature maps of the initial image, and the multiple sub-feature maps correspond to multiple regions of the initial image. The server classifies the multiple regions based on the multiple sub-feature maps through the image recognition model and outputs the at least one first foreground region in the multiple regions, and the first foreground region is the foreground region in the multiple regions.
[0094] Among them, when classifying the initial image through the image recognition model, the multiple regions of the initial image can be divided into a first foreground region and a background region. The image recognition model can both recognize the first foreground region in the multiple regions and recognize the background region in the multiple regions. When training the image recognition model, the first foreground region in the sample image can be used as a positive sample, and the background region can be used as a negative sample to train the image recognition model.
[0095] In some embodiments, the image recognition module is an object detection model that can label the object to be recognized in the image. In the embodiments of the present application, the object is also the first foreground region. The image recognition model is trained based on sample images including sample objects and the position information of the sample objects in the sample images. The purpose of training the image recognition model is that after an image including an object is input into the image recognition model, the image recognition model can output the position information of the object in the image. In some embodiments, the position information of the object is the coordinates of the object in the image, such as the pixel coordinates of the object in the image. In the embodiments of the present application, the object to be recognized is also the at least one foreground region. Taking the initial image as a PPT image including a chart as an example, after the PPT image is input into the image recognition model, the image recognition model can output the position information of the chart in the PPT image. Correspondingly, in the case where the initial image is a landscape image including a portrait, after the initial image is input into the image recognition model, the image recognition model can output the position information of the portrait in the landscape image.
[0096] In this implementation manner, inputting the initial image into the image recognition model can directly obtain at least one first foreground region in the initial image, without the need for the user to manually label the first foreground region in the initial image, which greatly improves the efficiency of human-computer interaction. At the same time, using the trained image recognition model to obtain at least one first foreground region in the initial image also has relatively high accuracy.
[0097] To illustrate the above implementation manner more clearly, the above implementation manner will be described in three parts below.
[0098] Part 1: The server inputs the initial image into the image recognition model, and the image recognition model extracts features from the initial image to obtain the feature map of the initial image.
[0099] In a possible implementation manner, the server inputs the initial image into the image recognition model, and the convolutional layer of the image recognition model performs convolution on the initial image to obtain the feature map of the initial image.
[0100] In this implementation manner, the server can extract features from the initial image through the convolutional layer of the image recognition model. Since the speed of convolution operation is relatively fast, the efficiency of obtaining the feature map of the initial image in this way is relatively high.
[0101] For example, the image recognition model includes a feature extraction unit, and the server inputs the initial image into the feature extraction unit of the image recognition model. The server performs convolution on the initial image through the convolutional layer of the feature extraction unit, that is, slides the convolutional kernel on the convolutional layer of the feature extraction unit over the initial image. During the sliding process, the convolutional kernel performs convolution with the area of the initial image covered by the convolutional kernel to obtain the feature map of the initial image. Here, the number of convolutional kernels on the convolutional layer of the feature extraction unit is one or more, and this application embodiment does not limit this. When the initial image includes multiple color channels, the number of convolutional kernels on the convolutional layer of the feature extraction unit is also multiple, and the multiple convolutional kernels are used to extract features from the multiple color channels of the initial image to obtain the color feature maps of the respective color channels. The server fuses the color feature maps of the multiple color channels through the feature extraction unit to obtain the feature map of the initial image. In some embodiments, when the server fuses the color feature maps of the multiple color channels through the feature extraction unit to obtain the feature map of the initial image, a weighted summation method can be used. In some embodiments, the feature extraction unit is a feature extractor based on Convolutional Neural Networks (CNN), such as a neural network Resnet-101 (Residual Network 101) pre-trained on the large-scale open-source dataset Imagenet (ImageNet).
[0102] In a possible implementation manner, the server inputs the initial image into the image recognition model, and through the image recognition model, encodes the initial image based on the attention mechanism to obtain the feature map of the initial image.
[0103] In this implementation manner, the server can perform feature extraction based on the attention mechanism through the image recognition model. The introduction of the attention mechanism enables the image recognition model to focus on the regions with a higher degree of importance in the initial image, and the obtained feature map can more accurately reflect the features of the initial image.
[0104] For example, the image recognition model includes a feature extraction unit, and the server inputs the initial image into the feature extraction unit of the image recognition model. The server encodes multiple parts of the initial image based on the attention mechanism through the feature extraction unit to obtain the feature map of the initial image. For instance, the server inputs the initial image into the feature extraction unit, and through the feature extraction unit, embeds and encodes multiple parts of the initial image to obtain multiple embedding vectors. One embedding vector corresponds to one part of the initial image, and the embedding vectors are used to represent the positions of the respective parts in the initial image and the content of the respective parts. The server inputs the multiple embedding vectors into the feature extraction unit, and through three linear transformation matrices of the feature extraction unit, performs linear transformation on the multiple embedding vectors to obtain the query vector, key vector, and value vector corresponding to each part of the initial image. The server obtains the attention weights of multiple parts of the initial image through the feature extraction unit based on the query vectors and key vectors corresponding to multiple parts of the initial image. The server obtains the attention encoding vectors of the respective parts of the initial image through the feature extraction unit based on the attention weights of the respective parts of the initial image and the value vectors of the respective parts of the initial image. The attention encoding vectors of the respective parts of the initial image are combined according to the positions of the respective parts in the initial image to obtain the feature map of the initial image.
[0105] In a possible implementation manner, the server inputs the initial image into the image recognition model, and performs a full connection on the initial image through the fully connected layer of the image recognition model to obtain the feature map of the initial image. In this implementation manner, the server can obtain the feature map of the initial image by performing a full connection on the initial image through the image recognition model. The full connection can extract the features of the initial image as a whole, and the obtained feature map can also reflect the features of the initial image as a whole.
[0106] For example, the image recognition model includes a feature extraction unit, and the server inputs the initial image into the feature extraction unit of the image recognition model. The server performs a full connection on the initial image through the fully connected layer of the feature extraction unit to obtain the feature map of the initial image. In some embodiments, the feature extraction model is a feature extractor based on Deep Neural Networks (DNN).
[0107] It should be noted that the server can obtain the feature map of the initial image through any of the above methods. Of course, with the development of science and technology, the server can also use other methods to obtain the feature map, and the embodiments of the present application do not limit this.
[0108] Second part: The server slides a sliding window on the feature map of the initial image through the image recognition model to obtain multiple sub-feature maps of the initial image.
[0109] Among them, the multiple sub-feature maps correspond to multiple regions of the initial image, and the multiple regions include both the background region in the initial image and the foreground region in the initial image.
[0110] In a possible implementation manner, the server slides a sliding window of a target size on the feature map of the initial image through the image recognition model, and collects multiple sub-feature maps covered by the sliding window during the sliding process.
[0111] Among them, the target size is also the size of the obtained sub-feature map. When the size of the feature map of the initial image remains unchanged, the larger the target size, the larger the size of the collected sub-feature map. The target size is set by the technical personnel according to the actual situation, and the embodiments of the present application do not limit this.
[0112] Third part: The server classifies the multiple regions based on the multiple sub-feature maps through the image recognition model, and outputs at least one first foreground region among the multiple regions.
[0113] In a possible implementation, the image recognition model includes a classification unit. The server inputs the multiple sub-feature maps into the classification unit of the image recognition model, and classifies the multiple sub-feature maps through the classification unit to obtain the types of the regions corresponding to the multiple sub-feature maps. The server determines the position information of at least one first foreground region in the multiple regions according to the types of the regions corresponding to the multiple sub-feature maps. For example, for any one of the multiple sub-feature maps on the multiple sub-feature maps, after the server inputs the sub-feature map into the classification unit, it fully connects the sub-feature map through the fully connected layer of the classification unit to obtain the classification feature of the sub-feature map. The server inputs the classification feature of the sub-feature map into the normalization layer of the classification unit, and normalizes the classification feature of the sub-feature map through the normalization layer of the classification unit to obtain the probability sequence of the region corresponding to the sub-feature map. The values in the probability sequence are used to represent the probabilities that the region corresponding to the sub-feature map belongs to different types of regions. The server determines the type of the region corresponding to the sub-feature map based on the probability sequence of the region corresponding to the sub-feature map. In some embodiments, the types of the regions corresponding to the sub-feature maps include foreground regions and background regions. In this case, the function of the image recognition model is to identify the foreground regions and background regions in the initial image. In the following description process, the first foreground region is used to represent the foreground region identified by the image recognition model in the initial image. The server determines the position information of at least one first foreground region in the multiple regions. In some embodiments, the position information is the coordinates of the at least one first foreground region in the initial image.
[0114] 303. The server performs pixel filling on the target foreground region in the initial image based on at least one reference background region surrounding the first foreground region in the initial image to obtain a first processed image, where the target foreground region is the first foreground region of the first type among the at least one first foreground region.
[0115] Among them, the reference background region surrounds the first foreground region, that is, the reference background region is the background region around the first foreground region. In the case where any first foreground region is a square region, the reference background region surrounding the first foreground region is a hollow square region surrounding the square region, and the shape and size of the hollow part of the hollow square region are the same as those of the square region. For example, see Figure 5 , the initial image 500 includes the first foreground region 501 and the reference background region 502 surrounding the first foreground region 501. The server performing pixel filling on the target foreground region means filling pixel points in the target foreground region to obtain a target foreground region composed of the refilled pixel points. For example, see Figure 6, the initial image is the PPT image 600, and the target foreground region is an icon 601 in the PPT image. After pixel-filling the target foreground region with pixels of the target color, the target foreground region becomes the region 602 of the target color, and the original icon 601 is completely replaced. The types of the first foreground regions include a first type and a second type. Among them, the first foreground regions of the first type can be processed based on the reference background region, and the first foreground regions of the second type are processed through the image restoration model in the subsequent steps.
[0116] In a possible implementation manner, the server determines the target foreground region from the at least one first foreground region based on the pixel values of the pixels in the at least one reference background region. The server fills the target foreground region in the initial image with a target pixel value to obtain a first processed image, and the target pixel value is determined based on the pixel values of the pixels in the reference background region corresponding to the target foreground region.
[0117] Among them, the reference background region corresponding to the target foreground region refers to the background region covered by the target foreground region on the initial image.
[0118] To explain the above implementation manner more clearly, the above implementation manner will be described in two parts below.
[0119] The first part: The server determines the target foreground region from the at least one first foreground region based on the pixel values of the pixels in the at least one reference background region.
[0120] In a possible implementation manner, the server performs a histogram statistics on the pixel values of the pixels in the at least one reference background region to obtain the pixel histogram of the at least one reference background region. The server determines the types of the at least one first foreground region based on the pixel histogram of the at least one reference background region. The server determines the first foreground regions of the first type in the at least one first foreground region as the target foreground region.
[0121] Among them, in the above implementation manner of determining the types of the at least one first foreground region based on the pixel histogram of the at least one reference background region, the at least one reference background region and the at least one first foreground region are in one-to-one correspondence, that is, one reference background region corresponds to the type of one first foreground region. The type of the first foreground region refers to the type of the background region covered by the first foreground region. In some embodiments, the types of the first foreground regions include a first type and a second type. The first type is also called a simple type, and the background region covered by the first foreground region of the simple type is a simple background region. The second type is also called a complex type, and the background region covered by the first foreground region of the complex type is a complex background region.
[0122] In this embodiment, the server can classify the first foreground region by means of histogram statistics. Since the pixel histogram obtained by histogram statistics can reflect the distribution of pixel values of pixel points in the reference background region, the classification of the first foreground region using the pixel histogram is also based on the distribution of pixel values, which facilitates subsequent processing of different types of first foreground regions in different ways.
[0123] For example, the server obtains at least one reference background region corresponding to the at least one first foreground region in the initial image. The server performs histogram statistics on the pixel values of pixel points in the at least one reference background region based on a target interval, and obtains the pixel histogram of the at least one reference background region. The target interval is a pixel value region, and the target interval is the bin width on the pixel histogram. The pixel histogram is used to show the distribution of pixel values of multiple pixel points in the reference background region. The target interval is set by those skilled in the art according to the actual situation, and the embodiments of the present application do not limit this. For any one of the at least one reference background regions, the server slides a sliding window on the pixel histogram of the reference background region. In response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram meeting the quantity condition, the server determines the type of the first foreground region surrounded by the reference background region as the first type. In response to the number of pixel points in the reference background region covered by the sliding window on the pixel histogram not meeting the quantity condition, the server determines the type of the first foreground region surrounded by the reference background region as the second type, and the second type is different from the first type. The server determines the first foreground region of the first type among the at least one first foreground region as the target foreground region.
[0124] For example, for any one of the at least one first foreground region, the server expands the edge of the first foreground region by a certain number of pixels and then performs cropping to obtain the cropped region of the first foreground region in the initial image. The cropped region includes the first foreground region and a reference background region surrounding the first foreground region. Here, the certain number of pixels is set by technicians according to the actual situation. For example, if the first foreground region is a square region with a size of 50 pixels × 50 pixels, the server expands the first foreground region by 20 pixels and then performs cropping, that is, obtains a cropped region with a size of 50 pixels × 50 pixels. The server performs a histogram statistics on the reference background region in the cropped region to obtain the pixel histogram of the reference background region. The pixel histogram shows the distribution of the pixel values of the pixel points in the reference background region. The server uses a sliding window to slide on the pixel histogram to determine the number of pixel points in the reference background region covered by the sliding window on the pixel histogram. Here, the number of pixel points in the reference background region covered by the sliding window on the pixel histogram refers to the total number of pixel points corresponding to the region covered by the sliding window on the pixel histogram. For example, the sliding window covers two statistical bars on the pixel histogram. Among them, the pixel value range corresponding to the first statistical bar is 100 - 150, and the number of pixel points is 50; the pixel value range corresponding to the second statistical bar is 150 - 200, and the number of pixel points is 65. In this case, the number of pixel points in the reference background region covered by the sliding window is 115. In response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram being greater than or equal to the pixel point number threshold, the server determines the type of the first foreground region surrounded by the reference background region as the first type. In response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram being less than the pixel point number threshold, the server determines the type of the first foreground region surrounded by the reference background region as the second type. The pixel point number threshold is related to the number of pixel points in the reference background region and is less than the number of pixel points in the reference background region, such as 80% of the number of pixel points in the reference background region, etc. In the case where the first foreground region is of the first type, the server determines the first foreground region as the target foreground region.
[0125] It should be noted that the above description is given by taking the initial image including one color channel as an example. In the case where the initial image includes multiple color channels, the server performs histogram statistics on the pixel values of the pixel points in the at least one reference background region under multiple color channels to obtain the pixel histograms of the at least one reference background region. One reference background region corresponds to multiple pixel histograms, and each pixel histogram corresponds to one color channel. The server determines the types of the at least one first foreground region based on the pixel histograms of the at least one reference background region. For any reference background region, that is, based on the multiple pixel histograms corresponding to the reference background region, the type of the first foreground region surrounded by the reference background region is determined. In the case where the types determined based on the multiple pixel histograms corresponding to the reference background region are all of the first type, the server determines the first foreground region surrounded by the reference background region as the target foreground region. The method for the server to determine the type of the first foreground region based on multiple pixel histograms belongs to the same inventive concept as the previous description, and the implementation process will not be elaborated here.
[0126] Second part: The server fills the target foreground region with a target pixel value in the initial image to obtain a first processed image.
[0127] Among them, filling the target foreground region with a target pixel value means replacing the pixel values of all pixel points in the target foreground region with the target pixel value, so that the color of the target foreground region as a whole becomes the color corresponding to the target pixel value.
[0128] For a clearer description of the above embodiments, the method for the server to determine the target pixel value will be described below.
[0129] In a possible implementation manner, the server divides the pixel points in the reference background region corresponding to the target foreground region into multiple pixel point sets, and the multiple pixel point sets correspond to multiple pixel value intervals. The server determines a target pixel point set from the multiple pixel point sets, and the target pixel point set is the pixel point set with the largest number of pixel points among the multiple pixel point sets. The server determines the median of the pixel value interval corresponding to the target pixel point set as the target pixel value.
[0130] In this implementation manner, the server can use the median of the pixel value interval corresponding to the target pixel point set as the target pixel value, that is, use the median of the pixel value interval corresponding to the pixel point set with the largest number of pixel points as the target pixel value, so that the target pixel value can more accurately represent the background region covered by the target foreground region.
[0131] For example, the server divides the pixel points in the reference background area corresponding to the target foreground area into multiple pixel point sets based on the pixel values of the pixel points in the reference background area. That is, for two pixel points in the reference background area, when the pixel values of the two pixel points both belong to the first pixel value interval, the two pixel points are divided into the pixel point set corresponding to the first pixel value interval. When the server performs a histogram statistics on the reference background area to obtain the pixel histogram of the reference background area, the multiple class intervals in the pixel histogram correspond to multiple pixel value intervals, and the multiple statistical bars correspond to multiple pixel point sets. The server determines the statistical bar with the largest number of corresponding pixel points in the pixel histogram as the target statistical bar, and the multiple pixel points corresponding to the target statistical bar constitute the target pixel point set. The server determines the median of the class interval corresponding to the target statistical bar as the target pixel value.
[0132] 304. The server performs image restoration on the second foreground area in the first processed image based on the first processed image and the image restoration model to obtain a second processed image, where the second foreground area is the first foreground area other than the target foreground area among the at least one first foreground area.
[0133] Among them, the second foreground area is the first foreground area other than the target foreground area among the at least one first foreground area, that is, the remaining first foreground area after the above step 303 is processed. The image restoration model is an image restoration model trained based on the first sample image and the second sample image, and has the ability to restore images. The second sample image is a sample image with a blank foreground area randomly generated in the first sample image. During the training process, with the first sample image as the supervision, the purpose of training the image restoration model is to eliminate the blank foreground area in the second sample image. The training process of the image restoration model will be described later. The second processed image is the image obtained by the initial image eliminating the second foreground area.
[0134] In a possible implementation manner, the server generates a mask image based on the second foreground area, and the mask image is used to represent the position of the second foreground area in the first processed image. The server superimposes the mask image on the first processed image to obtain a third processed image. The server splices the third processed image and the mask image to obtain a spliced image. The server inputs the spliced image into the image restoration model, and the image restoration model performs downsampling, feature extraction, and upsampling on the spliced image, and outputs the second processed image.
[0135] Among them, the mask image is used to cover the second foreground area in the first processed image. After superimposing the mask image and the first processed image, the second foreground area in the first processed image can be blocked. The size of the mask image is the same as that of the first processed image. The superposition of the mask image and the first processed image means placing the initial image and the mask image together with their edges completely coinciding.
[0136] In this embodiment, the server can process the second foreground area in the first processed image through an image inpainting model, so as to eliminate the second foreground image in the first processed image. The introduction of the model improves the efficiency of image processing.
[0137] To illustrate the above embodiment more clearly, the above embodiment will be described in the following three parts.
[0138] The first part: The server generates a mask image based on the second foreground area.
[0139] In a possible embodiment, the server generates the mask image based on the position information of the second foreground area in the first processed image. In some embodiments, the mask image is a binary image, that is, the mask image includes two colors, one is white and the other is black. In the mask image, white represents the area that does not need to be blocked, and black represents the area that needs to be blocked. In some embodiments, in the mask image, the position corresponding to the second foreground area is black, indicating that it needs to be blocked; the positions other than the second foreground area are white, indicating that they do not need to be blocked.
[0140] The second part: The server superimposes the mask image and the first processed image to obtain a third processed image.
[0141] Among them, in the third processed image, the second foreground area is covered, and the position where the second foreground area is located can be regarded as a "hole". The "hole" is also the target processed by the image processing model.
[0142] The third part: The server splices the third processed image and the mask image to obtain a spliced image.
[0143] Among them, the splicing here means combining the third processed image and the mask image into a spliced image. The spliced image includes all the information of the third processed image and the mask image. The size of the spliced image is the sum of the sizes of the third processed image and the mask image.
[0144] The fourth part: The server inputs the spliced image into the image inpainting model, and through the image inpainting model, downsamples, extracts features, and upsamples the spliced image, and outputs the second processed image.
[0145] In a possible implementation, the server performs Fourier convolution on the stitched image through the image inpainting model to obtain a downsampled image of the stitched image. The server performs Fourier convolution and residual connection on the downsampled image through the image inpainting model to obtain the image features of the downsampled image. The server performs deconvolution on the image features of the downsampled image through the image inpainting model and outputs the second processed image.
[0146] Among them, the image inpainting model can be regarded as an image generation model. By processing the stitched image through the image inpainting model, a background area can be generated at the position corresponding to the second foreground area in the first processed image, thereby eliminating the second foreground area in the first processed image. When using the stitched image for image processing, better image processing effects can be achieved. This is because if only the third processed image is used, there may be pixel points in the third processed image that interfere with image generation. For example, the mask image is a binary image, and there may be pixel points in the third processed image with the same color as the binary image. The existence of such pixel points may reduce the effect of image inpainting. Based on this, when using the image inpainting model for image inpainting, the stitched image obtained by stitching the third processed image and the mask image is used to ensure that the information in the third processed image and the mask image can be fully utilized to eliminate the above interference and improve the effect of image inpainting.
[0147] For example, the server processes the stitched image through the image inpainting model and uses the Fourier convolution unit to obtain the local feature map and global feature map of the stitched image. The server stitches the local feature map and global feature map of the stitched image through the image inpainting model to obtain the downsampled image of the stitched image. The server performs Fourier convolution on the downsampled image through the image inpainting model using the residual processing unit to obtain the feature map of the downsampled image. The server performs residual connection between the downsampled image and the feature map of the downsampled image through the image inpainting model using the residual processing unit to obtain the image features of the downsampled image. The server performs deconvolution on the image features of the downsampled image through the image inpainting model using the deconvolution unit and outputs the second processed image.
[0148] For example, see Figure 7 , the image inpainting model is divided into a downsampling part 701, a residual part 702, and an upsampling part 703. Among them, the downsampling (Down Sample) part 701 includes three Fourier convolution (FF Conv) units, the residual (ResBlocks) part 702 includes nine residual units (FFC ResBlock), and the upsampling (Up Sample) part 703 includes three deconvolution units. Among them, seeFigure 8 , the residual unit 800 includes two Fourier convolution sub-units 801 and 802 and a residual connection sub-unit 803. Refer to Figure 9 , the Fourier convolution unit 900 includes a local feature extraction branch and a global feature extraction branch. The local feature extraction branch is used to extract the local feature map 902 of the spliced image 901, and the global feature extraction branch is used to extract the global feature map 903 of the spliced image 901. After splicing the local feature map 902 and the global feature map 903, the downsampled image of the spliced image is obtained. Among them, both the local feature extraction branch and the global feature extraction branch include a convolution (Conv) sub-unit, a batch normalization (BN) sub-unit, and a rectified linear unit (ReLU) sub-unit. The global feature extraction branch further includes a frequency domain transformation sub-unit 904. The convolution sub-unit is used to perform convolution. In some embodiments, the size of the convolution kernel in the convolution unit is 3×3; the batch normalization sub-unit is used to perform batch normalization to eliminate the influence of the value range on the model result; the rectified linear unit sub-unit is used to perform activation using an activation function; the frequency domain transformation sub-unit 904 is used to extract features based on frequency domain transformation. Refer to Figure 10 , the frequency domain sub-transformation unit 1000 includes a convolution (Conv)-batch normalization (BN)-rectified linear unit (ReLU) layer 1001, a Fourier transform layer 1002, a convolution-batch normalization-rectified linear unit layer 1003, an inverse Fourier transform layer 1004, and a convolution layer 1005. Among them, the Fourier transform layer 1003 is used to perform a Fourier transform on the spliced image, that is, to implement a frequency domain transformation, converting the spliced image in the spatial domain into a spliced image in the frequency domain. The spliced image in the frequency domain is more sensitive to periodic structures. Adding the Fourier transform enables the convolution operation to explicitly obtain the global feature extraction ability, and the entire image inpainting model can obtain global features at an early stage, improving the inpainting ability for high-resolution large-hole images (such as PPT images). The inverse Fourier transform layer 1004 is used to perform an inverse Fourier transform on the image in the frequency domain to obtain the image in the spatial domain.
[0149] It should be noted that the above description of the image inpainting model structure is only an example. In other possible implementation manners, the image inpainting model may also have other structures, and the embodiments of the present application do not limit this.
[0150] 305. The server generates a target image based on the first processed image and the second processed image. The target image is the background image after removing the at least one first foreground region from the initial image.
[0151] In a possible implementation, the server fuses the first processed image and the second processed image based on the mask image to obtain the target image, where the mask image is generated based on the position of the second foreground region in the first processed image. The method for generating the mask image belongs to the same inventive concept as the content described above and will not be elaborated here.
[0152] For example, the server fuses the mask image with the second processed image to obtain a first reference image. The server fuses the inverse image of the mask image with the first processed image to obtain a second reference image. The server superimposes the first reference image and the second reference image to obtain the target image. The inverse image of the mask image is an image with pixel values opposite to those of the mask image. For example, if the mask image is a binary image, the pixel values of the pixel points in the mask image are 0 or 1. The pixel values of the pixel points in the inverse image of the mask image are opposite to the pixel values of the corresponding pixel points in the mask image. For example, if the pixel value of a certain pixel point in the mask image is 1, then in the inverse image, the pixel value of the pixel point at the same position is 0. In some embodiments, the process of the server generating the target image is represented by the following formula (1).
[0153] TI = mask * output + (1 - mask) * src (1)
[0154] Where TI is the target image, mask is the mask image, (1 - mask) is the inverse image of the mask image, output is the second processed image, that is, the image output by the image inpainting model, and src is the first processed image, also known as the policy erasure image. The policy erasure image is an image obtained by eliminating the first foreground region based on the target pixel value.
[0155] The following will combine Figure 11 to illustrate the above steps 301 - 305.
[0156] See Figure 11, the server identifies the initial image 1101 through an image recognition model, and obtains the position information 1102 of at least the first foreground area in the initial image 1101. Based on the position information 1102, the server processes the initial image 1101 in a policy erasure manner to obtain a first processed image 1103. Based on the position information 1104 of the second foreground area in the first processed image 1103, the server generates a mask image 1105. The server superimposes the mask image 1105 on the first processed image 1103 to obtain a third processed image 1106. The server inputs the third processed image 1106 and the mask image 1105 into an image restoration model 1107, and outputs a second processed image 1108 through the image restoration model 1107. Based on the mask image 1105, the server fuses the first processed image 1103 with the second processed image 1108 to obtain a target image 1109.
[0157] See Figure 12 , a comparison graph using the image generation method provided in the embodiments of the present application is provided. After processing the initial image 1201 using the technical solution provided in the embodiments of the present application, a target image 1202 is obtained. By Figure 12 It can be seen that after processing, the first foreground area 1203 in the initial image 1201 is eliminated.
[0158] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated here one by one.
[0159] Through the technical solution provided in the embodiments of the present application, when generating a target image, the initial image is identified to obtain a first foreground area to be processed, which is also the foreground area to be processed. Based on the reference background area adjacent to the first foreground area, pixel filling is performed on the first foreground area to obtain a first processed image for processing the target foreground area, and this processing method is also a policy-based processing method. The second foreground area in the first restored image is image-restored through an image restoration model to obtain a second processed image, and this processing method is also a model-based processing method. Finally, based on the first processed image and the second processed image, a target image that eliminates the first foreground area can be obtained. Combining the policy-based and model-based processing methods can improve the authenticity of the generated target image while ensuring the image generation efficiency.
[0160] To more clearly illustrate the technical solution provided in the embodiments of the present application, the training process of the image restoration model involved in the above steps 301-305 will be described below. See Figure 13 , in the following description process, one model iteration is taken as an example, and the method includes:
[0161] 1301. The server obtains a first sample image and a second sample image, where the second sample image is a sample image with a blank foreground area randomly generated in the first sample image.
[0162] In some embodiments, in addition to being able to randomly generate a blank foreground area in the first sample image to obtain the second sample image, the server can also perform data augmentation by randomly adding pictures, generating color blocks, etc. in the second sample image, so as to increase the number of sample images for training the image restoration model. When the server performs data augmentation by randomly adding pictures, generating color blocks, etc. in the second sample image, the server performs the same processing on the first sample image corresponding to the second sample image to ensure the consistency between the first sample image and the second sample image.
[0163] 1302. The server inputs the second sample image into the image restoration model, and the image restoration model repairs the blank foreground area in the second sample image and outputs a sample restored image.
[0164] Among them, the manner in which the server inputs the second sample image into the image restoration model and the image restoration model repairs the blank foreground area in the second sample image belongs to the same inventive concept as the manner in which the server repairs the second foreground area in the first processed image through the image restoration model in step 304 above. For the implementation process, refer to the relevant description in step 304 above and will not be elaborated here.
[0165] 1303. The server trains the image restoration model based on the difference information between the first sample image and the sample restored image.
[0166] In a possible implementation manner, the server inputs the first sample image and the sample restored image into a discriminant model, and the discriminant model outputs a authenticity parameter of the sample restored image based on the first difference information between the first sample image and the sample restored image. The authenticity parameter is used to represent the authenticity of the sample restored image. The server trains the image restoration model based on the authenticity parameter.
[0167] Among them, the authenticity of the sample restored image is used to represent the degree of difference between the sample restored image and the corresponding first sample image. The higher the authenticity of the sample restored image, the smaller the degree of difference between the sample restored image and the first sample image considered by the discriminant model; the lower the authenticity of the sample restored image, the greater the degree of difference between the sample restored image and the first sample image considered by the discriminant model. In some embodiments, the authenticity parameter is positively correlated with the authenticity of the sample restored image, that is, the larger the authenticity parameter, the higher the authenticity of the sample image; the smaller the authenticity parameter, the lower the authenticity of the sample image.
[0168] For example, the server inputs the first sample image and the sample restored image into a discriminant model, and the discriminant model extracts features from the first sample image and the sample restored image to obtain the feature map of the first sample image and the feature map of the sample restored image. Based on the first difference information between the feature map of the first sample image and the feature map of the sample restored image, the discriminant model outputs the authenticity parameter of the sample restored image. The server trains the image restoration model using the gradient descent method based on the authenticity parameter. In some embodiments, during the training process, the image restoration model and the discriminant model can form an "adversarial relationship", that is, the image restoration model and the discriminant model are trained simultaneously during training. The purpose of training the image restoration model is to maximize the authenticity parameter obtained through the discriminant model. The purpose of training the discriminant model is to distinguish as much as possible between the sample restored image output by the image restoration model and the first sample image. Through the "adversary" between the image restoration model and the discriminant model, the authenticity of the sample restored image output by the image restoration model is improved.
[0169] In a possible implementation, the server inputs the first sample image and the sample restored image into a feature extraction model, and the feature extraction model extracts features from the first sample image and the sample restored image and outputs the first sample image feature of the first sample image and the sample restored image feature of the sample restored image. The server trains the image restoration model based on the second difference information between the first sample image feature and the sample restored image feature.
[0170] Among them, the feature extraction model is a pre-trained feature extraction model, such as the neural network Resnet-101 (residual network 101) pre-trained on the large-scale open-source dataset Imagenet (image network).
[0171] It should be noted that the server can train the image restoration model using at least one of the above embodiments.
[0172] In some embodiments, the server trains the image restoration model based on the above first difference information and second difference information. For example, the server constructs a first loss function based on the first difference information, constructs a second loss function based on the second difference information, and uses the gradient descent method to train the image restoration model in combination with the first loss function and the second loss function. For example, the server trains the image restoration model through the following formula (2).
[0173] L = L adv + L pl + R (2)
[0174] Among them, L is the total loss function, Ladv is the first loss function, L pl is the second loss function, and R is a regularization term used to prevent overfitting. R is set by technicians according to actual situations, and this application embodiment does not make any limitations on it.
[0175] During the experiment, refer to Figure 14 , in the first column, the image 1401 is the first sample image input during training, in the second column, the image 1402 is the sample repaired image output by the image repair model, and in the third column, the image 1403 is the result after fusing the first sample image and the second sample image.
[0176] Figure 15 is a schematic structural diagram of an image generation device provided by an embodiment of this application. Refer to Figure 15 , the device includes: an image recognition module 1501, a pixel filling module 1502, an image repair module 1503, and an image generation module 1504.
[0177] The image recognition module 1501 is configured to recognize an initial image to obtain at least one first foreground area to be processed in the initial image.
[0178] The pixel filling module 1502 is configured to perform pixel filling on a target foreground area in the initial image based on at least one reference background area surrounding the first foreground area in the initial image to obtain a first processed image. The target foreground area is the first foreground area of the first type among the at least one first foreground area.
[0179] The image repair module 1503 is configured to perform image repair on a second foreground area in the first processed image based on the first processed image and an image repair model to obtain a second processed image. The second foreground area is the first foreground area among the at least one first foreground area except the target foreground area.
[0180] The image generation module 1504 is configured to generate a target image based on the first processed image and the second processed image. The target image is the background image after removing the at least one first foreground area from the initial image.
[0181] In a possible implementation manner, the pixel filling module 1502 is configured to determine the target foreground area from the at least one first foreground area based on the pixel values of the pixel points in the at least one reference background area. Fill the target foreground area in the initial image with a target pixel value to obtain a first processed image. The target pixel value is determined based on the pixel values of the pixel points in the reference background area corresponding to the target foreground area.
[0182] In a possible implementation, the pixel filling module 1502 is configured to perform a histogram statistics on the pixel values of the pixel points in the at least one reference background region to obtain the pixel histogram of the at least one reference background region. Based on the pixel histogram of the at least one reference background region, determine the types of the at least one first foreground region. Determine the first foreground region of the first type in the at least one first foreground region as the target foreground region.
[0183] In a possible implementation, for any one of the at least one reference background regions, the pixel filling module 1502 is configured to slide a sliding window on the pixel histogram of the reference background region. In response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram meeting the quantity condition, determine the type of the first foreground region surrounded by the reference background region as the first type. In response to the number of pixel points in the reference background region covered by the sliding window on the pixel histogram not meeting the quantity condition, determine the type of the first foreground region surrounded by the reference background region as the second type, where the second type is different from the first type.
[0184] In a possible implementation, the apparatus further includes:
[0185] A pixel value determination module, configured to divide the pixel points in the reference background region corresponding to the target foreground region into a plurality of pixel point sets, where the plurality of pixel point sets correspond to a plurality of pixel value intervals. Determine a target pixel point set from the plurality of pixel point sets, where the target pixel point set is the pixel point set with the largest number of pixel points among the plurality of pixel point sets. Determine the median of the pixel value interval corresponding to the target pixel point set as the target pixel value.
[0186] In a possible implementation, the image inpainting module 1503 is configured to generate a mask image based on the second foreground region, where the mask image is used to represent the position of the second foreground region in the first processed image. Superimpose the mask image and the first processed image to obtain a third processed image. Stitch the third processed image and the mask image to obtain a stitched image. Input the stitched image into the image inpainting model, and perform downsampling, feature extraction, and upsampling on the stitched image through the image inpainting model, and output the second processed image.
[0187] In a possible implementation, the image inpainting module 1503 is configured to perform Fourier convolution on the stitched image through the image inpainting model to obtain a downsampled image of the stitched image. Fourier convolution and residual connection are performed on the downsampled image through the image inpainting model to obtain image features of the downsampled image. Deconvolution is performed on the image features of the downsampled image through the image inpainting model to output the second processed image.
[0188] In a possible implementation, the image inpainting module 1503 is configured to process the stitched image through the image inpainting model by using a Fourier convolution unit to obtain a local feature map and a global feature map of the stitched image. The local feature map and the global feature map of the stitched image are stitched together to obtain a downsampled image of the stitched image.
[0189] In a possible implementation, the image generation module 1504 is configured to fuse the first processed image and the second processed image based on a mask image to obtain the target image, where the mask image is generated based on the position of the second foreground region in the first processed image.
[0190] In a possible implementation, the image recognition module 1501 is configured to input the initial image into an image recognition model, perform feature extraction on the initial image through the image recognition model to obtain a feature map of the initial image. The image recognition model is used to slide a sliding window on the feature map of the initial image to obtain a plurality of sub-feature maps of the initial image, and the plurality of sub-feature maps correspond to a plurality of regions of the initial image. The image recognition model is used to classify the plurality of regions based on the plurality of sub-feature maps and output at least one first foreground region in the plurality of regions, where the first foreground region is a foreground region in the plurality of regions.
[0191] In a possible implementation, the device further includes:
[0192] A model training module, configured to obtain a first sample image and a second sample image, where the second sample image is a sample image with a blank foreground region randomly generated in the first sample image. The second sample image is input into the image inpainting model, and the blank foreground region in the second sample image is repaired through the image inpainting model to output a sample repaired image. The image inpainting model is trained based on the difference information between the first sample image and the sample repaired image.
[0193] In a possible implementation, the model training module is configured to perform at least one of the following:
[0194] Input the first sample image and the sample restored image into a discriminative model. Based on the first difference information between the first sample image and the sample restored image, the discriminative model outputs a authenticity parameter of the sample restored image, and the authenticity parameter is used to represent the authenticity of the sample restored image. Based on the authenticity parameter, train the image restoration model.
[0195] Input the first sample image and the sample restored image into a feature extraction model. The feature extraction model extracts features from the first sample image and the sample restored image, and outputs a first sample image feature of the first sample image and a sample restored image feature of the sample restored image. Based on the second difference information between the first sample image feature and the sample restored image feature, train the image restoration model.
[0196] It should be noted that when the image generation device provided in the above embodiment generates an image, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the image generation device provided in the above embodiment and the image generation method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0197] Through the technical solution provided in the embodiment of the present application, when generating a target image, the initial image is recognized to obtain a first foreground area to be processed, which is also the foreground area to be processed. Pixel filling is performed on the first foreground area based on a reference background area adjacent to the first foreground area to obtain a first processed image of the target foreground area, and this processing method is also a strategy-based processing method. The second foreground area in the first restored image is image-restored through an image restoration model to obtain a second processed image, and this processing method is also a model-based processing method. Finally, based on the first processed image and the second processed image, a target image eliminating the first foreground area can be obtained. Combining the strategy plus model processing method can improve the authenticity of the generated target image while ensuring the image generation efficiency.
[0198] The embodiment of the present application provides a computer device for executing the above method. The computer device can be implemented as a terminal or a server. First, the structure of the terminal will be introduced below:
[0199] Figure 16 It is a schematic structural diagram of a terminal provided in the embodiment of the present application. Generally, the terminal 1600 includes one or more processors 1601 and one or more memories 1602.
[0200] The processor 1601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1601 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1601 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0201] The memory 1602 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1602 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1602 is used to store at least one computer program, and the at least one computer program is used to be executed by the processor 1601 to implement the image generation method provided in the method embodiments of this application.
[0202] In some embodiments, the terminal 1600 may further optionally include: a peripheral device interface 1603 and at least one peripheral device. The processor 1601, the memory 1602, and the peripheral device interface 1603 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1603 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607, and a power supply 1608.
[0203] The peripheral device interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1601 and the memory 1602. In some embodiments, the processor 1601, the memory 1602, and the peripheral device interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1601, the memory 1602, and the peripheral device interface 1603 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0204] The radio frequency circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1604 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1604 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on.
[0205] The display screen 1605 is used to display a UI (User Interface). The UI can include graphics, text, icons, videos, and any combination thereof. When the display screen 1605 is a touch display screen, the display screen 1605 also has the ability to collect touch signals on or above the surface of the display screen 1605. The touch signal can be input to the processor 1601 as a control signal for processing. At this time, the display screen 1605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard.
[0206] The camera assembly 1606 is used to collect images or videos. Optionally, the camera assembly 1606 includes a front camera and a rear camera. Generally, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal.
[0207] The audio circuit 1607 can include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 1601 for processing, or input them to the radio frequency circuit 1604 to achieve voice communication.
[0208] The power supply 1608 is used to supply power to each component in the terminal 1600. The power supply 1608 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0209] In some embodiments, the terminal 1600 further includes one or more sensors 1609. The one or more sensors 1609 include, but are not limited to: an acceleration sensor 1610, a gyroscope sensor 1611, a pressure sensor 1612, an optical sensor 1613, and a proximity sensor 1614.
[0210] The acceleration sensor 1610 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1600.
[0211] The gyroscope sensor 1611 can detect the body direction and rotation angle of the terminal 1600. The gyroscope sensor 1611 can cooperate with the acceleration sensor 1610 to collect the 3D actions of the user on the terminal 1600.
[0212] The pressure sensor 1612 can be disposed on the side frame of the terminal 1600 and / or the lower layer of the display screen 1605. When the pressure sensor 1612 is disposed on the side frame of the terminal 1600, it can detect the holding signal of the user on the terminal 1600, and the processor 1601 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1612. When the pressure sensor 1612 is disposed on the lower layer of the display screen 1605, the processor 1601 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1605.
[0213] The optical sensor 1613 is used to collect the ambient light intensity. In one embodiment, the processor 1601 can control the display brightness of the display screen 1605 according to the ambient light intensity collected by the optical sensor 1613.
[0214] The proximity sensor 1614 is used to collect the distance between the user and the front of the terminal 1600.
[0215] Those skilled in the art can understand that Figure 16 the structure shown in does not constitute a limitation on the terminal 1600, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.
[0216] The above computer device can also be implemented as a server. The structure of the server will be introduced below:
[0217] Figure 17It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1700 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1701 and one or more memories 1702. Among them, at least one computer program is stored in the one or more memories 1702, and the at least one computer program is loaded and executed by the one or more processors 1701 to implement the methods provided by the above-mentioned various method embodiments. Of course, the server 1700 may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server 1700 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0218] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a computer program. The above computer program can be executed by a processor to complete the image generation method in the above embodiment. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0219] In an exemplary embodiment, a computer program product or a computer program is also provided. The computer program product or the computer program includes program code, and the program code is stored in a computer-readable storage medium. The processor of the computer device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the above image generation method.
[0220] In some embodiments, the computer program involved in the embodiments of the present application may be deployed to be executed on one computer device, or on multiple computer devices located at one location. Or, it may be executed on multiple computer devices distributed at multiple locations and interconnected by a communication network. The multiple computer devices distributed at multiple locations and interconnected by a communication network may form a blockchain system.
[0221] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a magnetic disk, or an optical disc, etc.
[0222] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. An image generation method, characterized in that, The method includes: Identifying an initial image to obtain at least one first foreground region to be processed in the initial image; Based on at least one reference background region surrounding the first foreground region in the initial image, performing pixel filling on a target foreground region in the initial image to obtain a first processed image, where the target foreground region is a first foreground region of a first type among the at least one first foreground region; Generating a mask image based on a second foreground region in the first processed image, where the second foreground region is a first foreground region other than the target foreground region among the at least one first foreground region, and the mask image is used to represent the position of the second foreground region in the first processed image; Overlaying the mask image and the first processed image to obtain a third processed image; Stitching the third processed image and the mask image to obtain a stitched image; Inputting the stitched image into an image inpainting model, and performing downsampling, feature extraction, and upsampling on the stitched image through the image inpainting model to output a second processed image; Generating a target image based on the first processed image and the second processed image, where the target image is a background image after removing the at least one first foreground region from the initial image.
2. The method according to claim 1, wherein The performing pixel filling on a target foreground region in the initial image based on at least one reference background region surrounding the first foreground region in the initial image to obtain a first processed image includes: Determining the target foreground region from the at least one first foreground region based on the pixel values of pixel points in the at least one reference background region; Filling the target foreground region with a target pixel value in the initial image to obtain a first processed image, where the target pixel value is determined based on the pixel values of pixel points in the reference background region corresponding to the target foreground region.
3. The method according to claim 2, wherein The determining the target foreground region from the at least one first foreground region based on the pixel values of pixel points in the at least one reference background region includes: Performing histogram statistics on the pixel values of pixel points in the at least one reference background region to obtain a pixel histogram of the at least one reference background region; Determining the types of the at least one first foreground region based on the pixel histogram of the at least one reference background region; Determining the first foreground region of the first type among the at least one first foreground region as the target foreground region.
4. The method according to claim 3, characterized in that, The determining the types of the at least one first foreground region based on the pixel histogram of the at least one reference background region includes: For any reference background region among the at least one reference background region, sliding a sliding window on the pixel histogram of the reference background region; In response to the number of pixel points in the reference background region covered by the sliding window at any position on the pixel histogram meeting a quantity condition, determining the type of the first foreground region surrounded by the reference background region as the first type; In response to the number of pixel points in the reference background area covered by the sliding window on the pixel histogram not meeting the quantity condition, determine the type of the first foreground area surrounded by the reference background area as a second type, where the second type is different from the first type.
5. The method according to claim 2, wherein Before filling the target foreground area with a target pixel value in the initial image to obtain a first processed image, the method further includes: Dividing the pixel points in the reference background area corresponding to the target foreground area into a plurality of pixel point sets, where the plurality of pixel point sets correspond to a plurality of pixel value intervals; Determining a target pixel point set from the plurality of pixel point sets, where the target pixel point set is the pixel point set with the largest number of pixel points among the plurality of pixel point sets; Determining the median of the pixel value interval corresponding to the target pixel point set as the target pixel value.
6. The method according to claim 1, wherein The outputting of a second processed image by performing downsampling, feature extraction, and upsampling on the stitched image through the image inpainting model includes: Performing Fourier convolution on the stitched image through the image inpainting model to obtain a downsampled image of the stitched image; Performing Fourier convolution and residual connection on the downsampled image through the image inpainting model to obtain the image features of the downsampled image; Performing deconvolution on the image features of the downsampled image through the image inpainting model to output the second processed image.
7. The method according to claim 6, characterized in that The performing of Fourier convolution on the stitched image through the image inpainting model to obtain a downsampled image of the stitched image includes: Processing the stitched image through the image inpainting model using a Fourier convolution unit to obtain a local feature map and a global feature map of the stitched image; stitching the local feature map and the global feature map of the stitched image to obtain a downsampled image of the stitched image.
8. The method according to claim 1, wherein The generating of a target image based on the first processed image and the second processed image includes: Fusing the first processed image and the second processed image based on the mask image to obtain the target image.
9. The method according to claim 1, wherein The identifying of at least one first foreground area to be processed in the initial image includes: Inputting the initial image into an image recognition model, and performing feature extraction on the initial image through the image recognition model to obtain a feature map of the initial image; Sliding a sliding window on the feature map of the initial image through the image recognition model to obtain a plurality of sub-feature maps of the initial image, where the plurality of sub-feature maps correspond to a plurality of areas of the initial image; Classifying the plurality of areas based on the plurality of sub-feature maps through the image recognition model, and outputting the at least one first foreground area among the plurality of areas, where the first foreground area is a foreground area among the plurality of areas.
10. The method according to claim 1, wherein The method further includes: Obtaining a first sample image and a second sample image, where the second sample image is a sample image with a blank foreground area randomly generated in the first sample image; Input the second sample image into the image inpainting model, and repair the blank foreground area in the second sample image through the image inpainting model to output a sample repaired image; Train the image inpainting model based on the difference information between the first sample image and the sample repaired image.
11. The method according to claim 10, characterized in that, The training of the image inpainting model based on the difference information between the first sample image and the sample repaired image includes at least one of the following: Input the first sample image and the sample repaired image into a discriminant model, and output the authenticity parameter of the sample repaired image through the discriminant model based on the first difference information between the first sample image and the sample repaired image, where the authenticity parameter is used to represent the authenticity of the sample repaired image; train the image inpainting model based on the authenticity parameter; Input the first sample image and the sample repaired image into a feature extraction model, extract features from the first sample image and the sample repaired image through the feature extraction model, and output the first sample image feature of the first sample image and the sample repaired image feature of the sample repaired image; train the image inpainting model based on the second difference information between the first sample image feature and the sample repaired image feature.
12. An image generation device, characterized in that, The device includes: An image recognition module, configured to recognize an initial image to obtain at least one first foreground area to be processed in the initial image; A pixel filling module, configured to perform pixel filling on a target foreground area in the initial image based on at least one reference background area surrounding the first foreground area in the initial image to obtain a first processed image, where the target foreground area is the first foreground area of the first type among the at least one first foreground area; An image inpainting module, configured to generate a mask image based on a second foreground area in the first processed image, where the second foreground area is the first foreground area other than the target foreground area among the at least one first foreground area, and the mask image is used to represent the position of the second foreground area in the first processed image; superimpose the mask image and the first processed image to obtain a third processed image; splice the third processed image and the mask image to obtain a spliced image; input the spliced image into an image inpainting model, and perform downsampling, feature extraction, and upsampling on the spliced image through the image inpainting model to output a second processed image; An image generation module, configured to generate a target image based on the first processed image and the second processed image, where the target image is the background image after removing the at least one first foreground area from the initial image.
13. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, and at least one computer program is stored in the one or more memories, and the computer program is loaded and executed by the one or more processors to implement the image generation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the image generation method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the image generation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Image processing apparatus and image processing method
CN101902549A
Cavity filling method based on image sequence
CN104065946A