An image generation method, apparatus, electronic device, and storage medium

By performing fusion analysis on reference and synthesized images, the fusion region and intensity information are determined. Image fusion is then performed using an image generation model, which solves the data bias problem in image synthesis and generates a more realistic target fused image.

CN115131202BActive Publication Date: 2025-10-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210579016.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-10-31
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

The existing image synthesis methods cause data bias, resulting in data deviations when images are applied.

Method used

By acquiring the reference image and the synthesized image to be fused, fusion analysis is performed to determine the fusion region information and fusion intensity information, and image fusion is performed based on this information. A trained image generation model is used to mitigate data bias.

Benefits of technology

It mitigates image synthesis data bias caused by specific methods or features, generating more realistic target fusion images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131202B_ABST
    Figure CN115131202B_ABST
Patent Text Reader

Abstract

This invention discloses an image generation method, apparatus, electronic device, and storage medium, comprising: acquiring a reference image and a composite image to be fused, wherein the reference image includes a reference object and the composite image includes a composite object; performing fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image; and fusing the reference image and the composite image based on the fusion region information and fusion intensity information to obtain a target fused image. This invention can determine the fusion region information and fusion intensity information corresponding to different fusion methods, so that images can be fused according to different fusion strategies, mitigating data bias caused by fusion using specific methods or features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to an image generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of internet technology, the application of images in various fields is increasing. However, obtaining original images is extremely difficult. For example, images downloaded from the internet may vary in quality, are costly to acquire, and pose copyright risks. Therefore, existing technologies often use image synthesis to create a large number of images.

[0003] However, the general methods of image synthesis follow specific methods or features, which can lead to data bias due to the synthesis of specific features, resulting in data deviation when the image is applied. Summary of the Invention

[0004] To address the problems of existing technologies, embodiments of the present invention provide an image generation method, apparatus, electronic device, and storage medium. The technical solution is as follows:

[0005] On the one hand, an image generation method is provided, the method including:

[0006] Obtain the reference image and the composite image to be fused; wherein the reference image includes a reference object; and the composite image includes a composite object.

[0007] A fusion analysis is performed on the reference object in the reference image and the composite object in the composite image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the composite image.

[0008] The reference image and the synthetic image are fused based on the fusion region information and fusion intensity information to obtain the target fused image.

[0009] On the other hand, an image generation apparatus is provided, the apparatus comprising:

[0010] The image acquisition module is used to acquire the reference image and the composite image to be fused; wherein, the reference image includes a reference object; and the composite image includes a composite object.

[0011] The fusion information determination module is used to perform fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the composite image.

[0012] The synthesis module is used to fuse the reference image and the synthesized image based on the fusion region information and fusion intensity information to obtain the target fused image.

[0013] On the other hand, an electronic device is provided, including a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the image generation method described above.

[0014] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction or at least one program is stored therein, the at least one instruction or the at least one program being loaded and executed by a processor to implement the image generation method as described above.

[0015] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image generation method described above.

[0016] This invention provides embodiments of the invention that acquire a reference image and a composite image to be fused. The reference image includes a reference object, and the composite image includes a composite object. Fusion analysis is performed on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image. Based on the fusion region information and fusion intensity information, the reference image and the composite image are fused to obtain a target fused image. This embodiment of the invention can determine the fusion region information and fusion intensity information corresponding to different fusion methods, enabling images to be fused according to different fusion strategies and mitigating data bias caused by fusion using specific methods or features. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0019] Figure 2 This is a flowchart illustrating an image generation method provided in an embodiment of the present invention;

[0020] Figure 3 This is a schematic flowchart of an image acquisition method provided in an embodiment of the present invention;

[0021] Figure 4This is a schematic flowchart of a method for acquiring a synthesized image provided in an embodiment of the present invention;

[0022] Figure 5 This is a schematic diagram of the structure of a trained image generation model provided in an embodiment of the present invention;

[0023] Figure 6 This is a schematic diagram of an adversarial network structure provided in an embodiment of the present invention;

[0024] Figure 7 This is a schematic diagram of an adversarial network training process provided by an embodiment of the present invention;

[0025] Figure 8 This is a structural block diagram of an image generation device provided in an embodiment of the present invention;

[0026] Figure 9 This is a hardware structure block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0029] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0030] The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0031] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to cloud computing business models. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to backend systems for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing.

[0032] Cloud storage is a new concept that extends and develops from cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to aggregate a large number of various types of storage devices (also called storage nodes) in a network through application software or application interfaces to work together and provide data storage and business access functions. Currently, the storage method of a storage system is as follows: Logical volumes are created. When creating a logical volume, physical storage space is allocated to each logical volume. This physical storage space may consist of the disks of one or several storage devices. Clients store data on a logical volume, which means storing the data on the file system. The file system divides the data into many parts, each part being an object. Each object contains not only data but also additional information such as a data identifier (ID, IDentity). The file system writes each object to the physical storage space of the logical volume and records the storage location information of each object. Therefore, when a client requests access to data, the file system can allow the client to access the data based on the storage location information of each object. The process by which a storage system allocates physical storage space to a logical volume is as follows: the physical storage space is pre-divided into stripes according to the capacity estimate of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of independent redundant disk arrays (R_ID, Redundancy of Independent Disk). A logical volume can be understood as a stripe, thus allocating physical storage space to the logical volume.

[0033] A database (D_t_b_se) can be simply viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, with minimal redundancy, and independent of application programs.

[0034] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile devices; or according to the query language used, such as SQL (Structured Query Language) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages.

[0035] Please see Figure 1 The diagram shown is an implementation environment provided by an embodiment of the present invention. The implementation environment may include a client 110, a server 120, and a database 130.

[0036] The client 110 and the server 120, as well as the server 120 and the database 130, can communicate via a network.

[0037] The client 110 includes, but is not limited to, mobile phones, computers, smart voice interaction devices, smart home appliances, in-vehicle clients, and aircraft. The client 110 runs an application with human-computer interaction capabilities, which can launch virtual item distribution activities for different business scenarios, such as flash sales, lotteries, and reward-for-complete-task activities.

[0038] Server 120 can provide background services for the application in client 110. Specifically, the server can obtain a reference image and a composite image to be fused. The reference image includes a reference object, and the composite image includes a composite object. The server performs fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image. Based on the fusion region information and fusion intensity information, the server fuses the reference image and the composite image to obtain the target fused image. The composite image can be obtained from database 130.

[0039] Database 130 may include in-memory databases and relational databases. It should be noted that the servers, databases, nodes, etc. in the embodiments of the present invention can be independent physical servers, or server clusters or distributed systems composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0040] In one exemplary implementation, the client 110, server 120, and database 130 can all be node devices in a blockchain system, capable of sharing acquired and generated information with other node devices in the blockchain system, thus achieving information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which consists of multiple blocks, and adjacent blocks are related, ensuring that any data tampering in any block can be detected by the next block, thereby preventing data tampering in the blockchain and guaranteeing the security and reliability of the data in the blockchain.

[0041] Please see Figure 2 The diagram shown is a flowchart of an image generation method provided by an embodiment of the present invention. This method can be applied to... Figure 1 The implementation environment shown can be the subject executing this method. Figure 1 The server that generates the execution image. It should be noted that this specification provides the operational steps of the method as described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many, and does not represent the only execution order. In actual system or product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown... Figure 2 As shown, the method may include:

[0042] S201, Obtain the reference image and the composite image to be fused; wherein, the reference image includes a reference object; and the composite image includes a composite object.

[0043] In this embodiment, the server can obtain a reference image and a synthesized image to be fused. In an optional embodiment, where the image generation model in the following embodiments is a face image generation model, the reference image and the synthesized image can be face images. Where the image generation model in the following embodiments is an environment image generation model, the reference image and the synthesized image can be environment images. For ease of explanation, this embodiment will be described using a face generation model.

[0044] In this embodiment, the server can generate a fused image from a real face image and a fake face image, so that the image generation model that generates the fused image has better generalization ability and alleviates data bias. Based on this, the reference image in step S201 can be a real face image, and the synthesized image can be a fake face image. The reference object included in the reference image can be a real human face, and the synthesized object included in the synthesized image can be a synthesized fake human face.

[0045] In this embodiment, a real face and a fake face can include an outer face region and an inner face region. Optionally, the inner face region can include the area within the face outline, and may include various parts such as the mouth, nose, eyes, and eyebrows. The outer face region includes the area outside the face outline, and may include areas such as hair, ears, and neck.

[0046] In this embodiment of the application, in addition to including real faces and fake faces respectively, real face images and fake face images may also include the environmental part next to the face.

[0047] Please see Figure 3 The diagram shown is a flowchart illustrating an image acquisition method provided in an embodiment of the present invention. Figure 3 As shown, it includes:

[0048] S301, Obtain the reference image.

[0049] Continuing with the explanation using a reference image as the real face image, in one alternative embodiment, the server can acquire the reference image in various ways.

[0050] In some possible embodiments, the server can receive a reference image sent by the client. Optionally, the reference image may be a face captured by the client using a camera. Optionally, after capturing a human image, the client can crop the human image to obtain the reference image, i.e., a real face image including the face.

[0051] In some other possible embodiments, the server may receive reference images sent by a real face database.

[0052] S302, Determine the feature information of the reference object in the reference image.

[0053] In this embodiment of the application, in order to better integrate the reference object in the reference image with the synthetic object in the synthetic image, the server can determine the feature information of the reference object in the reference image, that is, the server can determine the feature information of the real face in the real face image.

[0054] Optionally, the feature information may include pose feature information. In one optional embodiment, the server can perform pose analysis on a reference object in a reference image to obtain pitch, yaw, and roll information corresponding to the reference image. Subsequently, the server can determine the pose feature information of the reference object based on the pitch, yaw, and roll information. When the reference image is a real face image, the pose feature information is used to characterize the pitch angle, yaw angle, and roll angle of the real face in the real face image, wherein the pitch information includes the pitch angle, the yaw information includes the yaw angle, and the roll information includes the roll angle.

[0055] Optionally, the feature information may include attribute feature information. Specifically, the server can perform attribute analysis on the reference object in the reference image to obtain the attribute feature information of the reference object. The attribute feature information is used to characterize the gender or the feature information of various parts of the facial features (including double eyelids, single eyelids, high nose bridge, etc.) of the real face in the real face image.

[0056] Optionally, the feature information may include pose feature information and attribute feature information. Specific implementation methods for obtaining pose feature information and attribute feature information can be found in the embodiments described above, and will not be repeated here.

[0057] S303, determine the feature information of the object in each candidate image in the candidate image library.

[0058] Optionally, if the feature information of the reference object in the reference image is attitude feature information, then the feature information of the object in each candidate image is also attitude feature information. In an optional embodiment, the server can perform attitude analysis on the object in each candidate image in the candidate image library to obtain the pitch, yaw, and roll information of the object in each candidate image. Subsequently, the attitude feature information of the object in each candidate image can be determined based on the pitch, yaw, and roll information of the object in each candidate image.

[0059] Optionally, if the feature information of the reference object in the reference image is attribute feature information, then the feature information of the object in each candidate image is also attribute feature information. In an optional embodiment, the server can perform attribute analysis on the object in each candidate image in the candidate image library to obtain the attribute feature information of the reference object.

[0060] Optionally, if the reference-specific feature information in the reference image consists of pose feature information and attribute feature information, then the feature information of the object in each candidate image also consists of pose feature information and attribute feature information. Specific implementation methods for obtaining pose feature information and attribute feature information can be found in the above embodiments, and will not be repeated here.

[0061] In this way, the server can determine the feature information of the object in each candidate image in the candidate image library based on the content included in the feature information.

[0062] S304, determine the feature distance value corresponding to each candidate image based on the feature information of the object in each candidate image and the feature information of the reference object.

[0063] Optionally, if the feature information is pose feature information, the server can determine the feature distance value between the object in each candidate image and the reference object by subtracting the pose feature information of the object in each candidate image from the pose feature information of the reference object.

[0064] Optionally, if the feature information is attribute feature information, the server can determine the feature distance value between the object in each candidate image and the reference object by subtracting the attribute feature information of the object in each candidate image from the attribute feature information of the reference object.

[0065] Optionally, if the feature information consists of pose feature information and attribute feature information, the server can determine a first feature distance value between the object in each candidate image and the reference object by subtracting the pose feature information of the object in each candidate image from the pose feature information of the reference object. The server can also determine a second feature distance value between the object in each candidate image and the reference object by subtracting the attribute feature information of the object in each candidate image from the attribute feature information of the reference object. The sum of the first and second feature distance values ​​is then used as the feature distance value between the object in each candidate image and the reference object. Alternatively, the average of the first and second feature distance values ​​can be used as the feature distance value between the object in each candidate image and the reference object.

[0066] S305, determine the synthesized image from the candidate image library based on the feature distance value corresponding to each candidate image.

[0067] In an alternative embodiment, the server may use the candidate image with the smallest feature distance value from the candidate image library as the synthesized image.

[0068] In another optional embodiment, the server can sort all candidate images in the candidate image library from smallest to largest by feature distance values, determine the candidate images corresponding to the top N feature distance values ​​as the composite image set, and determine any one image from the composite image set as the composite image. Optionally, N can be preset.

[0069] In an alternative embodiment, in order to better blend the reference object in the reference image with the composite object in the composite image, the server can adjust the hue of the reference object and the composite object to a uniform hue.

[0070] Please see Figure 4 The diagram shown is a flowchart illustrating a method for acquiring a synthesized image according to an embodiment of the present invention. Figure 4 As shown, it includes:

[0071] S401, candidate images are determined from the candidate image library based on the feature distance value corresponding to each candidate image.

[0072] In one alternative embodiment, the server may select the candidate image from the candidate image library with the smallest feature distance value as the candidate image.

[0073] In another optional embodiment, the server can sort the feature distance values ​​of all candidate images in the candidate image library from smallest to largest, determine the candidate images corresponding to the top N feature distance values ​​as the candidate image set, and determine any one image from the synthesized image set as the candidate image. Optionally, N can be preset.

[0074] S402, Determine the color feature information of the reference object in the reference image.

[0075] In one implementation of determining the color feature information of a reference object in a reference image, the server can determine the color feature information of the reference object, that is, determine the color feature information of a real human face.

[0076] Optionally, the server can recognize a real face, distinguishing between the inner and outer face regions. If the outer face region includes different areas, such as the hair region and the neck region, the server can also distinguish these different regions within the outer face region and calculate the color feature information for each region (including the outer face region and the different regions within it). The color feature information of each region can represent the color feature information of a reference object.

[0077] Optionally, the server can recognize a real human face, identify the inner face region, and then calculate the color feature information of the inner face region. The color feature information of the inner face region can represent the color feature information of a reference object.

[0078] In an embodiment for calculating color feature information, taking the calculation of color feature information of the inner face region as an example, the server can determine the color histogram of each pixel in the inner face region, determine the data corresponding to the color histogram of each pixel, determine the mean value based on the data corresponding to the color histogram of each pixel, and use the mean value as the color feature information of the inner face region.

[0079] S403, adjust the color feature information of the object in the candidate image based on the color feature information of the reference object to obtain a composite image, wherein the matching value between the color feature information of the adjusted object and the color feature information of the reference object in the composite image is greater than or equal to a preset threshold.

[0080] In this embodiment, where the color feature information of the inner face region can represent the color feature information of the reference object, the server can identify the inner face region from the object in the candidate image and determine the color histogram of each pixel in the inner face region of the object in the candidate image. Subsequently, the server can determine the mean value based on the data corresponding to the color histogram of each pixel and use the mean value as the color feature information of the inner face region of the object in the selected image.

[0081] Next, the server can adjust the color feature information of the objects in the candidate image based on the color feature information of the reference object to obtain the synthesized image.

[0082] In the synthesized image, a matching value between the color feature information of the adjusted object and the color feature information of the reference object that is greater than or equal to a preset threshold means that the color feature information of the adjusted object and the color feature information of the reference object are the same or similar.

[0083] S203, perform fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the composite image.

[0084] In one optional embodiment, the server can determine a first fusion strategy corresponding to the reference image and a second fusion strategy corresponding to the synthesized image. The first fusion strategy includes first fusion region information and first fusion intensity range information of the reference image, and the second fusion strategy includes second fusion region information and second fusion intensity range information of the synthesized image. The server can determine the fusion region information corresponding to the reference image and the synthesized image based on the first and second fusion region information, and determine the fusion intensity information corresponding to the reference image and the synthesized image based on the first and second fusion intensity range information.

[0085] In another alternative embodiment, the server can perform fusion analysis on objects in the reference image and the synthesized image based on the fusion method generator in the image generation model to obtain fusion region information and fusion intensity information corresponding to the reference image and the synthesized image.

[0086] Figure 5 This is a schematic diagram illustrating the structure of a trained image generation model according to an exemplary embodiment, such as... Figure 5 As shown, the image generation model includes a fusion generator and an image synthesizer.

[0087] In this embodiment, the server can input the reference image and the synthesized image into the fusion mode generator in the trained image generation model. That is, based on the fusion mode generator in the image generation model, the server performs fusion analysis on the objects in the reference image and the synthesized image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the synthesized image.

[0088] In this embodiment, the fusion intensity information can characterize the proportion of fake faces in the synthesized object during the fusion process. Optionally, the fusion intensity information can characterize the proportion of real faces in the reference object during the fusion process. The following will use the example of fusion intensity information referring to the proportion of fake faces in the synthesized object during the fusion process to illustrate this.

[0089] In this embodiment, the fusion region information refers to which region is merged during the fusion process between the fake face and the real face. Optionally, it can be the inner face region, the full face region, or the outer face region.

[0090] S205, the reference image and the synthetic image are fused based on the fusion region information and fusion intensity information to obtain the target fused image.

[0091] like Figure 5 As shown, the server can fuse the reference image and the synthesized image based on the fusion region information and fusion intensity information to obtain the target fused image.

[0092] In this embodiment, the server can directly determine whether to fuse the inner face region, the full face region, or the outer face region based on the fusion region information. Then, based on the fusion intensity information, it can determine the proportion of the fake face in the synthesized object during the fusion process. Based on, for example, the inner face region and the proportion occupied, the server can fuse the reference image and the synthesized image to obtain the target fused image.

[0093] In this embodiment, the server can input the reference image, the synthesized image, the fusion intensity information, and the fusion region information into the image synthesizer in the trained image generation model to obtain the target fused image.

[0094] In one optional embodiment, when fusing region information to represent the internal region of an object, the server can obtain a reference internal object from a reference object in a reference image and a synthesized internal object from a synthesized object in a synthesized image, based on the image synthesizer in the image generation model. Specifically, when fusing region information to represent the internal region of an object, the server can obtain the inner face region from a real face in a real face image and the inner face region from a fake face in a fake face image.

[0095] The server merges the reference internal object and the composite internal object based on the fusion strength information to obtain the merged internal object, i.e., the merged internal face region.

[0096] In this embodiment of the application, the server can determine the external object to be fused, and determine the target fused image containing the fused object based on the internal object to be fused and the external object to be fused.

[0097] In one optional embodiment, the server can determine that the external region to be fused is the outer face region of a real human face, and stitch the outer face region of the real human face and the inner face region to obtain a fused face. The fused face is then stitched together with the environmental region in the real face image or the fake face image to obtain a target fused image containing the fused face.

[0098] In another alternative embodiment, the server can determine that the external region to be fused is the outer face region of a fake face, and stitch the outer face region of the fake face and the inner face region to obtain a fused face. The fused face is then stitched together with the environmental region in either the fake face image or the real face image to obtain a target fused image containing the fused face.

[0099] In the embodiments of this application, when the fused region information is the internal region of the object, there can be multiple fusion methods in the actual fusion method, such as the alpha fusion method and the Poisson fusion method.

[0100] Optionally, if the fusion region information is the internal region of the object corresponding to the alpha fusion method, the server can obtain the synthesis formula (1) of the fusion internal region corresponding to the alpha method:

[0101] I out1 =αI f *M+(1-α)I r *M……Formula (1)

[0102] Among them, I f It's a fake face image, I r This is a real face image, α is the fusion intensity information, M is the mask image of the inner face region (1 for inner face region, 0 for non-inner face region), and I... out1 It is a fusion of the inner face area.

[0103] Subsequently, the server can stitch the inner face region to be fused with the outer object to be fused to obtain the fused face. The fused face is then stitched with the environmental region from either the real face image or the fake face image to obtain the target fused image containing the fused face.

[0104] Optionally, if the fusion region information is the internal region of the object corresponding to the Poisson blending method, the server can use the Poisson blending method to blend the reference internal object and the composite internal object to obtain the fused internal object. When splicing the fused internal object and the external object to be blended, the server can use the Poisson blending method to process the splicing edges, making the transition of the splicing edges more natural and smooth.

[0105] In another optional embodiment, when the fusion region information represents the object region, i.e., when the fusion region information is the full face region of the object, the server obtains a reference object (a real face) from the reference image and a synthetic object (a fake face) from the synthetic image. The reference object and the synthetic object are then fused according to the fusion intensity information to obtain a target fused image containing the fused objects.

[0106] Specifically, the server merges real and fake faces based on fusion strength information to obtain a merged face. Then, the server can stitch the merged face with environmental regions in the real or fake face images to obtain a target fused image containing the merged face.

[0107] In the embodiments of this application, when representing the object region with fused region information, the actual fusion method can correspond to a linear mixing method.

[0108] Optionally, if the fusion region information is the object region corresponding to the linear blending method, the server can obtain the synthesis formula (2) of the fused face corresponding to the linear blending method:

[0109] I out2 =αIf +(1-α)I r .........Formula (2)

[0110] Among them, I f It's a fake face image, I r It is a real face image, α is the fusion intensity information, and I out2 It's a face fusion.

[0111] In one alternative embodiment, when the fusion region information represents the reference image, i.e., the real face image, regardless of the amount of fusion intensity information, the target fused image synthesized by the image synthesizer is the reference image, i.e., the real face image.

[0112] In one alternative embodiment, when the fusion region information represents the synthesized image, i.e., the fake face image, regardless of the amount of fusion intensity information, the target fused image synthesized by the image synthesizer is the synthesized image, i.e., the fake face image.

[0113] In this embodiment of the application, the fusion strength information can be represented by a value between 0 and 1, such as 0, 0.1, 0.2, 0.3...0.9, 1.

[0114] In this way, the server can use the fusion method generator in the trained image generation model to determine the fusion region information and fusion intensity information corresponding to different fusion methods, so that the image synthesizer can fuse images according to different fusion strategies and mitigate the data bias caused by fusion of specific methods or specific features.

[0115] This application also includes a process for training an image generation model. In an optional embodiment, in order to make the target fused image generated by the image generation model more realistic, and in order to train a module that detects the target fused image, the image generation model can be trained in an adversarial network containing a discriminator.

[0116] Figure 6 This is a schematic diagram illustrating the structure of an adversarial network according to an exemplary embodiment, such as... Figure 6 As shown, it includes an image generation model and a discriminator, wherein the image generation model includes a fusion generator and an image synthesizer.

[0117] Figure 7 This is a schematic diagram illustrating a process for training an adversarial network according to an exemplary embodiment, such as... Figure 7 As shown, it includes:

[0118] S701, Obtain the training reference image and the training synthesized image to be fused; wherein, the training reference image includes the training reference object; the training synthesized image includes the training synthesized object.

[0119] In this embodiment, the server can obtain the training reference image and the training synthesized image to be fused. In an optional embodiment, when the image generation model is a face image generation model, the training reference image and the training synthesized image can be face images. When the image generation model is an environment image generation model, the training reference image and the training synthesized image can be environment images.

[0120] In this embodiment, the training reference image can be a real face image, and the training synthesized image can be a fake face image. The reference object in the reference image can be a real human face, and the synthesized object in the synthesized image can be a synthesized fake human face.

[0121] For the process of the body, please refer to the corresponding embodiment and extension of step S201, which will not be repeated here.

[0122] S702, input the training reference image and the training synthesized image into the fusion mode generator in the untrained image generation model to obtain the training fusion region information and the training fusion intensity information.

[0123] In this embodiment, the server can input the training reference image and the training synthesized image into the fusion method generator in the untrained image generation model to obtain training fusion region information and training fusion intensity information. For the specific process, please refer to the embodiments and extensions corresponding to step S203, which will not be repeated here.

[0124] S703 inputs the training reference image, the training synthesized image, the training fusion region information, and the training fusion intensity information into the image synthesizer in the untrained image generation model to obtain the training fused image.

[0125] In this embodiment, the server can input the training reference image, the training synthesized image, the training fusion region information, and the training fusion intensity information into the image synthesizer in the untrained image generation model to obtain the training fused image. For the specific process, please refer to the embodiment and extension corresponding to step S205, which will not be repeated here.

[0126] S704: Train the untrained image generation model and discriminator based on the training fused image to obtain the trained image generation model and discriminator.

[0127] In one optional embodiment, the server can input the training fused image into the discriminator to obtain discrimination information, determine a first target loss based on the discrimination information and the first annotation information of the training fused image corresponding to the discriminator, determine a second target loss based on the discrimination information and the second annotation information of the training fused image corresponding to the image generation model, train an untrained image generation model and discriminator based on the constraint loss function, the first target loss and the second target loss, and obtain a trained image generation model and discriminator when the iteration termination condition is met.

[0128] Optionally, the discriminator typically uses 0-1 to represent the probability of an image being real or fake. In practice, since the training fused image is generated by an image generation model and is not a real face image, and the image generation model is not yet fully trained, the discriminative information of the training fused image tends to be closer to 1 (the closer to 1, the more fake it is). For example, the discriminative information of the training fused image is 0.85.

[0129] The first annotation information of the training fused image corresponding to the discriminator and the second annotation information of the training fused image corresponding to the image generation model are determined. Since the training fused image is false relative to the discriminator, the first annotation information of the training fused image corresponding to the discriminator is 1. The training objective of the image generation model is to make the generated image approach the truth in the image generation model; therefore, the second annotation information of the training fused image corresponding to the image generation model is 0.

[0130] Optionally, the server can determine a first target loss based on the discrimination information 0.85 and the first annotation information 1 corresponding to the discriminator in the training fused image, and determine a second target loss based on the discrimination information 0.85 and the second annotation information 0 corresponding to the image generation model in the training fused image. Then, the server trains the untrained image generation model and the discriminator based on the constraint loss function, the first target loss, and the second target loss.

[0131] Optionally, the formula for the constraint loss function can be expressed as:

[0132] min w max θ L AMS (f(G θ (I f ,I r )),y;w)……(3)

[0133] Among them, I f It is a training synthetic image, I r It is the training reference image, G θ (I f ,I r To determine the training fusion region information and training fusion intensity information, f(G) θ(I f ,I r ) represents the generation of the training fused image, where θ is the model parameter in the fusion generator, w is the model parameter in the discriminator, and y is the discrimination information of the generated training fused image relative to the discriminator.

[0134] By this point, the image generation model and discriminator have completed one round of training, resulting in an updated image generation model and discriminator.

[0135] Optionally, the server can continue to train the image generation model and discriminator (i.e., adversarial network) for multiple rounds using the original training reference images and training synthetic images, or it can continue to train the image generation model and discriminator for multiple rounds using new training reference images and training synthetic images until the iteration termination condition is met. The trained image generation model is then identified as the trained image generation model, and the trained discriminator is identified as the trained discriminator.

[0136] Optionally, the iteration termination condition includes a preset number of iterations.

[0137] In an optional embodiment, the image generation model and discriminator described above can be various neural networks, such as neural convolutional networks, grouped convolutional networks, Xception networks, etc.

[0138] Image generation models and discriminators are types of machine learning models. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0139] In this embodiment, after obtaining the trained discriminator, the server can acquire the image to be detected and perform detection analysis on the image based on the trained discriminator. That is, the image to be detected is input into the trained discriminator to obtain the detection result corresponding to the image, such as true or false. Optionally, the image to be detected can be the target fusion image mentioned above.

[0140] Please see Figure 8The diagram shows a structural schematic of an image generation device provided in an embodiment of the present invention. This device has the function of implementing the image generation method described in the above-described method embodiments. This function can be implemented in hardware or by hardware executing corresponding software. Figure 8 As shown, the image generation apparatus 800 may include:

[0141] Image acquisition module 801 is used to acquire a reference image to be fused and a composite image; wherein, the reference image includes a reference object; and the composite image includes a composite object;

[0142] The fusion information determination module 802 is used to perform fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the composite image.

[0143] The synthesis module 803 is used to fuse the reference image and the synthesized image based on the fusion region information and the fusion intensity information to obtain the target fused image.

[0144] In some possible embodiments, the synthesis module is used for:

[0145] When fusing regional information to represent the internal region of an object, the image synthesizer in the image generation model obtains the reference internal object from the reference object of the reference image and the synthesized internal object from the synthesized object of the synthesized image.

[0146] The reference internal object and the synthesized internal object are fused based on the fusion strength information to obtain the fused internal object;

[0147] Identify the external objects to be merged;

[0148] The target fused image containing the fused objects is determined based on the fused internal objects and the external objects to be fused.

[0149] In some possible embodiments, the synthesis module is used for:

[0150] When representing the object region by fusing regional information, the image synthesizer in the image generation model obtains the reference object from the reference image and the synthesized object from the synthesized image.

[0151] The reference object and the composite object are fused based on the fusion intensity information to obtain a target fused image containing the fused object.

[0152] In some possible embodiments, the image acquisition module is used for:

[0153] Obtain a reference image;

[0154] Determine the feature information of the reference object in the reference image;

[0155] Determine the feature information of the object in each candidate image in the candidate image library;

[0156] The feature distance value corresponding to each candidate image is determined based on the feature information of the object in each candidate image and the feature information of the reference object;

[0157] A composite image is determined from the candidate image library based on the feature distance value corresponding to each candidate image.

[0158] In some possible embodiments, the image acquisition module is used for;

[0159] Perform attitude analysis on the reference object in the reference image to obtain the pitch, yaw and roll information of the reference object;

[0160] The attitude characteristics of the reference object are determined based on pitch, yaw, and roll information.

[0161] and / or;

[0162] Attribute analysis is performed on the reference object in the reference image to obtain the attribute feature information of the reference object.

[0163] In some possible embodiments, the image acquisition module is used for:

[0164] Candidate images are determined from the candidate image library based on the feature distance value corresponding to each candidate image;

[0165] Determine the color feature information of the reference object in the reference image;

[0166] The color feature information of objects in the candidate image is adjusted based on the color feature information of the reference object to obtain the composite image;

[0167] In the synthesized image, the matching value between the color feature information of the adjusted object and the color feature information of the reference object is greater than or equal to a preset threshold.

[0168] In some possible embodiments, the fusion information determination module is used for:

[0169] The fusion generator in the image generation model performs fusion analysis on objects in the reference image and the synthesized image to obtain fusion region information and fusion intensity information corresponding to the reference image and the synthesized image.

[0170] In some possible embodiments, the method further includes a network training module for:

[0171] Obtain the training reference image and the training synthesized image to be fused; wherein, the training reference image includes the training reference object; and the training synthesized image includes the training synthesized object;

[0172] The training reference image and the training synthesized image are input into the fusion method generator in the untrained image generation model to obtain the training fusion region information and the training fusion intensity information.

[0173] The training reference image, the training synthesized image, the training fusion region information, and the training fusion intensity information are input into the image synthesizer in the untrained image generation model to obtain the training fusion image.

[0174] The untrained image generation model and discriminator are trained based on the training fused images to obtain the trained image generation model and discriminator.

[0175] In some possible embodiments, the fusion information determination module is used for:

[0176] The fused training image is input into the discriminator to obtain discrimination information;

[0177] The first target loss is determined based on the discrimination information and the first annotation information of the training fused image corresponding to the discriminator;

[0178] The second target loss is determined based on the discriminant information and the second annotation information corresponding to the image generation model in the training fused image;

[0179] The untrained image generation model and discriminator are trained based on the constraint loss function, the first objective loss and the second objective loss;

[0180] If the iteration termination condition is met, the trained image generation model and discriminator are obtained.

[0181] In some possible embodiments, the method further includes a detection module for:

[0182] Acquire the image to be detected;

[0183] The trained discriminator is used to analyze the image to be detected, and the detection result corresponding to the image is obtained.

[0184] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0185] This invention provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program segment, which is loaded and executed by the processor to implement the image generation method provided in the above method embodiments.

[0186] The memory can be used to store software programs and modules. The processor executes various functional applications and image generation by running the software programs and modules stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for the functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0187] The methods and embodiments provided in this invention can be executed on a computer terminal, server, or similar computing device. Taking running on a server as an example... Figure 9 This is a hardware structure block diagram of a server running an image generation method provided in an embodiment of the present invention, such as... Figure 9 As shown, the server 1000 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1010 (CPUs 1010 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGs), a memory 1030 for storing data, and one or more storage media 1020 (e.g., one or more mass storage devices) for storing application programs 1023 or data 1022. The memory 1030 and storage media 1020 may be temporary or persistent storage. The program stored in the storage media 1020 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 1010 may be configured to communicate with the storage media 1020 and execute the series of instruction operations stored in the storage media 1020 on the server 1000. Server 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server™, macOS™, Unix™, Linux™, FreeBSD™, etc.

[0188] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 1000. In one example, the input / output interface 1040 includes a network adapter (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1040 may be a radio frequency (RF) module used for wireless communication with the Internet.

[0189] Those skilled in the art will understand that Figure 9 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 1000 may also include... Figure 9 The more or fewer components shown, or having the same Figure 9 The different configurations shown.

[0190] Embodiments of the present invention also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to implementing an image generation method, wherein the at least one instruction or the at least one program is loaded and executed by the processor to implement the image generation method provided in the above-described method embodiments.

[0191] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0192] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0193] Embodiments of the present invention also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the image generation method described above.

[0194] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0195] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0196] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An image generation method, characterized in that, The method includes: Obtain a reference image and a composite image to be fused; wherein the reference image includes a reference object; and the composite image includes a composite object. A fusion analysis is performed on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image. The fusion intensity information is used to characterize the proportion occupied by the reference object or the composite object during the fusion process, and the fusion region information is used to characterize the fusion region in the reference object and the composite object. In the case that the reference image and the composite image are face images, the fusion region includes an inner face region, an outer face region, or a full face region. The fusion region information and the fusion intensity information correspond to different fusion methods of the reference image and the composite image. The reference image and the synthesized image are fused based on the fusion region information and the fusion intensity information to obtain the target fused image.

2. The method according to claim 1, characterized in that, The step of fusing the reference image and the synthesized image based on the fusion region information and the fusion intensity information to obtain the target fused image includes: When the fused region information represents the internal region of the object, the image synthesizer in the image generation model obtains the reference internal object from the reference object of the reference image and the synthesized internal object from the synthesized object of the synthesized image. The reference internal object and the synthesized internal object are fused according to the fusion strength information to obtain a fused internal object; Identify the external objects to be merged; The target fused image containing the fused objects is determined based on the fused internal objects and the external objects to be fused.

3. The method according to claim 1, characterized in that, The step of fusing the reference image and the synthesized image based on the fusion region information and the fusion intensity information to obtain the target fused image includes: When representing the object region with the fused region information, the reference object is obtained from the reference image based on the image synthesizer in the image generation model, and the synthesized object is obtained from the synthesized image; The reference object and the composite object are fused according to the fusion intensity information to obtain the target fused image containing the fused object.

4. The method according to any one of claims 2-3, characterized in that, The process of obtaining the reference image and the composite image to be fused includes: Obtain a reference image; Determine the feature information of the reference object in the reference image; Determine the feature information of the object in each candidate image in the candidate image library; The feature distance value corresponding to each candidate image is determined based on the feature information of the object in each candidate image and the feature information of the reference object; The synthesized image is determined from the candidate image library based on the feature distance value corresponding to each candidate image.

5. The method according to claim 4, characterized in that, The step of determining the feature information of the reference object in the reference image includes: The attitude analysis of the reference object in the reference image is performed to obtain the pitch information, yaw information and roll information of the reference object. The attitude characteristic information of the reference object is determined based on the pitch information, the yaw information, and the roll information; and / or; Attribute analysis is performed on the reference object in the reference image to obtain the attribute feature information of the reference object.

6. The method according to claim 4, characterized in that, The step of determining the synthesized image from the candidate image library based on the feature distance value corresponding to each candidate image includes: Candidate images are determined from the candidate image library based on the feature distance value corresponding to each candidate image; Determine the color feature information of the reference object in the reference image; The color feature information of the objects in the candidate image is adjusted based on the color feature information of the reference object to obtain the synthesized image; In the synthesized image, the matching value between the color feature information of the adjusted object and the color feature information of the reference object is greater than or equal to a preset threshold.

7. The method according to claim 2, characterized in that, The step of performing fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image includes: The fusion generator in the image generation model performs fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain the fusion region information and fusion intensity information corresponding to the reference image and the composite image.

8. The method according to claim 7, characterized in that, The method further includes: Obtain the training reference image and the training synthesis image to be fused; wherein, the training reference image includes a training reference object; and the training synthesis image includes a training synthesis object; The training reference image and the training synthesized image are input into the fusion method generator in the untrained image generation model to obtain training fusion region information and training fusion intensity information. The training reference image, the training synthesized image, the training fusion region information, and the training fusion intensity information are input into the image synthesizer in the untrained image generation model to obtain the training fusion image. The untrained image generation model and discriminator are trained based on the training fused images to obtain the trained image generation model and discriminator.

9. The method according to claim 8, characterized in that, The step of training the untrained image generation model and discriminator based on the training fused image to obtain the trained image generation model and discriminator includes: The trained fused image is input into the discriminator to obtain discrimination information; The first target loss is determined based on the discrimination information and the first annotation information of the training fusion image corresponding to the discriminator; The second target loss is determined based on the discrimination information and the second annotation information of the training fused image corresponding to the image generation model; The untrained image generation model and discriminator are trained based on the constraint loss function, the first target loss, and the second target loss. If the iteration termination condition is met, the trained image generation model and discriminator are obtained.

10. The method according to claim 9, characterized in that, The method further includes: Acquire the image to be detected; The discriminator, after training, is used to perform detection analysis on the image to be detected, and the detection result corresponding to the image to be detected is obtained.

11. An image generation apparatus, characterized in that, The device includes: An image acquisition module is used to acquire a reference image to be fused and a composite image; wherein, the reference image includes a reference object; and the composite image includes a composite object; The fusion information determination module is used to perform fusion analysis on the reference object in the reference image and the composite object in the composite image to obtain fusion region information and fusion intensity information corresponding to the reference image and the composite image. The fusion intensity information is used to characterize the proportion occupied by the reference object or the composite object in the fusion process, and the fusion region information is used to characterize the fusion region in the reference object and the composite object. In the case that the reference image and the composite image are face images, the fusion region includes an inner face region, an outer face region, or a full face region. The fusion region information and the fusion intensity information correspond to different fusion methods of the reference image and the composite image. The fusion module is used to fuse the reference image and the fused image based on the fusion region information and the fusion intensity information to obtain a target fused image.

12. An electronic device, characterized in that, The method includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the image generation method as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the image generation method as described in any one of claims 1 to 10.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the image generation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processing method, device and equipment and storage medium

    CN111063008A

  • Image processing method and device, equipment and storage medium

    CN114049290A