Image synthesis method, device, electronic device and storage medium
By acquiring and converting high-definition preset background images and replacing the actual background images in live videos, the problem of background replacement in the prior art that is difficult to achieve high precision and fast in live broadcast scenarios is solved, and high-quality live broadcast screen effects are achieved.
Patent Information
- Application Number
- CN202210126942.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-02-11
AI Technical Summary
The prior art is difficult to meet the high-precision and fast-speed requirements of image background replacement in live broadcast scenarios, and it is impossible to effectively handle background replacement in live broadcast videos.
By obtaining the actual background image of the source image, obtaining a high-definition preset background image, and using the conversion model to convert it into a converted background image, replacing the actual background image in the source image, thereby obtaining a high-precision foreground image and synthesizing a high-quality synthetic image.
It realizes high-precision and fast image processing in live broadcast scenarios, improves the imaging effect of live broadcast images, and can meet the real-time background replacement requirements of live broadcast videos.
Smart Images

Figure CN114463241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image synthesis method, apparatus, electronic device, and storage medium. Background Art
[0002] Replacing the background of a person means extracting the person from the original image or video frame and pasting the person onto other preset backgrounds, so as to achieve the effect of background replacement.
[0003] However, most of the current background replacement methods are difficult to simultaneously meet the processing requirements of accuracy and speed, and cannot meet the background replacement processing requirements in the live broadcast scenario. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide an image synthesis method, apparatus, electronic device, and storage medium with high precision and high speed to at least partially solve the above problems.
[0005] According to one aspect of the present disclosure, there is provided an image synthesis method, including: obtaining a preset background image of the actual background image according to the actual background image of the source image; performing conversion processing on the preset background image to obtain a converted background image of the preset background image; obtaining a foreground image of the source image according to the source image and the converted background image; and synthesizing the preset background image and the foreground image to obtain a synthesized image of the source image.
[0006] According to another aspect of the present disclosure, there is provided an image synthesis apparatus, including: a background processing module, configured to obtain a preset background image of the actual background image according to the actual background image of the source image, and perform conversion processing on the preset background image to obtain a converted background image of the preset background image; a foreground processing module, configured to obtain a foreground image of the source image according to the source image and the converted background image; and a synthesis module, configured to synthesize the preset background image and the foreground image to obtain a synthesized image of the source image.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the image synthesis method described above.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the image synthesis method described above.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to execute the text recognition model training method as described in the first aspect.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the text recognition model training method as described in the first aspect.
[0011] The image synthesis solution provided by one or more embodiments of the present disclosure can simultaneously meet the requirements of fast speed and high-precision image processing, and can meet the image synthesis processing in the live video scenario to improve the imaging effect of the live broadcast screen. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features and advantages of the present disclosure are disclosed. In the drawings:
[0013] Figure 1 It is a schematic flowchart of the image synthesis method according to an exemplary embodiment of the present disclosure.
[0014] Figure 2 It is a schematic flowchart of the image synthesis method according to another exemplary embodiment of the present disclosure.
[0015] Figure 3 It is a schematic flowchart of the image synthesis method according to another exemplary embodiment of the present disclosure.
[0016] Figure 4 It is a schematic flowchart of the image synthesis method according to another exemplary embodiment of the present disclosure.
[0017] Figure 5 It is a schematic flowchart of the image synthesis method according to another exemplary embodiment of the present disclosure.
[0018] Figure 6 It is a block diagram of the structure of the image synthesis device according to an exemplary embodiment of the present disclosure.
[0019] Figure 7 It is a block diagram of the structure of the electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0021] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0022] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions executed by these devices, modules or units or their interdependent relationships.
[0023] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0024] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0025] With the continuous maturity of the character background replacement technology, it has been widely applied to scenarios such as film shooting and production, picture post-editing, and poster production.
[0026] Currently, the mainstream portrait matting technology usually has relatively high requirements for auxiliary input, making it difficult for these matting technologies to simultaneously meet the processing requirements of high precision and high speed, and thus unable to be applied to background replacement processing in live broadcast scenarios.
[0027] Specifically, most of the current mainstream portrait matting solutions are deep learning-based technologies that rely on auxiliary inputs outside the target image, such as trimaps, background images (e.g., an image taken from the same angle without the target object). Such auxiliary inputs mean additional annotation or shooting processes, resulting in the current portrait matting technology being unable to meet the real-time matting and background replacement processing requirements in live broadcast scenarios. However, matting processing without relying on auxiliary inputs has the problem of poor matting effects.
[0028] In view of this, the present disclosure provides an image synthesis processing solution that can solve various problems existing in the above-mentioned prior art.
[0029] The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0030] Figure 1 It is a processing flowchart of the image synthesis method according to an exemplary embodiment of the present disclosure. As shown in the figure, this embodiment mainly includes the following steps:
[0031] Step S102, according to the actual background image of the source image, obtain the preset background image of the actual background image.
[0032] Optionally, the source image can be obtained from live video data.
[0033] Optionally, the preset background image and the actual background image are two images with the same imaging content but different imaging styles.
[0034] In this embodiment, the preset background image is a high-definition background image generated using graphics software, and the actual background image is a background image with an actual imaging style (e.g., a real shooting picture style) generated by using a camera device to shoot the preset background image.
[0035] Specifically, in some cases, due to different actual imaging conditions (e.g., lighting environment conditions) of the camera device, there are significant differences in the colors between the generated actual background image and the preset background image. In other cases, due to a certain degree of deviation in the actual imaging angle of the camera device, there is an affine transformation between the actual background image and the preset background image. In view of this, in this step, according to the actual background image of the source image, obtaining a preset background image with higher imaging quality can help improve the subsequent matting processing effect.
[0036] Step S104, perform a conversion process on the preset background image to obtain the converted background image of the preset background image.
[0037] Optionally, a conversion model can be utilized to perform conversion processing on a preset background image according to a preset imaging style, so as to obtain a converted background image that meets the preset imaging style.
[0038] In this embodiment, the preset imaging style can be determined according to the actual imaging style of the source image, so that the generated converted background image fits the actual imaging style of the actual background image.
[0039] Optionally, the conversion model can include but is not limited to: a cyclic adversarial generation model (Cycle-GAN model).
[0040] Step S106: Obtain the foreground image of the source image according to the source image and the converted background image.
[0041] Optionally, the actual background image in the source image can be replaced with the converted background image, and a matting process can be performed based on the converted background image and the source image to obtain the foreground image of the source image (for example, the portrait part of the source image). By means of the above technical means, the matting effect of the foreground image can be improved, and a high-precision foreground image can be obtained.
[0042] Optionally, a matting model including a deep learning network structure can be utilized to obtain the foreground image in the source image.
[0043] In this embodiment, the deep learning network structure can include but is not limited to: the BGMv2 (Background Matting v2) network structure.
[0044] Step S108: Synthesize the preset background image and the foreground image to obtain a synthesized image of the source image.
[0045] Specifically, the foreground image obtained in step S106 can be synthesized with the preset background image obtained in step S102 to obtain a synthesized image with a better imaging effect.
[0046] In summary, the image synthesis method of this embodiment obtains a high-definition preset background image according to the actual background image of the source image, and uses the converted background image converted from the preset background image to replace the actual background image in the source image, so as to use the converted background image as an auxiliary input to obtain a high-precision foreground image from the source image, and by synthesizing the high-precision foreground image and the high-definition preset background image, a synthesized image with a better imaging quality can be obtained. Accordingly, the image synthesis processing means provided in this embodiment has the advantages of fast processing speed and high processing precision, and can be well applied to image synthesis in a live broadcast scenario to obtain a high-quality live broadcast picture effect.
[0047] Figure 2Schematic flowchart of an image synthesis method according to another exemplary embodiment of the present disclosure. This embodiment is a specific implementation of the above step S104. As shown in the figure, this embodiment mainly includes the following steps:
[0048] Step S202: Obtain a first training sample with a source domain and a second training sample with a target domain.
[0049] In this embodiment, the first training sample and the second training sample can be two training images with the same imaging content but different imaging styles.
[0050] For example, the first training sample is a training image with a source domain generated by using drawing software, and the second training sample is a training image with a target domain (for example, the actual imaging style under real shooting scene conditions) generated by using a camera device to photograph the first training sample.
[0051] Step S204: Use the first training sample as the input and the second training sample as the output to train the first generator of the conversion model to obtain a trained first generator.
[0052] Optionally, the first generator can be used to perform a first conversion prediction on the source domain of the first training sample according to the target domain of the second training sample to obtain a first predicted image, and the first discriminator of the conversion model can be used to compare and discriminate the first predicted image and the second training sample to obtain the discrimination result of the first discriminator, and the first generator can be optimized and updated based on the discrimination result of the first discriminator, and the steps of the first conversion prediction can be repeatedly executed by using the optimized first generator until the discrimination result of the first discriminator meets the first preset convergence condition to obtain a trained first generator.
[0053] In this embodiment, when the first discriminator cannot distinguish the difference between the first predicted image and the second training sample, or when the first generator completes the first conversion prediction of the first training sample in a preset batch, the discrimination result of the first discriminator meets the first preset convergence condition.
[0054] Step S206: Use the second training sample as the input and the first training sample as the output to train the second generator of the conversion model to obtain a trained second generator.
[0055] Optionally, a second generator may be utilized to perform a second conversion prediction on the target domain of the second training sample according to the source domain of the first training sample, obtain a second predicted image, and use the second discriminator of the conversion model to compare and discriminate according to the second predicted image and the first training sample to obtain the discrimination result of the second discriminator, and optimize and update the second generator based on the discrimination result of the second discriminator, and use the optimized second generator to repeat the steps of the second conversion prediction until the discrimination result of the second discriminator meets the second preset convergence condition to obtain a trained second generator.
[0056] In this embodiment, when the second discriminator cannot distinguish the difference between the second predicted image and the first training sample, or when the second generator completes the second conversion prediction of the second training sample in a preset batch, the discrimination result of the second discriminator meets the second preset convergence condition.
[0057] Step S208, determine a trained conversion model according to the trained first generator and the trained second generator.
[0058] Specifically, when both the first generator and the second generator are trained, it means that the training of the conversion model is completed.
[0059] In summary, the image synthesis method of this embodiment uses a conversion model with two generators and two discriminators, and learns the mapping relationship between the source domain and the target domain bidirectionally during the training process, so that the trained conversion model can adapt to various visual problem scenarios such as super-resolution, style transformation, and image enhancement, so as to achieve a better imaging style conversion effect, make the imaging style of the converted background image closer to the imaging style of the actual background image, and thus improve the matting effect of the subsequent foreground image.
[0060] Figure 3 It is a schematic flowchart of an image synthesis method according to another exemplary embodiment of the present disclosure. This embodiment is a specific implementation of the above step S106. As shown in the figure, this embodiment mainly includes the following steps:
[0061] Step S302, use the basic network of the matting model to perform a preliminary prediction according to the source image and the converted background image to obtain a preliminary prediction result.
[0062] Optionally, the converted background image may be used as an auxiliary input to be used as the input of the basic network together with the source image for performing region position localization prediction to obtain a preliminary prediction result including foreground probability (Alpha), foreground residual (ForegroundResidual), error map (Error Map), and hidden layer node features (Hidden).
[0063] Step S304: Use the refinement network of the matting model to perform refinement prediction based on the preliminary prediction result to obtain the foreground image.
[0064] Optionally, the refinement network of the matting model can be used to perform refinement prediction based on the preliminary prediction structure including the foreground probability, foreground residual, error map, and hidden layer node features output by the base network, so as to obtain the foreground image of the source image (such as the portrait part of the source image).
[0065] In summary, in this embodiment, the actual background image in the source image is replaced with a transformed background image as an auxiliary input for the matting model, which can not only enable the matting model to obtain a high-precision foreground image from the source image, but also assist in improving the imaging effect of the subsequent composite image.
[0066] Figure 4 It is a processing flow chart of an image synthesis method according to another exemplary embodiment of the present application. As shown in the figure, this embodiment mainly includes the following steps:
[0067] Step S402: Collect multiple original background samples, and use the trained transformation model to perform transformation processing on each original background sample based on a preset imaging style to obtain multiple target background samples that meet the preset imaging style.
[0068] Optionally, the original background sample can be a high-definition background image generated using graphics software.
[0069] Optionally, the trained transformation model can be used to perform transformation processing on the original background sample to obtain a target background sample with an actual imaging style. For example, a target background sample with a real shooting picture style that fits the shooting picture of the imaging device can be generated.
[0070] Step S404: Combine the foreground image obtained by the matting model with any one of the target background samples to obtain an augmented training sample.
[0071] Specifically, the foreground image obtained by the matting model (for example, the portrait in the source image) can be synthesized with any one of the target background samples generated based on the transformation model to obtain multiple augmented training samples with different background effects.
[0072] Step S406: Use the augmented training sample and the foreground image to train the matting model.
[0073] Specifically, the augmented training sample can be used as the input and the foreground image can be used as the output to train the matting model.
[0074] In summary, the present embodiment can utilize the trained conversion model to effectively expand the existing training samples for training the matte model, which can improve the robustness and prediction accuracy of the matte model.
[0075] Figure 5 It is a processing flowchart of an image synthesis method according to another exemplary embodiment of the present disclosure. As shown in the figure, this embodiment mainly includes the following steps:
[0076] Step S502: Obtain the video data to be processed.
[0077] Optionally, the video data to be processed may include pre-recorded video data or live video data generated in real time.
[0078] Step S504: Extract each video frame from the video data to be processed based on a preset frame interval to obtain each source image.
[0079] In this embodiment, the preset frame interval can be set based on actual video synthesis requirements, and the present disclosure does not limit this.
[0080] Step S506: Stitch the composite images of each source image to obtain the composite video data of the video data to be processed.
[0081] In this embodiment, the image synthesis method described in any one of Figures 1 to 4 can be used to obtain the composite image of the source image, and according to the timestamp of each source image, stitch the composite images of each source image to obtain the composite video data of the video data to be processed.
[0082] In summary, the present application performs frame-by-frame processing on the video data to be processed to obtain the corresponding composite images for each source image, and then stitches the composite images to obtain the composite video data of the video data to be processed. By means of this technical means, video data with better imaging quality can be obtained to improve the video viewing experience.
[0083] Figure 6 It shows a schematic architecture diagram of an image synthesis apparatus according to an exemplary embodiment of the present disclosure. As shown in the figure, the image synthesis apparatus 600 of this embodiment mainly includes:
[0084] A background processing module 602, configured to obtain a preset background image of the actual background image according to the actual background image of the source image, and perform conversion processing on the preset background image to obtain a converted background image of the preset background image.
[0085] A foreground processing module 604, configured to obtain the foreground image of the source image according to the source image and the converted background image.
[0086] A synthesis module 606, configured to synthesize a preset background image and a foreground image to obtain a synthesized image of the source image.
[0087] Optionally, the background processing module 602 may further be configured to: perform a conversion process on the preset background image according to a preset imaging style by using a conversion model to obtain a converted background image that meets the preset imaging style. Exemplarily, the preset imaging style may be determined according to the actual imaging style of the source image.
[0088] Optionally, the conversion model may be trained in the following manner: obtain a first training sample with a source domain and a second training sample with a target domain; use the first training sample as an input and the second training sample as an output to train a first generator of the conversion model to obtain a trained first generator, and use the second training sample as an input and the first training sample as an output to train a second generator of the conversion model to obtain a trained second generator, and determine a trained conversion model according to the trained first generator and the trained second generator.
[0089] Exemplarily, the first generator may be used to perform a first conversion prediction on the source domain of the first training sample according to the target domain of the second training sample to obtain a first predicted image, and a first discriminator of the conversion model may be used to perform discrimination according to the first predicted image and the second training sample. According to the discrimination result of the first discriminator, the first conversion prediction of the first generator is repeatedly performed until the first discriminator meets a first preset convergence condition to obtain a trained first generator.
[0090] Exemplarily, the second generator may be used to perform a second conversion prediction on the target domain of the second training sample according to the source domain of the first training sample to obtain a second predicted image, and a second discriminator of the conversion model may be used to perform discrimination according to the second predicted image and the first training sample. According to the discrimination result of the second discriminator, the second conversion prediction of the second generator is repeatedly performed until the second discriminator meets a second preset convergence condition to obtain a trained second generator.
[0091] Optionally, the foreground processing module 604 may further be configured to: perform a preliminary prediction on the source image and the converted background image by using a basic network of a matting model to obtain a preliminary prediction result, and perform a refined prediction on the preliminary prediction result by using a refined network of the matting model to obtain a foreground image.
[0092] Optionally, the image synthesis device 600 further includes: a sample augmentation module, configured to collect a plurality of original background samples, perform a conversion process on each original background sample based on a preset imaging style by using a trained conversion model to obtain a plurality of target background samples that meet the preset imaging style, and combine the foreground image obtained by using the matting model with any one of the target background samples to obtain an augmented training sample.
[0093] Optionally, a matting model can be trained using augmented training samples and foreground images.
[0094] Optionally, the image synthesis device 600 further includes: a video acquisition module, a processing module, and a splicing module. The video acquisition module is configured to acquire video data to be processed. The processing module is configured to: extract each video frame from the video data to be processed based on a preset frame interval to obtain each source image. The splicing module is configured to: splice the synthesized images of each source image to obtain synthesized video data of the video data to be processed.
[0095] In addition, the image synthesis device 600 according to the embodiments of the present disclosure can also be used to implement other steps in the foregoing embodiments of each image synthesis method, and has the beneficial effects of the corresponding method step embodiments, which will not be elaborated herein.
[0096] An exemplary embodiment of the present disclosure further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and when the computer program is executed by the at least one processor, it is configured to cause the electronic device to execute the method according to the embodiments of the present disclosure.
[0097] An exemplary embodiment of the present disclosure further provides a non-transitory computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor of a computer, it is configured to cause the computer to execute the method according to the embodiments of the present disclosure.
[0098] An exemplary embodiment of the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor of a computer, it is configured to cause the computer to execute the method according to the embodiments of the present disclosure.
[0099] Referring to Figure 7 , a block diagram of an electronic device 700 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0100] As Figure 7As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0101] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 can be any type of device capable of inputting information into the electronic device 700. The input unit 706 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 707 can be any type of device capable of presenting information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 704 can include but is not limited to a magnetic disk, an optical disk. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0102] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above. For example, in some embodiments, the image synthesis method of the foregoing embodiments can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 700 via the ROM 702 and / or the communication unit 709. In some embodiments, the computing unit 701 can be configured to execute the image synthesis method in any other appropriate manner (e.g., by means of firmware).
[0103] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0104] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0105] As used in the present disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0106] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0107] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0108] A computer system can include clients and servers. Clients and servers are generally far apart from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. An image synthesis method, comprising: obtaining a preset background image of the actual background image according to the actual background image of the source image; performing a conversion process on the preset background image by using a conversion model to obtain a converted background image of the preset background image; obtaining a foreground image of the source image according to the source image and the converted background image; synthesizing the preset background image and the foreground image to obtain a synthesized image of the source image; wherein, the conversion model is trained in the following manner: obtaining a first training sample with a source domain and a second training sample with a target domain; using the first training sample as an input and the second training sample as an output to train a first generator of the conversion model to obtain a trained first generator; using the second training sample as an input and the first training sample as an output to train a second generator of the conversion model to obtain a trained second generator; determining a trained conversion model according to the trained first generator and the trained second generator.
2. The image synthesis method according to claim 1, wherein, the performing a conversion process on the preset background image by using a conversion model to obtain a converted background image of the preset background image includes: performing a conversion process on the preset background image by using the conversion model according to a preset imaging style to obtain the converted background image that meets the preset imaging style; wherein, the preset imaging style can be determined according to the actual imaging style of the source image.
3. The image synthesis method according to claim 1, wherein, the using the first training sample as an input and the second training sample as an output to train a first generator of the conversion model to obtain a trained first generator includes: using the first generator to perform a first conversion prediction on the source domain of the first training sample according to the target domain of the second training sample to obtain a first predicted image; using a first discriminator of the conversion model to perform discrimination according to the first predicted image and the second training sample to obtain a discrimination result of the first discriminator; repeating the first conversion prediction of the first generator according to the discrimination result of the first discriminator until the discrimination result of the first discriminator meets a first preset convergence condition to obtain the trained first generator; wherein, the using the second training sample as an input and the first training sample as an output to train a second generator of the conversion model to obtain a trained second generator includes: using the second generator to perform a second conversion prediction on the target domain of the second training sample according to the source domain of the first training sample to obtain a second predicted image; using a second discriminator of the conversion model to perform discrimination according to the second predicted image and the first training sample to obtain a discrimination result of the second discriminator; According to the discrimination result of the second discriminator, repeat the second conversion prediction of the second generator until the discrimination result of the second discriminator meets the second preset convergence condition, so as to obtain the trained second generator.
4. The image synthesis method according to any one of claims 1 to 3, wherein, obtaining the foreground image of the source image according to the source image and the converted background image includes: using the basic network of the matting model to perform preliminary prediction according to the source image and the converted background image, and obtaining a preliminary prediction result; using the fine-tuning network of the matting model to perform fine-tuning prediction according to the preliminary prediction result, and obtaining the foreground image of the source image.
5. The image synthesis method according to claim 4, wherein, the method further includes: collecting a plurality of original background samples, and using the trained conversion model to perform conversion processing on each of the original background samples based on a preset imaging style to obtain a plurality of target background samples that meet the preset imaging style; combining the foreground image obtained by the matting model with any one of the target background samples to obtain an augmented training sample; using the augmented training sample and the foreground image to train the matting model.
6. The image synthesis method according to claim 1, wherein, the method further includes: obtaining video data to be processed; extracting each video frame in the video data to be processed based on a preset frame interval to obtain each source image; stitching the synthesized images of each source image to obtain the synthesized video data of the video data to be processed.
7. An image synthesis device, including: a background processing module, configured to obtain a preset background image of the actual background image according to the actual background image of the source image, and use a conversion model to perform conversion processing on the preset background image to obtain a converted background image of the preset background image; a foreground processing module, configured to obtain the foreground image of the source image according to the source image and the converted background image; a synthesis module, configured to synthesize the preset background image and the foreground image to obtain a synthesized image of the source image; wherein, the conversion model is trained in the following manner: obtaining a first training sample with a source domain and a second training sample with a target domain; using the first training sample as an input, using the second training sample as an output, training the first generator of the conversion model to obtain a trained first generator, and using the second training sample as an input, using the first training sample as an output, training the second generator of the conversion model to obtain a trained second generator, and determining the trained conversion model according to the trained first generator and the trained second generator.
8. An electronic device, including: a processor; and a memory storing a program, wherein, the program includes instructions that, when executed by the processor, cause the processor to execute the method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Method for replacing background in picture, equipment, storage medium and program product
CN113259698A