Image synthesis method, apparatus, device, and medium

By decomposing and fusing the intrinsic maps of the background and foreground, and combining them with background lighting information to render the image, the problem of visual disharmony between the foreground and background is solved, the visual effect of image synthesis is improved, and the need for manual image retouching is reduced.

CN116128777BActive Publication Date: 2026-05-12MASHANG CONSUMER FINANCE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2022-09-27
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, image synthesis methods result in visual disharmony between foreground and background images, and manual image retouching is time-consuming and labor-intensive.

Method used

Multiple background intrinsic images and background lighting information are obtained by decomposing the background image. The foreground image is also decomposed accordingly. Background intrinsic images and foreground intrinsic images of the same intrinsic image category are merged and combined with background lighting information to render a composite image.

Benefits of technology

It achieves consistency in lighting and shadow between the foreground and background, improves the visual effect of the composite image, and reduces the time and labor required for manual image retouching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116128777B_ABST
    Figure CN116128777B_ABST
Patent Text Reader

Abstract

The embodiments of the present specification provide an image synthesis method, device, equipment and medium. The image synthesis method can include: decomposing a background image to obtain a plurality of background eigenimages and background illumination information, each of the background eigenimages corresponding to an eigenimage category; decomposing a foreground image to obtain a plurality of foreground eigenimages corresponding to the eigenimage category; respectively fusing the background eigenimages and the foreground eigenimages corresponding to the same eigenimage category to obtain a plurality of fused eigenimages; and synthesizing an image according to the plurality of fused eigenimages and the background illumination information. The synthesized image has a better visual effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described in this specification relate to the field of image processing technology, specifically to an image synthesis method, apparatus, device, and medium. Background Technology

[0002] Currently, in the field of image processing, it's common to process two or more images to create a new image. Specifically, for example, image A might contain a tree, and image B might contain a grassland. It might be necessary to place the tree from image A as a foreground object into image B to obtain the composite image.

[0003] Therefore, in related technologies, the final image is usually synthesized by placing the foreground image in a specified position within the background image and then performing adaptive editing. However, the final image obtained by this image synthesis method often lacks visual harmony between the foreground and background images. Summary of the Invention

[0004] This specification provides a method, apparatus, device, and medium for image synthesis, through various embodiments. These methods can generate synthesized images with superior visual effects.

[0005] One embodiment of this specification provides an image synthesis method, comprising: decomposing a background image to obtain multiple background intrinsic images and background illumination information, each of the background intrinsic images corresponding to an intrinsic image category; decomposing a foreground image to obtain multiple foreground intrinsic images corresponding to the intrinsic image categories; fusing the background intrinsic images and foreground intrinsic images corresponding to the same intrinsic image category to obtain multiple fused intrinsic images; and synthesizing an image based on the multiple fused intrinsic images and the background illumination information.

[0006] One embodiment of this specification provides an image synthesis apparatus, comprising: a background decomposition unit for decomposing a background image to obtain multiple background intrinsic images and background illumination information; each of the background intrinsic images corresponds to an intrinsic image category; a construction unit for decomposing a foreground image to obtain multiple foreground intrinsic images corresponding to the intrinsic image categories; a fusion unit for fusing the background intrinsic images and foreground intrinsic images corresponding to the same intrinsic image category to obtain multiple fused intrinsic images; and a synthesis unit for synthesizing an image based on the multiple fused intrinsic images and the background illumination information.

[0007] One embodiment of this specification provides an electronic device, the electronic device including: a memory, and one or more processors communicatively connected to the memory; the memory stores instructions executable by the one or more processors, the instructions being executed by the one or more processors to cause the one or more processors to perform the method as described in any of the foregoing embodiments.

[0008] One embodiment of this specification provides a computer storage medium storing a computer program that, when executed by a processor, implements the method described in any of the foregoing embodiments.

[0009] The various embodiments provided in this specification involve fusing the background intrinsic image of a background image and the foreground intrinsic image of a foreground image, respectively, and then rendering the fused intrinsic image using the background lighting information of the background image to obtain a composite image. Since the final composite image uniformly uses the lighting information of the background image, the lighting and shadows between the foreground and background can have good consistency, resulting in a more visually harmonious composite image and improving its visual effect. Attached Figure Description

[0010] Figure 1 This is a schematic flowchart of an image synthesis method provided for one embodiment of this specification.

[0011] Figure 2 This is a schematic diagram of an image synthesis process provided for one embodiment of this specification.

[0012] Figure 3 This is a schematic diagram of an image synthesis apparatus provided for one embodiment of this specification.

[0013] Figure 4 A schematic diagram of an electronic device provided for one embodiment of this specification. Detailed Implementation

[0014] Image compositing refers to combining two or more images into a single image. Specifically, one image can be designated as the foreground image, and another as the background image. In some techniques, the foreground image can be simply pasted onto the background image. However, the resulting composite image often lacks visual harmony. Specifically, for example, there may be visual disharmony in the colors, lighting, etc., between the foreground and background images. Usually, image editing software can be used to process the composite image and improve its visual effect. However, in some cases, even after manual retouching, some disharmony may still exist in the final image. Furthermore, manual retouching of composite images is time-consuming and labor-intensive.

[0015] Therefore, it is necessary to provide an image synthesis method that can obtain multiple foreground intrinsic maps, multiple background intrinsic maps, and background illumination information by performing intrinsic decomposition on the foreground and background images. Then, the multiple foreground and background intrinsic maps can be fused together and combined with the background illumination information to obtain a synthesized image. Since the illumination information of the background image is uniformly used in the final synthesized image, the lighting and shadow between the foreground and background can have good consistency in the synthesized image, thus solving the technical problem of visual disharmony between the foreground and background in the synthesized image.

[0016] One embodiment of this specification provides an image compositing method. The image compositing method can be applied to a computer device. In some embodiments, the computer device may run a computer software application that provides a client for user interaction. This allows the user to control the computer device to perform the image compositing method by operating the client. In some embodiments, the computer device may be a desktop computer, tablet computer, laptop computer, etc. Please refer to [link to documentation]. Figure 1 and Figure 2 The image synthesis method may include the following steps.

[0017] Step S110: Decompose the background image to obtain multiple background intrinsic images and background illumination information; each of the background intrinsic images corresponds to an intrinsic image category.

[0018] The background image can be a pre-specified image used as a background. Typically, a foreground image is incorporated into this background image to obtain a composite image. The background image can also serve as a positional reference, specifying the relative position of the foreground image with respect to the background image.

[0019] A foreground image can refer to an image that needs to be added to a background image. It can be understood that a foreground image represents a target object. This target object can be a real-world object, or it can be an abstract graphic. The user may need to add the target object to the background image. Specifically, for example, if the background image is a photograph of seawater and a beach, and a sailboat needs to be added to the seawater, the sailboat can also be understood as the target object, and the image representing the sailboat can be the foreground image.

[0020] During the imaging process, changes in environmental factors can affect the image representation to varying degrees. These environmental factors can include: light intensity, incident light angle, shadow occlusion, and changes in object pose. These environmental factors significantly impact image synthesis. In this embodiment, by performing intrinsic decomposition on the background image to obtain multiple background intrinsic maps, the influence of environmental factors on image synthesis is reduced. Specifically, images possess essential characteristics that do not change with lighting-related factors. For example, these essential characteristics may be inherent properties of the object represented by the image. For instance, the essential characteristics of an object may include color, texture, and material.

[0021] In this embodiment, the background intrinsic map can be used to represent the essential features of the object represented by the background image. This allows the background intrinsic map to represent features of the object represented by the background image that are independent of lighting information. Each background intrinsic map can correspond to an intrinsic image category. Each intrinsic image category can be used to represent at least one essential feature of the object. Specifically, for example, the intrinsic image category may include: distance depth, albedo, normal, etc. According to the aforementioned intrinsic image categories, the background image can be decomposed into corresponding background intrinsic maps. Specifically, for example, the background intrinsic map may include a background distance depth map corresponding to distance depth, a background albedo map corresponding to albedo, and a background normal map corresponding to normal.

[0022] Background lighting information can be used to represent lighting-related factors in a background image. Specifically, for example, background lighting information can represent light intensity, light incidence angle, etc. In some implementations, background lighting information can be stored in the form of voxelization of the image. This can be understood as voxelizing the objects represented by the background image and storing the lighting information for each voxel. Of course, other methods can also be used to generate background lighting information; specifically, for example, background lighting information can be generated using spectral feature extraction.

[0023] In this embodiment, reverse rendering can be used to decompose the background image to obtain the background intrinsic image and background illumination information. Specifically, for example, an image decomposition model can be constructed using neural network algorithms. After inputting the background image into the image decomposition model, the background intrinsic image and background illumination information can be obtained.

[0024] Step S120: Construct multiple foreground intrinsic images for the foreground image according to the intrinsic image category.

[0025] In this embodiment, a foreground intrinsic map of the foreground image can be constructed corresponding to the intrinsic image category of the background image. The intrinsic image category of the foreground image is consistent with the intrinsic image category of the background image, so that each background intrinsic map can have a corresponding foreground intrinsic map. For example, the background intrinsic map can be a background depth map, and the corresponding foreground intrinsic map can be a foreground depth map.

[0026] A similar approach to decomposing a background image into a background intrinsic map can be used to decompose the foreground image into a foreground intrinsic map. This allows the foreground intrinsic map to reflect the lighting-independent characteristics of the object represented by the foreground image. Of course, those skilled in the art will understand that the methods for decomposing the foreground and background images can differ. In some cases, the foreground image can be a 3D image. After determining the pose and viewpoint of the 3D image within the background image, multiple foreground intrinsic maps corresponding to the corresponding intrinsic image categories can be rendered.

[0027] Step S130: Fuse the background intrinsic image and the foreground intrinsic image of the same intrinsic image category to obtain multiple fused intrinsic images.

[0028] In this embodiment, background and foreground intrinsic images of the same intrinsic image category are fused, so that the fused intrinsic image can better reflect the essential features of the foreground and background images in a certain intrinsic image category. Furthermore, by fusing the background and foreground intrinsic images, the combination of a single intrinsic image category is achieved under the influence of illumination, resulting in a more harmonious blend between the foreground and background images in that intrinsic image category.

[0029] In some implementations, the foreground intrinsic image may need to undergo a transformation before being integrated into the background intrinsic image. Specifically, this transformation may include scaling, angle adjustment, etc. Different foreground intrinsic images can undergo the same transformation, thus reducing the likelihood of mismatches in internal patterns, colors, or shapes in the final synthesized image.

[0030] In some implementations, the method of integrating the foreground intrinsic image into a background intrinsic image of the corresponding intrinsic image category may include: constructing a new layer on top of the background intrinsic image and placing it within the foreground intrinsic image. In some cases, the foreground intrinsic image can be placed into a layer before transformation. Alternatively, the foreground intrinsic image can be transformed first and then placed into a layer. In some cases, the transformed foreground intrinsic image can be directly placed into the background intrinsic image layer and used to replace the target area.

[0031] Step S140: Synthesize an image based on the plurality of fused intrinsic images and the background illumination information.

[0032] In this embodiment, after obtaining multiple fused intrinsic maps independent of lighting information, the fused intrinsic maps can be rendered using the background lighting information of the background image to obtain a composite image. Since only background lighting information is used to render the fused intrinsic maps, a more harmonious visual effect can be achieved between the foreground and background in the composite image. This significantly reduces the problem of foreground and background inconsistency in the composite image caused by differences in lighting between the foreground and background images.

[0033] The embodiments described in this specification involve fusing the background intrinsic image of the background image and the foreground intrinsic image of the foreground image, and then rendering the fused intrinsic image using the background lighting information of the background image to obtain a composite image. This results in a more visually harmonious relationship between the foreground and background in the composite image, enhancing its visual effect.

[0034] In some implementations, the specific way to fuse background intrinsic images and foreground intrinsic images of the same intrinsic image category to obtain a fused intrinsic image is to fuse foreground intrinsic images of the same intrinsic image category into the same target region of the background intrinsic image to obtain a fused intrinsic image.

[0035] In some cases, if the relative positions of the foreground and background intrinsic maps differ among different fused intrinsic maps, it may lead to misalignment or rendering failure in the final composite image due to issues such as color, lighting, and the shape and color of objects. It is necessary to pre-define the target area of ​​the foreground intrinsic map relative to the background intrinsic map, and to fuse different foreground intrinsic maps with the background intrinsic map within this target area to ensure a more accurate final composite image.

[0036] In this embodiment, the target region can be determined by specifying reference pixels. Specifically, a background reference point is specified in the background image, and a foreground reference point is specified in the foreground image, with their relative positions clearly defined. When integrating the foreground intrinsic image into the background intrinsic image, the relative position of the foreground intrinsic image with respect to the background intrinsic image can be determined based on the relative positions of the foreground reference point and the background reference point. Of course, the number of foreground reference points and background reference points can each be at least one. Specifically, for example, the number of foreground reference points and background reference points can be two, three, etc.

[0037] In some embodiments, before fusing background intrinsic images and foreground intrinsic images corresponding to the same intrinsic image category to obtain multiple fused intrinsic images, the image synthesis method may further include: performing the same transformation operation on each foreground intrinsic image; wherein the transformation operation includes: scaling and / or position transformation.

[0038] In some cases, if the transformation operations performed on the foreground intrinsic maps differ between different fused intrinsic maps, inconsistencies may arise, reducing the quality of the synthesized image. Therefore, applying the same transformation operation to foreground intrinsic maps of different intrinsic image categories can, to some extent, improve the quality of the synthesized image.

[0039] In this embodiment, the transformation operation may include, but is not limited to, scaling and / or position transformation. This allows the foreground intrinsic image to be adjusted to a desired state, thereby integrating the foreground intrinsic image into the background intrinsic image. Specifically, scaling may refer to changing the size of the foreground intrinsic image to fit the volume ratio of the foreground image within the background image. Position transformation may refer to changing the position, angle, etc., of the foreground intrinsic image to fit the foreground image within the target region of the background image.

[0040] In some implementations, the intrinsic image category includes a distance-depth map category, the background intrinsic image includes a background distance-depth map, and the foreground intrinsic image includes a foreground distance-depth map; the fused intrinsic image includes a fused distance-depth map; fusing foreground intrinsic images corresponding to the same intrinsic image category to the same target region of the background intrinsic image to obtain the fused intrinsic image may include: extracting a distance-depth sub-image of the target region in the background distance-depth map; determining a reference depth corresponding to the foreground distance-depth map based on the distance-depth sub-image; and fusing the foreground distance-depth map to the target region of the background distance-depth map based on the reference depth to obtain the fused distance-depth map.

[0041] Distance depth can be used to indicate the distance between a pixel representing a local object in an image and the photosensitive sensor that captured the image. Specifically, for example, in an image of a cube with a surface facing a camera lens, the distance between the pixels forming that surface and the camera lens can be represented by distance depth. In some implementations, grayscale values ​​can also be used to represent the distance between a pixel and the photosensitive sensor. Distance depth map types can include a class of images generated based on distance depth. In this implementation, a background distance depth map of the background image can be constructed by calculating the distance from each pixel representing a local object in the background image to the photosensitive sensor. A foreground distance depth map of the foreground image can be constructed by calculating the distance from each pixel representing a local object in the foreground image to the photosensitive sensor.

[0042] In this embodiment, the image generated based on distance and depth is taken as an intrinsic image category. By fusing the foreground distance and depth map and the background distance and depth map, a fused depth map is obtained, so that the foreground image and the background image can be well matched in terms of distance and depth in the final synthesized image.

[0043] In this embodiment, a distance-depth sub-map of the target region can be extracted from the background distance-depth map. This sub-map can be used as a reference to determine a baseline depth for the foreground distance-depth map, which can then be modified based on this baseline depth. Specifically, for example, assuming the target depth of a pixel in the foreground distance-depth map is z", the current distance depth of each pixel is z", and the minimum depth in the foreground distance-depth map is z' min Assume the reference depth determined based on the distance-depth submap of the target region is z. target Formula 1 can exist as follows.

[0044] Formula 1: z″=∣z―z min ±z target |

[0045] Among them, zz min This can represent the relative distance depth of pixels in the foreground image. Then, by summing the relative distance depth with the reference depth, the target depth z of the pixels in the foreground image can be obtained.

[0046] In some embodiments, the maximum distance depth of pixels in the target region can be used as the reference depth, or the average distance depth of pixels in the target region can be used as the reference depth. Of course, those skilled in the art can make other modifications based on the technical teachings of the embodiments described in this specification, which will not be elaborated further.

[0047] After determining the reference depth, the foreground distance depth map can be adjusted based on the reference depth. This achieves a better match between the foreground and background distance depth maps, which can improve the visual effect of the synthesized image to some extent. Specifically, for example, when the reference depth is used as the maximum distance depth value of pixels in the target region, the reference depth can be used as the maximum distance depth value in the foreground distance depth map, and the foreground distance depth map can be adjusted accordingly. In this case, Formula 1 can be adjusted as follows.

[0048] z″=z target ―(z―z min ).

[0049] In some cases, where the reference depth is the minimum distance depth of pixels in the target region, the reference depth can be used as the minimum distance depth in the foreground distance depth map, and the foreground distance depth map can be adjusted accordingly. In this case, Formula 1 can be used.

[0050] z″=z target +(z―z min ).

[0051] Regarding how to adjust the foreground distance depth map, those skilled in the art can make other changes based on the aforementioned implementation methods, which will not be elaborated further.

[0052] In some embodiments, before fusing the foreground distance depth map to the target region of the background distance depth map based on the reference depth to obtain a fused depth map, the image synthesis method further includes performing a transformation operation on the foreground distance depth map. The transformation operation may be scaling.

[0053] We can assume the scaling factor is s, the scaled depth of a pixel in the foreground distance-depth map is z', the current distance-depth of each pixel is z, and the minimum depth in the foreground distance-depth map is z'. min Formula 2 can exist as follows.

[0054] Formula 2:

[0055] Formula 2 adjusts the distance and depth of pixels in the foreground distance and depth map, ensuring accuracy of pixel distance and depth even when scaling the foreground image. Based on the adjusted foreground distance and depth map according to Formula 2 and the reference depth, the adjusted foreground distance and depth map can be fused to the target region of the background distance and depth map to obtain a fused depth map. This avoids distortion of the scaled foreground image caused by inaccurate pixel distance and depth.

[0056] In some implementations, the intrinsic image category includes an albedo map category; the background intrinsic image includes a background albedo map; the foreground intrinsic image includes a foreground albedo map; the fused intrinsic image includes a fused albedo map; fusing foreground intrinsic images corresponding to the same intrinsic image category to the same target region of the background intrinsic image to obtain the fused intrinsic image may include: fusing the foreground albedo map to the target region of the background albedo map to obtain the fused albedo map.

[0057] Albedo can be used to represent the ratio of total reflected radiant flux to incident radiant flux for a given surface. In this embodiment, the albedo of each pixel can be used to represent the ratio of total reflected radiant flux to incident radiant flux for the surface of the object represented by that pixel. Albedo map categories can include images generated based on the albedo of pixels in an image.

[0058] In this embodiment, different materials, textures, and shapes have a certain impact on albedo. By fusing the foreground albedo map to the background albedo map, the resulting fused albedo map can better represent the overall albedo of the synthesized image. Thus, when the fused albedo map is combined with background lighting information for rendering, the synthesized image will have a more harmonious visual effect overall.

[0059] In some embodiments, the intrinsic image category includes a normal map category; the background intrinsic map includes a background normal map; the foreground intrinsic map includes a foreground normal map; the fused intrinsic map includes a fused normal map; fusing foreground intrinsic maps corresponding to the same intrinsic image category to the same target region of the background intrinsic map to obtain a fused intrinsic map includes: fusing the foreground normal map to the target region of the background normal map to obtain the fused normal map.

[0060] The normal map category can include normal maps of images. A normal map can be used to represent a normal line drawn at every point on the surface of the original object, with the direction of the normal line indicated by RGB color channels.

[0061] In this embodiment, a background normal map of the background image and a foreground normal map of the foreground image can be constructed separately. By integrating the foreground normal map into the target area of ​​the background normal map, the merged normal maps can be used as a whole in the rendering of the synthesized image, resulting in a synthesized image with better scene consistency.

[0062] In some implementations, the foreground image includes a three-dimensional foreground image; decomposing the foreground image into multiple foreground intrinsic maps corresponding to the intrinsic image category includes: determining a target viewing angle of the three-dimensional foreground image relative to the background image; wherein the target viewing angle represents the viewing angle from which the three-dimensional foreground image can be observed after placing the object represented by the three-dimensional foreground image in the space represented by the background image; and rendering the three-dimensional foreground image with the target viewing angle corresponding to the intrinsic image category to generate multiple foreground intrinsic maps.

[0063] In some cases, the foreground image can also be a three-dimensional image. Specifically, the foreground image can be a three-dimensional modeled object. In this embodiment, the target viewpoint of the three-dimensional foreground image relative to the background image can be determined, and then the three-dimensional foreground image can be rendered according to the target viewpoint to obtain a foreground intrinsic map.

[0064] In some implementations, the viewpoint of the background image can be used as the target viewpoint of the 3D foreground image. In some cases, the target spatial position of the 3D foreground image relative to the background image can be determined first, and then the 3D foreground image can be rendered according to the target viewpoint to obtain a foreground intrinsic image. Specifically, for example, the 3D foreground image may include a sphere and a cylinder modeled in solid form. The target spatial positions of the sphere and the cylinder relative to the background image can be determined separately, and then the foreground intrinsic image can be obtained by rendering the sphere and the cylinder together according to the target viewpoint.

[0065] Please see Figure 3 One embodiment of this specification also provides an image synthesis apparatus, comprising: a background decomposition unit for decomposing a background image to obtain multiple background intrinsic images and background illumination information; each of the background intrinsic images corresponds to an intrinsic image category; a construction unit for decomposing a foreground image to obtain multiple foreground intrinsic images corresponding to the intrinsic image categories; a fusion unit for fusing the background intrinsic images and foreground intrinsic images corresponding to the same intrinsic image category to obtain multiple fused intrinsic images; and a synthesis unit for synthesizing an image based on the multiple fused intrinsic images and the background illumination information.

[0066] In some implementations, the fusion unit is used to fuse foreground intrinsic images corresponding to the same intrinsic image category into the same target region of the background intrinsic image to obtain a fused intrinsic image.

[0067] In some embodiments, the image synthesis apparatus further includes a transformation unit for performing the same transformation operation on each foreground intrinsic image; wherein the transformation operation includes scaling and / or position transformation.

[0068] In some embodiments, the intrinsic image category includes a distance-depth map category; the background intrinsic image includes a background distance-depth map; the foreground intrinsic image includes a foreground distance-depth map; the fused intrinsic image includes a fused distance-depth map; the fusion unit includes: an extraction module for extracting a distance-depth sub-image of the target region in the background distance-depth map; a reference depth determination module for determining a reference depth corresponding to the foreground distance-depth map based on the distance-depth sub-image; and a depth map fusion module for fusing the foreground distance-depth map to the target region of the background distance-depth map based on the reference depth to obtain the fused distance-depth map.

[0069] In some embodiments, the intrinsic image category includes an albedo map category; the background intrinsic image includes a background albedo map; the foreground intrinsic image includes a foreground albedo map; the fused intrinsic image includes a fused albedo map; and the fusion unit includes an albedo map fusion module, used to fuse the foreground albedo map into the target region of the background albedo map to obtain the fused albedo map.

[0070] In some embodiments, the intrinsic image category includes a normal map category; the background intrinsic map includes a background normal map; the foreground intrinsic map includes a foreground normal map; the fused intrinsic map includes a fused normal map; and the fusion unit includes a normal map fusion module, used to fuse the foreground normal map to the target region of the background normal map to obtain the fused normal map.

[0071] In some embodiments, the foreground image includes a three-dimensional foreground image; the construction unit includes: a determining module, configured to determine a target viewing angle of the three-dimensional foreground image relative to the background image; wherein the target viewing angle represents the viewing angle from which the three-dimensional foreground image can be observed after placing the object represented by the three-dimensional foreground image in the space represented by the background image; and a rendering module, configured to render the three-dimensional foreground image from the target viewing angle, corresponding to the intrinsic image category, to generate multiple foreground intrinsic images.

[0072] The specific functions and effects of the image compositing device can be explained by referring to other embodiments in this specification, and will not be repeated here. Each module in the image compositing device can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0073] Please see Figure 4 This specification also provides an electronic device, comprising: a memory, and one or more processors communicatively connected to the memory; the memory storing instructions executable by the one or more processors, the instructions being executed by the one or more processors to cause the one or more processors to implement the image compositing method as described in any of the above embodiments.

[0074] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image synthesis method of any of the above embodiments.

[0075] This specification also provides a computer program product containing instructions that, when executed by a computer, cause the computer to perform the image synthesis method described in any of the above embodiments.

[0076] It is understood that the specific examples in this document are only intended to help those skilled in the art better understand the embodiments described herein, and are not intended to limit the scope of the invention.

[0077] It is understood that in the various embodiments described in this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments described in this specification.

[0078] It is understood that the various implementation methods described in this specification can be implemented individually or in combination, and the implementation methods in this specification are not limited in this respect.

[0079] Unless otherwise stated, all technical and scientific terms used in the embodiments of this specification have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this specification. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0080] It is understood that the processor in the embodiments of this specification can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this specification can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0081] It is understood that the memory in the embodiments of this specification may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0082] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.

[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the aforementioned method implementations, and will not be repeated here.

[0084] In the several embodiments provided in this specification, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0085] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0086] In addition, the functional units in the various embodiments of this specification can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0087] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of this specification, in essence, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] The above description is merely a specific embodiment of this specification, but the scope of protection of this invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this specification should be included within the scope of protection of this specification. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. An image synthesis method, characterized in that, include: Decompose the background image to obtain multiple background intrinsic images and background illumination information, and each background intrinsic image corresponds to an intrinsic image category; The intrinsic image categories include distance-depth map category, albedo map category, and normal map category; According to the intrinsic image category, the foreground image is decomposed to obtain multiple foreground intrinsic images; The background intrinsic image and the foreground intrinsic image corresponding to the same intrinsic image category are fused to obtain multiple fused intrinsic images; An image is synthesized based on the multiple fused intrinsic maps and the background illumination information.

2. The method according to claim 1, characterized in that, The specific methods for fusing background and foreground intrinsic images that correspond to the same intrinsic image category to obtain a fused intrinsic image include: Foreground intrinsic images corresponding to the same intrinsic image category are fused into the same target region of the background intrinsic image to obtain a fused intrinsic image.

3. The method according to claim 2, characterized in that, Before fusing background and foreground intrinsic images belonging to the same intrinsic image category to obtain multiple fused intrinsic images, the process also includes: Perform the same transformation operation on each foreground eigenmap; wherein the transformation operation includes scaling and / or position transformation.

4. The method according to claim 2, characterized in that, The background intrinsic map includes a background distance-depth map, the foreground intrinsic map includes a foreground distance-depth map, and the fused intrinsic map includes a fused distance-depth map. The step of fusing foreground intrinsic images corresponding to the same intrinsic image category into the same target region of the background intrinsic image to obtain a fused intrinsic image includes: Extract the distance-depth sub-map of the target region from the background distance-depth map; Determine the reference depth corresponding to the foreground distance-depth map based on the distance-depth submap; Based on the reference depth, the foreground distance depth map is fused to the target region of the background distance depth map to obtain the fused distance depth map.

5. The method according to claim 2, characterized in that, The background intrinsic map includes a background albedo map; the foreground intrinsic map includes a foreground albedo map. The fused intrinsic map includes a fused albedo map; Foreground intrinsic maps corresponding to the same intrinsic image category are fused into the same target region of the background intrinsic map to obtain a fused intrinsic map, including: The foreground albedo map is fused to the target region of the background albedo map to obtain the fused albedo map.

6. The method according to claim 2, characterized in that, The background intrinsic map includes a background normal map; the foreground intrinsic map includes a foreground normal map; the blended intrinsic map includes a blended normal map; Foreground intrinsic maps corresponding to the same intrinsic image category are fused into the same target region of the background intrinsic map to obtain a fused intrinsic map, including: The foreground normal map is fused to the target region of the background normal map to obtain the fused normal map.

7. The method according to claim 1, characterized in that, The foreground image includes a three-dimensional foreground image; Corresponding to the intrinsic image category, the foreground image is decomposed to obtain multiple foreground intrinsic images, including: Determine the target viewing angle of the three-dimensional foreground image relative to the background image; wherein, the target viewing angle represents the viewing angle from which the three-dimensional foreground image can be observed after the object represented by the three-dimensional foreground image is placed in the space represented by the background image; From the target perspective, the three-dimensional foreground image is rendered to generate multiple foreground intrinsic images corresponding to the intrinsic image category.

8. An image synthesis apparatus, characterized in that, include: Background decomposition unit is used to decompose the background image to obtain multiple background intrinsic maps and background illumination information; Each of the aforementioned background intrinsic images corresponds to an intrinsic image category; The intrinsic image categories include distance-depth map category, albedo map category, and normal map category; The construction unit is used to decompose the foreground image into multiple foreground intrinsic images corresponding to the intrinsic image category; The fusion unit is used to fuse the background intrinsic image and the foreground intrinsic image that correspond to the same intrinsic image category, respectively, to obtain multiple fused intrinsic images; A synthesis unit is used to synthesize an image based on the plurality of fused intrinsic maps and the background illumination information.

9. An electronic device, characterized in that, The electronic device includes: A memory, and one or more processors communicatively connected to the memory; The memory stores instructions that can be executed by the one or more processors to cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.