Image fusion method and device, electronic equipment and readable storage medium
By generating a foreground portrait with a transparent channel and a low-resolution background image, and combining blurring with random and regular sampling strategies, the problem of balancing image fusion quality and real-time device performance in existing technologies is solved, achieving high-quality, real-time image fusion effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU HUYA TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing image fusion technologies cannot simultaneously balance image fusion quality and real-time device performance, and have poor adaptive processing capabilities for different background content.
A foreground portrait with an alpha channel and a low-resolution background image to be fused are generated. The images are then blurred in stages using a preset random sampling strategy and a regular sampling strategy to generate a fine-grained blurred background image. Finally, the images are fused.
It achieves a balance between high visual fusion quality and the real-time operation requirements of devices, and improves the adaptability to diverse background content and the feasibility of edge deployment.
Smart Images

Figure CN121961869A_ABST
Abstract
Description
Image fusion methods, apparatus, electronic devices and readable storage media Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to an image fusion method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the rapid development of computer vision and image processing technologies, technologies such as video image processing and virtual background compositing have been widely applied in scenarios such as video conferencing, virtual live streaming, and real-time shooting systems. These applications place strict requirements on the lighting matching and fusion effects between foreground portraits and background images, aiming to achieve a visually natural and harmonious lighting performance to enhance the realism and immersive experience of the images.
[0003] Existing image background fusion techniques can be mainly divided into two categories. The first category is image reconstruction and illumination modeling methods based on deep learning. Although these methods can achieve good fusion results in static or offline processing tasks, they face challenges in practical applications, such as high computational overhead in neural network inference processes, difficulty in meeting the needs of real-time processing, high requirements for the quality and consistency of input images, and insufficient robustness. The second category is traditional image fusion and color matching methods. Although these methods have relatively low computational complexity, their CPU-based implementation process is difficult to run efficiently on resource-constrained devices in practical applications, and their performance deteriorates in dynamic images or complex backgrounds.
[0004] Therefore, existing solutions cannot simultaneously balance image fusion quality and real-time device performance, and have poor adaptive processing capabilities for different background content. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide an image fusion method, apparatus, electronic device and readable storage medium that can simultaneously take into account the image fusion quality and the real-time performance of the device during the image fusion process, and has strong adaptive processing capability for different background content.
[0006] To achieve the above objectives, the technical solution adopted in the embodiments of the present invention is as follows: In a first aspect, the present invention provides an image fusion method, the method comprising: generating a foreground portrait to be fused and a low-resolution background image; wherein the foreground portrait has a transparency channel for supporting the fusion of the foreground portrait with different background images; blurring the low-resolution background image based on a preset random sampling strategy to obtain a coarse-grained blurred background image; blurring the coarse-grained blurred background image using a preset regular sampling strategy to obtain a fine-grained blurred background image; and performing image fusion on the foreground portrait to be fused and the fine-grained blurred background image.
[0007] Secondly, the present invention provides an image fusion apparatus, comprising: a generation module for generating a foreground portrait and a low-resolution background image to be fused; wherein the foreground portrait has a transparency channel for supporting fusion with different background images; a blurring module for blurring the low-resolution background image based on a preset random sampling strategy to obtain a coarse-grained blurred background image; the blurring module is further configured to blur the coarse-grained blurred background image using a preset rule sampling strategy to obtain a fine-grained blurred background image; and a fusion module for performing image fusion of the foreground portrait to be fused and the fine-grained blurred background image.
[0008] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing machine-executable instructions executable by the processor, the processor executing the machine-executable instructions to implement the image fusion method described in any of the foregoing embodiments.
[0009] Fourthly, the present invention provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image fusion method as described in any of the foregoing embodiments.
[0010] The image fusion method, apparatus, electronic device, and readable storage medium provided in this invention first generate a foreground portrait to be fused and a low-resolution background image with an alpha channel. This ensures that the foreground portrait can flexibly adapt to various background content, improving the system's adaptive processing capability in different scenarios. Then, based on a preset random sampling strategy, the low-resolution background is initially blurred to form a coarse-grained blurred background image. This step reduces computational complexity while preserving the main structural information of the background. Next, a regular sampling strategy is used to further optimize the coarse-grained blurring result, generating a fine-grained blurred background image with more natural detail transitions, enhancing the background's sense of layering and visual comfort, laying the foundation for high-quality fusion. Finally, the foreground portrait and the fine-grained blurred background image are fused. While ensuring edge clarity and accurate synthesis of the alpha channel, this achieves both high visual fusion quality and meets the real-time operation requirements of the device. The overall process, through a phased and differentiated blurring strategy, collaboratively optimizes image quality and processing efficiency, significantly improving adaptability to diverse background content and the feasibility of edge deployment.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 is a schematic flowchart of the image fusion method provided in an embodiment of the present invention; Figure 2 shows the original background image and the foreground portrait to be fused provided in an embodiment of the present invention; Figure 3 shows the coarse-grained blurred background image provided in an embodiment of the present invention; Figure 4 shows the fine-grained blurred background image provided in an embodiment of the present invention; Figure 5 shows the foreground portrait after soft light fusion provided in an embodiment of the present invention; Figure 6 shows the fusion effect diagram of the virtual scene background and the foreground portrait provided in an embodiment of the present invention; Figure 7 is an example diagram of the image fusion method provided in an embodiment of the present invention; Figure 8 is a functional block diagram of the image fusion device provided in an embodiment of the present invention; Figure 9 is a structural block diagram of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0015] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0016] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0017] Considering that existing image fusion schemes cannot simultaneously achieve both image fusion quality and real-time device performance, and have poor adaptive processing capabilities for different background content, this invention provides an image fusion method. Please refer to Figure 1, which is a schematic flowchart of the image fusion method provided by this invention, including steps S101 to S104, explained as follows: S101: Generate a foreground portrait to be fused and a low-resolution background image; In this invention, the foreground portrait to be fused is a digital image containing color and transparency information, used to accurately represent the position, shape, and edge transition effects of the foreground subject (such as a portrait) in the image. It is a key data format for achieving high-quality image fusion. The foreground portrait to be fused has a transparency channel to support the fusion of the foreground portrait with different background images. The low-resolution background image refers to the image obtained by downsampling the original background image, serving as the basic input for subsequent adaptive fusion processing.
[0018] S102: The low-resolution background image is blurred based on a preset random sampling strategy to obtain a coarse-grained blurred background image; In this embodiment of the invention, the random sampling strategy is to randomly select multiple sampling points in the neighborhood of each pixel, and the blurring effect based on this strategy can be natural.
[0019] S103: The coarse-grained blurred background image is blurred using a preset rule sampling strategy to obtain a fine-grained blurred background image. In this embodiment of the invention, the rule sampling strategy is to use a biomimetic-inspired golden spiral distribution for multiple sampling, with the sampling points arranged in a divergent pattern along the golden angle. Blurring based on this strategy can obtain a smooth blurring effect that is more in line with the human visual system.
[0020] S104: Perform image fusion on the foreground portrait and fine-grained blurred background image to be fused.
[0021] Unlike existing technologies, this invention generates a foreground portrait with a transparent channel and a low-resolution background image to be fused. First, it ensures the foreground portrait can flexibly adapt to various background content, improving the system's adaptive processing capabilities in different scenarios. Then, based on a preset random sampling strategy, the low-resolution background undergoes initial blurring to form a coarse-grained blurred background image. This step reduces computational complexity while preserving the main structural information of the background. Next, a regular sampling strategy further optimizes the coarse-grained blurring result, generating a fine-grained blurred background image with more natural detail transitions, enhancing the background's layering and visual comfort, laying the foundation for high-quality fusion. Finally, the foreground portrait and the fine-grained blurred background image are fused. While ensuring edge clarity and accurate synthesis of the transparent channel, this achieves both high visual fusion quality and meets the real-time operation requirements of the device. The overall process, through a phased and differentiated blurring strategy, collaboratively optimizes image quality and processing efficiency, significantly improving adaptability to diverse background content and the feasibility of edge deployment.
[0022] Next, the image fusion process shown in FIG1 will be described in detail with reference to the accompanying drawings.
[0023] In one embodiment of the present invention, the foreground portrait generated in step S101 with a transparency channel is a prerequisite step for achieving high-quality background fusion in this embodiment. Its purpose is to accurately separate the main subject from the original image while preserving its color and transparency information for subsequent adaptive lighting fusion processing on the GPU. Specifically, the implementation method for generating the foreground portrait in step S101 is as follows: Step a1: Obtain the original portrait image; In this embodiment, the source of the original portrait image can be a video frame captured in real-time by a camera, a frame from a pre-recorded video, or a still image. The image format is typically a digital image in RGB or RGBA format, with each pixel containing three color channels: red (R), green (G), and blue (B).
[0024] Step a2: Separate the human body region from the original human image and generate a human portrait segmentation mask; wherein, the grayscale value of the human portrait segmentation mask is used to characterize the degree to which a pixel belongs to the human body region; in this embodiment of the invention, it can be extracted from the original human image by human portrait segmentation technology to accurately separate the foreground human body region and the background region, and output the corresponding human portrait segmentation mask.
[0025] In one implementation, step a2 can be achieved using a lightweight portrait segmentation model based on deep learning. This involves using a lightweight neural network to perform semantic segmentation on the original portrait image, outputting a portrait segmentation mask. Pixels with grayscale values close to 1 represent the "foreground region," while pixels with grayscale values close to 0 represent the "background region." Intermediate transition regions (such as hair strands or semi-transparent clothing) can be preserved with soft edge values between 0 and 1 to achieve more natural edge synthesis.
[0026] In another implementation, step a2 can also be based on green screen keying technology, whereby the user takes a picture with a solid color background (usually green or blue), and the system detects color differences and removes a specified color gamut to generate a corresponding portrait segmentation mask.
[0027] Step a3: Construct the foreground portrait to be fused based on the color channels of the original portrait image and the grayscale values of the portrait segmentation mask.
[0028] Specifically, using the color information of the original image and the transparency information of the segmentation mask, an RGBA four-channel image with transparency information is constructed, which is the foreground portrait to be fused. The mathematical expression is as follows: ForegroundImage(x,y)={(R,G,B, = (I(x,y).rgb, M(x,y)) where: I(x,y).rgb refers to the RGB color value of the original image of the person at position (x,y); M(x,y) refers to the gray value of the image segmentation mask at position (x,y), which is used as the Alpha channel (value range [0.0, 1.0]). That is, the alpha channel; the final ForegroundImage is a four-channel image that can be used for blending operations in subsequent GPU shaders.
[0029] The RGBA texture of the foreground image generated by the above implementation method can be directly uploaded to the GPU memory for efficient reading and calculation by subsequent shader programs; moreover, the foreground extraction process is independent of the target background, and has good versatility and modularity.
[0030] The process of generating a low-resolution background image in step S101 of this embodiment of the invention will be described next.
[0031] In one embodiment of the present invention, the process of generating a low-resolution background image is a key optimization step for improving GPU processing efficiency and reducing computational load. This process generates a background image with lower resolution but retaining the main lighting and color features by downsampling the original background image, which serves as the input for subsequent blurring processing. This significantly reduces the amount of shader computation while maintaining visual quality. Specifically, the implementation process of generating a low-resolution background image is as follows: Step b1: Obtain the original background image; In this embodiment of the present invention, the original background image can be a static image, a sequence of dynamic video frames, or a virtual scene image generated in real time by a graphics engine.
[0032] Step b2: Create a rendering buffer of a preset size; wherein the preset size is smaller than the image size of the original background image; in this embodiment of the invention, a small off-screen rendering buffer can be created in the GPU, for example, scaling the original image from 1920×1080 to 480×270 (i.e., reducing it by a factor of 4), leaving only 1 / 16 of the original area. This rendering buffer will be used to receive the scaled result.
[0033] Step b3: Draw the original background image into the render buffer, and downsample the original background image in the render buffer to obtain a low-resolution background image.
[0034] In this embodiment of the invention, the original background image is directly drawn onto a small rendering buffer. Then, the linear interpolation filtering mechanism in the GPU rasterization process is used to achieve high-quality and efficient image scaling, resulting in a low-resolution background image. This low-resolution background image retains the main structural features and overall illumination distribution of the original background, but significantly reduces high-frequency details and the number of pixels, thereby effectively reducing the computational load of the subsequent blurring stage and improving the overall algorithm execution efficiency.
[0035] It should be understood that the above operation process not only helps to improve the GPU processing speed, but also provides a suitable input scale for subsequent blurring processing, achieving an optimized balance between resource consumption and processing performance while ensuring the quality of visual fusion.
[0036] Based on the low-resolution background image generated in step S101, this embodiment of the invention will then complete multiple blurring processes on the low-resolution background image through steps S102 and S103 to generate a fine-grained blurred background image with more natural detail transitions, thereby enhancing the sense of layering and visual comfort of the background and laying the foundation for high-quality fusion.
[0037] First, considering that in virtual background blending scenarios, the goal of image blending is not only to change the background, but also to make the foreground figures appear to be truly within the background environment, i.e., consistent lighting direction, harmonious brightness, and scattered light at the edges. Therefore, this invention first provides a blurring strategy based on random sampling. This strategy performs efficient and visually smooth blurring of low-resolution background images by simulating random sampling and statistical averaging on the GPU, thereby generating an approximate illumination field map that reflects the trend of ambient light distribution. This results in a natural blurring effect for the background image, simulating realistic optical defocusing effects, and is suitable for simulating natural blurring effects such as depth of field, motion blur, and halos.
[0038] In one embodiment of the present invention, the blurring process based on a preset random sampling strategy in step S102 is as follows: Step c1: Obtain blur parameters; In this embodiment of the present invention, the blur parameters include the number of random samplings and the blur radius. The number of random samplings is represented by N, which refers to randomly selecting N neighboring locations from the input image for sampling. For example, N=5 means randomly sampling 5 times around each pixel. The blur radius is represented by R, which is used to control the maximum diffusion distance of the blur (for example, it can be set to 5%~10% of the image width and height), which facilitates dynamic adjustment of the blur intensity. If R is very small, such as 0.001, then the randomly offset sampling point will be very close to the original pixel, and the blur effect is very light. If the radius is large, such as 0.01~0.05, then the offset will be farther away from the original pixel, which is equivalent to random sampling in a larger area, and the blur is more obvious.
[0039] Step c2: Perform multiple random samplings around each pixel in the low-resolution background image according to the random sampling number, and determine the sampling pixel position obtained in each random sampling. In this embodiment of the invention, the number of sampling pixel positions is the same as the number of random samplings, and the sampling pixel positions are located within the blur radius centered on the pixel. To determine which pixel positions to sample, the following implementation method can be used: First step: Generate random numbers based on the texture coordinates of the pixel; Second step: Generate random offsets based on the random numbers; Third step: Map the random offsets to the range of the blur radius to obtain the screen coordinate offsets relative to the pixel; Fourth step: Superimpose the screen coordinate offsets onto the pixel positions to obtain the sampling pixel positions.
[0040] In this embodiment of the invention, the texture coordinates (u, v) of a pixel can be used as input, and a random number can be generated using a non-linear hash function. This function outputs a random number in the range [0, 1] for each (u, v). In each random sampling, this random number is used to generate random offsets dX and dY. This random offset, which is a normalized offset in the range [-1, 1], needs to be multiplied by the blur radius to become the screen coordinate offset. Finally, the screen coordinate offset is superimposed on the pixel position to obtain the sampled pixel position.
[0041] Step c3: Accumulate the color values of each sampled pixel position corresponding to each pixel and calculate the average value to obtain the new color value for each pixel. After determining the sampled pixel position corresponding to a certain pixel through the above preliminary steps, the color value at that position can be directly read. Then, the color values at each position are accumulated and averaged to obtain the new color value for that pixel. This process is repeated to obtain the new color value for each pixel, as shown in the following formula:
[0042] in, This represents the final color value corresponding to the pixel at position x; i represents the index of the sampling number; N represents the number of samples. This represents the screen coordinate offset obtained from the i-th sample; Step c4: Extract color values; Step c5: Colorize pixels based on the new color values corresponding to each pixel to obtain a coarse-grained blurred background image.
[0043] It is understood that steps c1 to c4 above can be completed in the GPU shader, ultimately yielding a coarse-grained blurred background image. To visually demonstrate the blurring effect of the first stage in this embodiment of the invention, please refer to Figures 2 and 3. Figure 2 shows the original background image and foreground image provided in this embodiment of the invention, and Figure 3 shows the coarse-grained blurred background image provided in this embodiment of the invention.
[0044] Next, for the image after the above blurring process, the second stage of blurring process in this embodiment of the invention can be performed. The purpose is to make the background lighting information softer and more uniform, and to have a natural feel similar to lens defocus or light spiral diffusion, thereby improving the naturalness of the blending and avoiding screen flickering or ripple interference caused by repeated sampling modes. See step S103.
[0045] In one embodiment of the present invention, the coarse-grained blurred background image obtained in step S102 can be blurred according to the rule sampling strategy introduced in this embodiment to obtain a fine-grained blurred background image. Specifically, the implementation process of step S103 is as follows: Step d1 to d4: Obtain the preset number of iterations and determine the sampling direction and blur radius used in each iteration; in this embodiment of the present invention, the number of iterations M can be flexibly set by the user, for example, M=7, that is, 7 iterations can be performed, and each iteration uses a different sampling direction and blur radius. The angle between adjacent sampling directions is the golden angle. , The sampling direction is indicated by the included angle. If we express this as follows, then the sampling direction used in the j-th iteration can be represented as... , Let the fuzzy radius be denoted by r, then the fuzzy radius used in the j-th iteration can be expressed as: .
[0046] Alternatively, the fuzzy radius can be determined by the formula We obtain that s is a fixed parameter value that can be flexibly set by the user, or it can be flexibly set by the user according to the rule that the fuzzy radius decreases with the number of iterations. There is no limitation here.
[0047] Step d2: In each iteration, the sampling pixel position is determined based on the sampling direction and the blur radius, using each pixel in the coarse-grained blurred background image as the center. This embodiment of the invention introduces a bidirectional sampling mechanism in determining the sampling pixel position. Specifically, in the positive direction of the sampling direction, a positive offset is determined based on the blur radius and sampling direction, and then this positive offset is superimposed on the pixel position to obtain the first sampling pixel position. In the opposite direction of the sampling direction, a negative offset is determined based on a preset multiple of the blur radius and sampling direction, and then this negative offset is superimposed on the pixel position to obtain the second sampling pixel position. The preset multiple can be flexibly set by the user; for example, the sampling multiple is 0.7. In this way, there are two sampling positions in each iteration, resulting in a total of 2M output sampling pixel positions. For example, assuming M=7, the average of the 14 sampling points is taken to achieve a uniform spiral blur effect.
[0048] It should be noted that the "positive direction" mentioned above refers to the direction of outward radiation along the sampling direction determined by the current golden angle distribution, that is, from the center of the current pixel along the polar coordinate angle. The outward direction; "negative direction" refers to the vector direction opposite to the sampling direction and pointing in the opposite direction to the pixel center, that is, the direction after rotating 180° relative to the current sampling direction.
[0049] For ease of subsequent description, the position of the first sampled pixel in the j-th iteration can be represented as... The position of the second sampled pixel can be represented as Where u is the pixel's position coordinate; t is a preset multiplier.
[0050] Step d3: Extract the color value at the sampling pixel position in each iteration, then sum the color values obtained in each iteration and average them to obtain the final color value of each pixel; In this embodiment of the invention, as can be seen from step d2, each pixel can determine two corresponding sampling points in each iteration. At this time, the color values of the two sampling points can be summed as the sampling color value in this iteration, and then the sampling color values obtained in each iteration are summed to obtain the final color value of the pixel. The whole process can be represented by the following formula:
[0051] in, This represents the final color value of the pixel at position u; Step d4: Colorize the pixels based on the final color value of each pixel to generate a fine-grained blurred background image.
[0052] To visually demonstrate the blurring effect of the second stage described in this embodiment of the invention, please refer to Figure 4, which shows a fine-grained blurred background image provided by this embodiment. It can be seen that by setting offset sampling points with different weights in the forward and reverse sampling directions, the asymmetric blurring characteristics of light scattering in a real optical system can be simulated, enhancing the spatial continuity and visual smoothness of the blurring effect. This bidirectional asymmetric sampling strategy not only increases the effective sampling density but also reduces noise and banded artifacts with the same number of iterations, thereby improving the naturalness of the blur quality while maintaining high performance.
[0053] After the two stages of blurring described above, a fine-grained blurred background image is obtained. This background image enhances the sense of depth and visual comfort, laying the foundation for high-quality fusion. Based on this, step S104 can be performed to fuse the foreground portrait to be fused and the fine-grained blurred background image.
[0054] In one embodiment of the present invention, in step S104, a soft-light blending method can be used to fuse the fine-grained blurred background image and the foreground portrait. This blending method is a color mixing method based on conditional judgment and nonlinear calculation of pixel brightness values. Its output is not a simple interpolation of the input color, but dynamically determines the brightening or darkening behavior of the background according to the brightness of the foreground, thereby generating a more natural and more harmonious blended image, making the portrait appear to be truly in the current background lighting environment, achieving a natural, real-time, and efficient virtual background synthesis effect. For a clear visual demonstration of the above blending effect, please refer to Figure 5, which shows the foreground portrait after soft-light blending provided in an embodiment of the present invention.
[0055] Finally, the merged foreground portrait can be rendered in real time through the GPU graphics pipeline and output to the target device's display screen, video stream, or frame buffer, achieving efficient end-to-end image compositing. For example, by compositing the foreground portrait obtained in Figure 5 and the background image shown in Figure 2, the final result is shown in Figure 6. Please refer to Figure 6, which is a fusion effect diagram of the virtual scene background and foreground portrait provided by an embodiment of the present invention.
[0056] To facilitate a comprehensive understanding of the image fusion method provided in this embodiment of the invention, please refer to Figure 7, which is an example diagram of the image fusion method provided in this embodiment of the invention. Referring to Figure 7, the image fusion method provided in this embodiment of the invention has the following advantages: First, this embodiment of the invention provides multiple blurring processes through random sampling and regular sampling strategies. This effectively simulates the diffuse reflection characteristics of background lighting and introduces a soft light blending strategy, significantly alleviating the "cut-and-paste" effect and abrupt edge problems commonly found in traditional image matting and compositing. This makes the foreground portrait and background more harmonious and natural in terms of lighting direction, brightness matching, tone consistency, and edge transition, greatly enhancing the realism and visual appeal of the image.
[0057] Secondly, the entire image fusion process in this embodiment of the invention can fully utilize the parallel computing power of the GPU, deploying core modules such as blur processing, image fusion, and rendering output entirely within the GPU pipeline, greatly improving processing efficiency. It can stably support real-time operation at high frame rates of 60 FPS and above, meeting the low-latency-sensitive application requirements of live streaming, video conferencing, AR / VR, and other applications. Simultaneously, this solution does not rely on complex deep learning models or high-overhead graphics architectures, featuring a lightweight structure, low resource consumption, and excellent platform adaptability, allowing flexible deployment in various hardware environments such as mobile devices, desktops, and embedded devices. Thanks to background sampling and adaptive blur fusion mechanisms, the system maintains stable fusion quality even under dynamic background switching or complex lighting conditions, effectively suppressing visual artifacts such as jagged edges and inconsistent halo effects.
[0058] Furthermore, the image fusion method provided in this embodiment of the invention has good scalability and can be seamlessly integrated into multiple practical scenarios such as live video streaming, virtual anchors, remote teaching, and short video editing, significantly improving product experience and technological competitiveness.
[0059] Based on the same inventive concept as Figure 1, an implementation of the image fusion device 80 is also provided below in this embodiment of the invention. Please refer to Figure 8, which is a functional block diagram of the image fusion device provided in this embodiment of the invention. The image fusion device 80 includes: a generation module 801, a blurring processing module 802, and a fusion module 803.
[0060] The generation module 801 is used to generate a foreground portrait and a low-resolution background image to be fused; wherein, the foreground portrait to be fused has a transparency channel to support fusion with different background images; the blurring module 802 is used to blur the low-resolution background image based on a preset random sampling strategy to obtain a coarse-grained blurred background image; the blurring module 802 is also used to blur the coarse-grained blurred background image using a preset rule sampling strategy to obtain a fine-grained blurred background image; the fusion module 803 is used to fuse the foreground portrait to be fused and the fine-grained blurred background image.
[0061] It is understandable that the generation module 801, the fuzzing module 802, and the fusion module 803 can work together to execute the various steps in Figure 1 to achieve the corresponding technical effects.
[0062] It should be noted that the image fusion device 80 provided in this embodiment of the invention can be specific hardware on a device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this embodiment of the invention are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiments can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.
[0063] Optionally, the above-mentioned modules can be stored in the memory shown in FIG9 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 90, and can be executed by the processor in FIG9. At the same time, the data, program code, etc. required to execute the above-mentioned modules can be stored in the memory.
[0064] Please refer to Figure 9, which shows a structural block diagram of an electronic device provided in an embodiment of the present invention. The device includes a memory 901, a processor 902, and a communication interface 903. The memory 901, processor 902, and communication interface 903 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0065] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used in Figure 9, but this does not imply that there is only one bus or one type of bus.
[0066] In this embodiment of the invention, the processor 902 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The software modules may be located in the memory 901, and the processor 902 reads the program instructions from the memory 901 and, in conjunction with its hardware, completes the steps of the aforementioned methods.
[0067] In this embodiment of the invention, the memory 901 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as RAM. The memory can also be any other medium capable of carrying or storing desired executable program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in this embodiment of the invention can also be a circuit or any other device capable of implementing a storage function for storing instructions and / or data.
[0068] The memory 901 can be used to store software programs and modules, such as the instructions / modules of the image fusion device 80 provided in this embodiment of the invention. These can be stored in the memory 901 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 90. The processor 902 executes various functional applications and data processing by executing the software programs and modules stored in the memory 901. The communication interface 903 can be used for signaling or data communication with other node devices.
[0069] It is understood that the structure shown in Figure 9 is for illustrative purposes only, and the electronic device 90 may also include more or fewer components than shown in Figure 9, or have a different configuration than shown in Figure 9. The components shown in Figure 9 may be implemented in hardware, software, or a combination thereof.
[0070] Based on the above embodiments, the present invention also provides a storage medium storing a computer program. When the computer program is executed by a computer, it causes the computer to perform the image fusion method provided in the above embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0071] Based on the above embodiments, the present invention also provides a program product, which includes a computer program. The processor can execute the computer program to implement the image fusion method provided in the embodiments of the present invention. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0072] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0073] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs.
[0074] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0075] It should be noted that if the function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An image fusion method, characterized in that, The method includes: generating a foreground portrait and a low-resolution background image to be fused; wherein the foreground portrait has a transparency channel to support the fusion of the foreground portrait with different background images; blurring the low-resolution background image based on a preset random sampling strategy to obtain a coarse-grained blurred background image; blurring the coarse-grained blurred background image using a preset regular sampling strategy to obtain a fine-grained blurred background image; and performing image fusion on the foreground portrait to be fused and the fine-grained blurred background image.
2. The image fusion method according to claim 1, characterized in that, Generating a foreground image to be merged includes: obtaining an original human image; separating the human body region from the original human image to generate a human portrait segmentation mask; wherein the grayscale value of the human portrait segmentation mask is used to characterize the degree to which a pixel belongs to the human body region; and constructing the foreground human portrait to be merged based on the color channels of the original human image and the grayscale value of the human portrait segmentation mask.
3. The image fusion method according to claim 1, characterized in that, Generating a low-resolution background image includes: obtaining an original background image; creating a rendering buffer of a preset size; wherein the preset size is smaller than the image size of the original background image; drawing the original background image into the rendering buffer, and downsampling the original background image in the rendering buffer to obtain the low-resolution background image.
4. The image fusion method according to claim 1, characterized in that, The low-resolution background image is blurred based on a preset random sampling strategy to obtain a coarse-grained blurred background image. This includes: obtaining blur parameters; wherein the blur parameters include a number of random samplings and a blur radius; performing multiple random samplings around each pixel in the low-resolution background image according to the number of random samplings, and determining the sampling pixel position obtained in each random sampling; wherein the sampling pixel position is within the blur radius centered on the pixel; accumulating the color values of each sampling pixel position corresponding to each pixel and calculating the average value as the new color value for each pixel; and performing pixel coloring based on the new color values corresponding to each pixel to obtain the coarse-grained blurred background image.
5. The image fusion method according to claim 4, characterized in that, The process involves performing multiple random samplings around each pixel in the low-resolution background image according to the specified number of random samplings, and determining the position of the sampled pixel obtained in each random sampling. This includes: generating a random number based on the texture coordinates of the pixel; generating a random offset based on the random number; mapping the random offset to the range of the blur radius to obtain a screen coordinate offset relative to the pixel; and superimposing the screen coordinate offset onto the position of the pixel to obtain the sampled pixel position.
6. The image fusion method according to any one of claims 1-5, characterized in that, A fine-grained blurred background image is obtained by blurring the coarse-grained blurred background image using a preset rule sampling strategy, including: obtaining a preset number of iterations and determining the sampling direction and blur radius used in each iteration; wherein the blur radius decreases as the number of iterations increases; the angle between adjacent sampling directions is the golden angle; in each iteration, the sampling pixel position is determined based on the sampling direction and the blur radius, with each pixel in the coarse-grained blurred background image as the center; the color value at the sampling pixel position in each iteration is extracted, and then the color values obtained in each iteration are accumulated and averaged to obtain the final color value of each pixel; the pixel is colored based on the final color value of each pixel to generate the fine-grained blurred background image.
7. The image fusion method according to claim 6, characterized in that, For each pixel in the coarse-grained blurred background image, in each iteration, the sampling pixel position is determined based on the sampling direction and the blur radius, with the pixel as the center. This includes: determining a positive offset in the positive direction of the sampling direction based on the blur radius and the sampling direction, and superimposing the positive offset onto the pixel position to obtain a first sampling pixel position; determining a negative offset in the opposite direction of the sampling direction based on the blur radius and the sampling direction at a preset multiple, and superimposing the negative offset onto the pixel position to obtain a second sampling pixel position.
8. An image fusion apparatus, characterized in that, include: A generation module is used to generate a foreground portrait and a low-resolution background image to be fused; wherein the foreground portrait has a transparency channel to support fusion with different background images; The blurring module is used to blur the low-resolution background image based on a preset random sampling strategy to obtain a coarse-grained blurred background image; the blurring module is also used to blur the coarse-grained blurred background image using a preset rule sampling strategy to obtain a fine-grained blurred background image; the fusion module is used to perform image fusion of the foreground portrait to be fused and the fine-grained blurred background image.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor to implement the image fusion method according to any one of claims 1 to 7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image fusion method as described in any one of claims 1 to 7.