Image fusion method and device, computer equipment and storage medium

By using the original long frame as the reference frame in HDR image fusion to perform global alignment and ghost map generation, the image quality problem caused by only short frames in the moving area in the prior art is solved, and a higher quality image fusion effect is achieved.

CN120147145APending Publication Date: 2025-06-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311719659.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When using high dynamic range (HDR) technology to fuse long and short frames, the prior art only takes short frames in the moving area, resulting in poor image quality of the fused image.

Method used

Taking the original long frame as the reference frame, the original short frame is globally aligned with the original long frame, a global ghost map is generated, and the globally aligned short frame is fused with the original long frame based on the global ghost map in the motion area to obtain the first fusion frame.

Benefits of technology

By fusing globally aligned short frames and original long frames, the picture quality of the fused image is improved, and the picture quality decline caused by taking only short frames in the moving area in the prior art is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147145A_ABST
    Figure CN120147145A_ABST
Patent Text Reader

Abstract

The invention relates to an image fusion method and device, computer equipment and a storage medium. The method comprises the following steps: by taking an original long frame as a reference frame, carrying out global alignment on an original short frame and the original long frame to obtain a global alignment short frame; a global ghosting map is generated based on the global alignment short frame and the original long frame, and a first motion area of the global alignment short frame relative to the original long frame is marked in the global ghosting map; and in the first motion area, based on a global ghosting map, fusing the global alignment short frame and the original long frame to obtain a first fusion frame. The image quality of the fused image can be improved by adopting the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an image fusion method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of image processing technology, High-Dynamic Range (HDR) technology has emerged. The technical principle of HDR is to merge several photos with different exposures together to retrieve the high-light and shadow details in a high-contrast environment, so as to improve the brightness and contrast of the details in the picture.

[0003] Currently, when using HDR technology to fuse long frames and short frames, the short frame is used as the reference frame. After the long frame and the short frame are registered, the motion between the long frame and the short frame can be calculated to remove ghosts. In the motion area, usually only the short frame is taken during fusion to ensure no ghosts, but the quality of the fused image is poor in this way. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an image fusion method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the quality of the fused image.

[0005] In a first aspect, this application provides an image fusion method, including:

[0006] Using the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame;

[0007] Based on the globally aligned short frame and the original long frame, generating a global ghost map, where the first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map;

[0008] In the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame.

[0009] In one embodiment, the step of using the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame includes:

[0010] Using the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector;

[0011] Using the original long frame as the reference frame, globally aligning the original short frame based on the global motion vector to obtain the globally aligned short frame.

[0012] In one embodiment, taking the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector includes:

[0013] Performing brightness alignment processing on the original short frame according to the exposure ratio between the original long frame and the original short frame to obtain a brightness-aligned short frame;

[0014] Taking the original long frame as the reference frame, globally registering the brightness-aligned short frame with the original long frame to obtain a global motion vector.

[0015] In one embodiment, within the first motion area, fusing the globally aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame includes:

[0016] Performing brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame to obtain a target-aligned short frame;

[0017] Within the first motion area, fusing the target-aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame.

[0018] In one embodiment, within the first motion area, fusing the globally aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame includes:

[0019] For each pixel in the first motion area of the global ghost map, calculating an error parameter of the globally aligned short frame relative to the original long frame;

[0020] According to the error parameter of each pixel in the first motion area, within the first motion area, fusing the globally aligned short frame with the original long frame to obtain the first fused frame.

[0021] In one embodiment, after fusing the globally aligned short frame with the original long frame based on the global ghost map within the first motion area to obtain a first fused frame, the method further includes:

[0022] Determining a locally aligned short frame according to the original short frame and the globally aligned short frame;

[0023] Generating a local ghost map based on the locally aligned short frame and the original long frame, where the local ghost map is marked with a second motion area of the locally aligned short frame relative to the original long frame;

[0024] Within the second motion area, fuse the local aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame.

[0025] In one embodiment, the determining the local aligned short frame according to the original short frame and the globally aligned short frame includes:

[0026] Input the original short frame and the globally aligned short frame into an artificial intelligence optical flow network to obtain the output local motion vectors;

[0027] Taking the original long frame as the reference frame, locally align the original short frame based on the local motion vectors to obtain a local aligned short frame.

[0028] In one embodiment, the method further includes: obtaining a confidence map corresponding to the local motion vectors output by the artificial intelligence optical flow network;

[0029] The fusing the local aligned short frame and the first fused frame according to the local ghost map within the second motion area to obtain a second fused frame includes:

[0030] Determine a fusion weight according to the local ghost map and the confidence map;

[0031] Within the second motion area, fuse the local aligned short frame and the first fused frame according to the fusion weight to obtain a second fused frame.

[0032] In one embodiment, the determining the fusion weight according to the local ghost map and the confidence map includes

[0033] Perform image processing on the confidence map to obtain a weight map: wherein the image processing includes at least one of erosion processing, blurring processing, contrast and brightness adjustment:

[0034] Determine a fusion weight according to the local ghost map and the weight map.

[0035] In one embodiment, the inputting the original short frame and the globally aligned short frame into the artificial intelligence optical flow network includes:

[0036] Downsample the original short frame and the globally aligned short frame and input them into the artificial intelligence optical flow network;

[0037] The image processing further includes: upsampling processing.

[0038] In one embodiment, within the second motion area, fusing the locally aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame includes:

[0039] For each pixel in the second motion area of the local ghost map, calculating an error parameter of the locally aligned short frame relative to the first fused frame;

[0040] According to the error parameter of each pixel in the second motion area;

[0041] According to the error parameter of each pixel in the second motion area, within the second motion area, fusing the locally aligned short frame and the first fused frame to obtain a second fused frame.

[0042] In a second aspect, the present application further provides an image fusion device, including:

[0043] An alignment module, configured to globally align an original short frame with the original long frame using the original long frame as a reference frame to obtain a globally aligned short frame;

[0044] A generation module, configured to generate a global ghost map based on the globally aligned short frame and the original long frame, where a first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map;

[0045] A fusion module, configured to fuse the globally aligned short frame and the original long frame within the first motion area to obtain a first fused frame.

[0046] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0047] Taking the original long frame as a reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame;

[0048] Generating a global ghost map based on the globally aligned short frame and the original long frame, where a first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map;

[0049] Fusing the globally aligned short frame and the original long frame within the first motion area to obtain a first fused frame.

[0050] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0051] Taking the original long frame as the reference frame, globally align the original short frame with the original long frame to obtain a globally aligned short frame;

[0052] Based on the globally aligned short frame and the original long frame, generate a global ghost map, in which the first motion area of the globally aligned short frame relative to the original long frame is marked;

[0053] Within the first motion area, based on the global ghost map, fuse the globally aligned short frame with the original long frame to obtain a first fused frame.

[0054] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0055] Taking the original long frame as the reference frame, globally align the original short frame with the original long frame to obtain a globally aligned short frame;

[0056] Based on the globally aligned short frame and the original long frame, generate a global ghost map, in which the first motion area of the globally aligned short frame relative to the original long frame is marked;

[0057] Within the first motion area, fuse the globally aligned short frame with the original long frame to obtain a first fused frame.

[0058] In the above image fusion method, device, computer device, storage medium and computer program product, taking the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame, and based on the globally aligned short frame and the original long frame, generating a global ghost map, and within the first motion area marked in the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame. In this way, the first fused frame obtained will include some features in the original long frame, which can improve the image quality of the fused image compared with only taking the short frame in the motion area in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0060] Figure 1 Schematic flowchart of the image fusion method in an embodiment Figure 1 ;

[0061] Figure 2 Schematic of the image fusion process in one embodiment Figure 1 ;

[0062] Figure 3 Schematic diagram of the image fusion result in one embodiment;

[0063] Figure 4 Schematic flow of the image fusion method in one embodiment Figure 2 ;

[0064] Figure 5 Schematic of the image fusion process in one embodiment Figure 2 ;

[0065] Figure 6 Schematic diagram of the Sigmoid curve in one embodiment;

[0066] Figure 7 Structure block diagram of the image fusion device in one embodiment;

[0067] Figure 8 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0068] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0069] The technical principle of HDR is to merge several photos with different exposures together to retrieve the high - light and shadow details in a high - contrast environment, so as to improve the light - and - shade contrast of the picture details.

[0070] The goal of HDR technology is to capture underexposed frames, align and merge these frames to generate a single intermediate image with a high bit depth, and perform tone mapping on this image to generate a high - resolution photo.

[0071] Current HDR technology faces ghosting problems and image quality problems. Among them, the ghosting problem is manifested as the ghosting of moving objects (such as when a person waves quickly, there will be two hands or a broken hand, etc.), and the image quality problem is manifested as image quality jumps and uneven distribution of noise around the moving area, etc.

[0072] Taking two-frame HDR fusion as an example, due to dynamics, signal-to-noise ratio, motion, etc., for the two frames, a fusion map will be calculated to guide which positions to take the long frame and which positions to take the dark frame to ensure an increase in the dynamic range of the picture and ensure that no ghosting appears. In the scenario of a large backlight for a portrait, if the portrait moves or the camera shakes, it will cause texture differences between the long and short frames. To protect the dynamic range, the moving area will be selected as the short frame during fusion to be incorporated into the result. However, incorporating the short frame will lead to a sharp decline in image quality. Therefore, for the scenario of a backlit moving portrait, this paper designs an algorithm scheme for image quality restoration to protect the dynamics while improving the image quality of the fused frame.

[0073] Ghosting problem

[0074] Traditional HDR technology synthesizes one image based on multiple images with different exposures. However, ideal shooting scenarios are rare after all. When taking handheld photos, the camera may shake, or the object being photographed may move too fast, and the multiple pictures taken may be different. Forcing synthesis will cause ghosting phenomena.

[0075] Image quality problem

[0076] HDR fusion requires fusing the long and short frames to obtain the result. So, some materials in the result map come from the long frame and some from the short frame. However, there are differences in exposure between the long and short frames, that is, the generation time of the frames and the image gain (Gain) are different, and there will be a large difference in the signal-to-noise ratio between the two frames.

[0077] The brightness ratio relationship between the long and short frames is called the exposure ratio (Ratio). Then, the smaller the exposure ratio, the smaller the drop in the signal-to-noise ratio (Signal-to-Noise Ratio Drop, SNR Drop). If the exposure rate is 1, there is no SNR Drop; on the contrary, if the exposure rate is larger, the dynamic range is larger, and the SNR Drop is also larger.

[0078] In existing HDR technologies, after registering the long and short frames, ghosting (ghost) can be removed by calculating the motion between the long and short frames. In the moving area, in order to ensure no ghosting, only the short frame can be taken, which will result in poor image quality. In the scenario of a large backlight, when the Ratio is large, SNR Drop is very likely to occur between the long and short frames; also, in the night scene, it will be interfered by noise, resulting in the incorporation of more short frames, and the image quality will be significantly lost compared to the long frame, and the signal-to-noise ratio is large.

[0079] The image fusion method provided in the embodiments of this application can improve the image quality of the fused image, and can effectively improve the image quality of the fused image in night scenes, large backlight portraits, and non-portrait scenes.

[0080] The image fusion method provided by the embodiments of the present application can improve the image quality of the fused image and can be applied to a computer device or an image fusion device with image processing functions. The image fusion device can be a functional module or a functional entity in the computer device. The computer device can be a server or a terminal. The terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart TVs, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, head-mounted devices, etc. The server can be implemented by an independent server or a server cluster composed of multiple servers.

[0081] In an exemplary embodiment, as Figure 1 shown, a flowchart of an image fusion method is provided Figure 1 , and the method includes the following steps 101 to 103.

[0082] Step 101: Taking the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame.

[0083] In HDR technology, the long frame and the short frame refer to image frames obtained under different exposure times.

[0084] The long frame usually corresponds to a longer exposure time, such as several seconds or several minutes. Under long-frame exposure, the camera can capture more light, thus showing darker areas and more details in the image. However, long-frame exposure may also cause motion blur in the image because any moving object will leave a blurred track in the image during the exposure.

[0085] The short frame usually corresponds to a shorter exposure time, such as several milliseconds or dozens of milliseconds. Under short-frame exposure, the camera can capture faster moving objects and can reduce motion blur. However, short-frame exposure may also cause the image to be too dark because the camera cannot capture enough light.

[0086] In HDR technology, a combination of long frames and short frames is usually used to capture different parts of the image. For example, long frames can be used to capture dark details, short frames can be used to capture high-light details, and then these images are merged into an HDR image. This method can reduce the problems of motion blur and over-darkness while retaining image details.

[0087] Among them, the above-mentioned original long frame and original short frame are the long frame and short frame that have not undergone image processing (such as global alignment).

[0088] The above global alignment is an image alignment technology used to align two or more image frames to ensure that they have the same position and orientation in space.

[0089] In some embodiments, taking the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame may include, but is not limited to: taking the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector; taking the original long frame as the reference frame, globally aligning the original short frame based on the global motion vector to obtain a globally aligned short frame.

[0090] Among them, the original short frame and the original long frame can be registered by computer vision (CV) to calculate the global motion vector, and the global motion vector can be expressed as Global MV.

[0091] In CV registration, there can be two main steps: feature extraction and registration algorithm. Feature extraction refers to extracting features from an image for registration. These features can be points, lines, regions, or other forms of image features, which can be used to describe important information in the image. Common feature extraction methods include the Harris corner detection algorithm, Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), etc. The registration algorithm refers to the process of matching or aligning the extracted features. Common registration algorithms include feature point-based registration, gray value-based registration, transformation model-based registration, etc.

[0092] Feature point-based registration is a common registration method that determines the relative position and orientation between two images by comparing the positions and orientations of the feature points extracted from the two images. Gray value-based registration determines the relative position and orientation between two images by comparing the gray value distributions of the two images. Transformation model-based registration aligns the images by transforming the images.

[0093] In some embodiments, taking the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector includes: performing brightness alignment processing on the original short frame according to the exposure ratio between the original long frame and the original short frame to obtain a brightness-aligned short frame; taking the original long frame as the reference frame, globally registering the brightness-aligned short frame with the original long frame to obtain a global motion vector.

[0094] Through the original long frame and the original short frame, brightness alignment can be first completed. First, obtain the exposure time (Exp Time) and exposure gain (Exp Gain) of the original long frame and the original short frame. The exposure ratio between the original long frame and the original short frame is denoted as Ratio, and Ratio can be calculated using the following formula (1):

[0095] Ratio = ExpTime_long * ExpGain_long / (ExpTime_short * ExpGain_short) (1)

[0096] Wherein, ExpTime_long represents the exposure time of the original long frame, ExpGain_long represents the exposure gain of the original long frame, ExpTime_short represents the exposure time of the original short frame, and ExpGain_short represents the exposure gain of the original short frame.

[0097] In some embodiments, the luminance-aligned short frame is the product of the original short frame and Ratio.

[0098] In some embodiments, when calculating the luminance alignment of the original short frame, the black level needs to be deducted. The result after luminance alignment of the original short frame can be expressed by the following formula (2):

[0099] Sexp = (S_input - S_blc) * Ratio + S_blc (2)

[0100] Wherein, Sexp represents the luminance-aligned short frame, S_input represents the original short frame, and S_blc represents the parameter of the black level.

[0101] Performing the above luminance alignment on the original long frame and the original short frame can reduce the luminance difference between images, thereby improving the quality and clarity of the images, making the images more visually consistent, thus enhancing the visual effect and aesthetic feeling. In subsequent applications such as registration, luminance alignment can improve the matching accuracy of feature points, thereby improving the performance of the algorithm. And in some cases, luminance alignment can also reduce the computational amount of subsequent processing, thereby improving the efficiency of the algorithm.

[0102] Step 102: Generate a global ghost map based on the globally aligned short frame and the original long frame.

[0103] Wherein, the first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map.

[0104] The above-mentioned globally aligned short frame refers to the short frame image obtained by globally aligning the original short frame, which contains all the pixels of the original short frame, but their positions and directions have been aligned with the original long frame.

[0105] The above-mentioned global ghost map is a map used to describe the ghost phenomenon in an image. Ghosts refer to false, duplicate, or semi-transparent objects that appear in an image, usually caused by the movement of the camera or the object. The global ghost map is usually a grayscale image, where the value of each pixel represents the intensity of the ghost at that pixel position. In the global ghost map, regions with higher ghost intensity usually appear as darker pixels, while regions with lower ghost intensity appear as brighter pixels. The calculation of ghost intensity is usually based on the pixel differences between the globally aligned short frame and the original long frame, so the ghost map can reflect the position and intensity distribution of ghosts in the image.

[0106] Generating a global ghost map based on the globally aligned short frame and the original long frame means that by comparing the pixel positions and color values in the globally aligned short frame and the original long frame, it is possible to determine whether there is a ghost phenomenon in the image and generate a global ghost map. This map can be used for subsequent image processing operations, such as ghost removal, image enhancement, etc., to improve the quality and details of the image. The generation of the global ghost map usually involves techniques such as image alignment, pixel comparison, and ghost detection. The specific implementation method will vary depending on the algorithm used and the application scenario.

[0107] Step 103: In the first motion area, based on the global ghost map, fuse the globally aligned short frame and the original long frame to obtain a first fused frame.

[0108] In some embodiments, during the process of fusing the globally aligned short frame and the original long frame based on the global ghost map in the first motion area to obtain a first fused frame, only the content of the original long frame can be taken in the first motion area, and the content of the globally aligned short frame is not taken. That is to say, in the first motion area, the fusion weight corresponding to the original long frame is set to 1, and the fusion weight of the globally aligned short frame is set to 0.

[0109] In some embodiments, the fusion weight for fusing the globally aligned short frame and the original long frame can be determined according to the global ghost map. For example, the fusion weight of the original long frame can be set corresponding to the ghost intensity in the global ghost map.

[0110] Exemplarily, in the global ghost map, regions with higher ghost intensity usually appear as darker pixels. In these regions, the weight of the original long frame can be set larger, while regions with lower ghost intensity appear as brighter pixels, and the weight of the original long frame in these regions can be set smaller.

[0111] In the above image fusion method, the original long frame is used as the reference frame, the original short frame is globally aligned with the original long frame to obtain a globally aligned short frame, and based on the globally aligned short frame and the original long frame, a global ghost map is generated. And within the first motion area marked in the global ghost map, the globally aligned short frame is fused with the original long frame to obtain a first fused frame. In this way, the first fused frame obtained will include some features in the original long frame. Compared with only taking the short frame in the motion area in the prior art, the image quality of the fused image can be improved.

[0112] In some embodiments, within the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame may include, but is not limited to: performing brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame to obtain a target aligned short frame; within the first motion area, based on the global ghost map, fusing the target aligned short frame with the original long frame to obtain a first fused frame.

[0113] Among them, the manner of performing brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame to obtain a target aligned short frame is the same as that of performing brightness alignment processing on the original short frame according to the exposure ratio between the original long frame and the original short frame to obtain a brightness aligned short frame.

[0114] The manner of calculating the exposure ratio between the above original long frame and the globally aligned short frame is similar to the manner of calculating the exposure ratio between the above original long frame and the original short frame, and will not be elaborated here.

[0115] Performing the above brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame can reduce the brightness difference between images, thereby improving the quality and clarity of the images, making the images more visually consistent, thus enhancing the visual effect and aesthetic feeling. In image fusion, the image quality of the fused image after brightness alignment is better.

[0116] In some embodiments, within the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame includes: for each pixel in the first motion area of the global ghost map, calculating the error parameter of the globally aligned short frame relative to the original long frame; according to the error parameter of each pixel in the first motion area, within the first motion area, fusing the globally aligned short frame with the original long frame to obtain a first fused frame.

[0117] In some embodiments, the error parameter of each pixel in the above first motion area is determined according to the inter-frame difference and the brightness parameter. The error parameter of each pixel in the first motion area can be calculated according to the following formula (3).

[0118] SAD = DIFF / (DIFF + c * luma) (3)

[0119] Wherein, c is a control factor, which is a constant. SAD represents the error parameter of pixels and can also be used as a measure of the similarity between two images. DIFF represents the absolute value of the difference in grayscale values of corresponding pixels in two images, and luma represents the luminance parameter, which is the average value of the grayscale values of corresponding pixels in two images.

[0120] Wherein, according to the error parameter of each pixel in the first motion area, when fusing the globally aligned short frame and the original long frame within the first motion area to obtain the first fused frame, the fusion weight of the original long frame corresponding to each pixel can be determined according to the SAD of each pixel. For example, when the SAD is larger, it is determined that there is local motion in this area, and at this time, the fusion weight of the original long frame is set larger; when the SAD is smaller, it is determined that the possibility of local motion in this area is smaller, and at this time, the fusion weight of the original long frame is set smaller.

[0121] It should be noted that within the first motion area, when fusing the globally aligned short frame and the original long frame based on the global ghost map to obtain the first fused frame, for the first motion area, the globally aligned short frame and the original long frame will be fused based on the global ghost map. For the area outside the first motion area, the content of the globally aligned short frame will be incorporated into the overexposed area, and the content of the original long frame will be incorporated into the non-overexposed area.

[0122] Figure 2 The following shows the schematic of the image fusion process in an embodiment Figure 1 . As Figure 2 shown in (a) represents the original long frame. As Figure 2 shown in (b) represents the original short frame. It can be seen that there is a difference in the position height of the people in the original short frame and the original long frame, indicating that there is global motion between the original short frame and the original long frame. After globally aligning the original short frame and the original long frame with the original long frame as the reference frame, the globally aligned short frame as shown in (c) in Figure 2 can be obtained. It can be seen that Figure 2 in (c), the position of the people after global alignment is the same as the position height of the people in (a) in Figure 2 . The area indicated by the dotted circle in (c) in Figure 2 is the first motion area. The area to the left of the vertical line represents the overexposed area, and the area to the right of the vertical line represents the non-overexposed area. After image fusion, the first fused frame as shown in (d) in Figure 2 can be obtained. In the first motion area indicated by the circle in (d) in Figure 2 , the content of the fusion of the globally aligned short frame and the original long frame is taken; in Figure 2Outside the circle in (d) and in the overexposed area to the left of the vertical line, the content of the short frame is taken; in Figure 2 Outside the circle in (d) and in the non-overexposed area to the right of the vertical line, the content of the long frame is taken.

[0123] It should be noted that Figure 2 The first fused frame shown in (d) is the fused result under the ideal fusion effect. However, in practical applications, ghost problems may occur in the human arm with local movement. Therefore, after fusing the original short frame and the globally aligned short frame, the fused result of the image may be as shown in Figure 3 The image fused result shown. Figure 3 It is a schematic diagram of the image fused result in an embodiment. As shown in Figure 3 When there is local movement in the human arm, as shown in Figure 3 The first fused frame obtained by fusing the image shown may have a broken hand situation.

[0124] In an exemplary embodiment, as shown in Figure 4 A flowchart of an image fusion method is provided. Figure 2 This method includes the following steps 401 to step 406.

[0125] 401. Taking the original long frame as the reference frame, globally align the original short frame with the original long frame to obtain a globally aligned short frame.

[0126] 402. Based on the globally aligned short frame and the original long frame, generate a global ghost map, and the global ghost map is marked with the first motion area of the globally aligned short frame relative to the original long frame.

[0127] 403. In the first motion area, based on the global ghost map, fuse the globally aligned short frame with the original long frame to obtain the first fused frame.

[0128] For the descriptions of the above steps 301 to step 303, reference can be made to the relevant descriptions of the above steps 101 to step 103, which will not be elaborated here.

[0129] 404. Determine a locally aligned short frame according to the original short frame and the globally aligned short frame.

[0130] In some embodiments, determining a locally aligned short frame according to the original short frame and the globally aligned short frame may include, but is not limited to: inputting the original short frame and the globally aligned short frame into an artificial intelligence optical flow network to obtain the output local motion vector; taking the original long frame as the reference frame, and locally aligning the original short frame based on the local motion vector to obtain a locally aligned short frame.

[0131] Among them, the above artificial intelligence optical flow network is a pre-trained network model for determining local motion vectors.

[0132] The above-mentioned local motion vectors can be calculated by comparing the pixel values of adjacent frames. Specifically, the optical flow method or the block matching algorithm can be used to estimate the pixel offsets between adjacent frames, and then these offsets are combined into vectors to represent local motion.

[0133] The above-mentioned locally aligned short frames refer to the new short frames obtained by aligning the original short frames with local motion vectors. The purpose of local alignment is to make the object motion in the image sequence smoother and more continuous by aligning the pixel positions between adjacent frames.

[0134] In some embodiments, inputting the original short frame and the globally aligned short frame into an artificial intelligence optical flow network includes: downsampling the original short frame and the globally aligned short frame and inputting them into the artificial intelligence optical flow network.

[0135] Among them, downsampling is a common image processing technique that can reduce the resolution and data volume of an image while maintaining the basic features of the image. After downsampling the original short frame and the globally aligned short frame, the resolution and data volume of the image input to the artificial intelligence optical flow network for processing can be reduced.

[0136] 405. Generate a local ghost map based on the locally aligned short frame and the original long frame.

[0137] Among them, the second motion area of the locally aligned short frame relative to the original long frame is marked in the local ghost map.

[0138] The local ghost map is a map used to describe the object motion in an image sequence. It is usually obtained by aggregating and visualizing the local motion vectors between adjacent frames.

[0139] The generation of the above-mentioned local ghost map based on the locally aligned short frame and the original long frame may include but is not limited to: calculating local motion vectors, aggregating local motion vectors, and generating a local ghost map. Figure 3 These steps include: when calculating local motion vectors, the pixel offsets between adjacent frames can be calculated by comparing the pixel values between the locally aligned short frame and the original long frame, using methods such as the optical flow method or the block matching algorithm to obtain local motion vectors; when aggregating local motion vectors, the calculated local motion vectors can be aggregated, and usually methods such as the average value or the median value can be used to obtain the global motion vector of each pixel point; when generating the local ghost map, the aggregated global motion vectors can be visualized, and usually methods such as colors or gray values can be used to represent the motion direction and speed of each pixel point on the image to obtain the local ghost map.

[0140] 406. In the second motion area, fuse the locally aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame.

[0141] In some embodiments, the fusion weight for fusing the locally aligned short frame and the first fusion frame can be determined according to the local ghost map. For example, the fusion weight of the first fusion frame can be set according to the ghost intensity in the local ghost map.

[0142] Exemplarily, in the local ghost map, regions with higher ghost intensity usually appear as darker pixels, and the weight of the first fusion frame in these regions can be set larger, while regions with lower ghost intensity appear as brighter pixels, and the weight of the first fusion frame in these regions can be set smaller.

[0143] After obtaining the first fusion frame through the first fusion in the above image fusion method, the locally aligned short frame and the local ghost map can be further determined, so as to know the second motion region without local alignment. And for the second motion region, the locally aligned short frame and the first fusion frame are fused again according to the local ghost map to obtain the second fusion frame. The second fusion frame obtained in this way can not only include the content in the original long frame, but also achieve the effect of local alignment, avoiding problems such as broken hands and ghosting caused by local motion, thereby further improving the image quality of the fused image.

[0144] In an exemplary embodiment, Figure 5 is a schematic diagram of the image fusion process in an embodiment Figure 2 . As Figure 5 shown in Figure 5 , (1) in Figure 5 is the first fusion frame, Figure 5 , (2) in Figure 5 is the locally aligned short frame, and the dotted circles in (1) in Figure 5 and (2) in

[0145] are the second motion regions. Within the second motion regions, by aligning the locally aligned short frame and the first fusion frame, some content in the locally aligned short frame can be incorporated into the first fusion frame, so as to obtain the second fusion frame as shown in (3) in

[0146] The above-mentioned fusion of the locally aligned short frame and the first fusion frame according to the local ghost map within the second motion region to obtain the second fusion frame may include but is not limited to: determining the fusion weight according to the local ghost map and the above-mentioned confidence map; within the second motion region, fusing the locally aligned short frame and the first fusion frame according to the fusion weight to obtain the second fusion frame.

[0147] In some embodiments, when determining the fusion weights based on the local ghost map and the confidence map, image processing can be performed on the confidence map to obtain a weight map, and then the fusion weights can be determined based on the local ghost map and the weight map.

[0148] The above-mentioned confidence map is a visualization tool for representing data uncertainty, and usually uses grayscale values to represent the magnitude of confidence. Among them, 0 represents complete uncertainty, and 255 represents complete certainty. In the confidence map, the grayscale value of each pixel represents the confidence level of the area represented by the pixel. The higher the grayscale value, the higher the confidence level of the area, and vice versa. Usually, the confidence map divides the grayscale values into multiple levels to better display the confidence differences in different areas. For example, the grayscale value range can be divided into four levels: 0-64, 65-128, 129-192, and 193-255, representing uncertainty, low confidence, medium confidence, and high confidence respectively. In this way, users can quickly understand the uncertainty level of the data by observing the distribution of grayscale values.

[0149] It should be noted that the grayscale value range and division method of the confidence map can be adjusted according to the specific application scenario and data characteristics to better meet the needs of users.

[0150] Among them, the image processing includes at least one of the following processes:

[0151] Erosion processing (erode), blur processing (blur), contrast and brightness adjustment.

[0152] The above-mentioned erosion processing (erode): In digital image processing, erode is a morphological operation used to erode the target in the image. It gradually eliminates the boundary of the target by using a structuring element, thereby reducing the size of the target. In morphological processing, the structuring element is usually a small shape (such as a rectangle, a circle, etc.), which moves on the image and is compared with the target. If the structuring element overlaps with a part of the target, that part is deleted from the target. This process is repeated until the structuring element can no longer overlap with the target. Erode can be used to remove noise in the image, smooth the boundary, extract features of the target, etc.

[0153] The above-mentioned blurring: It is an image or signal processing technique used to reduce details and high-frequency information in an image or signal, making it appear smoother and softer. Blurring can be achieved through various methods, such as mean blurring, median blurring, Gaussian blurring, etc. In image processing, blurring is usually used to reduce noise in an image, remove details, or smooth edges. For example, in photography, blurring can be used to simulate the shallow depth-of-field effect, making the transition between the foreground and background more natural. In video processing, blurring can be used to reduce motion blur or remove noise. Blurring is used to improve the quality of images and signals, making them more suitable for specific application scenarios.

[0154] The above-mentioned contrast and brightness adjustment is achieved through the Sigmoid function. The confidence map can be adjusted through the Sigmoid function so that the gray values in the confidence map make the bright areas brighter and the dark areas darker.

[0155] In an exemplary embodiment, Figure 6 It is a schematic diagram of the Sigmoid curve in an embodiment. This curve is used to reflect the relationship between the input X of the Sigmoid function and the output Sigmoid(X) of the Sigmoid function. The confidence map can be adjusted through the Sigmoid function shown by this Sigmoid curve.

[0156] In some embodiments, if the original short frame and the globally aligned short frame are downsampled before being input into the artificial intelligence optical flow network, then correspondingly, after the artificial intelligence optical flow network outputs the local motion vector and the confidence map corresponding to this local motion vector, upsampling processing can also be performed on the local motion vector and the confidence map corresponding to this local motion vector. The above image processing can also include upsampling processing.

[0157] In some embodiments, within the second motion area, the local aligned short frame and the first fused frame are fused according to the local ghost map to obtain the second fused frame, including: for each pixel in the second motion area of the local ghost map, calculating the error parameter of the local aligned short frame relative to the first fused frame; according to the error parameter of each pixel in the second motion area; according to the error parameter of each pixel in the second motion area, within the second motion area, the local aligned short frame and the first fused frame are fused to obtain the second fused frame.

[0158] In some embodiments, the error parameter of each pixel in the second motion area is determined according to the inter-frame difference and the brightness parameter.

[0159] The calculation method of the error parameter of each pixel in the second motion area is similar to that of each pixel in the first motion area described above, and will not be elaborated here.

[0160] Among them, according to the error parameter of each pixel in the second motion area, when fusing the local alignment short frame and the first fusion frame within the second motion area to obtain the second fusion frame, the fusion weight of the local alignment short frame corresponding to each pixel can be determined according to the SAD of each pixel. For example, when the SAD is larger, it is determined that there is local motion in this area, and at this time, the fusion weight of the local alignment short frame is set larger; when the SAD is smaller, it is determined that the possibility of local motion in this area is smaller, and at this time, the fusion weight of the local alignment short frame is set smaller.

[0161] The error parameter can reflect the difference between the local alignment short frame and the first fusion frame. For those with a large difference, it is considered that there is local motion, and the content in the local motion short frame will be taken as much as possible through weight setting. For those with a small difference, it is considered that there is no local motion, and the content in the first fusion frame will be taken as much as possible to retain more features in the original long frame, so as to avoid the ghosting problem and improve the image quality at the same time.

[0162] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0163] Based on the same inventive concept, the embodiments of the present application also provide an image fusion device for implementing the image fusion method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the image fusion device provided below can refer to the limitations on the image fusion method in the above text, and will not be elaborated here.

[0164] In an exemplary embodiment, as Figure 7As shown in the figure, an image fusion device is provided, including: an alignment module 701, configured to globally align an original short frame with the original long frame using the original long frame as a reference frame to obtain a globally aligned short frame; a generation module 702, configured to generate a global ghost map based on the globally aligned short frame and the original long frame, where a first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map; and a fusion module 703, configured to fuse the globally aligned short frame with the original long frame based on the global ghost map within the first motion area to obtain a first fusion frame.

[0165] In some embodiments, the alignment module 701 is specifically configured to: the step of globally aligning the original short frame with the original long frame using the original long frame as a reference frame to obtain a globally aligned short frame includes: globally registering the original short frame with the original long frame using the original long frame as a reference frame to obtain a global motion vector; and globally aligning the original short frame based on the global motion vector using the original long frame as a reference frame to obtain the globally aligned short frame.

[0166] In some embodiments, the alignment module 701 is specifically configured to: the step of globally registering the original short frame with the original long frame using the original long frame as a reference frame to obtain a global motion vector includes: performing brightness alignment processing on the original short frame according to the exposure ratio between the original long frame and the original short frame to obtain a brightness-aligned short frame; and globally registering the brightness-aligned short frame with the original long frame using the original long frame as a reference frame to obtain a global motion vector.

[0167] In some embodiments, the fusion module 703 is specifically configured to: the step of fusing the globally aligned short frame with the original long frame based on the global ghost map within the first motion area to obtain a first fusion frame includes: performing brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame to obtain a target-aligned short frame; and fusing the target-aligned short frame with the original long frame based on the global ghost map within the first motion area to obtain a first fusion frame.

[0168] In some embodiments, the fusion module 703 is specifically configured to: the step of fusing the globally aligned short frame with the original long frame based on the global ghost map within the first motion area to obtain a first fusion frame includes: calculating an error parameter of the globally aligned short frame relative to the original long frame for each pixel in the first motion area of the global ghost map; and fusing the globally aligned short frame with the original long frame within the first motion area according to the error parameter of each pixel in the first motion area to obtain the first fusion frame.

[0169] In some embodiments, the alignment module 701 is specifically configured to: after the fusion module 703 fuses the global aligned short frame and the original long frame in the first motion area based on the global ghost map, determine a local aligned short frame according to the original short frame and the global aligned short frame; the generation module 702 is further configured to: generate a local ghost map based on the local aligned short frame and the original long frame, where the second motion area of the local aligned short frame relative to the original long frame is marked in the local ghost map; the fusion module 703 is specifically configured to: fuse the local aligned short frame and the first fusion frame in the second motion area according to the local ghost map to obtain a second fusion frame.

[0170] In some embodiments, the alignment module 701 is specifically configured to: the determining the local aligned short frame according to the original short frame and the global aligned short frame includes: inputting the original short frame and the global aligned short frame into an artificial intelligence optical flow network to obtain the output local motion vectors; using the original long frame as a reference frame, locally aligning the original short frame based on the local motion vectors to obtain a local aligned short frame.

[0171] In some embodiments, the alignment module 701 is further configured to: obtain a confidence map corresponding to the local motion vectors output by the artificial intelligence optical flow network; the fusion module 703 is specifically configured to: the fusing the local aligned short frame and the first fusion frame in the second motion area according to the local ghost map to obtain a second fusion frame includes: determining a fusion weight according to the local ghost map and the confidence map; fusing the local aligned short frame and the first fusion frame in the second motion area according to the fusion weight to obtain a second fusion frame.

[0172] In some embodiments, the fusion module 703 is specifically configured to: the determining the fusion weight according to the local ghost map and the confidence map includes performing image processing on the confidence map to obtain a weight map: where the image processing includes at least one of erosion processing, blur processing, contrast and brightness adjustment: determining the fusion weight according to the local ghost map and the weight map.

[0173] In some embodiments, the alignment module 701 is specifically configured to: the inputting the original short frame and the global aligned short frame into the artificial intelligence optical flow network includes: downsampling the original short frame and the global aligned short frame and inputting them into the artificial intelligence optical flow network; the image processing further includes: upsampling processing.

[0174] In some embodiments, the fusion module 703 is specifically configured to: in the second motion area, fuse the locally aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame, including: for each pixel in the second motion area of the local ghost map, calculate an error parameter of the locally aligned short frame relative to the first fused frame; according to the error parameter of each pixel in the second motion area; according to the error parameter of each pixel in the second motion area, fuse the locally aligned short frame and the first fused frame in the second motion area to obtain a second fused frame.

[0175] In some embodiments, the error parameter of each pixel in the second motion area is determined according to the inter-frame difference and the luminance parameter.

[0176] Each module in the above image fusion device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0177] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image fusion method.

[0178] Those skilled in the art can understand that Figure 8 the structure shown in

[0179] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: using the original long frame as a reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame; generating a global ghost map based on the globally aligned short frame and the original long frame, where a first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map; within the first motion area, fusing the globally aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame.

[0180] In one embodiment, the step of using the original long frame as a reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame includes: using the original long frame as a reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector; using the original long frame as a reference frame, globally aligning the original short frame based on the global motion vector to obtain the globally aligned short frame.

[0181] In one embodiment, the step of using the original long frame as a reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector includes: performing brightness alignment processing on the original short frame according to the exposure ratio between the original long frame and the original short frame to obtain a brightness-aligned short frame; using the original long frame as a reference frame, globally registering the brightness-aligned short frame with the original long frame to obtain a global motion vector.

[0182] In one embodiment, the step of within the first motion area, fusing the globally aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame includes: performing brightness alignment processing on the globally aligned short frame according to the exposure ratio between the original long frame and the globally aligned short frame to obtain a target-aligned short frame; within the first motion area, fusing the target-aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame.

[0183] In one embodiment, the step of within the first motion area, fusing the globally aligned short frame with the original long frame based on the global ghost map to obtain a first fused frame includes: for each pixel in the first motion area of the global ghost map, calculating an error parameter of the globally aligned short frame relative to the original long frame; according to the error parameter of each pixel in the first motion area, within the first motion area, fusing the globally aligned short frame with the original long frame to obtain the first fused frame.

[0184] In one embodiment, after fusing the global aligned short frame and the original long frame based on the global ghost map within the first motion area to obtain a first fused frame, the method further includes: determining a local aligned short frame according to the original short frame and the global aligned short frame; generating a local ghost map based on the local aligned short frame and the original long frame, where the second motion area of the local aligned short frame relative to the original long frame is marked in the local ghost map; and within the second motion area, fusing the local aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame.

[0185] In one embodiment, the determining a local aligned short frame according to the original short frame and the global aligned short frame includes: inputting the original short frame and the global aligned short frame into an artificial intelligence optical flow network to obtain the output local motion vectors; and taking the original long frame as the reference frame, locally aligning the original short frame based on the local motion vectors to obtain a local aligned short frame.

[0186] In one embodiment, the method further includes: obtaining a confidence map corresponding to the local motion vectors output by the artificial intelligence optical flow network; the fusing the local aligned short frame and the first fused frame according to the local ghost map within the second motion area to obtain a second fused frame includes: determining a fusion weight according to the local ghost map and the confidence map; and within the second motion area, fusing the local aligned short frame and the first fused frame according to the fusion weight to obtain a second fused frame.

[0187] In one embodiment, the determining a fusion weight according to the local ghost map and the confidence map includes performing image processing on the confidence map to obtain a weight map: where the image processing includes at least one of erosion processing, blur processing, contrast and brightness adjustment: and determining a fusion weight according to the local ghost map and the weight map.

[0188] In one embodiment, the inputting the original short frame and the global aligned short frame into the artificial intelligence optical flow network includes: downsampling the original short frame and the global aligned short frame and inputting them into the artificial intelligence optical flow network; and the image processing further includes: upsampling processing.

[0189] In one embodiment, within the second motion area, fusing the locally aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame includes: for each pixel in the second motion area of the local ghost map, calculating an error parameter of the locally aligned short frame relative to the first fused frame; according to the error parameter of each pixel in the second motion area; and within the second motion area, fusing the locally aligned short frame and the first fused frame according to the error parameter of each pixel in the second motion area to obtain a second fused frame.

[0190] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each process in the above image fusion method is implemented.

[0191] In one embodiment, a computer program product is provided, including a computer program that can implement each process in the above image fusion method when executed by a processor.

[0192] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0193] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0194] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An image fusion method, characterized in that, the method includes: Taking the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame; Based on the globally aligned short frame and the original long frame, generating a global ghost map, and the first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map; Within the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame.

2. The method according to claim 1, characterized in that, the step of taking the original long frame as the reference frame, globally aligning the original short frame with the original long frame to obtain a globally aligned short frame includes: Taking the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector; Taking the original long frame as the reference frame, globally aligning the original short frame based on the global motion vector to obtain the globally aligned short frame.

3. The method according to claim 2, characterized in that, the step of taking the original long frame as the reference frame, globally registering the original short frame with the original long frame to obtain a global motion vector includes: According to the exposure ratio between the original long frame and the original short frame, performing brightness alignment processing on the original short frame to obtain a brightness-aligned short frame; Taking the original long frame as the reference frame, globally registering the brightness-aligned short frame with the original long frame to obtain a global motion vector.

4. The method according to claim 1, characterized in that, the step of within the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame includes: According to the exposure ratio between the original long frame and the globally aligned short frame, performing brightness alignment processing on the globally aligned short frame to obtain a target-aligned short frame; Within the first motion area, based on the global ghost map, fusing the target-aligned short frame with the original long frame to obtain a first fused frame.

5. The method according to claim 2, characterized in that, the step of within the first motion area, based on the global ghost map, fusing the globally aligned short frame with the original long frame to obtain a first fused frame includes: For each pixel in the first motion area of the global ghost map, calculating the error parameter of the globally aligned short frame relative to the original long frame; According to the error parameter of each pixel in the first motion area, within the first motion area, fusing the globally aligned short frame with the original long frame to obtain the first fused frame.

6. The method according to any one of claims 1 to 5, characterized in that, after fusing the globally aligned short frame with the original long frame within the first motion area based on the global ghost map, the method further includes: Determining a locally aligned short frame according to the original short frame and the globally aligned short frame; Based on the locally aligned short frame and the original long frame, generate a local ghost map, where the second motion area of the locally aligned short frame relative to the original long frame is marked in the local ghost map; Within the second motion area, fuse the locally aligned short frame and the first fused frame according to the local ghost map to obtain a second fused frame.

7. The method according to claim 6, wherein, the determining the locally aligned short frame according to the original short frame and the globally aligned short frame includes: inputting the original short frame and the globally aligned short frame into an artificial intelligence optical flow network to obtain the output local motion vectors; using the original long frame as the reference frame, locally align the original short frame based on the local motion vectors to obtain a locally aligned short frame.

8. The method according to claim 7, wherein, the method further includes: obtaining a confidence map corresponding to the local motion vectors output by the artificial intelligence optical flow network; the fusing the locally aligned short frame and the first fused frame according to the local ghost map within the second motion area to obtain a second fused frame includes: determining a fusion weight according to the local ghost map and the confidence map; within the second motion area, fuse the locally aligned short frame and the first fused frame according to the fusion weight to obtain a second fused frame.

9. The method according to claim 8, wherein, the determining the fusion weight according to the local ghost map and the confidence map includes performing image processing on the confidence map to obtain a weight map: wherein the image processing includes at least one of erosion processing, blurring processing, contrast and brightness adjustment; determining the fusion weight according to the local ghost map and the weight map.

10. The method according to claim 9, wherein, the inputting the original short frame and the globally aligned short frame into the artificial intelligence optical flow network includes: downsampling the original short frame and the globally aligned short frame and inputting them into the artificial intelligence optical flow network; the image processing further includes: upsampling processing.

11. The method according to claim 6, wherein, the fusing the locally aligned short frame and the first fused frame according to the local ghost map within the second motion area to obtain a second fused frame includes: for each pixel in the second motion area of the local ghost map, calculating an error parameter of the locally aligned short frame relative to the first fused frame; according to the error parameter of each pixel in the second motion area; according to the error parameter of each pixel in the second motion area, within the second motion area, fuse the locally aligned short frame and the first fused frame to obtain a second fused frame.

12. The method according to claim 11, wherein, the error parameter of each pixel in the second motion area is determined according to the inter-frame difference and the brightness parameter.

13. An image fusion device, wherein, the device includes: An alignment module, configured to globally align an original short frame with the original long frame using the original long frame as a reference frame to obtain a globally aligned short frame; A generation module, configured to generate a global ghost map based on the globally aligned short frame and the original long frame, wherein a first motion area of the globally aligned short frame relative to the original long frame is marked in the global ghost map; A fusion module, configured to fuse the globally aligned short frame and the original long frame based on the global ghost map within the first motion area to obtain a first fused frame.

14. A computer device, comprising a memory and a processor, where the memory stores a computer program, characterized in that, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.