Training Method of Image Fusion Model, Image Fusion Method and Application

By building a fusion loss function and self-attention module, the accuracy and clarity of the image fusion model are improved, and the data dependence, calculation complexity and real-time problems of the image fusion method in the prior art are solved, thereby achieving high-quality image fusion effect.

CN119941531BActive Publication Date: 2025-07-08SUZHOU INS IMAGE SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510432562.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-08
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing image fusion method based on deep learning models relies on a large amount of labeled data, has high computational complexity and poor real-time performance, making it difficult to effectively apply in scenarios with high real-time performance requirements, and has poor sensitivity to structure and texture information, resulting in the inability to guarantee image clarity and details.

Method used

The fusion loss function is constructed, including the basic loss sub-item, the perceived loss sub-item and the structural similarity loss sub-item. The image features are extracted through the self-attention module, the attention weight map is generated, and the image weighting process is performed. Combined with the image overlap area cropping and registration technology, the accuracy and clarity of the image fusion model are improved.

Benefits of technology

The image fusion model's ability to restore edges, details and textures is enhanced, ensuring the brightness, contrast and local structure of the fusion result are consistent, improving the structural fidelity and overall visual effect of the image, and generating high-quality fusion images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941531B_ABST
    Figure CN119941531B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method for an image fusion model, an image fusion method and an application. Among them, in the training method for the image fusion model, the image fusion model is used to fuse images from the same perspective or different perspectives to generate a fused image, and the method includes: establishing a training sample set, where the training sample set includes image samples of a target object taken from the same perspective or different perspectives and a reference image of the target object; constructing a fusion loss function, where the fusion loss function includes at least two of a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term; training the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model. The training method for the image fusion model of the present invention can improve and strengthen the model's ability to restore pictures by constructing a fusion loss function, and enhance the structural fidelity of the overall image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a training method for an image fusion model, an image fusion method and an application thereof. Background Art

[0002] With the rapid development of photographic imaging technology and computational photography in recent years, and the continuous reduction of camera costs, multi-camera systems have gradually been widely used in various fields. Specifically, in the consumer electronics field, the main applications of multi-camera systems focus on expanding the field of view and enhancing the depth of field, while achieving computational photography effects such as portrait mode and zoom function by combining cameras with different focal lengths or apertures. In fields such as autonomous driving and robotics, multi-camera systems can organically fuse data from different fields of view, thereby effectively extracting depth information and enhancing the machine's perception ability of complex environments. In the fields of virtual reality (VR) and augmented reality (AR), multi-camera systems achieve the reconstruction of 3D objects or scenes by combining images from different perspectives.

[0003] Image fusion technology is a technology that synthesizes multiple images into one to present a more realistic visual effect of the photographed object. In the prior art, image fusion methods based on deep learning models can automatically learn to extract image features and then complete image fusion. However, in the above fusion methods, training depends on a large amount of labeled data. In addition, deep learning models have high computational complexity and large resource requirements, making it difficult to effectively implement and apply them in scenarios with high real-time requirements. In addition, when there are high requirements for image authenticity, the unexplainable characteristics of deep learning models also have a potential impact on their reliability and acceptability. Finally, deep learning models are less sensitive to structural and texture information, resulting in the inability to guarantee the clarity and details of the formed images.

[0004] Therefore, in view of the above technical problems, it is necessary to provide a training method for an image fusion model, an image fusion method and an application thereof. Summary of the Invention

[0005] The purpose of the present invention is to provide a training method for an image fusion model, an image fusion method and an application thereof, which can solve the problems of the above model's deficiencies in data dependence, computational efficiency, real-time performance, interpretability and detail retention.

[0006] In order to achieve the above purpose, a specific embodiment of the present invention provides a training method for an image fusion model, and the technical solution is as follows:

[0007] A training method for an image fusion model, the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image, and the method includes:

[0008] Establish a training sample set, where the training sample set includes image samples of the target object taken from the same perspective or different perspectives and the reference image of the target object;

[0009] Construct a fusion loss function, where the fusion loss function includes at least two of a basic loss sub - term, a perceptual loss sub - term, and a structural similarity loss sub - term. Among them, the basic loss sub - term includes a pixel loss sub - term and / or an image gradient loss sub - term. The pixel loss sub - term represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub - term represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub - term represents the high - frequency feature error of the fused image relative to the reference image; the structural similarity loss sub - term represents the structural similarity error of the fused image relative to the reference image;

[0010] Train the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model.

[0011] In one or more embodiments of the present invention, the method further includes pre - processing the image samples, specifically including:

[0012] Calculate the gradient map of the image samples respectively and normalize them;

[0013] Enhance the information of the corresponding sample images based on the normalized gradient maps.

[0014] In one or more embodiments of the present invention, the image fusion model includes a self - attention module, and the method further includes:

[0015] Extract the structural features of the image samples based on the self - attention module to generate corresponding attention weight maps, where the structural features include image edge features and / or texture features;

[0016] When fusing the feature maps of image samples from the same perspective or different perspectives, use the attention weight maps to perform weighted processing on them to generate a fused image.

[0017] In one or more embodiments of the present invention, the method further includes:

[0018] Obtain initial image samples of the target object taken from the same perspective or different perspectives;

[0019] Map the initial image samples to a reference plane and calculate the image overlapping area;

[0020] Crop the initial image samples based on the image overlapping area to obtain corresponding image samples.

[0021] In one or more embodiments of the present invention, the basic loss sub-item is constructed by weighting the pixel loss sub-item and the image gradient loss sub-item;

[0022] And / or, the structural similarity loss sub-item represents the similarity error of the fused image relative to the reference image in terms of brightness, contrast, and local structure;

[0023] And / or, the fusion loss function is constructed by weighting the basic loss sub-item, the perceptual loss sub-item, and the structural similarity loss sub-item.

[0024] A specific embodiment of the present invention further provides an image fusion method, and the technical solution is as follows:

[0025] An image fusion method, comprising:

[0026] Obtain a set of images to be fused, wherein the set of images to be fused includes images of a target object taken at different perspectives with different exposure levels;

[0027] Based on the image fusion model obtained by the above training method, fuse the images of the target object taken at the same perspective with different exposure levels in the set of images to be fused to obtain intermediate fused images for each perspective;

[0028] Register the intermediate fused images for each perspective and then fuse them to obtain the fused image of the target object.

[0029] In one or more embodiments of the present invention, based on the image fusion model obtained by the above training method, fuse the registered intermediate fused images for each perspective.

[0030] In one or more embodiments of the present invention, fuse the images of the target object taken at the same perspective with different exposure levels in the set of images to be fused by using a weighted average fusion strategy; and fuse the intermediate fused images for each perspective by using a maximum value fusion strategy.

[0031] A specific embodiment of the present invention further provides a training device for an image fusion model, and the technical solution is as follows:

[0032] A training device for an image fusion model, the image fusion model being used to fuse images of the same perspective or different perspectives to generate a fused image, the device comprising:

[0033] An acquisition module, configured to establish a training sample set, the training sample set including image samples of a target object taken at the same perspective or different perspectives and a reference image of the target object;

[0034] A construction module for constructing a fusion loss function, where the fusion loss function includes at least two of a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term. Among them, the basic loss sub-term includes a pixel loss sub-term and / or an image gradient loss sub-term. The pixel loss sub-term represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub-term represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-term represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-term represents the structural similarity error of the fused image relative to the reference image;

[0035] A determination module for training the image fusion model based on the fusion loss function and determining the model parameters of the image fusion model.

[0036] A specific embodiment of the present invention further provides an image fusion device, and the technical solution is as follows:

[0037] An image fusion device includes:

[0038] An acquisition module for acquiring a set of images to be fused, where the set of images to be fused includes images of a target object taken at different perspectives with different exposure degrees;

[0039] A fusion module for fusing the images of the target object taken at the same perspective with different exposure degrees in the set of images to be fused based on the above image fusion model to obtain intermediate fused images for each perspective;

[0040] A processing module for registering and then fusing the intermediate fused images for each perspective to obtain the fused image of the target object.

[0041] A specific embodiment of the present invention further provides an electronic device, and the technical solution is as follows:

[0042] An electronic device includes:

[0043] At least one processor; and

[0044] A memory that stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the training method of the image fusion model as described above or the image fusion method as described above.

[0045] A specific embodiment of the present invention further provides a machine-readable storage medium, and the technical solution is as follows:

[0046] A machine-readable storage medium stores executable instructions, and when the instructions are executed, the machine executes the training method of the image fusion model as described above or the image fusion method as described above.

[0047] Compared with the prior art, the present invention has at least one of the following beneficial technical effects:

[0048] 1. By constructing a fusion loss function, the training method of the image fusion model of the present invention can improve and strengthen the model's ability to restore the edges, details, textures, and local structures of pictures, ensure the consistency of the fusion results in terms of brightness, contrast, and local structure, and enhance the overall structural fidelity of the images.

[0049] 2. By calculating the gradient map of the image samples and performing normalization processing, the training method of the image fusion model of the present invention enhances the edge details and texture features, improves the sensitivity of feature extraction to textures and structures, enables it to better capture detailed information, and improves the accuracy of fusion.

[0050] 3. By adding a self-attention module, the training method of the image fusion model of the present invention enhances the model's ability to understand and capture key structural information and long-distance correlations, and thus can more accurately capture the significant structural information in complex scenes, significantly improving the clarity, texture fidelity, and overall visual effect of the fused images.

[0051] 4. The image fusion method of the present invention can perform fusion processing and registration processing on multiple pictures containing the target object, and can obtain high-quality 2D images of the target object with rich details. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0053] Figure 1 It is a diagram of the implementation environment of the training method of the image fusion model and the image fusion method in an embodiment of the present invention;

[0054] Figure 2 It is a flowchart of the training method of the image fusion model in an embodiment of the present invention;

[0055] Figure 3 It is a flowchart of the image fusion method in an embodiment of the present invention;

[0056] Figure 4 It is a layout diagram of a multi-camera system in an embodiment of the present invention;

[0057] Figure 5 It is a step diagram of the image fusion method in an embodiment of the present invention;

[0058] Figure 6 It is a module diagram of a training device for an image fusion model in an embodiment of the present invention;

[0059] Figure 7 It is a module diagram of an image fusion device in an embodiment of the present invention;

[0060] Figure 8 It is a hardware structure diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0061] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0062] With the rapid development of photographic imaging technology and computational photography in recent years, and the continuous reduction of camera costs, multi-camera systems have gradually been widely used in various fields. A multi-camera system is an imaging system composed of multiple cameras, which can capture image data from different angles or perspectives to meet the needs of various application scenarios. In the embodiments of the present invention, it is expected to integrate the image features of the images captured by multiple cameras, explore the unique detail advantages brought by different angles, and fuse images from different perspectives to improve the quality and clarity of the images and achieve rich detail presentation.

[0063] With reference to Figure 4 and Figure 5 , in a specific scenario example, a multi-camera system can capture the target object from different angles, and when capturing the target object at each angle, different exposure degrees can also be used according to the settings. The image fusion model first performs HDR fusion on the images with different exposure degrees at the same perspective to obtain an HDR picture, that is, the images at the same perspective are fused into one. Subsequently, a step-by-step registration operation is performed on the images from different perspectives, including global registration and local registration. Finally, the image fusion model (which can reuse the model in the previous fusion step) is used to fuse the registered multi-perspective images to obtain a single high-definition and detail-rich 2D image.

[0064] Refer to Figure 1, which shows a schematic diagram of the implementation environment provided by an exemplary embodiment of the present invention. The implementation environment includes a terminal and a server. Among them, data communication is carried out between the terminal and the server through a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network. The image fusion method disclosed in the present invention can be executed by the server. Correspondingly, the image fusion model can also be deployed in the server.

[0065] Alternatively, the terminal and the server can cooperate to run the image fusion method provided by the embodiments of the present application to complete the fusion of images.

[0066] In a system architecture where the terminal can provide matching computing power, the image fusion method disclosed in the present application can also be directly executed by the terminal. Correspondingly, the image fusion model can also be deployed on the terminal.

[0067] The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0068] Refer to Figure 2 , which introduces the training method of the image fusion model in an embodiment of the present invention. Among them, the image fusion model is used to fuse images from the same perspective or different perspectives to generate a fused image. The training method of the image fusion model includes:

[0069] S101. Obtain initial image samples of the target object taken from the same perspective or different perspectives; map the initial image samples to a reference plane, and calculate the image overlapping area; crop the initial image samples based on the image overlapping area to obtain corresponding image samples. Specifically, multiple exposure levels (such as low, medium, and high) can be used to synchronously capture the details of the target object under different lighting conditions by the same camera at the same time.

[0070] In this embodiment, the formula for mapping the initial image samples to the reference plane and calculating the image overlapping area is as follows:

[0071]

[0072] Among them, : The image capture area of the i-th camera; f i : The mapping function; K i : The camera internal parameter; R i : The camera external parameter rotation matrix; ti : The translational vector of the extrinsic camera parameters; (x , y): The two-dimensional image pixel coordinates; : The corresponding region on the reference plane.

[0073] Compared with simple image cropping, each image is first mapped to a unified reference plane, and then the overlapping regions are calculated uniformly on this plane, eliminating the errors caused by perspective distortion and ensuring the information consistency in the subsequent fusion process.

[0074] S102. Establish a training sample set, which includes image samples of the target object taken from the same perspective or different perspectives and the reference image of the target object.

[0075] The image samples in the training sample set are the image samples obtained by cropping in the above steps. And in this embodiment, the image samples can be further pre-processed, specifically including: calculating the gradient map of the image samples respectively and normalizing them; enhancing the information of the corresponding sample images based on the normalized gradient maps.

[0076] In image processing, the gradient map reflects the intensity and direction of the change in pixel values in the image (i.e., edge information). Using the gradient map for image enhancement can highlight the gradient information (edges, details) to improve the clarity, contrast or texture performance of the image. The gradient map can be calculated by first-order or second-order differential operators, and the gradient magnitude is weighted and fused with the original image to enhance the contrast of the edge region, or the enhancement intensity can be dynamically adjusted according to the local gradient, etc. In this embodiment, the gradient map of the image can be calculated first, and added to the corresponding sample image based on the normalized gradient map to enhance the information of the sample image.

[0077] S103. Construct a fusion loss function, which includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item. Among them, the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item. The pixel loss sub-item represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub-item represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-item represents the structural similarity error of the fused image relative to the reference image.

[0078] The fusion loss function can include three terms: a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term, or it can include only two of them. For example, the fusion loss function can include only the basic loss sub-term and the perceptual loss sub-term, or only the perceptual loss sub-term and the structural similarity loss sub-term, or only the basic loss sub-term and the structural similarity loss sub-term. It can be understood that the basic loss sub-term can include only a pixel loss sub-term or an image gradient loss sub-term, or it can include both the pixel loss sub-term and the image gradient loss sub-term.

[0079] Specifically, the basic loss sub-term is constructed by weighting the pixel loss sub-term and the image gradient loss sub-term; the structural similarity loss sub-term represents the similarity error between the fused image and the reference image in terms of brightness, contrast, and local structure; the fusion loss function is constructed by weighting the basic loss sub-term, the perceptual loss sub-term, and the structural similarity loss sub-term.

[0080] The above is the general concept of the fusion loss function. In this embodiment, the fusion loss function is further elaborated based on specific parameter settings, which does not limit the fusion loss function in this embodiment. Specifically as follows:

[0081] In a specific embodiment, the fusion loss function includes three terms: a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term, and the basic loss function includes two terms: a pixel loss sub-term and an image gradient loss sub-term.

[0082] The function of the basic loss sub-term is as follows:

[0083]

[0084] Among them, I p and I g respectively represent the predicted fused image and the reference image, i represents the RGB channel index, H g 、 W g respectively represent the height and width of the reference image, and λ represents the weight parameter. The above formula includes the pixel loss sub-term and the image gradient loss sub-term. By the image gradient loss, the gradient can be constrained, improving the ability to restore edge and detail features.

[0085] The function of the perceptual loss sub-term is as follows:

[0086]

[0087] Among them, f p and f grespectively represent the feature maps of the predicted fused image and the reference image, i denote the channel indices of the feature maps, C f 、H f and W f respectively represent the number of channels, height, and width of the feature maps. This item enhances the model's ability to restore texture and local structure through the comparison of high-frequency features.

[0088] The function of the structural similarity loss sub-item is as follows:

[0089]

[0090] where, I p and I g respectively represent the predicted fused image and the reference image. The structural similarity loss sub-item ensures the consistency of the fused result in terms of brightness, contrast, and local structure by optimizing the structural similarity index, further enhancing the overall structural fidelity of the image.

[0091] In summary, the fusion loss function of the model is expressed as:

[0092]

[0093] where, α、β and γ are the weight coefficients of each loss item. In this embodiment, by constructing the fusion loss function, the clarity, texture fidelity, and overall visual effect of the image are improved.

[0094] S104. Train the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model.

[0095] Furthermore, the image fusion model includes a self-attention module, and the training method of the image fusion model further includes:

[0096] S105. Extract the structural features of the image samples based on the self-attention module to generate corresponding attention weight maps, where the structural features include image edge features and / or texture features; when fusing the feature maps of image samples from the same or different perspectives, use the attention weight maps to perform weighted processing on them to generate the fused image.

[0097] In this embodiment, the features of the image can be extracted through a convolutional neural network, and the self-attention module uses its ability to capture long-range dependencies to generate attention weight maps based on structural features such as image edges and textures. Among them, the ability to capture long-range dependencies means that the self-attention module can obtain the structural features of image samples within a relatively large range.

[0098] In the feature fusion stage of the image sample, the feature map obtained by the convolutional neural network is weighted by the attention weight map to enhance the feature expression of the key region. At the same time, through the attention module, the model can achieve a comprehensive understanding of the image details and the overall structure, so it can capture the significant structure information in complex scenes more accurately, significantly improving the clarity, texture fidelity, and overall visual effect of the fused image.

[0099] Referring to Figure 3 , an image fusion method in an embodiment of the present invention is introduced. It should be noted in advance that the image fusion model involved in the image fusion method of this embodiment can be obtained based on the above-mentioned training method of the image fusion model.

[0100] The following specifically introduces an image fusion method in an embodiment of the present invention, which specifically includes:

[0101] S201. Obtain a set of images to be fused, where the set of images to be fused includes images of the target object taken at different perspectives with different exposure levels.

[0102] In this embodiment, images of the target object taken at different perspectives with different exposure levels can be obtained by shooting with a multi-camera system. The layout of the multi-cameras can obtain the details of the target image from different perspectives. With reference to Figure 4 , the number, position, distance from the target object, and angle of the cameras in this embodiment can be adjusted according to the actual situation to ensure that the detailed information of different perspectives of the target image is presented. Among them, the same camera can simultaneously shoot with multiple exposure levels (such as low, medium, and high) at the same moment to capture the details of the target object under different lighting conditions.

[0103] S202. Based on the image fusion model obtained by the above-mentioned training method of the image fusion model, fuse the images of the target object taken at different exposure levels from the same perspective in the set of images to be fused, and obtain the intermediate fused images for each perspective.

[0104] With reference to Figure 5 , through the multi-camera system, an LDR image sequence with different exposure levels for each perspective can be obtained. Inputting the LDR image sequence into the image fusion model can fuse and obtain an HDR image, that is, the intermediate fused images for each perspective can be obtained to ensure the quality of the images obtained from each perspective. Among them, LDR (Low Dynamic Range) refers to low dynamic range, and its images or videos have limited brightness and color ranges, which are suitable for ordinary display devices. HDR (High Dynamic Range) refers to high dynamic range, which provides a wider brightness and color range and can present more details. Usually, a device supporting HDR is required to show the best effect.

[0105] S203. After registering the intermediate fusion images of each perspective, perform fusion to obtain the fusion image of the target object.

[0106] Cooperate with the reference Figure 5 , perform registration processing on the obtained intermediate fusion images of multiple perspectives, which can align the features in the images under different perspectives and ensure the quality of subsequent image fusion. Among them, in this embodiment, global registration and local registration methods can be used, and during the local registration process, an image registration model can be cooperated to complete pixel-level registration of the images.

[0107] In this embodiment, a sparse map can be generated during global registration of the images. Demonstratively, features of image samples can be extracted based on ORB (Oriented FAST and Rotated BRIEF), and key feature information can be screened based on the RANSAC (Random Sample Consensus) robust matching algorithm to obtain a sparse map.

[0108] During local registration, a pixel registration method based on pyramid optical flow can be used and the aforementioned sparse map can be used to optimize the optical flow field to improve the registration accuracy. The above registration method can solve problems such as distortion and artifacts that occur when the perspective difference is large, and performs excellently when processing three-dimensional objects or images with insufficient local features.

[0109] Among them, the intermediate fusion images of each perspective after registration can be fused based on the image fusion model obtained by the above training method of the image fusion model. In the HDR fusion step for image samples under different exposure levels, a weighted average fusion strategy can be adopted, and for the fusion of the intermediate fusion images of each perspective, a maximum value fusion strategy can be adopted. The reason is that for the HDR fusion of image samples under different exposure levels, an image with appropriate exposure needs to be obtained, so the weighted average fusion strategy is adopted. And for the fusion of the intermediate fusion images of each perspective, the optimal details, textures, and unique information in all perspectives need to be integrated to finally form a clear and detailed high-quality image, so the maximum value fusion strategy is adopted.

[0110] In this embodiment, by reusing the same image fusion model, flexible solutions are provided for different tasks at different stages of image fusion, effectively improving the overall efficiency of the system.

[0111] The image fusion method in an embodiment of the present invention overcomes the defect that the prior art cannot effectively utilize the unique detail information of different perspectives in multi-perspective image fusion, and proposes a novel multi-perspective image fusion method. Different from traditional methods that only focus on expanding the field of view or depth information extraction, this method maximally retains and enhances the detail information under different perspectives through carefully designed image acquisition, image registration, and image fusion strategies, thus significantly improving the quality and clarity of the final image.

[0112] Meanwhile, compared with the prior art, the image fusion method in the present invention makes targeted improvements and optimizations for multi-view scenarios, especially complex situations such as large parallax, through the cooperation of image fusion and image registration, further enhancing the authenticity and consistency of the fused image.

[0113] Refer to Figure 6 to introduce the training device of the image fusion model in an embodiment of the present invention. The image fusion model is used to fuse images of the same view or different views to generate a fused image. In this embodiment, the training device of the image fusion model includes an acquisition module 301, a construction module 302, and a determination module 303.

[0114] The acquisition module 301 is used to establish a training sample set, and the training sample set includes image samples of the target object taken at the same view or different views and a reference image of the target object; the construction module 302 is used to construct a fusion loss function, and the fusion loss function includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub-item represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-item represents the structural similarity error of the fused image relative to the reference image; the determination module 303 is used to train the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model.

[0115] In an optional embodiment, the acquisition module 301 is further used for preprocessing the image samples, specifically including: calculating the gradient map of the image samples respectively and normalizing; enhancing the information of the corresponding sample images based on the normalized gradient map.

[0116] In an optional embodiment, the image fusion model includes a self-attention module, and extracts the structural features of the image samples based on the self-attention module to generate corresponding attention weight maps, wherein the structural features include image edge features and / or texture features; when fusing the feature maps of image samples of the same view with different exposure degrees, the attention weight maps are used to perform weighted processing on them to generate a fused image.

[0117] In an optional embodiment, it further includes a preprocessing module, which is used to obtain initial image samples of the target object taken at the same view with different exposure degrees; map the initial image samples to a reference plane and calculate the image overlapping area; crop the initial image samples based on the image overlapping area to obtain the corresponding image samples.

[0118] In an alternative embodiment, the construction module 302 is further configured such that the basic loss sub-item is constructed by weighting a pixel loss sub-item and an image gradient loss sub-item; and / or, the structural similarity loss sub-item represents the similarity error between the fused image and the reference image in terms of brightness, contrast, and local structure; and / or, the fusion loss function is constructed by weighting the basic loss sub-item, the perceptual loss sub-item, and the structural similarity loss sub-item.

[0119] Referring to Figure 7 , an image fusion device according to an embodiment of the present invention is introduced. In this embodiment, the image fusion device includes an acquisition module 401, a fusion module 402, and a processing module 403.

[0120] The acquisition module 401 is configured to acquire a set of images to be fused, where the set of images to be fused includes images of a target object taken at different perspectives and with different exposure levels; the fusion module 402 is configured to fuse the images of the target object taken at the same perspective and with different exposure levels in the set of images to be fused based on the image fusion model trained by the above method to obtain intermediate fused images for each perspective; the processing module 403 is configured to register the intermediate fused images for each perspective and then fuse them to obtain a fused image of the target object.

[0121] In an alternative embodiment, the processing module 403 also fuses the registered intermediate fused images for each perspective based on the image fusion model obtained by the above training method.

[0122] In an alternative embodiment, the images of the target object taken at the same perspective and with different exposure levels in the set of images to be fused are fused using a weighted average fusion strategy; and the intermediate fused images for each perspective are fused using a maximum value fusion strategy.

[0123] As described above with reference to Figures 1 to 5 , the training method and the image fusion method of the image fusion model according to the embodiments of this specification are described. The details mentioned in the above description of the method embodiments also apply to the image fusion device of the embodiments of this specification. The above image fusion device can be implemented in hardware, or in software, or in a combination of hardware and software.

[0124] Referring to Figure 8 , a hardware structure diagram of an electronic device according to an embodiment of this specification is shown. As Figure 8 shown, the electronic device 50 may include at least one processor 51, a memory 52 (such as a non-volatile memory), a memory 53, and a communication interface 54, and at least one processor 51, the memory 52, the memory 53, and the communication interface 54 are connected together via an internal bus 55. At least one processor 51 executes at least one computer-readable instruction stored or encoded in the memory 52.

[0125] It should be understood that the computer-executable instructions stored in the memory 52, when executed, cause at least one processor 51 to perform the various operations and functions described above in connection with the various embodiments of this specification. Figures 1 to 5 described above.

[0126] In an embodiment of this specification, the electronic device 50 may include, but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and so on.

[0127] According to one embodiment, there is provided a program product such as a machine-readable medium. The machine-readable medium may have instructions (i.e., the elements implemented in software as described above), which, when executed by the machine, cause the machine to perform the various operations and functions described above in connection with the various embodiments of this specification. Figures 1 - 5 Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above-described embodiments is stored, and the computer or processor of the system or device is caused to read and execute the instructions stored in the readable storage medium.

[0128] In this case, the program code read from the readable medium itself can implement the functions of any one of the above-described embodiments, and thus the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0129] Examples of the readable storage medium include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as a CD-ROM, a CD-R, a CD-RW, a DVD-ROM, a DVD-RAM, a DVD-RW, a DVD-RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0130] Those skilled in the art should understand that the various embodiments disclosed above can be variously modified and changed without departing from the essence of the invention. Therefore, the protection scope of this specification should be defined by the appended claims.

[0131] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined according to requirements. The device structures described in the above embodiments can be physical structures or logical structures, that is, some units may be implemented by the same physical entity, or some units may be implemented separately by multiple physical entities, or some components in multiple independent devices may be jointly implemented.

[0132] In the above embodiments, the hardware units or modules can be implemented mechanically or electrically. For example, a hardware unit, module, or processor can include permanent dedicated circuits or logic (such as a dedicated processor, FPGA, or ASIC) to perform corresponding operations. The hardware unit or processor can also include programmable logic or circuits (such as a general-purpose processor or other programmable processors), which can be temporarily configured by software to perform corresponding operations. The specific implementation method (mechanical method, or dedicated permanent circuit, or temporarily configured circuit) can be determined based on cost and time considerations.

[0133] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of the claims. The term "exemplary" used throughout this specification means "serving as an example, instance, or illustration", and does not mean "preferred" or "advantageous" compared to other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.

[0134] The above description of the present disclosure is provided to enable any ordinary person skilled in the art to implement or use the present disclosure. Various modifications to the present disclosure are obvious to those of ordinary skill in the art, and the general principles corresponding to this article can also be applied to other variations without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.

[0135] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An image fusion method, characterized in that, Including: Obtain a set of images to be fused, where the set of images to be fused includes images of a target object captured at different perspectives with different exposure levels; Based on an image fusion model obtained by a training method of the image fusion model, fuse the images of the target object captured at the same perspective with different exposure levels in the set of images to be fused to obtain intermediate fusion images for each perspective; Register the intermediate fusion images for each perspective and then fuse them to obtain a fused image of the target object; Wherein, the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image, and the training method of the image fusion model includes: Establish a training sample set, which includes image samples of the target object captured at the same perspective or different perspectives and a reference image of the target object; Construct a fusion loss function, which includes at least two of a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term. Among them, the basic loss sub-term includes a pixel loss sub-term and / or an image gradient loss sub-term. The pixel loss sub-term represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub-term represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-term represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-term represents the structural similarity error of the fused image relative to the reference image; Train the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model.

2. The image fusion method according to claim 1, wherein The training method of the image fusion model further includes preprocessing the image samples, specifically including: Calculate the gradient map of each image sample and normalize it respectively; Perform information enhancement on the corresponding image sample based on the normalized gradient map.

3. The image fusion method according to claim 1, wherein The image fusion model includes a self-attention module, and the training method of the image fusion model further includes: Extract the structural features of the image sample based on the self-attention module to generate a corresponding attention weight map, where the structural features include image edge features and / or texture features; When fusing the feature maps of image samples of the same perspective or different perspectives, use the attention weight map to perform weighted processing on them to generate a fused image.

4. The image fusion method according to claim 1, wherein The training method of the image fusion model further includes: Obtain initial image samples of the target object captured at the same perspective or different perspectives; Map the initial image samples to a reference plane and calculate the image overlapping region; Crop the initial image samples based on the image overlapping region to obtain corresponding image samples.

5. The image fusion method according to claim 1, wherein The basic loss sub-term is constructed by weighting the pixel loss sub-term and the image gradient loss sub-term; And / or, the structural similarity loss sub-term represents the similarity error of the fused image relative to the reference image in terms of brightness, contrast, and local structure; And / or, the fusion loss function is constructed by weighting the basic loss sub-term, the perceptual loss sub-term, and the structural similarity loss sub-term.

6. The image fusion method according to claim 1, wherein Based on the image fusion model obtained by the training method of the image fusion model, fuse the registered intermediate fusion images for each perspective.

7. The image fusion method according to claim 1, characterized in that The method specifically includes: Fusing the images of the target object captured at different exposure levels from the same perspective in the image set to be fused by a weighted average fusion strategy; and Fusing the intermediate fusion images of each perspective by a maximum value fusion strategy.

8. An image fusion device, characterized in that, Comprising: An acquisition module, configured to acquire an image set to be fused, where the image set to be fused includes images of the target object captured at different exposure levels from different perspectives; A fusion module, configured to fuse the images of the target object captured at different exposure levels from the same perspective in the image set to be fused based on an image fusion model obtained by a training method of the image fusion model, to obtain intermediate fusion images of each perspective; A processing module, configured to register and then fuse the intermediate fusion images of each perspective to obtain a fused image of the target object; Wherein, the image fusion model is used to fuse images from the same perspective or different perspectives to generate a fused image, and the training method of the image fusion model includes: Establishing a training sample set, where the training sample set includes image samples of the target object captured from the same perspective or different perspectives and a reference image of the target object; Constructing a fusion loss function, where the fusion loss function includes at least two of a basic loss sub-term, a perceptual loss sub-term, and a structural similarity loss sub-term. Among them, the basic loss sub-term includes a pixel loss sub-term and / or an image gradient loss sub-term. The pixel loss sub-term represents the pixel error of the fused image relative to the reference image, and the image gradient loss sub-term represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-term represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-term represents the structural similarity error of the fused image relative to the reference image; Training the image fusion model based on the fusion loss function to determine the model parameters of the image fusion model.

9. An electronic device, characterized in that, Comprising: At least one processor; And A memory, where the memory stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the image fusion method according to any one of claims 1 to 7.

10. A machine-readable storage medium, characterized in that, It stores executable instructions, and when the instructions are executed, the machine executes the image fusion method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on self-attention mechanism

    CN111709902A

  • Neural network model training method, vehicle view generation method and vehicle

    CN115565155A