Training method of image fusion model, image fusion method and application

By constructing a fusion loss function of multiple loss subitems and adding a self-attention module, the shortcomings of the image fusion model in the prior art in terms of data dependence, computing efficiency, real-time, interpretability and detail retention are solved, and a higher quality image fusion effect is achieved.

CN119941531AActive Publication Date: 2025-05-06SUZHOU INS IMAGE SOFTWARE TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510432562.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the prior art, the image fusion method based on the deep learning model relies on a large amount of labeled data, has high computational complexity, high resource requirements, and is difficult to achieve real-time performance, and has insufficient retention of image authenticity and details.

Method used

By constructing a fusion loss function, including basic loss sub-item, perceived loss sub-item and structural similarity loss sub-item, the image fusion model is trained to improve the model's ability to restore image edges, details, textures and local structures. At the same time, a self-attention module is added to enhance the model's understanding and capture of key structural information and long-distance correlation.

Benefits of technology

The image fusion model's ability to restore image edges, details, textures and local structures is improved, ensuring the consistency of fusion results in brightness, contrast and local structure, and enhancing the overall structural fidelity and clarity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941531A_ABST
    Figure CN119941531A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of an image fusion model, an image fusion method and application. In the training method of the image fusion model, the image fusion model is used for fusing images of the same visual angle or different visual angles to generate a fused image, and the method comprises the following steps: establishing a training sample set, the training sample set comprises image samples shot for a target object at the same visual angle or different visual angles and a reference image of the target object; constructing a fusion loss function, wherein the fusion loss function comprises at least two of a basic loss subitem, a perception loss subitem and a structural similarity loss subitem; and training the image fusion model based on the fusion loss function, and determining model parameters of the image fusion model. According to the training method of the image fusion model, the fusion loss function is constructed, so that the image reduction capability of the model can be improved and enhanced, and the overall structure fidelity of the image is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a training method of an image fusion model, an image fusion method and an application thereof. Background Art

[0002] With the rapid development of photographic imaging technology and computational photography in recent years, as well as the continuous reduction of camera costs, multi-camera systems have gradually been widely used in various fields. Specifically, in the field of consumer electronics, the main applications of multi-camera systems are focused on expanding the field of view and enhancing the depth of field, while achieving computational photography effects such as portrait mode and zoom function by combining cameras with different focal lengths or apertures. In fields such as autonomous driving and robotics, multi-camera systems can organically integrate data from different fields of view, thereby effectively extracting depth information and enhancing the machine's perception of complex environments. In the fields of virtual reality (VR) and augmented reality (AR), multi-camera systems achieve reconstruction of 3D objects or scenes by combining images from different perspectives.

[0003] Image fusion technology is a technology that combines multiple images into one to present a subject that is closer to the real visual effect. In the prior art, the image fusion method based on the deep learning model can automatically learn to extract image features and then complete the image fusion. However, the above fusion method relies on a large amount of labeled data for training. In addition, the deep learning model has high computational complexity and high demand for resources, making it difficult to effectively implement and apply it in scenarios with high real-time requirements. In addition, when there are high requirements for image authenticity, the unexplainable characteristics of the deep learning model also have a potential impact on its reliability and acceptability. Finally, the deep learning model has poor sensitivity to structural and texture information, resulting in the inability to guarantee the clarity and details of the formed image.

[0004] Therefore, in response to the above technical problems, it is necessary to provide an image fusion model training method, an image fusion method and an application. Summary of the invention

[0005] The purpose of the present invention is to provide a training method, an image fusion method and an application of an image fusion model, which can solve the problems of the above-mentioned model in terms of data dependence, computational efficiency, real-time performance, interpretability and detail retention.

[0006] In order to achieve the above object, a specific embodiment of the present invention provides a training method for an image fusion model, and the technical solution is as follows: A training method for an image fusion model, wherein the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image, the method comprising: Establishing a training sample set, the training sample set comprising image samples taken of a target object at the same viewing angle or at different viewing angles and a reference image of the target object; Constructing a fusion loss function, the fusion loss function includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents a pixel error of a fused image relative to a reference image, the image gradient loss sub-item represents an image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents a high-frequency feature error of the fused image relative to the reference image; and the structural similarity loss sub-item represents a structural similarity error of the fused image relative to the reference image; The image fusion model is trained based on the fusion loss function to determine model parameters of the image fusion model.

[0007] In one or more embodiments of the present invention, the method further includes preprocessing the image sample, specifically including: Calculating and normalizing the gradient maps of the image samples respectively; Information enhancement is performed on the corresponding sample image based on the normalized gradient map.

[0008] In one or more embodiments of the present invention, the image fusion model includes a self-attention module, and the method further includes: Extracting structural features of the image sample based on the self-attention module and generating a corresponding attention weight map, wherein the structural features include image edge features and / or texture features; When fusing feature maps of image samples of the same or different perspectives, the attention weight map is used to perform weighted processing on them to generate a fused image.

[0009] In one or more embodiments of the present invention, the method further comprises: Obtaining initial image samples of the target object taken from the same or different perspectives; Mapping the initial image samples to a reference plane and calculating image overlap areas; The initial image samples are cropped based on the image overlapping area to obtain corresponding image samples.

[0010] In one or more embodiments of the present invention, the basic loss sub-item is constructed by weighting a pixel loss sub-item and an image gradient loss sub-item; And / or, the structural similarity loss sub-item represents the similarity error of the fused image relative to the reference image in brightness, contrast and local structure; And / or, the fusion loss function is constructed by weighting the basic loss sub-item, the perceptual loss sub-item and the structural similarity loss sub-item.

[0011] A specific embodiment of the present invention also provides an image fusion method, and the technical solution is as follows: An image fusion method, comprising: Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; Based on the image fusion model obtained by the above training method, the images to be fused are images captured at different exposures of the target object at the same viewing angle to obtain intermediate fused images at each viewing angle; The intermediate fused images of the various viewing angles are registered and then fused to obtain a fused image of the target object.

[0012] In one or more embodiments of the present invention, the intermediate fused images of the respective perspectives after registration are fused based on the image fusion model obtained by the above-mentioned training method.

[0013] In one or more embodiments of the present invention, the images to be fused are images captured at different exposures of the target object at the same viewing angle using a weighted average fusion strategy; and the intermediate fused images of each viewing angle are fused using a maximum fusion strategy.

[0014] A specific embodiment of the present invention further provides a training device for an image fusion model, and the technical solution is as follows: A training device for an image fusion model, wherein the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image, the device comprising: An acquisition module, used to establish a training sample set, wherein the training sample set includes image samples taken of a target object at the same viewing angle or at different viewing angles and a reference image of the target object; A construction module is used to construct a fusion loss function, wherein the fusion loss function includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents a pixel error of a fused image relative to a reference image, the image gradient loss sub-item represents an image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents a high-frequency feature error of the fused image relative to the reference image; and the structural similarity loss sub-item represents a structural similarity error of the fused image relative to the reference image; A determination module is used to train the image fusion model based on the fusion loss function to determine model parameters of the image fusion model.

[0015] A specific embodiment of the present invention further provides an image fusion device, and the technical solution is as follows: An image fusion device, comprising: An acquisition module, used for acquiring a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; A fusion module is used to fuse the images to be fused, which are images taken of the target object at different exposures at the same viewing angle, based on the above-mentioned image fusion model, to obtain intermediate fused images at each viewing angle; The processing module is used to register the intermediate fused images of each viewing angle and then fuse them to obtain a fused image of the target object.

[0016] A specific embodiment of the present invention further provides an electronic device, and the technical solution is as follows: An electronic device, comprising: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to execute the above-mentioned image fusion model training method or the above-mentioned image fusion method.

[0017] A specific embodiment of the present invention further provides a machine-readable storage medium, and the technical solution is as follows: A machine-readable storage medium stores executable instructions, which, when executed, enable the machine to perform the above-mentioned image fusion model training method or the above-mentioned image fusion method.

[0018] Compared with the prior art, the present invention has at least one of the following beneficial technical effects: 1. The training method of the image fusion model of the present invention can improve and strengthen the model's ability to restore the edges, details, textures and local structures of the image by constructing a fusion loss function, ensure the consistency of the fusion results in brightness, contrast and local structure, and enhance the overall structural fidelity of the image.

[0019] 2. The training method of the image fusion model of the present invention enhances edge details and texture features by calculating the gradient map of image samples and normalizing them, thereby improving the sensitivity of feature extraction to texture and structure, enabling it to better capture detail information and improving the accuracy of fusion.

[0020] 3. The training method of the image fusion model of the present invention enhances the model's ability to understand and capture key structural information and long-range correlations by adding a self-attention module, thereby being able to more accurately capture significant structural information in complex scenes, and significantly improving the clarity, texture fidelity and overall visual effect of the fused image.

[0021] 4. The image fusion method of the present invention can perform fusion processing and registration processing on multiple images containing the target object, and can obtain a high-quality 2D image of the target object with rich details. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 A diagram of an implementation environment of an image fusion model training method and an image fusion method in one embodiment of the present invention; Figure 2 is a flow chart of a method for training an image fusion model in one embodiment of the present invention; Figure 3 is a flow chart of an image fusion method in one embodiment of the present invention; Figure 4 A schematic diagram of the layout of a multi-camera system in one embodiment of the present invention; Figure 5 A schematic diagram of the steps of an image fusion method in one embodiment of the present invention; Figure 6 It is a module diagram of a training device for an image fusion model in one embodiment of the present invention; Figure 7 It is a module diagram of an image fusion device in one embodiment of the present invention; Figure 8 It is a hardware structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0025] With the rapid development of photographic imaging technology and computational photography in recent years, as well as the continuous reduction of camera costs, multi-camera systems have gradually been widely used in various fields. A multi-camera system is an imaging system composed of multiple cameras that can capture image data from different angles or perspectives to meet the needs of various application scenarios. In an embodiment of the present invention, it is expected to integrate the image features of images taken by multiple cameras, explore the unique detail advantages brought by different angles, and improve the image quality and clarity by fusing images from different perspectives, and achieve rich detail presentation.

[0026] With reference Figure 4 and Figure 5 In a specific scene example, a multi-camera system can shoot the target object from different angles, and different exposures can be used according to the settings when shooting the target object at each angle. The image fusion model first performs HDR fusion on images of different exposures at the same perspective to obtain an HDR picture, that is, the images at the same perspective are fused into one. Subsequently, the images from different perspectives are registered in steps, including global registration and local registration. Finally, the image fusion model (the model in the previous fusion step can be reused) is used to fuse the registered multi-perspective images to obtain a single high-definition and detailed 2D image.

[0027] Reference Figure 1 , showing a schematic diagram of an implementation environment provided by an exemplary embodiment of the present invention. The implementation environment includes a terminal and a server. The terminal and the server communicate data via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network. The image fusion method disclosed in the present invention can be executed by a server, and accordingly, the image fusion model can also be deployed in the server.

[0028] Alternatively, the terminal and the server may cooperate to run the image fusion method provided in the embodiment of the present application to complete the fusion of the images.

[0029] In a system architecture where the terminal can provide matching computing power, the image fusion method disclosed in the present application can also be directly executed by the terminal, and accordingly, the image fusion model can also be deployed on the terminal.

[0030] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0031] Reference Figure 2 , introduces a training method for an image fusion model in an embodiment of the present invention, wherein the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image. The training method for the image fusion model includes: S101, obtaining initial image samples of the target object taken from the same or different perspectives; mapping the initial image samples to a reference plane and calculating the image overlap area; cropping the initial image samples based on the image overlap area to obtain corresponding image samples. Specifically, the same camera can be used to synchronously shoot at multiple exposures (e.g., low, medium, and high) at the same time to capture the details of the target object under different lighting conditions.

[0032] In this embodiment, the initial image samples are mapped to the reference plane, and the formula for calculating the image overlap area is as follows:

[0033] in, : The image area captured by the i-th camera; f i : mapping function; K i : Camera internal reference; R i : Camera extrinsic rotation matrix; t i : Camera extrinsic translation vector; (x , y): 2D image pixel coordinate; : The corresponding area on the reference plane.

[0034] Compared with simple image capture, each image is first mapped to a unified reference plane, and then the overlapping area is uniformly calculated on this plane, which eliminates the error caused by perspective deformation and ensures the consistency of information in the subsequent fusion process.

[0035] S102: Establish a training sample set, where the training sample set includes image samples taken of a target object at the same viewing angle or different viewing angles and a reference image of the target object.

[0036] The image samples in the training sample set adopt the image samples obtained by cropping in the above steps, and in this embodiment, the image samples can also be further preprocessed, specifically including: calculating the gradient map of the image samples respectively and normalizing them; and enhancing the information of the corresponding sample image based on the normalized gradient map.

[0037] In image processing, the gradient map reflects the intensity and direction of the change in pixel value in the image (i.e., edge information). Using the gradient map for image enhancement can highlight the gradient information (edges, details) to improve the clarity, contrast or texture performance of the image. The gradient map can be calculated by a first-order or second-order differential operator, and the gradient amplitude can be weighted and fused with the original image to enhance the contrast of the edge area, or the enhancement intensity can be dynamically adjusted according to the local gradient. In this embodiment, the gradient map of the image can be calculated first, and the normalized gradient map can be added to the corresponding sample image to enhance the information of the sample image.

[0038] S103. Construct a fusion loss function, which includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents the pixel error of the fused image relative to the reference image, the image gradient loss sub-item represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents the high-frequency feature error of the fused image relative to the reference image; and the structural similarity loss sub-item represents the structural similarity error of the fused image relative to the reference image.

[0039] The fusion loss function may include three items: a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, or may include only two of them. For example, the fusion loss function may include only the basic loss sub-item and the perceptual loss sub-item, or may include only the perceptual loss sub-item and the structural similarity loss sub-item, or may include only the basic loss sub-item and the structural similarity loss sub-item. It is understandable that the basic loss sub-item may include only the pixel loss sub-item or the image gradient loss sub-item, or may include the pixel loss sub-item and the image gradient loss sub-item.

[0040] Specifically, the basic loss sub-item is constructed by weighting the pixel loss sub-item and the image gradient loss sub-item; the structural similarity loss sub-item represents the similarity error of the fused image relative to the reference image in brightness, contrast and local structure; the fusion loss function is constructed by weighting the basic loss sub-item, the perceptual loss sub-item and the structural similarity loss sub-item.

[0041] The above is the general concept of the fusion loss function. In this embodiment, the fusion loss function is further explained based on specific parameter settings, which is not a limitation of the fusion loss function in this embodiment. The details are as follows: In a specific embodiment, the fusion loss function includes three items: a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, and the basic loss function includes two items: a pixel loss sub-item and an image gradient loss sub-item.

[0042] The function of the basic loss sub-item is as follows:

[0043] in, I p and I g denote the predicted fusion image and the reference image respectively, i Represents the RGB channel index, H g , W g Respectively represent the height and width of the reference image, and λ represents the weight parameter. The above formula includes pixel loss sub-items and image gradient loss sub-items. The image gradient loss can constrain the gradient and improve the ability to restore edge and detail features.

[0044] The function of the perceptual loss sub-item is as follows:

[0045] in, f p and f g Represent the feature maps of the predicted fusion image and the reference image respectively, i represents the channel index of the feature map, C f 、H f and W f Represents the number of channels, height and width of the feature map respectively. This item enhances the model's ability to restore texture and local structure by comparing high-frequency features.

[0046] The function of the structural similarity loss sub-item is as follows:

[0047] in, I p and I g Represent the predicted fused image and the reference image respectively. The structural similarity loss sub-item optimizes the structural similarity index to ensure the consistency of the fusion result in brightness, contrast and local structure, further enhancing the overall structural fidelity of the image.

[0048] In summary, the fusion loss function of the model is expressed as:

[0049] in, α、β and γ is the weight coefficient of each loss term. In this embodiment, by constructing a fusion loss function, the clarity, texture fidelity and overall visual effect of the image are improved.

[0050] S104: training the image fusion model based on the fusion loss function to determine model parameters of the image fusion model.

[0051] Furthermore, the image fusion model includes a self-attention module, and the training method of the image fusion model also includes: S105. Extract structural features of image samples based on the self-attention module and generate a corresponding attention weight map, wherein the structural features include image edge features and / or texture features; when fusing feature maps of image samples of the same perspective or different perspectives, perform weighted processing on them using the attention weight map to generate a fused image.

[0052] In this embodiment, the features of the image can be extracted through the convolutional neural network, and the self-attention module uses its ability to capture long-distance dependencies to generate an attention weight map based on structural features such as image edges and textures. Among them, the ability of long-distance dependencies means that the self-attention module can obtain the structural features of image samples in a relatively large range.

[0053] In the feature fusion stage of image samples, the feature map obtained by the convolutional neural network is weighted through the attention weight map to enhance the feature expression of key areas. At the same time, the attention module can realize the model's comprehensive understanding of image details and overall structure, so that it can more accurately capture significant structural information in complex scenes, significantly improving the clarity, texture fidelity and overall visual effect of the fused image.

[0054] Reference Figure 3 , an image fusion method in an embodiment of the present invention is introduced. It should be noted in advance that the image fusion model involved in the image fusion method of this embodiment can be obtained based on the training method of the image fusion model described above.

[0055] The following specifically introduces an image fusion method in an embodiment of the present invention, which specifically includes: S201. Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures.

[0056] In this embodiment, a multi-camera system can be used to obtain images of a target object at different viewing angles and exposures. The multi-camera layout can obtain details of the target image at different viewing angles. Figure 4 The number, position, distance and angle of the cameras in this embodiment can be adjusted according to the actual situation to ensure that the detailed information of the target image from different perspectives is presented. The same camera can use multiple exposures (such as low, medium and high) to shoot synchronously at the same time to capture the details of the target object under different lighting conditions.

[0057] S202, based on the image fusion model obtained by the training method of the above-mentioned image fusion model, fuse the images to be fused, which are images of the target object taken at different exposures at the same viewing angle, to obtain intermediate fused images at each viewing angle.

[0058] With reference Figure 5 , through a multi-camera system, an LDR image sequence with different exposures can be obtained for each viewing angle. The LDR image sequence can be input into the image fusion model to fuse the HDR image, that is, the intermediate fusion image of each viewing angle can be obtained to ensure the quality of the image obtained from each viewing angle. Among them, LDR (Low Dynamic Range) refers to low dynamic range, and its image or video has limited brightness and color range, which is suitable for ordinary display devices. HDR (High Dynamic Range) refers to high dynamic range, which provides a wider brightness and color range, can present more details, and usually requires a device that supports HDR to show the best effect.

[0059] S203 , registering the intermediate fused images of each viewing angle and fusing them to obtain a fused image of the target object.

[0060] With reference Figure 5 , the obtained multi-view intermediate fusion image is registered, and the features in the images at different viewpoints can be aligned to ensure the quality of subsequent image fusion. In this embodiment, the pixel-level registration of the image can be completed by global registration and local registration, and the image registration model can be used in the local registration process.

[0061] In this embodiment, a sparse graph may be generated when globally registering images. Exemplarily, features of image samples may be extracted based on ORB (Oriented FAST and Rotated BRIEF), and key feature information may be screened based on a RANSAC (Random Sample Consensus) robust matching algorithm to obtain a sparse graph.

[0062] In local registration, the pixel registration method based on pyramid optical flow can be used to optimize the optical flow field using the aforementioned sparse graph to improve the registration accuracy. The above registration method can solve problems such as distortion and artifacts that occur when the viewing angle difference is large, and performs well when processing three-dimensional objects or images with insufficient local features.

[0063] Among them, the image fusion model obtained based on the training method of the above-mentioned image fusion model can be used to fuse the intermediate fusion images of each perspective after registration. A weighted average fusion strategy can be used in the HDR fusion step of image samples at different exposures, and a maximum fusion strategy can be used for the fusion of intermediate fusion images of each perspective. The reason is that the HDR fusion of image samples at different exposures requires obtaining images with appropriate exposure, so a weighted average fusion strategy is adopted. The fusion of intermediate fusion images of each perspective needs to integrate the best details, textures and unique information from all perspectives, and finally form a clear and detailed high-quality image, so a maximum fusion strategy is adopted.

[0064] In this embodiment, by reusing the same image fusion model, flexible solutions are provided for different tasks at different stages of image fusion, effectively improving the overall efficiency of the system.

[0065] The image fusion method in one embodiment of the present invention overcomes the defect that the prior art cannot effectively utilize the unique detail information of different perspectives in multi-perspective image fusion, and proposes a novel multi-perspective image fusion method. Unlike the traditional method that only focuses on expanding the field of view or extracting depth information, this method maximizes the retention and enhancement of detail information under different perspectives through carefully designed image acquisition, image registration and image fusion strategies, thereby significantly improving the quality and clarity of the final image.

[0066] At the same time, compared with the existing technology, the image fusion method in the present invention further improves the authenticity and consistency of the fused image by coordinating image fusion and image registration, and makes targeted improvements and optimizations for multi-view scenes, especially complex situations such as large parallax.

[0067] Reference Figure 6 , introduces a training device for an image fusion model in an embodiment of the present invention. The image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image. In this embodiment, the training device for the image fusion model includes an acquisition module 301, a construction module 302, and a determination module 303.

[0068] An acquisition module 301 is used to establish a training sample set, which includes image samples taken of a target object at the same or different perspectives and a reference image of the target object; a construction module 302 is used to construct a fusion loss function, which includes at least two of a basic loss sub-item, a perceptual loss sub-item and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents the pixel error of the fused image relative to the reference image, the image gradient loss sub-item represents the image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents the high-frequency feature error of the fused image relative to the reference image; the structural similarity loss sub-item represents the structural similarity error of the fused image relative to the reference image; a determination module 303 is used to train an image fusion model based on the fusion loss function and determine the model parameters of the image fusion model.

[0069] In an optional embodiment, the acquisition module 301 is further used to preprocess the image samples, specifically including: respectively calculating the gradient maps of the image samples and normalizing them; and performing information enhancement on the corresponding sample images based on the normalized gradient maps.

[0070] In an optional embodiment, the image fusion model includes a self-attention module, and extracts structural features of image samples based on the self-attention module to generate a corresponding attention weight map, wherein the structural features include image edge features and / or texture features; when fusing feature maps of image samples with different exposure levels at the same viewing angle, the attention weight map is used to perform weighted processing on them to generate a fused image.

[0071] In an optional embodiment, a pre-processing module is also included, which is used to obtain initial image samples taken of the target object at the same viewing angle with different exposures; map the initial image samples to a reference plane and calculate the image overlapping area; and crop the initial image samples based on the image overlapping area to obtain corresponding image samples.

[0072] In an optional embodiment, the construction module 302 is also used to construct the basic loss sub-item by weighting the pixel loss sub-item and the image gradient loss sub-item; and / or, the structural similarity loss sub-item represents the similarity error of the fused image relative to the reference image in brightness, contrast and local structure; and / or, the fusion loss function is constructed by weighting the basic loss sub-item, the perceptual loss sub-item and the structural similarity loss sub-item.

[0073] Reference Figure 7 , an image fusion device in an embodiment of the present invention is introduced. In this embodiment, the image fusion device includes an acquisition module 401 , a fusion module 402 and a processing module 403 .

[0074] The acquisition module 401 is used to acquire a set of images to be fused, wherein the set of images to be fused includes images of the target object shot at different perspectives and with different exposures; the fusion module 402 is used to fuse the images of the target object shot at the same perspective and with different exposures in the image set to be fused based on the image fusion model trained by the above method, to obtain intermediate fused images of each perspective; the processing module 403 is used to align the intermediate fused images of each perspective and then fuse them to obtain a fused image of the target object.

[0075] In an optional embodiment, the processing module 403 further fuses the registered intermediate fused images of each perspective based on the image fusion model obtained by the above-mentioned training method.

[0076] In an optional embodiment, the images to be fused are images captured at different exposures of the target object at the same viewing angle using a weighted average fusion strategy; and the intermediate fused images of the viewing angles are fused using a maximum fusion strategy.

[0077] As above Figures 1 to 5 , the training method of the image fusion model and the image fusion method according to the embodiment of this specification are described. The details mentioned in the above description of the method embodiment are also applicable to the image fusion device of the embodiment of this specification. The above image fusion device can be implemented by hardware, or by software, or by a combination of hardware and software.

[0078] Reference Figure 8 , shows a hardware structure diagram of an electronic device according to an embodiment of this specification. Figure 8 As shown, the electronic device 50 may include at least one processor 51, a memory 52 (e.g., a non-volatile memory), a memory 53, and a communication interface 54, and the at least one processor 51, the memory 52, the memory 53, and the communication interface 54 are connected together via an internal bus 55. At least one processor 51 executes at least one computer-readable instruction stored or encoded in the memory 52.

[0079] It should be understood that the computer executable instructions stored in the memory 52, when executed, cause at least one processor 51 to perform the above combined operations in various embodiments of the present specification. Figures 1 to 5 Describes the various operations and functions.

[0080] In the embodiments of the present specification, the electronic device 50 may include, but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.

[0081] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the above-mentioned elements implemented in the form of software), which, when executed by a machine, causes the machine to perform the above-mentioned combination of various embodiments of this specification. Figure 1-Figure 5 Specifically, a system or device equipped with a readable storage medium may be provided, on which a software program code implementing the functions of any of the above-mentioned embodiments is stored, and a computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.

[0082] In this case, the program code itself read from the machine-readable medium can realize the function of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0083] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0084] Those skilled in the art should understand that the various embodiments disclosed above can be modified and altered in various ways without departing from the essence of the invention. Therefore, the protection scope of this specification should be defined by the appended claims.

[0085] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some units may be implemented by the same physical client, or some units may be implemented by multiple physical clients, or some components in multiple independent devices may be implemented together.

[0086] In the above embodiments, the hardware unit or module can be realized by mechanical or electrical means. For example, a hardware unit, module or processor can include permanent dedicated circuit or logic (such as special processor, FPGA or ASIC) to complete the corresponding operation. The hardware unit or processor can also include programmable logic or circuit (such as general-purpose processor or other programmable processor), which can be temporarily set by software to complete the corresponding operation. Specific implementation (mechanical method or dedicated permanent circuit or temporary circuit) can be determined based on cost and time consideration.

[0087] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "exemplary" used throughout this specification means "used as an example, instance or illustration" and does not mean "preferred" or "having advantages" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, in order to avoid making the concepts of the described embodiments difficult to understand, well-known structures and devices are shown in block diagram form.

[0088] The above description of the present disclosure is provided to enable any person of ordinary skill in the art to implement or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles corresponding to the present disclosure may be applied to other variations without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the widest range of principles and novel features disclosed herein.

[0089] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A training method for an image fusion model, wherein the image fusion model is used to fuse images of the same perspective or different perspectives to generate a fused image, characterized in that: The method comprises: Establishing a training sample set, wherein the training sample set includes image samples taken of a target object at the same viewing angle or at different viewing angles and a reference image of the target object; Constructing a fusion loss function, the fusion loss function includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents a pixel error of a fused image relative to a reference image, the image gradient loss sub-item represents an image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents a high-frequency feature error of the fused image relative to the reference image; and the structural similarity loss sub-item represents a structural similarity error of the fused image relative to the reference image; The image fusion model is trained based on the fusion loss function to determine model parameters of the image fusion model.

2. The image fusion model training method according to claim 1, characterized in that: The method further includes preprocessing the image sample, specifically including: Calculating and normalizing the gradient maps of the image samples respectively; Information enhancement is performed on the corresponding sample image based on the normalized gradient map.

3. The image fusion model training method according to claim 1, characterized in that: The image fusion model includes a self-attention module, and the method further includes: Extracting structural features of the image sample based on the self-attention module and generating a corresponding attention weight map, wherein the structural features include image edge features and / or texture features; When fusing feature maps of image samples of the same or different perspectives, the attention weight map is used to perform weighted processing on them to generate a fused image.

4. The method for training an image fusion model according to claim 1, characterized in that: The method further comprises: Obtaining initial image samples of the target object taken from the same or different perspectives; Mapping the initial image samples to a reference plane and calculating image overlap areas; The initial image samples are cropped based on the image overlapping area to obtain corresponding image samples.

5. The image fusion model training method according to claim 1, characterized in that: The basic loss sub-item is constructed by weighting the pixel loss sub-item and the image gradient loss sub-item; And / or, the structural similarity loss sub-item represents the similarity error of the fused image relative to the reference image in brightness, contrast and local structure; And / or, the fusion loss function is constructed by weighting a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item.

6. An image fusion method, characterized in that: include: Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; Based on the image fusion model obtained by the training method according to any one of claims 1 to 5, the images to be fused are images captured at different exposures of the target object at the same viewing angle to obtain intermediate fused images at each viewing angle; The intermediate fused images of the various viewing angles are registered and then fused to obtain a fused image of the target object.

7. The image fusion method according to claim 6, characterized in that: Based on the image fusion model obtained by the training method according to any one of claims 1 to 5, the intermediate fused images of each perspective after registration are fused.

8. The image fusion method according to claim 6, characterized in that: The method specifically comprises: Using a weighted average fusion strategy, the images to be fused are images captured at different exposures of the target object at the same viewing angle; and The intermediate fusion images of the various perspectives are fused using a maximum fusion strategy.

9. A training device for an image fusion model, wherein the image fusion model is used to fuse images of the same or different perspectives to generate a fused image, characterized in that: The device comprises: An acquisition module, used to establish a training sample set, wherein the training sample set includes image samples taken of a target object at the same viewing angle or at different viewing angles and a reference image of the target object; A construction module is used to construct a fusion loss function, wherein the fusion loss function includes at least two of a basic loss sub-item, a perceptual loss sub-item, and a structural similarity loss sub-item, wherein the basic loss sub-item includes a pixel loss sub-item and / or an image gradient loss sub-item, the pixel loss sub-item represents a pixel error of a fused image relative to a reference image, the image gradient loss sub-item represents an image gradient error of the fused image relative to the reference image; the perceptual loss sub-item represents a high-frequency feature error of the fused image relative to the reference image; and the structural similarity loss sub-item represents a structural similarity error of the fused image relative to the reference image; A determination module is used to train the image fusion model based on the fusion loss function to determine model parameters of the image fusion model.

10. An image fusion device, characterized in that: include: An acquisition module, used for acquiring a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; A fusion module, configured to fuse the images to be fused, which are images of the target object captured at different exposures at the same viewing angle, based on the image fusion model according to any one of claims 1 to 5, to obtain intermediate fused images at each viewing angle; The processing module is used to register the intermediate fused images of each viewing angle and then fuse them to obtain a fused image of the target object.

11. An electronic device, characterized in that: include: at least one processor; as well as A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to execute the image fusion model training method as described in any one of claims 1 to 5 or the image fusion method as described in any one of claims 6 to 8.

12. A machine-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable the machine to perform the image fusion model training method as described in any one of claims 1 to 5 or the image fusion method as described in any one of claims 6 to 8.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on self-attention mechanism

    CN111709902A

  • Neural network model training method, vehicle view generation method and vehicle

    CN115565155A

  • Binocular vision system stained image restoration method

    CN117058523A

  • Target perception model training method and application, unmanned vehicle and storage medium

    CN117786520A

  • Multi-view industrial image fusion method based on cross attention mechanism

    CN119693764A