Image fusion method and image fusion device based on spatial pyramid pooling

Through the image fusion method based on spatial pyramid pooling, the spatial pyramid model is used for feature extraction and similarity score, and the optimal image fusion method is selected, which solves the problems of many noises and unclear edges in image fusion, and achieves higher quality image fusion.

CN120070205APending Publication Date: 2025-05-30CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221808.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, there are problems such as many noises and unclear edges during image fusion.

Method used

The image fusion method based on spatial pyramid pooling is adopted. By acquiring images of different image acquisition methods, the spatial pyramid model is used to extract and similarity score the image data, and the optimal image fusion method is selected to reduce feature loss and distortion.

Benefits of technology

It effectively reduces feature loss and distortion during image fusion process, improves the quality of the fusion image, and solves the problems of many noises and unclear edges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070205A_ABST
    Figure CN120070205A_ABST
Patent Text Reader

Abstract

The invention provides an image fusion method and device based on spatial pyramid pooling, and the method comprises the steps: obtaining the data of a plurality of images in different collection modes, and carrying out the processing of a plurality of image fusion methods, and obtaining a plurality of groups of fusion images; performing feature extraction on the fused image and the source image thereof by using a spatial pyramid model to form multiple groups of target features, and calculating a similarity score between the fused image and the source image based on each feature group; and selecting the fusion method corresponding to the maximum score value as a target fusion method to perform fusion processing on the new image. According to the method, the problems of more noise points and unclear edges of the fused image in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data communication technologies, and more particularly to the field of image processing technologies, and in particular to an image fusion method, an image fusion device, a computer-readable storage medium, and an image processing system based on spatial pyramid pooling. Background Art

[0002] In the field of digital image processing, image fusion technology, as a key component, has been significantly developed and widely applied. Its core is to fuse different expressions of the same scene from images from different channels, such as infrared images, visible light images, multi-focus images, etc., and finally generate an image containing all the useful information of the source images. The fused image can provide a wider observation range and richer details, which is very important for enhancing image quality, improving night monitoring effects, optimizing remote sensing applications, and in medical image analysis.

[0003] In the prior art, operations such as cropping and scaling during the image fusion process will cause problems such as image object cropping and shape distortion, that is, there is a loss of image features in the preprocessing process, resulting in problems of more noise and unclear edges in the fused image. Summary of the Invention

[0004] The main purpose of the present application is to provide an image fusion method, an image fusion device, a computer-readable storage medium, and an image processing system based on spatial pyramid pooling, so as to at least solve the problems of more noise points and unclear edges in the fused image in the prior art.

[0005] To achieve the above object, according to one aspect of the present application, an image fusion method based on spatial pyramid pooling is provided, including: obtaining images collected by different image acquisition methods to obtain a plurality of first image data, processing the first image data based on different image fusion methods to obtain a plurality of second image data, where the second image data corresponds one-to-one with the image fusion method, and any one of the second image data corresponds to all the first image data; using a spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data to obtain a plurality of target feature groups, where the target feature groups correspond one-to-one with the second image data; using the spatial pyramid model to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively according to each target feature group to obtain a first target score; determining the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and processing the image to be fused based on the target fusion method.

[0006] Optionally, before extracting features from each second image data and the first image data that are fused to form the second image data using a spatial pyramid model to obtain multiple target feature groups, it includes: obtaining a second target score, where the second target score is the relative subjective evaluation score corresponding to the second image data; using a preset training set as input data and the second target score as the output label to perform transfer training on a first target model to obtain a spatial pyramid model, and the first target model includes at least a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a third convolutional layer, and a spatial pyramid pooling layer.

[0007] Optionally, using the spatial pyramid model to calculate the similarity score between each target feature group and the corresponding second image data and the first image data that are fused to form the second image data respectively to obtain a first target score, includes: calculating the similarity score between the second image data and the corresponding first image data according to a preset formula to obtain a second target score: Score = ∑SSIM(A, B, F) / 128; where Score is the similarity score, SSIM is the calculation of the SSIM index, A and B are the features of the first image data, and F is the feature of the second image data; processing the second target score based on the scale corresponding to the second target score to obtain a third target score, and calculating the mean value based on the second target score to obtain the first target score.

[0008] Optionally, processing the second target score based on the scale corresponding to the second target score to obtain a third target score, includes: in the case where the scale corresponding to the second target score is equal to a preset value, determining the second target score as the third target score; in the case where the scale corresponding to the second target score is greater than the preset value, determining the minimum value of the second target score as the third target score under the current scale.

[0009] Optionally, processing the first image data using different image fusion methods to obtain multiple second image data, includes: processing the first image data through Laplacian pyramid fusion to obtain second image data; processing the first image data through dual-tree complex wavelet transform fusion to obtain second image data; processing the first image data through curvelet transform fusion to obtain second image data; processing the first image data through non-subsampled contourlet transform fusion to obtain second image data; processing the first image data through contrast image fusion to obtain second image data; processing the first image data through guided filter fusion to obtain second image data; processing the first image data through pulse coupled neural network fusion to obtain second image data; processing the first image data through convolutional neural network fusion to obtain second image data.

[0010] Optionally, after processing the first image data by different image fusion methods to obtain multiple second image data, the method further includes: sequentially rotating each second image data by multiple preset angles and recording the second image data after each rotation.

[0011] Optionally, a spatial pyramid model is used to extract features from each second image data and the first image data that is fused to form the second image data, obtaining multiple target feature groups, including: obtaining multiple preset scales, and respectively dividing the second image data and the corresponding first image data based on each preset scale to obtain multiple target image groups, where the target image groups correspond one-to-one with the preset scales; respectively performing feature extraction according to each target image group to obtain the target feature groups.

[0012] According to another aspect of the present application, there is provided an image fusion device based on spatial pyramid pooling. The device includes: a first acquisition unit, configured to acquire images collected by different image acquisition methods to obtain multiple first image data, and process the first image data based on different image fusion methods to obtain multiple second image data, where the second image data corresponds one-to-one with the image fusion methods, and any one of the second image data corresponds to all of the first image data; a first processing unit, configured to use a spatial pyramid model to extract features from each second image data and the first image data that is fused to form the second image data, obtaining multiple target feature groups, where the target feature groups correspond one-to-one with the second image data; a second processing unit, configured to use a spatial pyramid model to calculate the similarity scores between the corresponding second image data and the first image data that is fused to form the second image data respectively according to each target feature group, obtaining a first target score; a third processing unit, configured to determine the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and process the image to be fused based on the target fusion method.

[0013] According to still another aspect of the present application, there is provided a computer-readable storage medium. The computer-readable storage medium includes a stored program, where, when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods.

[0014] According to yet another aspect of the present application, there is provided an image processing system, including: one or more processors, a memory, and one or more programs, where, one or more programs are stored in the memory and are configured to be executed by one or more processors, and the one or more programs include those for executing any one of the methods.

[0015] Applying the technical solution of the present application in the above-mentioned image fusion method based on spatial pyramid pooling, first, images acquired by different image acquisition methods are obtained to get multiple first image data. The first image data is processed based on different image fusion methods to obtain multiple second image data. The second image data corresponds one-to-one with the image fusion method, and any one of the second image data corresponds to all the first image data. Then, a spatial pyramid model is used to extract features from each second image data and the first image data that are fused to form the second image data, obtaining multiple target feature groups. The target feature groups correspond one-to-one with the second image data. After that, the spatial pyramid model is used to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively, obtaining the first target score. Finally, the image fusion method corresponding to the maximum value of the first target score is determined as the target fusion method, and the image to be fused is processed based on the target fusion method. The present application compares the source image and the fused image at different scales based on the spatial pyramid pooling technology to ensure the authenticity of the similarity analysis between the fused image and the source image. Furthermore, the method of selecting the image fusion method based on the similarity analysis can minimize the feature loss and distortion of the image during the image fusion process, solving the problems of more noise points and unclear edges in the fused image in the prior art. Description of the Drawings

[0016] Figure 1 Shows a hardware structure block diagram of a mobile terminal for an image fusion method based on spatial pyramid pooling provided in an embodiment of the present application;

[0017] Figure 2 Shows a schematic flowchart of an image fusion method based on spatial pyramid pooling provided in an embodiment of the present application;

[0018] Figure 3 Shows a structure block diagram of an image fusion device based on spatial pyramid pooling provided in an embodiment of the present application.

[0019] Among them, the above-mentioned drawings include the following reference numerals:

[0020] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed Embodiments

[0021] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.

[0022] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of this application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0024] For the convenience of description, the following explains some nouns or terms related to the embodiments of this application:

[0025] Image Fusion: Image fusion refers to the process of taking the image data of the same target collected from multiple source channels, through image processing and computer technology, etc., to extract the beneficial information in each channel to the greatest extent, and finally synthesize high-quality images to improve the utilization rate of image information, improve the accuracy and reliability of computer interpretation, enhance the spatial resolution and spectral resolution of the original image, and facilitate monitoring.

[0026] Image Fusion Quality Assessment: Image fusion quality assessment mainly refers to the method of subjectively evaluating the quality of the fused image by relying on the human eye. This method is simple and intuitive, and can provide an intuitive and fast evaluation of obvious image information. Image objective quality assessment refers to evaluating the quality of image fusion according to various algorithms.

[0027] Structural Similarity Index: According to the principle of the human visual system, the human eye's perception of images is highly structured. Therefore, calculating the degree of loss of image structural information can measure the approximate perceptual image distortion.

[0028] Spatial Pyramid Pooling: Spatial Pyramid Pooling is an improved network based on convolutional neural network. Specifically, the feature map obtained in the conv5 layer is 256 layers, and spatial pyramid pooling is performed once for each layer.

[0029] Subjective Image Quality Assessment: Subjective image quality assessment is divided into two categories: absolute subjective assessment and relative subjective assessment methods. Absolute subjective assessment of images is to directly grade and score images according to visual perception without a standard reference. The relative subjective assessment method is to classify a batch of images from good to bad by observers in the presence of a standard image, compare them with each other to determine good or bad, and give corresponding scores.

[0030] As introduced in the background art, operations such as cropping and scaling in the prior art image fusion process will cause problems such as image object cropping and completion and shape distortion, resulting in loss of image features, and causing problems of more noise and unclear edges in the fused image. To solve the problems of more noise points and unclear edges in the fused image in the prior art, embodiments of the present application provide an image fusion method, an image fusion device, a computer-readable storage medium, and an image processing system based on spatial pyramid pooling.

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0032] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of an image fusion method based on spatial pyramid pooling according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown, or have a different configuration from

[0033] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the display method of device information in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0034] In this embodiment, an image fusion method based on spatial pyramid pooling running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0035] Figure 2 It is a flowchart of an image fusion method based on spatial pyramid pooling according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0036] Step S201, obtain images collected by different image acquisition methods to obtain a plurality of first image data, process the first image data based on different image fusion methods to obtain a plurality of second image data, the second image data corresponds to the image fusion method one by one, and any one of the second image data corresponds to all the first image data;

[0037] Specifically, first, obtain the first image data for the same scene from different image acquisition systems, such as visible light images, infrared images, etc. Then, use a variety of known image fusion algorithms, such as LP, DTCWT, DC, etc., to perform fusion processing on these source image data to generate multiple second image data, that is, fused images. Each fusion method generates a fused image, and all fused images correspond one-to-one with the source image data set.

[0038] Step S202: Use the spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data, obtaining multiple target feature groups, and the target feature groups correspond one-to-one with the second image data;

[0039] Then, input each fused image into the SPP model (spatial pyramid model). The model first extracts the feature map from the fused image, and then divides the feature map into three levels: global, medium scale, and small scale, that is, direct pooling, 2x2 sub-region pooling, and 4x4 sub-region pooling, respectively obtaining feature vectors of 1xN, 4xN, and 16xN, where N is the number of channels, such as 128. These feature vectors are then concatenated into a fixed-length 21xN feature vector for subsequent evaluation.

[0040] Step S203: Use the spatial pyramid model to calculate the similarity score between each second image data and the first image data that are fused to form the second image data according to each target feature group, obtaining the first target score;

[0041] Through the feature vectors extracted by SPP, calculate the SSIM index between the fused image and the source image respectively to obtain the fusion quality score at different scales. The similarity scores at different scales are used to reflect the retention of global information and local details of the fused image. The fusion quality scores at different scales are synthesized to obtain the overall score of each fusion method, that is, the above-mentioned first target score.

[0042] Step S204: Determine the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and process the image to be fused based on the target fusion method.

[0043] Select the fusion method with the highest score as the target fusion method, and process the new image data to be fused based on this method to obtain the optimal fused image quality.

[0044] It is understandable that in the field of image fusion, it is fundamental to obtain image data under different image acquisition modes. These data may come from different devices such as visible light cameras, infrared cameras, multi-focus cameras, etc., each capturing different features of the same scene. The goal of image fusion is to combine the advantages of these source images to produce a fused image with more comprehensive information and richer details. Before fusion, choosing a suitable fusion method is crucial to the quality of the final image. Different image fusion methods, such as Laplace pyramid fusion (LP), dual tree complex wavelet transform fusion (DTCWT), contrast image fusion (DC), etc., each have their own unique advantages and limitations, suitable for different scenarios and needs. Furthermore, in the above embodiment, the spatial pyramid pooling (SPP) technology is used, combined with the structural similarity (SSIM) index, to objectively evaluate the quality of the fused image, so as to intelligently select the optimal fusion method. Among them, the spatial pyramid pooling technology can process input images of any size. By dividing the image feature map into grids of different scales for pooling, multi-scale features can be extracted. This happens to solve the limitation that traditional convolutional neural networks (CNNs) need fixed-size inputs when processing images, allowing the model to capture image information more comprehensively, including global features and local details. The SSIM indicator is based on the characteristics of the human visual system. By comparing the similarity of brightness, contrast and structural information between the fused image and the source image, it provides a standard for measuring image distortion and can more accurately evaluate the quality of image fusion.

[0045] In a specific embodiment, taking infrared-visible light image fusion as an example: 10 pairs of infrared images and visible light images are selected from a known image database as source image data. The eight image fusion methods mentioned above are used to fuse these source images respectively to obtain 80 fused images. The SPP model is used to extract features from each fused image and the corresponding source image, and then the SSIM score between the source image and the fused image is calculated, including the global feature score Score1, the mid-scale feature score Score2, and the small-scale feature score Score3. Finally, the overall score Score of each fused image is obtained. The overall scores under different fusion methods are analyzed, and the fusion method with the highest score is identified, assuming it is the DTCWT fusion method. This means that the DTCWT method can produce fused images with the highest structural similarity and the best quality when processing infrared-visible light image fusion.

[0046] Through this embodiment, first, images acquired by different image acquisition methods are obtained to get a plurality of first image data. The first image data is processed based on different image fusion methods to obtain a plurality of second image data. The second image data corresponds one-to-one with the image fusion method, and any second image data corresponds to all the first image data. Then, a spatial pyramid model is used to extract features from each second image data and the first image data that are fused to form the second image data, obtaining a plurality of target feature groups. The target feature groups correspond one-to-one with the second image data. After that, the spatial pyramid model is used to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively, obtaining a first target score. Finally, the image fusion method corresponding to the maximum value of the first target score is determined as the target fusion method, and the image to be fused is processed based on the target fusion method. This application compares the source image and the fused image at different scales based on the spatial pyramid pooling technology to ensure the authenticity of the similarity analysis between the fused image and the source image. Furthermore, the method of selecting the image fusion method based on the similarity analysis can minimize the feature loss and distortion of the image during the image fusion process, solving the problems of more noise points and unclear edges in the fused image in the prior art.

[0047] In order to train the above-mentioned spatial pyramid model, in an optional implementation manner, before using the spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data to obtain a plurality of target feature groups, it includes:

[0048] Step S301, obtain a second target score, where the second target score is the relative subjective evaluation score corresponding to the second image data;

[0049] Before feature extraction using the SPP model, subjective evaluation data needs to be collected. One achievable way is to invite a certain number of observers to conduct subjective quality evaluation on the fused images to obtain the relative subjective evaluation scores (second target scores) of each fused image. Specifically, the source images and the fused images are sent to different staff terminals, and the received relative subjective evaluation scores from the staff are used as the target output labels for training the SPP model.

[0050] Step S302, use the preset training set as the input data, and use the second target score as the output label to perform transfer training on the first target model to obtain the spatial pyramid model. The first target model includes at least a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a third convolutional layer, and a spatial pyramid pooling layer.

[0051] Specifically, a pre-trained neural network model is selected as the first target model, which at least includes a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a third convolutional layer, and a spatial pyramid pooling layer. Using these second image data (i.e., fused images) obtained by the image fusion method as the input and the corresponding second target scores (relative subjective evaluation scores) as the output labels, transfer training is performed on the first target model, that is, by adjusting the network weights, enabling the model to predict the correct quality scores based on the features. The introduction of the spatial pyramid pooling layer allows the model to process input images of different sizes without causing image distortion or information loss. During the model fine-tuning process, the SPP layer divides the feature map into grids of different scales, performs pooling operations, and extracts multi-scale image features. These features will be concatenated into a fixed-length vector for subsequent classification or regression tasks, that is, the objective evaluation of image fusion quality.

[0052] It can be understood that in the image fusion quality evaluation method based on spatial pyramid pooling (SPP), using a neural network for feature extraction and fusion quality assessment is its core. In this process, in order to make the SPP model adapt to the actual application scenario, the pre-trained neural network model is trained through transfer learning. Transfer learning is a machine learning strategy, and its core idea is to pre-train a model in one domain (source domain) and then apply it to another domain (target domain). Through adjustment and fine-tuning, the model can be made to adapt to the requirements of the new task. In this application, based on the convolutional layer, pooling layer, etc. in the pre-trained model (the first target model), the spatial pyramid pooling layer is introduced, and then the model is fine-tuned using the labeled training data (i.e., the fused images and their corresponding relative subjective evaluation scores), enabling it to accurately understand the features of image fusion quality and make evaluations.

[0053] In a specific embodiment, taking infrared and visible light image fusion as an example: 50 observers were invited to subjectively evaluate the quality of the fused images generated by 8 different fusion methods. The evaluation criteria included factors such as image sharpness, detail retention, and noise suppression. The average score of each fused image would be used as the relative subjective evaluation score of the fused image. A model containing at least three convolutional layers and two pooling layers was selected from existing pre-trained models (such as VGG, ResNet, etc.) as the first target model. Using the collected fused images (the second image data) and the corresponding subjective evaluation scores (the second target scores), transfer training was performed on the model. During the training process, the weights of the model were adjusted so that the SPP layer could extract multi-scale features contributing to the quality evaluation from the fused images, and through regression training, the output of the model was made close to the quality scores given by the observers. After training was completed, the constructed spatial pyramid model was used to extract features from the new fused image data and then predict its quality score. The results showed that the SPP model trained through transfer learning could accurately predict the image quality of different fusion methods, and its predicted scores were highly consistent with the relative subjective evaluation scores of the observers, demonstrating the effectiveness of the model in image fusion quality evaluation.

[0054] Through the above solution, the spatial pyramid model can not only process input images of any size, avoiding information loss caused by image scaling in traditional methods, but also, through transfer learning and the guidance of relative subjective evaluation scores, the model can understand the subtle features of image fusion quality and provide objective and accurate scoring results.

[0055] In order to calculate the above first target score, in an alternative embodiment, step S203 includes:

[0056] Step S2031, calculate the similarity score between the second image data and the corresponding first image data according to a preset formula to obtain the second target score:

[0057] Score = ∑SSIM(A, B, F) / 128;

[0058] Where Score is the similarity score, SSIM is the calculation of the SSIM metric, A and B are the features of the first image data, and F is the feature of the second image data;

[0059] Specifically, first, for the fused images (the second image data) generated by each image fusion method, the SSIM metric was used for similarity scoring. Here, the SSIM metric was applied between the features (F) of the fused images and the features (A and B) of the source images, and their structural similarity was calculated through a preset formula. It can be understood that the source image features can be more than just A and B.

[0060] In one embodiment, SSIM can be compatible with three factors: brightness, contrast, and structure, and evaluate similarity by calculating the mean, variance, and covariance of local regions of the image. For features of different scales (1xN, 4xN, 16xN), the SSIM scores are calculated respectively to obtain the second target scores at different scales, and these scores reflect the similarity between the fused image and the source image at different levels of detail.

[0061] Step S2032: Process the second target scores based on the scales corresponding to the second target scores to obtain the third target scores, and calculate the mean value based on the second target scores to obtain the first target scores.

[0062] Specifically, subsequently, scale processing is performed on the second target scores, which means appropriate weighting or processing of the scores at different scales to reflect the importance of different scale features. For example, at the 1xN scale (global features), the average value of the scores may be directly used; at the 4xN scale (medium-scale features), the worst value of the scores may be focused on to ensure the quality of local regions of the image; at the 16xN scale (small-scale features), the worst scores are also focused on, but more weight may be given in the calculation because the human visual system is more sensitive to small-scale details of images. Finally, based on the processing results of the second target scores (the third target scores), the mean value of the scores at all scales is calculated to obtain the first target scores, which is an evaluation value of the image fusion quality that comprehensively combines global, medium-scale, and small-scale features.

[0063] It can be understood that image fusion quality evaluation is an important link in image fusion technology, which needs to accurately measure the similarity between the fused image and the source image to ensure that the fused image can retain the details and features of the source image. The structural similarity (SSIM) index is a widely used method for image quality evaluation. Based on the characteristics of the human visual system, it evaluates the degree of image distortion by comparing the brightness, contrast, and structural information of images. A more comprehensive image fusion quality evaluation system can be constructed. The SPP technology allows the model to capture global and local features at different scales, thereby evaluating image similarity in multiple dimensions.

[0064] In a specific embodiment, taking the fusion of infrared and visible light images as an example, the spatial pyramid pooling technology is used to extract features from the fused image (the second image data) and the source infrared image and visible light image (the first image data), respectively obtaining global features (1xN), medium-scale features (4xN), and small-scale features (16xN). Based on the features extracted by SPP, the SSIM metrics under the above different scale features are calculated respectively. When calculating, the features of the source images A and B are compared with the features of the fused image F to obtain the second target score corresponding to the corresponding scale. Scale processing is performed on the second target score. For example, for the 1xN scale, the average SSIM value of all 128-dimensional features is calculated (Score1); for the 4xN scale, the worst SSIM value is determined using the cask principle as the scale quality score (Score2); for the 16xN scale, the worst SSIM value is also determined (Score3). Finally, the average value of the scores under all scales is calculated: Score = (Score1 + Score2 + Score3) / 3 to obtain the first target score.

[0065] Through the above embodiments, not only the global structure of the image is considered, but also the medium-scale and small-scale features are deeply analyzed, which are crucial for the retention of image details and textures. Through scale processing (such as using the cask principle), it can be ensured that even if the fused image performs poorly in some local areas, it will not be ignored, thus improving the objectivity and accuracy of the evaluation. Moreover, the image fusion quality evaluation provided by the above solution not only technically avoids the defect of fixed-size scaling of the image, reduces information loss, but also, in terms of the formulation of evaluation metrics, is more in line with the characteristics of human visual perception, making the entire evaluation process more accurate and reliable.

[0066] In order to make the above similarity score adapt to different scales, in an alternative embodiment, the above step S2032 includes:

[0067] Step S20321, when the scale corresponding to the second target score is equal to the preset value, determining the second target score as the third target score;

[0068] Specifically, for each fused image, the spatial pyramid pooling (SPP) model generates feature vectors of different scales, such as 1xN, 4xN, and 16xN. The preset value can be set to a specific scale here, such as 1xN, representing the importance of global features. For the second target score under each scale, we adopt the following strategy for processing:

[0069] When the scale is equal to the preset value (for example, 1xN), the second target score is directly used as the third target score, which means that the score of the global features is not further adjusted and directly reflects the overall fusion quality of the image.

[0070] Step S20322, when the scale corresponding to the second target score is greater than a preset value, determine the minimum value of the second target score as the third target score at the current scale.

[0071] Specifically, when the scale is greater than the preset value (such as 4xN and 16xN), we focus on the local area of the image. Therefore, the minimum value in the second target score is used as the third target score at the current scale. This can ensure that even if there is a quality decline in the local area of the image, it can be reflected in the final score, avoiding the problem that a high global score masks local defects. It not only evaluates the global quality of the image but also carefully examines the local details of the image, ensuring the comprehensiveness and accuracy of the evaluation.

[0072] It can be understood that the "second target score" represents the structural similarity score (SSIM) between the fused image and the source image, while the "third target score" is the score after scale processing, which is used to more accurately reflect the fusion quality of the image at different scales. The preset value is used here to distinguish the importance of the scale. When the scale is equal to the preset value, it usually means that the features at this scale have an important impact on the overall image quality. Therefore, the second target score at this scale is directly used as the third target score. When the scale is greater than the preset value, it indicates that we focus on more local features, and the minimum value of the score is used to represent the fusion quality at the current scale because local defects in the image may have a significant impact on the overall visual effect. Even a very small area with a quality reduction may be a key factor in the evaluation.

[0073] Through the above embodiments, the quality performance of the fused image at different scales can be more accurately identified, especially when dealing with possible local defects in the image. In the embodiments, we found that when the global features of the fused image are well maintained, local quality defects may not significantly affect the global score, but the strategy of calculating the minimum value can ensure that the quality problems in these areas are also fully considered, thus improving the comprehensiveness and objectivity of the evaluation.

[0074] In order to achieve different ways of image fusion, in an optional embodiment, the above step S201 includes:

[0075] Step S2011, process the first image data through Laplacian pyramid fusion to obtain the second image data;

[0076] Specifically, Laplace Pyramid Fusion (LP): LP is a fusion technique based on decomposing an image into multiple resolution layers. It first decomposes the source image into a Laplace pyramid to obtain a series of low-frequency and high-frequency detail images. Then, these detail images are selectively fused, and finally, a fused image is obtained through an inverse pyramid reconstruction process. This method is suitable for detail preservation and contrast enhancement.

[0077] Step S2012: Process the first image data through dual-tree complex wavelet transform fusion to obtain second image data.

[0078] Specifically, Dual-tree Complex Wavelet Transform Fusion (DTCWT): DTCWT is a fusion technique based on wavelet transform. It decomposes an image through two independent wavelet trees, enabling more effective capture of the complex structure and edge details of the image. During fusion, details with higher local energy are selected to maintain the clarity and texture of the image.

[0079] Step S2013: Process the first image data through curve wave transform fusion to obtain second image data.

[0080] Specifically, Curve Wave Transform Fusion (CVT): CVT is a method for analyzing high-dimensional signals, especially suitable for processing curves and edges in images. It can accurately locate the structures in an image. During fusion, details with significant structures are preferentially retained, which has an advantage in fusing complex scenes.

[0081] Step S2014: Process the first image data through nonsubsampled contour wave transform fusion to obtain second image data.

[0082] Specifically, NonSubsampling Contour wave Transform Fusion (NSCT): The NSCT fusion method preserves the resolution of the image and avoids information loss caused by downsampling. It performs multi-resolution analysis to retain edges and textures and is suitable for occasions where the original resolution of the image needs to be maintained.

[0083] Step S2015: Process the first image data through contrast image fusion to obtain second image data.

[0084] Specifically, Contrast Image Fusion (DC): The DC method focuses on the contrast of the image and improves the image quality by enhancing the contrast and retaining details, which is suitable for the fusion of low-contrast images.

[0085] Step S2016, process the first image data through guided filter fusion to obtain second image data;

[0086] Specifically, Guided Filter Fusion (GFF): GFF is a filter-based fusion technique that uses a guided filter to smooth the image and retain the boundaries. When fusing, the filtered boundary information is combined with the source image to keep the boundaries clear.

[0087] Step S2017, process the first image data through pulse coupled neural network fusion to obtain second image data;

[0088] Specifically, Pulse Coupled Neural Network Fusion (PCNN): The PCNN fusion method simulates the pulse coupling mechanism of the human brain and selects the pixels to be fused through the interaction between networks. It is suitable for the fusion of complex images and can handle non-linear relationships.

[0089] Step S2018, process the first image data through convolutional neural network fusion to obtain second image data.

[0090] Specifically, Convolutional Neural Network Fusion (CNN): The CNN fusion method uses deep learning technology to automatically extract the features of the image and make fusion decisions by training a neural network. This method can handle large-scale data and complex patterns and is suitable for automated image fusion scenarios.

[0091] Through the above embodiments, each fusion method independently processes the source image data to generate its own fused image; these fused images will then be used for subsequent quality evaluation to determine which fusion method produces the highest-quality image. By combining multiple image fusion methods, it not only provides a basis for evaluating the quality of fused images but also makes it possible to intelligently select the most suitable image fusion method, thus achieving a better image fusion effect in specific application scenarios.

[0092] For data augmentation, in an alternative embodiment, after processing the first image data with different image fusion methods to obtain multiple second image data, the method further includes:

[0093] Step S401: Rotate each piece of second image data at multiple preset angles and record the second image data after each rotation.

[0094] Specifically, after using different image fusion methods to process the source image data to obtain multiple fused images (second image data), these fused images will be further rotated. The specific steps are as follows:

[0095] For each generated fused image, rotate it at multiple preset angles, such as 60°, 120°, 180°, 240°, 300°, and 360°. This not only includes common angles but also covers a complete rotation cycle, ensuring that all possible perspectives are considered. After each rotation, the rotated fused image data will be recorded. This step generates a more diverse image dataset for subsequent feature extraction and fusion quality evaluation.

[0096] It can be understood that by rotating the image, the imaging conditions under multiple perspectives can be simulated, enabling the image fusion quality evaluation model to maintain a high level of accuracy and stability when facing images with unknown or changing angles. This technique is based on the concept of data augmentation in deep learning, which believes that by increasing the diversity of model training data, the robustness and adaptability of the model can be improved, making it more reliable in practical applications.

[0097] Through the above embodiments, the accuracy and stability of image fusion quality evaluation can be significantly improved. Specifically, this method can bring the following effects: improving the generalization ability of the model: The rotated images provide more diverse inputs for the model, enabling the model to learn the impact of angle changes on image fusion quality, and thus being able to more accurately evaluate the quality of images with angle changes in practical applications. Enhancing robustness and invariance: Image rotation enhances the model's robustness to angle changes, making the evaluation results independent of a specific perspective, which is particularly important for application scenarios that require image fusion and quality assessment at non-fixed angles. Optimizing the selection of fusion methods: By analyzing the quality of fused images at different rotation angles, the advantages and disadvantages of various fusion methods can be more comprehensively understood, helping users intelligently select the most suitable fusion method in practical applications to meet the image fusion requirements at different angles. In summary, the data augmentation technique of image rotation provides angle invariance and robustness for the deep learning model of image fusion and quality evaluation, helps optimize the selection of fusion methods, and improves the overall performance of image fusion applications.

[0098] In an alternative embodiment, in order to extract the features of the source image and the fused image, the above step S202 includes:

[0099] Step S2021: Obtain multiple preset scales, and respectively divide the second image data and the corresponding first image data based on each preset scale to obtain multiple target image groups, where each target image group corresponds to a preset scale one by one;

[0100] Specifically, before using the SPP model for feature extraction, we first need to define multiple preset scales. These scales are usually a series of grid sizes, such as 1x1, 2x2, 4x4, 8x8, etc., which will be used to divide the fused image (second image data) and its corresponding source image (first image data). Each preset scale corresponds to an image partition of a different size, thereby achieving multi-scale extraction of image features.

[0101] Step S2022: Respectively perform feature extraction according to each target image group to obtain a target feature group.

[0102] Specifically, for each target image group under each preset scale (i.e., image partitions at different scales), the SPP model will perform corresponding feature extraction operations. Specifically, the model will perform a max pooling operation on the image feature map processed by pyramid pooling at each scale to obtain feature representations at different scales. These feature representations will be concatenated into a feature vector of a fixed length, which will be used as the input for subsequent quality evaluation.

[0103] It can be understood that the Spatial Pyramid Pooling (SPP) technique is widely used in deep learning to process input images of different sizes. Its core lies in being able to convert an image feature map of any size into a vector of a fixed length, thereby achieving unified representation and processing of images. In the evaluation of image fusion quality, the SPP technique also plays an important role. It can capture the features of the fused image and the source image at different scales, and thus provide a more comprehensive and robust quality assessment.

[0104] Through the above embodiments, the multi-scale feature extraction scheme using the SPP technique can not only provide a more comprehensive evaluation of image fusion quality, but also enhance the robustness and generalization ability of the model, which has important value for optimizing the selection of image fusion methods and improving the overall performance of image fusion applications.

[0105] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0106] The embodiments of the present application also provide an image fusion device based on spatial pyramid pooling. It should be noted that the image fusion device based on spatial pyramid pooling in the embodiments of the present application can be used to execute the image fusion method based on spatial pyramid pooling provided by the embodiments of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0107] The following introduces the image fusion device based on spatial pyramid pooling provided by the embodiments of the present application.

[0108] Figure 3 It is a structural block diagram of the image fusion device based on spatial pyramid pooling according to the embodiments of the present application. As Figure 3 shown, the device includes:

[0109] A first acquisition unit 10, configured to acquire images collected by different image acquisition methods to obtain a plurality of first image data, and process the first image data based on different image fusion methods to obtain a plurality of second image data. The second image data corresponds one-to-one with the image fusion method, and any one of the second image data corresponds to all the first image data;

[0110] A first processing unit 20, configured to extract features from each second image data and the first image data that are fused to form the second image data by using a spatial pyramid model to obtain a plurality of target feature groups, and the target feature groups correspond one-to-one with the second image data;

[0111] A second processing unit 30, configured to calculate the similarity score between the corresponding second image data and the first image data that are fused to form the second image data respectively according to each target feature group by using a spatial pyramid model to obtain a first target score;

[0112] A third processing unit 40, configured to determine the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and process the image to be fused based on the target fusion method.

[0113] Through this embodiment, the first acquisition unit acquires images collected by different image acquisition methods to obtain a plurality of first image data, processes the first image data based on different image fusion methods to obtain a plurality of second image data, and the second image data corresponds one-to-one with the image fusion methods. Any second image data corresponds to all the first image data. The first processing unit uses a spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data, obtaining a plurality of target feature groups, and the target feature groups correspond one-to-one with the second image data. The second processing unit uses the spatial pyramid model to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively, obtaining a first target score. The third processing unit determines the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and processes the images to be fused based on the target fusion method. This application compares the source image and the fused image at different scales based on the spatial pyramid pooling technology to ensure the authenticity of the similarity analysis between the fused image and the source image. Furthermore, the method of selecting the image fusion method based on the similarity analysis can minimize the feature loss and distortion of the image during the image fusion process, solving the problems of more noise points and unclear edges in the fused image in the prior art.

[0114] In an optional implementation manner, in order to train the above spatial pyramid model, the above device further includes:

[0115] A second acquisition unit, configured to acquire a second target score before using the spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data, obtaining a plurality of target feature groups, where the second target score is the relative subjective evaluation score corresponding to the second image data;

[0116] A training unit, configured to use a preset training set as input data, use the second target score as an output label, and perform transfer training on a first target model to obtain a spatial pyramid model, where the first target model at least includes a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a third convolutional layer, and a spatial pyramid pooling layer.

[0117] In an optional implementation manner, in order to calculate the above first target score, the above second processing unit includes:

[0118] A first calculation module, configured to calculate the similarity score between the second image data and the corresponding first image data according to a preset formula, obtaining a second target score:

[0119] Score = ∑SSIM(A, B, F) / 128;

[0120] Among them, Score is the similarity score, SSIM is the SSIM metric calculation, A and B are the features of the first image data, and F is the feature of the second image data;

[0121] The first processing module is used to process the second target score based on the scale corresponding to the second target score to obtain a third target score, and calculate the mean value based on the second target score to obtain a first target score.

[0122] In order to make the above similarity score adapt to different scales, in an optional implementation manner, the above first processing module includes:

[0123] The first determination sub-module is used to determine the second target score as the third target score when the scale corresponding to the second target score is equal to the preset value;

[0124] The second determination sub-module is used to determine the minimum value of the second target score as the third target score at the current scale when the scale corresponding to the second target score is greater than the preset value.

[0125] In order to implement image fusion in different ways, in an optional implementation manner, the above first acquisition unit includes:

[0126] The second processing module is used to process the first image data through Laplacian pyramid fusion to obtain the second image data;

[0127] The third processing module is used to process the first image data through dual-tree complex wavelet transform fusion to obtain the second image data;

[0128] The fourth processing module is used to process the first image data through curvelet transform fusion to obtain the second image data;

[0129] The fifth processing module is used to process the first image data through non-subsampled contourlet transform fusion to obtain the second image data;

[0130] The sixth processing module is used to process the first image data through contrast image fusion to obtain the second image data;

[0131] The seventh processing module is used to process the first image data through guided filter fusion to obtain the second image data;

[0132] The eighth processing module is used to process the first image data through pulse-coupled neural network fusion to obtain the second image data;

[0133] The ninth processing module is used to process the first image data through convolutional neural network fusion to obtain the second image data.

[0134] For data augmentation, in an alternative embodiment, the above-mentioned apparatus further includes:

[0135] A fourth processing unit, configured to process the first image data by different image fusion methods to obtain multiple second image data, and then sequentially rotate each second image data by multiple preset angles and record the second image data after each rotation.

[0136] To extract the features of the source image and the fused image, in an alternative embodiment, the above-mentioned first processing unit includes:

[0137] A first acquisition module, configured to acquire multiple preset scales, and respectively divide the second image data and the corresponding first image data based on each preset scale to obtain multiple target image groups, where the target image groups correspond to the preset scales one by one;

[0138] Specifically, before using the SPP model for feature extraction, we first need to define multiple preset scales. These scales are usually a series of grid sizes, such as 1x1, 2x2, 4x4, 8x8, etc., which will be used to divide the fused image (second image data) and its corresponding source image (first image data). Each preset scale corresponds to a different-sized image partition, thereby realizing multi-scale extraction of image features.

[0139] A second acquisition module, configured to perform feature extraction according to each target image group respectively to obtain a target feature group.

[0140] The above-mentioned image fusion apparatus based on spatial pyramid pooling includes a processor and a memory. The above-mentioned first acquisition unit, first processing unit, second processing unit, third processing unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned each module is located in different processors in any combination form.

[0141] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the similarity between the fused image and the source image can be improved.

[0142] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0143] An embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned image fusion method based on spatial pyramid pooling.

[0144] An embodiment of the present invention provides a processor. The processor is used to run a program. When the program runs, it executes the above-mentioned image fusion method based on spatial pyramid pooling.

[0145] An embodiment of the present invention provides an image processing system. The image processing system includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the steps of the above-mentioned image fusion method based on spatial pyramid pooling.

[0146] This application also provides a computer program product. When executed on a data processing device, it is adapted to execute a program initialized with at least the steps of the above-mentioned image fusion method based on spatial pyramid pooling.

[0147] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0148] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0149] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the specified function in one or more flows Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the specified function in one or more blocks.

[0150] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the specified function in one or more flows Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the specified function in one or more blocks.

[0151] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified function in one or more flows Figure 1 or more flows and / or blocks Figure 1 or a means for implementing the specified function in one or more blocks.

[0152] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0153] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0154] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0155] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0156] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0157] 1), The image fusion method based on spatial pyramid pooling of the present application. First, images acquired by different image acquisition methods are obtained to get multiple first image data. The first image data are processed based on different image fusion methods to obtain multiple second image data. The second image data correspond one-to-one with the image fusion methods, and any one of the second image data corresponds to all the first image data. Then, a spatial pyramid model is used to extract features from each second image data and the first image data that are fused to form the second image data, obtaining multiple target feature groups. The target feature groups correspond one-to-one with the second image data. After that, the spatial pyramid model is used to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively, obtaining a first target score. Finally, the image fusion method corresponding to the maximum value of the first target score is determined as the target fusion method, and the image to be fused is processed based on the target fusion method. The present application compares the source image and the fused image at different scales based on the spatial pyramid pooling technology to ensure the authenticity of the similarity analysis between the fused image and the source image. Furthermore, the method of selecting the image fusion method based on the similarity analysis can minimize the feature loss and distortion of the image during the image fusion process, solving the problems of more noise points and unclear edges in the fused image in the prior art.

[0158] 2), The image fusion device based on spatial pyramid pooling of the present application. The first acquisition unit acquires images acquired by different image acquisition methods to get multiple first image data. The first image data are processed based on different image fusion methods to obtain multiple second image data. The second image data correspond one-to-one with the image fusion methods, and any one of the second image data corresponds to all the first image data. The first processing unit uses a spatial pyramid model to extract features from each second image data and the first image data that are fused to form the second image data, obtaining multiple target feature groups. The target feature groups correspond one-to-one with the second image data. The second processing unit uses the spatial pyramid model to calculate the similarity scores between the corresponding second image data and the first image data that are fused to form the second image data respectively, obtaining a first target score. The third processing unit determines the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and processes the image to be fused based on the target fusion method. The present application compares the source image and the fused image at different scales based on the spatial pyramid pooling technology to ensure the authenticity of the similarity analysis between the fused image and the source image. Furthermore, the method of selecting the image fusion method based on the similarity analysis can minimize the feature loss and distortion of the image during the image fusion process, solving the problems of more noise points and unclear edges in the fused image in the prior art.

[0159] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An image fusion method based on spatial pyramid pooling, characterized in that: include: Acquire images acquired by different image acquisition methods to obtain a plurality of first image data, process the first image data based on different image fusion methods to obtain a plurality of second image data, wherein the second image data correspond to the image fusion methods one by one, and any one of the second image data corresponds to all the first image data; Using a spatial pyramid model, extracting features from each second image data and the first image data fused to form the second image data, to obtain a plurality of target feature groups, wherein the target feature groups correspond to the second image data one by one; Using the spatial pyramid model, respectively calculating a similarity score between the second image data corresponding to each of the target feature groups and the first image data fused to form the second image data, to obtain a first target score; The image fusion method corresponding to the maximum value of the first target score is determined as the target fusion method, and the image to be fused is processed based on the target fusion method.

2. The method according to claim 1, characterized in that Before extracting features from each second image data and the first image data fused to form the second image data using the spatial pyramid model to obtain a plurality of target feature groups, the method includes: Obtaining a second target score, where the second target score is a relative subjective evaluation score corresponding to the second image data; The preset training group is used as input data, and the second target score is used as the output label to perform migration training on the first target model to obtain the spatial pyramid model, wherein the first target model at least includes a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a third convolutional layer and a spatial pyramid pooling layer.

3. The method according to claim 1, characterized in that The spatial pyramid model is used to calculate the similarity score between the second image data corresponding to each target feature group and the first image data fused to form the second image data to obtain a first target score, including: The similarity score between the second image data and the corresponding first image data is calculated according to a preset formula to obtain a second target score: Score=ΣSSIM(A,B,F) / 128; Wherein, Score is the similarity score, SSIM is the SSIM index calculation, A and B are the features of the first image data, and F is the feature of the second image data; The second target score is processed based on a scale corresponding to the second target score to obtain a third target score, and a mean is calculated based on the second target score to obtain the first target score.

4. The method according to claim 3, characterized in that Processing the second target score based on the scale corresponding to the second target score to obtain a third target score includes: When the scale corresponding to the second target score is equal to a preset value, determining the second target score as the third target score; When the scale corresponding to the second target score is greater than the preset value, the minimum value of the second target score is determined as the third target score under the current scale.

5. The method according to claim 1, characterized in that The first image data is processed based on different image fusion methods to obtain a plurality of second image data, including: Processing the first image data by Laplacian pyramid fusion to obtain the second image data; Processing the first image data by dual-tree complex wavelet transform fusion to obtain the second image data; Processing the first image data by curvelet transform fusion to obtain the second image data; Processing the first image data by non-subsampled contourlet transform fusion to obtain the second image data; Processing the first image data by contrast image fusion to obtain the second image data; Processing the first image data by guided filter fusion to obtain the second image data; Processing the first image data by pulse coupled neural network fusion to obtain the second image data; The first image data is processed by convolutional neural network fusion to obtain the second image data.

6. The method according to claim 1, characterized in that After the first image data is processed by different image fusion methods to obtain a plurality of second image data, the method further includes: The second image data are rotated sequentially by a plurality of preset angles and the second image data after each rotation is recorded.

7. The method according to claim 1, characterized in that The spatial pyramid model is used to extract features from each second image data and the first image data fused to form the second image data, to obtain a plurality of target feature groups, including: Acquire multiple preset scales, and divide the second image data and the corresponding first image data based on each of the preset scales to obtain multiple target image groups, where the target image groups correspond to the preset scales one by one; Feature extraction is performed respectively according to each of the target image groups to obtain the target feature group.

8. An image fusion device based on spatial pyramid pooling, characterized in that: The device comprises: A first acquisition unit is used to acquire images acquired by different image acquisition methods to obtain a plurality of first image data, and process the first image data based on different image fusion methods to obtain a plurality of second image data, wherein the second image data corresponds to the image fusion methods one by one, and any one of the second image data corresponds to all the first image data; a first processing unit, configured to extract features from each second image data and the first image data fused to form the second image data using a spatial pyramid model, to obtain a plurality of target feature groups, wherein the target feature groups correspond to the second image data one by one; a second processing unit, configured to respectively calculate a similarity score between the second image data corresponding to the second image data and the first image data fused to form the second image data according to each of the target feature groups using the spatial pyramid model, to obtain a first target score; The third processing unit is used to determine the image fusion method corresponding to the maximum value of the first target score as the target fusion method, and process the image to be fused based on the target fusion method.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.

10. An image processing system, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include methods for executing any one of claims 1 to 7.