Transform domain heterogeneous image fusion method, device and equipment and storage medium

By optimizing the parameters of the PCNN model and introducing a saliency mechanism, combined with an improved bilateral filtering method, the problem of low fusion quality between ToF images and visible light images was solved, achieving adaptive termination and better image segmentation results, thus improving the image fusion quality in orchard environments.

CN115937646BActive Publication Date: 2025-11-18GANSU AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310024673.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-11-18
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

Existing pulse-coupled neural network (PCNN) models suffer from problems such as insufficient empirical parameter setting, inability to adaptively terminate, and susceptibility to oversegmentation in image fusion, resulting in low fusion quality of ToF images and visible light images in orchard environments.

Method used

A saliency-guided pulse-coupled neural network (SMPCNN) is used to perform multi-scale decomposition and fusion of ToF images and visible light images by optimizing parameters such as iteration termination information, dynamic threshold amplification factor, link channel feedback term and link strength, combined with an improved bilateral filtering method.

Benefits of technology

It achieves adaptive termination and reasonable image segmentation, and has a gray-scale clustering illumination mechanism, which improves the fusion quality of ToF images and visible light images, especially the fusion effect of heterogeneous images in orchard environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937646B_ABST
    Figure CN115937646B_ABST
Patent Text Reader

Abstract

The application provides a transform domain heterogeneous image fusion method, device and equipment and storage medium, wherein the method comprises: performing multi-scale decomposition on a time-of-flight image and a visible light image to obtain a first low-frequency sub-band image, a second low-frequency sub-band image, a plurality of first high-frequency sub-band images and a plurality of second high-frequency sub-band images; then performing a significant function, a KL divergence and a momentum driven multi-objective artificial bee colony algorithm on the first low-frequency sub-band image to obtain parameter information of an SMPCNN model; next, performing fusion on the low-frequency sub-band image by the SMPCNN model to obtain a low-frequency fusion image; performing bilateral filtering fusion on the high-frequency sub-band image to obtain a high-frequency fusion image; and finally performing multi-scale decomposition inverse transformation on the low-frequency fusion image and the high-frequency fusion image to obtain a final fusion image. The method has the advantages of gray clustering lighting mechanism and same gray attribute priority lighting, and can achieve better image fusion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image fusion technology, and more specifically, to a transform domain heterogeneous image fusion method, apparatus, device, and storage medium. Background Technology

[0002] Orchard apple images acquired solely through visible light suffer from uneven illumination, fruit occlusion, and overlapping, leading to reduced effectiveness in target recognition and posing numerous challenges to automated harvesting. Currently, a key research area in the vision of harvesting robots is the fusion of heterogeneous image information from visible light images and time-of-flight (ToF) images. ToF images, with their characteristics of invariant illumination, spatial hierarchy, near-infrared sensing, and reliable data analysis, can overcome the limitations of a single visible light data source and have been widely applied in agricultural harvesting research.

[0003] Current heterogeneous image information fusion technologies mainly include spatial domain fusion models and transform domain fusion models. Among them, the most typical transform domain model is the multi-scale non-subsampled shearlet transform (NSST) method, and among the NSST methods, the pulse coupled neural network (PCNN) is widely used.

[0004] However, current PCNN models suffer from drawbacks such as empirical parameter setting, inability to adaptively terminate, and susceptibility to oversegmentation. Furthermore, the different imaging mechanisms between ToF images acquired by a binocular acquisition system in an orchard environment and visible light heterogeneous images also lead to low fusion quality. Therefore, improving the image fusion quality in PCNN models has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this application is to address the shortcomings of the prior art by providing a transform domain heterogeneous image fusion method, apparatus, device, and storage medium to solve the problem of low image fusion quality in the PCNN model in the prior art.

[0006] To achieve the above objectives, the technical solution adopted in this application is as follows:

[0007] In a first aspect, this application provides a transform domain heterogeneous image fusion method, the method comprising:

[0008] Acquire time-of-flight images and visible light images of the target scene, and determine a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image;

[0009] Based on the first low-frequency sub-band image, the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network is determined. The parameter information includes: iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor.

[0010] The pulse-coupled neural network guided by the saliency mechanism determines the low-frequency fusion image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image;

[0011] An improved bilateral filtering method is used to fuse the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images to obtain a high-frequency fused image;

[0012] The fused image of the target scene is determined based on the low-frequency fused image and the high-frequency fused image.

[0013] Optionally, determining the first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determining the second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image, includes:

[0014] The time-of-flight image is decomposed into multiple scales to obtain the first low-frequency sub-band image and multiple first high-frequency sub-band images;

[0015] The visible light image is decomposed into multiple scales to obtain the second low-frequency sub-band image and multiple second high-frequency sub-band images.

[0016] Optionally, determining the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency subband image includes:

[0017] The iteration termination information is determined based on the saliency function and the first low-frequency sub-band image;

[0018] The dynamic threshold amplification factor is obtained by calculating the relative entropy divergence of the first low-frequency sub-band image.

[0019] Momentum-driven multi-target artificial bee colony calculation is performed on the first low-frequency sub-band image to obtain the link channel feedback term, link strength, and dynamic threshold attenuation factor.

[0020] Optionally, determining the iteration termination information based on the saliency function and the first low-frequency subband image includes:

[0021] Based on the saliency function, the first low-frequency sub-band image is segmented into multiple illuminations to obtain multiple ignition segmentation maps. The iteration termination information is determined based on the multiple ignition segmentation maps.

[0022] Optionally, the pulse-coupled neural network guided by the saliency mechanism determines the low-frequency fused image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image, including:

[0023] The dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold attenuation factor are input as input parameters into the saliency mechanism-guided pulse-coupled neural network. The iteration termination information is used as the iteration termination condition of the saliency mechanism-guided pulse-coupled neural network. The saliency mechanism-guided pulse-coupled neural network fuses the first low-frequency sub-band image and the second low-frequency sub-band image based on the input parameters and the iteration termination condition to obtain the low-frequency fused image.

[0024] Optionally, the step of fusing the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images using an improved bilateral filtering method to obtain a high-frequency fused image includes:

[0025] An improved bilateral filtering method is used to fuse the multiple first high-frequency sub-band images and the multiple second high-frequency sub-band images using a spatial neighborhood Gaussian function and a high-frequency component gray value similarity Gaussian function to obtain the high-frequency fused image.

[0026] Optionally, determining the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image includes:

[0027] The low-frequency fused image and the high-frequency fused image are subjected to multi-scale decomposition inverse transformation to obtain the fused image of the target scene.

[0028] Secondly, this application provides a transform domain heterogeneous image fusion apparatus, the apparatus comprising:

[0029] The acquisition module is used to: acquire time-of-flight images and visible light images of a target scene, and determine a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image;

[0030] The determination module is used to determine the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency subband image. The parameter information includes: iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor.

[0031] A low-frequency determination module is used to determine a low-frequency fused image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image by a pulse-coupled neural network guided by the saliency mechanism.

[0032] The high-frequency determination module is used to fuse the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images using an improved bilateral filtering method to obtain a high-frequency fused image;

[0033] The image determination module is used to determine the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image.

[0034] Optionally, the acquisition module is specifically used for:

[0035] The time-of-flight image is decomposed into multiple scales to obtain the first low-frequency sub-band image and multiple first high-frequency sub-band images;

[0036] The visible light image is decomposed into multiple scales to obtain the second low-frequency sub-band image and multiple second high-frequency sub-band images.

[0037] Optionally, the determining module is specifically used for:

[0038] The iteration termination information is determined based on the saliency function and the first low-frequency sub-band image;

[0039] The dynamic threshold amplification factor is obtained by calculating the relative entropy divergence of the first low-frequency sub-band image.

[0040] Momentum-driven multi-target artificial bee colony calculation is performed on the first low-frequency sub-band image to obtain the link channel feedback term, link strength, and dynamic threshold attenuation factor.

[0041] Optionally, the determining module is further specifically used for:

[0042] Based on the saliency function, the first low-frequency sub-band image is segmented into multiple illuminations to obtain multiple ignition segmentation maps. The iteration termination information is determined based on the multiple ignition segmentation maps.

[0043] Optionally, the low-frequency determination module is specifically used for:

[0044] The dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold attenuation factor are input as input parameters into the saliency mechanism-guided pulse-coupled neural network. The iteration termination information is used as the iteration termination condition of the saliency mechanism-guided pulse-coupled neural network. The saliency mechanism-guided pulse-coupled neural network fuses the first low-frequency sub-band image and the second low-frequency sub-band image based on the input parameters and the iteration termination condition to obtain the low-frequency fused image.

[0045] Optionally, the high-frequency determination module is specifically used for:

[0046] An improved bilateral filtering method is used to fuse the multiple first high-frequency sub-band images and the multiple second high-frequency sub-band images using a spatial neighborhood Gaussian function and a high-frequency component gray value similarity Gaussian function to obtain the high-frequency fused image.

[0047] Optionally, the image determination module is specifically used for:

[0048] The low-frequency fused image and the high-frequency fused image are subjected to multi-scale decomposition inverse transformation to obtain the fused image of the target scene.

[0049] Thirdly, this application provides an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the transform domain heterogeneous image fusion method described above.

[0050] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the transform domain heterogeneous image fusion method described above.

[0051] The beneficial effects of this application are as follows: By optimizing parameters such as iteration termination information, dynamic threshold amplification coefficient, link channel feedback term, link strength, and dynamic threshold attenuation factor based on the first low-frequency sub-band image, the resulting saliency mechanism-guided PCNN model combines the saliency mechanism and the PCNN clustering segmentation mechanism. This model can achieve adaptive termination and reasonable image segmentation, and has the advantages of gray-level clustering illumination mechanism and priority illumination of the same gray-level attribute. Thus, it can achieve better image fusion effect for low-frequency sub-band images. By using an improved bilateral filtering method to fuse high-frequency sub-band images, and determining the fused image of the target scene based on the obtained high-frequency fused image and low-frequency fused image, the fusion quality of ToF image and visible light heterogeneous source image can be improved. Attached Figure Description

[0052] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A schematic diagram illustrating an application scenario provided by an embodiment of this application is shown;

[0054] Figure 2A flowchart of a transform domain heterogeneous image fusion method provided in an embodiment of this application is shown;

[0055] Figure 3 This document illustrates a flowchart of a multi-scale decomposition of an image according to an embodiment of this application.

[0056] Figure 4 This document illustrates a flowchart of a method for determining parameter information according to an embodiment of this application.

[0057] Figure 5 This paper shows a schematic diagram of the structure of an SMPCNN model provided in an embodiment of this application;

[0058] Figure 6 This paper presents a flowchart illustrating yet another transform domain heterogeneous image fusion method provided in an embodiment of this application.

[0059] Figure 7 This paper shows a schematic diagram of the structure of a transform domain heterogeneous image fusion device provided in an embodiment of this application;

[0060] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0062] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0063] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0064] Heterogeneous image fusion techniques mainly include spatial domain fusion models and transform domain fusion models. Spatial domain models, because they study spatiotemporal pixel grayscale values, have the disadvantage of not easily identifying the texture and boundary features of the source image, thus limiting their application scope. The most typical transform domain model is the NSST method. Compared to other multi-scale transforms such as pyramid transform, discrete wavelet transform (DWT), contourlet transform, and nonsubsampled contourlet transform (NSCT), the NSST transform has multi-scale and multi-directional characteristics, the decomposition result has translation invariance, and it overcomes the spectral aliasing problem, resulting in fusion results with better edge and detail information.

[0065] PCNN models can fully utilize local pixel information and have been used in the research of high-frequency and low-frequency fusion rules in transform domain fusion models. Current research on PCNN parameter settings and fusion rule improvements mainly falls into four categories: Improving link strength: In Locally Nonsubsampled Shearlet Transform (LNSST), an adaptive dual-channel pulse-coupled neural network with three link strengths is used. Improved average gradient and modified Laplacian operator are used as adaptive link strengths, and an L2-norm-based optimization model is used to merge output coefficients, effectively fusing images with large spectral differences. Improving the input domain: The responses of PCNN neurons are adjusted based on the basic global information of the source images. After multi-scale decomposition, the fusion rule is based on the number of pixel responses. Low-frequency subbands are fused using the basic PCNN model, while high-frequency subbands are fused using a complex Weighted Modified Spatial Frequency (WMSF) excitation, which better preserves detail information. Parameter adaptation: The adaptive Dual PCNN model can simultaneously reflect information from two source images, simplifying many peripheral parameters, adaptively setting link strengths, and improving fusion accuracy. We utilize swarm intelligence algorithms to optimize parameters, employing a nature-inspired optimal feature selection method based on ant colony optimization to reduce the complexity of PCNN fusion of infrared and visible light images. We use NSCT to independently decompose the intensity-hue saturation of images, use PCNN to fuse high-frequency subband images and low-frequency images, and use a hybrid frog-jumping algorithm to optimize the PCNN network parameters.

[0066] However, traditional spatial domain fusion methods create fusion models in the grayscale space of images, which have the disadvantage of not being able to easily find the texture and boundary features of the source image.

[0067] The PCNN model has drawbacks such as empirical parameter setting, inability to adaptively terminate, and susceptibility to oversegmentation. During the ignition process, it ignores the impact of image changes and fluctuations on the results, leading to pixel artifacts, region blurring, and unclear edges. Furthermore, the different imaging mechanisms between ToF and visible light heterogeneous images acquired by the binocular acquisition system in the orchard environment also result in low fusion quality.

[0068] Therefore, improving the quality of image fusion in PCNN models has become an urgent problem to be solved.

[0069] To address the aforementioned problems, this application proposes a transform domain heterogeneous image fusion method, which can be applied to scenarios involving the heterogeneous fusion of ToF images and visible light images, such as... Figure 1 The diagram shown is an application scenario illustration provided in this application. The electronic device first acquires a ToF image and a visible light image in the same scene. Through the transform domain heterogeneous image fusion method of this application, a fused image of the ToF image and the visible light image in the scene can be obtained.

[0070] This application introduces a saliency mechanism into the PCNN model and optimizes some parameters in the PCNN model, which can better highlight the confidence region in the ToF image, thereby achieving a better image fusion effect.

[0071] Next, combine Figure 2 This paper describes the transform domain heterogeneous image fusion method of this application. The execution subject of this method can be an electronic device, such as... Figure 2 As shown, the method includes:

[0072] S201: Acquire time-of-flight images and visible light images of the target scene, and determine the first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine the second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image.

[0073] Optionally, the time-of-flight image can be a ToF image, which can be a ToF image and a visible light image of the same scene at the same time. The visible light image can be, for example, an RGB image.

[0074] Optionally, the first low-frequency sub-band image can be the low-frequency component of the ToF image, and the first high-frequency sub-band image can be the high-frequency component of the ToF image. The low-frequency component of the ToF image has the characteristics of extracting targets at a certain distance and separating the background.

[0075] Optionally, the second low-frequency sub-band image can be the low-frequency component of the visible light image, and the second high-frequency sub-band image can be the high-frequency component of the visible light image.

[0076] For example, NSST decomposition can be performed on the ToF image and the visible light image respectively to obtain the low-frequency subband image and the high-frequency subband image of the ToF image and the visible light image respectively.

[0077] S202: Determine the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency subband image. The parameter information includes: iteration termination information, dynamic threshold amplification coefficient, link channel feedback term, link strength, and dynamic threshold decay factor.

[0078] Optionally, when processing a scene, the saliency mechanism can automatically process the region of interest while selectively ignoring the region of no interest. The saliency mechanism-guided pulse-coupled neural network can process key regions in the image. Taking an orchard scene as an example, the key region in the image can be the fruit region.

[0079] Optionally, the saliency mechanism-guided pulse coupled neural network (SMPCNN) can be a saliency mechanism-guided PCNN model that combines the saliency mechanism, saliency function and PCNN clustering segmentation mechanism. It has the advantages of gray-level clustering lighting mechanism and priority lighting of the same gray-level attribute, and is suitable for heterogeneous image fusion in complex orchard environments.

[0080] It is worth noting that the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network may not be limited to the iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength and dynamic threshold decay factor given in this application. Other parameter information can be set according to empirical values.

[0081] The PCNN model comprises a feedback input domain, a coupling link domain, and a pulse generation domain, which can be described by the mathematical equations shown in equations (1)-(5). PCNN has the characteristic of grouping similar pattern features into one class based on the principle of similarity clustering and capture characteristics, and has a clustered point lighting segmentation mechanism.

[0082] F ij (n)=I ij (1)

[0083] L ij (n)=exp(-α L )L ij (n-1)+V L ∑ kl W ij,kl Y kl(n-1) (2)

[0084] U ij (n)=F ij (n)(1+βL ij (n)) (3)

[0085] θ ij (n)=exp(-α θ )θ ij (n-1)+V θ Y ij (n) (4)

[0086] Y ij (n)=step(U ij (n)-θ ij (n)) (5)

[0087] Where, I ij It is the external stimulus to the neuron, represented by the grayscale value of the input image; F ij (n) is the feedback input domain, L ij (n) is the link input field; W ij,kl U represents the link coefficient; β represents the link strength, which determines the weight of the coupling link channel; ij (n) represents the internal state signal of the model; θ ij V is the dynamic threshold of the neuron. θ V L α is the dynamic threshold amplification factor, which controls the increase in the threshold value after neuron activation; L and α θ These factors respectively determine the decay rate of the link channel feedback term and the dynamic threshold; Y ij (n) represents the current neuron's impulse output, which is the response result after comparing the internal activity term with the dynamic threshold in the impulse generator. When U is satisfied... ij (n)>θ ij When (n) is reached, the ignition condition will be met, and the output Y will be achieved. ij (n) = 1. Step represents the step function, whose output is 0 or 1, and n represents the nth neuron in the image.

[0088] S203: The pulse-coupled neural network guided by the saliency mechanism determines the low-frequency fused image based on parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image.

[0089] Optionally, the low-frequency fused image can be obtained by fusing the low-frequency components of the ToF image and the visible light image.

[0090] Optionally, the parameter information of the SMPCNN model can be the iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength and dynamic threshold decay factor determined in step S202 above, as well as other parameter information in the PCNN model that are set empirically, such as feedback input domain, link input domain, model internal state signal, etc.

[0091] Optionally, a weighted average fusion criterion can be used to fuse the first low-frequency sub-band image and the second low-frequency sub-band image to obtain a low-frequency fused image.

[0092] For example, suppose the first low-frequency subband image of the ToF image is represented as The second low-frequency subband image of the visible light image is represented as follows: The calculation method for low-frequency fused images can be expressed as shown in equation (6).

[0093]

[0094] S204: Multiple first high-frequency sub-band images and multiple second high-frequency sub-band images are fused using an improved bilateral filtering method to obtain a high-frequency fused image.

[0095] Optionally, an improved bilateral filtering method can be used to fuse the high-frequency components of the ToF image (first high-frequency sub-band image) and the high-frequency components of the visible light image (second high-frequency sub-band image) to obtain a high-frequency fused image of the ToF image and the visible light image.

[0096] It is worth noting that bilateral filtering is a local, nonlinear, and non-iterative technique. In bilateral filtering, high-frequency fusion rules can be introduced to measure the similarity between the ToF image and the visible light image at the corresponding positions of the high-frequency components after decomposition. The improved bilateral filtering method in this application can be a bilateral filtering method that introduces a Gaussian function of spatial neighborhood and a Gaussian function of high-frequency component gray value similarity.

[0097] S205: Determine the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image.

[0098] It is worth noting that the fused image of the target scene determined by the low-frequency fused image and the high-frequency fused image can effectively solve the problem of low target brightness and poor clarity caused by backlighting, and can better highlight the target area in the ToF image, thereby achieving a better visual effect.

[0099] For example, the low-frequency fused image and the high-frequency fused image can be subjected to the inverse NSST transform to obtain the fused image of the target scene.

[0100] In this embodiment, by optimizing parameters such as iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold attenuation factor based on the first low-frequency sub-band image, the resulting saliency mechanism-guided PCNN model combines the saliency mechanism and the PCNN clustering segmentation mechanism. This model can achieve adaptive termination and reasonable image segmentation, and has the advantages of gray-level clustering illumination mechanism and priority illumination of the same gray-level attribute. Thus, it can achieve better image fusion effect for the low-frequency sub-band image. By using an improved bilateral filtering method to fuse the high-frequency sub-band image, and determining the fused image of the target scene based on the obtained high-frequency fused image and low-frequency fused image, the fusion quality of the ToF image and the visible light heterogeneous source image can be improved.

[0101] The following describes the steps for determining the first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and for determining the second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image. Figure 3 As shown, the above step S201 includes:

[0102] S301: Perform multi-scale decomposition on the time-of-flight image to obtain the first low-frequency sub-band image and multiple first high-frequency sub-band images.

[0103] Optionally, multi-scale decomposition can be performed on the ToF image using a multi-scale nonsubsampled shearlet transform (NSST).

[0104] For example, the NSST transform domain decomposition method can be used to perform a nonsubsampled pyramid filter bank (NSP) u-level transform on the two precisely registered heterogeneous images to obtain one low-frequency sub-band image and u high-frequency sub-band images, achieving translation invariance. The high-frequency sub-band images are then further decomposed using a shearlet filter bank (SF) at the v-level in multiple directions to form 2... v High-frequency sub-bands in each direction.

[0105] S302: Perform multi-scale decomposition on the visible light image to obtain a second low-frequency sub-band image and multiple second high-frequency sub-band images.

[0106] Optionally, the method for multi-scale decomposition of visible light images can be the same as the method for multi-scale decomposition of ToF images, which will not be elaborated here.

[0107] It is worth noting that traditional spatial domain fusion methods create fusion models in the image grayscale space, which have the disadvantage of not easily identifying the texture and boundary features of the source image. Compared with traditional spatial domain fusion, the NSST transform domain decomposition method produces subbands of the same size as the source image, exhibiting high sparsity, accurate representation of fusion information, and multi-scale and multi-directional characteristics. The decomposition results are translation-invariant and overcome the spectral aliasing problem, resulting in fusion results with better edge and detail information.

[0108] The following is a description of the steps involved in determining the parameter information of the pulse-coupled neural network guided by the saliency mechanism based on the first low-frequency subband image. Figure 4 As shown, the above includes:

[0109] S401: Determine the iteration termination information based on the saliency function and the first low-frequency subband image.

[0110] Optionally, the PCNN model can perform multiple iterative segmentations on the first low-frequency sub-band image to present an ignition segmentation state. The two ignition segmentation maps at time t and time t+1 are correlated, but are not related to the ignition segmentation map at an earlier time. Based on this, the ignition segmentation map at a time interval of two time intervals can be defined as a first-order Markov situation.

[0111] The mathematical representation of the significance function will be explained next. For example, after ignition and splitting at time t, the model is in state s. u Under the given conditions, at time t+1, after ignition and splitting, the model transitions to state s. v The probability is defined as the significance one-step transition probability, expressed as equation (7).

[0112] p uv =p(S t+1 =s v |S t =s u )=p(s v |s u ),s u ,s v ∈S (7)

[0113] In a significant first-order Markov situation, the model is in any state s u Transition to state s under condition ∈S v The average uncertainty when ∈S is defined as the significant conditional entropy, expressed as Equation (8).

[0114]

[0115] The overall uncertainty of the sequence formed by the ignition segmentation diagram in a significant first-order Markov situation is defined as the entropy of the significant first-order Markov source, expressed as Equation (9).

[0116]

[0117] The amount of information transmitted in the model state transition of the ignition segmentation diagram at different times is defined as the significant first-order Markov mutual information, expressed as Equation (10).

[0118] I(U;V)=H(U)-H(U|V) (10)

[0119] PCNN performs multiple iterations of segmentation, and the ignition segmentation images at two time intervals have significant feature differences, representing the maximum information transmission rate, which is numerically expressed as the maximum mutual information. Since mutual information has a maximum value under certain conditions, the saliency function can be numerically defined as the saliency first-order Markov mutual information. Substituting the above equations (8) and (9) into equation (10), we can obtain the mathematical representation of the saliency function, namely equation (11).

[0120]

[0121] The above saliency function is used as the basis for terminating the PCNN model iteration, expressed as equation (12), where δ can be a user-defined parameter. For a low-frequency ignition segmentation map, the larger the first-order Markov mutual information of the saliency, the better the consistency within the region.

[0122] I saliency (u,v)>δ (12)

[0123] S402: Calculate the relative entropy divergence of the first low-frequency sub-band image to obtain the dynamic threshold amplification factor.

[0124] Relative entropy is also known as Kullback-Leibler divergence (KLD), information divergence, and information gain.

[0125] KL divergence is a measure of the asymmetry of the difference between two probability distributions P and Q. It measures the number of extra bits required to encode, on average, samples from P using Q-based coding. Typically, P represents the true distribution of the data, and Q represents the theoretical distribution, model distribution, or approximate distribution of P.

[0126] In Time-of-Flight (ToF) images, the target region often appears as a region with high grayscale values ​​and a normal distribution. Two fire-segmentation images are used, with states s... u ∈S and state s v The probability distribution p(s) corresponding to ∈S u ) and p(s vThe Kullback-Leibler divergence between the two is calculated to measure the dynamic threshold amplification factor of the PCNN model, expressed as (13). This formula is used to measure the similarity between the probability distributions of the two ignition segmentation maps. Since the more similar the probability distributions of the two ignition segmentation maps are, the smaller the dynamic threshold amplification factor will be, which will enable the PCNN model to ignite and output when the target region tends to stabilize during continuous iteration.

[0127]

[0128] S403: Perform momentum-driven multi-target artificial bee colony calculation on the first low-frequency subband image to obtain the link channel feedback term, link strength, and dynamic threshold decay factor.

[0129] The Artificial Bee Colony Algorithm (ABC) is a swarm intelligence optimization algorithm that simulates the characteristics of bee colonies. It has the advantages of strong global optimization ability, few parameters, high accuracy, and strong robustness. However, its optimization strategy has the defects of singularity and randomness, which makes the algorithm suffer from problems such as premature convergence and convergence stagnation.

[0130] Optionally, momentum-driven multi-objective artificial bee colony computation can be a momentum-driven multi-objective artificial bee colony algorithm obtained by introducing the concept of momentum in deep learning on the basis of the ABC algorithm.

[0131] The following section explains the principle of the momentum-driven multi-objective artificial bee colony algorithm for parameter optimization of link channel feedback terms, link strength, and dynamic threshold decay factor.

[0132] Link channel feedback item α L Link strength β, dynamic threshold decay factor α θ The three parameters serve as the initial population for the momentum-driven multi-objective artificial bee colony algorithm, and NP food source information is randomly generated according to the following formula (14).

[0133] X={x ij |x ij =(x i1 ,x i2 ,…x ij ,…,x id )i=1,2,…,NP; j=1,2,…,d=3}

[0134] x ij =min j +rand(0,1)×(max j -min j (14)

[0135] NP food source information is randomly generated. In a single food update evolution, a food source X to which a mercenary bee is attached is randomly selected from the bee colony. k =(x k1 ,x k2 ,…,x kd In d-dimensional space, for each food source X in the food source information spatial database... i =(x i1 ,x i2 ,…x ij ,…,x id Randomly select the j-th dimension component x ij Evolution is carried out using the following hired bee momentum update strategy, as shown in equations (15)-(16), to obtain a new food source. Where i,k∈[1,2,…,NP],i≠k,j∈[1,2,…,d],r∈[-1,1]. a ij This indicates the update step size of the previous evolution. This represents the update step size obtained after the current momentum update evolution, where γ represents momentum and has a value of 0.9.

[0136]

[0137]

[0138] In a single food update evolution, the selection probability of the observation bee can be calculated using formula (17), and a food source X to which the observation bee depends can be randomly selected in the bee colony. t =(x t1 ,x t2 ,…,x td In d-dimensional space, for each food source X in the food source information spatial database... i =(x i1 ,x i2 ,…x ij ,…,x id Randomly select the j-th dimension component x ij A new food source is obtained by observing the Nesterov momentum update strategy of bees and performing evolution as shown in equations (18)-(19). Where i, t∈[1,2,…,NP], i≠t, j∈[1,2,…,d], r∈[-1,1]. b ij This indicates the update step size of the previous evolution. This represents the update step size obtained after the current Nesterov momentum update evolution, where γ represents momentum and has a value of 0.9. Target = 2, j = 1, 2, ..., NP.

[0139]

[0140]

[0141]

[0142] In multi-objective optimization problems, the quality of individual solutions is determined by dominance relationships and density information. This application uses a lattice density construction method to ensure that the distribution of optimal solutions in the Pareto optimal solution set is not too dense. The lattice is a dynamic, n-grid equally divided interval within the range of (-inf, +inf).

[0143] Determine the maximum and minimum values ​​for each dimension of the non-dominated solution. The current interval can be divided into nGrid+1 parts using a predefined nGrid. The smallest interval starts at negative infinity - inf, and the largest interval ends at positive infinity + inf, to prevent non-dominated solutions from going out of bounds and to ensure that all non-dominated solutions fall within the grid.

[0144] For example, the formula for solving the grid index value can be shown in (20). Where, low i Represents the minimum boundary value of the grid, and Target represents the number of objective functions. i = 1, Targret, Targret = 2, j = 1, ..., nGrid.

[0145]

[0146] After determining the grid index values, a Pareto optimal solution set can be constructed to perform a uniform operation on all optimization objectives fairly and obtain solutions with a relatively fair probability of being deleted.

[0147] Constructing the optimal solution set requires a certain probability of randomly deleting redundant non-dominated solutions. The method for constructing the deletion selection probability is to use the absolute value of the difference between the grid index values ​​of non-dominated solutions in the same dimension for calculation. The mathematical expression can be shown in the following equations (21)-(22).

[0148]

[0149] poss i =1 / (poss) i +1) (22)

[0150] In addition, to address the issue of the diversity of image fusion quality evaluation functions, two image fusion quality evaluation functions, CrossEntropy (CE) and Mutual Information (MI), can be selected for multi-target fitness calculation. The calculation formula is shown in Equation (23) below.

[0151] fitness_pareto=max{CE,MI} (23)

[0152] Through the above momentum-driven multi-objective artificial bee colony calculation, the optimal link channel feedback term, link strength, and dynamic threshold decay factor can be determined as the final parameters.

[0153] The following describes the steps for determining the iteration termination information based on the saliency function and the first low-frequency sub-band image, including:

[0154] Based on the saliency function, the first low-frequency sub-band image is segmented by lighting up multiple times to obtain multiple ignition segmentation maps. The iteration termination information is determined based on the multiple ignition segmentation maps.

[0155] Optionally, the first low-frequency sub-band image can be lit up and segmented multiple times based on the saliency function of the aforementioned equation (12) to obtain multiple ignition segmentation maps, and the iteration termination information can be determined based on the ignition segmentation maps in any two states.

[0156] The following describes the steps of the pulse-coupled neural network guided by the saliency mechanism to determine the low-frequency fused image based on parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image. Step S203 includes:

[0157] The dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor are input as input parameters into the saliency mechanism-guided pulse-coupled neural network. The iteration termination information is used as the iteration termination condition of the saliency mechanism-guided pulse-coupled neural network. The saliency mechanism-guided pulse-coupled neural network fuses the first low-frequency sub-band image and the second low-frequency sub-band image based on the input parameters and the iteration termination condition to obtain a low-frequency fused image.

[0158] Optional, such as Figure 5 The diagram shown is a structural diagram of an SMPCNN model provided in this application. Figure 5 In this application, the dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor can be used as input parameters of the SMPCNN model, and the iteration termination information can be used as the iteration termination condition of the SMPCNN model. Other input parameters can be set to fixed values ​​by the user based on experience. Finally, the SMPCNN model outputs a low-frequency fused image of the first low-frequency sub-band image and the second low-frequency sub-band image.

[0159] For multiple first high-frequency sub-band images and multiple second high-frequency sub-band images, this application can use bilateral filtering technology to fuse high-frequency components. The above-mentioned step S204 includes:

[0160] An improved bilateral filtering method is used to fuse multiple first high-frequency sub-band images and multiple second high-frequency sub-band images using a spatial neighborhood Gaussian function and a high-frequency component gray value similarity Gaussian function to obtain a high-frequency fused image.

[0161] Optionally, assume that the high-frequency components of the ToF image decomposed by NSST are: The high-frequency components of a color image decomposed by NSST are The Gaussian function of the spatial neighborhood is w Neighborhood As shown in equation (24), the Gaussian function for the similarity of gray values ​​of high-frequency components is w. Similarity As shown in equation (25), the calculation method of the high-frequency fusion image can be shown in equation (26).

[0162]

[0163]

[0164]

[0165] After obtaining the low-frequency fused image and the high-frequency fused image, this application can further determine the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image. The above-mentioned step S205 includes:

[0166] Multi-scale decomposition and inverse transformation are performed on the low-frequency fused image and the high-frequency fused image to obtain the fused image of the target scene.

[0167] In this application, the low-frequency fused image and the high-frequency fused image can be subjected to NSST inverse transform to obtain the fused image of the target scene.

[0168] Next, combine Figure 6 The transform domain heterogeneous image fusion method of this application will be further explained.

[0169] First, the ToF image and the visible light image are decomposed using NSST to obtain a first low-frequency sub-band image, a second low-frequency sub-band image, multiple first high-frequency sub-band images, and multiple second high-frequency sub-band images. Then, the parameters of the SMPCNN model are obtained by applying a saliency function, KL divergence, and momentum-driven multi-objective artificial bee colony algorithm to the first low-frequency sub-band images. Next, the SMPCNN model fuses the low-frequency sub-band images to obtain a low-frequency fused image. An improved bilateral filtering method is used to fuse the high-frequency sub-band images to obtain a high-frequency fused image. Finally, the low-frequency fused image and the high-frequency fused image are subjected to inverse NSST transform to obtain the final fused image.

[0170] In the orchard image fusion scenario, after experimental comparison with other fusion models, such as the improved discrete wavelet model, the fusion model based on NSCT and local average gradient, the improved non-subsampled contourlet transform model, the simplified pulse-coupled neural network model, and the fusion model based on NSST and Dual-PCNN, the brightness of the foreground fruit target in the fused image of this application is significantly improved, and the background details are clearer and closer to the visible light image, resulting in better visual effects. Compared with the results of the other five models, the fusion result of the model in this application solves the problem of low brightness and unclear target target caused by backlighting, and effectively highlights the target area in the ToF image, resulting in better visual effects.

[0171] The results of the first set of samples from Shunguang show that the entropy, average gradient, peak signal-to-noise ratio, edge strength, spatial frequency, and image sharpness of the proposed model are 7.19, 9.34, 20.49, 92.52, 22.35, and 11.91, respectively, with a recognition fusion rate of 100.00%, higher than the results of the other five models, and a running time of 8.12s. The results of the second set of samples from Shunguang show that the entropy, average gradient, peak signal-to-noise ratio, edge strength, spatial frequency, image sharpness, and structural similarity of the proposed model are 7.25, 9.95, 19.41, 97.22, 27.89, 13.01, and 0.52, respectively, with a recognition fusion rate of 85.71%, higher than the results of the other five models, and a running time of 8.02s. The results of the third set of backlight samples show that the entropy, average gradient, peak signal-to-noise ratio, edge intensity, spatial frequency, image sharpness, and structural similarity of the proposed model are 7.51, 12.05, 16.23, 117.96, 32.03, 15.64, and 0.53, respectively, with a recognition fusion rate of 100.00%, higher than the results of the other five models, and a running time of 8.09s. The results of the fourth set of backlight samples show that the entropy, average gradient, peak signal-to-noise ratio, edge intensity, spatial frequency, image sharpness, and structural similarity of the proposed model are 7.70, 16.07, 13.79, 153.15, 43.44, 21.74, and 0.49, respectively, with a recognition fusion rate of 100.00%, higher than the results of the other five models, and a running time of 8.32s. Compared with existing image fusion models, the quality of the fused image obtained by this application is significantly improved.

[0172] Based on the same inventive concept, this application also provides a transform domain heterogeneous image fusion device corresponding to the transform domain heterogeneous image fusion method. Since the principle of the device in this application is similar to the transform domain heterogeneous image fusion method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0173] Reference Figure 7The diagram shown is a schematic of a transform domain heterogeneous image fusion device provided in an embodiment of this application. The device includes: an acquisition module 701, a determination module 702, a low-frequency determination module 703, a high-frequency determination module 704, and an image determination module 705, wherein:

[0174] The acquisition module 701 is used to: acquire time-of-flight images and visible light images of the target scene, and determine a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image;

[0175] The determination module 702 is used to determine the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency sub-band image. The parameter information includes: iteration termination information, dynamic threshold amplification coefficient, link channel feedback term, link strength, and dynamic threshold attenuation factor.

[0176] The low-frequency determination module 703 is used to determine the low-frequency fused image based on parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image by a pulse-coupled neural network guided by a saliency mechanism.

[0177] The high-frequency determination module 704 is used to fuse multiple first high-frequency sub-band images and multiple second high-frequency sub-band images using an improved bilateral filtering method to obtain a high-frequency fused image;

[0178] The image determination module 705 is used to determine the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image.

[0179] Optionally, module 701 is specifically used for:

[0180] Multi-scale decomposition of the time-of-flight image yields a first low-frequency sub-band image and multiple first high-frequency sub-band images;

[0181] The visible light image is decomposed into multiple scales to obtain a second low-frequency sub-band image and multiple second high-frequency sub-band images.

[0182] Optionally, module 702 is specifically used for:

[0183] The iteration termination information is determined based on the saliency function and the first low-frequency subband image;

[0184] The dynamic threshold amplification factor is obtained by calculating the relative entropy divergence of the first low-frequency sub-band image.

[0185] Momentum-driven multi-target artificial bee colony calculations were performed on the first low-frequency sub-band image to obtain the link channel feedback term, link strength, and dynamic threshold decay factor.

[0186] Optionally, the determining module 702 is also specifically used for:

[0187] Based on the saliency function, the first low-frequency sub-band image is illuminated and segmented multiple times to obtain multiple ignition segmentation maps. The iteration termination information is determined based on the multiple ignition segmentation maps.

[0188] Optionally, the low-frequency determination module 703 is specifically used for:

[0189] The dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor are input as input parameters into the saliency mechanism-guided pulse-coupled neural network. The iteration termination information is used as the iteration termination condition of the saliency mechanism-guided pulse-coupled neural network. The saliency mechanism-guided pulse-coupled neural network fuses the first low-frequency sub-band image and the second low-frequency sub-band image based on the input parameters and the iteration termination condition to obtain a low-frequency fused image.

[0190] Optionally, the high-frequency determination module 704 is specifically used for:

[0191] An improved bilateral filtering method is used to fuse multiple first high-frequency sub-band images and multiple second high-frequency sub-band images using a spatial neighborhood Gaussian function and a high-frequency component gray value similarity Gaussian function to obtain a high-frequency fused image.

[0192] Optionally, the image determination module 705 is specifically used for:

[0193] Multi-scale decomposition and inverse transformation are performed on the low-frequency fused image and the high-frequency fused image to obtain the fused image of the target scene.

[0194] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0195] This application embodiment optimizes parameters such as iteration termination information, dynamic threshold amplification coefficient, link channel feedback term, link strength, and dynamic threshold attenuation factor based on the first low-frequency sub-band image. The resulting saliency mechanism-guided PCNN model combines the saliency mechanism and the PCNN clustering segmentation mechanism, achieving adaptive termination and reasonable image segmentation. It also has the advantages of gray-level clustering illumination mechanism and priority illumination of the same gray-level attribute, thus achieving better image fusion effect for low-frequency sub-band images. By using an improved bilateral filtering method to fuse high-frequency sub-band images and determining the fused image of the target scene based on the obtained high-frequency fused image and low-frequency fused image, the fusion quality of ToF image and visible light heterogeneous source image can be improved.

[0196] This application also provides an electronic device, such as... Figure 8The diagram shown is a schematic representation of an electronic device structure provided in an embodiment of this application, including: a processor 801, a memory 802, and a bus. The memory 802 stores machine-readable instructions executable by the processor 801 (e.g., ...). Figure 7 The device includes the execution instructions corresponding to the acquisition module 701, determination module 702, low-frequency determination module 703, high-frequency determination module 704, and image determination module 705. When the computer device is running, the processor 801 and the memory 802 communicate via a bus. When the machine-readable instructions are executed by the processor 801, the above-mentioned transform domain heterogeneous image fusion method is performed.

[0197] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the transform domain heterogeneous image fusion method described above.

[0198] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0199] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0200] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A transform domain heterogeneous image fusion method, characterized in that, include: Acquire time-of-flight images and visible light images of the target scene, and determine a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image; Based on the first low-frequency sub-band image, the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network is determined. The parameter information includes: iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor. The pulse-coupled neural network guided by the saliency mechanism determines the low-frequency fusion image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image; An improved bilateral filtering method is used to fuse the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images to obtain a high-frequency fused image; The fused image of the target scene is determined based on the low-frequency fused image and the high-frequency fused image.

2. The method according to claim 1, characterized in that, The process of determining a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determining a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image, includes: The time-of-flight image is decomposed into multiple scales to obtain the first low-frequency sub-band image and multiple first high-frequency sub-band images; The visible light image is decomposed into multiple scales to obtain the second low-frequency sub-band image and multiple second high-frequency sub-band images.

3. The method according to claim 1, characterized in that, The step of determining the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency sub-band image includes: The iteration termination information is determined based on the saliency function and the first low-frequency sub-band image; The dynamic threshold amplification factor is obtained by calculating the relative entropy divergence of the first low-frequency sub-band image. Momentum-driven multi-target artificial bee colony calculation is performed on the first low-frequency sub-band image to obtain the link channel feedback term, link strength, and dynamic threshold attenuation factor.

4. The method according to claim 3, characterized in that, Determining the iteration termination information based on the saliency function and the first low-frequency sub-band image includes: Based on the saliency function, the first low-frequency sub-band image is segmented into multiple illuminations to obtain multiple ignition segmentation maps. The iteration termination information is determined based on the multiple ignition segmentation maps.

5. The method according to claim 1, characterized in that, The pulse-coupled neural network guided by the saliency mechanism determines the low-frequency fused image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image, including: The dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold attenuation factor are input as input parameters into the saliency mechanism-guided pulse-coupled neural network. The iteration termination information is used as the iteration termination condition of the saliency mechanism-guided pulse-coupled neural network. The saliency mechanism-guided pulse-coupled neural network fuses the first low-frequency sub-band image and the second low-frequency sub-band image based on the input parameters and the iteration termination condition to obtain the low-frequency fused image.

6. The method according to claim 1, characterized in that, The step of fusing the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images using an improved bilateral filtering method to obtain a high-frequency fused image includes: An improved bilateral filtering method is used to fuse the multiple first high-frequency sub-band images and the multiple second high-frequency sub-band images using a spatial neighborhood Gaussian function and a high-frequency component gray value similarity Gaussian function to obtain the high-frequency fused image.

7. The method according to claim 1, characterized in that, Determining the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image includes: The low-frequency fused image and the high-frequency fused image are subjected to multi-scale decomposition inverse transformation to obtain the fused image of the target scene.

8. A transform domain heterogeneous image fusion device, characterized in that, include: The acquisition module is used to: acquire time-of-flight images and visible light images of a target scene, and determine a first low-frequency sub-band image and multiple first high-frequency sub-band images of the time-of-flight image, and determine a second low-frequency sub-band image and multiple second high-frequency sub-band images of the visible light image; The determination module is used to determine the parameter information corresponding to the saliency mechanism-guided pulse-coupled neural network based on the first low-frequency subband image. The parameter information includes: iteration termination information, dynamic threshold amplification factor, link channel feedback term, link strength, and dynamic threshold decay factor. A low-frequency determination module is used to determine a low-frequency fused image based on the parameter information, the first low-frequency sub-band image, and the second low-frequency sub-band image by a pulse-coupled neural network guided by the saliency mechanism. The high-frequency determination module is used to fuse the plurality of first high-frequency sub-band images and the plurality of second high-frequency sub-band images using an improved bilateral filtering method to obtain a high-frequency fused image; The image determination module is used to determine the fused image of the target scene based on the low-frequency fused image and the high-frequency fused image.

9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus. The processor executes the program instructions to perform the steps of the transform domain heterogeneous image fusion method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the transform domain heterogeneous image fusion method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image fusion method based on improved pulse coupling neural network

    CN109934887A

  • Infrared and visible light image perception fusion method

    CN112017139A