Infrared and visible light video mimicry fusion salient frame detection method based on possibility distribution evidence synthesis

Through the mimicry fusion significant frame detection method of infrared and visible videos based on the probability distribution evidence, the problem of poor video fusion quality in dynamic scenes is solved, and the accurate recognition and adaptive optimization of significant frames are achieved, which improves the stability and accuracy of the fusion effect.

CN120259934APending Publication Date: 2025-07-04ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510178458.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing infrared and visible light video fusion models are difficult to respond to scene changes in time in dynamic scenarios, resulting in unsatisfactory fusion phenomena such as target blur, information loss and artifacts. The existing mimic fusion models lack targeted frame selection optimization mechanisms, resulting in unstable fusion effect.

Method used

The infrared and visible video mimicry fusion significant frame detection method based on possibility distribution evidence synthesis is adopted. Through the differential feature space-time joint characterization module, the possibility distribution synthesis module with significant change in different characteristics and the D-S evidence synthesis decision module, the significant frame is accurately identified and optimized to achieve adaptive fusion in dynamic scenarios.

Benefits of technology

The quality of infrared and visible video fusion in dynamic scenes is significantly improved, the accuracy of significant frame detection and the robustness of the fusion process are improved, and the poor fusion performance caused by structural fixation in traditional methods is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259934A_ABST
    Figure CN120259934A_ABST
Patent Text Reader

Abstract

The invention relates to the field of multi-modal video fusion, in particular to an infrared and visible light video mimicry fusion salient frame detection method based on possibility distribution evidence synthesis, which fully characterizes the amplitude and frequency distribution of features in a bimodal video through a difference feature spatio-temporal joint characterization module. Through a possibility distribution function and a distribution synthesis rule, flexible quantification of difference characteristic changes in a complex scene is realized, and limitations of a traditional threshold-based method in the aspects of capturing uncertainty and significant changes are overcome. And sorting the difference characteristic change significance by using a multi-attribute decision, and driving the optimization selection of a mimicry fusion strategy. And based on a D-S evidence synthesis rule, carrying out fusion decision on the multi-source difference feature information, and realizing accurate detection of the significant frame. According to the invention, through accurate identification of the significant frame and adaptive optimization of the fusion strategy, the quality of infrared and visible light video fusion in a dynamic scene is improved, and the limitation of poor fusion performance caused by a fixed structure in a traditional method is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal video fusion, and specifically to a method for detecting significant frames of infrared and visible light video mimicry fusion based on evidence synthesis of possibility distributions. Background Art

[0002] Infrared and visible light video fusion, with the temporal continuity and information richness of its dynamic information, can not only significantly improve the efficiency and accuracy of target detection, but also greatly enhance the spatial resolution and anti-interference ability of the detection system. By giving full play to the complementary advantages of bimodal videos, it has shown extensive application value in complex scenarios that require dynamic response, such as night monitoring, intelligent driving, and personnel detection.

[0003] The changes of features in videos usually have a high degree of complexity, and the environment often fluctuates violently due to factors such as movement and weather changes. This makes the video fusion task face more severe challenges. However, most of the existing fusion models independently process static image information frame by frame based on predefined fusion rules. This method performs well in static environments, but in actual video fusion tasks, the preset fusion strategy is difficult to respond to dynamically changing scenarios in a timely manner, resulting in unsatisfactory fusion phenomena such as blurred targets, information loss, and artifacts in video fusion. In severe cases, it may even cause fusion failure, as Figure 1 shown. Although the fusion model based on deep learning has the ability to learn complex features from a large number of samples, its performance is extremely dependent on the coverage and quality of the training data set. When encountering new scenarios outside the training set or when the environment changes violently, due to the fixed weights after the model training is completed, it is often unable to adjust the strategy in a timely manner to adapt to these changes. At the same time, in the case of scarce data or limited computing resources, the application of deep learning is also greatly restricted.

[0004] As an emerging intelligent bionic fusion method, mimicry fusion effectively improves the adaptive performance of the model. However, most of the existing mimicry fusion models adopt a full-frame processing strategy and lack a targeted frame selection optimization mechanism. For those frames that have achieved good fusion effects through a single algorithm, additional mimicry fusion not only cannot significantly improve the fusion quality, but may also introduce errors and unnecessary computational overhead. In addition, the existing models mainly rely on pre-computed inter-modal difference features to drive mimicry changes. However, this driving method fails to consider the temporal changes between frames, lacks dynamic continuity and pertinence, and may lead to unstable fusion effects and decreased adaptability.

[0005] In summary, it is urgent to introduce a more accurate and targeted mechanism to improve the performance of the mimicry fusion model, so as to effectively solve the problem of poor video fusion quality caused by the fixed structure of traditional fusion algorithms. Summary of the Invention

[0006] In order to solve the problem of poor video fusion quality caused by the fixed structure of traditional fusion algorithms, the present invention provides a method for detecting significant frames of infrared and visible light video mimic fusion based on the synthesis of evidence of possibility distribution.

[0007] The present invention is implemented by the following technical solutions: A method for detecting significant frames of infrared and visible light video mimic fusion based on the synthesis of evidence of possibility distribution includes a spatio-temporal joint characterization module for differential features, a possibility distribution synthesis module for the significance of differential feature changes, and a D-S evidence synthesis decision module; First, in the spatio-temporal joint characterization module for differential features, inter-frame and intra-frame differential features are respectively extracted for bimodal complementary information and comprehensively characterized; Secondly, in the possibility distribution synthesis module for the significance of differential feature changes, the quantification of differential feature changes is regarded as an uncertain problem, the distribution characteristics of differential feature data are analyzed to construct a possibility distribution function to flexibly quantify the degree of feature changes, and by designing a significance weight distribution matrix and a possibility distribution synthesis rule, the possibility values of differential features are synthesized; Then, in the D-S evidence synthesis decision module, based on the constructed possibility mass function, a comprehensive ranking of the significance of differential feature changes is obtained by combining the preference order index and its net flow. Finally, the final decision result is obtained by combining the discount coefficient method and its D-S evidence synthesis rule.

[0008] In the above method for detecting significant frames of infrared and visible light video mimic fusion based on the synthesis of evidence of possibility distribution, six types of features, namely, gray mean, standard deviation, edge intensity, average gradient, contrast, and roughness, are selected as differential features.

[0009] In the above method for detecting significant frames of infrared and visible light video mimic fusion based on the synthesis of evidence of possibility distribution, in the spatio-temporal joint characterization module for differential features, first, an m×n smoothing window is used to perform non-overlapping partitioning on each frame of the video, and the feature amplitude of each pixel block within the i-th frame is calculated respectively and d i,r , where r represents 6 types of differential features, represents the r-th type of feature value of the i-th frame infrared image, represents the r-th type of feature value of the i-th frame visible light image, D i,r represents the bimodal intra-frame differential feature amplitude, and these feature amplitudes constitute the initial sample set of each frame; interpolation expansion is performed on each pixel block within the initial sample set of each frame with a moving step size st to obtain a denser sample set {Z} for that frame, Z ∈ {F V , F I , D}, and a single sample point in the sample set is represented by Z k , k = 1, 2,..., M. Then, the formula is used to calculate each sample point Z in the sample set {Z} kThe probability density, where is the total number of samples after interpolation expansion, and Z R and Z L are the left and right boundaries of the sample set. represents the number of nearest neighbors. is the sample point that is the k k th nearest to the sample point Z M ; Combining the characteristic amplitude of each sample point Z k and its probability density ρ(Z k ), calculate the comprehensive weight ω′ of each sample point and normalize it. The calculation is as follows: ω′(Z k ) = Z k ×ρ(Z k ). Finally, according to the formula Calculate the comprehensive characterization values of the three difference features of the i - th frame of visible - light inter - frame, infrared inter - frame, and dual - mode intra - frame. Among them, represents the characteristic amplitude of the k - th sample point in the i - th frame of infrared image. represents the characteristic amplitude of the k - th sample point in the i - th frame of visible - light image. represents the difference - feature amplitude of the k - th sample point in the i - th frame of infrared image.

[0010] In the above - mentioned infrared and visible - light video mimicry fusion significant - frame detection method based on the synthesis of evidence of possibility distribution, in the possibility - distribution synthesis module of the difference - feature change significance, first use K - means clustering to obtain the segmentation points of the possibility - distribution function, and use the Manhattan distance to divide the difference - feature data into three clusters and obtain the centroid μ j of each cluster respectively. C j = {s i : |s i -μ j | = min(|s i -μ j |)}, where are the three difference features of the i - th frame. Secondly, according to the clustering centroid μ j , construct the following possibility - distribution function to dynamically map the difference - feature value of each frame to the significance possibility value of its change degree.

[0011] Then, a new non - linear weighted synthesis method is proposed, and the pairwise - synthesis method is used to obtain the possibility - distribution synthesis result that comprehensively reflects the significance of the difference - feature changes between different frames and different modalities.

[0012] Among them, α is the weight corresponding to the significance level combination, and β = 1 - α.

[0013] For the above-mentioned infrared and visible light video mimicry fusion significant frame detection method based on possibility distribution evidence synthesis, in the D-S evidence synthesis decision module, the recognition framework is set as Θ = {low, medium, high}, and each type of difference feature is used as an evidence body, with the possibility distribution synthesis result As the input x, construct the following possibility mass function Normalize it According to the discount coefficient method, for the possibility mass function of the evidence body Update it to obtain the corrected mass function m l ′(A), Among them, m l ′(Θ) represents the uncertain part of the evidence body, and a l represents the importance weight of the evidence body; finally, according to the D-S evidence synthesis rule, fuse the possibility mass functions m′ l (low), m′ l (medium), m′ l (high) and m′ l (Θ) of each evidence body l to obtain the synthetic mass value of each decision category. Finally, determine the decision result according to the maximum mass value, and judge the frame with the decision result of high as the significant frame.

[0014] For the above-mentioned infrared and visible light video mimicry fusion significant frame detection method based on possibility distribution evidence synthesis, in the D-S evidence synthesis decision module, two types of criteria, within the evidence body and between evidence bodies, are adopted, and the importance of each evidence body is sorted using the PEOMETHEE-II method based on entropy weight; first, construct the criterion matrix g p (l), where l = 1, 2,..., L represents the number of evidence bodies, p = 1, 2,..., N represents the number of criteria, L = N = 6. Normalize the criterion matrix to c lp and calculate the entropy value E p , Then, calculate the difference degree G p of each criterion, convert the entropy value to a positive information difference degree measure, and obtain the weight Ω p of each criterion, Among them, G p = 1 - E p , which defines the possibility distribution of the preference function Among them,

[0015] g p (q), the evidential body m is further calculated l and m q preference order index Then, by calculating the net flow Φ l , the relative importance of the evidential body in decision-making is obtained Since the importance weight a of the evidential body l is an increasing function of the net flow, a right-skewed possibility distribution is defined where a = min(η l ), b = max(η l ).

[0016] Through the precise recognition of significant frames and the adaptive optimization of the fusion strategy, the present invention significantly improves the quality of infrared and visible light video fusion in dynamic scenes, breaking through the limitation of poor fusion performance caused by fixed structures in traditional methods. Aiming at the complex changes of scene features in infrared and visible light videos, the present invention designs a spatio-temporal joint representation module for differential features, and realizes the flexible measurement of differential information changes by constructing a possibility distribution function and designing a synthesis rule. In addition, based on multi-attribute decision-making and the D-S evidence synthesis rule, the significance ranking of differential features and the fusion decision-making of multi-source information are realized, thus significantly improving the accuracy of significant frame detection and the robustness of the fusion process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of problems in infrared and visible light video fusion. Fusion j shows obvious blurring and artifact phenomena, which are caused by the preset fusion strategy failing to adapt to dynamic scene changes in time, resulting in poor fusion effect.

[0018] Figure 2 It is a schematic diagram of a heuristic model for significant frame detection based on the mimic octopus decision-making mechanism.

[0019] Figure 3 It is the overall flowchart of the method of the present invention.

[0020] Figure 4 It is a schematic diagram of the possibility evidence distribution synthesis module.

[0021] Figure 5 It is a schematic diagram of the possibility distribution function.

[0022] Figure 6 It is a significance weight distribution matrix diagram, and each grid cell represents a different combination of significance levels, and the value is its corresponding weight α.

[0023] Figure 7 Schematic diagram of the scores of the fusion algorithm for each feature.

[0024] Figure 8 Schematic diagram of the neutrality comparison of the fusion algorithm.

[0025] Figure 9 Schematic diagram of the objective annotation results of significant frames.

[0026] Figure 10 Schematic diagram of the visualization of the distribution of significant frames.

[0027] Figure 11 Schematic diagram of the visual comparison of 11 fusion methods Figure 1 .

[0028] Figure 12 Schematic diagram of the quantitative comparison of the fusion results in 9 objective evaluation indicators Figure 1 .

[0029] Figure 13 Schematic diagram of the visual comparison of 11 fusion methods Figure 2 .

[0030] Figure 14 Schematic diagram of the quantitative comparison of the fusion results in 9 objective evaluation indicators Figure 2 .

[0031] Figure 15 Schematic diagram of the quantitative comparison of the overall video fusion effect. Detailed implementation manners

[0032] Studies have shown that the multi-mimicry behavior of the mimic octopus does not occur in all situations, but shows a high degree of selectivity. Therefore, inspired by its selective mimicry behavior, the present invention introduces a significant frame detection and targeted driving mechanism into video mimicry fusion, aiming to optimize the selection of mimicry fusion strategies for frames with drastic dynamic changes. The research on significant frame detection technology in the field of video fusion is still relatively blank. Especially how to effectively detect and target the processing of significant frames in dynamic scenes to improve the video fusion effect has not received extensive attention. However, significant research progress has been made in the fields of video compression, video retrieval classification, and video summary generation. However, existing significant frame detection methods mainly focus on extracting representative frames of video content, rather than specifically capturing frames with drastic dynamic changes, and do not have the ability to provide guidance for subsequent fusion strategies, which greatly limits their application in video fusion. In addition, traditional significant frame detection methods mostly rely on hard threshold settings. However, due to the differences in the dual-modal imaging mechanism and the real-time changes of targets and backgrounds in dynamic scenes, the method of fixed thresholds becomes unreliable. And the present invention's significant frame detection method for infrared and visible light video mimicry fusion based on the synthesis of evidence of possibility distribution analogizes the decision-making mechanism of the mimic octopus to make a decision on whether to perform mimicry transformation after evaluating the degree of environmental threat to the process of judging whether a frame is a significant frame according to the change of differential features in the dual-modal video fusion task, and uses a dynamic decision-making strategy to cleverly overcome the complexity of significant frame detection in dynamic scenes. Figure 2 Shows the analogical relationship between the decision-making mechanism of the mimic octopus and significant frame detection.

[0033] The overall block diagram of the significant frame detection method for infrared and visible light video mimicry fusion based on the synthesis of evidence of possibility distribution is as Figure 3 shown, mainly including three modules: the spatio-temporal joint characterization module of differential features, the possibility distribution synthesis module of the significance of differential feature changes, and the D-S evidence synthesis decision module. First, six inter-frame and intra-frame differential features are extracted for the dual-modal complementary information respectively and comprehensively characterized. Secondly, the quantification of differential feature changes is regarded as an uncertain problem, and the distribution characteristics of differential feature data are analyzed to construct a possibility distribution function to flexibly quantify the degree of feature changes. By designing a significance weight distribution matrix and a possibility distribution synthesis rule, the possibility values of differential features are synthesized. Then, based on the constructed possibility mass function, combined with the preference order index and its net flow, a comprehensive ranking of the significance of differential feature changes is obtained. Finally, the final decision result is obtained by combining the discount coefficient method and its D-S evidence synthesis rule.

[0034] Spatio-temporal joint characterization module of differential features

[0035] Due to the different imaging mechanisms of infrared and visible light videos, the complementary information between the two is mainly manifested in three aspects: brightness, edges, and texture details. To effectively characterize these three types of differences, the present invention selects six types of features: gray mean, standard deviation, edge strength, average gradient, contrast, and roughness.

[0036] The amplitude of the difference feature refers to the absolute difference degree of the amplitude between different frames or different modalities of the same feature. The present invention extracts the amplitude of the difference feature between infrared frames, between visible light frames, and within dual-modal frames respectively, as shown in the following formula.

[0037]

[0038] Where r represents the 6 types of difference features, represents the r-th type of feature value of the i-th infrared image, represents the r-th type of feature value of the i-th visible light image.

[0039] The frequency of the difference feature reflects the distribution of the amplitude of the difference feature in the frame. The present invention obtains its probability density distribution based on the K-nearest neighbor non-parametric estimation method, and then obtains the frequency distribution of the amplitude of the difference feature. First, use a smoothing window of m×n to perform non-overlapping block division on the video frame by frame, and calculate the feature amplitude and D i,r of each pixel block within the i-th frame. These feature amplitudes constitute the initial sample set of each frame. To improve the accuracy of subsequent probability density estimation, each pixel block in the initial sample set of each frame is interpolated and expanded with a moving step size st to obtain a denser sample set {Z} (Z∈{F V ,F I ,D}) of that frame. A single sample point in the sample set is represented by Z k , k = 1, 2, …, M. Then, use formula (4) to calculate the probability density of each sample point Z k in the sample set {Z}.

[0040]

[0041] Among them, is the total number of samples after interpolation and expansion, Z R and Z L are the left and right boundaries of the sample set. represents the number of nearest neighbors. is the sample point that is the k k -th nearest to the sample point Z M .

[0042] Combining the feature amplitude of each sample point Z k and its probability density ρ(Z k ), calculate the comprehensive weight ω′ of each sample point and normalize it, as calculated below.

[0043] ω′(Z k ) = Z k × ρ(Z k ) (5)

[0044]

[0045] Finally, the comprehensive characterization values of the three difference features between visible light frames, between infrared frames, and within dual - modal frames of the i - th frame are calculated according to the following formula. The formula is as follows.

[0046]

[0047] Among them, represents the feature amplitude of the k - th sample point in the i - th infrared image, represents the feature amplitude of the k - th sample point in the i - th visible - light image, represents the difference - feature amplitude of the k - th sample point in the i - th infrared image.

[0048] Possibility - distribution synthesis module for the significance of difference - feature changes

[0049] In a dynamic scene, the change of difference features is often fuzzy and uncertain. If the significance degree of difference - feature changes is directly judged by setting a threshold, the details of feature dynamic changes are likely to be ignored. Therefore, the method of possibility - distribution synthesis is adopted in the present invention, and the process is as Figure 4 shown.

[0050] First, the segmentation points of the possibility - distribution function are obtained by using K - means clustering. It can be known from statistical analysis that the difference - feature values are right - skewed data. If the Euclidean distance is directly used, the influence of extreme values is likely to be amplified, causing the clustering center to deviate from the main body of the data. Therefore, the Manhattan distance is adopted to divide the difference - feature data into three clusters and obtain the centroid μ j of each cluster. The formula is as follows.

[0051] C j ={s i :|s i - μ j | = min(|s i - μ j |)} (10)

[0052] Among them, are the three difference - feature values of the i - th frame. Secondly, according to the clustering centroid μ j , the following possibility - distribution function is constructed (see Figure 5 and formula (11)), and the difference - feature values of each frame are dynamically mapped to the significance possibility values of their change degrees.

[0053]

[0054] To effectively fuse the significant information of the differential features between different modalities, a possibility distribution synthesis rule based on the weight distribution matrix and non-linear weighting is designed. First, according to the combination method of the significant differences of different modality features, combined with the clustering clusters and their centroids obtained by formula (10), a 3×3 significant weight distribution matrix (as shown in Figure 6 (a) and 6(b)) is constructed to map the weights of different significant combinations. Secondly, in order to make the values with higher possibility have a greater impact on the synthesis result, a new non-linear weighted synthesis method is proposed and defined as follows.

[0055]

[0056] where α is the weight corresponding to the combination of significant degrees, and β = 1 - α. Finally, combined with the weight matrix and the synthesis formula, the pairwise synthesis method is used to obtain the possibility distribution synthesis result that synthesizes the significant changes of the differential features between different frames and different modalities D-S evidence synthesis decision module

[0057] To effectively fuse the changes of six types of differential features to obtain the decision result, the present invention adopts the method of multi-attribute decision-making and D-S evidence synthesis. Let the recognition framework be Θ = {low, medium, high}, and each type of differential feature is regarded as an evidence body. To avoid the situation of multi-level fuzziness, the events of cross-combination are not considered. Using the above possibility distribution synthesis result as the input x, the following possibility mass function is constructed.

[0058]

[0059] To ensure that the sum of the mass functions of all events is 1, it is normalized.

[0060]

[0061] Two types of criteria, within the evidence body and between the evidence bodies (see Table 1 and Table 2), are adopted, and the importance of each evidence body is ranked using the PEOMETHEE-II method based on entropy weight.

[0062] Table 1 Evaluation criteria within the evidence body

[0063]

[0064] Table 2 Evaluation criteria between the evidence bodies

[0065]

[0066] First, the criterion matrix g is constructed according to the evaluation criteriap (l), where l = 1, 2, …, L represents the number of evidence bodies, p = 1, 2, …, N represents the number of criteria, and L = N = 6. To measure the uncertainty of each criterion, the criterion matrix is normalized to c lp , and the entropy value E of each criterion is calculated p .

[0067]

[0068] Then, the difference degree G of each criterion is calculated p , the entropy value is converted into a positive information difference degree measure, and the weight Ω of each criterion is obtained p , to measure the importance of each criterion in the decision-making process.

[0069]

[0070] Among them, G p = 1 - E p . To distinguish the differences between different evidence bodies under the same criterion, the following possibility distribution of the preference function is defined

[0071]

[0072] Among them,

[0073] Furthermore, the preference order index H of the evidence bodies m l and m q is calculated lq , to quantify their relative preference degrees in the decision-making.

[0074]

[0075] Then, by calculating the net flow Φ l , the relative importance of the evidence bodies in the decision-making is obtained. The larger Φ i , the higher the importance of the evidence body m l . The calculation is as follows.

[0076]

[0077] Since the importance weight a l of the evidence body is an increasing function of the net flow, the following right-skewed possibility distribution is defined.

[0078]

[0079] Among them, a = min(η l ), b = max(η l) Update the possibility mass function of the evidence body according to the discount coefficient method to obtain the revised mass function m l ′(A).

[0080]

[0081] Among them, m l ′(Θ) represents the uncertain part of the evidence body. Finally, according to the D-S evidence combination rule of formula (29), fuse the possibility mass functions (m′ l (low), m′ l (medium), m′ l (high), m l ′(Θ)) of each evidence body l to obtain the combined mass value of each decision category. Finally, determine the decision result according to the maximum mass value, and judge the frame with the decision result of high as the significant frame.

[0082]

[0083] Verification experiment

[0084] Source video dataset

[0085] The present invention uses the following three infrared and visible light video datasets to test the algorithm performance.

[0086] The BEPMS dataset describes the scene of a soldier shuttling through the jungle, with the camera following the target and the background changing dynamically and complexly. The image size is 576×480.

[0087] The VLIRVDIF dataset Hangar scene describes a group of people walking indoors, accompanied by the switching on and off of lights and the shaking of flashlights. The image size is 720×480.

[0088] The VLIRVDIF dataset Camouflage scene describes a man with a long gun hiding in the grass, and after a while, there is smoke occlusion. The image size is 720×480.

[0089] Evaluation index

[0090] To evaluate the effect of significant frame detection, the present invention selects three commonly used evaluation indexes: recall rate, accuracy rate, and F1 score. In addition, the interval density coverage D ratio is defined to measure the density of detection points of the algorithm within the significant frame interval manually annotated. The definition is as follows.

[0091]

[0092] Among them, N is the total number of manually marked significant frame intervals, is the number of points detected by the algorithm within the i-th significant frame interval, is the number of manually marked points within the i-th significant frame interval.

[0093] To evaluate the fusion quality, the present invention selects the following nine evaluation indicators: average gradient (AG), edge retention (EI), visual information fidelity (VIF), entropy (EN), Q W , root mean square error (RMSE), Q AB / F , spatial frequency (SF), and standard deviation (SD). Among them, EN, Q W , VIF, RMSE, Q AB / F belong to the overall evaluation indicators, and mainly evaluate the fusion quality from aspects such as information richness, global information retention, visual fidelity, and pixel error; AG, EI, SF, and SD mainly evaluate from local details such as edges, textures, and contrasts. Except for RMSE, the higher the value of other indicators, the better the fusion effect.

[0094] Comparison methods

[0095] To evaluate the performance of the method of the present invention, the proposed method is compared with three methods, including traditional methods and deep learning-based methods: significant frame detection based on inter-frame difference, significant frame detection based on CNN, and significant frame detection based on clustering, which are respectively represented as S1 to S3.

[0096] To measure the improvement of the video fusion effect by this method, the present invention is compared with 10 fusion methods, including WPT, SWT, NSST, NSCT, LP, DTCWT, CVT, NestFuse, RFN-Nest, and the mimic fusion method without using significant frame difference features as a driver, denoted as A1 to A10. The parameters of all comparison methods are configured according to the original literature.

[0097] Experimental cases

[0098] Taking the BEPMS dataset as an example to illustrate the overall process of the method of the present invention, due to position limitations, only the experimental results of 3 frames are shown. First, non-overlapping block processing is performed frame by frame using a 16×16 smoothing window, and the amplitude and frequency information of the corresponding block features are extracted. Then, the comprehensive weight is calculated to obtain the comprehensive value of the difference features between video frames and within bimodal frames, as shown in Table 3.

[0099] Table 3 Comprehensive characterization of difference features

[0100]

[0101] Taking the cluster centroid as the segmentation point, the possibility distribution functions of various difference features in different modalities are constructed respectively based on formula (11). Then, according to the significance weight and the non-linear synthesis rule, the synthesized value of the possibility distribution of the difference features is obtained, as shown in Table 4. For example, the synthesized value of the possibility of the difference feature GM in the 93rd frame is 0.4646, that is, the possibility of GM being a significant change is 46.46%.

[0102] Table 4 Synthesis results of the possibility distribution of difference features between different modalities

[0103]

[0104] First, the possibility mass functions of each difference feature are constructed through formulas (14-16), and the relative importance of each criterion in the decision-making is calculated based on formula (20). Secondly, the preference degree of each evidence body is evaluated based on formulas (21) and (22), and the importance weight of each evidence body is determined through the net flow and formula (26). Then, the possibility mass functions of each evidence body are corrected by using formulas (27) and (28), so as to obtain the significance ranking of the changes of each difference feature. Finally, the final evidence synthesis decision result is obtained based on formula (29). Table 5 shows the corrected possibility mass function values and evidence synthesis results of the 93rd frame, where m(high) represents the support degree of each difference feature change being highly significant, and the frames with the decision result of "high" are judged as significant frames. Table 6 lists the significance discrimination results of 3 frames, and the highlighted parts are the difference features with significant changes, which are used to drive the selection of the mimicry fusion strategy.

[0105] Table 5 Evidence synthesis results of the 93rd frame

[0106]

[0107]

[0108] Table 6 Detection results of significant frames

[0109]

[0110] Manual annotation of significant frames in the dataset

[0111] Since the infrared and visible light video datasets lack significant frame labels, and traditional manual annotation methods mostly rely on the subjective judgment of annotators, there are problems of insufficient objectivity and bias. Therefore, the present invention designs a subjective and objective combined annotation method to improve the accuracy and reliability of annotation.

[0112] First, select a neutral fusion method as the benchmark tool to ensure that no bias towards certain features is introduced during the fusion process. In this invention, seven classic fusion algorithms (WPT, SWT, NSST, NSCT, LP, DTCWT, and CVT) are fused on multiple datasets, and the evaluation index scores of each algorithm for each feature are calculated ( Figure 7 ). Through the analysis of the radar chart ( Figure 8 ), it is found that the NSST algorithm performs relatively evenly in all evaluation indexes without obvious deviation. Therefore, the NSST algorithm is selected as the benchmark tool, and a sequence of fused images is generated.

[0113] To accurately evaluate the change in video fusion quality, this invention combines the image structure, brightness, edge texture, and the degree of preservation of visual global information, and selects three indexes, SSIM, VIFF, and Q AB / F for evaluation. Due to the content differences between video frames, the index values will show a fluctuating trend. To effectively identify significantly changed frames, this invention calculates the change rate of each frame's index value, and statistically calculates the mean and standard deviation of the change rates of all frames. Based on the 95% confidence interval, a change threshold is set, and the frames with a change rate lower than this threshold are marked as significant frames, indicating that there may be a large quality decline or dynamic change in the scene for this frame. The formula is as follows. Figure 9 Shows the objective annotation results of the first 300 frames of the BEPMS dataset.

[0114] ΔM = M i - M i-1 (31)

[0115] T M = μ M - 1.96×σ M (32)

[0116] Finally, an artificial subjective confirmation link is introduced based on the objective annotation. The annotator uses the objective annotation results as candidate frames, and combines criteria such as scene change and visual saliency to determine the final significant frame annotation results.

[0117] Analysis of experimental results

[0118] Analysis of significant frame detection effect

[0119] To verify the significant frame detection effect of this invention in complex dynamic scenes, comparative experiments are conducted with three significant frame detection methods (S1, S2, and S3) on the BEPMS dataset and the VLIRVDIF dataset respectively, and the significant frame labels of the datasets are obtained through manual annotation.

[0120] Figure 10(a) and Table 7 show the experimental results under the BEPMS dataset. It can be seen that the detection results of S1 and S2 have a wide coverage range, but the distribution is too uniform and scattered, resulting in too many redundant frames. The S3 method has serious missed detections, and the Recall is only 0.1. In contrast, the significant frames detected by the present invention are closest to the results of manual annotation in terms of quantity and position, reflecting higher accuracy and interval matching degree. Moreover, the scores of each index of the present invention are significantly better than other methods, reflecting its excellent performance in the task of detecting significant frames in video fusion.

[0121] The detection results of significant frames in the Hangar subset of the VLIRVDIF dataset are as Figure 10 (b) and Table 8 show. It can be seen from this that there are still obvious redundancies in the detection results of S1. The D ratio of S2 is only 0.1, indicating that the position deviation of its detection results is relatively large. The Precision value of S3 is 0.09, and the detection performance is poor. In contrast, the comprehensive performance of the present invention in terms of coverage rate and accuracy rate is more balanced, further proving the adaptability and robustness of this method.

[0122] Table 7 Quantitative comparison of significant frame detection under the BEPMS dataset

[0123]

[0124] Table 8 Quantitative comparison of significant frame detection under the VLIRVDIF dataset

[0125]

[0126] Fusion Performance on the BEPMS dataset

[0127] To verify the key role of significant frames in driving the mimicry fusion strategy and its improvement of the fusion effect, a comparative experiment on the single-frame fusion effect was designed. Six frames were randomly selected from the detection results of significant frames in the BEPMS dataset for analysis.

[0128] In the method (A-Prop) of the present invention, the significant frames drive the mimicry fusion strategy selection through the differential features of significant changes within the frame; while the non-significant frames keep the fusion strategy unchanged. Figure 11The visual fusion effects of 11 methods are shown. It can be observed that a single fusion algorithm shows obvious deficiencies in frames with large dynamic changes. For example, in frame 986, A6 shows obvious loss of texture and edge information, as well as bright and dark flickering problems. A10 retains better details in frames 121 and 499, but due to the lack of targeted differential feature driving, the overall fusion effect is not stable enough. In contrast, the A-Prop method shows excellent fusion effects in all these 6 frame scenarios, and the fusion results present clearer and more realistic textures.

[0129] The 9 objective evaluation results are as Figure 12 and shown in Table 9. As deep learning-based fusion algorithms, A8 and A9 perform better in overall metrics such as Q AB / F and SSIM, but they are insufficient in adaptability and stability in dynamic scenarios. Especially for A9, significant decreases occur in multiple metrics in frame 694, resulting in a lower average metric value. In contrast, A-Prop achieves the optimal values in multiple key metrics such as VIFF, RMSE, EN, Q W etc., and is significantly superior to other algorithms especially in local detail evaluation metrics such as EI, AG, SF, which is consistent with the above qualitative analysis. This fully demonstrates the targeted driving ability of significant frames and their key role in improving video fusion effects.

[0130] Table 9 Average values of fusion results in 9 objective evaluation metrics.

[0131] Italic bold: Best effect

[0132]

[0133] Fusion Performance on the VLIRVDIF dataset

[0134] To verify the robustness and effectiveness of the proposed method, further tests were carried out in the Camouflage scenario of the VLIRVDIF dataset, and the fusion results are as Figure 13 shown. The visual effects of A1 - A7 in 6 frames have little difference, but the contrast between the target and the background is low. A8 retains better texture details of visible light, but fails to effectively fuse infrared image information in some frames. For example, in the smoke scenario, the contour information of the target is severely lost. In contrast, the A-Prop method can better fuse the complementary information of infrared and visible light images, and its visual effect is superior to other algorithms in terms of contrast and edge retention.

[0135] Combined with Figure 14According to the quantitative result analysis in Table 10, the A-Prop algorithm achieves the optimal values in local detail evaluation metrics such as AG, EI, SF, and SD, demonstrating excellent detail retention ability. Although A10 is close to A-Prop in some metrics, its performance in dynamic scenarios is not stable enough, such as a significant performance decline in the 50th frame. Comprehensive analysis shows that the significant-frame-driven mimicry fusion method exhibits higher robustness and adaptability in dynamic scenarios, and the fusion results are significantly better than those of other single algorithms.

[0136] Average performance of the fusion results in 9 objective evaluation metrics.

[0137] Italic bold: Best effect

[0138]

[0139] Analysis of the overall video fusion effect

[0140] To further verify the improvement effect of significant frames on the overall video fusion effect, 50 consecutive infrared and visible light videos were intercepted from the Camouflage scenarios of the BEPMS dataset and the VLIRVDIF dataset for experiments, focusing on evaluating whether the fusion algorithm can maintain stable fusion performance in video scenarios. The experimental results are as Figure 15 shown. The histogram represents the average normalized value of each metric in consecutive frames, and the line graph represents the cumulative value of the metric. The higher the score, the more stable the fusion performance. The results show that the A-Prop algorithm exhibits the highest cumulative score and balanced metric performance on both datasets, fully demonstrating its stable performance in dynamic scenarios and effectively solving the problem of the decline in fusion performance caused by the fixed structure of traditional fusion algorithms in dynamic video scenarios.

[0141] Inspired by the selective mimicry behavior and its decision-making mechanism of the mimic octopus, the present invention proposes an infrared and visible light video mimic fusion salient frame detection method based on the synthesis of evidence of possibility distribution. Through the accurate recognition of salient frames and the adaptive optimization of fusion strategies, the quality of infrared and visible light video fusion in dynamic scenes is significantly improved, breaking through the limitation of poor fusion performance caused by fixed structures in traditional methods. Aiming at the complex changes of scene features in infrared and visible light videos, the present invention designs a spatio-temporal joint representation module for differential features, and realizes the flexible measurement of differential information changes by constructing a possibility distribution function and designing a synthesis rule. In addition, based on multi-attribute decision-making and the D-S evidence synthesis rule, the saliency ranking of differential features and the fusion decision of multi-source information are realized, thus significantly improving the accuracy of salient frame detection and the robustness of the fusion process. Experimental results show that the present invention is significantly superior to existing methods in key indicators such as Recall, Precision, F1-score, and D_ratio, and at the same time demonstrates excellent performance in terms of video fusion quality and fusion stability in dynamic scenes.

Claims

1. An infrared and visible light video mimicry fusion significant frame detection method based on likelihood distribution evidence synthesis, characterized in that: It includes a spatio-temporal joint representation module for differential features, a likelihood distribution synthesis module for the significance of differential feature changes, and a D-S evidence synthesis decision module. First, in the spatio-temporal joint representation module for differential features, inter-frame and intra-frame differential features are respectively extracted from the bimodal complementary information and comprehensively represented. Second, in the likelihood distribution synthesis module for the significance of differential feature changes, the quantification of differential feature changes is regarded as an uncertain problem, the distribution characteristics of differential feature data are analyzed to construct a likelihood distribution function for flexible quantification of the degree of feature changes, and the likelihood values of differential features are synthesized by designing a significance weight distribution matrix and a likelihood distribution synthesis rule. Then, in the D-S evidence synthesis decision module, based on the constructed likelihood mass function, a comprehensive ranking of the significance of differential feature changes is obtained by combining the preference order index and its net flow. Finally, the final decision result is obtained by combining the discount coefficient method and its D-S evidence synthesis rule.

2. The method for detecting significant frames of infrared and visible light video mimicry fusion based on evidence synthesis of possibility distribution according to claim 1, wherein: Six types of features, namely grayscale mean, standard deviation, edge intensity, average gradient, contrast, and roughness, are selected as differential features.

3. The infrared and visible light video mimicry fusion significant frame detection method based on the evidence synthesis of possibility distribution according to claim 2, characterized in that: In the spatio-temporal joint representation module of differential features, first, an m×n smoothing window is used to non-overlappingly divide each frame of the video into blocks, and the feature amplitudes of each pixel block within the i-th frame are calculated respectively and D i,r , where r represents 6 types of differential features, represents the r-th type of feature value of the i-th frame infrared image, represents the r-th type of feature value of the i-th frame visible light image, and D i,r represents the differential feature amplitude within the dual-modal frame. These feature amplitudes constitute the initial sample set of each frame; Interpolate and expand each pixel block in the initial sample set of each frame with a moving step size st to obtain a denser sample set {Z} for that frame, where Z ∈ {F V , F I , D}, and a single sample point in the sample set is represented by Z k , where k = 1, 2, …, M. Then, use the formula to calculate the probability density of each sample point Z k in the sample set {Z}, where is the total number of samples after interpolation and expansion, Z R and Z L are the left and right boundaries of the sample set, represents the number of nearest neighbors, is the sample point that is the k k -th nearest to the sample point Z M ; Combine the feature amplitude of each sample point Z k and its probability density ρ(Z k ) to calculate and normalize the comprehensive weight ω′ of each sample point, and the calculation is as follows: ω′(Z k ) = Z k ×ρ(Z k ), Finally, according to the formula Calculate the comprehensive representation value of three difference features between visible light frames, between infrared frames, and within dual-modal frames for the i-th frame. Among them, represents the feature amplitude of the k-th sample point in the i-th infrared image, represents the feature amplitude of the k-th sample point in the i-th visible light image, represents the difference feature amplitude of the k-th sample point in the i-th infrared image.

4. The method for detecting significant frames of infrared and visible light video mimicry fusion based on evidence synthesis of possibility distribution according to claim 3, wherein: In the likelihood distribution synthesis module for the significance of differential feature changes, the segmentation points of the likelihood distribution function are first obtained using K-means clustering, and the Manhattan distance is adopted to divide the differential feature data into three clusters and the centroid μ of each cluster is obtained respectively. j , C j ={s i : |s i -μ j | = min(|s i -μ j |)}, where are the three difference features of the i-th frame. Secondly, based on the clustering centroid μ j , the following possibility distribution function is constructed to dynamically map the difference feature value of each frame to the significance possibility value of its degree of change. Then a new non-linear weighted synthesis method is proposed, and the pairwise synthesis method is used to obtain the synthesis result of the possibility distribution that comprehensively reflects the significant changes in the difference features between different frames and different modalities. Among them, α is the weight corresponding to the significance level combination, and β = 1 - α.

5. The method for detecting significant frames of infrared and visible light video mimicry fusion based on evidence synthesis of possibility distribution according to claim 4, wherein: In the D-S evidence synthesis decision module, the recognition framework is set as Θ = {low, medium, high}, and each type of differential feature is regarded as an evidence body, with the possibility distribution synthesis result as the input x, and the following possibility mass function is constructed. Normalize it. According to the discount coefficient method, the possibility mass function of the evidence body is updated to obtain the revised mass function m l ′(A), where m l ′(Θ) represents the uncertain part of the evidence body, and a l represents the importance weight of the evidence body; finally, according to the D-S evidence synthesis rule, the possibility mass functions m l ′(low), m l ′(medium), m l ′(high) and m l l(Θ) of each evidence body l are fused to obtain the synthetic mass value of each decision category. Finally, the decision result is determined according to the maximum mass value, and the frame with the decision result of high is judged as a significant frame.

6. The infrared and visible light video mimicry fusion significant frame detection method based on likelihood distribution evidence synthesis according to claim 5, characterized in that: In the D-S evidence synthesis decision-making module, two types of criteria, namely, within-evidence body and between-evidence body criteria, are adopted, and the PEOMETHEE-II method based on entropy weight is used to rank the importance of each evidence body; first, a criterion matrix g p (l) is constructed according to the evaluation criteria, where l = 1, 2, …, L represents the number of evidence bodies, p = 1, 2, …, N represents the number of criteria, L = N = 6, and the criterion matrix is normalized to c lp , and the entropy value E p of each criterion is calculated. Then, calculate the difference degree G of each criterion p , convert the entropy value into a positive information difference degree measure, and obtain the weight Ω of each criterion p , where G p = 1 - E p , which defines the possibility distribution of the preference function Among them, The preference order index of the evidence body \(m\) l and \(m\) q is further calculated. Then, by calculating the net flow \(\varPhi\) l , the relative importance of the evidence body in decision-making is obtained. Since the importance weight \(a\) of the evidence body l is an increasing function of the net flow, a right-skewed possibility distribution is defined Among them, \(a = \min(\eta\) l ), \(b = \max(\eta\) l ).