Infrared and visible light image fusion method based on T-WavKAN and C-Corr
By using T-WavKAN and C-Corr methods in infrared and visible image fusion, multi-resolution text semantic features and calculating visual feature correlations, the shortcomings of existing methods in processing deep semantic information and balancing visual features are solved, and more accurate traffic object detection and higher quality image fusion are achieved.
Patent Information
- Application Number
- CN202510053369.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-09
AI Technical Summary
The existing infrared and visible light image fusion methods are insufficient in processing text semantic information and balancing visual features, making it difficult to effectively capture deep semantic information in complex traffic scenes, and it is difficult to achieve the balance of infrared and visible visual features in the fusion results.
The infrared and visible image fusion method based on T-WavKAN and C-Corr is adopted to extract multi-resolution text semantic features through T-WavKAN, and the correlation of visual features is calculated in combination with the C-Corr module to achieve image fusion.
This method can effectively capture the complex text semantics in the image fusion context and establish local visual feature correlation between infrared and visible images, improving the accuracy of traffic object detection and the quality of the fused image.
Smart Images

Figure CN119963965A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to an infrared and visible light image fusion method based on T-WavKAN and C-Corr. Background Art
[0002] Infrared and visible image fusion (IVIF) involves integrating the same scene content captured by sensors of different modalities to produce a single composite image that contains comprehensive and detailed information. This process can integrate relevant data, eliminate redundancy, and improve the information quality and perception ability of the image, thereby facilitating high-level tasks such as traffic object detection. Notably, IVIF emphasizes salient objects in infrared images while preserving structural features and texture details in visible images. This fused image is particularly beneficial for traffic object detection because it mitigates the impact of weather and lighting interference and improves the accuracy of subsequent processing tasks. This capability is critical to ensure reliable performance in dynamic traffic environments, where visibility can be significantly affected under different conditions. Currently, IVIF tasks face two major limitations in image fusion: (1) they process textual semantic information superficially and fail to extract deeper semantic information that is helpful for identifying traffic objects; (2) they rely solely on a loss function to balance the relationship between visual features from infrared and visible images.
[0003] To address the first limitation, existing fusion methods usually utilize semantic information extracted by text encoders to guide visual features. However, the traditional linear modeling paradigm limits the extraction of core semantic information, which in turn restricts its ability to handle complex traffic scenes and weakens the effectiveness of its guided visual features.
[0004] Regarding the second limitation, current methods tend to simply concatenate the visual features obtained from the image encoder to generate a fused output. These methods mainly focus on the differences between the two modalities and tend to ignore the connection between shared and specific features, which is crucial for accurate detection of traffic targets. As a result, they find it difficult to achieve a balanced fusion of infrared and visible light visual features in the fusion results. Therefore, effectively processing textual semantics and visual features has become a key area for IVIF for traffic target detection.
[0005] To address the above challenges, this paper proposes a method for fusion of infrared and visible light images based on T-WavKAN and C-Corr. This method aims to effectively capture the complex text semantics in the context of image fusion and establish the correlation of local visual features between the two modalities. Summary of the invention
[0006] In order to overcome the problems in the background technology, the present invention provides an infrared and visible light image fusion method based on T-WavKAN and C-Corr.
[0007] To achieve the above object, the present invention is implemented by the following technical solution: a method for fusion of infrared and visible light images based on T-WavKAN and C-Corr, the symbols used are described as shown in Table 1, and the technical solution includes the following steps:
[0008] Table 1
[0009]
[0010]
[0011] S1. Capture corresponding images through infrared and RGB cameras of mobile vehicles to obtain infrared and visible light image datasets in different complex scenes, and generate corresponding text descriptions by using the large language model GPT4;
[0012] S2. Process the text description Text and use the text encoder ′E_T(·) to obtain the text feature t0: t0 = ′E_T(Tokenizer(Text));
[0013] S3. Based on the processing of t0 in S2, the structure potential space representation t is obtained by using the learnable activation function Φ in T-WavKAN. L-1 :
[0014]
[0015] S4, according to S3 L-1 Processing, using the separation operation S(·) to obtain the scale factor s and bias factor b: {s,b}=S(t L-1 );
[0016] S5. According to {I, V} processing, use the image encoder {′E_I(·), ′E_V(·)} to obtain visual features
[0017] {Atten_I (k) ,Atten_V (k)}:
[0018] Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V); where k=0,1,...,K-1;
[0019] S6. According to S5, (k) ,Atten_V (k)} is processed and input into the cross attention mechanism CA(·) to further extract the mutual guided visual features, and {atten_I (k) ,atten_V (k)}:
[0020] {atten_I (k) ,atten_V (k) =CA(Atten_I (k) ,Atten_V (k) );
[0021] S7, according to S6 {atten_I (k) ,atten_V (k)} and input it into the correlation layer C(·) to obtain the correlation weight graph
[0022]
[0023] S8. According to S7 Processing, using operation F(·) to fuse them, to obtain relevant deep representation features
[0024] S9. According to S8 Processing, using the self-attention mechanism SA(·) to enhance its representation ability, get f (k) :
[0025] S10, according to {f (k) ,s,b}, and use the semantic guidance operation G(·) to obtain the fusion feature f (k+1) :f (k+1) =G(f (k) ,s,b)=(E+f (k) )·s+b;
[0026] S11, according to S10 (k+1) Processing, through the fusion operation F(·) to obtain the final fusion representation feature
[0027] S12. According to S11 Processing, using Restormer decoder D(·) Decode and get the final fused image F:
[0028] S13, according to {I, V, F} processing, in the fusion loss function θ←▽ θ {L f}, a gradient descent operation is performed on it to obtain the optimal fusion result, where:
[0029] L f =α1·L c +α2·L i +α3·L g +α4·L s
[0030] sL c =||T(F)-T(V)||1 / HW
[0031] L i =||F-max(I,V)||1 / HW
[0032] L g =||▽F-max(▽I,▽V)||1 / HW
[0033] L s =(1-ssim(F,V))+λ·(1-ssim(F,I))
[0034] S14. Apply the fused results to the traffic target detection task for verification.
[0035] Furthermore, according to the infrared and visible light image fusion method based on T-WavKAN and C-Corr according to claim 1, it is characterized in that the shape structure of T-WavKAN in S3 is represented as: [n0,n1,...,n L-1 ]; when the learnable activation function of the (l,i) position unit and the (l+1,j) position unit is expressed as:
[0036] ψ l,j,i ,l=0,1,...,L-1,i=1,2,...,n l ,j=1,2,...,n l+1
[0037] Where n i for i th The number of nodes in the layer, (l,i) represents l th i within layer th Node; the activation value of the neuron at position (l+1,j) is simply the sum of all previous activation values. This process includes the basic operations of T-WavKAN:
[0038] Among them, t(l,i) is the activation feature of the neuron at position (l,i) in T-WavKAN, and its matrix form is:
[0039]
[0040] where Φ l Indicates the l in T-WavKAN th The function matrix of the layer.
[0041] Furthermore, the T-WavKAN combines wavelet transform to dynamically adapt to data points with different densities of text features; by expanding the wavelet basis in areas with dense data points, complex high-frequency details are captured; while shrinking the wavelet basis in areas with sparse data points, global information is used to discern overall low-frequency trends; T-WavKAN decomposes text semantics into multi-resolution features, which helps to fully extract structured latent space representations:
[0042]
[0043] Among them, E represents an all-one matrix, σ is the standard deviation of the Gaussian distribution, and ω is a learnable parameter that facilitates the mother wavelet to accurately approximate the shape of the target function.
[0044] After processing by T-WavKAN, we get the text feature t L-1 :
[0045]
[0046] Furthermore, the S5 provides a cross-perceptual visual feature correlation calculation C-Corr; Restormer is used as the encoder {′E_I(·),′E_V(·)}; Restormer extracts features while maintaining temporal consistency:
[0047] Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V)
[0048] Among them, k=0,1,...,K-1.
[0049] Furthermore, a correlation layer C(·) is designed in the C-Corr and applied to the S7 visual feature correlation calculation, as follows:
[0050]
[0051] Furthermore, after S13 is completed, the fusion method is integrated into the intelligent traffic monitoring system and directly embedded in the application module inside the existing traffic camera for use.
[0052] Beneficial effects of the present invention:
[0053] 1. The present invention proposes a text semantic extraction method T-WavKAN based on an improved WavKAN architecture, which introduces a trainable wavelet activation function for performing multi-resolution analysis of text features. In this way, T-WavKAN not only improves the adaptability to various text structures, but also optimizes the processing efficiency of complex scene content. Specifically, the present invention can effectively capture multi-level information in text data, thereby enhancing the flexibility and accuracy of text processing in different application scenarios. In addition, the design of T-WavKAN allows it to be adaptively adjusted according to specific task requirements, ensuring consistent high performance in a wide range of application environments.
[0054] 2. The present invention proposes a visual feature fusion module C-Corr based on cross-correlation perception. This module generates a correlated feature representation by calculating the correlation between feature maps from two different modalities. Subsequently, these correlated feature representations are fused with the original feature maps of their respective modalities, which not only effectively retains the shared feature information, but also highlights the unique characteristics of each modality. Specifically, the C-Corr module first evaluates and quantifies the relationship between the feature maps of the two modalities to capture the potential synergy between the two. Then, by combining these correlated features with the original feature maps, the expressiveness of the fused image, which has both common features and information specific to each modality, is enhanced. This method ensures that the fusion result contains both rich common features and unique details of each modality, thereby providing better quality basic data for subsequent processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0056] Figure 1 It is a schematic diagram of the process of the infrared and visible light image fusion method provided by the present invention;
[0057] Figure 2 It is a schematic diagram of the network structure of the present invention provided by the present invention;
[0058] Figure 3 This is an example diagram of the fusion method provided by the present invention applied to a traffic target detection task;
[0059] Figure 4 This is a comparison chart of the fusion effects provided by the present invention and 18 most advanced infrared and visible light image fusion methods. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] See also Figures 1 to 4 The present invention discloses a method for fusion of infrared and visible light images based on T-WavKAN and C-Corr, comprising the following steps, wherein the symbols used are described as shown in Table 1:
[0062] Table 1
[0063]
[0064]
[0065] S1. Capture corresponding images through infrared and RGB cameras of mobile vehicles to obtain infrared and visible light image datasets in different complex scenes, and generate corresponding text descriptions by using the large language model GPT4;
[0066] S2, according to the text description Text processing, use the text encoder 'E_T(·) to obtain the text feature t0: t0 = 'E_T(Tokenizer(Text)); The principle of extracting the text description feature in S2 is: first, the text description is converted into a LongTensor representation, which encapsulates the token sequence in the text. Then, the text encoder 'E_T(·) processes the tensor representation to generate the text feature t0 for subsequent analysis:
[0067] t0=′E_T(Tokenizer(Text))
[0068] S3. Based on the processing of t0 in S2, the structure potential space representation t is obtained by using the learnable activation function Φ in T-WavKAN. L-1 : The principle of the T-WavKAN multi-scale learnable activation function in S3 is: The shape structure of T-WavKAN is represented as: [n0,n1,...,n L-1 ]. Assume that the learnable activation function connecting the (l,i) position unit and the (l+1,j) position unit is expressed as:
[0069] ψ l,j,i ,l=0,1,...,L-1,i=1,2,...,n l ,j=1,2,...,nl+1
[0070] Where n i for i th The number of nodes in the layer, (l,i) represents l th i within layer th Node. The activation value of the neuron at position (l+1,j) is simply the sum of all previous activation values. This process includes the basic operations of T-WavKAN:
[0071]
[0072] Among them, t(l,i) is the activation feature of the neuron at position (l,i) in T-WavKAN, and its matrix form is:
[0073]
[0074] where Φ l Indicates the l in T-WavKAN th The function matrix of the layer.
[0075] T-WavKAN combines wavelet transform to dynamically adapt to data points with different densities of text features. This method captures complex high-frequency details by expanding the wavelet basis in areas with dense data points, and shrinking the wavelet basis in areas with sparse data points, using global information to discern overall low-frequency trends. T-WavKAN decomposes text semantics into multi-resolution features, which helps to fully extract structured latent space representations:
[0076]
[0077] Where E represents an all-one matrix, σ is the standard deviation of the Gaussian distribution, and ω is a learnable parameter that allows the mother wavelet to accurately approximate the shape of the target function.
[0078] After processing by T-WavKAN, we get the text feature t L-1 :
[0079]
[0080] S4, according to S3 L-1 Processing, using the separation operation S(·) to obtain the scale factor s and bias factor b: {s,b}=S(t L-1 ); The principle of the text feature separation operation in S4 is: generating a scale factor s and a deviation factor b, {s, b} is used to enhance the interaction between semantic vision in the fusion framework and improve the expression ability of the fusion feature:
[0081] {s,b}=S(t L-1 )
[0082] S5. According to {I, V} processing, use the image encoder {′E_I(·), ′E_V(·)} to obtain visual features
[0083] {Atten_I (k) ,Atten_V (k)}:
[0084] Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V);
[0085] Where k = 0, 1, ..., K-1;
[0086] The principle of image feature extraction in S5 is: in the multimodal image fusion task, there are significant differences in the intrinsic structures between different modalities, and the content representation is diverse. In order to balance the modal similarities and differences between infrared and visible light images and improve the fusion effect, the present invention designs a cross-perception visual feature correlation calculation (C-Corr). Restormer is used as the encoder {′E_I(·),′E_V(·)}. Restormer is good at effectively extracting features while maintaining temporal consistency:
[0087] Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V)
[0088] Among them, k=0,1,...,K-1.
[0089] S6. According to S5, (k) ,Atten_V (k)} and input it into the cross attention mechanism CA(·) to further extract the mutual guided visual features.
[0090] {atten_I (k) ,atten_V (k)}:
[0091] {atten_I (k) ,atten_V (k) =CA(Atten_I (k) ,Atten_V (k) );
[0092] The principle of cross-perception of visual features in S6 is: shallow features {Atten_I (k) ,Atten_V (k)} are fed into a cross-attention mechanism that aims to promote interaction between features, allowing infrared and visible light images to mutually guide the extraction of visual features, ensuring that the fused representation captures complementary information from both modalities:
[0093] {atten_I (k) ,atten_V (k) =CA(Atten_I (k) ,Atten_V (k) )
[0094] S7, according to S6 {atten_I (k) ,atten_V (k)} and input it into the correlation layer C(·) to obtain the correlation weight graph
[0095]
[0096] The principle of calculating the correlation of the S7 visual features is: considering {atten_I (k) ,atten_V (k)}, the present invention designs a correlation layer C(·) in C-Corr. C(·) aims to learn and capture the complex correlation between the two modalities, thereby enhancing the overall feature representation and improving its robustness and discrimination ability, as follows:
[0097]
[0098] Where m atten ∈[-r,r] 2 Atten_I (k) or Atten_V (k) m iterations in the domain [-r]×[r].
[0099] S8. According to S7 Processing, using operation F(·) to fuse them, to obtain relevant deep representation features
[0100] The principle of the fusion operation in S8 is to use the fusion operation (concatenate) F(·) to strengthen the local correlation between infrared and visible light features and emphasize the local differences in the resulting deep representation:
[0101]
[0102] S9. According to S8 Processing, using the self-attention mechanism SA(·) to enhance its representation ability, get f (k) :
[0103] The principle of the visual feature self-attention mechanism in S9 is: the features extracted by C-Corr Under the action of the self-attention mechanism SA(·), it can undergo an initial self-enhancement process, which aims to enhance the representation ability of features and lay the foundation for the subsequent visual language interaction stage:
[0104]
[0105] S10, according to {f (k) ,s,b}, and use the semantic guidance operation G(·) to obtain the fusion feature f (k+1) :f (k+1) =G(f (k) ,s,b)=(E+f (k) )·s+b;
[0106] The principle of the text semantics-guided visual feature fusion mechanism in S10 is: using the semantic factors {s, b} derived from T-WavKAN, visual fusion features can be guided from multiple perspectives to inject text semantics into the fusion representation f (k) middle:
[0107] f (k+1) =G(f (k) ,s,b)=(E+f (k) )·s+b
[0108] S11, according to S10 (k+1) Processing, through the fusion operation F(·) to obtain the final fusion representation feature
[0109] The principle of the fusion representation feature operation in S11 is: the final fusion representation feature is obtained by the fusion operation F(·) This method can improve the comprehensiveness of fusion information from multiple perspectives of text and images:
[0110]
[0111] S12. According to S11 Processing, using Restormer decoder D(·) Decode and get the final fused image F:
[0112] The principle of the decoder in S12 is: the Restormer block D(·) is used to fuse the representation Decode to generate the final fused image F:
[0113]
[0114] S13, according to {I, V, F} processing, in the fusion loss function θ←▽ θ {L f}, a gradient descent operation is performed on it to obtain the optimal fusion result, where:
[0115] L f =α1·L c +α2·L i +α3·L g +α4·L s
[0116] sL c =||T(F)-T(V)||1 / HW
[0117] L i =||F-max(I,V)||1 / HW
[0118] L g =||▽F-max(▽I,▽V)||1 / HW
[0119] L s =(1-ssim(F,V))+λ·(1-ssim(F,I))
[0120] The principle of the fusion loss function in S13 is: using color consistency loss L c , strength loss L i , gradient loss L g , structural similarity loss L s as a comprehensive loss function.
[0121] Wherein, the color consistency loss: In order to maintain the color information and facilitate further processing, the present invention uses operation T(·) to convert the fused image and the visible light image from the RGB space to the YCbCr space:
[0122] L c =||T(F)-T(V)||1 / HW
[0123] The intensity loss is intended to highlight important objects and is specifically defined as follows:
[0124] L i =||F-max(I,V)||1 / HW
[0125] The gradient loss is able to determine the optimal texture distribution in regions where pixel intensity varies significantly and is defined as follows:
[0126] L g=||▽F-max(▽I,▽V)||1 / HW
[0127] The structural similarity loss has the ability to quantify the similarity between images according to the brightness, structure and contrast dimensions, and is defined as follows:
[0128] L s =(1-ssim(F,V))+λ·(1-ssim(F,I))
[0129] Here, λ is the balancing factor.
[0130] Finally, the fusion loss is defined as follows:
[0131] L f =α1·L c +α2·L i +α3·L g +α4·L s
[0132] Among them, {α1, α2, α3, α4} represent hyperparameters.
[0133] S14, applying the fusion result to the traffic target detection task for verification; the principle of the high-level task verification in S14 is: after the completion of S13, the fusion method is integrated into the intelligent traffic monitoring system and directly embedded into the application module inside the existing traffic camera to ensure that the fusion method is applied in various real traffic environments.
[0134] It should be noted that, see Figures 1 to 4 As shown, the above steps S2-S13 obtain Figure 2 , which is a schematic diagram of the network structure of the present invention, and step S14 obtains Figure 3 The example diagram shows that the fusion method provided by the present invention is applied to the traffic target detection task, which effectively captures the multi-level information in the text data, thereby enhancing the flexibility and accuracy of text processing in different application scenarios. In addition, the design of T-WavKAN allows it to be adaptively adjusted according to specific task requirements, ensuring consistent high performance in a wide range of application environments. By calculating the fusion results of different advanced methods and applying them to different objective indicators, we can obtain Figure 4 ; It is a comparison chart of the fusion effect with 18 most advanced infrared and visible light image fusion methods. Compared with 18 existing most advanced image fusion technologies, the present invention shows significant performance advantages on multiple data sets. In particular, on three publicly released data sets containing infrared and visible light images (MSRS, FLIR and M 3In addition, the present invention is further demonstrated to be effective and superior through a series of high-level visual tasks, including but not limited to object detection, semantic segmentation, and instance segmentation.
[0135] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for fusion of infrared and visible light images based on T-WavKAN and C-Corr, characterized in that: The following steps are included, wherein the symbols used are described as shown in Table 1: Table 1 S1. Capture corresponding images through infrared and RGB cameras of mobile vehicles to obtain infrared and visible light image datasets in different complex scenes, and generate corresponding text descriptions by using the large language model GPT4; S2. Process the text description Text and use the text encoder E_T(·) to obtain the text feature t0: t0 = E_T(Tokenizer(Text)); S3. Based on the processing of t0 in S2, the structure potential space representation t is obtained by using the learnable activation function Φ in T-WavKAN. L-1 : S4, according to S3 L-1 Processing, using the separation operation S(·) to obtain the scale factor s and bias factor b: {s,b}=S(t L-1 ); S5. According to {I, V} processing, use the image encoder {′E_I(·), ′E_V(·)} to obtain visual features: {Atten_I (k) ,Atten_V (k) }: Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V); Where k = 0, 1, ..., K-1; S6. According to S5, (k) ,Atten_V (k) } and input it into the cross attention mechanism CA(·) to further extract the mutual guided visual features. {atten_I (k) ,atten_V (k) }: {atten_I (k) ,atten_V (k) }=CA(Atten_I (k) ,Atten_V (k) ); S7, according to S6 {atten_I (k) ,atten_V (k) } and input it into the correlation layer C(·) to obtain the correlation weight graph S8. According to S7 Processing, using operation F(·) to fuse them, to obtain relevant deep representation features S9. According to S8 Processing, using the self-attention mechanism SA(·) to enhance its representation ability, get f (k) : S10, according to {f (k) ,s,b}, and use the semantic guidance operation G(·) to obtain the fusion feature f (k+1) :f (k+1) =G(f (k) ,s,b)=(E+f (k) )·s+b; S11, according to S10 (k+1) Processing, through the fusion operation F(·) to obtain the final fusion representation feature S12. According to S11 Processing, using Restormer decoder D(·) Decode and get the final fused image F: S13, according to {I, V, F} processing, in the fusion loss function The gradient descent operation is performed on to obtain the optimal fusion result, where: L f =α1·L c +α2·L i +α3·L g +α4·L s s.t.L c =||T(F)-T(V)||1 / HW L i =||F-max(I,V)||1 / HW L s =(1-ssim(F,V))+λ·(1-ssim(F,I)) S14. Apply the fused results to the traffic target detection task for verification.
2. According to claim 1, the infrared and visible light image fusion method based on T-WavKAN and C-Corr is characterized in that: The shape structure of T-WavKAN in S3 is represented as: [n0,n1,...,n L-1 ]; when the learnable activation function of the (l,i) position unit and the (l+1,j) position unit is expressed as: ψ l,j,i ,l=0,1,...,L-1,i=1,2,...,n l ,j=1,2,...,n l+1 Where n i for i th The number of nodes in the layer, (l,i) represents l th i within layer th Node; the activation value of the neuron at position (l+1,j) is simply the sum of all previous activation values. This process includes the basic operations of T-WavKAN: Among them, t(l,i) is the activation feature of the neuron at position (l,i) in T-WavKAN, and its matrix form is: where Φ l Indicates l in T-WavKAN th The function matrix of the layer.
3. The infrared and visible light image fusion method based on T-WavKAN and C-Corr according to claim 2, characterized in that: The T-WavKAN combines wavelet transform to dynamically adapt to data points with different densities of text features; it captures complex high-frequency details by expanding the wavelet basis in areas with dense data points; and shrinks the wavelet basis in areas with sparse data points, using global information to discern overall low-frequency trends; T-WavKAN decomposes text semantics into multi-resolution features, which helps to fully extract structured latent space representations: Among them, E represents an all-one matrix, σ is the standard deviation of the Gaussian distribution, and ω is a learnable parameter that facilitates the mother wavelet to accurately approximate the shape of the target function. After processing by T-WavKAN, we get the text feature t L-1 :
4. The infrared and visible light image fusion method based on T-WavKAN and C-Corr according to claim 1, characterized in that: The S5 provides a cross-perceptual visual feature correlation calculation C-Corr; Restormer is used as the encoder {′E_I(·),′E_V(·)}; Restormer extracts features while maintaining temporal consistency: Atten_I (k) =′E_I(I);Atten_V (k) =′E_V(V) Among them, k=0,1,...,K-1.
5. The infrared and visible light image fusion method based on T-WavKAN and C-Corr according to claim 4, characterized in that: A correlation layer C(·) is designed in the C-Corr and applied to the S7 visual feature correlation calculation, as follows:
6. The infrared and visible light image fusion method based on T-WavKAN and C-Corr according to any one of claims 1 to 5, characterized in that: After S13 is completed, the fusion method is integrated into the intelligent traffic monitoring system and directly embedded in the application module inside the existing traffic camera for use.
Citation Information
Cited By
Chemical adding control method, system and equipment in water treatment process and medium
CN121974416A
A method, system, apparatus, and medium for chemical dosage control of a water treatment process
CN121974416B