A method and system for infrared-visible light person re-identification based on wavelet representation

Through the cross-modal person re-identification model represented by wavelet, the wavelet adaptive encoder and frequency enhancement module are used to process high and low frequency features, and combined with the wavelet frequency domain attention mechanism and enhanced similarity distribution clustering loss, the problem of inter-modal spectrum differences in VI-ReID is solved, and the accuracy and efficiency of infrared-visible light person re-identification are improved.

CN120599667BActive Publication Date: 2025-10-03JIANGNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511089689.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-03
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing VI-ReID methods ignore the multi-scale information contained in the frequency domain features of visible light and infrared images and the inherent spectral differences between modalities, resulting in poor cross-modal feature extraction and alignment.

Method used

A cross-modal person re-identification model based on wavelet representation is adopted. The wavelet adaptive encoder and frequency enhancement module are used to process high and low frequency features respectively, and the wavelet frequency domain attention mechanism and enhanced similarity distribution clustering loss are used to optimize the feature space structure.

Benefits of technology

It significantly improves the accuracy and efficiency of infrared-visible light person re-identification, enhances the model's robustness to different lighting conditions and imaging quality, and solves the problem of spectral differences between modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599667B_ABST
    Figure CN120599667B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of person re-identification. The present invention provides an infrared-visible light person re-identification method and system based on wavelet representation. A wavelet adaptive encoder is designed, and features are decomposed into different frequency sub-bands through discrete wavelet transform. High-frequency and low-frequency information are processed respectively by high- and low-wave perception units, and the importance of features in different frequency bands is adaptively adjusted using a wavelet frequency domain attention mechanism. A frequency enhancement module is introduced, and global spectral characteristics are modeled using Fourier transform. Deep fusion of frequency domain and spatial domain information is achieved through Einstein matrix multiplication, thereby enhancing the model's robustness to different lighting conditions and imaging qualities. An enhanced similarity distribution clustering loss is proposed, which not only brings cross-modal features of the same identity closer together, but also explicitly pushes away features of different identities within the same modality, constructing a more robust feature space, effectively solving the problem of infrared image homogeneity, and enhancing the discriminative ability of the feature space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of person re-identification, and in particular to an infrared-visible light person re-identification method and system based on wavelet representation. Background Art

[0002] Visible-infrared person re-identification (VI-ReID) aims to match images of people from different spectral modalities (visible light and infrared) and enable mutual retrieval between visible and infrared images. This technology holds broad application prospects in all-weather surveillance, public security, and smart city development. However, due to significant differences in the imaging principles and conditions of visible and infrared images: visible light images record the visible spectrum reflected by objects, containing rich color and texture details; while infrared images capture the thermal radiation emitted by objects, primarily displaying temperature distribution. VI-ReID still faces significant challenges. The core challenge lies in mapping the features of the two modalities into a common latent space and effectively aligning and extracting features from the two modalities.

[0003] Existing VI-ReID research falls into two main paths: feature alignment and auxiliary information generation. Feature alignment methods map cross-modal features into a unified semantic space through metric learning or enhanced feature extraction components. For example, some researchers have designed complex metric learning techniques, while others focus on improving network architectures. Mapping comprehensive cross-modal features into a common semantic space reduces cross-modal discrepancies, but neglects to leverage modality-specific and shared cues, inevitably leading to performance bottlenecks. Auxiliary information generation methods provide supplementary knowledge through additional models to compensate for modality-specific or modality-shared features at the embedding or pixel level. Examples include using GANs to generate compensatory features at the image or embedding level, introducing X-modality as an intermediate representation, or simultaneously reducing discrepancies at both the image and feature levels through style alignment and feature adjustment. While these methods have made progress, they inevitably introduce losses and noise during the generation process or require additional models for data processing, reducing efficiency and convenience. Therefore, the VI-ReID community urgently needs more comprehensive and efficient feature extraction and modality alignment methods.

[0004] Existing VI-ReID methods mainly focus on network architecture design or auxiliary information generation, but ignore the multi-scale information contained in the frequency domain features of visible light and infrared images and the inherent spectral differences between modalities. Figure 1 As shown in (a) of Figure 1, the difference between visible light and infrared images is particularly evident in the frequency domain: the spectrum of visible light images contains more high-frequency information, while infrared images are dominated by low-frequency information. Traditional methods typically address these differences directly in the spatial domain, making it difficult to effectively model the inherent spectral gap between the modalities.

[0005] Frequency domain analysis offers a new perspective for overcoming modal differences. By decomposing an image into the frequency domain, modality-specific spectral characteristics can be more accurately processed. However, existing frequency domain methods, such as the frequency domain adaptive framework, primarily focus on single-scale frequency domain analysis and fail to fully exploit multi-scale frequency domain characteristics. Wavelet transforms, on the other hand, provide both frequency and spatial information, making them particularly well-suited for analyzing non-stationary signals and capturing multi-scale features. Figure 1 (b) in the figure shows how the wavelet transform decomposes the image into frequency bands in different directions, enabling the model to process high- and low-frequency information in a targeted manner, thereby more effectively capturing cross-modal shared features.

[0006] As a powerful signal processing tool, the wavelet transform (WMT) can simultaneously provide both frequency and spatial information, making it particularly well-suited for analyzing non-stationary signals and capturing multi-scale features. In the field of computer vision, numerous studies have utilized the WMT to enhance visual representation learning, and this approach has been extended to tasks such as style transfer and image generation. The WMT's multi-scale decomposition capability and reversibility make it particularly effective in image restoration. The multi-level wavelet CNN (MWCNN) proposed by Liu et al. utilizes the WMT to reduce feature map size, maintaining performance while improving computational efficiency. The DWSR method proposed by Guo et al. utilizes low-resolution wavelet subbands and recovers high-resolution details by predicting subband residuals. In image deraining and dehazing research, wavelet decomposition has been shown to effectively separate information at different scales while preserving image integrity. However, the application of the WMT in visible-infrared person re-identification (VIR) remains limited, and existing VI-ReID methods rarely explore its potential for handling spectral differences between modalities.

[0007] In summary, how to use wavelet transform to effectively solve the modal difference problem in infrared-visible light person re-identification method, thereby improving the accuracy and efficiency of infrared-visible light person re-identification. Summary of the Invention

[0008] To this end, the technical problem to be solved by the present invention is to overcome the fact that the VI-ReID method in the existing technology mainly focuses on network architecture design or auxiliary information generation, ignoring the multi-scale information contained in different frequency domain features and the inherent spectral differences between modalities, and therefore does not effectively utilize wavelet transform to solve the modal difference problem in the infrared-visible light person re-identification method.

[0009] In order to solve the above technical problems, the present invention provides an infrared-visible light person re-identification method based on wavelet representation, comprising: inputting a visible light image and an infrared image into a cross-modal person re-identification model based on wavelet representation, wherein the cross-modal person re-identification model based on wavelet representation includes a feature extractor and a backbone network, wherein the backbone network includes multiple groups of alternately arranged wave adaptive encoders and frequency enhancement modules; using the feature extractor to extract the initial features of the visible light image and the infrared image, splicing them and then inputting them into the backbone network; for each wave adaptive encoder, respectively extracting the high-frequency features and low-frequency features of the input features of the wave adaptive encoder; The features are concatenated, discrete wavelet transform is performed on the concatenated features, and the wave frequency domain attention mechanism is used to obtain the attention coefficients of different wavelet coefficients. After generating the attention weight map, the inverse discrete wavelet transform is used to obtain the output features of the wave adaptive encoder; for each frequency enhancement module, the global frequency domain features and spatial domain weight features of the input features of the frequency enhancement module are obtained respectively through fast Fourier transform and depthwise separable convolution, and after fusing the global frequency domain features and spatial domain weight features, the output features of the frequency enhancement module are obtained through inverse fast Fourier transform; the person re-identification result is generated based on the output features of the backbone network.

[0010] Preferably, for each wave adaptive encoder:

[0011] A high-frequency sensing unit is used to extract the high-frequency features of the input features of the wavelet adaptive encoder; a low-frequency sensing unit is used to extract the low-frequency features of the input features of the wavelet adaptive encoder; the high-frequency features are spliced ​​with the low-frequency features, and the spliced ​​features are decomposed into different frequency sub-bands using discrete wavelet transform to obtain multiple wavelet coefficients; the wavelet frequency domain attention mechanism is used to obtain the attention coefficients of different wavelet coefficients, generate fused attention features, and obtain an attention weight map; based on the attention weight map, the spatial domain features are reconstructed using inverse discrete wavelet transform to obtain the output features of the wavelet adaptive encoder.

[0012] Preferably, the high-wave perception unit includes a residual network layer, a discrete wavelet transform unit, a depthwise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer connected in series along the propagation direction; wherein the output of the residual network layer is jump-connected to the output of the layer normalization layer;

[0013] The low-wave perception unit includes a discrete wavelet transform unit, a depth-wise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer, which are sequentially connected in series along the propagation direction; wherein, the input of the low-wave perception unit is jump-connected with the output of the layer normalization.

[0014] Preferably, the expression of the high-frequency sensing unit is:

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] ;

[0020] ;

[0021] in, is the input feature of the high-wave perception unit; is the feature after processing by the residual network layer, is the residual network layer processing function; 、 、 and They are Low-frequency approximation coefficients, horizontal high-frequency detail coefficients, vertical high-frequency detail coefficients, and diagonal high-frequency detail coefficients obtained by discrete wavelet transform decomposition; is the discrete wavelet transform function; 、 、 and They are 、 、 、 The low-frequency approximation coefficients, horizontal high-frequency detail coefficients, vertical high-frequency detail coefficients, and diagonal high-frequency detail coefficients after processing by the depthwise separable convolution layer; It is a 3×3 depth-separable convolution operation function; is the inverse discrete wavelet transform; For 、 、 and The spatial domain features output after inverse discrete wavelet transform; is the layer normalization function; For Features after normalization; It is the high-frequency feature output by the high-frequency perception unit.

[0022] Preferably, the expression of the wave frequency domain attention mechanism is:

[0023] ;

[0024] ;

[0025] ;

[0026] in, Represent the horizontal high-frequency wavelet coefficients, low-frequency wavelet coefficients, and vertical high-frequency wavelet coefficients of the input wave frequency domain attention mechanism respectively; is the depth-separable convolution operation function; They are the horizontal high-frequency attention coefficient, low-frequency attention coefficient, and vertical high-frequency attention coefficient; is the fused attention feature; is a 1×1 convolution operation function; is the feature concatenation operator, connecting along the channel dimension; is the Sigmoid activation function; This is the attention weight map output by the frequency domain attention mechanism.

[0027] Preferably, for each frequency enhancement module:

[0028] The input features of the frequency enhancement module are subjected to fast Fourier transform to obtain the global spectrum features; the spatial domain weight features of the input features of the frequency enhancement module are obtained through depthwise separable convolution; the global spectrum features and spatial domain weight features are fused using Einstein matrix multiplication, and the features are reconstructed through inverse Fourier transform to obtain the output features of the frequency enhancement module.

[0029] Preferably, after fusing the global spectrum features and the spatial domain weight features using Einstein matrix multiplication, the features are reconstructed by inverse Fourier transform to obtain the output features of the frequency enhancement module, including:

[0030] The imaginary features and real features of the global frequency domain features and the imaginary weights and real weights of the spatial domain weight features are input into the first Einstein matrix multiplication layer, and the first imaginary features, the first real features, the first imaginary weights and the first real weights are output; the first imaginary features and the first real features are processed by a nonlinear activation function to obtain enhanced imaginary features and enhanced real features; the first imaginary weights and the first real weights are processed by depthwise separable convolution to obtain enhanced imaginary weights and enhanced real weights; the enhanced imaginary features, enhanced real features, enhanced imaginary weights and enhanced real weights are input into the second Einstein matrix multiplication layer, and the target imaginary features, target real features, target imaginary weights and target real weights are output; the target imaginary features and target real features are inverse Fourier transformed to obtain the output features of the frequency enhancement module.

[0031] Preferably, the loss function of the cross-modal person re-identification model based on wavelet representation is an enhanced similarity distribution clustering loss, including: similarity distribution clustering loss, visible light modality KL divergence loss and infrared modality KL divergence loss;

[0032] The construction process of the visible light modal KL divergence loss and the infrared modal KL divergence loss includes:

[0033] Calculate the visible light mode The sample and The similarity probability of samples :

[0034] ;

[0035] Calculate the infrared modality The sample and The similarity probability of samples :

[0036] ;

[0037] in, The first The sample and Similarity matrix of samples; The first The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; is the temperature coefficient; is the total number of samples;

[0038] Calculate the visible light mode The sample and The ideal similarity distribution of samples :

[0039] ;

[0040] Calculate the infrared modality The sample and The ideal similarity distribution of samples :

[0041] ;

[0042] in, The first The sample and The identity label indicator of each sample, The first The sample and Identity tag indicator for each sample; Infrared mode The sample and The identity label indicator of each sample, Infrared mode The sample and The identity label indicator of the samples; when the identity label indicator is 0, it indicates that the two samples do not belong to the same identity; when the identity label indicator is 1, it indicates that the two samples belong to the same identity;

[0043] Constructing visible light modality KL divergence loss :

[0044] ;

[0045] Constructing infrared modality KL divergence loss :

[0046] ;

[0047] in, is a constant.

[0048] Preferably, the loss function of the cross-modal person re-identification model based on wavelet representation is:

[0049] ;

[0050] ;

[0051] in, is the total loss function; To enhance the similarity distribution clustering loss; Classify loss for identity; is the triplet loss; is the cross-modal pairing loss; is the similarity clustering distribution loss; is the KL divergence loss of the visible light mode; Infrared modality KL divergence loss, is a constant.

[0052] The present invention also provides an infrared-visible light person re-identification system based on wavelet representation, comprising: a memory for storing a computer program; and a processor for implementing the steps of the infrared-visible light person re-identification method based on wavelet representation as described above when executing the computer program.

[0053] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0054] To achieve effective multi-scale frequency domain feature extraction, the present invention designs a wavelet-adaptive encoder that decomposes features into different frequency subbands through discrete wavelet transforms, capturing multi-scale feature representations. The high / low wavelet perception units in the wavelet-adaptive encoder process high- and low-frequency information separately, while the wavelet frequency domain attention mechanism adaptively adjusts the importance of features in different frequency bands. The wavelet-adaptive encoder design enables the model to adopt specialized processing strategies for the different frequency characteristics of infrared and visible light images, effectively adapting to the different frequency characteristics of the two modalities.

[0055] Taking into account the limitations of wavelet transform in global spectrum representation, the present invention further designs a frequency enhancement module, which uses fast Fourier transform to model global spectrum characteristics to make up for the limitations of wavelet transform in spectrum representation. The frequency enhancement module uses deep separable convolution to extract spatial domain weight features, and realizes deep fusion of frequency domain and spatial domain feature information through Einstein matrix multiplication, thereby enhancing the robustness of the model to different lighting conditions and imaging quality. The frequency enhancement module can capture global frequency domain information that wavelet transform cannot effectively represent, provide more comprehensive spectrum support for cross-modal feature learning, and significantly improve the global spectrum modeling capability and environmental adaptability of the model.

[0056] To address the problem that traditional similarity distribution clustering loss focuses only on cross-modal alignment while ignoring intramodal differentiation, this paper proposes an enhanced similarity distribution clustering loss. This loss constructs a more robust feature space by simultaneously optimizing cross-modal feature alignment and intramodal feature differentiation. The enhanced similarity distribution clustering loss implements bidirectional constraints through the KL divergence framework. It not only considers the close proximity of cross-modal features of the same identity, but also explicitly pushes away features of different identities within the same modality. This effectively addresses the homogeneity issue of infrared images, enhances the discriminative power of the feature space, and significantly improves recognition performance on large-scale datasets.

[0057] In summary, the present invention provides a cross-modal person re-identification model based on wavelet representation, which realizes targeted multi-scale frequency domain feature decomposition through the wavelet adaptive encoder, and adaptively processes different frequency band information in combination with the wavelet frequency domain attention mechanism. The wavelet adaptive encoder works together with the frequency enhancement module to achieve comprehensive frequency domain feature modeling, and optimizes the feature space structure through the enhanced similarity distribution clustering mechanism, effectively improving the accuracy and efficiency of infrared-visible light person re-identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0059] Figure 1(a) is a comparison of the frequency domain capability distribution of visible light and infrared images;

[0060] Figure 1 (b) is a schematic diagram of the decomposition of an RGB image into four frequency bands: diagonal high frequency, vertical high frequency, horizontal high frequency, and low frequency after wavelet transform;

[0061] Figure 2 A schematic diagram of the structure of a cross-modal person re-identification model based on wavelet representation provided by an embodiment of the present invention;

[0062] Figure 3 Schematic diagram of the structure of the wave adaptive encoder;

[0063] Figure 4 It is a structural diagram of the high / low wave sensing unit;

[0064] Figure 5 It is a structural diagram of the frequency enhancement module;

[0065] Figure 6 Schematic diagram of the composition of enhanced similarity distribution clustering loss;

[0066] Figure 7 Comparison chart of retrieval results between the WaveRID model and the baseline model. DETAILED DESCRIPTION

[0067] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0068] Because infrared and visible light images differ significantly in the frequency domain: infrared images primarily contain low-frequency information, while visible light images are rich in high-frequency details, the multi-scale frequency domain decomposition characteristics of the wavelet transform provide a new technical approach for solving cross-modal person re-identification. This paper recognizes that modality-shared information (such as pedestrian outlines and motion characteristics) can be considered basic features, while modality-specific information (color and texture details of the RGB modality and thermal characteristics of the IR modality) are detail features. The integration and optimization of these two features are crucial. Based on this insight, this paper introduces wavelet transform into the VI-ReID task and proposes a cross-modal person re-identification model based on wavelet representation (Wavelet Representation for Infrared-Visible Domain Re-Identification, WaveRID). By combining wavelet transform and frequency domain processing, this paper addresses the modality difference problem in infrared-visible light person re-identification.

[0069] The core concept of the WaveRID model is to perform feature extraction and fusion simultaneously in the multi-scale wavelet domain and frequency domain to effectively capture the shared feature representation of infrared and visible light images. Figure 2 As shown in Figure 1, the WaveRID model consists of a feature extractor and a backbone network consisting of multiple alternating sets of Wave Adaptive Encoders (WAEs) and Frequency Enhancement Blocks (FEBs). The input infrared and visible light images are first passed through a shared feature extractor to generate initial feature representations. These features are then processed by the backbone network to extract multi-scale features. This hierarchical feature extraction approach significantly outperforms single-domain processing methods, capturing both local texture and global structural information simultaneously and providing a more comprehensive feature representation for cross-modal matching.

[0070] Since there are significant modal differences between infrared images and visible light images, it is necessary to gradually narrow the modal differences through multi-level feature transformation. Preferably, in an embodiment of the present invention, the backbone network includes three groups of wavelet adaptive encoders and frequency enhancement modules that are alternately arranged, realizing a progressive feature extraction process from coarse to fine. In the backbone network, each group of wavelet adaptive encoders and frequency enhancement modules undertakes a specific feature transformation task, thereby realizing a complete conversion process from the bottom-level pixel-level difference to the high-level semantic-level alignment, and ultimately achieving effective cross-modal feature alignment. Among them, the first group is mainly responsible for capturing the basic frequency domain feature decomposition and global spectrum information, and establishing a preliminary cross-modal feature mapping relationship. The second group performs feature refinement processing on this basis, and further enhances the feature alignment capability between modalities through deeper wavelet decomposition and frequency domain enhancement. The third group further optimizes the feature representation, extracts cross-modal shared features with high discriminability, and ensures the quality and robustness of the final feature representation.

[0071] Three alternating sets of wave adaptive encoders and frequency enhancement modules provide sufficient network depth for modal alignment, enabling the model to learn the complex mapping relationship from modality-specific features to modality-shared features layer by layer. From the perspective of training stability, the three-set setting ensures that the network has sufficient expressive power while effectively avoiding the gradient vanishing or exploding problems that may be caused by an overly deep network. The residual connection mechanism and attention mechanism within each group further ensure the stable propagation of the gradient and ensure the effective training of the network. This architectural design enables the WaveRID model to achieve a deep understanding and precise extraction of cross-modal features while maintaining computational efficiency, providing the optimal network depth configuration for infrared-visible light person re-identification.

[0072] The infrared-visible light person re-identification method based on wavelet representation provided in the embodiment of the present invention specifically includes the following steps:

[0073] S1: Input the infrared image and visible light image into the shared feature extractor, extract the initial features of the infrared image and visible light image respectively, and splice them as the output of the feature extractor and input them into the backbone network.

[0074] The feature extractor uses parallel feature extraction branches to extract the initial features of the infrared image and the visible light image respectively, and then splices them together, and uses the spliced ​​features as the output of the feature extractor.

[0075] The parallel feature extraction branches include convolutional layer, BN layer (Batch Normalization), nonlinear activation function, and maximum pooling layer.

[0076] S2: The high-frequency and low-frequency features of the input features of the wavelet adaptive encoder are extracted using the high-frequency perception unit and the low-frequency perception unit of the wavelet adaptive encoder respectively, and then spliced. The spliced ​​features are subjected to discrete wavelet transform, and the wavelet frequency domain attention mechanism is used to obtain the attention coefficients of different wavelet coefficients. After generating the attention weight map, the output features of the wavelet adaptive encoder are obtained using the inverse discrete wavelet transform.

[0077] like Figure 3 As shown, the wavelet adaptive encoder includes a high wavelet perception unit (High WPU), a low wavelet perception unit (LowWPU), a discrete wavelet transform (DWT) unit, a wavelet frequency domain attention mechanism, and an inverse discrete wavelet transform (IDWT) unit.

[0078] S21: Use the high-frequency perception unit to extract the high-frequency characteristic information of the input features:

[0079] ;

[0080] in, is the output feature tensor of the high-frequency perception unit, is the input feature tensor of the high-wave perception unit; It is the high-frequency perception unit processing function, used to extract high-frequency feature information; It is a depth-separable convolution operation function for high-frequency features.

[0081] S22: extracting low-frequency feature information of the input feature using the low-frequency perception unit;

[0082] ;

[0083] in, is the output feature tensor of the low-wave perception unit; is the input feature tensor of the low-wave perception unit; It is the low-frequency perception unit processing function, used to extract low-frequency feature information; It is a depth-separable convolution operation function for low-frequency features.

[0084] The adaptive encoder divides input features into high-frequency and low-frequency channels, processing them through high-frequency perception units and low-frequency perception units, respectively. This design enables the model to simultaneously capture local details in the high-frequency channel and global semantic information in the low-frequency channel. The high-frequency and low-frequency perception units use a residual structure to facilitate information transfer and gradient flow.

[0085] like Figure 4 As shown, the high-wavelength perception unit provided by the present invention includes a residual network layer, a discrete wavelet transform unit, a depthwise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer connected in series along the propagation direction; wherein the output of the residual network layer is jump-connected to the output of the layer normalization. The low-wavelength perception unit includes a discrete wavelet transform unit, a depthwise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer connected in series along the propagation direction; wherein the input of the low-wavelength perception unit is jump-connected to the output of the layer normalization.

[0086] The high-frequency perception unit processes the input features through the residual network layer (ResNet layer) and then performs a discrete wavelet transform. The output features of the ResNet layer are decomposed into low-frequency approximation coefficients (LL) and high-frequency detail coefficients in three directions (LH, HL, and HH), corresponding to detail information in the horizontal, vertical, and diagonal directions, respectively. After processing the low-frequency approximation coefficients and high-frequency detail coefficients using depthwise separable convolution, the processed wavelet coefficients are reconstructed into spatial domain features using an inverse discrete wavelet transform, resulting in the output features of the high-frequency perception unit.

[0087] The mathematical expression of the high-wave perception unit is:

[0088] ;

[0089] ;

[0090] ;

[0091] ;

[0092] ;

[0093] ;

[0094] in, is the input feature tensor of the high / low wave perception unit, It is the feature tensor after processing by the ResNet layer; It is the residual network layer processing function, used for deep feature extraction; is the discrete wavelet transform function, The feature decomposition is divided into four sub-bands: LL is the low-frequency approximation coefficient tensor, LH is the horizontal high-frequency detail coefficient tensor, HL is the vertical high-frequency detail coefficient tensor, and HH is the diagonal high-frequency detail coefficient tensor; 、 、 and They are the low-frequency approximation coefficient tensor, horizontal high-frequency detail coefficient tensor, vertical high-frequency detail coefficient tensor, and diagonal high-frequency detail coefficient tensor after processing by the depthwise separable convolution layer; It is a 3×3 depth-separable convolution operation function that enhances feature representation capabilities; is the inverse discrete wavelet transform; For 、 、 and The spatial domain features output after inverse discrete wavelet transform; It is the layer normalization function to ensure the stability of feature distribution; For The normalized feature tensor; It is the high-frequency feature tensor output by the high-frequency perception unit.

[0095] The mathematical expression of the low-wave perception unit is:

[0096] ;

[0097] ;

[0098] in, is the feature tensor after layer normalization of the low-wave perception unit; It is the low-frequency feature tensor output by the low-frequency perception unit.

[0099] High-frequency features contain rich edge, texture, and detail information. This information is easily affected by factors such as illumination variations and image quality differences in cross-modal scenarios, requiring deeper feature extraction and abstraction capabilities. Furthermore, the processing of high-frequency features involves more nonlinear transformations, which can easily lead to the vanishing gradient problem. Therefore, the high-wave perception unit combines ResNet layers with wavelet transforms. Through the residual connection mechanism of the ResNet layer, a more complex high-frequency feature representation is learned, effectively capturing subtle texture changes and edge information. The skip connection mechanism of the ResNet layer ensures effective gradient propagation, avoiding the difficulties of training deep networks. Low-frequency features primarily contain the overall outline and global structure of the image. This information is relatively stable and not easily affected by environmental factors. Effective low-frequency feature representations can be obtained directly through wavelet transforms and depthwise separable convolutions, without the need for additional deep feature extraction. Excessive feature transformations can lead to the loss of important global structural information.

[0100] In summary, the embodiments of the present invention achieve an optimal balance between feature quality and computational efficiency by using ResNet layers in high-frequency sensing units to enhance feature extraction capabilities while maintaining a lightweight design in low-frequency sensing units. This asymmetric design enables the model to control computational complexity while ensuring performance. This differentiated design strategy enables the adaptive encoder to adopt the most appropriate processing method for the characteristics of different frequency domain features, ensuring the effective extraction of high-frequency details while maintaining the integrity of the low-frequency structure.

[0101] S23: After concatenating the high-frequency feature information and the low-frequency feature information, the discrete wavelet transform is performed to decompose the information into different frequency domain sub-bands. The attention coefficients of different wavelet coefficients are obtained by using the frequency domain attention operation to generate the fused attention features and obtain the attention weight map.

[0102] ;

[0103] in, is the output feature tensor of the frequency domain attention mechanism; It is the frequency domain attention processing function, which is used to fuse different frequency domain features; is the discrete wavelet transform function, which decomposes the features into frequency subbands.

[0104] Because the frequency domain subbands after the wavelet transform have different physical meanings, such as the LL subband containing low-frequency approximate information, and the LH, HL, and HH subbands containing high-frequency details in the horizontal, vertical, and diagonal directions, respectively, traditional spatial or channel-domain attention mechanisms cannot distinguish the differences in the importance of these directional features. The complex nature of frequency domain features and the real nature of spatial domain features have representation differences, and directly applying traditional spatial domain attention mechanisms will lose the phase information of frequency domain features. Furthermore, the information distribution of different modalities in each frequency domain subband differs significantly: infrared images are primarily concentrated in low-frequency subbands, while visible light images contain rich texture information in high-frequency subbands. Using fixed weights with traditional attention mechanisms cannot effectively advance the image features of different modalities.

[0105] To address the aforementioned technical difficulties, the present invention designs a wavelet frequency-domain attention mechanism that is applied directly to the frequency subbands of the wavelet transform. By designing independent, depthwise separable convolution branches for the three key subbands (LH, LL, and HL), the importance weights of frequency-domain features in each direction can be specifically learned. Feature concatenation and a 1×1 convolution fusion strategy are employed to effectively integrate multi-directional frequency-domain information. A sigmoid activation function is used to generate a normalized attention weight map to ensure the proper distribution of weights across subbands. This design enables the model to adaptively adjust the weights of each subband, adaptively strengthening frequency-domain features that are beneficial for cross-modal recognition while suppressing interference from noise and modality-specific information.

[0106] The expression of the wave frequency domain attention mechanism is:

[0107] ;

[0108] ;

[0109] ;

[0110] in, They represent the horizontal high-frequency wavelet coefficient tensor, low-frequency wavelet coefficient tensor, and vertical high-frequency wavelet coefficient tensor of the input wave frequency domain attention mechanism respectively; is the depth-separable convolution operation function; They are the horizontal high-frequency attention coefficient tensor, the low-frequency attention coefficient tensor, and the vertical high-frequency attention coefficient tensor respectively; is the fused attention feature tensor; It is a 1×1 convolution operation function used for feature fusion and dimension adjustment; is the feature concatenation operator, connecting along the channel dimension; Sigmoid activation function normalizes the attention weight to the range of 0-1; This is the attention weight map output by the frequency domain attention mechanism.

[0111] The wave frequency domain attention mechanism provided by the present invention can accurately capture the frequency response differences between different modalities, especially through the three submodules LH, LL, and HL to process feature changes in the horizontal, vertical, and diagonal directions respectively. This directional perception ability enables the model to perform well in cross-modal re-identification tasks with complex lighting conditions and perspective changes.

[0112] S24: Based on the attention weight map, spatial features are reconstructed using inverse discrete wavelet transform;

[0113] ;

[0114] is the reconstructed feature tensor after inverse wavelet transform; As an auxiliary feature tensor involved in reconstruction; is the inverse discrete wavelet transform function, which converts the frequency domain features back to the spatial domain; It is the element-wise dot product operator.

[0115] S3: The global frequency domain features and spatial domain weight features of the input features of the frequency enhancement module are obtained respectively through fast Fourier transform and depthwise separable convolution. After fusing the global frequency domain features and spatial domain weight features using Einstein matrix multiplication, the output features of the frequency enhancement module are obtained through inverse fast Fourier transform.

[0116] Although the wavelet transform provides good time-frequency localization capabilities, its spectral representation is limited and lacks global spectrum modeling capabilities, resulting in the model being unable to effectively capture the global frequency domain characteristic differences between cross-modal images. In particular, when dealing with complex lighting conditions and different imaging qualities, the model's robustness is significantly insufficient. Therefore, in order to compensate for the limitations of the wavelet transform in global spectrum representation, this application further designs a frequency enhancement module, such as Figure 5 As shown in Figure 2, after the frequency domain attention mechanism, the frequency enhancement module further enhances the feature representation in the frequency domain through Fourier transform. Fourier transform provides complete spectral analysis, mapping spatial domain features to the frequency domain. This transforms the structural differences between infrared and visible light images into processable spectral differences, thereby enriching the feature representation.

[0117] The frequency enhancement module, through a two-layer Einstein Matrix Multiplication (EMM) structure, achieves a deep fusion of frequency and spatial domain information. This ensures that the output of the frequency enhancement module incorporates both the spatial texture of the visible light image extracted by DWConv and the frequency domain profile of the infrared image obtained through Fourier transform, improving the accuracy of cross-modal matching. The first-layer EMM combines frequency and spatial features, while the second-layer EMM further enhances feature representation after sigmoid activation. By adding nonlinear transformations, the two-layer EMM can better adapt to the differences between modal data and extract more discriminative features. This design enables the model to highlight key frequency components while preserving complementary information between modalities. The two-layer EMM uses different weight matrices to more flexibly adjust feature dimensions. The first-layer EMM performs preliminary screening and weighting of input features to highlight key feature dimensions. The second-layer EMM further refines these key feature dimensions to enhance feature representation. For example, the first-layer EMM can preliminarily weight the channels of frequency domain features based on the information of spatial domain features, retaining the frequency components related to the pedestrian's identity; the second-layer EMM further adjusts the weights of these frequency components based on the results of the first layer, so that the final feature representation more accurately reflects the pedestrian's identity information.

[0118] Perform fast Fourier transform on the input features of the frequency enhancement module to obtain global spectrum features; obtain spatial domain weight features of the input features of the frequency enhancement module through deep separable convolution. Input the imaginary features and real features of the global frequency domain features, and the imaginary weights and real weights of the spatial domain weight features into the first Einstein matrix multiplication layer, and output the first imaginary feature, the first real feature, the first imaginary weight, and the first real weight; use the nonlinear activation function to process the first imaginary feature and the first real feature to obtain enhanced imaginary feature and enhanced real feature. Use deep separable convolution to process the first imaginary weight and the first real weight to obtain enhanced imaginary weight and enhanced real weight; input the enhanced imaginary feature, enhanced real feature, enhanced imaginary weight, and enhanced real weight into the second Einstein matrix multiplication layer, and output the target imaginary feature, target real feature, target imaginary weight, and target real weight; perform inverse Fourier transform on the target imaginary feature and target real feature to obtain the output features of the frequency enhancement module. The mathematical expression of the frequency enhancement module is:

[0119] ;

[0120] ;

[0121] ;

[0122] ;

[0123] ;

[0124] in, Represents the input feature tensor of the frequency enhancement module; They are respectively the imaginary feature tensor and real feature tensor after fast Fourier transform of the input features of the frequency enhancement module; is the fast Fourier transform function, which converts the spatial domain features into the frequency domain; The imaginary component tensor and real component tensor of the spatial domain weights generated by depth-wise separable convolution respectively; Generate spatial domain weights for depth-wise separable convolution operation function; They are The first imaginary feature tensor, the first real feature tensor, the first imaginary weight tensor, and the first real weight tensor after the first layer of EMM processing; It is the Einstein matrix multiplication function to achieve feature fusion; They are the target imaginary feature tensor, target real feature tensor, target imaginary weight tensor and final real weight tensor after the second layer of EMM processing; Sigmoid activation function provides nonlinear transformation; is the output feature tensor of the frequency enhancement block; is the inverse fast Fourier transform function, which converts the frequency domain features back to the spatial domain.

[0125] The wavelet-based infrared-visible light person re-identification method provided by the present invention uses wavelet decomposition and a wavelet-adaptive encoder to effectively address the homogeneity of infrared images and extract more discriminative frequency domain features. Furthermore, the high- and low-wave channel separation processing strategy enables the model to adopt different processing strategies for different frequency band features, better adapting to the different frequency characteristics of the two modalities. Through the wavelet-domain attention mechanism, the model can also adaptively adjust the importance of different frequency band features to enhance cross-modal feature consistency. Compared with traditional convolutional networks, the wavelet-adaptive encoder demonstrates significant advantages in processing multi-scale features and reducing modal differences. Considering the limitations of the wavelet transform in global spectral representation, the present invention further designs a frequency enhancement module that uses the Fourier transform to model global spectral characteristics, thereby compensating for these limitations of the wavelet transform in spectral representation. The Fourier transform (FEB) achieves a deep fusion of frequency and spatial domain information through Einstein matrix multiplication, enhancing the model's robustness to varying lighting conditions and imaging quality, and providing more comprehensive spectral support for cross-modal feature learning.

[0126] Existing model optimization strategies primarily focus on aligning cross-modal features of the same identity, while paying insufficient attention to distinguishing between different identities within the same modality. This one-sided optimization strategy is particularly inadequate when processing large datasets, easily leading to samples with visually similar but different identities being incorrectly mapped to similar feature space locations. This problem is particularly exacerbated by the homogeneity of infrared images, as exemplified by the typical Similarity Distribution Clustering (SDC) strategy.

[0127] To solve the above problems, the embodiment of the present invention proposes an enhanced similarity distribution clustering loss (EnhancedSDC, ESDC), such as Figure 6 As shown in the figure, by introducing the intra-modal repulsion mechanism, a more comprehensive feature learning framework is constructed. Figure 6 The figure shows the distribution of two pedestrians with different IDs in the feature space: the blue area on the left is the male ID feature, and the red area on the right is the female ID feature. ) only focuses on clustering cross-modal features of the same ID, ignoring the distinction between features of different IDs within the same modality. ESDC loss not only considers bringing together cross-modal features of the same identity, but more importantly, explicitly pushes away features of different identities within the same modality. This bidirectional constraint mechanism, implemented through the KL divergence framework, effectively solves the homogeneity problem of infrared images, enhances the discriminability of the feature space, constructs a more robust feature representation, and improves the recognition accuracy of the model in complex scenarios, especially for samples with visual similarity but different identities.

[0128] After the visible light image and infrared image pass through the feature extractor, they are processed by multiple sets of alternately set wave adaptive encoders and frequency enhancement modules to obtain the visible light feature vector matrix and infrared feature vector matrix as the final output features of the WaveRID model backbone network. For the input visible light image, after the complete WaveRID model forward propagation, global average pooling and feature normalization are performed to obtain a dimension of [N, D] Matrix, where N is the number of visible light samples in the batch and D is the feature dimension. Similarly, the infrared image undergoes the same network processing flow to obtain the corresponding These normalized feature vectors retain the cross-modal discriminative information after the frequency domain decomposition of the wave adaptive encoder and the global spectrum modeling of the frequency enhancement module, providing a high-quality feature representation basis for subsequent similarity calculation and loss optimization.

[0129] Based on the features output by the backbone network, the final feature vector is obtained through global average pooling and L2 normalization:

[0130] ;

[0131] ;

[0132] in, is the normalized visible light eigenvector matrix, is the normalized infrared eigenvector matrix, and They are the output feature maps of visible light and infrared images after passing through the WaveRID model backbone network; It is a global average pooling operation that compresses the feature map of the spatial dimension into a feature vector; This is an L2 normalization operation to ensure that the modulus of the feature vector is 1, which facilitates subsequent similarity calculations.

[0133] The construction process of the enhanced similarity distribution clustering loss includes:

[0134] Calculate the feature similarity matrix within the visible light modality:

[0135] ;

[0136] Compute the feature similarity matrix within the infrared modality:

[0137] ;

[0138] in, is the similarity matrix within the visible light modality; is the similarity matrix within the infrared modality; is the normalized visible light eigenvector matrix; is the normalized infrared eigenvector matrix; is the matrix transpose operator; is the matrix multiplication operator.

[0139] Calculate the visible light mode The sample and The similarity probability of samples , value range [0,1]:

[0140] ;

[0141] Calculate the infrared modality The sample and The similarity probability of samples , value range [0,1]:

[0142] ;

[0143] in, The first The sample and Similarity matrix of samples; The first The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; is the temperature coefficient, which controls the smoothness of the probability distribution; is the total number of samples.

[0144] Ideally, features of the same identity should be highly similar, and features of different identities should be separated. Based on this, we define the ideal label distribution:

[0145] Calculate the ideal similarity distribution within the visible light modality , value range [0,1]:

[0146] ;

[0147] Calculating ideal similarity distribution within infrared modalities , value range [0,1]:

[0148] ;

[0149] in, The first The sample and The identity label indicator (0 or 1) of each sample, The first The sample and Identity tag indicator for each sample; Infrared mode The sample and The identity label indicator of each sample, Infrared mode The sample and The identity label indicator of the two samples; when the identity label indicator is 0, it indicates that the two samples do not belong to the same identity; when the identity label indicator is 1, it indicates that the two samples belong to the same identity.

[0150] The KL divergence is used to calculate the difference between the actual similarity distribution and the ideal similarity distribution in the visible light modality, that is, the KL divergence loss of the visible light modality:

[0151] ;

[0152] The KL divergence is used to calculate the difference between the actual distribution and the ideal distribution in the infrared modality, that is, the infrared modality KL divergence loss:

[0153] ;

[0154] in, It is a numerical stability constant to prevent division by zero errors and is usually set to 1e-8.

[0155] The expression of the intra-modal repulsion loss is: ; The expression of enhanced similarity distribution clustering loss is: .

[0156] In other embodiments of the present invention, in addition to using ESDC loss to optimize the WaveRID model, identity classification loss can also be used. , triplet loss and CPM loss The model is optimized with multiple losses, and the network is jointly optimized in an end-to-end manner by minimizing the sum of these four losses. Multiplied by a constant Adding it to the total loss, through the following ablation experiment, it can be seen that the optimal setting of λ is 0.1.

[0157] The expression of the total loss function is: .

[0158] To demonstrate the performance of the WaveRID model provided by the present invention, we comprehensively evaluate the performance of the proposed method on two challenging infrared-visible cross-modal person retrieval datasets:

[0159] The RegDB dataset consists of 412 characters, each captured simultaneously with 10 visible light images and 10 infrared images using a pair of overlapping cameras. The dataset includes 254 women and 158 men, 156 of whom were photographed from the front and 256 from the back, providing multi-angle pose data.

[0160] The SYSU-MM01 dataset contains 491 identities, captured by four visible-light cameras and two infrared cameras. It supports both full search and indoor search modes. In full search mode, all visible-light camera images are used as the search library; in indoor search mode, only images captured by two indoor visible-light cameras are used to construct the search library, providing a more focused evaluation scenario.

[0161] This paper uses the widely recognized Rank-k metric (k = 1, 10, 20) as the primary evaluation criterion. Rank-k represents the probability of finding at least one corresponding infrared person image in the top k candidate lists when given a visible light image as a query. For the RegDB dataset, this paper evaluates both visible to infrared (VIS to IR) and infrared to visible (IR to VIS) query directions. For a comprehensive evaluation, this paper also introduces mean average precision (mAP) as a supplementary retrieval metric. For all metrics, higher values ​​indicate better performance.

[0162] During model training, we use random horizontal flipping, random cropping with padding, and random erasing techniques to enhance image data and improve model robustness. All input images are first resized to a uniform size of 3×384×144. The experiment trains for a total of 150 epochs with a batch size of 8. The training process uses the AGW optimizer with an initial learning rate of 1×10 -2 , and increased to 1×10 after the first 10 epochs under the warm-up strategy -1 Subsequently, the learning rate was reduced to 1×10 at the 20th, 60th, and 120th epochs, respectively. -2 , 1×10 -3 and 1×10 -4 , achieving effective gradient descent. In the ESDC loss function, the clustering factor M is set to 2, and the temperature coefficient The loss weight λ is set to 0.05 to balance the model's discriminative ability and generalization performance. The loss weight λ is set to 0.1. Ablation experiments show the performance of the model with different loss weights. All experiments are trained and evaluated on a single NVIDIA RTX 3090 GPU.

[0163] Table 1 shows a comparison of the WaveRID model proposed in this paper with state-of-the-art methods. The proposed WaveRID model achieves state-of-the-art performance on the RegDB dataset. In the visible to infrared (VIS to IR) retrieval mode, WaveRID achieves 94.8% Rank-1 accuracy, 99.6% Rank-20 accuracy, and 89.2% mAP. In the infrared to visible (IR to VIS) retrieval mode, WaveRID also performs exceptionally well, achieving 94.8% Rank-1 accuracy, 99.6% Rank-20 accuracy, and 89.1% mAP, demonstrating the method's bidirectional retrieval capabilities. The best performance is highlighted in bold, and the second-best performance is underlined.

[0164] Table 1 Comparison results of WaveRID model and recent state-of-the-art methods on RegDB dataset

[0165]

[0166] As shown in Table 2, compared with the baseline model, the method provided by the present invention also shows excellent performance advantages on the more challenging SYSU dataset. It has reached the current state-of-the-art level in all evaluation indicators. In the All Search mode, it achieved a Rank-1 accuracy of 75.6%, a Rank-20 accuracy of 99.3%, and a mAP of 74.7%; in the IndoorSearch mode, it achieved a Rank-1 accuracy of 81.7%, a Rank-20 accuracy of 99.9%, and a mAP of 87.9%. This further verifies the robustness and generalization ability of the proposed model architecture in different scenarios. It also proves the effectiveness and advantages of the ESDC optimization strategy proposed in this invention on complex datasets.

[0167] Table 2 Comparison results of the WaveRID model and recent state-of-the-art methods on the SYSU-MM01 dataset

[0168]

[0169] To further visualize the effectiveness of WaveRID, Figure 7 A set of retrieval results in complex lighting scenes and scenes with changing perspectives are shown. For each retrieval case, the retrieved image with a green frame represents the correct match corresponding to the given query, while the one with a red frame represents an incorrect match. Figure 7 As can be seen, WaveRID is able to rank more correctly matched images at the top, particularly when dealing with challenging scenarios such as illumination changes, pose differences, and complex backgrounds. WaveRID's ability to effectively improve ranking results validates the effectiveness of this architecture and further demonstrates the effectiveness of multi-scale wavelet representation and enhanced similarity learning in cross-modal recognition tasks.

[0170] Ablation Analysis of Each Component of the WaveRID Model: To fully demonstrate the impact of the various components of the proposed method, detailed ablation experiments on each component of the WaveRID model are performed in All Search mode on the SYSU-MM01 dataset. As shown in Table 3, compared to using only WAE, adding the FEB module improves Rank-1% by 1.0%. Adding the ESDC loss further improves Rank-1% by 1.3%. The complete WaveRID model achieves state-of-the-art performance, demonstrating the superiority of the proposed method.

[0171] Table 3 Ablation test results of each component of WaveRID

[0172]

[0173] Comparative analysis of the effects of ESDC and SDC losses: SDC loss may cause samples with visual similarity but different identities to be incorrectly mapped to similar positions, and performs poorly on large-scale datasets. Compared with the original SDC loss, the ESDC loss proposed in the present invention introduces a new intra-modal repulsion mechanism, which effectively enhances the discriminability of the feature space, thereby enhancing the performance of the model on large-scale datasets. As shown in Table 4, the model using ESDC loss has better performance than the model using SDC loss. On the SYSU-MM01 dataset, the Rank-1 accuracy is improved by 2.8% in the All-Search mode and by 1.4% in the Indoor-Search mode. This illustrates the limitations of the original SDC loss on large-scale datasets, and verifies the effectiveness of the improvement of the present invention.

[0174] Table 4 Comparison of the performance of the WaveRID model on the SYSU-MM01 dataset using ESDC and SDC losses

[0175]

[0176] ESDC loss weighting experiments: The ESDC loss is multiplied by a constant λ and added to the total loss. Table 5 shows the impact of different values ​​of this constant on the model's performance on the All-Search mode of the SYSU-MM01 dataset when all other loss hyperparameters are set to 1.0. When λ is equal to 0.1, the model achieves the highest Rank-1 and mAP accuracy.

[0177] Table 5 Performance of WaveRID model on SYSU-MM01 dataset under different ESDC loss weights

[0178]

[0179] Loss function weight analysis: In addition to ESDC loss, this paper systematically analyzes the weight coefficients of each item in the total loss function. We set α, β, and γ to be Table 6 shows the impact of different loss function weights on model performance when the ESDC loss hyperparameter λ is set to 0.1. Experimental results show that when α=1.0, β=1.0, and γ=1.0, the model achieves optimal performance on all evaluation metrics. Setting the identity loss weight too small (α=0.5) weakens identity discrimination, while setting it too large (α=2.0) may lead to overfitting. The performance is better in the range of 1.0-1.5, and there is a slight but not significant improvement when the weight is 1.5. Weights between 1.0 and 1.5 have a positive impact on performance, but the effect begins to decline after exceeding 1.5. These results verify that the weight ratio (1.0, 1.0, 1.0) selected by this invention can achieve a good balance between the various loss functions and provide the best training effect for the model.

[0180] Table 6 Performance analysis of different loss function weight ratios on the SYSU-MM01 dataset

[0181]

[0182] In summary, this application provides a novel cross-modal model WaveRID based on wavelet representation, which effectively solves the modal difference problem in infrared-visible light person re-identification. By designing a wave adaptive encoder (WAE) to achieve multi-scale frequency domain feature decomposition, using a frequency enhancement module (FEB) for global spectrum modeling, and introducing an enhanced similarity distribution clustering mechanism (ESDC) to simultaneously optimize cross-modal feature alignment and same-modal feature differentiation, WaveRID constructs a more robust feature space. Extensive experiments on the RegDB and SYSU-MM01 datasets show that the proposed method achieves state-of-the-art performance in all evaluation indicators. The present invention provides a new research perspective for cross-modal feature extraction and infrared-visible light person re-identification. Future work will explore more lightweight network design and its application potential in other cross-modal tasks.

[0183] The present invention also provides an infrared-visible light person re-identification system based on wavelet representation, comprising: a memory for storing a computer program; and a processor for implementing the steps of the infrared-visible light person re-identification method based on wavelet representation as described above when executing the computer program.

[0184] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0185] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0186] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0188] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A wavelet-based infrared-visible light person re-identification method, characterized in that: include: Inputting the visible light image and the infrared image into a cross-modal person re-identification model based on wavelet representation, wherein the cross-modal person re-identification model based on wavelet representation includes a feature extractor and a backbone network, wherein the backbone network includes multiple groups of wavelet adaptive encoders and frequency enhancement modules arranged alternately; The feature extractor is used to extract the initial features of the visible light image and the infrared image, and then the initial features are stitched together and input into the backbone network; For each wavelet adaptive encoder, the high-frequency and low-frequency features of the input features of the wavelet adaptive encoder are extracted and spliced, the spliced ​​features are discrete wavelet transformed, and the wavelet frequency domain attention mechanism is used to obtain the attention coefficients of different wavelet coefficients. After generating the attention weight map, the output features of the wavelet adaptive encoder are obtained by inverse discrete wavelet transform. For each frequency enhancement module, the global frequency domain features and spatial domain weight features of the input features of the frequency enhancement module are obtained respectively through fast Fourier transform and depthwise separable convolution. After fusing the global frequency domain features and spatial domain weight features, the output features of the frequency enhancement module are obtained through inverse fast Fourier transform. Generate person re-identification results based on the output features of the backbone network.

2. The infrared-visible light person re-identification method based on wavelet representation according to claim 1 is characterized in that: For each wave adaptive encoder: The high-frequency features of the input features of the wave adaptive encoder are extracted using the high-frequency perception unit; The low-frequency features of the input features of the adaptive encoder are extracted using the low-frequency sensing unit; Concatenate high-frequency features with low-frequency features, and use discrete wavelet transform to decompose the concatenated features into different frequency sub-bands to obtain multiple wavelet coefficients; Using the wavelet frequency domain attention mechanism, we obtain the attention coefficients of different wavelet coefficients, generate the fused attention features, and obtain the attention weight map; Based on the attention weight map, the spatial domain features are reconstructed using inverse discrete wavelet transform to obtain the output features of the wavelet adaptive encoder.

3. The infrared-visible light person re-identification method based on wavelet representation according to claim 2 is characterized in that: The high-wave perception unit includes a residual network layer, a discrete wavelet transform unit, a depthwise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer, which are connected in series along the propagation direction. The output of the residual network layer is jump-connected to the output of the layer normalization layer. The low-wave perception unit includes a discrete wavelet transform unit, a depth-wise separable convolution layer, an inverse discrete wavelet transform unit, and a layer normalization layer, which are sequentially connected in series along the propagation direction; wherein, the input of the low-wave perception unit is jump-connected with the output of the layer normalization.

4. The infrared-visible light person re-identification method based on wavelet representation according to claim 2 is characterized in that: The expression of the high-frequency sensing unit is: ; ; ; ; ; ; in, is the input feature of the high-wave perception unit; is the feature after processing by the residual network layer, is the residual network layer processing function; 、 、 and They are Low-frequency approximation coefficients, horizontal high-frequency detail coefficients, vertical high-frequency detail coefficients, and diagonal high-frequency detail coefficients obtained by discrete wavelet transform decomposition; is the discrete wavelet transform function; 、 、 and They are 、 、 、 The low-frequency approximation coefficients, horizontal high-frequency detail coefficients, vertical high-frequency detail coefficients, and diagonal high-frequency detail coefficients after processing by the depthwise separable convolution layer; It is a 3×3 depth-separable convolution operation function; is the inverse discrete wavelet transform; For 、 、 and The spatial domain features output after inverse discrete wavelet transform; is the layer normalization function; For Features after normalization; It is the high-frequency feature output by the high-frequency perception unit.

5. The infrared-visible light person re-identification method based on wavelet representation according to claim 1 is characterized in that: The expression of the wave frequency domain attention mechanism is: ; ; ; in, Represent the horizontal high-frequency wavelet coefficients, low-frequency wavelet coefficients, and vertical high-frequency wavelet coefficients of the input wave frequency domain attention mechanism respectively; is the depth-separable convolution operation function; They are the horizontal high-frequency attention coefficient, low-frequency attention coefficient, and vertical high-frequency attention coefficient; is the fused attention feature; is a 1×1 convolution operation function; is the feature concatenation operator, connecting along the channel dimension; is the Sigmoid activation function; This is the attention weight map output by the frequency domain attention mechanism.

6. The infrared-visible light person re-identification method based on wavelet representation according to claim 1 is characterized in that: For each frequency boost module: Perform fast Fourier transform on the input features of the frequency enhancement module to obtain global spectrum features; The spatial domain weight features of the input features of the frequency enhancement module are obtained through depth-wise separable convolution; After using Einstein matrix multiplication to fuse the global spectrum features and spatial domain weight features, the features are reconstructed through inverse Fourier transform to obtain the output features of the frequency enhancement module.

7. The infrared-visible light person re-identification method based on wavelet representation according to claim 6 is characterized in that: After using Einstein matrix multiplication to fuse global spectrum features and spatial domain weight features, the features are reconstructed through inverse Fourier transform to obtain the output features of the frequency enhancement module, including: Input the imaginary and real features of the global frequency domain features, and the imaginary weights and real weights of the spatial domain weight features into the first Einstein matrix multiplication layer, and output the first imaginary feature, the first real feature, the first imaginary weight, and the first real weight; The first imaginary feature and the first real feature are processed by using a nonlinear activation function to obtain an enhanced imaginary feature and an enhanced real feature; Processing the first imaginary weight and the first real weight using a depthwise separable convolution to obtain an enhanced imaginary weight and an enhanced real weight; Input the enhanced imaginary feature, enhanced real feature, enhanced imaginary weight and enhanced real weight into the second Einstein matrix multiplication layer, and output the target imaginary feature, target real feature, target imaginary weight and target real weight; The target imaginary part features and the target real part features are inverse Fourier transformed to obtain the output features of the frequency enhancement module.

8. The infrared-visible light person re-identification method based on wavelet representation according to claim 1 is characterized in that: The loss function of the cross-modal person re-identification model based on wavelet representation is an enhanced similarity distribution clustering loss, including: similarity distribution clustering loss, visible light modality KL divergence loss and infrared modality KL divergence loss; The construction process of the visible light modal KL divergence loss and the infrared modal KL divergence loss includes: Calculate the visible light mode The sample and The similarity probability of samples : ; Calculate the infrared modality The sample and The similarity probability of samples : ; in, The first The sample and Similarity matrix of samples; The first The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; Infrared mode The sample and Similarity matrix of samples; is the temperature coefficient; is the total number of samples; Calculate the visible light mode The sample and The ideal similarity distribution of samples : ; Calculate the infrared modality The sample and The ideal similarity distribution of samples : ; in, The first The sample and The identity label indicator of each sample, The first The sample and Identity tag indicator for each sample; Infrared mode The sample and The identity label indicator of each sample, Infrared mode The sample and The identity label indicator of the samples; when the identity label indicator is 0, it indicates that the two samples do not belong to the same identity; when the identity label indicator is 1, it indicates that the two samples belong to the same identity; Constructing visible light modality KL divergence loss : ; Constructing infrared modality KL divergence loss : ; in, is a constant.

9. The infrared-visible light person re-identification method based on wavelet representation according to claim 1 is characterized in that: The loss function of the cross-modal person re-identification model based on wavelet representation is: ; ; in, is the total loss function; To enhance the similarity distribution clustering loss; Classify loss for identity; is the triplet loss; is the cross-modal pairing loss; is the similarity clustering distribution loss; is the KL divergence loss of the visible light mode; Infrared modality KL divergence loss, is a constant.

10. An infrared-visible light person re-identification system based on wavelet representation, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the infrared-visible light person re-identification method based on wavelet representation as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Infrared and visible light image perception enhancement fusion method and system based on deep learning

    CN119722492A

  • Wood water paint surface defect detection method based on computer vision

    CN119831958A