Unsupervised cross-modal video person re-identification method based on reinforcement semantic consistency

By employing a trajectory diversity modeling and Fourier transform-enhanced asymmetric cross-contrast learning strategy, the problems of feature distribution discreteness and noise influence in unsupervised cross-modal video pedestrian re-identification are solved, achieving more efficient cross-modal feature learning and recognition.

CN121545187BActive Publication Date: 2026-03-27ROCKET FORCE UNIV OF ENG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-21
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing unsupervised cross-modal video pedestrian re-identification methods face challenges in constructing reliable track representations and cross-modal correspondences, especially under the influence of feature distribution discreteness and noise between visible and infrared modes, making it difficult to effectively learn modality-invariant features.

Method used

We adopt a track diversity modeling structure, filter and extract simple and difficult frames, and model reliable and information-rich track features respectively. We use Fourier transform to enhance the high-level semantic consistency between modes, and establish a reliable cross-modal correspondence between visible light and infrared modes through an asymmetric cross-contrast learning strategy.

Benefits of technology

It improves the model's cross-modal person re-identification performance under unsupervised conditions, significantly enhances the accuracy and robustness of feature learning without relying on manual annotation, and strengthens the reliability of cross-modal correspondence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121545187B_ABST
    Figure CN121545187B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on reinforcement semantic consistency Unsupervised Cross-Modal Video Pedestrian Re-Identification Method, it is related to image recognition technical field, including the image of multiple re-identification pedestrians is input into trained pedestrian re-identification model, obtains re-identification result;Pedestrian re-identification model training: distinguish each frame image in each track of each mode as difficult frame or simple frame, obtain difficult track feature and simple track feature;Reliable memory dictionary is generated;Simple frame is Fourier enhanced and cross-modal matching relationship is suggested;Utilize asymmetric cross comparison learning strategy to promote the training of pedestrian re-identification model.The application utilizes Fourier transform to strengthen the advanced semantic consistency between visible light and infrared mode, promote the construction reliable cross-modal corresponding relationship, utilize difficult memory dictionary to learn the discriminative feature of reliable track, while utilizing reliable memory dictionary to learn the discriminative feature of difficult track, without manual annotation can realize performance improvement.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image recognition, and particularly relates to an unsupervised cross-modal video person re-identification method based on reinforced semantic consistency. BACKGROUND

[0002] Person re-identification (ReID) aims to find a person with a specific identity from a large-scale image library. Research related to ReID has the potential to promote the progress of intelligent monitoring and effectively reduce the labor cost related to person identification and retrieval. With the introduction of deep learning technology, early research has achieved outstanding results in ReID in the visible light modality. As the demand for intelligent security systems continues to increase, 24-hour intelligent monitoring sensor devices have been widely used. The imaging principles of sensor devices in the daytime and at night differ significantly, resulting in a large cross-modal gap between visible light data and infrared data. Visible-infrared person re-identification (VI-ReID) aims to match a specific person with the same identity between visible light and infrared light modalities, and therefore has received increasing attention. This challenge of cross-modal gap has prompted scholars to pay more attention to research on visible-infrared person re-identification.

[0003] In recent years, multiple image-based visible-infrared person re-identification methods have been proposed, which establish cross-modal correspondence and promote the model to learn modality-invariant features, but they are not suitable for more common video data scenarios. Unlike static images, video data provides more rich spatiotemporal information, enabling more accurate and robust person retrieval. Recently, some methods have explored video-based visible-infrared person re-identification (VVI-ReID), aiming to learn consistent semantic representations from information-rich but complex cross-modal video data through supervised learning methods. However, the huge cost involved in cross-modal labeling hinders the practical deployment of these methods. This limitation highlights the need to develop an unsupervised video-based visible-infrared person re-identification technology (USL-VVI-ReID). USL-VVI-ReID is a highly challenging research direction, mainly due to the characteristic of discrete distribution of frame features within each track segment. This discreteness hinders the model from learning modality-invariant representations without human labeling.

[0004] Currently, unsupervised pedestrian re-identification algorithms are mainly divided into three categories: the first category is unsupervised single-modal image pedestrian re-identification algorithm; the second category is unsupervised cross-modal image pedestrian re-identification algorithm; the third category is unsupervised single-modal video pedestrian re-identification algorithm.

[0005] Unsupervised single-modal image pedestrian re-identification algorithm does not consider the feature distribution difference between modalities, and only aims at the blur, low contrast, camera angle difference and occlusion of pedestrian images captured by multiple cameras, extracts discriminative features of images through network structure design, and predicts more accurate pseudo-labels to guide model iterative training. The existing unsupervised single-modal image pedestrian re-identification technology method integrates clustering algorithm into the contrast learning framework based on memory dictionary. They usually first predict pseudo-labels through clustering algorithm (such as DBSCAN, K-means), and then optimize the model using contrast loss based on pseudo-labels. However, the noise instance features in outliers will hinder the learning of feature representation.

[0006] Unsupervised cross-modal image pedestrian re-identification algorithm mainly adopts a dual-branch contrast learning framework, which uses clustering algorithm to predict pseudo-labels for pedestrian images in visible light modality and infrared modality respectively, and establishes cross-modal correspondence based on clustering center prototype based on these pseudo-labels to promote the model to learn modality-independent features. However, in the USL-VVI-ReID task, the random sampling strategy for track segments will cause the deviation of clustering centers and establish unreliable cross-modal correspondence.

[0007] Unsupervised single-modal video pedestrian re-identification algorithm usually adopts random sampling strategy, randomly samples frames in each track, and uses average pooling to sample frame features as track representation. This random sampling strategy simplifies the track feature modeling process, but may reduce the quality of track representation. In order to reduce the influence of occluded frames or frames with significant appearance changes, some studies introduce noise pruning module to reduce the influence of noise frames, so as to realize reliable track segment representation modeling. However, this noise pruning strategy may lose important but information valuable frame features, and the model's ability to extract discriminative features of frames is limited, which limits the reliability of constructing cross-modal correspondence in the USL-VVI-ReID task.

[0008] Unsupervised single-modal video person re-identification algorithms focus on constructing reliable track representations from discrete frames, which usually contain noise or hard samples caused by view changes and illumination condition differences. Unsupervised image cross-modal person re-identification algorithms take into account the inherent differences in color, texture and sensor-specific properties, aiming to facilitate the model to learn modal-invariant features between visible light and infrared modalities. In contrast, the core challenge of unsupervised cross-modal video person re-identification algorithms lies in two interrelated tasks: constructing reliable track representations to establish reliable cross-modal correspondence, and extracting difficult but informative frames from discrete distribution to enhance the model's ability to learn discriminative features. SUMMARY

[0009] The technical problem to be solved by the present application is to provide an unsupervised cross-modal video person re-identification method based on reinforced semantic consistency to solve the problems in the prior art.

[0010] To solve the above technical problems, the technical solution adopted by the present application is: an unsupervised cross-modal video person re-identification method based on reinforced semantic consistency, comprising:

[0011] Obtaining images of the person to be re-identified, the images of the person to be re-identified including visible light images and infrared images, multiple visible light images forming a visible light modality, and multiple infrared images forming an infrared modality;

[0012] Inputting the multiple images of the person to be re-identified into a trained person re-identification model to obtain a re-identification result, the training process of the person re-identification model comprising:

[0013] Step 1, calculating the feature mean value of each frame image in each track of each modality, and the cosine similarity between the feature value of each frame image and the feature mean value in the corresponding track, setting a cosine similarity threshold parameter in each track, and setting a cosine similarity threshold parameter in each track, the frame with a cosine similarity greater than the cosine similarity threshold parameter in the corresponding track is a simple frame, and the frame with a cosine similarity not greater than the cosine similarity threshold parameter in the corresponding track is a difficult frame;

[0014] Constructing difficult frame features and easy frame features for each track;

[0015] Performing average pooling processing on the difficult frame features and the easy frame features respectively to obtain difficult track features and simple track features;

[0016] Step 2, clustering the simple track features of the visible light modality and the infrared modality respectively to predict pseudo-labels in the visible light modality and pseudo-labels in the infrared modality; generating a reliable memory dictionary of the visible light modality and a reliable memory dictionary of the infrared modality from the clustering centers of the simple track features according to the predicted pseudo-labels in the visible light modality and the predicted pseudo-labels in the infrared modality;

[0017] the difficult track features of the visible light modal and the infrared modal are respectively clustered, and a difficult memory dictionary of the visible light modal and a difficult memory dictionary of the infrared modal are respectively generated based on the cluster centers of the difficult track features;

[0018] Step 3, Fourier enhancement is performed on the simple frames in the simple track features;

[0019] An optimal transport strategy is applied to establish a reliable cross-modal matching relationship between the Fourier-enhanced visible light modal and infrared modal, and to minimize the transport plan cost, so as to obtain an optimal transport scheme, and based on the optimal transport scheme, two groups of pseudo-labels matched with each other are obtained;

[0020] Step 4, an asymmetric cross-contrast learning strategy is used to perform interactive contrast learning between the simple track features of the visible light modal and the difficult memory dictionary of the visible light modal, and to perform interactive contrast learning between the difficult track features of the visible light modal and the reliable memory dictionary of the visible light modal, and at the same time, interactive contrast learning is performed between the simple track features of the infrared modal and the difficult memory dictionary of the infrared modal, and interactive contrast learning is performed between the difficult track features of the infrared modal and the reliable memory dictionary of the infrared modal, so as to obtain a first loss function;

[0021] An asymmetric cross-contrast learning strategy is used to perform interactive contrast learning between the Fourier-enhanced cross-modal matching simple track features and the reliable memory dictionary of the visible light modal and the reliable memory dictionary of the infrared modal, respectively, so as to obtain a second loss function;

[0022] The first loss function and the second loss function are used to promote the learning of the modal-invariant features of the human re-identification model and the discriminability.

[0023] In step 2 of the above unsupervised cross-modal video pedestrian re-identification method based on reinforced semantic consistency, the DBSCAN algorithm is used to cluster the simple track features of the visible light modal and the infrared modal, respectively.

[0024] In step 3 of the above unsupervised cross-modal video pedestrian re-identification method based on reinforced semantic consistency, the process of Fourier enhancement on the simple frames in the simple track features is as follows:

[0025] The simple frames in the simple track features are subjected to inverse Fourier transform, the amplitude components after the inverse Fourier transform retain the low-level modal perception information of the simple frames, and the phase components after the inverse Fourier transform retain the high-level semantic information of the simple frames, and the phase component information is integrated back into the original simple frames to enhance the high-level semantic information in the same modal;

[0026] that is, the simple frames after Fourier enhancement , wherein, is the simple frames in the simple track features, is a phase component after inverse Fourier transform, is a proportional coefficient.

[0027] The unsupervised cross-modal video pedestrian re-identification method based on reinforced semantic consistency has the following advantages: ; wherein, is an asymmetric cross-contrast learning strategy, is a simple track feature of the visible light mode, is a difficult memory dictionary of the visible light mode, is a difficult track feature of the visible light mode, is a reliable memory dictionary of the visible light mode; is a simple track feature of the infrared mode, is a difficult memory dictionary of the infrared mode, is a difficult track feature of the infrared mode, is a reliable memory dictionary of the infrared mode;

[0028] The asymmetric cross-contrast learning strategy is used to perform interactive contrast learning on the simple track features matched by the Fourier enhancement cross-modal by respectively interacting with the reliable memory dictionary of the visible light mode and the reliable memory dictionary of the infrared mode, to obtain a second loss function ; wherein, is a simple track feature matched by the Fourier enhancement cross-modal;

[0029] According to the formula , a comprehensive loss function of the pedestrian re-identification model is obtained , wherein, is a first weight coefficient, is a second weight coefficient and .

[0030] Compared with the prior art, the present application has the following advantages:

[0031] 1. The present application designs a track diversity modeling structure aiming at the discrete characteristics of track features in the frame, detects and extracts simple frames and difficult frames, and respectively models reliable and information-rich track features. Among them, the former is used to predict accurate pseudo labels, and the latter is used to learn global discriminative features of discrete frames.

[0032] 2, The present application uses the Fourier transform to enhance the pedestrian images in the visible light and infrared modalities, which are mapped in a similar feature space, thereby enhancing the high-level semantic consistency between the two modalities and establishing a more reliable cross-modal correspondence.

[0033] 3, The method steps of the present application are simple, first, the reliable memory dictionary is initialized by the clustering prototype of simple track features, and the difficult memory dictionary is initialized by the clustering prototype of difficult track features, then in an asymmetric collaborative learning manner, the cross comparison learning of simple track features and difficult memory dictionary, and difficult track features and reliable memory dictionary is designed, which can realize interactive extraction of global feature embedding from reliable and information-rich track features.

[0034] In summary, the present application aims to learn robust features from frame feature distribution discrete tracks and visible-infrared modal features in an unsupervised manner. First, in order to improve the feature quality of the track, the method screens and extracts simple frames and difficult frames in each track, and models simple and difficult track features respectively. Specifically, the simple track features are used to predict more accurate pseudo labels, and the difficult track features are used to initialize a diverse memory dictionary to promote the model to learn discriminative features. Secondly, in order to establish a reliable cross-modal correspondence, the method uses the Fourier transform to enhance the high-level semantic consistency between the visible light and infrared modalities, and promotes the construction of a reliable cross-modal correspondence. In addition, the method uses the difficult memory dictionary to learn the discriminative features of reliable tracks, and uses the reliable memory dictionary to learn the discriminative features of difficult tracks, which provides a robust way to solve the inherent complexity of USL-VVI-ReID without using artificial labels, and significantly improves the performance without using artificial labels, facilitating popularization and use.

[0035] The technical solutions of the present application will be further described in detail below with the aid of the accompanying drawings and examples. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 The method flowchart of the present application.

[0037] Figure 2 The training process framework diagram of the pedestrian re-identification model of the present application.

[0038] Figure 3 The pedestrian feature space distribution visualization effect comparison chart corresponding to the baseline method, the baseline method combined with DRM, the baseline method combined with DRM and FSAM, and the present method in the present embodiment.

[0039] Figure 4 The similarity visualization effect comparison chart of the pedestrian features corresponding to the baseline method and the present method in the present embodiment.

[0040] Figure 5 The cross-modal pedestrian retrieval ranking visualization effect comparison chart of the benchmark method and the method corresponding to the visible light modal to the infrared modal in the embodiment.

[0041] Figure 6 The cross-modal pedestrian retrieval ranking visualization effect comparison chart of the benchmark method and the method corresponding to the infrared modal to the visible light modal in the embodiment.

[0042] Figure 7 The heat map visualization effect comparison chart of the benchmark method and the method of randomly selecting 5 consecutive frames of pedestrian pictures in the visible light modal in the embodiment.

[0043] Figure 8 The heat map visualization effect comparison chart of the benchmark method and the method of randomly selecting 5 consecutive frames of pedestrian pictures in the infrared modal in the embodiment. DETAILED DESCRIPTION

[0044] As shown in Figure 1 and Figure 2 The unsupervised cross-modal video pedestrian re-identification method based on reinforced semantic consistency of the application comprises:

[0045] An image of a re-identified pedestrian is obtained, wherein the image of the re-identified pedestrian comprises visible light images and infrared images, a plurality of visible light images form a visible light modal, and a plurality of infrared images form an infrared modal;

[0046] The plurality of images of the re-identified pedestrian are input into a trained pedestrian re-identification model to obtain a re-identification result, and the training process of the pedestrian re-identification model comprises:

[0047] Step 1, the feature mean value of each frame image in each track of each modal is calculated, and the cosine similarity between the feature value of each frame image and the feature mean value in the corresponding track is calculated, a cosine similarity threshold parameter in each track is set, the frame with a cosine similarity greater than the cosine similarity threshold parameter in the corresponding track is a simple frame, and the frame with a cosine similarity not greater than the cosine similarity threshold parameter in the corresponding track is a difficult frame;

[0048] The difficult frame feature and the easy frame feature are constructed for each track;

[0049] The difficult frame feature and the easy frame feature are respectively subjected to average pooling processing to obtain a difficult track feature and a simple track feature;

[0050] It should be noted that the simple frame and the difficult frame in each track are identified and screened in the embodiment. The simple frame is used to predict a pseudo label and construct a cross-modal corresponding relationship, and the difficult frame is used to enhance the learning of global discriminative features of the model.

[0051] Step 2, cluster simple orbit features of the visible light modality and the infrared modality respectively, and predict pseudo-labels in the visible light modality and the infrared modality; and generate a reliable memory dictionary of the visible light modality and a reliable memory dictionary of the infrared modality respectively according to the cluster centers of the simple orbit features and the predicted pseudo-labels in the visible light modality and the infrared modality;

[0052] cluster difficult orbit features of the visible light modality and the infrared modality respectively, and generate a difficult memory dictionary of the visible light modality and a difficult memory dictionary of the infrared modality respectively according to the cluster centers of the difficult orbit features;

[0053] Step 3, Fourier enhance simple frames in the simple orbit features;

[0054] apply an optimal transport strategy to establish a reliable cross-modality matching relationship between the Fourier-enhanced visible light modality and the infrared modality, and at the same time minimize the transport plan cost, to obtain an optimal transport scheme, and based on the optimal transport scheme, obtain two groups of pseudo-labels matched with each other;

[0055] It should be noted that the present embodiment emphasizes the advanced semantic information in each modality and maps it into the same phase space, thereby improving the cross-modality matching accuracy. For any input simple frame (visible light or infrared light), the method can enhance the advanced semantic information in the same modality, and only needs to integrate the phase component information back into the original frame without including the amplitude component.

[0056] Step 4, use an asymmetric cross-contrast learning strategy to perform interactive contrast learning between the simple orbit features of the visible light modality and the difficult memory dictionary of the visible light modality, interactive contrast learning between the difficult orbit features of the visible light modality and the reliable memory dictionary of the visible light modality, interactive contrast learning between the simple orbit features of the infrared modality and the difficult memory dictionary of the infrared modality, and interactive contrast learning between the difficult orbit features of the infrared modality and the reliable memory dictionary of the infrared modality, to obtain a first loss function;

[0057] use the asymmetric cross-contrast learning strategy to perform interactive contrast learning between the Fourier-enhanced cross-modality matching simple orbit features and the reliable memory dictionary of the visible light modality and the reliable memory dictionary of the infrared modality respectively, to obtain a second loss function;

[0058] use the first loss function and the second loss function to promote the learning of the common and discriminative modality-invariant features of the human re-identification model.

[0059] In the present embodiment, in Step 2, DBSCAN algorithm is used to cluster the simple orbit features of the visible light modality and the infrared modality respectively.

[0060] In this embodiment, in step 3, the process of Fourier enhancement on the simple frame in the simple track feature is as follows:

[0061] The inverse Fourier transform is performed on the simple frame in the simple track feature, the amplitude component after the inverse Fourier transform retains the low-level modality perception information of the simple frame, and the phase component after the inverse Fourier transform retains the high-level semantic information of the simple frame. The phase component information is integrated back into the original simple frame to enhance the high-level semantic information under the same modality for the input simple frame;

[0062] that is, the simple frame after Fourier enhancement , wherein, is the simple frame in the simple track feature, is the phase component after the inverse Fourier transform, is a proportional coefficient.

[0063] In this embodiment, in step 4, by using the asymmetric cross-contrast learning strategy, the simple track feature of the visible light modality is interactively compared and learned with the difficult memory dictionary of the visible light modality, the difficult track feature of the visible light modality is interactively compared and learned with the reliable memory dictionary of the visible light modality, and at the same time, the simple track feature of the infrared modality is interactively compared and learned with the difficult memory dictionary of the infrared modality, and the difficult track feature of the infrared modality is interactively compared and learned with the reliable memory dictionary of the infrared modality, to obtain a first loss function ; wherein, is the asymmetric cross-contrast learning strategy, is the simple track feature of the visible light modality, is the difficult memory dictionary of the visible light modality, is the difficult track feature of the visible light modality, is the reliable memory dictionary of the visible light modality; is the simple track feature of the infrared modality, is the difficult memory dictionary of the infrared modality, is the difficult track feature of the infrared modality, is the reliable memory dictionary of the infrared modality;

[0064] By using the asymmetric cross-contrast learning strategy, the simple track feature after the Fourier enhancement and the cross-modal matching is interactively compared and learned with the reliable memory dictionary of the visible light modality and the reliable memory dictionary of the infrared modality respectively, to obtain a second loss function ; wherein, is the simple track feature after the Fourier enhancement and the cross-modal matching;

[0065] According to the formula , a comprehensive loss function of the pedestrian re-identification model is obtained , wherein, is a first weight coefficient, is a second weight coefficient and .

[0066] In use, the application first distinguishes and filters simple frames and difficult frames in each trajectory segment, wherein the simple frames model simple track features for clustering to predict more accurate pseudo labels and establish cross-modal correspondence, and the difficult frames model valuable track features for learning global discriminative features. In the clustering stage, the method uses a clustering algorithm to predict pseudo labels for simple track features, and at the same time, initializes a reliable memory dictionary with its clustering prototype, and on the other hand, initializes a difficult memory dictionary with the clustering prototype of valuable track features. In the matching stage, a Fourier-based semantic enhancement matching module is proposed to establish a reliable cross-modal correspondence between Fourier-enhanced visible light and infrared modalities using an optimal transport strategy. An asymmetric collaborative learning strategy is introduced in the training, which compares and learns the reliable simple track features and valuable difficult track features with the two memory dictionaries respectively, promotes the model to learn common and discriminative modal invariant features, and aims to learn robust features from the frame feature distribution discrete track and visible light-infrared modal features in an unsupervised manner. To solve the inherent complexity of USL-VVI-ReID, a robust way is provided to achieve significant performance improvement without using artificial labels.

[0067] The experimental platform used in the embodiment is a desktop workstation equipped with Intel(R) Xeon(R) CPU E5-2650 v2, 256GB of memory, and 4 NVIDIA GeForce RTX 2080 GPUs. A 64-bit Ubuntu 22.04.2LTS operating system and a Pytorch 1.12.1 deep learning framework are used for implementation.

[0068] The embodiment adopts AGW pre-trained on ImageNet as the backbone network. All video frame images are adjusted to 288144, and random flipping, random cropping, random erasing and channel augmentation are used as data enhancement means. The video sequence sampling batch size of the clustering stage is set to 8, and 12 frames are randomly sampled for each video sequence. The video sequence sampling batch size of the training stage is set to 16, and 4 frames are randomly sampled for each video sequence. All frames of each video sequence are used for testing in the test stage, and the sampling batch size is set to 1. In the training stage, the Adam optimizer is used to optimize the network, and the weight decay is 0.0005. In the first 50 cycles, the double-branch contrast learning network (DCL) is used as a benchmark method to train the model to learn specific modal information. In the last 50 cycles, the method is used for training. The initial learning rate is set to 0.00035, and then it is reduced to 1 / 10 of the original every 20 training cycles. The momentum update factor of the memory dictionary is set to 0.1, and the fusion proportion parameter is set to 0.5. The experiment uses two test modes for evaluation: Visible2Infrared and Infrared2Visible.

[0069] Since there are currently few works on unsupervised learning-based cross-modal video pedestrian re-identification methods, the method is compared with unsupervised cross-modal image pedestrian re-identification methods in the HITSZ-VCM dataset, including: ADCA, DOTLA, MBCCM, CCLNet, PGM and GUR. The random sampling frame-level features extracted by the unsupervised cross-modal image pedestrian re-identification method are compared and the video sequence representation is modeled by average pooling. Table 1 shows the comparison results of the method and the above methods.

[0070] Table 1

[0071]

[0072] As shown in Table 1, the method is significantly better than the existing unsupervised method. Specifically, the mAP accuracy of the method is improved by 5.4% and the Rank-1 accuracy is improved by 4.8% in the Visible2Infrared mode, and the mAP accuracy is improved by 5.2% and the Rank-1 accuracy is improved by 5.3% in the Infrared2Visible mode. The above results show that the method has better temporal information modeling and cross-modal consistent feature learning ability, significantly improving the performance of the unsupervised cross-modal video pedestrian re-identification model.

[0073] The key modules in the method, i.e., the video sequence modeling based on difficult frames and simple frames (DRM) module, the Fourier enhanced matching (FSAM) module, and the asymmetric cross-contrast learning (UCL) module, are further evaluated through ablation experiments. The ablation experiments take the double-branch-based contrast learning framework as the baseline method, and Table 2 shows the comparison results of the ablation experiments in the Visible2Infrared and Infrared2Visible modes.

[0074] Table 2

[0075]

[0076] To qualitatively evaluate the performance of the method, four groups of visualization experiments are performed, including pedestrian feature space distribution visualization, pedestrian feature similarity visualization, pedestrian retrieval ranking visualization, and feature map visualization.

[0077] Pedestrian feature space distribution visualization: The experiment randomly selects 7 pedestrian video sequences, and uses the t-SNE method to embed the merged and clustered feature of the visible light mode and infrared mode feature distribution into a two-dimensional space. As shown in Figure 3 , a figure is the pedestrian feature space distribution visualization effect diagram corresponding to the baseline method in the embodiment; b figure is the pedestrian feature space distribution visualization effect diagram corresponding to the baseline method combined with DRM in the embodiment; c figure is the pedestrian feature space distribution visualization effect diagram corresponding to the baseline method combined with DRM and FSAM in the embodiment; d figure is the pedestrian feature space distribution visualization effect diagram corresponding to the method in the embodiment. Different colors represent different pedestrian identities, and different shapes represent different modalities. As can be seen from the figure, compared with the baseline model, the intra-class feature distribution of the same pedestrian under different modalities extracted by the method is more compact (see the blue dashed line circle), and the inter-class feature distribution overlapping different pedestrians is more obvious (see the red dashed line circle). This shows that the method significantly improves the model's ability to learn cross-modal retrieval by establishing a reliable cross-modal correspondence.

[0078] Pedestrian feature similarity visualization: As shown in Figure 4 , a figure is the pedestrian feature similarity visualization effect diagram corresponding to the baseline method in the embodiment; b figure is the pedestrian feature similarity visualization effect diagram corresponding to the method in the embodiment. a figure and b figure show the comparison of the pedestrian feature similarity visualization results of the baseline method and the method, where purple is the intra-class distance, and green is the inter-class distance. As can be seen from the figure, the intra-class similarity of the extracted features is further enhanced by the method, and the inter-class distance is relatively reduced compared with the baseline. In addition, the gap between the intra-class and inter-class average distances is effectively increased, i.e., the gap between the intra-class and inter-class average distances of the baseline method is less than the gap between the intra-class and inter-class average distances of the method The results show that, compared with the intra-class distance of baseline features, the method can effectively increase the intra-class compactness of the same pedestrian feature distribution and reduce the difference between modalities.

[0079] Cross-modal pedestrian retrieval ranking visualization: As shown in Figure 5 , a figure is the cross-modal pedestrian retrieval ranking visualization effect diagram of the visible light modality to the infrared modality corresponding to the benchmark method in the embodiment; b figure is the cross-modal pedestrian retrieval ranking visualization effect diagram of the visible light modality to the infrared modality corresponding to the method in the embodiment; the cross-modal pedestrian retrieval ranking visualization result of the visible light modality to the infrared modality is visualized. As shown in Figure 6 , a figure is the cross-modal pedestrian retrieval ranking visualization effect diagram of the infrared modality to the visible light modality corresponding to the benchmark method in the embodiment; b figure is the cross-modal pedestrian retrieval ranking visualization effect diagram of the infrared modality to the visible light modality corresponding to the method in the embodiment; the cross-modal pedestrian retrieval ranking visualization result of the infrared modality to the visible light modality is visualized, and the first 10 video sequences in the gallery set most similar to it are listed. In order to simplify the representation, only the first frame in each video sequence is shown in the figure, the first column on the left is the video sequence to be retrieved, and the right side is the comparison of the pedestrian retrieval ranking visualization results of the benchmark method and the method, wherein the correct and incorrect retrieval results are represented by green numbers and red numbers respectively. From the results, it can be seen that the method is more robust to environmental factors such as occlusion, and the results show that the method can more accurately model the pedestrian video sequence features, promote the model to learn the discriminative features of the diversity frame feature distribution, and at the same time, the bidirectional cross-modal feature alignment method further alleviates the modal domain difference, effectively improving the cross-modal retrieval performance of the model.

[0080] Feature map visualization: the experiment uses the GradCAM method to compare the heat map visualization results of the benchmark method and the method, and randomly selects 5 consecutive frames of pedestrian pictures for each modality video sequence, as shown in Figure 7 , a figure is the effect diagram of randomly selecting 5 consecutive frames of pedestrian pictures in the visible light modality in the embodiment; b figure is the heat map visualization effect diagram of a figure using the benchmark method; c figure is the heat map visualization effect diagram of a figure using the method; the comparison results are as shown in Figure 7 , the heat map comparison results of the visible light modality can be seen, the benchmark method can accurately capture part of the pedestrian region features, but there are still many background interference frames, which affect the model to learn the discriminative feature information of the pedestrian; the method further focuses on the pedestrian region, and the degree of influence by the background factor is the lowest, indicating that the video sequence modeling method based on dynamic representation proposed by the method can robustly represent the global discriminative features of the video frame. As shown in Figure 8As shown, a figure is a random selection of 5 consecutive pedestrian picture effect diagram under infrared modal in the embodiment; b figure is the heat map visualization effect diagram of the reference method of a figure; c figure is the heat map visualization effect diagram of the method of a figure; the comparison results are as follows Figure 8 The comparison results of the heat map of the infrared modal can be seen that, due to the lack of color information, the reference method can accurately determine the key features of the pedestrian to a certain extent, but at the same time, the video frame is interfered by the environmental factors to different degrees, and the attention to the pedestrian area is single (for example, the head), which is not conducive to the model learning the diversity distribution information of the video frame features; the method can accurately identify the pedestrian area, and can capture the diversity discriminative feature information (for example, the head, the hand, the foot and the like) of different frames, and effectively improve the robustness of the model.

[0081] The above is only a preferred embodiment of the present application, not any limitation on the present application, any simple modification, change and equivalent structure change of the above embodiment according to the technical essence of the present application are still within the protection scope of the technical solution of the present application.

Claims

1. An unsupervised cross-modal video person re-identification method based on reinforcement semantic consistency, characterized in that, The method comprises the following steps: Obtain the image of the re-identified pedestrian, which includes visible light images and infrared images. Multiple visible light images form a visible light mode, and multiple infrared images form an infrared mode. Input the multiple images of the re-identified pedestrian into a trained pedestrian re-identification model to obtain a re-identification result. The training process of the pedestrian re-identification model comprises the following steps: Step 1: Calculate the feature mean value of each frame image in each track of each mode, and the cosine similarity between the feature value of each frame image and the feature mean value in the corresponding track. Set the cosine similarity threshold parameter in each track. The frame with a cosine similarity greater than the cosine similarity threshold parameter in the corresponding track is a simple frame, and the frame with a cosine similarity not greater than the cosine similarity threshold parameter in the corresponding track is a difficult frame. Construct the difficult frame feature and the easy frame feature for each track. Perform average pooling processing on the difficult frame feature and the easy frame feature to obtain the difficult track feature and the simple track feature. Step 2: Cluster the simple track features of the visible light mode and the infrared mode respectively to predict the pseudo-labels in the visible light mode and the infrared mode. According to the predicted pseudo-labels in the visible light mode and the infrared mode, generate a reliable memory dictionary of the visible light mode and a reliable memory dictionary of the infrared mode respectively based on the clustering centers of the simple track features. Cluster the difficult track features of the visible light mode and the infrared mode respectively to generate a difficult memory dictionary of the visible light mode and a difficult memory dictionary of the infrared mode respectively based on the clustering centers of the difficult track features. Step 3: Fourier enhance the simple frames in the simple track features. Establish a reliable cross-modal matching relationship between the Fourier-enhanced visible light mode and infrared mode by applying an optimal transport strategy, and at the same time minimize the transport plan cost to obtain an optimal transport scheme. Based on the optimal transport scheme, two groups of mutually matched pseudo-labels are obtained. Step 4: Use an asymmetric cross-contrast learning strategy to perform interactive contrast learning between the simple track features of the visible light mode and the difficult memory dictionary of the visible light mode, and between the difficult track features of the visible light mode and the reliable memory dictionary of the visible light mode. At the same time, perform interactive contrast learning between the simple track features of the infrared mode and the difficult memory dictionary of the infrared mode, and between the difficult track features of the infrared mode and the reliable memory dictionary of the infrared mode to obtain a first loss function. Use an asymmetric cross-contrast learning strategy to perform interactive contrast learning between the Fourier-enhanced cross-modal matching simple track features and the reliable memory dictionary of the visible light mode and the reliable memory dictionary of the infrared mode respectively to obtain a second loss function. Use the first loss function and the second loss function to promote the learning of the modal invariant features of the pedestrian re-identification model.

2. The unsupervised cross-modal video person re-identification method based on reinforced semantic consistency according to claim 1, characterized in that: In step 2, DBSCAN algorithm is used to cluster the simple track features of the visible light mode and the infrared mode respectively.

3. The unsupervised cross-modal video person re-identification method based on reinforced semantic consistency according to claim 1, characterized in that: In step 3, the process of Fourier enhancing the simple frames in the simple track features is as follows: The inverse Fourier transform is performed on the simple frame in the simple track feature, the amplitude component after the inverse Fourier transform retains the low-level modality perception information of the simple frame, the phase component after the inverse Fourier transform retains the high-level semantic information of the simple frame, and the phase component information is integrated back into the original simple frame to enhance the high-level semantic information under the same modality for the input simple frame; simple frame after Fourier enhancement wherein, is a simple frame in the simple track feature, is a phase component after inverse Fourier transform, is a scale factor.

4. The unsupervised cross-modal video person re-identification method based on reinforced semantic consistency according to claim 1, characterized in that: In step 4, by using the asymmetric cross contrast learning strategy, the first loss function is obtained by interactive contrast learning of the simple track features of the visible light modality with the difficult memory dictionary of the visible light modality, interactive contrast learning of the difficult track features of the visible light modality with the reliable memory dictionary of the visible light modality, and at the same time, by interactive contrast learning of the simple track features of the infrared light modality with the difficult memory dictionary of the infrared light modality, interactive contrast learning of the difficult track features of the infrared light modality with the reliable memory dictionary of the infrared light modality ; wherein, the asymmetric cross contrast learning strategy, the simple track features of the visible light modality, the difficult memory dictionary of the visible light modality, the difficult track features of the visible light modality, the reliable memory dictionary of the visible light modality; the simple track features of the infrared light modality, the difficult memory dictionary of the infrared light modality, the difficult track features of the infrared light modality, the reliable memory dictionary of the infrared light modality; By using an asymmetric cross-contrast learning strategy, the second loss function is obtained by respectively interacting and contrastively learning the Fourier-enhanced cross-modality matching simple orbit features with a reliable memory dictionary of the visible light modality and a reliable memory dictionary of the infrared modality ; wherein is the Fourier-enhanced cross-modality matching simple orbit feature According to the formula , the comprehensive loss function of the pedestrian re-identification model is obtained , wherein, is a first weight coefficient, is a second weight coefficient and .

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on self-supervised learning and pre-training model

    CN116052057A

  • Semantic analysis of video data for event detection and validation

    US12456299B1