Electric power dangerous personnel identity identification method and system based on multi-modal data
Through multimodal data processing and tag association optimization, the problem of low recognition accuracy in low light environments in traditional identity recognition technology is solved, and higher cross-modal personnel re-identification accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510297328.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional single-modal personnel identity recognition technology has low recognition accuracy in low light or night environments, and relies on large-scale labeling data, manual labeling is expensive, making it difficult to effectively capture the fine-grained characteristics of individuals.
The identification method of power hazard personnel based on multimodal data is adopted. By obtaining pedestrian image data of visible light and infrared modes, fine-grained features are extracted, Jaccard distance and similarity matrix are calculated, pseudo-labels are generated using clustering algorithms, pseudo-labels are optimized, cross-modal tag associations are established, and cross-modal tag associations are introduced into model training as supervisory signals.
The accuracy of cross-modal personnel re-identification is improved, the limitations of traditional methods in detail capture and inter-modal mapping are overcome, and the generalization ability and robustness of the model in cross-modal scenarios are enhanced.
Smart Images

Figure CN120164235A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and pattern recognition, and particularly to a method and system for identifying the identities of power dangerous personnel based on multi-modal data. Background Art
[0002] With the development of computer vision technology, personnel identity recognition, as an important visual task, is widely applied to various industrial scenarios. The goal of personnel identity recognition is to confirm and verify the identity of a specific individual in a given scenario or application.
[0003] In view of the complex and changeable grid-connected entity scenarios in the power system, how to quickly and accurately identify the identities of dangerous personnel remains an urgent problem to be solved. Traditional single-modal personnel identity recognition technologies can only retrieve visible light images captured under visible light, so they can only accurately identify in an environment with good lighting conditions. However, in practical applications, especially in the power system, night or low-light environments are often high-frequency scenarios for dangerous personnel activities. In such an environment, the recognition effect relying solely on visible light images will be greatly reduced, and the recognition accuracy cannot be effectively guaranteed.
[0004] Traditional personnel identity recognition algorithms often rely on large-scale labeled data. With the popularization of intelligent surveillance cameras, especially cameras equipped with infrared functions, visible light-infrared personnel identity recognition tasks have gradually become a research hotspot. Although a large number of studies have achieved remarkable results in this task, these methods usually rely on large-scale labeled cross-modal data sets. However, manually labeling such a large amount of cross-modal data is not only extremely difficult but also extremely costly, which limits the feasibility of these methods in practical industrial applications.
[0005] Traditional personnel identity recognition methods often have difficulty effectively capturing the fine-grained features of individuals. Due to relying too much on rough global features, these methods are easily interfered by modal differences and are difficult to solve the differences between the two modalities, namely intra-modal differences and inter-modal differences, resulting in a relatively coarse granularity of the single global feature matching method and being unable to accurately represent individuals. Therefore, the cross-modal similarity obtained based on this method lacks sufficient reliability, ultimately affecting the accuracy of personnel identity recognition. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide a method and system for identifying the identities of power dangerous personnel based on multi-modal data to solve or at least partially solve the above problems existing in the prior art.
[0007] To achieve the above purpose, the first aspect of the present invention provides a method for identifying the identities of power dangerous personnel based on multi-modal data, and the method specifically includes the following steps:
[0008] Obtain a pedestrian identity image dataset in a multi-modal scenario, collect pedestrian images in visible light modality and infrared modality respectively, and construct a visible light dataset and an infrared dataset. The visible light dataset is a set of visible light images, and the infrared dataset is a set of infrared images;
[0009] Extract features from the obtained pedestrian identity image dataset, extract the fine-grained information of the images, obtain the image feature representation, and capture the diverse features of the images;
[0010] Use the diverse features to calculate the Jaccard distance between visible light images and infrared images respectively, obtain the similarity matrix for each modality, and use the clustering algorithm to cluster the visible light images and infrared images respectively according to the calculated similarity matrix, and generate the pseudo-labels for each cluster;
[0011] Introduce multiple class markers to represent the embedded representations of visible light images and infrared images. For different class markers, first calculate the similarity matrix with other class markers to obtain the nearest neighbors of each class marker, and then perform an intersection operation on the nearest neighbors of each class marker to obtain a refined neighborhood set. Finally, combine the camera labels to filter the neighborhood set and optimize the pseudo-labels to ensure that the images in the same set come from the same camera;
[0012] Establish cross-modal label associations according to the optimized pseudo-labels, obtain the cross-modal label associations between the visible light modality and the infrared modality, and introduce them as supervision signals into the subsequent training process of the model;
[0013] Based on the above steps, train the model, and at the same time evaluate the model. The evaluation process will measure the convergence and learning progress of the model by calculating the loss function and evaluation metrics according to the learning effect of the model during the training process, and perform tuning;
[0014] Divide and obtain a test set from the pedestrian identity image dataset, and divide it into a query set and a gallery set. Extract the image features of the query set and the gallery set, input them into the model for model testing, calculate the corresponding evaluation metrics to obtain the test results, and reflect the comprehensive performance of the model.
[0015] Further, the step S101 specifically includes the following steps:
[0016] Collect pedestrian images in visible light and infrared modalities in a multi-modal scenario, and construct a visible light dataset and an infrared dataset respectively;
[0017] Perform data augmentation operations on the constructed visible light dataset and infrared dataset respectively to generate a multi-modal pedestrian identity image dataset;
[0018] The generated multi-modal pedestrian identity image dataset is divided into a training set and a test set.
[0019] Furthermore, the pedestrian identity image dataset obtained is subjected to feature extraction, which specifically includes the following steps:
[0020] Taking the obtained pedestrian identity image dataset as the input image, performing patch embedding on the input image, and extracting features from the image through multi-layer convolution operations to obtain a feature map;
[0021] Using a convolutional layer to map the feature map to the target embedding dimension, ensuring that the size of each patch in the feature space adapts to the requirements of the model, and flattening the feature map so that the features of each patch become a feature vector.
[0022] Furthermore, the visible light images and infrared images are clustered using a clustering algorithm, which specifically includes the following steps:
[0023] Determining the clustering parameters for the visible light dataset and the infrared dataset, setting the eps parameter values for the two data sources, setting the minimum number of samples for clustering, and using the clustering algorithm to cluster the visible light images and infrared images respectively, so as to assign a pseudo-label to each cluster;
[0024] Initializing two memory modules for the clustered visible light images and infrared images respectively, using the memory modules to store the central features of the clusters. After initialization, assigning the normalized central features of the clusters to the memory modules, and storing the memory modules in the trainer object.
[0025] Furthermore, obtaining a neighborhood set by performing an intersection operation on the nearest neighbors of various labels, and screening the neighborhood set in combination with the camera labels, specifically including:
[0026] Instantiating two memory objects respectively for storing the feature maps of RGB and IR, extracting the class label features from the obtained feature vectors, calculating the nearest neighbors of each class label feature one by one, and merging the nearest neighbors calculated for different class labels to fuse the neighborhood structures generated by different semantic information to obtain a more comprehensive neighborhood relationship;
[0027] Loading the pseudo-label data of the camera, for each camera, loading the pseudo-label data file, and clustering the images under each camera according to the pseudo-labels;
[0028] Neighbor screening, respectively processing the nearest neighbor lists of infrared images, visible light images, infrared-to-visible light images, and visible light-to-infrared images, checking whether the neighbor is taken by the same camera as the current image, and obtaining a more refined neighborhood set.
[0029] Furthermore, establishing cross-modal label associations specifically includes:
[0030] Based on the extracted feature maps, the visible light modality and the infrared modality respectively obtain corresponding multiple token features. For a pair of cross-modal clusters, the cross-modal similarity between them is calculated;
[0031] A matching relationship is established by selecting a pair of cross-modal clusters with the highest similarity from the similarity matrix. When the number of visible light clusters corresponding to the infrared clusters reaches a preset threshold, the matching process stops, ensuring that clusters representing the same identity can maintain consistent cross-modal labels.
[0032] Furthermore, the model training specifically includes:
[0033] Obtain the infrared dataset and the visible light dataset, extract the labels and indices of the infrared dataset and the visible light dataset, input the visible light images and the infrared images into the model to obtain their feature representations, map the input images to the feature space, and calculate the similarity matrix between the visible light images and the infrared images;
[0034] Further aggregate different visible light clusters with similarity, so that visible light images and infrared images with the same identity are close in the embedding space, while visible light images and infrared images with different identities are far away;
[0035] Calculate the within-modal contrast loss, the cross-modal contrast loss, the Dissimilar loss, the neighborhood loss, and the mate loss, and calculate the total loss according to the above losses, which is expressed as follows:
[0036] L = L id + L SDC + β1L scl + β2L neighbor + β3L mate
[0037] where L is the total loss, L id is the within-modal contrast loss, L SDC is the cross-modal contrast loss, L scl is the Dissimilar loss, L neighbor is the neighborhood contrast loss, L mate is the homogeneous fusion loss, and β1, β2, and β3 are the weighted parameters for coordinating L scl 、L neighbor 、L mate of the three different losses.
[0038] Furthermore, a computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method according to any one of claims 2-7.
[0039] In a second aspect of the present invention, a power dangerous personnel identity recognition system based on multimodal data is provided. The system includes:
[0040] A diversified token matching module: used to represent the embedded representation of an image using multiple class tokens, thereby establishing cross-modal label associations between data of different modalities. The cross-modal label associations will be introduced as supervision signals into the subsequent training process of the model;
[0041] A diversified token neighborhood learning module: used to mine the intra-modal consistency and inter-modal consistency between diversified tokens in each training. The intra-modal consistency reduces the differences brought by illumination and pose changes under the same modality, and the inter-modal consistency reduces the feature differences between different modalities, enhancing the generalization ability of the model in cross-modal scenarios. For each visible light image and infrared image, for different class tokens, first calculate the similarity matrix with other class tokens to obtain the nearest neighbors of each class token, and then obtain the domain set by performing an intersection operation on the nearest neighbors of each class token. Each class token represents different local information of the image, and each class token has a different neighborhood set. Finally, introduce the camera label in the diversified token domain learning module to jointly optimize the neighborhood set;
[0042] A homogeneous fusion module: used to aggregate different visible light clusters with the same identity. By calculating the similarity matrix between tokens in each modality, use the clustering algorithm to cluster the visible light image and the infrared image respectively, and then further aggregate different visible light clusters with similarity, so that visible light images with the same identity are associated in the feature space. The different visible light clusters with similarity are considered to belong to the same identity;
[0043] A joint module: used to jointly use the diversified token matching module, the diversified token neighborhood learning module, and the homogeneous fusion module to train the model, and continuously optimize during the model training process, so as to achieve the alignment of cross-modal labels, establish the association between visible light modality and infrared modality data, and thus achieve the effect of searching for the corresponding visible light image according to the infrared image, realizing cross-modal re-identification of dangerous personnel.
[0044] Compared with the prior art, the beneficial effects of the present invention are:
[0045] The present invention proposes a method and system for identifying the identities of power - dangerous personnel based on multi - modal data, which improves the accuracy of cross - modal person re - identification, overcomes the limitations of traditional methods in detail capture and inter - modal mapping. By mining the intra - modal consistency and inter - modal consistency between class labels in different modalities, it can accurately identify a reliable neighborhood set. By combining camera information and neighborhood features, the model can better capture modality - invariant discriminative features, improving the robustness and accuracy of cross - modal matching; by optimizing the clustering alignment of the same identity, the clusters of the same identity are pulled closer and the clusters of different identities are pushed farther apart. Through this method, the model can achieve a closer identity representation in cross - modal data, further improving the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following - described drawings are only the preferred embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic flowchart of a method for identifying the identities of power - dangerous personnel based on multi - modal data provided by an embodiment of the present invention;
[0048] Figure 2 It is a schematic structural diagram of a system for identifying the identities of power - dangerous personnel based on multi - modal data provided by an embodiment of the present invention;
[0049] Figure 3 It is a schematic framework diagram of a diverse token matching module provided by an embodiment of the present invention;
[0050] Figure 4 It is a schematic framework diagram of a diverse token neighborhood learning module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following describes the principles and features of the present invention with reference to the drawings. The listed embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.
[0052] Refer to Figure 1 , this embodiment provides a method for identifying the identities of power - dangerous personnel based on multi - modal data. The method specifically includes the following steps:
[0053] Obtain a pedestrian identity image data set in a multi - modal scenario. Specifically, collect pedestrian data sets in the visible - light modality and the infrared modality respectively, and construct a visible - light data set and an infrared data set. The visible - light data set is a set of visible - light images, and the infrared data set is a set of infrared images. It specifically includes the following steps:
[0054] Collect visible light and infrared modality pedestrian images in a multi-modal scenario, and construct a visible light dataset and an infrared dataset respectively;
[0055] Perform data augmentation operations on the constructed visible light dataset and infrared dataset respectively to generate a multi-modal pedestrian identity image dataset. The data augmentation operations include: random channel augmentation, color augmentation, image scaling, image padding, random cropping, and random horizontal flipping;
[0056] In this embodiment, the channel augmentation method is adopted, including: randomly erasing some regions in the image to simulate the occlusion of objects or the loss of some information in the image; randomly swapping channels to enhance the diversity of the image; converting some channels of the image into grayscale images or simulating different grayscale transformations, so that the image can still show features without color information and improve the model's learning ability for grayscale images;
[0057] Divide the generated multi-modal pedestrian identity image dataset into a training set and a test set according to a certain ratio for subsequent model training and evaluation.
[0058] Extract features from the obtained pedestrian identity image dataset, extract the fine-grained information of the image, obtain a more refined feature representation of the image, and capture the diverse features of the image. The specific steps are as follows:
[0059] Take the obtained pedestrian identity image dataset as the input image, perform patch embedding on the input image, and extract features from the image through multi-layer convolution operations to obtain a feature map. Specifically, each input image is divided into 16 small blocks, and each image block is mapped to an embedding vector. To preprocess the input image, a convolutional block consisting of three convolutional layers is used. The first convolutional layer uses a 7x7 convolutional kernel to extract the preliminary features of the input image, and the two side convolutional layers use 3x3 convolutional kernels to further extract the high-level features of the image;
[0060] In addition, BatchNorm and ReLU activation functions are used during the convolution process to normalize and activate the features to enhance the feature expression ability. After the convolution operation is completed, an additional convolutional layer is used to map the feature map of each image block to a feature space with a fixed dimension.
[0061] Use a convolutional layer to map the feature map to the target embedding dimension to ensure that the size of each patch in the feature space meets the requirements of the model. Flatten the feature map so that the features of each patch become a feature vector. Use the Transformer block to process the generated embedding vectors to capture the relationships between different parts of the image, and at the same time add position encoding to help the Transformer understand the spatial structure of the image blocks.
[0062] Calculate the Jaccard distance of visible light images and infrared images respectively using diverse features to obtain the similarity matrix for each modality. Then, use the DBSCAN algorithm to cluster the visible light images and infrared images based on the calculated similarity matrix, generating pseudo-labels for each cluster. For each cluster generated by DBSCAN, calculate the mean of all image features in the cluster as the central feature of the cluster. The cluster central feature is the average of all image features in the cluster, representing the "representative" feature of the class. The specific steps are as follows:
[0063] Determine the clustering parameters for visible light and infrared image data. The clustering algorithm uses DBSCAN, and set the eps parameter for the two data sources to 0.6, which means the density threshold for clustering is 0.6; the minimum number of samples for clustering is set to 4, indicating that each cluster should have at least 4 clusters; use the DBSCAN clustering algorithm to cluster the visible light images and infrared images respectively, so as to assign a pseudo-label to each cluster.
[0064] Initialize two memory modules for the clustered visible light images and infrared images respectively. Use the memory modules to store the central features of the clusters, and use the momentum update strategy to update the two settings during the subsequent training process. After initialization, assign the normalized cluster central features to the memory modules and store the memory modules in the trainer object for use in the subsequent training process.
[0065] Introduce multiple class labels to represent the embedded representations of visible light images and infrared images. For different class labels, first calculate the similarity matrix with other class labels to obtain the k-nearest neighbors of each class label, then perform an intersection operation on the k-nearest neighbors of each class label to obtain a refined neighborhood set. Finally, filter the neighborhood set in combination with the camera labels to optimize the pseudo-labels, ensuring that the images in the same set come from the same camera, thereby enhancing the robustness across cameras. Specifically, it includes:
[0066] Instantiate two memory objects, which are used to store the feature maps of RGB and IR respectively. Extract the class label features from the obtained feature vectors, calculate the nearest neighbor of each class label feature one by one, and merge the nearest neighbors calculated from different class labels to fuse the neighborhood structures generated by different semantic information to obtain a more comprehensive neighborhood relationship;
[0067] Load the pseudo-label data of the camera. For each camera, load the pseudo-label data file and cluster the images under each camera according to the pseudo-labels;
[0068] Neighbor screening, separately process the neighbor lists of infrared images, visible light images, infrared-to-visible light images, and visible light-to-infrared images, check whether the neighbor is taken by the same camera as the current image, obtain a finer neighborhood set, extract the N-class label feature vectors of the image from the features of infrared and visible light images, and use the consistency of the N embedding spaces to define the neighborhood set. Taking the visible light image as an example, its within-modal neighborhood set is represented as follows:
[0069]
[0070] Among them, is the within-modal neighborhood set obtained from the embedding space of the query image , is the within-modal neighborhood set obtained from the nth embedding space of the query image , The definition is as follows:
[0071]
[0072] Among them, is the feature representation of the query feature sample i in the nth embedding space in the visible light modality, is the feature of the candidate neighbor set of the nth class label for the query feature in the visible light modality (labeled as v), j is used to distinguish other samples relative to the query sample i, that is, j represents the sample index relative to the query sample i, N ν is the neighbor set of the nth embedding space for a certain query feature in the visible light modality, and γ is the neighbor selection range;
[0073] Similarly, the within-modal neighborhood set of the infrared image can also be defined in the same way, and then the neighborhood set of the infrared image in its embedding space is obtained With the mechanism of cross-modal alignment, the neighborhood set from the infrared image to the visible light image can be constructed and the neighborhood set from the visible light image to the infrared image
[0074] Given the query image , the neighborhood set and the corresponding local clustering generated by in-camera training and the corresponding camera pseudo-label c are obtained. On this basis, the refinement process of the neighborhood set is represented as follows:
[0075]
[0076] Among them, is the refinement process of the neighborhood set, For all samples except those from the pseudo-label c of the camera sample set. Similarly, the neighborhood set and can be refined in the same way. By removing redundant information, the refined neighborhood set focuses more on the key features across cameras and modalities, thereby enhancing the information transfer efficiency and the accuracy of cross-modal matching in the subsequent learning process.
[0077] Establish reliable cross-modal label associations based on the optimized pseudo-labels, obtain the cross-modal label associations between the visible light modality and the infrared modality, and introduce them as supervision signals into the subsequent training process of the model, specifically including:
[0078] According to the extracted feature maps, the visible light modality and the infrared modality respectively obtain corresponding N token features. For a pair of cross-modal clusters, an N×N cross-modal similarity can be obtained. The total similarity is obtained by calculating the sum of the similarities in the corresponding embedding spaces under different modalities. In this embodiment, S = {S(i,j)} is used to represent the similarity matrix, where each element represents the similarity between clusters φ i and φ j as follows:
[0079]
[0080] where N is the number of token features obtained after the feature extraction process for each modality's image, representing the features of different embedding spaces respectively, n is used to identify different embedding spaces, is the clustering feature for the nth embedding space in the visible light modality (labeled as v), where i is used to distinguish different clusters, is the clustering feature for the nth embedding space in the infrared modality (labeled as r), where j is used to distinguish different clusters, is the memory feature of the kth cluster, and the subscript e(n) can take {v,r}, representing the visible light modality and the infrared modality respectively, is the number of instances in the kth cluster in the visible light modality or the infrared modality, is the instance feature of the nth class label. The total similarity is calculated by aggregating the similarities from the same embedding space under different modalities;
[0081] To address the problem that the same pedestrian may be divided into different clusters due to camera differences, in this embodiment, a pair of cross-modal clusters with the highest similarity is selected from the similarity matrix to establish a matching relationship. When the number of visible light clusters of the infrared cluster reaches a preset threshold, the matching process stops, ensuring that the clusters representing the same identity can maintain consistent cross-modal labels. The most similar cross-modal cluster pair is selected by the following formula:
[0082] (i, j) = argmax C, i = 1, ..., Y v , j = 1, ..., Y r
[0083] where argmax obtains the row and serial number (i, j) of the global minimum value in the similarity matrix, and Y v is the maximum row value of the similarity matrix, specifically referring to the total number of clusters in the visible light modality, v is the visible light modality, and Y r is the maximum column value of the similarity matrix, specifically referring to the total number of clusters in the infrared modality, r is the infrared modality. By selecting the most similar cross-modal clusters from the current similarity matrix each time, it is ensured that multiple clusters associated with the same cross-modal cluster have a high degree of similarity, increasing the possibility that clusters from the same identity are assigned the same cross-modal label.
[0084] Based on the above steps, the model is trained, and at the same time, the model is evaluated. The evaluation process will be based on the learning effect of the model during the training process, and the convergence and learning progress of the model will be measured by calculating the loss function and evaluation metrics, and necessary tuning will be carried out, specifically including:
[0085] In this embodiment, the model is evaluated only when the training round >= 6 and any of the following conditions is met:
[0086] (1) The model is evaluated whenever "the current training round + 1" is divisible by the "evaluation step size". In this embodiment, the "evaluation step size" is set to 2;
[0087] (2) Evaluation is performed in the last epoch of training.
[0088] Obtain the current batch of infrared dataset and visible light dataset, perform processing such as normalization and channel enhancement on the infrared dataset and visible light dataset, and extract the labels and indices of the infrared dataset and visible light dataset. Initialize the set memory storage according to the strategy at each stage;
[0089] During the model training process, especially for application scenarios with a large training dataset, the dataset needs to be divided into multiple batches for training. A batch refers to a subset of the training dataset in one iteration (epoch), which contains multiple samples. In the specific implementation of the present invention, the batch size is set to 64, and this process is dynamically executed during each iteration to ensure that the model can effectively utilize the data in the current batch for training;
[0090] The visible light images and infrared images in the infrared dataset and visible light dataset are respectively input into the model. The model is used to extract the image feature representations in the current batch, map the input images to the feature space to load data, and obtain the high-dimensional feature vectors of the images through the convolutional layers and Transformer modules in the network. Using the extracted image features, the similarity matrix between the current visible light image and the infrared image is calculated;
[0091] Due to cross-modal label association, there is a high similarity between multiple visible clusters related to the same infrared cluster. These clusters are likely to be segmented clusters with the same identity. On this basis, the cross-modal label association is constrained, and different visible clusters with similarity are further aggregated, making the visible light images and infrared images with the same identity closer in the embedding space, while the visible light images and infrared images with different identities are farther apart, further alleviating the gap between modalities;
[0092] Given a visible query, assume len(R2V[V2R[y v ) = M, and M > 1 is satisfied. Therefore, multiple different visible clusters with the same cross-modal label can be obtained, denoted as where len(*) is the length of the obtained object, here referring to the number of visible light clusters with the same cross-modal label, R2V is the association relationship of cross-modal labels, referring to the label mapping from the infrared modality to the visible light modality, y v is the pseudo-label in the visible light modality, and M is the rule representing homogeneous fusion. M is used to specify when the homogeneous fusion process is performed. When the number of visible light clusters with the same cross-modal label satisfies M > 1, the homogeneous fusion will be performed;
[0093] An alternating contrast learning method is adopted, and specific-modal memories are used for homogeneous fusion, which is expressed as follows:
[0094]
[0095] where, is the homogeneous fusion process between visible light clusters. i and j are used to indicate two visible light clusters that need to be fused. The goal in this process is to merge visible light clusters with the same cross-modal label, q v (i) are different visible light clusters with multiple same cross-modal labels, and i is used to identify different clusters. N v is the number of clusters in the visible light modality, that is, the total number of clusters finally obtained through the feature extraction process in the visible light modality. k is the kth cluster, is the memory feature of different visible light clusters with the same cross-modal label, It is the memory feature of the k-th cluster in the visible light modality, and τ is the temperature coefficient, which is used to control the smoothness in the loss function;
[0096] Calculate the loss, which is expressed as follows:
[0097] Intra-modal contrast loss: For image pairs within the same modality, calculate the similarity between them and make a comparison;
[0098] Cross-modal contrast loss: Constrain the mapping relationship between infrared images and visible light images to ensure consistency in the cross-modal feature space;
[0099] Dissimilar loss: The dissimilar loss increases the interval between images of different classes, making their feature distances far apart;
[0100] Neighborhood loss: Enhance the model's ability to learn local structures;
[0101] Mate loss: By constraining and screening the labels of infrared images and RGB images, optimize the label information of infrared images so that each infrared image can establish a reasonable correspondence with multiple RGB images.
[0102] Calculate the total loss according to the above losses, which is expressed as follows:
[0103] L = L id + L SDC + β1L scl + β2L neighbor + β3L mate
[0104] Among them, L is the total loss, L id is the intra-modal contrast loss, L SDC is the cross-modal contrast loss, L neighbor is the neighborhood contrast loss, L mate is the homogeneous fusion loss, and β1, β2, and β3 are the weighted parameters for coordinating L scl 、L neighbor 、L mate The three different losses; perform backpropagation and optimize the network parameters, and at the same time conduct model evaluation at specific training epochs. In this embodiment, the performance evaluation of the model depends on three common metrics: CMC, mAP, and mINP.
[0105] Partition and obtain the test set from the pedestrian identity image dataset, and divide it into a query set and a gallery set. Extract the image features of the query set and the gallery set, input them into the model for model testing, calculate the corresponding evaluation metrics to obtain the test results, and reflect the comprehensive performance of the model.
[0106] A computer-readable storage medium stores a computer program which, when executed by a processor, implements any one of the above methods.
[0107] Referring to Figures 2 - 4 , another embodiment of the present invention provides a power hazard personnel identity recognition system based on multimodal data, and the system includes:
[0108] Diversified token matching module: It is used to reduce the impact of cross-modal differences on feature matching by using multiple class tokens to represent the embedded representation of an image, so as to establish a reliable cross-modal label association between data of different modalities. The cross-modal label association will be introduced as a supervision signal into the subsequent training process of the model. In Visual Transformer, class tokens are usually used to aggregate the global feature information of an image, and the design of introducing multiple class tokens enables each class token to focus on different feature dimensions, deeply mine the detailed information in the image, and then establish a highly reliable cross-modal label association.
[0109] Diversified token neighborhood learning module: It is used to mine the intra-modal consistency and inter-modal consistency between diversified tokens in each training. The intra-modal consistency reduces the differences brought by changes such as illumination and pose under the same modality, and the inter-modal consistency reduces the feature differences between different modalities, enhancing the generalization ability of the model in cross-modal scenarios. In order to accurately identify a reliable neighborhood set, for each visible light image and infrared image, for different class tokens, the similarity matrix with other class tokens is calculated one by one and the k nearest neighbors are screened. Among them, each class token represents different local information of the image, so each class token will have a different neighborhood set. Finally, the camera label is introduced into the diversified token domain learning module to jointly optimize the neighborhood set to avoid cross-camera mis-matching, thereby promoting robust neighborhood learning.
[0110] Homogeneous fusion module: It is used to aggregate different visible light clusters with the same identity. By calculating the similarity matrix between tokens in each modality, clustering algorithms are used for clustering visible light images and infrared images respectively, and then different visible light clusters with higher similarity are further aggregated. These different visible light clusters with high similarity are considered likely to belong to the same identity, that is, they have the same pedestrian ID in the dataset. Through further aggregation, the visible light images with the same identity are made more compact in the feature space, reducing the influence of noise.
[0111] Combined Module: It is used to jointly train a model using a diverse token matching module, a diverse token neighborhood learning module, and a homogeneous fusion module, and continuously optimize it during the model training process, so as to achieve precise alignment of cross-modal state labels, establish the correlation between visible light modality and infrared modality data, thereby achieving the effect of searching for the corresponding visible light image according to the infrared image, and realizing cross-modal re-identification of dangerous personnel.
[0112] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for identifying dangerous personnel in power based on multimodal data, characterized in that: The method specifically comprises the following steps: Acquire a pedestrian identity image dataset in a multimodal scenario, collect pedestrian images in visible light mode and infrared mode respectively, and construct a visible light dataset and an infrared dataset, wherein the visible light dataset is a collection of visible light images, and the infrared dataset is a collection of infrared images; Perform feature extraction on the obtained pedestrian identity image dataset, extract fine-grained information of the image, obtain image feature representation, and capture the diverse features of the image; The Jaccard distance of visible light images and infrared images is calculated using diversified features to obtain the similarity matrix of each modality. The visible light images and infrared images are clustered using clustering algorithms based on the calculated similarity matrix to generate pseudo labels for each cluster. Multiple class labels are introduced to represent the embedded representation of visible light images and infrared images. For different class labels, the similarity matrix with other class labels is first calculated to obtain the nearest neighbors of each class label. Then, a neighborhood set is obtained by performing an intersection operation on the nearest neighbors of each class label. Finally, the neighborhood set is filtered in combination with the camera label, and the pseudo label is optimized to ensure that the images in the same set come from the same camera. A cross-modal label association is established based on the optimized pseudo-labels to obtain the cross-modal label association of the visible light modality and the infrared modality, and it is introduced as a supervisory signal into the subsequent training process of the model; Based on the above steps, the model is trained and evaluated. The evaluation process will measure the convergence and learning progress of the model and perform optimization based on the learning effect of the model during the training process by calculating the loss function and evaluation indicators. A test set is obtained from the pedestrian identity image dataset, and then divided into a query set and a gallery set. The image features of the query set and the gallery set are extracted and input into the model for model testing. The corresponding evaluation indicators are calculated to obtain the test results, which reflect the comprehensive performance of the model.
2. The method for identifying dangerous personnel in electrical power based on multimodal data according to claim 1 is characterized in that: The step of obtaining a pedestrian identity image dataset in a multimodal scenario specifically includes the following steps: Collect visible light and infrared pedestrian images in multimodal scenes, and construct visible light datasets and infrared datasets respectively; Perform data enhancement operations on the constructed visible light dataset and infrared dataset respectively to generate a multimodal pedestrian identity image dataset; The generated multimodal pedestrian identity image dataset is divided into training set and test set.
3. The method for identifying dangerous personnel in power situations based on multimodal data according to claim 1 is characterized in that: The feature extraction of the obtained pedestrian identity image dataset specifically includes the following steps: The obtained pedestrian identity image dataset is used as the input image, the input image is patch-embedded, and the image is feature extracted through multi-layer convolution operations to obtain a feature map; A convolutional layer is used to map the feature map to the target embedding dimension, ensuring that the size of each patch in the feature space adapts to the requirements of the model, and the feature map is flattened so that the features of each patch become a feature vector.
4. The method for identifying dangerous personnel in power based on multimodal data according to claim 1 is characterized in that: The clustering of the visible light image and the infrared image using a clustering algorithm specifically includes the following steps: Determine the clustering parameters of the visible light data set and the infrared data set, set the EPS parameter values of the two data sources, set the minimum number of samples for clustering, use the clustering algorithm to cluster the visible light images and infrared images respectively, and assign a pseudo label to each cluster; Two memory modules are initialized for the clustered visible light image and infrared image respectively, and the memory modules are used to store the central features of the clusters. After initialization, the normalized central features of the clusters are assigned to the memory modules, and the memory modules are stored in the trainer object.
5. The method for identifying dangerous personnel in electrical power based on multimodal data according to claim 3 is characterized in that: The method obtains a neighborhood set by performing an intersection operation on the nearest neighbors of various tags, and filters the neighborhood set in combination with the camera tag, specifically including: Instantiate two memory objects to store RGB and IR feature maps respectively, extract class marker features from the obtained feature vectors, calculate the nearest neighbors of each class marker feature one by one, merge the nearest neighbors calculated for different class markers, and fuse the neighborhood structures generated by different semantic information to obtain a more comprehensive neighborhood relationship; Load the pseudo-label data of the camera. For each camera, load the pseudo-label data file and cluster the images under each camera according to the pseudo-label. Neighbor screening, processes the neighbor lists of infrared images, visible light images, infrared to visible light images, and visible light to infrared images respectively, checks whether the neighbor is taken by the same camera as the current image, and obtains a more refined neighborhood set.
6. The method for identifying dangerous personnel in electrical power based on multimodal data according to claim 3 is characterized in that: The establishing of cross-modal tag association specifically includes: According to the extracted feature graph, the visible light modality and the infrared modality each obtain corresponding multiple token features, and for a pair of cross-modal clusters, the cross-modal similarity between them is calculated; A pair of cross-modal clusters with the highest similarity is selected from the similarity matrix to establish a matching relationship. When the number of visible light clusters with infrared clusters reaches a preset threshold, the matching process stops, ensuring that clusters representing the same identity can maintain consistent cross-modal labels.
7. The method for identifying dangerous personnel in power based on multimodal data according to claim 1 is characterized in that: The model training specifically includes: Obtain infrared data sets and visible light data sets, extract labels and indexes of infrared data sets and visible light data sets, input visible light images and infrared images into the model, obtain their feature representations, map the input images to feature space, and calculate the similarity matrix between visible light images and infrared images; Different visible light clusters with similarity are further aggregated, so that visible light images and infrared images with the same identity are close in the embedding space, while visible light images and infrared images with different identities are far away from each other; Calculate the same-modality contrast loss, trans-membrane contrast loss, dissimilar loss, neighborhood loss, and mate loss, and calculate the total loss based on the above losses, as shown below: L=L id +L SDC +β1L scl +β2L neighbor +β3L mate Among them, L is the total loss, L id is the same-modality contrast loss, L SDC is the cross-modal contrast loss, L scl is the dissimilarity loss, L neighbor is the neighborhood contrast loss, L mate is the homogeneous fusion loss, β1, β2, β3 are the coordination L scl , L neighbor , L mate Weighting parameters for three different losses.
8. A computer-readable storage medium, characterized in that: A computer program is stored, and when the program is executed by a processor, any method in claims 2-7 is implemented.
9. A system for identifying dangerous personnel in power based on multimodal data, characterized in that: The system comprises: Diversified Token Matching Module: It is used to use multiple class tags to represent the embedded representation of the image and establish cross-modal label associations between data of different modalities. The cross-modal label associations will be introduced as supervisory signals into the subsequent training process of the model. Diversified token neighborhood learning module: used to mine the intra-modal consistency and inter-modal consistency between diversified tokens in each training. For each visible light image and infrared image, for different class labels, the similarity matrix with other class labels is calculated to obtain the nearest neighbors of each class label, and then the neighborhood set is obtained by performing an intersection operation on the nearest neighbors of each class label. Each class label represents different local information of the image, and each class label has a different neighborhood set. Finally, the camera label is introduced in the diversified token neighborhood learning module to jointly optimize the neighborhood set. Homogeneous fusion module: used to aggregate different visible light clusters with the same identity. By calculating the similarity matrix between tokens under each modality, clustering is performed on visible light images and infrared images using clustering algorithms respectively. Then, different visible light clusters with similarity are further aggregated so that visible light images with the same identity are associated in the feature space. The different visible light clusters with similarity are considered to belong to the same identity. Joint module: used to jointly use the diversified token matching module, the diversified token neighborhood learning module, and the homogeneous fusion module to train the model, realize the alignment of cross-modal labels, establish the association between visible light modality and infrared modality data, and realize cross-modal re-identification of dangerous persons.
Citation Information
Cited By
Electrical equipment fault detection method and system based on image processing
CN121147208A
Unsupervised visible light-infrared person re-identification method based on cyclic pairwise identity learning
CN121746749A