Miner safety monitoring method and system based on cross-modal pedestrian re-identification

By adopting cross-modal pedestrian re-identification technology in miner safety monitoring, using shared dual-stream ResNet50 backbone network and feature modules to calculate cross-modal and global losses, the problem of poor identification effect of traditional monitoring methods in complex environments is solved, and more efficient and accurate miner safety monitoring is achieved.

CN120014550AActive Publication Date: 2025-05-16CHINA UNIV OF MINING & TECH

Patent Information

Application Number
CN202510096783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Traditional miner safety management methods based on video surveillance are difficult to achieve accurate identification and rapid response in low-light, dusty and complex environments, and there are modal differences between visible light and infrared images, affecting the recognition effect.

Method used

A method based on cross-modal pedestrian re-identification is adopted to extract modal specific and shared features through a shared dual-stream ResNet50 backbone network, combined with local feature modules and global feature modules, and SP_Net modules and bilinear fully connected layer to calculate cross-modal alignment loss and global loss to achieve cross-modal features alignment and improve recognition performance.

Benefits of technology

Effectively eliminate the differences between the two modalities, extract more discriminant features, improve image recognition efficiency and cross-modal pedestrian re-identification performance, significantly improving the accuracy and rapid response capabilities of miners' safety monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014550A_ABST
    Figure CN120014550A_ABST
Patent Text Reader

Abstract

The invention discloses a miner safety monitoring method based on cross-modal pedestrian re-identification, and the method comprises the following steps: S1, inputting a data set into a shared double-flow ResNet50 backbone network, and obtaining feature maps preliminarily extracted from two modals; s2, inputting the feature maps of the two modes into a local feature module, equally dividing the feature maps into three local feature maps according to the heights of the feature maps, taking every two feature maps as a pair, inputting the same pair of local feature maps into the same SPNet module to obtain local feature vectors, and carrying out cross-mode alignment loss on the two local feature vectors; s3, inputting the feature maps of the two modes into a global feature module, obtaining a global feature vector through a bilinear full connection layer, and calculating a loss function; and S4, adjusting network parameters to enable the network to be optimal, and performing pedestrian matching on the to-be-identified cross-modal data set. The cross-modal pedestrian re-identification method can eliminate the difference between the two modals to extract the features with higher discrimination, improves the image identification efficiency, and improves the cross-modal pedestrian re-identification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and in particular relates to a miner safety monitoring method and system based on cross-modal pedestrian re-identification. Background Art

[0002] In the coal mining industry, miner safety monitoring is a key link to ensure the safety of miners and stable production operations. With the continuous increase in the depth of mine mining and the complexity of the operating environment, traditional monitoring methods face many challenges, including the existence of monitoring blind spots, difficulty in information collection, and severe environmental interference. In particular, the environment under the mine is usually characterized by low light, dust, and dense personnel. Traditional safety management methods based on video monitoring can no longer meet the needs of accurate identification and rapid response.

[0003] In recent years, with the rapid development of computer vision and artificial intelligence technology, cross-modal pedestrian re-identification technology has provided a new idea for solving the problem of miner safety monitoring. Cross-modal pedestrian re-identification technology can be used for rapid cross-view identification and matching retrieval of coal mine workers, and then accurately identify the location information of miners at a certain point in time, and identify whether they are in the coal mining face, water pump room, substation, excavation face, etc., to determine the regional location of miners, and facilitate the rapid positioning of personnel according to the time and location after matching and identification when a mine safety accident occurs, thereby providing accurate location information for rescue.

[0004] One of the main challenges faced by this technology in underground mine safety monitoring is the modal difference between visible light and infrared images. Due to the influence of human posture, viewing angle, lighting, low image resolution, occlusion and background between the same modalities, the network only focuses on the global area and ignores the local area. The feature extraction network cannot extract more discriminative features, resulting in poor recognition effect. Summary of the invention

[0005] The purpose of the present invention is to provide a miner safety monitoring method and system based on cross-modal pedestrian re-identification, which can eliminate the differences between the two modalities to extract more discriminative features, improve image recognition efficiency, and improve the performance of cross-modal pedestrian re-identification.

[0006] To achieve the above object, the present invention provides a miner safety monitoring method based on cross-modal pedestrian re-identification, comprising the following steps:

[0007] S1: Input the dataset into the shared two-stream ResNet50 backbone network to obtain the feature maps of the two modalities of modality-specific information and modality-shared information;

[0008] S2: Input the feature maps of the two modalities into the local feature module, and divide them into three parts of local feature maps according to their heights. Two of them are regarded as a pair, and the same pair of local feature maps is input into the same SP_Net module to obtain the local feature vector. The two local feature vectors are subjected to cross-modal alignment loss.

[0009] S3: Input the feature maps of the two modalities into the global feature module, obtain the global feature vector through the bilinear fully connected layer, and calculate the loss function;

[0010] S4: Adjust the network parameters to optimize the network and perform pedestrian matching on the cross-modal dataset to be identified.

[0011] As a further solution of the present invention: the shared dual-stream ResNet50 backbone network in S1 includes five network layers, each layer is a residual structure block, the parameters of the first two network layers are independent and used to extract modality-specific features, and the parameters of the last three network layers are shared and used to extract modality-shared features.

[0012] As a further solution of the present invention: the local feature module in S2 includes three independent SP_Net modules and three cross-modal alignment loss functions. The feature map output by S1 is evenly divided into three local feature maps according to its height. The local feature maps corresponding to different modalities are combined into a local feature map pair. The local feature map pair is input into the SP_Net module, and the SP_Net module outputs two local feature vectors of different modalities. The cross-modal alignment loss function is calculated for the local feature vectors of these two different modalities, and finally three cross-modal alignment loss functions are obtained, which are added to obtain the final local loss.

[0013] As a further solution of the present invention: the SP_Net module uses strip global average pooling to expand the field of view of the convolutional neural network. The process is as follows:

[0014] The global average pooling of the strips along the width direction is recorded as GAP 1*W , the global average pooling in the height direction is recorded as GAP H*1 , global average pooling is denoted as GAP, one-dimensional convolution layer is denoted as Conv, batch normalization is denoted as BN, input feature map is denoted as z, feature vector is denoted as f, visible light is denoted as RGB, and infrared is denoted as IR;

[0015] Perform one-dimensional convolution on the feature map z to reduce the dimension, and perform strip global average pooling on the feature map z obtained by the convolution layer in two directions to obtain the feature vector f:

[0016] f H*1 =GAP 1*W (Conv(z))

[0017] f1*W =GAP H*1 (Conv(z))

[0018] The obtained feature vector is inner-producted to obtain a new feature map z H*W :

[0019] z H*W =f H*1 *f 1*w

[0020] The feature map z H*W Perform one-dimensional convolution Conv to increase the dimension, and then pass Sigmoid to obtain a new feature matrix S:

[0021] S = Sigmoid(Conv(z H*W ))

[0022] Multiply the feature matrix S and the original feature map z element by element to get the new feature matrix Z H*W :

[0023] Z H*W =S*z

[0024] For the feature matrix Z H*W Perform GAP and BN to obtain the feature vector:

[0025] f=BN(GAP(Z H*W ))

[0026] The eigenvectors f of the two modes RGB and f IR Perform cross-modal alignment loss:

[0027]

[0028] The final local loss is obtained by summing the three parts of the cross-modal alignment loss:

[0029] L local =L alignment1 +L alignment2 +L alignment3 .

[0030] As a further solution of the present invention: the global feature module in S3 includes a bilinear fully connected layer and a global loss function, the bilinear fully connected layer includes a fully connected layer with a bias term, a batch normalization layer, after the batch normalization layer there are ReLu and Dropout activation functions, and there is also a fully connected layer without a bias term.

[0031] As a further solution of the present invention: the global loss function includes an identity loss function, a triple loss function, and a center loss function;

[0032] The formula of identity loss function is:

[0033]

[0034] Where N is the total number of categories, y i is the true label, p i is the probability that the model predicts that the image belongs to the i-th category;

[0035] The formula for the triplet loss function is:

[0036] L triplet =max(d p -d n +α,0)

[0037] Among them, d p is the distance between positive sample pairs, d n is the distance between negative sample pairs, α is a hyperparameter indicating the allowed distance difference;

[0038] The formula of the center loss function is:

[0039]

[0040] Among them, y i is the label of the i-th image in the mini-batch, The yth yth represents the deep feature i The class center, B is the batch size;

[0041] So the global loss function is:

[0042] L=L ID +L triplet +β Lcenter .

[0043] A miner safety monitoring system based on cross-modal pedestrian re-identification is based on the above-mentioned miner safety monitoring method, including a shared dual-stream ResNet50 module, a global feature module and a local feature module, wherein the local feature module contains three independent SP_Nets, the global feature module contains a bilinear fully connected layer and a global loss function, and the global feature module and the local feature module perform joint loss training.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] The present invention proposes a bilinear fully connected layer, which can make the cross-modal ID embedding converge and improve the discriminability of cross-modal features; introduces a central loss function to solve the problem that the triple loss only considers the difference between the positive sample pair and the negative sample pair but ignores the absolute value of the positive and negative sample pairs; proposes three independent SP_Net modules, and uses strip global average pooling to expand the perception field of the convolutional neural network to collect richer contextual information, and can also capture the remote relationship of isolated areas, and further use the inner product and Sigmoid function, which helps to model the remote dependency relationship at the high-level semantic level. In general, the global feature module and local feature module proposed in the present invention can significantly improve the recognition ability and robustness of the model, eliminate the difference between the two modalities to extract more discriminative features, improve image recognition efficiency, and improve the performance of cross-modal pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 This is a flow chart of the underground mine safety monitoring method based on cross-modal pedestrian re-identification of the present invention.

[0047] Figure 2 This is the algorithm framework diagram of the present invention.

[0048] Figure 3 It is a flow chart of the underground mine safety monitoring system based on cross-modal pedestrian re-identification of the present invention.

[0049] Figure 4 It is a schematic diagram of the SP_Net (strip pooling network) module in the present invention.

[0050] Figure 5 It is a schematic diagram of the algorithm effect in the present invention. DETAILED DESCRIPTION

[0051] The present invention will be further described below by way of examples.

[0052] like Figure 1 As shown, a miner safety monitoring method based on cross-modal pedestrian re-identification includes the following steps:

[0053] S1: Input the dataset into the shared dual-stream ResNet50 backbone network to obtain the feature maps of the two modalities of modality-specific information and modality-shared information. Use the shared dual-stream ResNet50 backbone network to extract modality-specific information and modality-shared information, so that the network can extract more feature information from different modalities.

[0054] S2: Input the feature maps of the two modalities into the local feature module, and divide them into three parts of local feature maps according to their heights. Two of them are regarded as a pair, and the same pair of local feature maps is input into the same SP_Net module to obtain the local feature vector. The two local feature vectors are subjected to cross-modal alignment loss.

[0055] S3: Input the feature maps of the two modalities into the global feature module, obtain the global feature vector through the bilinear fully connected layer, and calculate the loss function;

[0056] S4: Adjust the network parameters to optimize the network and perform pedestrian matching on the cross-modal dataset to be identified.

[0057] like Figure 2 As shown in the figure, the two modal data VisibleImages (visible light pictures) and InfraredImages (infrared light pictures) in the dataset are input into a shareable two-stream ResNet50 backbone network, and FeatureMaps (feature maps) are output. The obtained FeatureMaps are divided into three equal parts by the split function, and the FeatureMaps after different modal divisions are combined in pairs and input into three independent SP_Net modules. Three Alignmentloss (alignment loss) functions are used to train the local feature module. In the global feature module, FeatureMaps are first subjected to GAP (global average pooling) to obtain a feature vector, and the feature vector is then passed through FC (fully connected layer) and BN (normalization layer), and two activation functions Relu and Dropout are used. Finally, a layer of FC (fully connected layer) is used to obtain the final feature vector, and Tripletloss (triplet loss) and Centerloss (center loss) and ID loss (identity loss) are used to train the global feature module.

[0058] Furthermore, the shared dual-stream ResNet50 backbone network in S1 includes five network layers, each of which is a residual structure block. The parameters of the first two network layers are independent and used to extract modality-specific features, and the parameters of the last three network layers are shared and used to extract modality-shared features. The feature maps of visible light and infrared are obtained through the output of the shared dual-stream ResNet50 backbone network.

[0059] Furthermore, the local feature module in S2 includes three independent SP_Net modules and three cross-modal alignment loss functions. The SP_Net module is used to obtain more discriminative features. The feature map output by S1 is evenly divided into three local feature maps according to its height through the split function. The corresponding local feature maps under different modalities are combined into local feature map pairs, which are input into the SP_Net module. The SP_Net module outputs two local feature vectors of different modalities, and the cross-modal alignment loss function is calculated for the local feature vectors of these two different modalities. Finally, three cross-modal alignment loss functions are obtained, which are added to obtain the final local loss.

[0060] Further, such as Figure 4 As shown in the figure, the SP_Net module uses strip global average pooling to expand the field of view of the convolutional neural network, extract features in different directions from the feature map, perform inner products on the feature vectors extracted in different directions to obtain the feature matrix, and calculate the feature matrix using the Sigmoid function. The calculated feature matrix and the original feature map are element-by-element multiplied, and the matrix after the product is subjected to global average pooling and batch normalization. Therefore, the SP_Net module mainly includes two strip global average pooling layers, one batch normalization layer, one one-dimensional convolution layer, and one Sigmoid layer. In addition, in order to accelerate the convergence of the model and reduce the amount of calculation of the model, we use a one-dimensional convolution layer to reduce the dimension of the feature map. The process is as follows:

[0061] The global average pooling of the strips along the width direction is recorded as GAP 1*W , the global average pooling in the height direction is recorded as GAP H*1 , global average pooling is denoted as GAP, one-dimensional convolution layer is denoted as Conv, batch normalization is denoted as BN, input feature map is denoted as z, feature vector is denoted as f, visible light is denoted as RGB, and infrared is denoted as IR;

[0062] Perform one-dimensional convolution on the feature map z to reduce the dimension, and perform strip global average pooling on the feature map z obtained by the convolution layer in two directions to obtain the feature vector f:

[0063] f H*1 =GAP 1*W (Conv(z))

[0064] f 1*W =GAP H*1 (Conv(z))

[0065] The obtained feature vector is inner-producted to obtain a new feature map z H*W :

[0066] z H*W =fH*1 *f 1*w

[0067] The feature map z H*W Perform one-dimensional convolution Conv to increase the dimension, and then pass Sigmoid to obtain a new feature matrix S:

[0068] S = Sigmoid(Conv(z H*W ))

[0069] Multiply the feature matrix S and the original feature map z element by element to get the new feature matrix Z H*W :

[0070] Z H*W =S*z

[0071] For the feature matrix Z H*W Perform GAP and BN to obtain the feature vector:

[0072] f=BN(GAP(Z H*W ))

[0073] The eigenvectors f of the two modes RGB and f IR Perform cross-modal alignment loss:

[0074]

[0075] Since our local feature module has three independent SP_Net networks, we will get three cross-modal alignment losses. The final local loss is obtained by summing the three parts of the cross-modal alignment losses:

[0076] L local =L alignment1 +L alignment2 +L alignment3 .

[0077] Furthermore, the global feature module in S3 includes a bilinear fully connected layer and a global loss function. In the global feature module, the feature map is globally averaged and pooled. The purpose of the bilinear fully connected layer is to solve the problem of low discriminative cross-modal feature representation due to the difficulty of gradient vanishing shared networks to converge with cross-modal ID embedding. The bilinear fully connected layer includes a fully connected layer with a bias term, a batch normalization layer, and after the batch normalization layer, there are ReLu and Dropout activation functions, and a fully connected layer without a bias term. The two functions mainly prevent the problem of overfitting. A fully connected layer without a bias term is followed, and finally the feature vector is output, and the obtained feature vector is subjected to global loss calculation.

[0078] Furthermore, the global loss function includes identity loss function, triplet loss function and center loss function. The center loss is introduced to solve the problem that triplet loss only considers the difference between positive sample pairs and negative sample pairs, but ignores the absolute value of positive sample pairs and negative sample pairs.

[0079] The formula of identity loss function is:

[0080]

[0081] Where N is the total number of categories, y i is the true label, p i is the probability that the model predicts that the image belongs to the i-th category;

[0082] The formula for the triplet loss function is:

[0083] L triplet =max(d p -d n +α,0)

[0084] Among them, d p is the distance between positive sample pairs, d n is the distance between negative sample pairs, α is a hyperparameter indicating the allowed distance difference;

[0085] The formula of the center loss function is:

[0086]

[0087] Among them, y i is the label of the i-th image in the mini-batch, The yth yth represents the deep feature i The class center, B is the batch size;

[0088] So the global loss function is:

[0089] L=L ID +L triplet +β Lcenter .

[0090] The global loss and local loss are trained together, and the parameters are tuned during the training process to make the model reach the optimal state. In order to prove the performance of the present invention, the algorithm is tested on the data set, and the results are as follows Figure 5 As shown, the algorithm can significantly improve image recognition capabilities.

[0091] like Figure 3As shown, a miner safety monitoring system based on cross-modal pedestrian re-identification is based on the above-mentioned miner safety monitoring method, including a shared dual-stream ResNet50 module, a global feature module and a local feature module, wherein the local feature module contains three independent SP_Net modules, the global feature module contains a bilinear fully connected layer and a global loss function, and the global feature module and the local feature module perform joint loss training.

Claims

1. A miner safety monitoring method based on cross-modal pedestrian re-identification, characterized in that: The following steps are involved: S1: Input the dataset into the shared two-stream ResNet50 backbone network to obtain the feature maps of the two modalities of modality-specific information and modality-shared information; S2: Input the feature maps of the two modalities into the local feature module, and divide them into three parts of local feature maps according to their heights. Two of them are regarded as a pair, and the same pair of local feature maps is input into the same SP_Net module to obtain the local feature vector. The two local feature vectors are subjected to cross-modal alignment loss. S3: Input the feature maps of the two modalities into the global feature module, obtain the global feature vector through the bilinear fully connected layer, and calculate the loss function; S4: Adjust the network parameters to optimize the network and perform pedestrian matching on the cross-modal dataset to be identified.

2. According to claim 1, a miner safety monitoring method based on cross-modal pedestrian re-identification is characterized in that: The shared dual-stream ResNet50 backbone network in S1 consists of five network layers, each of which is a residual structure block. The parameters of the first two network layers are independent and used to extract modality-specific features, and the parameters of the last three network layers are shared and used to extract modality-shared features.

3. According to the method of claim 1, the method is characterized in that: The local feature module in S2 includes three independent SP_Net modules and three cross-modal alignment loss functions. The feature map output by S1 is evenly divided into three local feature maps according to its height. The corresponding local feature maps under different modalities are combined into a local feature map pair. The local feature map pair is input into the SP_Net module. The SP_Net module outputs two local feature vectors of different modalities. The cross-modal alignment loss function is calculated for the local feature vectors of these two different modalities. Finally, three cross-modal alignment loss functions are obtained, which are added together to get the final local loss.

4. According to claim 3, a miner safety monitoring method based on cross-modal pedestrian re-identification is characterized in that: The SP_Net module uses strip global average pooling to expand the field of view of the convolutional neural network. The process is as follows: The global average pooling of the strips along the width direction is recorded as GAP 1*W , the global average pooling in the height direction is recorded as GAP H*1 , global average pooling is denoted as GAP, one-dimensional convolution layer is denoted as Conv, batch normalization is denoted as BN, input feature map is denoted as z, feature vector is denoted as f, visible light is denoted as RGB, and infrared is denoted as IR; Perform one-dimensional convolution on the feature map z to reduce the dimension, and perform strip global average pooling on the feature map z obtained by the convolution layer in two directions to obtain the feature vector f: f H*1 =GAP 1*W (Conv(z)) f 1*W =GAP H*1 (Conv(z)) The obtained feature vector is inner-producted to obtain a new feature map z H*W : z H*W =f H*1 *f 1*w The feature map z H*W Perform one-dimensional convolution Conv to increase the dimension, and then pass Sigmoid to obtain a new feature matrix S: S=Sigmoid(Conv(z H*W )) Multiply the feature matrix S and the original feature map z element by element to get the new feature matrix Z H*W : Z H*W =S*z For the feature matrix Z H*W Perform GAP and BN to obtain the feature vector: f=BN(GAP(Z H*W )) The eigenvectors f of the two modes RGB and f IR Perform cross-modal alignment loss: The final local loss is obtained by summing the three parts of the cross-modal alignment loss: L local =L alignment1 +L alignment2 +L alignment3 .

5. The method for safety monitoring of miners based on cross-modal pedestrian re-identification according to claim 1 is characterized in that: The global feature module in S3 includes a bilinear fully connected layer and a global loss function. The bilinear fully connected layer contains a fully connected layer with a bias term, a batch normalization layer, and after the batch normalization layer, there are ReLu and Dropout activation functions, and a fully connected layer without a bias term.

6. A miner safety monitoring method based on cross-modal pedestrian re-identification according to claim 5, characterized in that: The global loss function includes identity loss function, triple loss function, and center loss function; The formula of identity loss function is: Where N is the total number of categories, y i is the true label, p i is the probability that the model predicts that the image belongs to the i-th category; The formula for the triplet loss function is: L triplet =max(d p -d n +α,0) Among them, d p is the distance between positive sample pairs, d n is the distance between negative sample pairs, α is a hyperparameter indicating the allowed distance difference; The formula of the center loss function is: Among them, y i is the label of the i-th image in the mini-batch, The yth yth represents the deep feature i The class center, B is the batch size; So the global loss function is: L=L ID +L triplet +βL center 。 7. A miner safety monitoring system based on cross-modal pedestrian re-identification, characterized in that: The miner safety monitoring method based on any one of claims 1-6 includes a shared dual-stream ResNet50 module, a global feature module and a local feature module, wherein the local feature module includes three independent SP_Net modules, the global feature module includes a bilinear fully connected layer and a global loss function, and the global feature module and the local feature module perform joint loss training.

Citation Information

Patent Citations

  • Video pedestrian re-identification method based on channel attention mechanism and application

    CN112836646A

  • Pedestrian re-identification method based on global-local feature dynamic alignment

    CN113408492A

  • Near infrared-visible light cross-modal double-current pedestrian re-identification method and system

    CN114220124A

  • Visible light-infrared cross-modal pedestrian image retrieval method based on partition sensing network

    CN118015694A

  • Device and method for displaying fetal positions and fetal biological signals using portable technology

    US20140228653A1

Cited By

  • Cross-modal pedestrian re-identification method and system for deep ground emergency rescue

    CN121147979A

  • A cross-modal pedestrian re-identification method and system for deep earth emergency rescue

    CN121147979B