Power grid operation site face recognition method based on deep learning

Through the deep learning dynamic adversarial multimodal network DAM-Net, combined with multispectral image processing and feature fusion, the problems of low accuracy and poor adaptability of occluded face recognition at power grid operation sites are solved, and efficient occluded face identity recognition is achieved.

CN120635960AActive Publication Date: 2025-09-12DONGGUAN TRANSMISSION & TRANSFORMATION ENG CO
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510696878.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Traditional face recognition methods have a low recognition accuracy at power grid operation sites due to workers wearing masks, helmets and other obstructions. They also lack dynamic adaptability and single physiological feature judgment, making them unable to effectively identify obscured faces.

Method used

The deep learning-based dynamic adversarial multimodal network DAM-Net is adopted, combined with multispectral image processing and dynamic occlusion detection, adversarial feature generation, and multimodal identification. Through multispectral image alignment, occlusion generation and feature fusion, the identity recognition of occluded faces is realized.

Benefits of technology

The accuracy and robustness of occlusion detection are significantly improved, the integrity and authenticity of the repaired features are ensured, and the reliability and accuracy of occlusion feature recognition are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635960A_ABST
    Figure CN120635960A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power grid operation, and particularly relates to a power grid operation site face recognition method based on deep learning, which comprises the following specific steps: S1, collecting and preprocessing a face image pair of a power grid operator to obtain a multispectral image pair data set D1 of the power grid operator; s2, for any second face image pair P2 = lt; i ''1, I '2gt; belongs to D1, and a multispectral image pair data set D2 of the power grid operating personnel is obtained; s3, constructing a face recognition model dynamic adversarial multi-modal network DAM-Net under the condition that the operating personnel are shielded in the power grid site, wherein the face recognition model dynamic adversarial multi-modal network DAM-Net is used for identity recognition of the operating personnel in the power grid operating site under the condition that the face is shielded; and S4, training and updating parameters of each layer of the DAM-Net network constructed in the S3 to obtain a trained DAM-Net network. According to the DAM-Net, the deformable convolution fusion sub-module and the space-time attention mechanism sub-module are introduced through DODM, and the accuracy and robustness of occlusion detection can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power grid operations, and in particular to a face recognition method at a power grid operation site based on deep learning. Background Art

[0002] In recent years, facial recognition technology has been widely used in fields such as identity verification and security monitoring. However, in practice, workers in industries such as power grid operators and security personnel often need to wear protective gear such as masks, helmets, goggles, and scarves. This makes it difficult for traditional methods based on complete facial feature extraction to maintain high recognition accuracy.

[0003] Existing occluded face recognition algorithms mainly rely on generating and restoring facial information of the occluded part to fully recognize the facial features. However, existing methods have the following problems:

[0004] 1. Traditional methods use predetermined geometric occlusions (such as rectangles and ellipses), which cannot adapt to the contour characteristics of different target objects, resulting in excessive occlusion of the effective area or missed detection of irregular occlusions.

[0005] 2. Traditional methods lack dynamic adaptability. Existing mask parameters cannot be adjusted in real time according to input features. Mask positioning errors increase significantly in scenarios such as lighting changes and posture deviations.

[0006] 3. The discrimination mechanism is single and physiological characteristics are missing.

[0007] Based on the above, the present invention proposes a face recognition method for power grid operation site based on deep learning. Summary of the Invention

[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:

[0009] A method for face recognition at a power grid operation site based on deep learning includes the following specific steps:

[0010] S1: Collect and preprocess facial image pairs of power grid workers to obtain a multispectral image pair dataset D1 of power grid workers;

[0011] S2: For any second face image pair P2=<I″1,I'2> ∈D1 is processed to obtain the multispectral image pair dataset D2 of power grid workers;

[0012] S3: Build a dynamic adversarial multimodal network (DAM-Net) for face recognition in the presence of occluded faces at power grid sites. This model is used to identify workers at power grid sites when their faces are obscured.

[0013] S4: Train and update the parameters of each layer of the DAM-Net network built in S3 to obtain a trained DAM-Net network;

[0014] S5: Use the trained DAM-Net network to identify people with obscured faces in power grid operation scenes and output the identity label and confidence level of the target person.

[0015] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S1 are as follows:

[0016] S11: Use multispectral imaging equipment to collect facial images of workers at various power grid operation sites. For any worker, obtain the first facial image pair P1 =<I1,I2> , where I1 is the visible light image of the operator's face, and I2 is the infrared image of the operator's face;

[0017] S12: Perform spatial alignment and image normalization on I1 and I2 to obtain the second face image pair P2 = <I″1,I′2>;

[0018] S13: Pair all the second face images P2 =<I″1,I′2> A multispectral image pair dataset D1 of power grid workers is constructed.

[0019] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S12 are as follows:

[0020] S121: P1 spatially aligns I1 to I2 using an affine transformation matrix to obtain a spatially aligned visible light face image I'1, so that I'1 and I2 are spatially aligned. The calculation formula for I'1 is as follows:

[0021] I′1=warpAffine(I1,M)

[0022] Where warpAffine represents affine transformation; M is the affine transformation matrix, which is a 2×3 matrix in the form of Among them, a, b, c, d control image rotation, scaling, and shearing; c, f control translation;

[0023] S122: performing channel normalization on I'1 to obtain a normalized visible light face image I"1; performing linear normalization on I2 to obtain a normalized infrared face image I'2, thereby improving the accuracy of subsequent processing and analysis.

[0024] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S2 are as follows:

[0025] S21: Perform geometric occlusion on I″1 to obtain a visible light simulated occluded face image I m1 ; Semantic occlusion generation is performed on I″1 through 3D projection technology and GAN texture generation network to obtain visible light simulated occluded face image I m2 ; Perform adversarial occlusion generation based on the fast gradient sign method on I″1 to obtain the visible light simulated occluded face image I m3 ;

[0026] S22: Use the person’s name as a label to simulate the visible light occlusion of the face image I m1 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m1 ; Simulate visible light to block the face image I m2 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m2 ; Simulate visible light to block face image I m3 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m3 ; Use the person's name as a label, select and mark the face area of ​​​​I'2, and obtain the marked infrared face image y;

[0027] S23: By I' m1 and y construct the third sub-worker face image pair S1 = <I' m1 ,y>; by I' m2 and y construct the fourth sub-worker face image pair S2 = <I' m2 ,y>; by I' m3 and y construct the fifth sub-worker face image pair S3 = <I' m3 ,y>; Collect all S1, S2, and S3 to obtain a multispectral image pair dataset D2 of power grid workers.

[0028] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S3 are as follows:

[0029] S31: Design a dynamic occlusion detection module DODM, ​​which converts P3=<x,y> Input into DODM to obtain the face occlusion prediction heat map M _occ ;

[0030] S32: Construct an adversarial feature generator AFG, M _occand x are input into AFG to generate the face restoration feature map F face ;

[0031] S33: Construct a multimodal identification network MDN, which transforms x, M _occ and F face Input into MDN to obtain the result matrix result, which includes the final identity label and its corresponding confidence.

[0032] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S31 are as follows:

[0033] S311: Input x into the first three residual block groups of the ResNet18 encoder to obtain the simulated occluded face feature map F1; input y into the ViT encoder to obtain the infrared face feature map F2;

[0034] S312: Design a deformable convolution fusion submodule DCFS, input F1 and F2 into DCFS, and obtain the dimension of feature fusion The fusion feature map F4;

[0035] S313: Design a spatiotemporal attention mechanism submodule SAMS, input F4 into SAMS, and obtain a face occlusion prediction heat map M with a dimension of 1×H×W _occ ;

[0036] The specific steps of S312 are as follows:

[0037] S3121: F1 and F2 are concatenated through bimodal concatenation to obtain a dimension of The splicing feature map F3;

[0038] S3122: Input F3 to the offset prediction layer, perform a convolution operation with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, then input it to the ReLU layer and a convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, perform nonlinear activation operations and convolution operations, and obtain a dimension of The offset characteristic map △p;

[0039] S3123: Input F1, F2 and △p into the deformable convolution layer and upsampling layer, perform deformable convolution and upsampling operations, and obtain a dimension of The fusion feature map F4;

[0040] F4=Upsample(DCNv1(F1,F2,Δp))

[0041] Among them, DCNv1 represents the deformable convolution layer operation; Upsample represents the upsampling operation;

[0042] The specific steps of S313 are as follows:

[0043] S3131: Design a channel attention network, input F4 into the channel attention network, and obtain the channel-by-channel multiplication feature map F c ;

[0044] S3132: Design a spatial attention network to transform F c Input into the spatial attention network to get M _occ ;

[0045] The specific steps of S3131 are as follows:

[0046] S31311: Input F4 into the global average pooling layer and perform average pooling operation to obtain the dimension of Average pooling feature map F5; then input F5 into the shared multi-layer perceptron to obtain a dimension of The average weighted feature map F7;

[0047] S31312: Input F4 into the global maximum pooling layer and perform the maximum pooling operation to obtain a dimension of The maximum pooled feature map F6 is then input into the shared multi-layer perceptron to obtain the maximum weighted feature map F8;

[0048] S31313: Add F7 and F8 element by element, and the dimension is Pooling splicing feature map F9; then input F9 into Sigmoid and perform activation operation to obtain the dimension Pooled splicing activation feature map F 10 ;

[0049] F9=add(F7,F8)

[0050] F 10 =Sigmoid(F9)

[0051] Where add(F7,F8) represents the add operation, which is to add F7 and F8 element by element, and Sigmoid represents the nonlinear activation operation;

[0052] S31314: For F4 and F 10 Perform channel-by-channel multiplication to obtain a dimension of The channel-by-channel multiplication feature map F c ;

[0053]

[0054] in Represents a channel-by-channel multiplication operation;

[0055] The specific steps of S3132 are as follows:

[0056] S31321: F c Input to the channel average pooling layer, perform channel average pooling operation, and obtain the average pooling feature map F with a dimension of 1×H×W 11 ;

[0057] S31322: F C Input to the channel maximum pooling layer, perform the channel maximum pooling operation, and obtain the channel pooling feature map F with a dimension of 1×H×W 12 ;

[0058] S31323: F 11 and F 12 Perform the Concat operation to obtain a pooled splicing feature map F with a dimension of 2×H×W 13 ; Then F 13 Input to the convolution layer with a convolution kernel size of 7×7, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a spatial attention feature map F with a dimension of 1×H×W 14 ; Then F 14 Input into Sigmoid and perform activation operation to obtain M with dimension 1×H×W _occ ;

[0059] F 13 =Concat(F 11 ,F 12 )

[0060] F 14 =Conv 7×7 (F 13 )

[0061] M _occ =Sigmoid(F 14 )

[0062] Concat(F 11 ,F 12 ) represents F 11 and F 12 Perform the Concat operation.

[0063] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S32 are as follows:

[0064] S321: x and M_occ Perform the Concat splicing operation to obtain the face fusion feature map F with a dimension of (C+1)×H×W 15 ;

[0065] F 15 =Concat(x,M _occ )

[0066] Where Concat(x,M _occ ) represents the pair x and M _occ Perform Concat splicing operation;

[0067] S322: Construct a structure recovery channel submodule SRC to convert F 15 Input into SRC to obtain the structure recovery feature map F struct ;

[0068] S323: Construct a texture restoration channel submodule TRC, and convert M _occ And the randomly generated noise vector z is input into the texture recovery channel to obtain the texture recovery feature map F with a dimension of C×H×W texture ;

[0069] S324: Construct a dynamic region fusion submodule DRFM to transform F struct and F texture Input into the dynamic region fusion submodule DRFM to obtain F face ;

[0070] The specific steps of S322 are as follows:

[0071] S3221: F 15 Input to the ReLU layer and the convolution layer with a kernel size of 7×7, a stride of 2, and a zero padding of 3, the dimension is Face fusion enhanced feature map F 16 ;

[0072] F 16 =Conv 7×7 (ReLU(F 15 ))

[0073] S3222: F 16 Input into 4 series-connected hole residual blocks in sequence to obtain the fusion structure recovery feature map F 17 , wherein the structures of the first void residual block, the second void residual block, the third void residual block, and the fourth void residual block are the same;

[0074] S3223: F 17Input to the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, the ReLU layer, and the upsampling layer, perform convolution operations, activation operations, and upsampling operations with an upsampling factor of 2, and obtain an F with a dimension of C×H×W struct ;

[0075] F struct =Upsample(ReLU(Conv 3×3 (F 17 )))

[0076] Upsample represents the upsampling operation;

[0077] The specific steps of S323 are as follows:

[0078] S3231: Randomly generate a noise vector z and input z into the multilayer perceptron to obtain the noise feature vector f1;

[0079] S3232: M _occ Input to the global average pooling layer, perform average pooling operation, and obtain the face pooling scalar g; then input g into the multi-layer perceptron to obtain the face feature vector f2;

[0080] S3233: Concatenate f1 and f2 to obtain the noise face concatenation feature vector f3; input f3 into the multi-layer perceptron MLP to obtain the style feature vector w;

[0081] S3234: W and M _occ Perform AdaIN modulation operation to obtain F texture ;

[0082] The specific steps of S324 are as follows:

[0083] S3241: F struct Input to the global average pooling layer and the convolution layer with a convolution kernel size of 1×1, a stride of 1, and zero padding of 0, perform global average pooling and convolution operations, and obtain a structural weight vector α with a dimension of C×1×1;

[0084] α=Conv 1×1 (GAP(F struct ))

[0085] S3242: F texture Input it into the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a texture weight feature map β with a dimension of 1×H×W;

[0086] β=Conv 3×3 (F texture )

[0087] S3243: F struct 、F texture , α, β perform weighted fusion operation to obtain F with dimensions of C×H×W face ;

[0088] F face =α·F struct +(1-α)·(β⊙F texture )

[0089] Among them, ⊙ represents the product operation; · represents the dot product; + represents the vector addition.

[0090] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S33 are as follows:

[0091] S331: Construct an identity authentication submodule to combine x and F face Input to the identity identification submodule and output the identity similarity score vector S id ;

[0092] S332: Construct a repair identification submodule to face and M _occ Input to the repair identification submodule and output the authenticity score vector D real ;

[0093] S333: S id and D real Perform a decision fusion operation and then a max operation to obtain the result matrix result, where result is a two-dimensional matrix containing the final identification identity label and its corresponding confidence;

[0094] result=max(W id ·S id +W real ·D real )

[0095] Where W id Represents the weight of the identity authentication branch; W real Represents the weight of generating the identification branch, and max represents the maximum value and its index;

[0096] The specific steps of S331 are as follows:

[0097] S3311: F face Perform feature extraction based on the ArcFace algorithm to obtain the repaired face extraction feature vector f gen ;

[0098] S3312: Perform feature extraction on x based on the ArcFace algorithm to obtain the occluded face extraction feature vector f real ;

[0099] S3313: f gen and f real Perform cosine similarity calculation to obtain the identity similarity score vector S id ;

[0100] S id =1-cos({f gen |f real})

[0101] Its cos is a cosine similarity calculation function;

[0102] The specific steps of S332 are as follows:

[0103] S3321: F face Input into the PatchGAN encoder and get the dimension The repaired face image identification feature map F PGAN ;

[0104] S3322: M _occ Input into the spatial attention layer, the dimension is The spatial attention weighted feature map F space ;

[0105] S3323: F PGAN and F space Perform the Concat operation to obtain the dimension Multi-scale feature map F mul ; Then F mul Input it into the convolution layer with a convolution kernel size of 1×1, a step size of 1, and a zero padding of 1, and perform a convolution operation to obtain the authenticity score D real ;

[0106] F mul =Concat(F PGAN ,F space )

[0107] D real =conv 1×1 (F mul )

[0108] Concat(F PGAN ,F space ) represents the Concat operation.

[0109] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S4 are as follows:

[0110] S41: Initialize the hyperparameters required for training DAM-Net;

[0111] S42: Divide the multispectral image pair dataset of power grid workers constructed in S2 into a training set and a test set according to a certain ratio, ensuring that the two do not overlap; then divide the training set into multiple batches, input one batch of training set into the network for training each time, and calculate the loss value of the batch;

[0112] S43: After traversing all batches of the entire training set in each round, the performance of the model is evaluated in stages using the test set data; the test set is fed into the model in sequence according to the same batch size as the training stage, the prediction error of each batch is calculated respectively, and the test set loss value of the current round is summarized; the test set loss value can be used to dynamically track the performance changes of the network on non-training data; by continuously observing the changing trend of the test set loss value, it is determined whether the model shows signs of overfitting; when the test set loss value no longer decreases in multiple consecutive training cycles, or shows an upward trend, it can be determined that the generalization ability of the model has decreased, and the preset strategy adjustment mechanism is triggered; when the test set loss value tends to stabilize with the training process and meets the expected performance standard, the network training can be considered completed, and the final network parameters will be used for the identity recognition of workers in subsequent power grid operations when their faces are obscured.

[0113] As a preferred solution of the method for face recognition at a power grid operation site based on deep learning described in the present invention, the specific steps of S5 are as follows:

[0114] S51: After the DAM-Net network training is completed, it is applied to the face recognition task under human occlusion in actual power grid operation scenarios;

[0115] S52: The collected identity images of power grid workers are input into the trained DAM-Net network. The model identifies each power grid worker in the image one by one and outputs the corresponding identity label and confidence level of each worker in the image.

[0116] Compared with existing technologies:

[0117] The DAM-Net of the present invention introduces the deformable convolution fusion submodule and the spatiotemporal attention mechanism submodule through DODM, ​​which can significantly improve the accuracy and robustness of occlusion detection; in addition, through AFG, the structure recovery channel submodule and the texture recovery channel submodule are adopted in combination with the dynamic area fusion submodule, which can ensure the integrity and authenticity of the repaired features; at the same time, the dual mechanisms of the identity identification submodule and the repair identification submodule of MDN are combined with decision fusion, which effectively improves the reliability and accuracy of occlusion feature recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0118] Figure 1 It is a schematic diagram of the process of the present invention;

[0119] Figure 2 This is the structural diagram of the face recognition network DAM-Net under occlusion conditions of the present invention;

[0120] Figure 3 This is a structural diagram of the DODM module of the present invention;

[0121] Figure 4 This is the structural diagram of the deformable convolution fusion submodule DCFS of the present invention;

[0122] Figure 5 This is a structural diagram of the offset prediction layer of the present invention;

[0123] Figure 6 This is the structure diagram of the spatiotemporal attention mechanism submodule SAMS of the present invention;

[0124] Figure 7 This is the channel attention network structure diagram of the present invention;

[0125] Figure 8 This is the structure diagram of the spatial attention network of the present invention;

[0126] Figure 9 This is the structure diagram of the adversarial feature generator AFG of the present invention;

[0127] Figure 10 This is the structural diagram of the channel submodule SRC of the present invention;

[0128] Figure 11 This is a structural diagram of the texture recovery channel submodule TRC of the present invention;

[0129] Figure 12 This is a structural diagram of the dynamic region fusion submodule DRFM of the present invention;

[0130] Figure 13 This is the MDN module structure diagram of the present invention;

[0131] Figure 14 This is a structural diagram of the identity authentication submodule of the present invention;

[0132] Figure 15 This is the structural diagram of the repair identification submodule of the present invention. DETAILED DESCRIPTION

[0133] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0134] The present invention provides a method for face recognition on power grid operation site based on deep learning. Figures 1-15 , including the following specific steps:

[0135] S1: Collect and preprocess facial image pairs of power grid workers to obtain a multispectral image pair dataset D1 of power grid workers;

[0136] The specific steps of S1 are as follows:

[0137] S11: Use multispectral imaging equipment to collect facial images of workers at various power grid operation sites. For any worker, obtain the first facial image pair P1 =<I1,I2> , where I1 is the visible light image of the operator's face, and I2 is the infrared image of the operator's face;

[0138] S12: Perform spatial alignment and image normalization on I1 and I2 to obtain the second face image pair P2 = <I″1,I′2>;

[0139] The specific steps of S12 are as follows:

[0140] S121: P1 spatially aligns I1 to I2 using an affine transformation matrix to obtain a spatially aligned visible light face image I'1, so that I'1 and I2 are spatially aligned. The calculation formula for I'1 is as follows:

[0141] I′1=warpAffine(I1,M)

[0142] Where warpAffine represents affine transformation; M is the affine transformation matrix, which is a 2×3 matrix in the form of Among them, a, b, c, d control image rotation, scaling, and shearing; c, f control translation;

[0143] Matching refers to aligning the visible light image and the infrared image of the face in a unified spatial coordinate system to ensure that the position, shape, and proportion of the same scene or object in the two images are consistent;

[0144] Among them, the affine transformation matrix is ​​a matrix used to represent the linear transformation and displacement of an image or spatial object in a two-dimensional or three-dimensional space;

[0145] S122: performing a channel normalization operation on I'1 to obtain a normalized visible light face image I"1; performing a linear normalization operation on I2 to obtain a normalized infrared face image I'2, thereby improving the accuracy of subsequent processing and analysis;

[0146] Among them: Channel normalization refers to standardizing each channel, which helps improve training stability and accelerate convergence;

[0147]

[0148] μ=Apool(I′1)

[0149]

[0150] where X' i represents the result after normalization of the i-th pixel value; X i Represents the i-th pixel value; μ is the result of global average pooling of the image, that is, the average value of all pixel values, Apool represents global average pooling; σ is the standard deviation of the image, that is, the standard deviation of all pixel values; N represents the total number of pixels in the image;

[0151] Linear normalization refers to a common data normalization method that scales the data to a fixed range (usually [0,1]);

[0152]

[0153] where X' I2 Represents X I2 The result after linear normalization, that is, any pixel value of I'2; X I2 Represents any pixel value in I2; represents the minimum pixel value in I2, Represents the maximum pixel value in I2;

[0154] S13: Pair all the second face images P2 =<I″1,I'2> Construct a multispectral image pair dataset D1 of power grid workers;

[0155] S2: For any second face image pair P2=<I″1,I'2> ∈D1 is processed to obtain the multispectral image pair dataset D2 of power grid workers;

[0156] The specific steps of S2 are as follows:

[0157] S21: Perform geometric occlusion on I″1 to obtain a visible light simulated occluded face image I m1; Semantic occlusion generation is performed on I″1 through 3D projection technology and GAN texture generation network to obtain visible light simulated occluded face image I m2 ; Perform adversarial occlusion generation based on the Fast Gradient-Sign Method FGSM on I″1 to obtain the visible light simulated occluded face image I m3 ;

[0158] Among them: Geometric occlusion generation usually refers to using simple geometric shapes (such as rectangles, circles, polygons, etc.) to generate occlusions on an image; it is mainly used to simulate visual occlusions caused by opaque objects and is usually used to enhance data diversity;

[0159] 3D projection technology: Using 3D modeling tools (such as Blender) to project real objects (such as masks and glasses) onto facial images to create semantic occlusions. This method can accurately simulate the effect of occlusions on facial images based on the geometric shape and texture of the objects.

[0160] GAN texture generation network: Use a cycle-consistent generative adversarial network (CycleGAN) or similar generative adversarial networks to fuse the texture of the occluded area with the background image to ensure that the lighting and texture of the occluded object in the image are consistent with the actual situation;

[0161] The FGSM adversarial attack algorithm calculates the gradient information of the image, generates a small perturbation, and adds it to the original image. Specifically, the FGSM algorithm calculates the gradient of the image according to the loss function to obtain the sensitivity (gradient) of each pixel in the image to the loss function. Then, by symbolically processing this gradient information, a perturbation is generated and added to the original image to generate an adversarial sample. This perturbation has a disruptive effect, which can cause the deep learning model to incorrectly identify occluded areas, but at the same time has little impact on the overall visual effect of the image.

[0162] S22: Use the person’s name as a label to simulate the visible light occlusion of the face image I m1 The face area is framed and labeled to obtain the labeled visible light simulated occluded face image I′ m1 ; Simulate visible light to block face image I m2 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m2 ; Simulate visible light to block face image I m3 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m3 ; Use the person's name as a label, select and mark the face area of ​​​​I'2, and obtain the marked infrared face image y;

[0163] S23: From I' m1 and y, construct the third sub-operator face image pair S1 = <I' m1 , y>; from I' m2 and y, construct the fourth sub-operator face image pair S2 = <I' m2 , y>; from I' m3 and y, construct the fifth sub-operator face image pair S3 = <I' m3 , y>; collect all S1, S2, S3 to obtain a dataset D2 of multi-spectral image pairs of grid operators;

[0164] S3: Construct a dynamic adversarial multi-modal network DAM-Net (Dynamic Adversarial Multimodal Net) for face recognition in the case of operator occlusion in the grid field, which is used for identity recognition of operators in the grid operation field under face occlusion;

[0165] The structure of the dynamic adversarial multi-modal network DAM-Net constructed by the present invention is as Figure 2 shown;

[0166] For any operator face image pair P3 = <x, y> ∈ D2, where the visible light simulated occluded face image x refers to I' m1 , I' m2 or I' m3 ; input P3 = <x, y> into the dynamic occlusion detection module DODM (Dynamic Occlusion Detection Module) to obtain the face occlusion prediction heat map M _occ ; then input x and M _occ into the adversarial feature generator AFG (Adversarial Feature Generator) to obtain the face repair feature map F face ; then input x, M _occ and F face into the multi-modal discrimination network MDN (Multimodal Detection Network) to obtain the result matrix result, and result includes the final identity label and its corresponding confidence;

[0167] The specific steps of S3 are as follows:

[0168] S31: Design a dynamic occlusion detection module DODM (Dynamic Occlusion Detection Module), input P3 = <x, y> into DODM to obtain the face occlusion prediction heat map M _occ ;

[0169] The overall structure of the dynamic occlusion detection module DODM is as follows: Figure 3 As shown, the execution process is:

[0170] Input x with the dimension of C×H×W into the first three residual blocks of ResNet18 encoder, and get the dimension of The simulated occluded face feature map F1 is input into the ViT encoder with a dimension of 1×H×W and the labeled infrared face image y is obtained. The infrared face feature map F2 is obtained; then F1 and F2 are input into the deformable convolution fusion submodule DCFS to obtain a dimension of The fusion feature map F4 is then input into the spatiotemporal attention mechanism submodule SAMS to obtain the face occlusion prediction heat map M with a dimension of 1×H×W. _occ ;

[0171] The ResNet18 encoder processes images through multiple convolutional layers and residual blocks. The image input to the ResNet18 encoder first undergoes low-level feature extraction through a convolutional layer, then undergoes deep feature extraction through a series of residual blocks, and finally generates a high-level feature representation of the image through a fully connected layer. In the present invention, the first three residual blocks are mainly used to hierarchically extract shallow and mid-level semantic features of the image, while gradually compressing the spatial scale and increasing the number of feature channels through layer-by-layer downsampling, providing a foundation for subsequent high-level semantic modeling.

[0172] The ViT (Vision Transformer) encoder does not rely on convolution operations, but instead uses a self-attention mechanism to process images; compared with convolutional neural networks (CNNs), it can capture richer global information;

[0173] Infrared images have only one channel (single channel), that is, each pixel contains only one value (temperature or infrared radiation intensity);

[0174] The specific steps of S31 are as follows:

[0175] S311: Input x into the first three residual block groups of the ResNet18 encoder to obtain the simulated occluded face feature map F1; input y into the ViT encoder to obtain the infrared face feature map F2;

[0176] S312: Design a deformable convolution fusion submodule DCFS (Deformable convolution fusion submodule), input F1 and F2 into DCFS, and obtain the dimension of feature fusion The fusion feature map F4; the network structure of the deformable convolution fusion submodule DCFS designed by the present invention is as follows Figure 4 As shown;

[0177] The specific steps of S312 are as follows:

[0178] S3121: F1 and F2 are concatenated through bimodal concatenation to obtain a dimension of The splicing feature map F3;

[0179] Among them, bimodal stitching is to fuse data from two different modalities (such as visible light and infrared images); the stitching operation is to connect the feature maps of the two modalities along a certain dimension (usually the channel dimension) so that the model can learn the features of the two modalities simultaneously in a higher-dimensional space;

[0180] S3122: Input F3 to the offset prediction layer, perform a convolution operation with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, then input it to the ReLU layer and a convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, perform nonlinear activation operations and convolution operations, and obtain a dimension of The offset feature map △p; the structure of the offset prediction layer is as follows Figure 5 As shown;

[0181] S3123: Input F1, F2 and △p into the deformable convolution layer and upsampling layer, perform deformable convolution and upsampling operations, and obtain a dimension of The fusion feature map F4;

[0182] F4=Upsample(DCNv1(F1,F2,Δp))

[0183] Among them, DCNv1 represents the deformable convolution layer operation; Upsample represents the upsampling operation;

[0184] Among them, Deformable Convolution (DConv) is a technology designed to enhance the spatial adaptability of convolutional neural networks;

[0185] S313: Design a spatiotemporal attention mechanism submodule SAMS (Spatiotemporal Attention Mechanism Submodule), input F4 into SAMS, and obtain a face occlusion prediction heat map M with a dimension of 1×H×W _occ The network structure of the spatiotemporal attention mechanism submodule SAMS designed in this invention is as follows Figure 6 As shown;

[0186] The specific steps of S313 are as follows:

[0187] S3131: Design a channel attention network, input F4 into the channel attention network, and obtain the channel-by-channel multiplication feature map F c ;

[0188] The specific steps of S3131 are as follows:

[0189] S31311: Input F4 into the global average pooling layer and perform average pooling operation to obtain the dimension of Average pooling feature map F5; then input F5 into the shared multi-layer perceptron to obtain a dimension of The average weighted feature map F7;

[0190] Among them, shared multi-layer perceptron means applying the same MLP parameters to each element when processing multiple input elements (such as each point in a point cloud); this parameter sharing mechanism ensures that the model is insensitive to the order of the input elements;

[0191] S31312: Input F4 into the global maximum pooling layer and perform the maximum pooling operation to obtain a dimension of The maximum pooled feature map F6 is then input into the shared multi-layer perceptron to obtain the maximum weighted feature map F8;

[0192] S31313: Perform element-by-element addition operation on F7 and F8 (i.e. Figure 7 The add operation in ), the dimension is Pooling splicing feature map F9; then input F9 into Sigmoid and perform activation operation to obtain the dimension Pooled splicing activation feature map F 10 ;

[0193] F9=add(F7,F8)

[0194] F 10 =Sigmoid(F9)

[0195] Where add(F7,F8) represents the add operation, which is to add F7 and F8 element by element, and Sigmoid represents the nonlinear activation operation;

[0196] S31314: For F4 and F 10 Perform channel-by-channel multiplication to obtain a dimension of The channel-by-channel multiplication feature map F c ;

[0197]

[0198] in Represents a channel-by-channel multiplication operation;

[0199] Example:

[0200] Assume that the size of the fused feature map input to the channel attention network is C×H×W, where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map;

[0201] The fusion feature map F4 has a size of 256×224×224, which is input into the global average pooling layer to obtain the average pooling feature map F5 with a size of 256×1×1; F5 is input into the shared MLP to obtain the average weighted feature map F7 with a size of 256×1×1;

[0202] Input F4 to the global maximum pooling layer for maximum pooling operation to obtain the maximum pooling feature map F6 with a size of 256×1×1; then input F6 to the shared MLP to obtain the maximum weighted feature map F8 with a size of 256×1×1;

[0203] Add F7 and F8 to obtain a pooled splicing feature map F9 with a size of 256×1×1; then input F9 into Sigmoid and perform activation operation to obtain a pooled splicing activation feature map F with a size of 256×1×1. 10 ;

[0204] For F4 and F 10 Perform channel-by-channel multiplication operation to obtain a channel-by-channel multiplication feature map F with a size of 256×224×224 c ;

[0205] S3132: Design a spatial attention network to transform F c Input into the spatial attention network to get M _occ The overall structure of a spatial attention network designed by the present invention is as follows: Figure 8 As shown;

[0206] The specific steps of S3132 are as follows:

[0207] S31321: F c Input to the channel average pooling layer, perform channel average pooling operation, and obtain the average pooling feature map F with a dimension of 1×H×W 11 ;

[0208] S31322: F C Input to the channel maximum pooling layer, perform the channel maximum pooling operation, and obtain the channel pooling feature map F with a dimension of 1×H×W 12 ;

[0209] S31323: F 11and F 12 Perform the Concat operation to obtain a pooled splicing feature map F with a dimension of 2×H×W 13 ; Then F 13 Input to the convolution layer with a convolution kernel size of 7×7, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a spatial attention feature map F with a dimension of 1×H×W 14 ; Then F 14 Input into Sigmoid and perform activation operation to obtain M with dimension 1×H×W _occ ;

[0210] F 13 =Concat(F 11 ,F 12 )

[0211] F 14 =Conv 7×7 (F 13 )

[0212] M _occ =Sigmoid(F 14 )

[0213] Concat(F 11 ,F 12 ) represents F 11 and F 12 Perform Concat splicing operation;

[0214] Example:

[0215] Input channel multiplication feature map F of spatial attention network c The size is 256×224×224, which is input to the channel average pooling layer to obtain the average pooling feature map F with a size of 1×224×224 11 ; F c Input to the channel maximum pooling layer to obtain the average pooling feature map F with a size of 1×224×224 12 ;

[0216] F 11 and F 12 Perform the Concat operation to obtain a pooled splicing feature map F with a size of 2×224×224 13 ; Then F 13 Input to the convolution layer with a convolution kernel size of 7×7, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a spatial attention feature map F with a size of 1×224×224 14 ; Then F 14Input into Sigmoid and perform activation operation to obtain the face occlusion prediction heat map M with a size of 1×224×224 _occ ;

[0217] S32: Construct an adversarial feature generator AFG (Adversarial Feature Generator) to convert M _occ and x are input into AFG to generate the face restoration feature map F face ;

[0218] The specific steps of S32 are as follows:

[0219] S321: x and M _occ Perform the Concat splicing operation to obtain the face fusion feature map F with a dimension of (C+1)×H×W 15 ;

[0220] F 15 =Concat(x,M _occ )

[0221] Where Concat(x,M _occ ) represents the pair x and M _occ Perform Concat splicing operation;

[0222] Among them, the Concat splicing operation: This is a way to increase the number of channels, which means that the number of channels of the image itself has increased, but the information under each feature has not increased, mainly realizing the superposition in the horizontal or vertical space;

[0223] S322: Construct a structural recovery channel submodule SRC (Structural Recovery Channel) to convert F 15 Input into SRC to obtain the structure recovery feature map F struct The present invention designs a structural recovery channel submodule SRC (Structural Recovery Channel), whose structure is as follows Figure 10 As shown;

[0224] The specific steps of S322 are as follows:

[0225] S3221: F 15 Input to the ReLU layer and the convolution layer with a kernel size of 7×7, a stride of 2, and a zero padding of 3, the dimension is Face fusion enhanced feature map F 16 ;

[0226] F 16 =Conv 7×7 (ReLU(F15 ))

[0227] S3222: F 16 Input into 4 series-connected hole residual blocks in sequence to obtain the fusion structure recovery feature map F 17 , wherein the structures of the first void residual block, the second void residual block, the third void residual block, and the fourth void residual block are the same;

[0228] Among them, the Dilated Residual Block (DRB) is a deep learning module that combines dilated convolution and residual connection, which is a publicly available technology.

[0229] S3223: F 17 Input to the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, the ReLU layer, and the upsampling layer, perform convolution operations, activation operations, and upsampling operations with an upsampling factor of 2, and obtain an F with a dimension of C×H×W struct ;

[0230] F struct =Upsample(ReLU(Conv 3×3 (F 17 )))

[0231] Upsample represents the upsampling operation;

[0232] Example:

[0233] Assume that the size of the face fusion feature map input to the structure recovery channel submodule SRC is C×H×W, where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map;

[0234] Face fusion feature map F 15 The size is 4×224×224, which is input into the ReLU layer and the convolution layer with a convolution kernel size of 7×7, a stride of 2, and a zero padding of 3 to obtain the face fusion enhanced feature map F with a size of 192×112×112 16 ;

[0235] F 16 Input into 4 hole residual blocks in sequence to obtain the fusion structure recovery feature map F with a size of 192×112×112 17 ;

[0236] The feature map F 17Input to the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, the ReLU layer, and the upsampling layer, perform convolution operations, activation operations, and upsampling operations with an upsampling factor of 2, and obtain a structure recovery feature map F with a size of 3×224×224 struct ;

[0237] S323: Construct a texture restoration channel submodule TRC (Texture Restoration Channel) to _occ And the randomly generated noise vector z is input into the texture recovery channel to obtain the texture recovery feature map F with a dimension of C×H×W texture The present invention designs a texture restoration channel submodule TRC, whose structure is as follows Figure 11 As shown;

[0238] The specific steps of S323 are as follows:

[0239] S3231: Randomly generate a noise vector z and input z into the multilayer perceptron to obtain the noise feature vector f1;

[0240] Among them, Multilayer Perceptron (MLP) is a classic feedforward artificial neural network model, which is widely used in tasks such as classification, regression, and feature extraction.

[0241] S3232: M _occ Input to the global average pooling layer, perform average pooling operation, and obtain the face pooling scalar g; then input g into the multi-layer perceptron to obtain the face feature vector f2;

[0242] S3233: Concatenate f1 and f2 to obtain the noise face concatenation feature vector f3; input f3 into the multi-layer perceptron MLP to obtain the style feature vector w;

[0243] S3234: W and M _occ Perform AdaIN modulation operation to obtain F texture ;

[0244] Among them, AdaIN modulation (Adaptive Instance Normalization) is a feature normalization operation used for style transfer and generative adversarial networks;

[0245] Example: Assume that the M input to the texture recovery channel submodule TRC _occ The size is C×H×W, where C is the number of channels, H is the height, and W is the width;

[0246] Randomly generate a 512-dimensional noise vector z; input z into the multi-layer perceptron to obtain a 256-dimensional noise feature vector f1;

[0247] The size of M is 1×224×224 _occ Input to the global average pooling layer, perform average pooling operation, and obtain the face pooling scalar g; then input g into the multi-layer perceptron to obtain a 256-dimensional face feature vector f2;

[0248] Concatenate f1 and f2 to obtain a 512-dimensional noise face feature vector f3; input f3 into the multi-layer perceptron MLP to obtain a 512-dimensional style feature vector w;

[0249] W and M _occ Perform AdaIN modulation operation to obtain an F with a size of 3×224×224 texture ;

[0250] S324: Construct a dynamic region fusion submodule DRFM to transform F struct and F texture Input into the dynamic region fusion submodule DRFM to obtain F face The present invention designs a dynamic region fusion submodule DRFM; uses DRFM to repair the face occlusion area and obtains F face ; The structure of DRFM is as follows Figure 12 As shown:

[0251] F struct Input to the global average pooling layer and 1×1 convolution layer, perform global average pooling and convolution operations to obtain the structural weight vector α; F texture Input it into the 3×3 convolution layer, perform convolution operation, and obtain the texture weight feature map β; then F struct 、F texture , α and β are input into weighted fusion to obtain F face ;

[0252] The specific steps of S324 are as follows:

[0253] S3241: F struct Input to the global average pooling layer and the convolution layer with a convolution kernel size of 1×1, a stride of 1, and zero padding of 0, perform global average pooling and convolution operations, and obtain a structural weight vector α with a dimension of C×1×1;

[0254] α=Conv 1×1 (GAP(F struct ))

[0255] S3242: F textureInput it into the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a texture weight feature map β with a dimension of 1×H×W;

[0256] β=Conv 3×3 (F texture )

[0257] S3243: F struct 、F texture , α, β perform weighted fusion operation to obtain F with dimensions of C×H×W face ;

[0258] F face =α·F struct +(1-α)·(β⊙F texture )

[0259] Among them, ⊙ represents the product operation; · represents the dot product; + represents the vector addition;

[0260] S33: Construct a multimodal detection network MDN (Multimodal Detection Network) to transform x, M _occ and F face Input into MDN to obtain the result matrix result, which includes the final identity label and its corresponding confidence;

[0261] The present invention constructs a multimodal identification network MDN as follows: Figure 13 As shown; the execution process is:

[0262] x and F face Input into the identity identification submodule to obtain the identity similarity score vector S id ; F face Input into the repair identification submodule to obtain the authenticity score vector D real ; S id 、D real Input into the decision fusion to obtain the result matrix result, which includes the final identity label and its corresponding confidence;

[0263] The specific steps of S33 are as follows:

[0264] S331: Construct an identity authentication submodule to combine x and F face Input to the identity authentication submodule and output S id The present invention designs an identity authentication submodule, whose structure is as follows Figure 14 As shown;

[0265] The specific steps of S331 are as follows:

[0266] S3311: F face Perform feature extraction based on the ArcFace algorithm to obtain the repaired face extraction feature vector f gen ;

[0267] ArcFace is a face recognition algorithm based on deep learning, which is mainly used for facial feature extraction and matching. It extracts facial feature vectors through an optimized deep convolutional neural network (CNN) and then uses Euclidean distance or cosine similarity for matching.

[0268] S3312: Perform feature extraction on x based on the ArcFace algorithm to obtain the occluded face extraction feature vector f real ;

[0269] S3313: f gen and f real Perform cosine similarity calculation to obtain the identity similarity score vector S id ;

[0270] S id =1-cos({f gen |f real})

[0271] Its cos is a cosine similarity calculation function;

[0272] S332: Construct a repair identification submodule to face and M _occ Input to the repair identification submodule and output the authenticity score vector D real The present invention designs a repair identification submodule, whose structure is as follows Figure 15 As shown;

[0273] The specific steps of S332 are as follows:

[0274] S3321: F face Input into the PatchGAN encoder and get the dimension The repaired face image identification feature map F PGAN ;

[0275] Among them, the PatchGAN encoder is a publicly available neural network structure for image processing tasks, which is used to encode images to extract features;

[0276] S3322: M _occ Input into the spatial attention layer, the dimension is The spatial attention weighted feature map F space ;

[0277] Among them, the Spatial Attention mechanism (SA) enhances the feature expression of key areas and suppresses irrelevant information by calculating the spatial dimension weights of the feature map;

[0278] S3323: F PGAN and F space Perform the Concat operation to obtain the dimension Multi-scale feature map F mul ; Then F mul Input it into the convolution layer with a convolution kernel size of 1×1, a step size of 1, and a zero padding of 1, and perform a convolution operation to obtain the authenticity score D real ;

[0279] F mul =Concat(F PGAN ,F space )

[0280] D real =conv 1×1 (F mul )

[0281] Concat(F PGAN ,F space ) represents the Concat splicing operation;

[0282] S333: S id and D real Perform decision fusion operation and then perform max operation (such as Figure 13 As shown), the result matrix result is obtained, where result is a two-dimensional matrix containing the final identification identity label and its corresponding confidence;

[0283] result=max(W id ·S id +W real ·D real )

[0284] Where W id Represents the weight of the identity authentication branch; W real Represents the weight of generating the identification branch, and max represents the maximum value and its index;

[0285] S4: Train and update the parameters of each layer of the DAM-Net network built in S3 to obtain a trained DAM-Net network;

[0286] The specific steps of S4 are as follows:

[0287] S41: Initialize the hyperparameters required for training DAM-Net, such as the batch size, initial learning rate, weight decay coefficient, and training rounds.

[0288] S42: Divide the multispectral image pair dataset of power grid workers constructed in S2 into a training set and a test set according to a certain ratio, ensuring that the two do not overlap; then divide the training set into multiple batches, input one batch of training set into the network for training each time, and calculate the loss value of the batch;

[0289] Example: In the experimental scheme of this aspect, the ratio of the training set to the test set is 7:3;

[0290] S43: After traversing all batches of the entire training set in each round, the performance of the model is evaluated in stages using the test set data; the test set is fed into the model in sequence according to the same batch size as the training stage, the prediction error of each batch is calculated respectively, and the test set loss value of the current round is summarized; the test set loss value can be used to dynamically track the performance changes of the network on non-training data; by continuously observing the changing trend of the test set loss value, it is determined whether the model shows signs of overfitting; when the test set loss value no longer decreases in multiple consecutive training cycles, or shows an upward trend, it can be determined that the generalization ability of the model has decreased, and then the preset strategy adjustment mechanism is triggered; for example, the training process is stopped or the learning rate is adaptively reduced to optimize the model convergence path; when the test set loss value tends to stabilize with the training process and meets the expected performance standard, the network training can be considered completed, and the final network parameters will be used for the subsequent identification of workers in the power grid operation site under face occlusion;

[0291] S5: Use the trained DAM-Net network to identify people with obscured faces in power grid operation scenes and output the target person's identity label and confidence score;

[0292] The specific steps of S5 are as follows:

[0293] S51: After the DAM-Net network training is completed, it is applied to the face recognition task under human occlusion in actual power grid operation scenarios;

[0294] S52: The collected identity images of power grid workers are input into the trained DAM-Net network. The model identifies each power grid worker in the image one by one and outputs the corresponding identity label and confidence level of each worker in the image.

[0295] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A face recognition method for power grid operation site based on deep learning, characterized in that: The specific steps are as follows: S1: Collect and preprocess facial image pairs of power grid workers to obtain a multispectral image pair dataset D1 of power grid workers; S2: For any second face image pair P2=<I”1,I'2> ∈D1 is processed to obtain the multispectral image pair dataset D2 of power grid workers; S3: Build a dynamic adversarial multimodal network (DAM-Net) for face recognition in the presence of occluded faces at power grid sites. This model is used to identify workers at power grid sites when their faces are obscured. S4: Train and update the parameters of each layer of the DAM-Net network built in S3 to obtain a trained DAM-Net network; S5: Use the trained DAM-Net network to identify people with obscured faces in power grid operation scenes and output the identity label and confidence level of the target person.

2. A method for face recognition at power grid operation site based on deep learning according to claim 1, characterized in that: The specific steps of S1 are as follows: S11: Use multispectral imaging equipment to collect facial images of workers at various power grid operation sites. For any worker, obtain the first facial image pair P1 =<I1,I2> , where I1 is the visible light image of the operator's face, and I2 is the infrared image of the operator's face; S12: Perform spatial alignment and image normalization on I1 and I2 to obtain the second face image pair P2 = <I”1,I’2>; S13: Pair all the second face images P2 =<I”1,I'2> A multispectral image pair dataset D1 of power grid workers is constructed.

3. A method for face recognition at power grid operation site based on deep learning according to claim 2, characterized in that: The specific steps of S12 are as follows: S121: P1 spatially aligns I1 to I2 using an affine transformation matrix to obtain a spatially aligned visible light face image I'1, so that I'1 and I2 are spatially aligned. The calculation formula for I'1 is as follows: I'1=warpAffine(I1,M) Among them, warpAffine represents affine transformation; M is the affine transformation matrix, which is a 2×3 matrix in the form of Among them, a, b, c, d control image rotation, scaling, and shearing; c, f control translation; S122: Perform channel normalization on I'1 to obtain a normalized visible light face image I"1; perform linear normalization on I2 to obtain a normalized infrared face image I'2, thereby improving the accuracy of subsequent processing and analysis.

4. The method for face recognition at power grid operation site based on deep learning according to claim 1, characterized in that: The specific steps of S2 are as follows: S21: Perform geometric occlusion on I″1 to obtain a visible light simulated occluded face image I m1 ; Semantic occlusion generation is performed on I”1 through 3D projection technology and GAN texture generation network to obtain visible light simulated occluded face image I m2 ; Perform adversarial occlusion generation based on the fast gradient sign method on I"1 to obtain the visible light simulation occluded face image I m3 ; S22: Use the person’s name as a label to simulate the visible light occlusion of the face image I m1 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m1 ; Simulate visible light to block face image I m2 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m2 ; Simulate visible light to block face image I m3 The face area is framed and marked to obtain the marked visible light simulation occluded face image I' m3 ; Use the person's name as a label, select and mark the face area of ​​​​I'2, and obtain the marked infrared face image y; S23: By I' m1 and y construct the third sub-worker face image pair S1 = <I' m1 ,y>; by I' m2 and y construct the fourth sub-worker face image pair S2 = <I' m2 ,y>; by I' m3 and y construct the fifth sub-worker face image pair S3 = <I' m3 ,y>; Collect all S1, S2, and S3 to obtain a multispectral image pair dataset D2 of power grid workers.

5. The method for face recognition at power grid operation site based on deep learning according to claim 1, characterized in that: The specific steps of S3 are as follows: S31: Design a dynamic occlusion detection module DODM, ​​which converts P3=<x,y> Input into DODM to obtain the face occlusion prediction heat map M _occ ; S32: Construct an adversarial feature generator AFG, M _occ and x are input into AFG to generate the face restoration feature map F face ; S33: Construct a multimodal identification network MDN, which transforms x, M _occ and F face Input into MDN to obtain the result matrix result, which includes the final identity label and its corresponding confidence.

6. The method for face recognition at power grid operation site based on deep learning according to claim 5, characterized in that: The specific steps of S31 are as follows: S311: Input x into the first three residual block groups of the ResNet18 encoder to obtain the simulated occluded face feature map F1; input y into the ViT encoder to obtain the infrared face feature map F2; S312: Design a deformable convolution fusion submodule DCFS, input F1 and F2 into DCFS, and obtain the dimension of feature fusion The fusion feature map F4; S313: Design a spatiotemporal attention mechanism submodule SAMS, input F4 into SAMS, and obtain a face occlusion prediction heat map M with a dimension of 1×H×W _occ ; The specific steps of S312 are as follows: S3121: F1 and F2 are concatenated through bimodal concatenation to obtain a dimension of The splicing feature map F3; S3122: Input F3 to the offset prediction layer, perform a convolution operation with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, then input it to the ReLU layer and a convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, perform nonlinear activation operations and convolution operations, and obtain a dimension of The offset characteristic map △p; S3123: Input F1, F2 and △p into the deformable convolution layer and upsampling layer, perform deformable convolution and upsampling operations, and obtain a dimension of The fusion feature map F4; F4=Upsample(DCNv1(F1,F2,Δp)) Among them, DCNv1 represents the deformable convolution layer operation; Upsample represents the upsampling operation; The specific steps of S313 are as follows: S3131: Design a channel attention network, input F4 into the channel attention network, and obtain the channel-by-channel multiplication feature map F c ; S3132: Design a spatial attention network to transform F c Input into the spatial attention network to get M _occ ; The specific steps of S3131 are as follows: S31311: Input F4 into the global average pooling layer and perform average pooling operation to obtain the dimension of Average pooling feature map F5; then input F5 into the shared multi-layer perceptron to obtain a dimension of The average weighted feature map F7; S31312: Input F4 into the global maximum pooling layer and perform the maximum pooling operation to obtain a dimension of The maximum pooled feature map F6 is then input into the shared multi-layer perceptron to obtain the maximum weighted feature map F8; S31313: Add F7 and F8 element by element, and the dimension is Pooling splicing feature map F9; then input F9 into Sigmoid and perform activation operation to obtain the dimension Pooled splicing activation feature map F 10 ; F9=add(F7,F8) F 10 =Sigmoid(F9) Where add(F7,F8) represents the add operation, which is to add F7 and F8 element by element, and Sigmoid represents the nonlinear activation operation; S31314: For F4 and F 10 Perform channel-by-channel multiplication to obtain a dimension of The channel-by-channel multiplication feature map F c ; in Represents a channel-by-channel multiplication operation; The specific steps of S3132 are as follows: S31321: F c Input to the channel average pooling layer, perform channel average pooling operation, and obtain the average pooling feature map F with a dimension of 1×H×W 11 ; S31322: F C Input to the channel maximum pooling layer, perform the channel maximum pooling operation, and obtain the channel pooling feature map F with a dimension of 1×H×W 12 ; S31323: F 11 and F 12 Perform the Concat operation to obtain a pooled splicing feature map F with a dimension of 2×H×W 13 ; Then F 13 Input to the convolution layer with a convolution kernel size of 7×7, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a spatial attention feature map F with a dimension of 1×H×W 14 ; Then F 14 Input into Sigmoid and perform activation operation to obtain M with dimension 1×H×W _occ ; F 13 =Concat(F 11 ,F 12 ) F 14 =Conv 7×7 (F 13 ) M _occ =Sigmoid(F 14 ) Concat(F 11 ,F 12 ) represents F 11 and F 12 Perform the Concat operation.

7. The method for face recognition at power grid operation site based on deep learning according to claim 5, characterized in that: The specific steps of S32 are as follows: S321: x and M _occ Perform the Concat splicing operation to obtain the face fusion feature map F with a dimension of (C+1)×H×W 15 ; F 15 =Concat(x,M _occ ) Where Concat(x,M _occ ) represents the pair x and M _occ Perform Concat splicing operation; S322: Construct a structure recovery channel submodule SRC to convert F 15 Input into SRC to obtain the structure recovery feature map F struct ; S323: Construct a texture restoration channel submodule TRC, and convert M _occ And the randomly generated noise vector z is input into the texture recovery channel to obtain the texture recovery feature map F with a dimension of C×H×W texture ; S324: Construct a dynamic region fusion submodule DRFM to transform F struct and F texture Input into the dynamic region fusion submodule DRFM to obtain F face ; The specific steps of S322 are as follows: S3221: F 15 Input to the ReLU layer and the convolution layer with a kernel size of 7×7, a stride of 2, and a zero padding of 3, the dimension is Face fusion enhanced feature map F 16 ; F 16 =Conv 7×7 (ReLU(F 15 )) S3222: F 16 Input into 4 series-connected hole residual blocks in sequence to obtain the fusion structure recovery feature map F 17 , wherein the structures of the first void residual block, the second void residual block, the third void residual block, and the fourth void residual block are the same; S3223: F 17 Input to the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, the ReLU layer, and the upsampling layer, perform convolution operations, activation operations, and upsampling operations with an upsampling factor of 2, and obtain an F with a dimension of C×H×W struct ; F struct =Upsample(ReLU(Conv 3×3 (F 17 ))) Upsample represents the upsampling operation; The specific steps of S323 are as follows: S3231: Randomly generate a noise vector z and input z into the multilayer perceptron to obtain the noise feature vector f1; S3232: M _occ Input to the global average pooling layer, perform average pooling operation, and obtain the face pooling scalar g; then input g into the multi-layer perceptron to obtain the face feature vector f2; S3233: Concatenate f1 and f2 to obtain the noise face concatenation feature vector f3; input f3 into the multi-layer perceptron MLP to obtain the style feature vector w; S3234: W and M _occ Perform AdaIN modulation operation to obtain F texture ; The specific steps of S324 are as follows: S3241: F struct Input to the global average pooling layer and the convolution layer with a convolution kernel size of 1×1, a stride of 1, and zero padding of 0, perform global average pooling and convolution operations, and obtain a structural weight vector α with a dimension of C×1×1; α=Conv 1×1 (GAP(F struct )) S3242: F texture Input it into the convolution layer with a convolution kernel size of 3×3, a stride of 1, and a zero padding of 1, and perform a convolution operation to obtain a texture weight feature map β with a dimension of 1×H×W; β=Conv 3×3 (F texture ) S3243: F struct 、F texture , α, β perform weighted fusion operation to obtain F with dimensions of C×H×W face ; F face =α·F struct +(1-α)·(β⊙F texture ) Among them, ⊙ represents the product operation; · represents the dot product; + represents the vector addition.

8. The method for face recognition at power grid operation site based on deep learning according to claim 5, characterized in that: The specific steps of S33 are as follows: S331: Construct an identity authentication submodule to combine x and F face Input to the identity identification submodule and output the identity similarity score vector S id ; S332: Construct a repair identification submodule to face and M _occ Input to the repair identification submodule and output the authenticity score vector D real ; S333: S id and D real Perform a decision fusion operation and then a max operation to obtain the result matrix result, where result is a two-dimensional matrix containing the final identification identity label and its corresponding confidence; result=max(W id ·S id +W real ·D real ) Where W id Represents the weight of the identity authentication branch; W real Represents the weight of generating the identification branch, and max represents the maximum value and its index; The specific steps of S331 are as follows: S3311: F face Perform feature extraction based on the ArcFace algorithm to obtain the repaired face extraction feature vector f gen ; S3312: Perform feature extraction on x based on the ArcFace algorithm to obtain the occluded face extraction feature vector f real ; S3313: f gen and f real Perform cosine similarity calculation to obtain the identity similarity score vector S id ; S id =1-cos({f gen |f real }) Its cos is a cosine similarity calculation function; The specific steps of S332 are as follows: S3321: F face Input into the PatchGAN encoder and get the dimension The repaired face image identification feature map F PGAN ; S3322: M _occ Input into the spatial attention layer, the dimension is The spatial attention weighted feature map F space ; S3323: F PGAN and F space Perform the Concat operation to obtain the dimension Multi-scale feature map F mul ; Then F mul Input it into the convolution layer with a convolution kernel size of 1×1, a step size of 1, and a zero padding of 1, and perform a convolution operation to obtain the authenticity score D real ; F mul =Concat(F PGAN ,F space ) D real =conv 1×1 (F mul ) Concat(F PGAN ,F space ) represents the Concat operation.

9. The method for face recognition at power grid operation site based on deep learning according to claim 1, characterized in that: The specific steps of S4 are as follows: S41: Initialize the hyperparameters required for training DAM-Net; S42: Divide the multispectral image pair dataset of power grid workers constructed in S2 into a training set and a test set according to a certain ratio, ensuring that the two do not overlap; then divide the training set into multiple batches, input one batch of training set into the network for training each time, and calculate the loss value of the batch; S43: After traversing all batches of the entire training set in each round, the performance of the model is periodically evaluated using the test set data. The test set is fed into the model in batches of the same size as in the training phase, and the prediction error of each batch is calculated. The test set loss value for the current round is then aggregated to obtain the test set loss value. The test set loss value can be used to dynamically track changes in the network's performance on non-training data. By continuously observing the changing trend of the test set loss value, it is possible to determine whether the model shows signs of overfitting. When the test set loss value no longer decreases over multiple consecutive training cycles, or shows an upward trend, it can be determined that the model's generalization ability has decreased, thereby triggering the preset strategy adjustment mechanism. When the test set loss value tends to stabilize with the training process and meets the expected performance standards, the network training can be considered complete, and the final network parameters will be used for subsequent identification of workers at the power grid operation site when their faces are obscured.

10. The method for face recognition at power grid operation site based on deep learning according to claim 1, characterized in that: The specific steps of S5 are as follows: S51: After the DAM-Net network training is completed, it is applied to the face recognition task under human occlusion in actual power grid operation scenarios; S52: The collected identity images of power grid workers are input into the trained DAM-Net network. The model identifies each power grid worker in the image one by one and outputs the corresponding identity label and confidence level of each worker in the image.

Citation Information

Patent Citations

  • Face recognition method and system based on artificial intelligence, electronic equipment and medium

    CN114299568A

  • Shielded face recognition method and device, electronic equipment and storage medium

    CN116129499A

  • Thermal infrared face image texture enhancement method and system for recognition process, and storage medium

    CN118918024A

  • Deep learning-based face feature point detection method

    WO2022151535A1