Workshop personnel presence detection method based on multi-view gait spatiotemporal restoration network

Through multi-view generation of adversarial networks and occlusion gait space-time repair networks, the problem of reduced gait recognition accuracy caused by viewing angle changes and occlusion in the workshop environment is solved, and efficient on-the-job detection and identity recognition are achieved.

CN119206787BActive Publication Date: 2025-07-29青岛冠成软件有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411330334.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-07-29
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

In workshop environment, the problem of viewing angle changes and occlusion leads to a reduction in gait recognition accuracy, which is difficult for existing algorithms to effectively deal with, especially when identifying personnel under protective clothing, traditional methods consume a lot of resources and are limited in robustness.

Method used

Multi-view generation adversarial network is used to generate occlusion gait images from the target perspective, and an occlusion gait space-time repair network is designed, including semantic segmentation networks and gait space-time repair networks. The space-time features are extracted through an end-to-end method to solve the problem of view angle changes and occlusion.

Benefits of technology

It significantly improves the accuracy and consistency of identity identification of on-the-job inspections by workshop personnel, can effectively deal with different occlusion degrees and types, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206787B_ABST
    Figure CN119206787B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of gait recognition and personnel on-duty detection, and specifically discloses a method for detecting personnel on duty in a workshop based on a multi-view gait spatio-temporal repair network. The method of the present invention first preprocesses the input image to obtain a gait energy map in an occluded state, and uses a multi-view generative adversarial network to obtain an occluded gait image in the target view. Then, an occluded gait spatio-temporal repair network is used to repair the occluded gait sequence of the personnel, and an encoder-decoder structure is used to maintain the spatio-temporal coherence between frames while repairing the spatial information of each frame of the gait image. Finally, a feature extraction network is used to extract gait features, maintain the consistency between its identity information and the registered information, and accurately detect the on-duty state of the personnel. The present invention has a high repair effect for different occlusion degrees and different occlusion types, and improves the detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of gait recognition and personnel on-duty detection, and particularly relates to a method for detecting personnel on duty in a workshop based on a multi-view gait spatio-temporal repair network. Background Technique

[0002] In a workshop or production line, it is necessary to monitor in real time whether personnel are working at their designated posts. At the same time, different production environments have different requirements for personnel's clothing. For some special working scenarios that require wearing protective clothing, it is relatively difficult to identify the identity of personnel under the protective clothing. Compared with traditional recognition technologies such as face, iris, and fingerprint, gait recognition has great application potential in personnel on-duty detection technology because it has low requirements for image resolution, can identify at a long distance, and does not require the cooperation of the subject, with good versatility and flexibility. Since gait metrics have significant differences, by analyzing features such as the posture, step length, step width, and walking speed of an individual during walking, the identity can be recognized and confirmed.

[0003] Gait recognition is sensitive to view changes, and different camera angles may affect the accuracy of recognition. At the same time, gait occlusion under different views will also affect the recognition accuracy. Currently, most gait recognition algorithms use deep convolutional network models, and the improvements usually focus on adjusting the network structure, changing the convolution kernel size, extracting and fusing multi-scale features, adding attention mechanisms, and improving loss functions. Through the above improvements, the network is helped to learn more fine-grained deep features to improve the detection performance of the algorithm. In addition, some researchers also combine traditional algorithms with deep learning algorithms. Based on the original deep learning model, the traditional algorithms are embedded into the overall detection process in a pre-processing or post-processing manner to improve the detection accuracy of the algorithm. There are also some researchers who combine gait recognition with other recognition methods, such as face and iris, and propose a multi-modal fusion method. The two recognition methods complement each other to improve the recognition accuracy and precision.

[0004] Although the above several improved and combined algorithms can improve the personnel recognition accuracy on the premise of knowing the gait contour, the gait contour sequence is unknown in the case of occlusion, so it is impossible to effectively cope with the personnel occlusion situation in the workshop environment.

[0005] Specifically, in a workshop environment, the change in perspective has a significant impact on gait recognition, reducing the recognition accuracy. The gait features observed from the front and side are very different, which poses a challenge to the recognition accuracy of gait recognition algorithms. To address the issue of perspective change, although many methods obtain gait sequences from different angles by increasing the number of cameras, it requires a large amount of resources. There are also some methods that use skeletal information, but they require additional equipment and have limited robustness. At the same time, the occlusion problem caused by perspective change also needs to be solved. For example, when part of a pedestrian's body (such as the arm or leg) is occluded, the gait recognition system may not be able to obtain complete gait features, affecting the accuracy of recognition. When a person is occluded, some important spatio-temporal information will be lost, such as the limb movement trajectory and the body swing amplitude, which affects the recognition accuracy. Summary of the Invention

[0006] The purpose of the present invention is to propose a method for detecting the presence of workshop personnel based on a multi-view gait spatio-temporal repair network. This method uses a multi-view generative adversarial network to obtain occluded gait images from the target perspective and designs an occluded gait spatio-temporal repair network model to reconstruct and recognize the gait contours of employees from different perspectives, effectively solving the problems of perspective change and occlusion, thereby improving the accuracy of detecting the presence of workshop personnel and the accuracy of personnel identity.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A method for detecting the presence of workshop personnel based on a multi-view gait spatio-temporal repair network includes the following steps:

[0009] Step 1. First, for the input occluded gait energy maps from different perspectives, use a multi-view generative adversarial network for perspective conversion to generate occluded gait images from the target perspective;

[0010] Step 2. Send the occluded gait images from the target perspective into the occluded gait spatio-temporal repair network model, which includes a semantic segmentation network and a gait spatio-temporal repair network;

[0011] The semantic segmentation network obtains the occluded gait area and identifies the gait contours from different perspectives according to the semantic information of the image. The gait spatio-temporal repair network repairs the occluded gait area to obtain a complete gait image repair sequence;

[0012] Step 3. The gait image repair sequence is subjected to feature extraction by a gait feature extraction network to obtain employee identity information, and the detected employee identity is compared with the pre-stored employee identity file information to achieve the detection of the presence of workshop personnel.

[0013] The present invention has the following advantages:

[0014] As described above, the present invention relates to a method for detecting the presence of workshop personnel based on a multi-view gait spatio-temporal repair network. The method first introduces a multi-view generative adversarial network (Multi-View GAN) for generating occluded gait images from the target view, and then designs an occluded gait spatio-temporal repair network model, which includes a semantic segmentation network and a gait spatio-temporal repair network. Among them, the occluded gait images from the target view are segmented by the semantic segmentation model, and then the semantic segmentation model and the feature extraction model are connected by the gait spatio-temporal repair network. The gait spatio-temporal repair network constructed by using the Transformer and the encoder-decoder structure is responsible for extracting the spatial information and temporal information of the gait images respectively. The method of the present invention extracts spatio-temporal features in an end-to-end manner. Compared with the simple extraction of gait features, the gait features in the occluded situation are first repaired for the gait contour and then the features are extracted, which can significantly improve the consistency of personnel identity. The method of the present invention has a high repair effect for different occlusion degrees and different occlusion types, and improves the accuracy of detecting the presence of workshop personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the method for detecting the presence of workshop personnel based on the multi-view gait spatio-temporal repair network in the embodiment of the present invention;

[0016] Figure 2 is a multi-view generative adversarial network diagram in the embodiment of the present invention;

[0017] Figure 3 is a generator network structure diagram in the embodiment of the present invention;

[0018] Figure 4 is a gait spatio-temporal repair network diagram in the embodiment of the present invention;

[0019] Figure 5 is a U-Net semantic segmentation network structure diagram in the embodiment of the present invention;

[0020] Figure 6 is a multi-scale Transformer encoder-decoder structure diagram in the embodiment of the present invention;

[0021] Figure 7 is a feature extraction network GaitSet structure diagram in the embodiment of the present invention;

[0022] Figure 8 is a multi-layer full-process pipeline structure diagram in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present invention will be further described in detail below in conjunction with the drawings and the specific embodiments:

[0024] Embodiment 1

[0025] In view of the problems of personnel on-duty detection and identity recognition in the workshop or production line environment, the present invention proposes a method for detecting personnel on-duty in the workshop based on a multi-view gait spatio-temporal repair network, which is convenient for obtaining personnel gait images from different perspectives for identity recognition, and solves the influence of perspective changes and the resulting occlusion problems in different production environments on the identity of employees.

[0026] To solve the above problems, the present invention first introduces a multi-view generative adversarial network to generate occluded gait images in the target perspective, which can effectively overcome the misdetection and missed detection of personnel in a single perspective and significantly improve the accuracy of personnel gait recognition. Secondly, a semantic segmentation network is used to generate a pixel-level occlusion region map, which can provide more accurate occlusion information compared with the object detection network, and combined with global information, improve the repair effect when the occlusion degree is large. In addition, by introducing multi-scale features, by deepening the convolutional layer and adding pooling operations in the encoder-decoder of the gait spatio-temporal repair network, it is possible to better capture the detailed information and global information in the gait sequence, improve the gait repair effect, and use the encoder-decoder structure to maintain the spatio-temporal coherence between frames while repairing the spatial information of each frame of gait image. Finally, the binary cross-entropy loss function is introduced to better combine the feature extraction network and the gait spatio-temporal repair network, effectively distinguish the background and the employee body, focus on the pedestrian motion information, and significantly improve the consistency of the identity recognition of on-duty personnel in the workshop. The present invention is beneficial to solving the problem of personnel identity detection under perspective changes and occlusion conditions in the existing workshop environment, and improving the accuracy and efficiency of gait recognition.

[0027] As Figure 1 shown, the method for detecting personnel on-duty in the workshop based on the multi-view gait spatio-temporal repair network includes the following steps:

[0028] Step 1. First, for the input occluded gait energy maps from different perspectives, use the multi-view generative adversarial network for perspective conversion to generate occluded gait images in the target perspective that are easier to detect.

[0029] Since perspective changes will change the gait contour of pedestrians, resulting in changes in features, and existing gait datasets usually only contain limited perspectives, it is difficult to meet the needs of cross-perspective gait recognition. If the gait contour in an extreme perspective (such as 0°) is directly used for recognition, the recognition accuracy will be reduced.

[0030] Therefore, the present invention introduces a multi-view gait generation network model, namely a multi-view generative adversarial network, whose network structure is as Figure 2 shown. This network can generate gait images in different target perspectives through the combination of a generator and a discriminator, thereby expanding the existing dataset and providing more training samples while improving the recognition accuracy.

[0031] The processing flow of the multi-view generative adversarial network is as follows:

[0032] First, the input original gait images are processed to obtain occluded gait energy maps, and the occluded gait energy maps and target view labels are used as the inputs of the multi-view generative adversarial network. Each image corresponds to a walking phase, and the target view label specifies the target view of the generated gait image. For example, 0° represents the original view, 90° represents the target view, etc.

[0033] Taking the gait energy map and the target view label as the inputs of the multi-view generative adversarial network, and then passing through the decoder, it learns to convert the gait energy map into an image under the target view, generating a pseudo-image under the target view.

[0034] As Figure 3 shown, the encoder of the generator consists of four convolutional layers with a stride of 2, which is used to extract image feature information, and the decoder consists of four transposed convolutional layers with a stride of 2, which is used to generate an image under the target view according to the features.

[0035] The encoder consists of four convolutional layers with a stride of 2. Each convolutional layer uses the RELU activation function. Except for the first layer, each convolutional layer uses batch normalization (BN) to improve the performance of the model. The image feature information extracted by the encoder is input into the decoder. The decoder consists of four transposed convolutional layers. The first three convolutional layers use batch normalization and the RELU activation function, and the last layer uses the Tanh activation function to make the output image under the target view more standardized.

[0036] The discriminator is used to judge the authenticity, view angle, and identity of the images input into the discriminator, that is, to judge whether the input image is a real image or a fake image generated by the generator, whether the view angle of the generated image is the target view, and whether the generated image maintains the human identity information in the original image; then according to the output result of the discriminator, the parameters of the generator and the discriminator are updated to enable the generator to generate more realistic, more in line with the target view, and more able to maintain identity information images.

[0037] Since a single generator generates false gait images under different walking conditions, or even under different datasets, there should theoretically be a distribution difference between real and false gait images. If true and false samples are directly merged during the training process of the multi-view generative adversarial network, the distribution difference will affect the performance of the test samples. The present invention proposes a domain alignment method to reduce the distribution difference and guide false gait samples to improve the generalization ability of the multi-view generative adversarial network.

[0038] Assume that the real gait image feature set is The false gait image feature set is

[0039] where {x ri , y ri}, {x fj , y fj} both represent tuples containing feature vectors, and the label space Y r = Y f , Y r is the label set corresponding to the feature vectors in the real gait image feature set, and Y f is the label set corresponding to the feature vectors in the fake gait image feature set.

[0040] Domain alignment reduces the distribution difference between real and fake gait samples by learning a feature mapping function.

[0041] To achieve image generation covering the same and different data sets, consider the distribution alignment of gait images for marginal distribution alignment and conditional distribution alignment. The distribution function D F (D r , D f ) is defined as:

[0042]

[0043] where μ ∈ [0, 1] represents an adaptive factor that balances the importance of the marginal distribution and the conditional distribution, c ∈ 1,..., C represents the subject identity of the gait image, and D F (P r , P f ) represents the marginal distribution alignment; represents the conditional distribution alignment of subject c.

[0044] The distribution difference between fake gait images and real gait images is calculated by the projected maximum mean difference, and the alignment method of the distribution difference is written as:

[0045]

[0046] where E[·] represents the average of the embedded features, H k represents the reproducing kernel Hilbert space, F(·) represents the feature mapping function, X r represents the first feature vector in the real gait image feature set, X f represents the first feature vector in the fake gait image feature set, represents the set of gait image feature vectors belonging to the c-th target person in the real gait image set, represents the set of gait image feature vectors belonging to the c-th target person in the fake gait image set.

[0047] By introducing a multi-view generative adversarial network, the present invention inputs the original gait energy map and the target view label, and can generate occluded gait images of different views with high quality and high diversity. By generating diverse occluded gait images and performing domain adaptation, the problem of insufficient data in cross-view gait recognition is effectively solved, thereby improving the performance of gait recognition.

[0048] Step 2. Send the occluded gait image under the target view into the occluded gait spatio-temporal restoration network model, which includes a semantic segmentation network and a gait spatio-temporal restoration network, as Figure 4 shown.

[0049] The semantic segmentation network obtains the occluded gait area, identifies the gait contours under different views according to the semantic information of the image, and the gait spatio-temporal restoration network repairs the occluded gait area to obtain a complete gait image repair sequence.

[0050] The semantic segmentation network adopts a pre-trained U-Net network. In practical applications, the scenario of gait recognition is based on RGB images, and the recognition scenario may be relatively complex. The gait contour maps generated by different pedestrian detection and segmentation algorithms are also different. Using the semantic segmentation network U-Net can more accurately obtain the occluded area, thereby improving the effect of gait restoration.

[0051] Specifically, the trained U-Net network is used to detect the occluded part of each frame in the occlusion sequence as prior knowledge for repair, and accurate occluded area information is obtained to repair the part of the occluded area.

[0052] The semantic segmentation network includes a backbone segmentation network and an enhanced segmentation network, as Figure 5 shown.

[0053] The backbone segmentation network includes four groups of convolutional blocks. Each group of convolutional blocks contains two 3×3 convolutional layers and a pooling layer. Adding the pooling layer is to downsample the image. After passing through the backbone segmentation network, image features retaining semantic information are generated.

[0054] The obtained image features enter the enhanced segmentation network after passing through two 3×3 convolutional layers.

[0055] The enhanced segmentation network includes four groups of convolutional blocks. Each group of convolutional blocks contains an upsampling module and two 3×3 convolutional layers.

[0056] First, the occluded gait image under the target view is sent into the U-Net network. The input image contains occluders and occlusion information, which helps the semantic segmentation network focus on the target area.

[0057] Then, during the upsampling process of the enhanced segmentation network, continuously fuse the downsampling process, and transfer the feature information in the backbone segmentation network to the enhanced segmentation network, which helps to restore the detailed information.

[0058] Finally, the output of the U-Net network is the occluded gait segmentation image.

[0059] The semantic segmentation network obtains the occluded gait area. According to the semantic information of the image, it provides pixel-level occluded area information, identifies the gait contours from different perspectives, and can identify the detailed information of the occluded area when the occlusion degree is large, improving the robustness of gait recognition.

[0060] The gait spatio-temporal repair network takes the Transformer encoder-decoder as the backbone network. By deepening the convolutional layer and adding pooling operations, multi-scale features are formed, making it suitable for the gait repair model. Batch normalization (BN) and the ReLU non-linear activation function are used in each convolutional layer. Adding a pooling layer can involve every element during image downsampling, retaining more detailed information. At the same time, a convolution is performed first before each downsampling to retain more temporal information. The multi-scale features can better capture the detailed information and global information in the gait sequence.

[0061] The input sequence of occluded gait segmentation images passes through the encoder, spatio-temporal transformer, and decoder network in sequence, and a complete gait image repair sequence is obtained from the decoder output by combining prior knowledge.

[0062] The fused gait map not only completes the occlusion repair but also maximally retains the original gait contour data.

[0063] The improved multi-scale Transformer encoder-decoder structure is as Figure 6 shown. The encoder of the gait spatio-temporal repair network has four convolutional layers. The structure of each convolutional layer is uniformly a 3×3 convolutional kernel with a stride of 1. Except for the last layer, each convolutional layer enters the next convolutional layer after passing through the max-pooling layer; the kernel size of each max-pooling layer is kernel = 2, and the stride is stride = 2. The number of convolutional kernels also increases from 64 to 512; the occluded gait image sequence is fed into the encoder for convolutional operations to obtain a low-dimensional feature f with spatio-temporal information, including the temporal and spatial information of the gait image.

[0064] The low-dimensional features f containing spatiotemporal information generated by the encoder are used as input to the Transformer. A spatiotemporal transformer (STT) is added between the encoder and decoder. STT comprehensively repairs the missing parts of the gait image sequence, ensuring that the restored images are consistent with the original images in terms of spatial layout and temporal order. The introduction of STT also helps reduce the number of parameters required by the model, making it more efficient. Through STT, global and local restoration features of the occluded gait are obtained.

[0065] STT consists of 4 Transformer blocks, which are divided into two sub-layers: multi-head attention MultiHead and multi-layer perceptron (MLP).

[0066] MultiHead is a multi-scale self-attention module. i (h i ) corresponds to attention operations of different scales, and designs attention of different scales to obtain spatiotemporal information from local and global perspectives to repair the occluded parts.

[0067] The MLP sublayer uses two 2D convolution residual blocks with kernel 3×3 and stride 1 to process the spatiotemporal features of multi-head attention. In each MLP sublayer, the low-dimensional features output by the previous layer are used as the input of this layer.

[0068] Among them, t, h, w, and c represent the time dimension, height, width, and number of channels of the low-dimensional features respectively; the c dimension is divided into n features n is the number of heads in the multi-head attention.

[0069] Each feature As the h in the corresponding MultiHead i input.

[0070] In h i middle First it is mapped to q i 、v i and k i .

[0071] where q i 、v i 、k i They correspond to the query, value and key in the attention mechanism respectively.

[0072] Then, q i and k i Perform matrix multiplication to obtain attention, and finally use attention as weight and v iMatrix multiplication serves as the output of h i , specifically as follows:

[0073]

[0074] where σ represents the softmax activation function and size represents the number of elements in q o .

[0075] After concatenating the features of all h i outputs and performing a residual connection with , the output f of the MultiHead is obtained through convolution and activation functions 1 ∈R t×h×w×c / n , where the values of t, h, w, and c are the same as those of , specifically as follows:

[0076]

[0077] where ρ represents the Leaky ReLU (leaky rectified linear unit) activation function, 3_2C represents a two-dimensional convolutional neural network with a kernel of 3, and concat represents concatenation in the channel dimension.

[0078] As Figure 4 shown, the decoder of the gait spatio-temporal repair network includes four convolutional layers; except for the last layer in the decoder structure, each convolutional layer uses a 3×3 convolutional kernel with a stride of 1, and at the same time, alternating convolutional layers and upsampling operations are used. There is a ReLU non-linear activation function after each convolution to slow down the loss of details during downsampling in the encoding process and minimize the impact of the vanishing gradient problem; the low-dimensional features output by the spatio-temporal transformer are decoded to gradually restore the spatial information of the gait image and maintain temporal coherence, obtaining a reconstructed complete gait image repair sequence.

[0079] In the present invention, by improving the encoder-decoder in the gait spatio-temporal repair network, multi-scale features are introduced. Higher-dimensional features are obtained by increasing the depth of the convolutional layer, and lower-dimensional features are obtained by adding pooling operations, so as to better capture global and local information to repair the gait contour.

[0080] Step 3. The gait image repair sequence undergoes feature extraction through the gait feature extraction network to obtain employee identity information, and the detected employee identity is compared with the pre-stored employee identity profile information to achieve on-duty detection of workshop personnel.

[0081] Use the gait image repair sequence after being repaired in step 2 as the input of the feature extraction network Gaitset to obtain more discriminative gait feature information, which is convenient for identifying the identity of personnel. Its structure is as Figure 7 shown.

[0082] Gaitset is a set-based gait recognition network rather than an ordered sequence, and it can better handle the variable length and irregularity in the gait sequence. The detection model designed based on the convolutional network has a good effect on the extraction of gait features. Among them, the GaitSet feature extraction model has a greater influence on gait recognition. A custom end-to-end deep learning model is used to implement gait recognition. Convolutional network (CNN) and pooling operations are used to regard the gait sequence as a set, and the maximum function is used to compress the frame-level spatial feature sequence, which is extremely simple and effective. It takes the repaired multi-frame contour sequence as the input and mainly consists of two parts, namely the MultlayerGlobal Pipeline (MGP) and the Horizontal Pyramid Matching (HPM).

[0083] MGP is mainly divided into two branches: one is the main branch, and the other is the auxiliary branch. Its structure is as Figure 8 shown.

[0084] The main branch is used to process the multi-frame data separated from the video based on the features of all pictures, and feature calculation and dimensionality reduction processing are carried out by using two convolutional operations and one downsampling operation.

[0085] The auxiliary branch is used to process the multi-frame data separated from the video based on the features of a certain frame.

[0086] The processing of the auxiliary branch is synchronized with that of the main branch, and feature extraction is carried out on the data after each downsampling. The extracted frame features are fused into the feature processing result of the main branch, and set pooling is applied to aggregate the frame-level features into set-level features.

[0087] Then, the set-level features are divided into different scales through horizontal pyramid mapping to extract features at different spatial positions, and the pooling results are pooled together, so as to enrich the discriminative features of the data, make it more distinguishable and convenient for calculating similarity.

[0088] After feature extraction by the GaitSet gait feature extraction network, the feature data of the employee is obtained, and then it is compared with the personnel identity in the feature database (the pre-stored employee identity file information) to obtain the on-duty personnel identity information.

[0089] In the function optimization stage after feature extraction, a binary cross-entropy loss function is introduced, which is mainly used to handle the consistency of gait image sequence features. By further distinguishing the workshop environment and employees, it focuses on the motion information of pedestrians and better identifies the identities of employees. By introducing the cross-entropy loss function, the present invention better combines the gait spatio-temporal repair network and the feature extraction network, repairs the gait of personnel from different perspectives and extracts gait features, can significantly distinguish the workshop environment and employees, focuses on the motion information of pedestrians, and can better maintain the consistency of personnel identities. Compared with other models (in the case of gait cycle determination or integrity), by first performing perspective transformation and then repairing and extracting features, the problems of perspective change and occlusion detection in the existing workshop environment can be solved.

[0090] The present invention uses the reconstruction loss L of the repair sequence r , the triplet loss L t and the cross-entropy loss function L c as the joint loss to repair the network; the calculation formula of the reconstruction loss L r is as follows:

[0091] L r =γ h L h +γ v L v ;

[0092]

[0093] In the formula, is the original unoccluded sequence of the occluded sequence , is the repaired gait sequence, L r is divided into the hole area loss L and the non-hole area loss L h by the prior knowledge v , γ h and γ v are the loss weights of L h and L v respectively; · represents element-wise multiplication.

[0094] The calculation formulas of the triplet loss L t and the cross-entropy loss function L c are as follows:

[0095] L t =max(d(a,p)-d(a,q)+margin,0);

[0096]

[0097] where \(d(a,p)\) and \(d(a,n)\) are the distances between the anchor sample \(a\) and the positive sample \(p\), and between the anchor sample \(a\) and the negative sample \(q\) in the embedding space, respectively; margin is a preset threshold; \(m\) represents the number of samples; \(n\) represents the number of classes; \(p(x ij ) represents the true label that the \(i\)-th sample belongs to the \(j\)-th class; \(q(x ij ) represents the probability that the model predicts the \(i\)-th sample belongs to the \(j\)-th class.

[0098] By optimizing \(L r \), the repaired gait sequence is made as close as possible to the original unoccluded sequence in content, so as to maintain the consistency in content between the repaired sequence and the original sequence. Then, \(L r \) and \(L c \) are used to keep the personal identity information unchanged, so that the features of the repaired sequence are consistent with those of the original sequence. The calculation of the combined loss \(L\) is as follows:

[0099] \(L=\omega r L r +\omega t L t +\omega c L c ;\)

[0100] where \(\omega r \), \(\omega t \), \(\omega c \) are the loss weights of \(L r \), \(L t \) and \(L c \) respectively.

[0101] The present invention uses three loss functions and combines prior knowledge (occlusion area information) for optimization. Among them, the reconstruction loss is used to measure the difference between the repaired gait image sequence and the original unoccluded sequence to ensure consistency in content, and the triplet loss and cross-entropy loss are used to ensure the feature consistency between the repaired gait image sequence and the original sequence. The present invention uses the Gaitset network to extract features and optimizes the network through the triplet loss and cross-entropy loss functions, so that the repaired sequence has similar features to the original sequence, thereby correctly judging the employee identity features and determining whether the employee is on duty normally.

[0102] Of course, the above description is only the preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A method for detecting the presence of workshop personnel based on a multi-view gait spatio-temporal repair network, characterized in that: It includes the following steps: Step 1. First, use a multi-view generative adversarial network to perform perspective conversion on the input occluded gait energy maps from different perspectives to generate occluded gait images in the target perspective; Among them, the occluded gait energy map and the target perspective label are used as the input of the multi-view generative adversarial network, and the gait images in different target perspectives are generated through the combination of a generator and a discriminator; the discriminator is used to judge the authenticity, perspective, and identity recognition of the images input into the discriminator; a domain alignment method is designed in the multi-view generative adversarial network to reduce the distribution difference and guide the false gait samples to improve the generalization ability of the multi-view generative adversarial network model; By introducing a multi-view generative adversarial network, the original gait energy map and the target perspective label are input to generate high-quality and highly diverse occluded gait images in different perspectives. By generating diverse occluded gait images and performing domain adaptation, the problem of insufficient data in cross-view gait recognition is solved, thereby improving the performance of gait recognition; Step 2. Send the occluded gait image in the target perspective into the occluded gait spatio-temporal repair network model, and the occluded gait spatio-temporal repair network model includes a semantic segmentation network and a gait spatio-temporal repair network; The semantic segmentation network uses a pre-trained U-Net network; The semantic segmentation network obtains the occluded gait area, identifies the gait contours in different perspectives according to the semantic information of the image, and the gait spatio-temporal repair network repairs the occluded gait area to obtain a complete gait image repair sequence; The gait spatio-temporal repair network uses a Transformer encoder-decoder as the backbone network; The occluded gait segmentation image sequence output by the U-Net network passes through the encoder, spatio-temporal transformer, and decoder network in sequence, and combines prior knowledge to obtain the reconstructed complete gait image repair sequence from the decoder output; the number of convolutional layers is increased in the encoder and decoder of the gait spatio-temporal repair network, and pooling operations are added to capture global and local information changes; Use the reconstruction loss, triplet loss, and cross-entropy loss function of the repair sequence as the joint loss to repair the network; Step 3. The gait image repair sequence passes through the gait feature extraction network to extract the employee identity information, and compares the detected employee identity with the pre-stored employee identity file information to achieve the detection of the presence of workshop personnel.

2. The method for detecting the presence of workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 1, wherein, In the said Step 1, the multi-view generative adversarial network generates gait images in different target perspectives through the combination of a generator and a discriminator, and the generator adopts an encoder-decoder structure; the processing flow of the multi-view generative adversarial network is as follows: First, the input original gait image is processed to obtain the occluded gait energy map. The occluded gait energy map and the target perspective label are used as the input of the multi-view generative adversarial network. The gait feature information of the employee image is generated through the encoder, and then through the decoder, it learns to convert the original gait image into an image in the target perspective to generate a pseudo-image in the target perspective; The discriminator is used to judge the authenticity, perspective, and identity of the images input into it, to determine whether the input image is a real image or a fake image generated by the generator, whether the perspective of the generated image is the target perspective, and whether the generated image maintains the human identity information in the original image; then, according to the output results of the discriminator, the parameters of the generator and the discriminator are updated, so that the generator can generate more realistic, more in line with the target perspective, and more identity-preserving images.

3. The in-work detection method for workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 2, wherein, The encoder of the generator is used to extract image feature information; the decoder of the generator is used to generate an occluded gait image in the target perspective based on the features; where: The encoder consists of four convolutional layers with a stride of 2. Each convolutional layer uses the RELU activation function. Except for the first layer, each convolutional layer uses batch normalization to improve the performance of the model. The image feature information extracted by the encoder is input into the decoder; the decoder consists of four transposed convolutional layers with a stride of 2. The first three convolutional layers use batch normalization and the RELU activation function, and the last layer uses the Tanh activation function to make the output image in the target perspective more normalized.

4. The method for detecting the presence of workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 2, wherein Suppose the real gait image feature set is The fake gait image feature set is where {x ri , y ri}, {x fj , y fj} both represent tuples containing feature vectors, and the label space Y r = Y f , Y r is the label set corresponding to the feature vectors in the true gait image feature set, and Y f is the label set corresponding to the feature vectors in the false gait image feature set; Domain alignment reduces the distribution difference between real and fake gait samples by learning a feature mapping function; Implement image generation covering the same and different datasets, consider the distribution alignment of gait images with marginal distribution alignment and conditional distribution alignment, and the distribution function D F (D r ,D f ) is defined as: where, μ∈[0,1] represents an adaptive factor for balancing the importance of the marginal distribution and the conditional distribution, c∈1,...,C represents the subject identity of the gait image, D F (P r ,P f ) represents the marginal distribution alignment; represents the conditional distribution alignment of subject c; Real gait images and fake gait images calculate the distribution difference between real gait images and fake gait images through projected maximum mean discrepancy, and the alignment method of the distribution difference is written as: Among them, E[·] represents the average value of the embedded features, and H k represents the reproducing kernel Hilbert space, F(·) represents the feature mapping function, and X r represents the first eigenvector in the real gait image feature set, and X f represents the first eigenvector in the fake gait image feature set, represents the set of gait image feature vectors of the c-th target person in the real gait image set, represents the set of gait image feature vectors of the c-th target person in the fake gait image set.

5. The method for detecting the presence of workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 1, characterized in that, In step 2, the trained U-Net network is used to detect the occluded part of each frame in the occlusion sequence as prior knowledge for repair, and accurate occlusion region information is obtained to repair the occluded part of the region; The semantic segmentation network includes a backbone segmentation network and an enhanced segmentation network; The backbone segmentation network includes four groups of convolutional blocks. Each group of convolutional blocks contains two 3×3 convolutional layers and a pooling layer. Adding the pooling layer downsamples the image. After passing through the backbone segmentation network, image features that retain semantic information are generated; The obtained image features enter the enhanced segmentation network after passing through two 3×3 convolutional layers; The enhanced segmentation network includes four groups of convolutional blocks. Each group of convolutional blocks contains an upsampling module and two 3×3 convolutional layers; First, the occluded gait image in the target perspective generated by the multi-view generative adversarial network is sent into the U-Net network. The input image contains occluders and occlusion information, which helps the semantic segmentation network focus on the target area; Then, during the upsampling process of the enhanced segmentation network, the downsampling process is continuously fused, and the feature information in the backbone segmentation network is transmitted to the enhanced segmentation network, which helps to restore the detailed information; Finally, the output of the U-Net network is the occluded gait segmentation image.

6. The in-work detection method of workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 1, characterized in that The encoder of the gait spatio-temporal repair network has four convolutional layers. The structure of each convolutional layer is uniformly a 3×3 convolutional kernel with a stride of 1. Except for the last layer, each convolutional layer enters the next convolutional layer after passing through a max-pooling layer; the kernel size of each max-pooling layer is kernel = 2, the stride is stride = 2, and the number of convolutional kernels also increases from 64 to 512; the occluded gait segmentation image sequence is sent into the encoder for convolutional operation to obtain low-dimensional features with spatio-temporal information, including the time information and spatial information of the gait images.

7. The method for detecting the presence of workshop personnel based on the multi-perspective gait spatio-temporal repair network according to claim 6, wherein The decoder of the gait spatio-temporal repair network includes four convolutional layers; in the decoder structure, except for the last layer, each convolutional layer uses a uniformly structured 3×3 convolutional kernel with a stride of 1, and at the same time, convolutional layers and upsampling operations are alternately arranged, and there is a ReLU non-linear activation function after each convolution to slow down the loss of details during downsampling in the encoding process and minimize the impact of the vanishing gradient problem; the low-dimensional features output by the spatio-temporal transformer are decoded to gradually restore the spatial information of the gait images and maintain temporal coherence, obtaining a reconstructed complete gait image repair sequence.

8. The in-work detection method for workshop personnel based on the multi-view gait spatio-temporal repair network according to claim 1, characterized in that Reconstruction loss L using the repair sequence r , triplet loss L t and cross-entropy loss function L c are used as the combined loss to repair the network; Reconstruction loss L r The calculation formula is as follows: L r = γ h L h + γ v L v ; In the formula, is the occlusion sequence of the original unoccluded sequence, is the restored gait sequence, L r is divided by the prior knowledge into the hole region loss L h and the non-hole region loss L v , γ h and γ v are the loss weights of L h and L v respectively; · represents element-wise multiplication; Triplet loss L t and cross-entropy loss function L c are calculated as follows: L t = max(d(a, p) - d(a, q) + margin, 0); Wherein, d(a, p) and d(a, q) are the distances between the anchor sample a and the positive sample p, and between the anchor sample a and the negative sample q in the embedding space, respectively; margin is a preset threshold; m represents the number of samples; n represents the number of categories; p(x ii ) represents the true label of the i-th sample belonging to the j-th category; q(x ii ) represents the probability that the model predicts the i-th sample belongs to the j-th category; The calculation of the joint loss L is as follows: L = ω r L r + ω t L t + ω c L c ; Among them, ω r , ω t , ω c are respectively L r , L t and L c are the loss weights.