Personnel behavior detection and identity recognition method, equipment and storage medium
By improving the YOLOv5 network and InsightFace model, and combining PP-LCNet and CycleGAN, the problem of human behavior detection and identity recognition in dark environments was solved, achieving efficient and accurate nighttime recognition results.
Patent Information
- Application Number
- CN202310924556.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-26
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-07-26
AI Technical Summary
Traditional visible light behavior detection algorithms are difficult to work in dark environments, and traditional face recognition algorithms cannot achieve accurate identity recognition in dark environments, making it difficult to detect and recognize people's behavior at night.
By employing an improved YOLOv5 network model and an optimized InsightFace model, and by constructing a behavior dataset and a dual-light face dataset, combined with the PP-LCNet network and CycleGAN, we can achieve human behavior detection and identity recognition.
It achieves efficient behavior detection and identity recognition in dark environments, improving the model's recognition accuracy and real-time performance while reducing computational costs.
Smart Images

Figure CN117275083B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent recognition technology, and in particular relates to a method, device and storage medium for personnel behavior detection and identity recognition. Background Technology
[0002] Identity recognition and behavioral pose detection are important applications in computer vision, capable of identifying targets and analyzing whether they exhibit abnormal or dangerous behavior. They have wide applications in security, transportation, manufacturing, and healthcare, among many other fields. However, traditional visible light-based algorithms for personnel behavior detection and identity recognition are increasingly unable to meet the demands of complex application environments and are susceptible to environmental factors and inclement weather, failing to function effectively in dark environments. Currently, research on infrared thermal imaging-based personnel behavior detection and identity recognition is scarce, making the achievement of intelligent personnel behavior detection and identity recognition in dark environments a significant challenge. Therefore, researching infrared thermal imaging-based personnel behavior detection and identity recognition technologies has substantial research significance and application value in addressing the difficulties of detection in dark environments.
[0003] Currently, human behavior detection and identification in dark environments mainly face two major challenges:
[0004] First, traditional visible light behavior detection algorithms are difficult to perform in dark environments and are easily affected by complex environments and severe weather. Furthermore, human behavior is diverse, short-term, and irregular, so behavior detection and recognition have high requirements for the real-time performance and reliability of the algorithm.
[0005] Second: Traditional visible light facial recognition algorithms cannot accurately identify people in dark environments. Summary of the Invention
[0006] The purpose of this invention is to provide a method, device and storage medium for personnel behavior detection and identity recognition, so as to solve the problem that traditional algorithms cannot accurately recognize personnel behavior and identity in dark or complex environments.
[0007] This invention solves the above-mentioned technical problems through the following technical solution: a method for personnel behavior detection and identity recognition, the method comprising the following steps:
[0008] A behavior dataset is constructed based on thermal images of human behavior, and a dual-light face dataset is constructed based on corresponding infrared and visible light images of faces.
[0009] The backbone network of the YOLOv5 network model is changed to the PP-LCNet network to obtain the improved YOLOv5 network model; wherein, the PP-LCNet network includes a CBS module, a first depthwise separable convolutional module, a second depthwise separable convolutional module, a third depthwise separable convolutional module, a fourth depthwise separable convolutional module, and a depthwise separable convolutional module with an attention mechanism connected in sequence.
[0010] The improved YOLOv5 network model was trained using the aforementioned behavior dataset to obtain a human behavior detection model;
[0011] By adding CycleGAN before the input layer of the InsightFace optimized model, an improved InsightFace optimized model is obtained. The improved InsightFace optimized model is then trained using the dual-light face dataset to obtain a face recognition model.
[0012] The system acquires a full-body image of a person in real time, uses the person behavior detection model to detect the full-body image of the person, and obtains the behavior detection result; it then uses the face recognition model to recognize the full-body image of the person, and obtains the face recognition result.
[0013] The behavior detection results and face recognition results are fused to obtain the final recognition result.
[0014] Furthermore, the specific construction process of the behavioral dataset includes:
[0015] Thermal imagers were used to capture behavioral thermal images of different people at different times.
[0016] The behavioral thermal images are labeled using annotation tools, and the labeled behavioral thermal images constitute a behavioral dataset.
[0017] The specific construction process of the dual-light face dataset includes:
[0018] Infrared and visible light images of faces of different people with different facial expressions are captured from different angles, and the infrared and visible light images of faces are in one-to-one correspondence.
[0019] The infrared and visible light images of the face are labeled using a labeling tool, and the labeled infrared and visible light images of the face constitute a dual-light face dataset.
[0020] Furthermore, the CBS module consists of convolutional layers, batch normalization (BN) layers, and activation functions;
[0021] The first, second, and third depthwise separable convolutional modules are each composed of two depthwise separable convolutional layers stacked together; the fourth depthwise separable convolutional module is composed of five depthwise separable convolutional layers stacked together.
[0022] The depthwise separable convolutional module with an attention mechanism is composed of two stacked depthwise separable convolutional layers with an attention mechanism.
[0023] Furthermore, the depthwise separable convolutional layer with an attention mechanism includes a depthwise convolution (DW), a pointwise convolution (PW), and an attention mechanism module located between the depthwise convolution (DW) and the pointwise convolution (PW). The attention mechanism module consists of a global average pooling layer, two fully connected layers, and activation functions corresponding to the fully connected layers.
[0024] Furthermore, the 5×5 kernel size deep convolution DW in the fourth depthwise separable convolution module and the depthwise separable convolution module with attention mechanism is replaced with three parallel branches. Each branch consists of a deep convolution DW and a BN layer, and the kernel sizes of the deep convolution DW in each branch are 5×5, 3×3, and 1×1, respectively.
[0025] Furthermore, the ReLU activation function of the depthwise separable convolutional layer in the first depthwise separable convolutional module, the second depthwise separable convolutional module, the third depthwise separable convolutional module, the fourth depthwise separable convolutional module, and the depthwise separable convolutional module with the introduction of the attention mechanism is replaced with the H-Swish activation function.
[0026] Furthermore, during the training of the improved YOLOv5 network model, the specific expression of the loss function is as follows:
[0027]
[0028] Among them, L EIOU The loss value for the improved YOLOv5 network model is given by IoU, where IoU is the alternation ratio between the predicted and ground truth boxes, ρ() is the Euclidean distance calculation function, and b is the center point of the predicted box. gt Let w be the center point of the ground truth bounding box and w be the width of the predicted bounding box. gt h is the width of the ground truth bounding box, and h is the height of the predicted bounding box. gt Let c be the height of the ground truth bounding box, and c be the diagonal length of the smallest closed box that covers both the predicted and ground truth bounding boxes. w To determine the width of the minimum closed box that covers both the predicted and ground truth boxes, c h The height of the minimum closed box that covers both the predicted and ground truth boxes.
[0029] Furthermore, during the training of the improved InsightFace optimized model, the loss function of CycleGAN is added to the loss function of the InsightFace optimized model at a certain ratio, and CycleGAN and the InsightFace optimized model are trained synchronously.
[0030] Furthermore, the loss function of the improved InsightFace optimization model is:
[0031] L=αL3+βL total
[0032]
[0033] L total (G,F,D x D y ) = L GAN (G,D y ,X,Y)+L GAN (G,D x ,Y,X)+λ1L cyc (G,F)+λ2L Identity (G,F)
[0034] Where L is the loss value of the improved InsightFace optimization model, α and β are adjustable scaling coefficients, and L3 is the loss value of the InsightFace optimization model. total Here, θ represents the loss value of CycleGAN, N is the batch size, and θ is the loss value. yi Let θ be the target angle, s be the feature norm, m be the additional side-angle distance, and θ be the angle between the target angles. j Let x be the depth feature of the i-th sample. i and the weights W of the fully connected layer j The angle between y i Let be the actual class of the i-th sample, n be the number of classes of the sample, G be the generator for the forward transformation, F be the generator for the reverse transformation, and D be the generator for the inverse transformation. x For the discriminator of the positive transition, D y For the discriminator of the inverse transformation, X is the input graph in domain A, Y is the input graph in domain B, and L... GAN To counteract the loss, λ1 is the adjustment parameter for the cycle-consistent loss, L cyc The loss is the cycle-consistent loss, λ² is the adjustment parameter for the identity loss, and L... Identity This results in a loss of Identity.
[0035] Based on the same concept, the present invention also provides an electronic device, the device comprising:
[0036] Memory, used to store computer programs;
[0037] A processor is used to implement the personnel behavior detection and identity recognition method as described above when executing the computer program.
[0038] Based on the same concept, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the personnel behavior detection and identity recognition method as described above.
[0039] Beneficial effects
[0040] Compared with the prior art, the advantages of the present invention are as follows:
[0041] This invention improves the backbone network of the YOLOv5 network model by using the PP-LCNet network. The depthwise separable convolutional layer significantly reduces the number of parameters and computational complexity of the model by decomposing the convolution operation, greatly reducing the computational cost while maintaining model performance. At the same time, the introduction of the H-Swish activation function and attention mechanism SE further improves the performance of the model. Furthermore, the use of three parallel branches to replace the depthwise convolution DW enables the model to capture features at different scales, which helps to improve the performance of the model.
[0042] This invention utilizes CycleGAN to improve the InsightFace optimization model, while incorporating the loss function of CycleGAN into the loss function of the InsightFace optimization model at a certain ratio, thus merging the training of the two networks into the training of a single model and improving the model's recognition accuracy. Attached Figure Description
[0043] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only one embodiment of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart of the personnel behavior detection and identity recognition method in an embodiment of the present invention;
[0045] Figure 2 This is a structural block diagram of the improved YOLOv5 network model in this embodiment of the invention;
[0046] Figure 3 This is a diagram of the depth-separable convolutional layer architecture in an embodiment of the present invention;
[0047] Figure 4 This is an architecture diagram of a depthwise separable convolutional layer that incorporates an attention mechanism in an embodiment of the present invention;
[0048] Figure 5 This is a schematic diagram of replacing the depthwise convolution (DW) with three parallel branches in an embodiment of the present invention;
[0049] Figure 6 This is a structural block diagram of the improved InsightFace optimization model in this embodiment of the invention;
[0050] Figure 7 This is a schematic diagram of the cycle consistency principle in an embodiment of the present invention;
[0051] Figure 8 This is a framework diagram of the motion detection system in an embodiment of the present invention;
[0052] Figure 9 This is a loss curve diagram of the control experiment in the embodiments of the present invention;
[0053] Figure 10 This is the average precision (mAP) curve of the control experiment in the embodiments of the present invention;
[0054] Figure 11 This is a schematic diagram of the average accuracy mAP for different categories in the embodiments of the present invention;
[0055] Figure 12 This is a diagram showing the recognition effect of the test set in an embodiment of the present invention;
[0056] Figure 13 This refers to the ROC curve in an embodiment of the present invention;
[0057] Figure 14 This describes the effect of the improved InsightFace optimization model on the test set in this embodiment of the invention. Detailed Implementation
[0058] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0059] The technical solutions of this application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0060] like Figure 1 As shown in the figure, the personnel behavior detection and identity recognition method provided by the embodiment of the present invention includes the following steps:
[0061] Step 1: Construct a behavior dataset based on thermal images of human behavior, and construct a dual-light face dataset based on the corresponding infrared and visible light images of faces;
[0062] Step 2: Replace the backbone network of the YOLOv5 network model with a PP-LCNet network to obtain an improved YOLOv5 network model; train the improved YOLOv5 network model using a behavior dataset to obtain a human behavior detection model.
[0063] Step 3: Add CycleGAN (i.e., recurrent generative adversarial network) before the input layer of the InsightFace optimized model to obtain the improved InsightFace optimized model. Train the improved InsightFace optimized model using the dual-light face dataset to obtain the face recognition model.
[0064] Step 4: Real-time acquisition of overall images of people; use a people behavior detection model to detect the overall images of people and obtain behavior detection results; use a face recognition model to recognize the overall images of people and obtain face recognition results.
[0065] Step 5: Fuse the behavior detection results and face recognition results to obtain the final recognition result.
[0066] In step 1, a JUGE infrared thermal imager was used to collect data from multiple employees in a company's factory workshop. The behavioral dataset contained a total of 11,000 behavioral thermal images, which included nine behavioral categories: playing phone, making phone calls, fighting, smoking, carrying dangerous items, lying down, operating, picking, and processing. The first six categories were considered abnormal or dangerous behaviors, while the last three were considered normal work behaviors.
[0067] 3000 images were captured from different angles, including 1500 infrared images and 1500 visible light images of the faces of 30 volunteers with different facial expressions. The infrared and visible light images correspond to each other, representing a pair of images of the same person and the same facial expression captured from the same angle. Each person's 50 images include variations in expression, angle, and occlusion. Shooting scenarios included both with and without glasses obstructing the view. Facial expressions included normal, smiling, and frowning. Shooting angles included left half-profile, right half-profile, head up, and head down.
[0068] All image data were obtained after video frame extraction, deduplication, filtering, and annotation. The images in the behavior dataset and the two-light face dataset were labeled using the PaddleLabel annotation tool, resulting in XML format label files. The behavior heatmap images in the behavior dataset are 640×512 pixels in size, and the face images in the two-light face dataset are 512×512 pixels in size. Both the behavior dataset and the two-light face dataset are divided into training, validation, and test sets in an 8:1:1 ratio.
[0069] Considering both the real-time performance of behavior detection and the accuracy of identity recognition (i.e., face recognition) in dark environments, this paper proposes a method for human behavior detection and identity recognition in dark environments. This method employs a lightweight network PP-LCNet to improve the YOLOv5 network model and CycleGAN to improve the InsightFace optimization model. The two improved models are then cascaded to form a method for human behavior detection and identity recognition in dark environments. The trained model is then deployed to an embedded mobile device (such as an NVIDIA embedded platform development board or a tablet computer). The Jugo thermal imager SDK is improved to enable the thermal imager to connect to a PC or embedded mobile device. The captured images can be transmitted to the model in real time. By connecting the thermal imager to the embedded mobile device, a convenient and fast mobile detection system is formed, enabling real-time human behavior detection and identity recognition in dark environments.
[0070] The traditional YOLOv5 network model mainly consists of a backbone network, an enhancement feature extraction network (Neck), and a detection layer (Head). The backbone network of the traditional YOLOv5 network model uses CSPDarknet53, combined with a residual network (Residual) and a cross-stage local network (CSPNet) to form the BottleneckCSP module. The enhancement feature extraction network includes a Feature Pyramid Network (FPN) and a Path Aggregation Network (PAN). The FPN fuses high-level features with low-level features through upsampling from top to bottom, while the PAN further fuses and enhances these features from bottom to top. The detection layer uses the YOLO Head to predict the location, category, and probability of an object. Figure 2 As shown, the improved YOLOv5 network model of this invention is also divided into a backbone network, an enhanced feature extraction network (Neck), and a detection layer (Head). The backbone network adopts the lightweight PP-LCNet network, while the enhanced feature extraction network and detection layer still adopt the enhanced feature extraction network and detection layer of the traditional YOLOv5 network model. The improved YOLOv5 network model is trained using a behavior dataset, and the model parameters are optimized to obtain a human behavior detection model with excellent performance.
[0071] PP-LCNet (i.e., a lightweight CPU network based on the MKLDNN acceleration strategy) is a lightweight convolutional neural network that uses the depthwise separable convolutional layer DepthSepConv proposed in MobileNetv1 as its basic block. The DepthSepConv layer consists of depthwise convolutions (DW) and pointwise convolutions (PW), as follows: Figure 3 As shown. Figure 2 As shown, the PP-LCNet network includes a CBS module, a first depthwise separable convolutional module, a second depthwise separable convolutional module, a third depthwise separable convolutional module, a fourth depthwise separable convolutional module, and a depthwise separable convolutional module with an attention mechanism, all connected in sequence. The CBS module consists of a Convolutional layer (Conv), a Batch Normalization (BN) layer, and a SiLU activation function. The first, second, and third depthwise separable convolutional modules are each composed of two stacked depthwise separable convolutional layers with a kernel size of 3×3. The fourth depthwise separable convolutional module is composed of five stacked depthwise separable convolutional layers with a kernel size of 5×5. The depthwise separable convolutional module with an attention mechanism is composed of two stacked depthwise separable convolutional layers with a kernel size of 5×5 and an attention mechanism (SE).
[0072] like Figure 4 As shown, the depthwise separable convolutional layer with an attention mechanism includes a depthwise convolution (DW), a pointwise convolution (PW), and an attention mechanism module (SE) placed between the DW and PW. The attention mechanism module SE consists of a global average pooling layer (GAP), two fully connected layers (FC), and the corresponding activation functions ReLU and Sigmoid for the FC. Placing the attention mechanism module SE in the last two DepthSepConv layers effectively improves the model's feature extraction capability without increasing inference time. The specific principle of introducing the attention mechanism module SE into the DepthSepConv layer is as follows: the attention mechanism module SE is placed between DW and PW. The output of DW is fed into the attention mechanism module SE. First, global average pooling is performed on the input feature layer, followed by two fully connected layers. The first fully connected layer has fewer neurons, while the second fully connected layer has the same number of neurons as the input feature layer. After completing the two fully connected layers, a Sigmoid function is applied to fix the value between 0 and 1. At this point, the weight (between 0 and 1) of each channel of the input feature layer is obtained. This weight is then multiplied by the original input feature layer.
[0073] Both the fourth depthwise separable convolutional module and the depthwise separable convolutional module with an attention mechanism consist of a 5×5 DepSepConv layer. This invention further improves upon this by introducing the concept of structural reparameterization: [e.g.] Figure 5As shown, the 5×5 kernel size DW in the fourth depthwise separable convolutional module and the depthwise separable convolutional module with attention mechanism is replaced with three parallel branches. Each branch consists of a DW and a Batch Normalization (BN) layer. The kernel sizes of the DWs in each branch are 5×5, 3×3, and 1×1, respectively. For each DepSepConv layer, the DWs with kernel sizes of 5×5, 3×3, and 1×1 are combined, that is, the outputs of the three branches are added channel by channel, and the number of output channels is consistent with the original 5×5 kernel size DW. This allows the model to capture features at different scales and helps improve model performance. During the inference stage, to avoid affecting model efficiency, the DWs with different kernel sizes in each layer are combined into a single 5×5 kernel size DW, which can reduce the computational cost during inference while maintaining model performance. To further reduce computational and memory requirements, the Data Wrapper (DW) is fused with a Batch Normalized (BN) layer. This is achieved by incorporating the parameters of the BN layer into the weights of the DW, which can be considered as lossless compression. This fusion can maintain model performance while reducing computational costs.
[0074] The input-output mapping relationship of DW is as follows:
[0075] y dw-conv =W·x in +B (1)
[0076] The input-output mapping relationship of a batch normalized BN layer can be represented as follows:
[0077]
[0078] The input-output mapping relationship after fusing the DW and the batch normalized BN layer is represented as follows:
[0079]
[0080] Where, x in and y dw-conv Let W and B represent the input and output of the DW, respectively. Let W and B represent the weights and biases of the DW, respectively. Let μ and σ represent the mean and variance of the input for a given batch size, respectively. Let γ and β represent the normalization coefficients used for scaling and translation, respectively. out The output after passing through the DW and Batch Normalized BN layers, w fuse and b fuse These represent the weights and biases after fusing the DW and Batch Normalized BN layers, respectively.
[0081] After the improvement, the model outputs feature maps of three feature layers with dimensions of 80*80*(C+4+1), 40*40*(C+4+1), and 20*20*(C+4+1), respectively, which means the feature map depth is C class parameters, 4 location parameters, and 1 confidence parameter.
[0082] The ReLU activation function in the depthwise separable convolutional layers of the first, second, third, and fourth depthwise separable convolutional modules, as well as the depthwise separable convolutional module with an attention mechanism, is replaced with the H-Swish activation function. Using the H-Swish activation function can achieve better network performance without increasing inference time by replacing the ReLU activation function. Its expression is:
[0083]
[0084] The traditional YOLOv5 network model uses the CIOU loss function, which adds an aspect ratio loss to the predicted and actual bounding boxes based on DIOU, making the predicted and actual bounding boxes more consistent. However, the aspect ratio described by CIOU has some ambiguity and does not consider the balance of hard samples. CIOU It can be represented as:
[0085] L CIOU =1-IoU+ρ 2 (b,b gt ) / c 2 +av (5)
[0086]
[0087] Where, 'a' represents the balance ratio coefficient, 'v' measures the proportional consistency between the width and height of the predicted bounding box and the ground truth bounding box, 'ρ' represents the Euclidean distance between the center points of the predicted and ground truth bounding boxes, 'c' represents the diagonal length of the minimum closed box covering both the predicted and ground truth bounding boxes, and 'b' and 'b' represent the values of 'a', 'v ... gt Let w and h represent the center points of the predicted bounding box and the ground truth bounding box, respectively. Let w and h represent the width and height of the predicted bounding box, respectively. gt h gt These represent the width and height of the ground truth bounding box, respectively, and IoU represents the intersection-union ratio between the predicted bounding box and the ground truth bounding box.
[0088] Multiple experiments have demonstrated that, compared to equations (5) and (6), the EIOU loss function exhibits the best performance. It considers the overlap area, center point distance, and the true differences in length and width, resolving the fuzzy definition of aspect ratio, and incorporates Focal Loss to address sample imbalance in bounding box regression. The EIOU loss function can be expressed as:
[0089]
[0090] Among them, c w and c h These represent the width and height of the smallest closed box (or outer box) that covers the ground truth box and the predicted box, respectively.
[0091] When training the improved YOLOv5 network model, the training parameters are: image size 640, training times 100, batch size 128, learning rate 0.001, optimizer Adam, learning rate adjustment strategy CosineAnnealingLR, and loss function as shown in equation (7).
[0092] Compared to traditional convolutions, depthwise separable convolutional layers can significantly reduce the number of model parameters and computational complexity by decomposing the convolution operation, greatly reducing computational costs while maintaining similar performance. Furthermore, the introduction of the H-Swish activation function and the SE module further improves network performance. Building upon this, the concept of structural reparameterization (Rep) is introduced to improve the PP-LCNet network, simplifying the multi-branch structure used in the training phase to a unidirectional structure in the inference phase, thus achieving the fusion of convolutional layers and batch normalization layers. The improved YOLOv5 network model is shown below. Figure 2 As shown, Inputs represents the input image; CBS is an abbreviation for convolutional layers, BN layers, and the SiLU activation function; DepthSepConv represents a depthwise separable convolutional layer; DepthSepConv_SE represents a depthwise separable convolutional layer with an attention mechanism module SE; C3 represents a module consisting of three convolutional layers (Conv), with the first Conv having a stride of 2 and the second and third Convs having a stride of 1; Upsample represents an upsampling layer that enlarges the image size; Concat is an operation that concatenates two or more tensors along a certain dimension; GAP represents a global average pooling layer; FC represents a fully connected layer; ReLU, Sigmoid, and H-swish are three different activation functions; Stage represents the components of PP-LCNet, consisting of five parts, or five stages; ×2 and ×5 next to DepthSepConv represent the number of DepthSepConv layers. Figure 2The training process of the improved YOLOv5 network model is as follows: First, the image is input into PP-LCNet. After passing through the middle layer, lower middle layer, and bottom layer of PP-LCNet, i.e., the third to fifth stages, feature layers are output respectively. The shapes of the three feature layers (i.e., feature layer width, feature layer height, and number of feature layer channels) are (80,80,125), (40,40,256), and (20,20,512) respectively. After obtaining the three feature layers, the pyramid network FPN and path aggregation network PAN are constructed using them. The three prediction feature layer shapes (i.e., feature layer width, feature layer height, and number of feature layer channels) are obtained through the YOLO Head detection layer. They are (80,80,14), (40,40,14), and (20,20,14) respectively. The number of feature layer channels can be divided into 9+1+4, i.e., 9 class prediction probability parameters, 4 location prediction parameters, and 1 target confidence prediction parameter.
[0093] The traditional InsightFace optimization model, also known as ArcFace, uses a deep convolutional neural network (DCNN) to embed a face representation. The dot product between the DCNN features and the last fully connected layer is equal to the cosine distance between the features and the weights after normalization. The traditional InsightFace optimization model includes a feature extraction network, a loss layer, and a classification layer. The input image data is first processed by a ResNet feature extraction network to extract features, which are then normalized and max-pooled. After that, the features are fed into the loss layer to calculate the ArcFace Loss function. Finally, the resulting cross-entropy loss is fed into the classification layer to classify different faces.
[0094] This invention adds a CycleGAN before the input layer of the traditional InsightFace optimization model. The input is changed from a visible light image to an infrared image and a visible light image of the face. The infrared image is converted into a visible light image by CycleGAN and then input into the InsightFace optimization model. The visible light image is converted into an infrared image by CycleGAN. Subsequent processing is consistent with the traditional InsightFace optimization model, namely feature extraction, loss calculation, and face classification. Figure 6 As shown, the improved InsightFace optimized model includes CycleGAN, a feature extraction network, a loss layer, and a classification layer.
[0095] The InsightFace optimization model uses the inverse cosine function to calculate the angle between the current feature and the target weight, then adds an additional corner distance to the target angle, and finally obtains the new target logit using the cosine function. Subsequent steps are completely consistent with the Softmax loss. CycleGAN is a generative adversarial network that implements image style transfer. It consists of two generators and two discriminators, forming a ring structure. The purpose of CycleGAN is to complete style transfer between two domains (domain A and domain B) while ensuring that the geometry and spatial relationships of objects in the image remain unchanged. CycleGAN can be used to convert infrared images to visible light images and visible light images to infrared images. By adding CycleGAN to the front end of the InsightFace optimization model and simplifying the training process by incorporating the CycleGAN loss function into the InsightFace optimization model's loss function in a certain proportion, the two training steps are merged into one, enabling face recognition while converting the input infrared image to a visible light image.
[0096] The traditional InsightFace optimization model uses a deep convolutional neural network to embed a face representation, comprising three parts: a feature extraction network, a loss layer, and a classification layer. The feature extraction network includes a deep convolutional neural network (DCNN) and subsequent feature normalization operations. The loss layer involves max pooling, inverse cosine calculation, and adding an additional edge distance *m* to the cosine similarity of the subclasses obtained by the inner product of the normalized features and the weights of the normalized subcenters. This is then converted to cosine, rescaled using the feature norm *s*, and finally obtained through a softmax loss function to achieve cross-entropy loss. The ultimate goal of the loss layer is to obtain ArcFace Loss. The classification layer uses cross-entropy loss to classify face images.
[0097] according to Figure 6 The process of optimizing the InsightFace model is as follows: The DCNN extracts features from the visible light image to obtain embedded features x, which are then normalized. The inner product of the normalized embedded features and the weights W (also called normalized subcenters) of the last fully connected layer of the DCNN yields the cosine distance (i.e., subclass cosine similarity) between the normalized features and the normalized weights. Subsequently, max pooling is performed on the obtained subclass cosine similarity to obtain the cosine similarity of the classes.
[0098] Let the depth feature x of the i-th sample be... i The true category is the yth i Classes, then:
[0099]
[0100] pass Calculate the target angle Then add the custom additional corner spacing m to Obtain the adjusted target angle Adjusted target angle Calculate the cosine value to obtain the new target Logit (Logit is the predicted value of a batch), i.e. Then, all Logit (except the target Logit) are rescaled using the feature norm s. In addition, the remaining original Logit remains cosθ. j We obtain the new Logit, i.e., s*cosθ j Finally, the Softmax Loss is calculated for the new Logit. This involves first converting the predicted value into a probability value using the Softmax activation function, then encoding it using One-Hot, and finally calculating the cross-entropy loss. The form of the cross-entropy loss is... (f represents a function). Cross-entropy loss is often used in single-label multi-class classification tasks. In the InsightFace optimization model, it is used to classify multiple faces.
[0101] The loss function of the InsightFace optimization model is derived from Softmax. The loss function of Softmax is as follows:
[0102]
[0103] Where, x i ∈R d This represents the depth feature of the i-th sample, which actually belongs to the y-th sample. i Class; Let the embedding feature dimension d = 512, W j ∈R d The weights W∈R of the fully connected layer are represented as follows: d×n The j-th column, b j ∈R n This represents the bias term; the batch size and the number of categories are N and n, respectively. For ease of representation, the bias b is fixed. j =0, convert Logit to Where, θ j It is a depth feature x i and weight W j The included angle is fixed by the L2 norm ||W j ||and embedded features||x i ||, and rescale the Logit to s. The normalization steps on the features and weights make the prediction depend only on the angle θ between the features and weights. j The loss function is as follows:
[0104]
[0105] Given that the embedded features are distributed around each feature center in the hypersphere, and also in the depth feature x i and target weight An additional side-angle distance m was added between them to obtain This further enhances intra-class compactness and inter-class diversity. The resulting loss function for the InsightFace optimized model is:
[0106]
[0107] CycleGAN is a deep learning-based style transfer network that can be trained without paired datasets. The key idea behind CycleGAN is to use cycle consistency to ensure that the generator's output maintains content similarity to the original image. Figure 7 As shown, cycle consistency can be represented as follows: given two images x and y with different style domains, image x can be used to obtain y through generator GAB, and then to obtain image x' with the same domain as A through GBA. If x and x' are consistent, it is equivalent to making the images cycle and consistent.
[0108] CycleGAN's loss function L total From the counter-loss L GAN Cyclic consistency loss L cyc And Identity loss L Identity It consists of three parts. Countermeasures against loss L GAN This is the game-theoretic loss between the generator and the discriminator, essentially consistent with the adversarial losses described in other GAN networks. Its purpose, like traditional GANs, is to allow the generator to deceive the discriminator as much as possible, while the discriminator tries to expose the generator's flaws. Cyclic consistency loss L cyc The calculation is based on the image obtained after a loop and the original image. The goal is to make the image as similar as possible to the original image after passing through domain A to domain B and then back to domain A, resulting in a more natural and realistic generated image. Identity loss L Identity This ensures that the input and output remain consistent to a certain extent, preserving as many features of the original image as possible. The loss function of CycleGAN can be expressed as:
[0109] L total (G,F,D x D y ) = L GAN (G,D y ,X,Y)+L GAN (G,D x ,Y,X)+λ1L cyc(G,F)+λ2L Identity (G,F)(12)
[0110] Among them, L GAN This represents adversarial losses, including both A2B and B2A adversarial losses; L cyc For cycle-consistent loss, L Identity For Identity loss; G and D x The generator and discriminator for the A2B forward conversion are F and D, respectively. y X is the generator and discriminator for the B2A inverse transformation, X is the input graph in the A domain, Y is the input in the B domain, λ1 is the adjustment parameter for the cycle consistency loss, and λ2 is the adjustment parameter for the identity loss.
[0111] The improved InsightFace optimization model adds a modality transformation module (GAB) trained by CycleGAN to the input of the traditional InsightFace optimization model, converting the input infrared image into a visible light image. Then, a visible light face recognition step is performed, where features are extracted and normalized from the obtained visible light image to obtain normalized embedded features x. These x are then multiplied by the normalized weights W to obtain the subclass cosine similarity (cosθ). j Then, max pooling is used to obtain the cosine similarity. The obtained cosine similarity is inverse cosine (arccos) and an additional side-angle distance m is added. Then, the feature norm s is used to rescale s*cosθ. j Finally, the Softmax Loss is calculated for the new Logit. This involves first converting the predicted value into a probability value using the Softmax activation function, then encoding it using One-Hot, and finally calculating the cross-entropy loss. The cross-entropy loss is input into the classification layer to classify faces based on different facial features, thereby indirectly achieving infrared face recognition. The loss function of the improved InsightFace optimized model can be expressed as:
[0112] L=αL3+βL total (13)
[0113] Where L3 is the loss value of the InsightFace optimization model; α and β are both adjustable scaling coefficients, and their loss functions are calculated according to a certain ratio.
[0114] The training parameters for the improved InsightFace optimized model are: image size 512×512, training iterations 100, batch size 128, learning rate 0.001, optimizer Adam, classifier LargeScaleClassifier, and loss function ArcFace+GANLoss.
[0115] After training the improved YOLOv5 network model and the improved InsightFace optimized model, the model was tested and evaluated using the corresponding test set to verify the model performance and obtain the human behavior detection model and the face recognition model.
[0116] Deploy personnel behavior detection and face recognition models to embedded mobile devices (such as NVIDIA embedded platform development boards) or PCs. Improve the JUGEE thermal imager SDK based on a multi-threaded architecture to connect the thermal imager to the PC or embedded mobile device. This allows for real-time transmission of captured images, importing them into the personnel behavior detection model for behavior detection, and simultaneously using the face recognition model to detect facial images (the images must meet clarity requirements, i.e., the detected facial images must be at least 512*512 resolution to proceed with face recognition). Then, perform cross-modal identity recognition and display the results. Connect the thermal imager to the embedded development board and a mobile display (tablet) to create a convenient and efficient mobile detection system. Figure 8 As shown.
[0117] The specific operation process is as follows: First, the TensorRT deployment tool is used to deploy the personnel behavior detection model and the face recognition model to the NVIDIA embedded development board. The model weights are converted into .wts files, and then compiled to generate an .engine file for inference. The mobile detection system of this invention includes an infrared thermal imager module, a data transmission module, a data parsing and processing module, a detection and recognition module, and a mobile display module. In the infrared thermal imager module, real-time digital thermal image data output is achieved through secondary development of the original SDK. First, the infrared thermal imager device information is obtained: through the interface provided by the SDK, the device information of the infrared thermal imager, such as model and serial number, is obtained; then, the infrared thermal imager is initialized according to the device information to ensure its normal operation. The output data of the infrared thermal imager is captured in real time through the interface provided by the SDK and digitized, and the captured digital thermal image data is output to the data transmission module. The data transmission module uses a wired method to transmit the digital thermal image data output by the infrared thermal imager to a PC or embedded mobile device in real time. The data parsing and processing module receives digitized thermal image data from the data transmission module in real time. It parses the received data, extracting pixel values and temperature information from the thermal image, and reconstructs the thermal image based on this information. This data is then imported into the personnel behavior detection model and the face recognition model in real time. Real-time detection and recognition are achieved using a multi-threaded architecture: thread 1 transmits the real-time data collected by the infrared thermal imager to a PC or embedded mobile device; thread 2 imports the collected data into the personnel behavior detection model for behavior detection and recognition; and thread 3 imports the collected data into the face recognition model for identity verification. Finally, the recognition results are output to the mobile display module for real-time display.
[0118] The improved InsightFace optimized model is trained and optimized on a dual-light face dataset to obtain an infrared face recognition model with high recognition accuracy. The human behavior detection model and face recognition model are deployed on an NVIDIA embedded development board via TensorRT, and real-time behavior detection and face recognition are achieved based on a multi-threaded architecture.
[0119] To better verify the superiority of the improved YOLOv5 network model of this invention, a comparative experiment was conducted on the original YOLOv5 (CSPDarkNet) and YOLOv5 backbone networks using PP-LCNet, MobileNetv3, ShuffleNetv2, GhostNet, and EfficientNet under the same parameter settings. The training parameters were completely consistent with the settings in step 2 of the training process, and the loss curve during the training process is shown in Figure 1. Figure 9 As shown, the average accuracy (mAP) curve is as follows: Figure 10 As shown, epoch represents the number of training iterations.
[0120] from Figure 9 and Figure 10 As can be seen, the improved YOLOv5 network model of this invention exhibits better performance. With the increase of training iterations, both the model loss value and mean precision (mAP) are superior to other networks. The performance of each model was quantitatively evaluated using the following metrics on the test set: number of parameters, mean precision (mAP), and detection frame rate (FPS). The number of parameters indicates the complexity of the model; a larger number of parameters results in greater computational cost and memory consumption. Mean precision (mAP) specifically refers to the average precision at different recall rates, measuring the model's detection performance across all categories. Detection frame rate (FPS) represents the number of image frames the model can process per second, and can be used to measure the model's real-time detection capability. The performance metric comparison is shown in Table 1.
[0121] Table 1 Performance Indicator Comparison Table
[0122]
[0123]
[0124] As shown in Table 1, compared with the traditional YOLOv5 model, this invention reduces the number of parameters by 56.4%, reduces model training time by 60.6%, and improves average accuracy and inference speed by 5.6% and 32.2%, respectively, demonstrating excellent performance. Compared with other lightweight network improvement algorithms, this invention outperforms other models in terms of recognition accuracy, model complexity, and inference speed, making it very suitable for deployment in environments using NVIDIA development boards as hardware platforms. The detection accuracy for each category of data is as follows: Figure 11 As shown, by Figure 11 It can be seen that the working behavior processing and operation category detection accuracy of the present invention reaches a maximum of 99.5%, the smoking category detection accuracy is a minimum of 80.9%, and the average accuracy reaches 94.7%, which is sufficient to support the actual engineering testing needs.
[0125] The test set was tested, and the recognition results were as follows: Figure 12 As shown, from Figure 12 It is evident that the present invention can detect and identify personnel behavior using infrared thermal imaging, and has a good identification effect on personnel working behavior, abnormal and dangerous behavior under infrared conditions, indicating that the present invention can realize the function of personnel behavior detection in dark environments.
[0126] To verify the superiority of the improved InsightFace optimization model of this invention, the original InsightFace infrared model and InsightFace visible light model were trained using infrared face images and visible light images in dark environments, respectively. These models were used to verify the recognition performance of the original InsightFace model for infrared and visible light faces in dark environments. Simultaneously, to demonstrate the superiority of the improved InsightFace optimization model in dark environments, it was trained using a visible light face dataset under the same parameters. Additionally, an experiment was set up to recognize well-lit visible light images using the InsightFace visible light model as a control group.
[0127] The performance of each model was quantitatively evaluated using the following metrics on the test set: True Positive Rate (TPR), False Positive Rate (FPR), False Acceptance Rate (FAR), and Accuracy (hereinafter referred to as Acc). The results of classification algorithms are typically represented by a confusion matrix, which includes four categories: True Positive Examples (TP), True Negative Examples (TN), False Positive Examples (FP), and False Negative Examples (FN). TPR represents the ratio of actual positives predicted as positive; FPR represents the ratio of actual negatives predicted as positive; and FAR represents the false acceptance rate, i.e., the proportion of images of different people that are mistakenly identified as belonging to the same person—a lower FAR is better. Accuracy represents the proportion of correctly predicted samples out of all samples. Table 2 shows the performance metrics obtained during testing of different models with the same threshold.
[0128] Table 2. Performance indicators obtained during the testing of different models.
[0129]
[0130]
[0131] As shown in Table 2, compared to the original InsightFace model, the improved InsightFace optimized model significantly improves its recognition accuracy in infrared images and visible light images in dark environments, approaching the recognition accuracy of the control group in well-lit environments. A higher True Positive Rate (TPR) and a lower False Positive Rate (FPR) indicate better classification algorithm performance. The improved InsightFace optimized model significantly outperforms the original infrared model and the visible light / dark model, and its False Acceptance Rate (FAR) is also significantly reduced. A graph showing the TPR and FPR values at different thresholds, with FPR (False Positive Rate) on the x-axis and TPR (True Positive Rate) on the y-axis, yields the ROC (Receiver Operating Characteristic) curve, as shown below. Figure 13 As shown. The closer the ROC curve is to the top left corner, the better the classification algorithm's performance; the area under the curve and the coordinate axes (AUC) is used to measure the algorithm's performance, with a larger AUC indicating better performance. Figure 13 As can be seen, the AUC of the improved InsightFace model is 0.999, which is greater than that of the original InsightFace model. The improved InsightFace model has achieved a significant performance improvement, approaching the recognition effect in well-lit environments, and can meet the needs of engineering environments for face recognition in dark environments.
[0132] The improved InsightFace optimized model performs as follows on the test set: Figure 14 As shown, the recognition process can be broken down into two steps: first, converting the infrared image into a visible light image, and then performing face recognition. The improved InsightFace optimized model successfully solved the problem of poor recognition performance of InsightFace in dark environments, improving the accuracy of InsightFace for infrared face recognition.
[0133] The traditional InsightFace optimized model can achieve a recognition speed of 125 FPS, while the improved InsightFace optimized model of this invention has a reduced recognition speed of 65 FPS due to the addition of CycleGAN. This significantly improves the accuracy of face recognition in dark environments while still meeting the real-time requirements of industrial applications.
[0134] To address the problem of dark environment detection, a dataset of multiple human behaviors using infrared thermal imaging and a dataset of dual-light faces were constructed. Data acquisition was performed using a JUGU infrared thermal imager, supporting data saving in image or video format. For infrared human behavior detection, PaddleLabel annotation was used to annotate bounding boxes. Similarly, for dual-light faces, PaddleLabel was used to annotate and crop the faces, forming a dual-light face dataset. For intelligent detection of multiple human behaviors, the backbone network of the YOLOv5 network model was improved using the PP-LCNet network to obtain the lightweight detection network PPLCNet-YOLOv5. Model training and network optimization were performed on the infrared human behavior dataset, resulting in an infrared human behavior detection model with excellent performance. For intelligent multi-person face recognition, the InsightFace optimized model was improved using the CycleGAN recurrent generative adversarial network. Model training and network optimization were performed on the dual-light face dataset, resulting in an infrared face recognition model with high recognition accuracy.
[0135] This system implements real-time behavior detection and face recognition using a multi-threaded architecture. Thread 1 transmits real-time data acquired by an infrared thermal imager to a PC or embedded mobile device. Thread 2 imports this data into a behavior detection model for behavior detection and recognition. Thread 3 imports the data into a face recognition model for identity verification. The trained behavior detection and face recognition models are then deployed to an NVIDIA embedded development board using TensorRT, enabling behavior detection and face recognition on a mobile embedded device.
[0136] This invention utilizes a YOLOv5 network model improved by PP-LCNet and an InsightFace optimized model improved by CycleGAN to construct a method for personnel behavior detection and identity recognition. The trained model is deployed to an embedded mobile device (NVIDIA embedded platform development board). The domestic Jugo thermal imager SDK is improved to enable the thermal imager to connect to a PC or embedded mobile device, and the captured images can be transmitted and imported into the above model in real time. By connecting the thermal imager to the embedded development board and a mobile display screen (tablet computer), a convenient and fast mobile detection system is formed, realizing real-time personnel behavior detection and identity recognition in dark environments.
[0137] This invention also provides an electronic device, which includes a processor and a memory storing a computer program, wherein the processor is configured to execute the computer program to implement the personnel behavior detection and identity recognition method as described above.
[0138] Although not shown, the electronic device includes a processor that can perform various appropriate operations and processes based on programs and / or data stored in read-only memory (ROM) or loaded from a storage portion into random access memory (RAM). The processor can be a multi-core processor or may include multiple processors. In some embodiments, the processor may include a general-purpose main processor and one or more specialized coprocessors, such as a central processing unit, graphics processing unit (GPU), neural network processor (NPU), digital signal processor (DSP), etc. Various programs and data required for the operation of the electronic device are also stored in the RAM. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0139] The processor and memory described above are used together to execute programs stored in the memory. When the program is executed by a computer, it can implement the methods, steps, or functions described in the above embodiments.
[0140] Although not shown, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the personnel behavior detection and identity recognition method as described above.
[0141] Storage media in embodiments of the present invention include articles that are permanent or non-permanent, removable or non-removable, and can store information by any method or technology. Examples of storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information that can be accessed by a computing device.
[0142] The above description only discloses specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or modifications that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for personnel behavior detection and identity recognition, characterized in that, The method includes the following steps: A behavior dataset is constructed based on thermal images of human behavior, and a dual-light face dataset is constructed based on corresponding infrared and visible light images of faces. The backbone network of the YOLOv5 network model is changed to the PP-LCNet network to obtain the improved YOLOv5 network model; wherein, the PP-LCNet network includes a CBS module, a first depthwise separable convolutional module, a second depthwise separable convolutional module, a third depthwise separable convolutional module, a fourth depthwise separable convolutional module, and a depthwise separable convolutional module with an attention mechanism connected in sequence. The improved YOLOv5 network model was trained using the aforementioned behavior dataset to obtain a human behavior detection model; By adding CycleGAN before the input layer of the InsightFace optimized model, an improved InsightFace optimized model is obtained. The improved InsightFace optimized model is then trained using the dual-light face dataset to obtain a face recognition model. The system acquires a full-body image of a person in real time, uses the person behavior detection model to detect the full-body image of the person, and obtains the behavior detection result; it then uses the face recognition model to recognize the full-body image of the person, and obtains the face recognition result. The behavior detection results and face recognition results are fused to obtain the final recognition result; During the training of the improved InsightFace optimized model, the loss function of CycleGAN is added to the loss function of the InsightFace optimized model at a certain ratio, and the CycleGAN and InsightFace optimized models are trained synchronously. The loss function of the improved InsightFace optimized model is as follows: L=αL3+βL total L total (G,F,D x ,D y )=L GAN (G,D y ,X,Y)+L GAN (G,D x ,Y,X)+λ1L cyc (G,F)+λ2L Identity (G,F) Where L is the loss value of the improved InsightFace optimization model, α and β are adjustable scaling coefficients, and L3 is the loss value of the InsightFace optimization model. total Here, θ represents the loss value of CycleGAN, N is the batch size, and θ is the loss value. yi Let θ be the target angle, s be the feature norm, m be the additional side-angle distance, and θ be the angle between the target angles. j Let x be the depth feature of the i-th sample. i and the weights W of the fully connected layer j The angle between y i Let be the actual class of the i-th sample, n be the number of classes of the sample, G be the generator for the forward transformation, F be the generator for the reverse transformation, and D be the generator for the inverse transformation. x For the discriminator of the positive transition, D y For the discriminator of the inverse transformation, X is the input graph in domain A, Y is the input graph in domain B, and L... GAN To counteract the loss, λ1 is the adjustment parameter for the cycle-consistent loss, L cyc The loss is the cycle-consistent loss, λ² is the adjustment parameter for the identity loss, and L... Identity This results in a loss of Identity.
2. The personnel behavior detection and identity recognition method according to claim 1, characterized in that, The specific construction process of the behavioral dataset includes: Thermal imagers were used to capture behavioral thermal images of different people at different times. The behavioral thermal images are labeled using annotation tools, and the labeled behavioral thermal images constitute a behavioral dataset. The specific construction process of the dual-light face dataset includes: Infrared and visible light images of faces of different people with different facial expressions are captured from different angles, and the infrared and visible light images of faces are in one-to-one correspondence. The infrared and visible light images of the face are labeled using a labeling tool, and the labeled infrared and visible light images of the face constitute a dual-light face dataset.
3. The method for personnel behavior detection and identity recognition according to claim 1, characterized in that, The CBS module consists of convolutional layers, batch normalization layers, and activation functions. The first, second, and third depthwise separable convolutional modules are each composed of two depthwise separable convolutional layers stacked together; the fourth depthwise separable convolutional module is composed of five depthwise separable convolutional layers stacked together. The depthwise separable convolutional module with an attention mechanism is composed of two stacked depthwise separable convolutional layers with an attention mechanism.
4. The personnel behavior detection and identity recognition method according to claim 3, characterized in that, The depthwise separable convolutional layer with an attention mechanism includes a depthwise convolution (DW), a pointwise convolution (PW), and an attention mechanism module located between the DW and the PW. The attention mechanism module consists of a global average pooling layer, two fully connected layers, and activation functions corresponding to the fully connected layers.
5. The method for personnel behavior detection and identity recognition according to claim 1, characterized in that, The 5×5 kernel size of the depthwise separable convolutional module and the depthwise separable convolutional module with attention mechanism in the fourth depthwise separable convolutional module are replaced with three parallel branches. Each branch consists of a depthwise convolutional DW and a BN layer. The kernel sizes of the depthwise convolutional DW in each branch are 5×5, 3×3, and 1×1, respectively.
6. The method for personnel behavior detection and identity recognition according to claim 1, characterized in that, Replace the ReLU activation function of the depthwise separable convolutional layer in the first depthwise separable convolutional module, the second depthwise separable convolutional module, the third depthwise separable convolutional module, the fourth depthwise separable convolutional module, and the depthwise separable convolutional module with an attention mechanism with the H-Swish activation function.
7. The method for personnel behavior detection and identity recognition according to any one of claims 1 to 6, characterized in that, During the training of the improved YOLOv5 network model, the specific expression of the loss function is as follows: Among them, L EIOU The loss value for the improved YOLOv5 network model is given by IoU, where IoU is the alternation ratio between the predicted and ground truth boxes, ρ() is the Euclidean distance calculation function, and b is the center point of the predicted box. gt Let w be the center point of the ground truth bounding box and w be the width of the predicted bounding box. gt h is the width of the ground truth bounding box, and h is the height of the predicted bounding box. gt Let c be the height of the ground truth bounding box, and c be the diagonal length of the smallest closed box that covers both the predicted and ground truth bounding boxes. w To determine the width of the minimum closed box that covers both the predicted and ground truth boxes, c h The height of the minimum closed box that covers both the predicted and ground truth boxes.
8. An electronic device, characterized in that, The device includes: Memory, used to store computer programs; A processor is configured to implement the personnel behavior detection and identity recognition method as described in any one of claims 1 to 7 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the personnel behavior detection and identity recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Image processing method, and deep learning model training method and device
CN116206131A
Thermal imaging personnel identity recognition method, terminal equipment and storage medium
CN116343310A