Face recognition detection method based on improved YOLOv5 algorithm
By improving the YOLOv5 algorithm, using the depth separation convolution and SENet attention mechanism, the problem of poor detection of small faces and tilted faces is solved, and high-precision, fast and robust face recognition detection is achieved.
Patent Information
- Application Number
- CN202510028062.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
AI Technical Summary
The original YOLOv5 algorithm has poor detection effect on small or tilted faces in face recognition tasks, and may cause false detection or missed detection in complex scenarios.
By improving the YOLOv5 algorithm, deep separable convolution is used to replace the ordinary convolution of CBL layer in the backbone network, and a SENet attention mechanism is added before the SPPF module of the YOLOv5 backbone network to enhance feature representation capabilities.
It realizes efficient and real-time face recognition detection, has high detection accuracy and robustness, can adapt to complex scenarios, and significantly improves the accuracy and generalization capabilities of the model.
Smart Images

Figure CN119964216A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a face recognition detection method based on an improved YOLOv5 algorithm, and belongs to the technical field of face recognition. Background Art
[0002] With the rapid development of artificial intelligence and Internet of Things related technologies, face recognition technology has become a research topic that has attracted much attention in the field of computer vision, and has broad application prospects in all walks of life. Face verification technology is a method of comparing two face images to see if they are of the same person, which is common in identity authentication scenarios; face search is also called face recognition, which searches for images that match the face to be matched in a specific database to determine the identity of the person to be identified. In power plants in the power industry, face recognition technology strengthens security management, accurately identifies people entering and leaving, and prevents illegal intrusions; at the same time, it improves operational efficiency and automates the management of attendance and access control; in addition, the technology can also realize intelligent recognition and alarm, and combined with mobile applications, it provides a solid technical guarantee for the safe and efficient operation of power plants.
[0003] Given that the original YOLOv5 algorithm still has some shortcomings when it comes to specific tasks of face recognition, for example, the detection effect on small faces or tilted faces is not ideal, and there may be false detection or missed detection in some complex scenes. Therefore, researchers began to explore improvements to the YOLOv5 algorithm to improve its accuracy and robustness in face recognition tasks. Therefore, it is necessary to propose a new face recognition method with high accuracy, fast detection speed, strong robustness and small model data volume to better perform face recognition detection. Summary of the invention
[0004] The present invention proposes a face recognition detection method based on an improved YOLOv5 algorithm to solve the problems that the original YOLOv5 algorithm has an unsatisfactory detection effect on small faces or tilted faces in face recognition tasks and is prone to false detection and missed detection in complex scenes.
[0005] A face recognition detection method based on an improved YOLOv5 algorithm, the face recognition detection method based on the improved YOLOv5 algorithm comprising the following steps:
[0006] S100, selecting a publicly available face dataset and preprocessing the face dataset;
[0007] S200, constructing a face recognition detection model based on the improved YOLOv5 algorithm, using the preprocessed face data set to train the model, and obtaining the optimal face recognition detection model;
[0008] S300, inputting the preprocessed training set into the face recognition detection model based on the improved YOLOv5 algorithm constructed in S200 for training, and continuously adjusting the model parameters during the training process, and finally selecting the model with the smallest loss function during the training process;
[0009] S400, input the preprocessed test set into the model with the smallest loss function during the training process obtained in S300 to perform face recognition detection, and obtain the face recognition detection result, so as to evaluate the performance of the model in practical applications.
[0010] Furthermore, in S100, the publicly available face dataset is selected from the public face dataset of WiderFace, which is divided into 61 event categories and has a total of 393,703 face images.
[0011] Furthermore, in S100, the preprocessing is: dividing the face data set according to event categories, selecting 80% of the images for each event as a training set, and the remaining 20% as a test set.
[0012] Furthermore, in S200, the improvements to the YOLOv5 algorithm include:
[0013] Backbone network improvement: Based on YOLOv5 as the basic model architecture, the CBL layer ordinary convolution in the backbone network is replaced by depthwise separable convolution;
[0014] Add attention mechanism: Add SENet attention mechanism before the SPPF module of YOLOv5 backbone network, and combine with SE module to enhance the representation ability of convolutional neural network through compression and excitation operations, thereby improving the accuracy of face recognition detection. The specific operations include convolution to obtain feature map, compression and excitation operations to generate weight vector, and apply the weight vector to the original feature map for recalibration.
[0015] Furthermore, in the backbone network improvement, the following steps are included:
[0016] S210, depthwise separable convolution adopts a combination of depthwise convolution and pointwise convolution. The depthwise convolution performs layer-by-layer convolution operations on feature maps under different channels through 3×3 depthwise convolution kernels. Each convolution kernel is only responsible for extracting the feature map of one channel, and the number of channels of the output feature map remains unchanged.
[0017] S220, using a point-by-point convolution kernel of size 1×1×N, performs weighted combination of the feature maps of the previous layer input in the channel dimension to generate a new feature map, and the number of final output feature maps is the same as the number of convolution kernels.
[0018] Furthermore, in adding the attention mechanism, the following steps are included:
[0019] S230, input feature map X∈R h×w×c′ After the convolution operation F tr Get the feature map U∈R h′×w′×c , as shown below:
[0020] U=F tr (X)
[0021] Among them, F tr is a convolution operation with c convolution kernels;
[0022] S240, by squeeze (F sq ) compresses the two-dimensional features of U into real numbers 1×1×c to capture global feature information. The calculation formula is:
[0023]
[0024] Where h and w represent the height and width of the feature map respectively; c represents the number of channels; x ijc Represents the eigenvalue of the (i, j)th position on the cth channel;
[0025] S250, incentive operation F ex The dependencies between channels are captured by the fully connected layer with reduced dimension and the ReLU activation function, and the weight of each channel is generated by the fully connected layer with increased dimension and the sigmoid activation function. The weight vector after excitation is as follows:
[0026] s=F ex (z,W)=σf2=σ(W2·ReLU(f1))=σ(W2·ReLU(W1z))
[0027] Among them, f1 and f2 represent the feature maps after the first fully connected layer and the second fully connected layer respectively. f2∈R 1×1×c , r is the dimensionality reduction ratio, which is used to reduce parameters and calculations; W is the weight matrix, W1 is the weight matrix of the first fully connected layer, and W2 is the weight matrix of the second fully connected layer. It can be automatically back-propagated for training, and the values of W1 and W2 are adjusted by the gradient descent method to optimize the performance of the model; σ is the sigmoid activation function;
[0028] S260, applying the stimulated weight vector s to each channel of the original feature map U to recalibrate the features, and performing the operation F scale , SENet adaptively adjusts the feature response of each channel, emphasizes important features and suppresses unimportant features, thereby enhancing the representation ability of the model. The operation is expressed as:
[0029] X′=F scale (U) = U·s
[0030] Where X′∈R h′×w′×c It is the final output feature map of the SE module.
[0031] Furthermore, the model training process in S300 uses an early stopping method, and continuously monitors the loss function value on the validation set during the training process. When the validation set loss no longer decreases within several consecutive training rounds, the training is terminated in advance to avoid overfitting of the model, and the model parameters at this time are saved as the final training result.
[0032] A storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned face recognition detection method based on the improved YOLOv5 algorithm.
[0033] A computer device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-mentioned face recognition detection method based on the improved YOLOv5 algorithm.
[0034] Beneficial effects of the present invention: A face recognition detection method based on an improved YOLOv5 algorithm of the present invention can realize efficient and real-time face recognition detection and has high detection accuracy, specifically including the following technical effects:
[0035] (1) The public dataset WiderFace is used for face recognition detection. This dataset can provide large-scale, diverse, and richly annotated face data, which can significantly improve the accuracy and generalization ability of the model.
[0036] (2) Using depthwise separable convolution to replace the ordinary convolution of the CBL layer in the backbone network of the YOLOv5 algorithm can significantly reduce the amount of calculation and parameters while maintaining high detection accuracy, which helps to achieve lightweight and efficient model.
[0037] (3) The SENet attention mechanism is used to enable the network to automatically focus on the internal relationship between the channel information and position information of the infrared image of the photovoltaic string, solving the problem of photovoltaic string edge information loss caused by the traditional convolution and downsampling process. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 A method flow chart of a face recognition detection method based on an improved YOLOv5 algorithm of the present invention;
[0039] Figure 2 It is a structural diagram of the depth-separable convolution in the backbone network of the present invention;
[0040] Figure 3 It is a structural diagram of CBL in the backbone network of the YOLOv5 algorithm to be improved in the present invention;
[0041] Figure 4 This is a structural diagram of the SENet attention mechanism of the present invention. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0043] Reference Figure 1 As shown, a face recognition detection method based on an improved YOLOv5 algorithm, the face recognition detection method based on the improved YOLOv5 algorithm comprises the following steps:
[0044] S100, selecting a publicly available face dataset and preprocessing the face dataset;
[0045] S200, constructing a face recognition detection model based on the improved YOLOv5 algorithm, using the preprocessed face data set to train the model, and obtaining the optimal face recognition detection model;
[0046] S300, inputting the preprocessed training set into the face recognition detection model based on the improved YOLOv5 algorithm constructed in S200 for training, and continuously adjusting the model parameters during the training process, and finally selecting the model with the smallest loss function during the training process;
[0047] S400, input the preprocessed test set into the model with the smallest loss function during the training process obtained in S300 to perform face recognition detection, and obtain the face recognition detection result, so as to evaluate the performance of the model in practical applications.
[0048] Specifically, the face recognition detection method based on the improved YOLOv5 algorithm of the present invention has many beneficial effects. By selecting a rich and diverse public face data set such as WiderFace and preprocessing it, a sufficient and high-quality data foundation is provided for the model. The constructed improved model uses deep separable convolution to replace the ordinary convolution in the backbone network, reduces the amount of calculation and the amount of parameters to achieve lightweight and high efficiency, and at the same time adds the SENet attention mechanism to the backbone network to enhance the feature representation ability and improve the detection accuracy. During the training process, the model with the smallest loss function is selected by continuously adjusting the parameters, and the preprocessed test set is used to detect and evaluate the performance in the test stage, which effectively improves the accuracy and robustness of face recognition detection as a whole, and can adapt to complex scenes, and has broad application prospects and practical value in practical applications.
[0049] Furthermore, in S100, the publicly available face dataset is selected from the public face dataset of WiderFace, which is divided into 61 event categories and has a total of 393,703 face images.
[0050] Specifically, this embodiment selects WiderFace's public face dataset (including 61 event categories and 393,703 face images) in S100. This dataset is rich and diverse, covering face images of various postures, expressions, lighting and occlusion conditions, providing massive and comprehensive data samples for model training, greatly enhancing the model's ability to learn different types of facial features, so that the model can more accurately identify faces with different features when facing complex and diverse actual face recognition scenarios, thereby effectively improving the accuracy, robustness and generalization ability of the model, and laying a solid data foundation for the subsequent construction of a high-precision face recognition detection model.
[0051] Furthermore, in S100, the preprocessing is: dividing the face data set according to event categories, selecting 80% of the images for each event as a training set, and the remaining 20% as a test set.
[0052] Specifically, in this embodiment, this reasonable data division method ensures the richness and diversity of the training set data, enables the model to fully learn the facial features in different event scenarios, helps to improve the generalization ability of the model, and enables it to better adapt to various practical application scenarios. At the same time, retaining a certain proportion of the test set data can effectively evaluate the performance of the model on unseen data, provide an accurate basis for the optimization and adjustment of the model, ensure the reliability and accuracy of the model in the actual face recognition detection task, and promote the continuous iterative optimization of the model to achieve more accurate and efficient face recognition detection.
[0053] Furthermore, in S200, the improvements to the YOLOv5 algorithm include:
[0054] Backbone network improvement: Based on YOLOv5 as the basic model architecture, the CBL layer ordinary convolution in the backbone network is replaced by depthwise separable convolution;
[0055] Add attention mechanism: Add SENet attention mechanism before the SPPF module of YOLOv5 backbone network, and combine with SE module to enhance the representation ability of convolutional neural network through compression and excitation operations, thereby improving the accuracy of face recognition detection. The specific operations include convolution to obtain feature map, compression and excitation operations to generate weight vector, and apply the weight vector to the original feature map for recalibration.
[0056] Specifically, the backbone network uses deep separable convolution to replace the ordinary convolution of the CBL layer. The deep separable convolution decouples the spatial and channel correlations, greatly reduces the amount of calculation parameters while ensuring less accuracy loss, realizes model lightweight and high efficiency, improves detection efficiency, and enables the model to maintain good performance under limited resources, enhancing its applicability in practical applications. The SENet attention mechanism is added before the SPPF module of the backbone network and combined with the SE module. Through compression and excitation operations, the ability to represent features can be effectively enhanced, prompting the network to automatically pay attention to the internal relationship between the channels and position information of the face image, reducing the information loss caused by traditional convolution and downsampling, thereby significantly improving the accuracy of face recognition detection, making face recognition more accurate and reliable, especially in complex scenes, it can more accurately capture and identify key features, and improve the overall performance of the model.
[0057] Furthermore, in the backbone network improvement, the following steps are included:
[0058] S210, depthwise separable convolution adopts a combination of depthwise convolution and pointwise convolution. The depthwise convolution performs layer-by-layer convolution operations on feature maps under different channels through 3×3 depthwise convolution kernels. Each convolution kernel is only responsible for extracting the feature map of one channel, and the number of channels of the output feature map remains unchanged.
[0059] S220, using a point-by-point convolution kernel of size 1×1×N, performs weighted combination of the feature maps of the previous layer input in the channel dimension to generate a new feature map, and the number of final output feature maps is the same as the number of convolution kernels.
[0060] Specifically, the main network of YOLOv5 consists of four parts, namely the Input module for the image input, the Backbone module for the backbone network, the Neck module for feature fusion, and the Prediction module for the prediction part. Based on the YOLOv5 model architecture, the CBL layer (including a convolution layer, a BN layer, and a LeakyReLU activation function) in the backbone network of the YOLOv5 algorithm is replaced with a deep separable convolution, which reduces the amount of calculation and parameters of the model and improves the detection efficiency of the model. Figure 3 It is a CBL module structure;
[0061] Compared with conventional convolution kernels, depthwise separable convolution uses a combination of depthwise convolution and pointwise convolution to convolve feature maps under different channels. The structure diagram is as follows: Figure 2 As shown in the figure, the spatial and channel correlations of the convolution layer are decoupled, and the number of computational parameters of the backbone network is reduced while ensuring less accuracy loss. The depth-wise separable convolution is specifically divided into two processes: channel-by-channel convolution and point-by-point convolution. The channel-by-channel convolution performs layer-by-layer convolution operations on feature maps under different channels through a 3×3 depth convolution kernel. Each convolution kernel is only responsible for the feature map extraction operation of one channel, and the number of channels of the output feature map does not change. Then, a 1×1×N convolution kernel is used for point-by-point convolution, where N is the number of channels of the feature map of the previous layer. The point-by-point convolution is similar to the conventional convolution kernel operation. The feature map of the previous layer input is weighted and combined in the channel dimension to generate a new feature map. The number of final output feature maps is the same as the number of convolution kernels. The calculation amount comparison formula between conventional convolution and depth-wise separable convolution is as follows:
[0062]
[0063] Among them, C and C′ represent the computational complexity of a feature extraction operation of a conventional convolution and a depth-separable convolution kernel, respectively, and F i represents the size of the input feature, F k Denotes the size of the convolution kernel, M and N denote the number of channels of the input and output features, respectively. It can be seen that using feature maps with the same number of channels, the amount of computation generated by the depthwise separable convolution is greatly reduced. The specific structure of the depthwise separable convolution kernel is as follows: Figure 2 shown.
[0064] Furthermore, in adding the attention mechanism, the following steps are included:
[0065] S230, input feature map X∈R h×w×c′ After the convolution operation F tr Get the feature map U∈R h′×w′×c , as shown below:
[0066] U=Ftr (X)
[0067] Among them, F tr is a convolution operation with c convolution kernels;
[0068] S240, by squeeze (F sq ) compresses the two-dimensional features of U into real numbers 1×1×c to capture global feature information. The calculation formula is:
[0069]
[0070] Where h and w represent the height and width of the feature map respectively; c represents the number of channels; x ijc Represents the eigenvalue of the (i, j)th position on the cth channel;
[0071] S250, incentive operation F ex The dependencies between channels are captured by the fully connected layer with reduced dimension and the ReLU activation function, and the weight of each channel is generated by the fully connected layer with increased dimension and the sigmoid activation function. The weight vector after excitation is as follows:
[0072] s=F ex (z,W)=σf2=σ(W2·ReLU(f1))=σ(W2·ReLU(W1z))
[0073] Among them, f1 and f2 represent the feature maps after the first fully connected layer and the second fully connected layer respectively. f2∈R 1×1×c , r is the dimensionality reduction ratio, which is used to reduce parameters and calculations; W is the weight matrix, W1 is the weight matrix of the first fully connected layer, and W2 is the weight matrix of the second fully connected layer. It can be automatically back-propagated for training, and the values of W1 and W2 are adjusted by the gradient descent method to optimize the performance of the model; σ is the sigmoid activation function;
[0074] S260, applying the stimulated weight vector s to each channel of the original feature map U to recalibrate the features, and performing the operation F scale , SENet adaptively adjusts the feature response of each channel, emphasizes important features and suppresses unimportant features, thereby enhancing the representation ability of the model. The operation is expressed as:
[0075] X′=F scale (U) = U·s
[0076] Where X′∈R h′×w′×c It is the final output feature map of the SE module.
[0077] Specifically, in order to improve the detection accuracy of the model, the SENet (Squeeze-and-Excitation Network) attention mechanism is added before the SPPF module (full name Spatial Pyramid Pooling-Fast, which is an improved spatial pyramid pooling technology, which is used for multi-scale feature extraction to enhance the model's detection ability for targets of different sizes) of the YOLOv5 backbone network, combined with the SE module that can improve the detection accuracy of the algorithm. This module enhances the convolutional neural network's ability to represent features through compression and excitation operations. The attention mechanism optimizes the model, thereby improving the detection accuracy of face recognition. The network structure diagram is shown in the figure below: Figure 4 As shown. The specific operation is: First, input feature map X∈R h×w×c′ After the convolution operation F tr Get the feature map U∈R h′×w′×c , as shown below:
[0078] U=F tr (X)
[0079] Among them, F tr is a convolution operation with c convolution kernels.
[0080] Next, squeeze(F sq ) compresses the two-dimensional features of U into real numbers 1×1×c to capture global feature information.
[0081] The calculation formula is:
[0082]
[0083] Where h and w represent the height and width of the feature map respectively; c represents the number of channels; x ijc Represents the eigenvalue of the (i, j)th position on the cth channel.
[0084] Stimulus Operation F ex The dependencies between channels are captured by the fully connected layer with reduced dimension and the ReLU activation function, and the weight of each channel is generated by the fully connected layer with increased dimension and the sigmoid activation function. The weight vector after excitation is as follows:
[0085] s=F ex (z,W)=σf2=σ(W2·ReLU(f1))=σ(W2·ReLU(W1z))
[0086] Among them, f1 and f2 represent the feature maps after the first fully connected layer and the second fully connected layer respectively. f2∈R 1×1×c, r is the dimensionality reduction ratio, which is used to reduce parameters and calculations; W is the weight matrix, W1 is the weight matrix of the first fully connected layer, and W2 is the weight matrix of the second fully connected layer. It can be automatically back-propagated for training, and the values of W1 and W2 are adjusted by the gradient descent method to optimize the performance of the model; σ is the sigmoid activation function.
[0087] Finally, the activated weight vector s is applied to each channel of the original feature map U to recalibrate the features, through the operation F scale , SENet can adaptively adjust the feature response of each channel, emphasize important features and suppress unimportant features, thereby enhancing the representation ability of the model. The specific operation can be expressed as:
[0088] X′=F scale (U) = U·s
[0089] Where X′∈R h′×w′×c It is the final output feature map of the SE module.
[0090] After the feature map is obtained through convolution, the squeeze operation compresses the features to obtain global information, and the excitation operation uses the dimension reduction fully connected layer and activation function to capture channel dependencies and generate weights, and finally the weight vector is used to recalibrate the original feature map channel. This process enables the model to adaptively adjust feature responses, emphasize key features and suppress secondary features, enhance feature representation capabilities, and effectively improve the accuracy of face recognition detection. Especially in complex environments, it can better capture key facial features, improve the model's detection accuracy for faces of different sizes, postures, expressions, etc., and enhance the overall performance and adaptability of the model.
[0091] Furthermore, the model training process in S300 uses an early stopping method, and continuously monitors the loss function value on the validation set during the training process. When the validation set loss no longer decreases within several consecutive training rounds, the training is terminated in advance to avoid overfitting of the model, and the model parameters at this time are saved as the final training result.
[0092] Specifically, this embodiment continuously monitors the loss function value on the validation set, and terminates the training early once it finds that the validation set loss no longer decreases for several consecutive training rounds. This can effectively avoid the overfitting problem of the model and ensure that the model can be better generalized to unknown data. At the same time, the model parameters at this time are saved as the final training result, ensuring the reliability and stability of the model.
[0093] A storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned face recognition detection method based on the improved YOLOv5 algorithm.
[0094] Specifically, when the computer program stored in the storage medium is executed by the processor, it can implement a face recognition detection method based on the improved YOLOv5 algorithm, and can conveniently and quickly rely on corresponding hardware to implement the advanced face recognition detection method, so that the method is not limited to specific equipment or environment, and is easy to promote and use. It provides strong support for efficient and accurate face recognition detection work in different scenarios, further broadens its application scope, and helps to improve the overall efficiency and accuracy of face recognition in related fields.
[0095] A computer device comprises: a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-mentioned face recognition detection method based on the improved YOLOv5 algorithm.
[0096] Specifically, this computer device integrates a memory, a processor and a corresponding computer program, and implements a face recognition detection method based on the improved YOLOv5 algorithm by executing the program through the processor, integrating advanced face recognition detection technology into the computer device, so that the device has efficient and accurate face recognition detection capabilities, and can operate stably in different actual application scenarios, providing reliable technical support for many fields such as security, access control, identity authentication, etc., effectively improving the convenience, accuracy and overall efficiency of related business development, and is easy to operate and manage, which is conducive to large-scale deployment and application.
[0097] The present invention discloses a face recognition detection method based on an improved YOLOv5 algorithm. The WiderFace public data set is selected and the training set and the test set are reasonably divided to provide a rich and diverse data basis for the model, thereby improving the generalization ability and the evaluation accuracy. At the model level, the backbone network adopts deep separable convolution to reduce the number of parameters and improve efficiency, and the SENet attention mechanism is added to enhance the feature representation ability and improve the detection accuracy. The early stopping method is used in the training process to avoid overfitting and ensure the performance of the model. The application of storage media and computer equipment makes the method easy to implement, providing efficient and accurate face recognition support for security and other fields, facilitating large-scale deployment, and improving business efficiency and accuracy.
[0098] It should be noted that, in this article, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In addition, "front", "back", "left", "right", "upper" and "lower" in this article are all referenced to the placement state shown in the accompanying drawings.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A face recognition detection method based on an improved YOLOv5 algorithm, characterized in that: The face recognition detection method based on the improved YOLOv5 algorithm comprises the following steps: S100, selecting a publicly available face dataset and preprocessing the face dataset; S200, constructing a face recognition detection model based on the improved YOLOv5 algorithm, using the preprocessed face data set to train the model, and obtaining the optimal face recognition detection model; S300, inputting the preprocessed training set into the face recognition detection model based on the improved YOLOv5 algorithm constructed in S200 for training, and continuously adjusting the model parameters during the training process, and finally selecting the model with the smallest loss function during the training process; S400, input the preprocessed test set into the model with the smallest loss function during the training process obtained in S300 to perform face recognition detection, and obtain the face recognition detection result, so as to evaluate the performance of the model in practical applications.
2. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 1, characterized in that: In S100, the publicly available face dataset is selected from the public face dataset of WiderFace, which is divided into 61 event categories and has a total of 393,703 face images.
3. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 2, characterized in that: In S100, the preprocessing is: dividing the face data set according to event categories, selecting 80% of the images for each event as a training set, and the remaining 20% as a test set.
4. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 3, characterized in that: In S200, the improvements to the YOLOv5 algorithm include: Backbone network improvement: Based on YOLOv5 as the basic model architecture, the CBL layer ordinary convolution in the backbone network is replaced by depthwise separable convolution; Add attention mechanism: Add SENet attention mechanism before the SPPF module of YOLOv5 backbone network, and combine with SE module to enhance the representation ability of convolutional neural network through compression and excitation operations, thereby improving the accuracy of face recognition detection. The specific operations include convolution to obtain feature map, compression and excitation operations to generate weight vector, and apply the weight vector to the original feature map for recalibration.
5. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 4, characterized in that: The backbone network improvement includes the following steps: S210, depthwise separable convolution adopts a combination of depthwise convolution and pointwise convolution. The depthwise convolution performs layer-by-layer convolution operations on feature maps under different channels through 3×3 depthwise convolution kernels. Each convolution kernel is only responsible for extracting the feature map of one channel, and the number of channels of the output feature map remains unchanged. S220, using a point-by-point convolution kernel of size 1×1×N, performs weighted combination of the feature maps of the previous layer input in the channel dimension to generate a new feature map, and the number of final output feature maps is the same as the number of convolution kernels.
6. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 5, characterized in that: In adding the attention mechanism, the following steps are included: S230, input feature map X∈R h×w×c′ After the convolution operation F tr Get the feature map U∈R h′×w′×c , as shown below: U=F tr (X) Among them, F tr is a convolution operation with c convolution kernels; S240, by squeeze (F sq ) compresses the two-dimensional features of U into real numbers 1×1×c to capture global feature information. The calculation formula is: Where h and w represent the height and width of the feature map respectively; c represents the number of channels; x ijc Represents the eigenvalue of the (i, j)th position on the cth channel; S250, incentive operation F ex The dependencies between channels are captured by the fully connected layer with reduced dimension and the ReLU activation function, and the weight of each channel is generated by the fully connected layer with increased dimension and the sigmoid activation function. The weight vector after excitation is as follows: s=F ex (z,W)=σf2=σ(W2·ReLU(f1))=σ(W2·ReLU(W1z)) Among them, f1 and f2 represent the feature maps after the first fully connected layer and the second fully connected layer respectively. f2∈R 1×1×c , r is the dimensionality reduction ratio, which is used to reduce parameters and calculations; W is the weight matrix, W1 is the weight matrix of the first fully connected layer, and W2 is the weight matrix of the second fully connected layer. It can be automatically back-propagated for training, and the values of W1 and W2 are adjusted by the gradient descent method to optimize the performance of the model; σ is the sigmoid activation function; S260, applying the stimulated weight vector s to each channel of the original feature map U to recalibrate the features, and performing the operation F scale , SENet adaptively adjusts the feature response of each channel, emphasizes important features and suppresses unimportant features, thereby enhancing the representation ability of the model. The operation is expressed as: X′=F scale (U)=U·s Where X′∈R h′×w′×c is the final output feature map of the SE module.
7. A face recognition detection method based on an improved YOLOv5 algorithm according to claim 6, characterized in that: The model training process in S300 uses the early stopping method, and continuously monitors the loss function value on the validation set during the training process. When the validation set loss no longer decreases within several consecutive training rounds, the training is terminated in advance to avoid model overfitting, and the model parameters at this time are saved as the final training result.
8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the face recognition detection method based on the improved YOLOv5 algorithm described in any one of claims 1 to 7 is implemented.
9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a face recognition detection method based on an improved YOLOv5 algorithm as described in any one of claims 1 to 7.
Citation Information
Cited By
Expression recognition method and device, storage medium and electronic equipment
CN121305638A
Unmanned aerial vehicle signal detection method and system based on deep learning
CN122333115A