A face localization and detection method, application, and storage medium
By improving the YOLOv3-tiny model, combining MobileNetV3 and lightweight SE attention module, the network structure is optimized, and the accuracy of face detection under the influence of factors such as lighting, angle and size is solved, and efficient face positioning detection on embedded devices is achieved.
Patent Information
- Application Number
- CN202210614109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing face detection technology is insufficiently accurate under the influence of factors such as lighting, angle and size, and the deep learning method is slow, making it difficult to efficiently implement on embedded devices.
The MobileNetV3 network was introduced to improve the YOLOv3-tiny model, combining the lightweight SE attention module and h-swish activation function, optimize the network structure, and transform multi-objective regression into a single target, search for the optimal network architecture through NAS and NetAdapt algorithms, and face positioning detection is performed.
It improves the accuracy and accuracy of face detection, reduces the amount of calculation, makes the face positioning detection method more efficient on embedded devices, and expands application scenarios.
Smart Images

Figure CN115035575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and particularly relates to a face localization detection method, application, and storage medium. Background Art
[0002] With the development of computer technology, especially pattern recognition technology, face detection has emerged in people's vision as a technical direction. Face detection technology can serve as a basic task in various application projects in the fields of image processing and video analysis, such as face recognition, face image retrieval, and driver fatigue state detection, etc.
[0003] Currently, there are mainly two major directions in face detection. One is the traditional face detection method using artificial features plus a classifier, such as the widely used Viola-Jones face detection method; the other is the face detection method based on a deep learning framework. The traditional face detection method using artificial features plus a classifier is mainly affected by various factors such as illumination, angle, and size. The face detection method based on a deep learning framework has greatly improved in detection effect. Nowadays, advanced detection methods can detect faces smaller than 20 pixels × 20 pixels while controlling false detections, and have good robustness to angle changes and image quality. However, this method is usually slow. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a face localization detection method, application, and storage medium, which can quickly and accurately locate the face area, providing an accurate face image basis for subsequent processing algorithms.
[0005] The present invention adopts the following technical solutions to solve the technical problems:
[0006] The present invention provides a face localization detection method. The method introduces the MobileNetV3 network as a feature extraction network to improve the YOLOv3-tiny model, simplifies the network structure and controls the model calculation amount, and converts multi-object regression into a single object to train the face localization detection algorithm model, specifically including the following steps:
[0007] Configure the NAS algorithm and the NetAdapt algorithm as network structure complementary search algorithms, which are respectively used for the search of the overall and local network structures;
[0008] Combine the lightweight SE attention module with the inverted residual structure with a linear bottleneck in MobileNetV2 to process and generate the final attention weight feature map;
[0009] Use the h-swish activation function to replace the swish activation function in MobileNetV2, where the h-swish activation function formula is:
[0010]
[0011] Based on the improved YOLOv3-tiny model, the images of the face dataset are segmented, adjusted and arranged into grid cells according to several different sizes, the positions of the faces are located and detected on the non-overlapping grid cells, and they are classified.
[0012] Preferably, the overall and local search of the network structure specifically includes:
[0013] Use the NAS algorithm to construct a set of candidate neural network structure in the search space, and search for the optimal network structure from it through performance evaluation as the initial seed network architecture;
[0014] Use the NetAdapt algorithm to fine-tune the network layers in the initial seed network architecture until the target latency is reached.
[0015] Preferably, the use of the NetAdapt algorithm to fine-tune the network layers in the initial seed network architecture until the target latency is reached specifically includes:
[0016] Generate a new set of suggestions based on the initial seed network architecture, where each suggestion is independently configured as a modification to the network architecture, and the latency reduction amount after each modification is not less than 10ms;
[0017] Execute the suggestions sequentially. After each modification to the network architecture, fill in the newly proposed architecture;
[0018] Judge whether the existing architecture meets the preset resource requirement limit. If not, prune the filters based on the gradient information, and perform short-term fine-tuning on the currently pruned layer to restore its accuracy;
[0019] After all layers are pruned, select the one that meets the resource requirement limit and has the best network performance after fine-tuning as the network architecture scheme with the highest accuracy;
[0020] Iteratively execute the above steps until the network architecture reaches the target latency. At this time, perform long-term fine-tuning on the network architecture as the output of the adaptive network architecture.
[0021] Preferably, the combination of the lightweight SE attention module and the inverted residual structure with a linear bottleneck in MobileNetV2 to generate the final attention weight feature map specifically includes:
[0022] Perform a dimension increase operation using 1×1 convolution, then use the 3×3 depthwise separable convolution in MobileNetV1 to extract features from the dimension-increased feature map. Adjust the weights of each feature map channel for the extracted feature map through the activation functions of the pooling layer and the fully connected layer. Finally, perform a matrix multiplication operation on the adjusted weights and the feature map extracted by the 3×3 depthwise separable convolution to obtain the final attention weight feature map.
[0023] Preferably, the images of the face dataset are segmented and adjusted according to 10 different sizes, and the scale features of the grid cells include 13*13 and 26*26.
[0024] Preferably, for each grid cell, output the bounding box, the corresponding confidence, and the conditional probability of the face. Use the non-maximum suppression algorithm to suppress redundant bounding boxes, where the confidence formula is:
[0025]
[0026] In the formula, P T (Object) is the conditional probability of the face. If a face is included, then P T (Object) = 1; otherwise P T (Object) = 0, is the intersection over union between the bounding box and the actual box.
[0027] Preferably, the algorithm loss function in the face localization detection algorithm model consists of the center error term of the bounding box, the width and height error terms of the bounding box, the error term of the predicted confidence, and the error term of the predicted category.
[0028] The present invention also provides a fatigue driving detection algorithm, including extracting the facial area of the driver, extracting the eye and mouth coordinates according to the facial area coordinates, calculating the eye feature vector and the mouth feature vector in real time, determining the states of the eyes and the mouth, classifying the verified identity status, calculating the blink frequency and yawn frequency within a certain time period, and performing fatigue judgment. Among them, the above-mentioned face localization detection method is used to extract and locate the facial area of the driver from the complex background.
[0029] The present invention also provides an electronic device, including:
[0030] At least one processor; and
[0031] A memory communicatively connected to the at least one processor; wherein,
[0032] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the foregoing face localization detection method.
[0033] The present invention also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the foregoing face localization and detection method.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] The face localization and detection method of the present invention greatly improves the accuracy and precision of face detection. It optimizes the network structure and effectively reduces the computational amount, and can provide technical support for subsequent face recognition, feature extraction, etc. The face detection algorithm model completed based on offline training can not only accurately locate the face area, but also be more convenient to be transplanted on mobile devices such as embedded devices, effectively expanding the application scenarios and scope.
[0036] Regarding the present invention relative to the prior art, other prominent substantive features and remarkable progress are further introduced in detail in the embodiment part. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0038] Figure 1 It is a schematic diagram of the network structure of YOLOv3-tiny;
[0039] Figure 2 It is a schematic diagram of the improved YOLOv3-tiny network structure in Embodiment 1;
[0040] Figure 3 It is a flowchart of the face localization and detection method in Embodiment 1;
[0041] Figure 4 It is a flowchart of the NAS algorithm adopted in Embodiment 1;
[0042] Figure 5 It is a flowchart of the NetAdapt algorithm adopted in Embodiment 1;
[0043] Figure 6 It is a structure diagram of the SE attention module in Embodiment 1;
[0044] Figure 7 It is a comparison diagram of the swish activation function and the h-swish activation function adopted in Embodiment 1;
[0045] Figure 8 It is a schematic diagram of IOU in Embodiment 1;
[0046] Figure 9 It is the accuracy curve of face localization and detection in Embodiment 1. Detailed implementation manners
[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0048] It should be noted that certain names are used to refer to specific components in the specification and claims. It should be understood that those of ordinary skill in the art may use different names to refer to the same component. The specification and claims of this application do not use the difference in names as a way to distinguish components, but use the substantial difference in functions of components as the criterion for distinguishing components. For example, the terms "comprising" or "including" used in the specification and claims of this application are open-ended terms, which should be interpreted as "comprising but not limited to" or "including but not limited to". The embodiments described in the detailed implementation manners section are the preferred embodiments of the present invention and are not intended to limit the scope of the present invention.
[0049] In addition, those skilled in the art know that various aspects of the present invention can be implemented as a system, a method, or a computer program product. Therefore, various aspects of the present invention can be specifically implemented in a form of combination of software and hardware, which can be collectively referred to as "circuit", "module", or "system" here. In addition, in some embodiments, various aspects of the present invention can also be implemented in the form of a computer program product in one or more microcontroller-readable media, which contain program codes readable by the microcontroller.
[0050] Embodiment 1
[0051] The YOLO model is a fast object detection model based on deep learning. This model establishes a separate end-to-end network system, transforming object detection into a regression problem. More specifically, the regression method replaces the traditional sliding window of object detection with CNN in traditional object detection to achieve feature extraction of faces. Using this extraction method, the target features can be quickly extracted, reducing the influence of the external environment on the proposed method. YOLO uses an end-to-end convolutional neural network to extract the features of the input image and regards object detection as a regression problem to be solved. It can directly obtain some information about the target in the image, such as position, size, and category information. YOLOv3 is an improved version of the YOLO algorithm and is one of the best algorithms currently shown in the field of object detection. Based on YOLO, YOLOv3 incorporates many excellent research results. When compared with the SSD algorithm, when using it to detect a 320×320 image, the detected image accuracy is the same, but the speed is three times faster than SSD. The backbone feature extraction network used by YOLOv3 is Darknet53.
[0052] YOLOv3-tiny is a lightweight object detection model based on YOLOv3. Compared with YOLOv3, the YOLOv3-tiny algorithm has the following advantages:
[0053] (1) Faster detection speed. The detection results can be obtained by running the neural network once for each test image and can be used for real-time detection.
[0054] (2) Global understanding of the image. Information around the target can be learned during the training process, and the background error rate is less than half of that of the Faster-Rcnn algorithm. YOLOv3-tiny can be used in embedded devices with lower computing power.
[0055] By simplifying the network structure of YOLOv3, the network structure of YOLOv3-tiny is obtained. As Figure 1 shown, YOLOv3-tiny contains 13 convolutional layers, 6 max-pooling layers, 1 upsampling layer, 1 fully connected layer, and 2 output layers, for a total of 23 layers of network.
[0056] Although the detection speed of the YOLOv3-tiny model has been greatly improved compared with YOLOv3, the detection accuracy and precision have decreased significantly. Based on the problems of its detection accuracy and precision, this embodiment has improved the YOLOv3-tiny model, as Figure 2As shown, according to the regression design idea of the YOLO object detection model, in this embodiment, multi-object regression is transformed, the network is simplified and the model calculation is avoided from increasing, and it is transformed into a single object; secondly, the model structure of YOLOv3-tiny is improved, and the MobileNetV3 network is introduced as the feature extraction network to better locate suspicious face regions, such as Figure 3 As shown, the face localization detection method of this embodiment specifically includes the following steps:
[0057] Configure the NAS algorithm and the NetAdapt algorithm as network structure complementary search algorithms, which are respectively used for the overall and local search of the network structure, specifically including:
[0058] Use the NAS algorithm to construct a set of candidate neural network structure sets in the search space, and search for the optimal network structure from them through performance evaluation as the initial seed network architecture; as Figure 4 As shown, in each iteration of the search process, a "sample" is generated from the search space, that is, a neural network structure is obtained, which is called a "sub-network". The sub-network is trained on the training sample set, and then its performance is evaluated on the validation set. Gradually optimize the network structure until the optimal sub-network is found. More specific explanation: for models with small parameter quantities, the change of accuracy in terms of latency is more significant; therefore, a smaller weight factor is required, such as the weight factor w = 0.15, to compensate for the greater accuracy change brought by different latencies. Under the enhancement of this new weight factor w, a new architecture search is carried out from scratch to find the initial seed model, and then NetAdapt and other optimizations are applied to obtain the final network model.
[0059] Use the NetAdapt algorithm to fine-tune the network layers in the initial seed network architecture until the target latency is reached, specifically including:
[0060] Generate a new set of suggestions based on the initial seed network architecture, where each suggestion is independently configured as a modification to the network architecture, and the latency reduction amount after each modification is not less than 10ms;
[0061] Execute the suggestions sequentially. After each modification to the network architecture, fill in the newly proposed architecture;
[0062] Judge whether the existing architecture meets the preset resource requirement limit. If not, clip the filters based on the gradient information, and perform short-term fine-tuning on the currently clipped layer to restore its accuracy;
[0063] After all layers are clipped, select the one that meets the resource requirement limit and has the best network performance after fine-tuning as the network architecture scheme with the highest accuracy;
[0064] Iteratively execute the above steps until the network architecture reaches the target latency. At this time, perform long-term fine-tuning on the network architecture to determine the final network architecture;
[0065] As Figure 5 shown, this figure specifically shows the NetAdapt algorithm process in this embodiment. In each iteration, NetAdapt reduces resource consumption by simplification (i.e., deleting filters from a layer). To maximize accuracy, it tries to simplify each layer one by one and selects the simplified network with the highest accuracy. Once the target budget is reached, the selected network will be fine-tuned again until convergence. The NetAdapt algorithm is used to optimize the number of kernels in each layer. For the global network structure search, an RNN-based controller and a hierarchical search space are used, and accuracy-latency balance optimization is performed for a specific hardware platform. Search within the target latency range, and then use the NetAdapt method to optimize each layer in a sequential manner. While optimizing the model latency as much as possible, maintain the accuracy and reduce the size of the expansion layer and the bottleneck in each layer.
[0066] Combine the lightweight SE attention module with the inverted residual structure with a linear bottleneck in MobileNetV2 to process and generate the final attention weight feature map. See Figure 6 , specifically including:
[0067] Use a 1×1 convolution for dimension increase operation, and then use the 3×3 depthwise separable convolution in MobileNetV1 to extract features from the dimension-increased feature map. Pass the extracted feature map through the pooling layer and the activation function of the fully connected layer to adjust the weights of each feature map channel. Finally, perform a matrix multiplication operation on the adjusted weights and the feature map extracted by the 3×3 depthwise separable convolution to obtain the final attention weight feature map;
[0068] Use the h-swish activation function to replace the swish activation function in MobileNetV2, where the h-swish activation function formula is:
[0069] The expression shows that the activation function h-swish[x] is equal to the function value multiplied by the activation function ReLu6 (function value + 3) divided by 6; As Figure 7 shown is the comparison chart of the swish activation function and the h-swish activation function;
[0070] Refer to Figure 2, based on the improved YOLOv3-tiny model, images of a wider face dataset are segmented, adjusted, and arranged into grid cells according to several different sizes. The positions of detected faces are located on non-overlapping grid cells and classified. For each grid cell, a bounding box, the corresponding confidence, and the conditional probability of the face are output. The non-maximum suppression algorithm is used to suppress redundant bounding boxes. The confidence formula is as follows:
[0071]
[0072] In the formula, P T (Object) is the conditional probability of the face. If the face is included, P T (Object) = 1; otherwise, P T (Object) = 0. is the intersection over union between the bounding box and the actual box;
[0073] The images of the above face dataset are segmented and adjusted according to 10 different sizes. The scale features of the grid cells include 13*13 and 26*26. 10 is the optimized number of segments, and 13*13 and 26*26 are scale features suitable for detecting larger and medium-sized objects.
[0074] In the face localization and detection algorithm model of this embodiment, the algorithm loss function consists of the center error term of the bounding box, the width and height error terms of the bounding box, the error term of the predicted confidence, and the error term of the predicted category.
[0075] To further verify the performance of the face localization and detection method of the present invention, this embodiment conducts a quantitative evaluation on the WIDER FACE dataset.
[0076] In this embodiment, accuracy is selected as the evaluation index, and the intuitive evaluation index of its model performance is shown in the following formula:
[0077]
[0078] Among them, N d is the number of correctly detected images, and N t is the total number of images.
[0079] During the training and validation of the improved YOLOv3-tiny network, the intersection over union parameter (IOU) is introduced to measure the similarity between the face detection area and the marked ground truth area. Refer to Figure 8 , face_d is the face area detected by the model, and face is the marked actual area. The calculation formula is:
[0080]
[0081] Among them, S(face_d∩face) is the area of face_d∩face, and S(face_d∪face) is the area of face_d∪face.
[0082] The intersection ratio represents the degree of overlap between the model prediction area and the actual area. Figure 8 It shows that this value is in a direct proportional relationship with the detection accuracy. In the case of IOU = 1, the predicted box overlaps with the ground truth box. Generally speaking, in the task of object detection, it is considered that when IOU > 1, the object is correctly detected. In the face detection task of this embodiment, considering that the face detection result directly affects the accuracy of the subsequent algorithm, a relatively high threshold is set in this embodiment. When IOU > 0.75, it is considered that the face is correctly detected. Figure 9 It shows the accuracy curve of face detection during the training process of the network in this embodiment. It can be concluded that there is also a certain direct proportional relationship between the number of training times and the accuracy of face detection. The accuracy rate of the face localization and detection method of the present invention is 97.9%.
[0083] Embodiment 2
[0084] This embodiment provides a fatigue driving detection algorithm, including extracting the facial area of the driver, extracting the eye and mouth coordinates according to the facial area coordinates, calculating the eye feature vector and mouth feature vector in real time, determining the states of the eyes and mouth, classifying the verified identity status, calculating the blink frequency and yawn frequency within a certain time period, and performing fatigue judgment. This embodiment uses the face localization and detection method in Embodiment 1 to extract and locate the facial area of the driver from a complex background.
[0085] Embodiment 3
[0086] This embodiment provides an electronic device, including:
[0087] At least one processor; and
[0088] A memory communicatively connected to the at least one processor; wherein,
[0089] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the face localization and detection method described in Embodiment 1.
[0090] Embodiment 4
[0091] This embodiment provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the face localization and detection method described in Embodiment 1.
[0092] The various embodiments of the systems and techniques described in the above embodiments can be implemented in digital electronic circuitry, integrated circuit systems, application specific ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] These computing programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal for providing machine instructions and / or data to a programmable processor.
[0094] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0095] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0096] A computer system can include clients and servers. The clients and servers are generally remote from each other and typically interact through a communication network. The client - server relationship is created by computer programs running on respective computers and having a client - server relationship with each other.
[0097] It will be apparent to those skilled in the art that the present invention is not limited to the details of the above - described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non - restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claim concerned.
[0098] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment contains only one independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A face location detection method, characterized in that, The method uses the MobileNetV3 network as the feature extraction network to improve the YOLOv3-tiny model, simplifies the network structure and controls the model's computational complexity, and transforms multi-object regression into single-object to train the face localization detection algorithm model. The specific steps are as follows: Configure the NAS algorithm and the NetAdapt algorithm as network structure complementary search algorithms, which are used for the overall and local search of the network structure respectively, specifically including: Use the NAS algorithm to construct a set of candidate neural network structure in the search space, and search for the optimal network structure from it through performance evaluation as the initial seed network architecture; Use the NetAdapt algorithm to fine-tune the network layers in the initial seed network architecture until the target latency is reached; Combine the lightweight SE attention module with the inverted residual structure with linear bottleneck in MobileNetV2 to generate the final attention weight feature map, specifically including: Use 1×1 convolution for dimensionality increase operation, then use the 3×3 depthwise separable convolution in MobileNetV1 to extract features from the dimensionality-increased feature map, adjust the weights of each feature map channel through the activation functions of the pooling layer and the fully connected layer, and finally perform matrix multiplication operation on the adjusted weights and the feature map extracted by the 3×3 depthwise separable convolution to obtain the final attention weight feature map; Use the h-swish activation function to replace the swish activation function in MobileNetV2, where the formula of the h-swish activation function is: where x is the function value and ReLu6() is the activation function; Based on the improved YOLOv3-tiny model, the images of the face dataset are segmented, adjusted and arranged into grid cells according to several different sizes, the positions of the faces are located and detected on non-overlapping grid cells, and they are classified.
2. The face localization and detection method according to claim 1, characterized in that The step of using the NetAdapt algorithm to fine-tune the network layers in the initial seed network architecture until the target latency is reached specifically includes: Generate a new set of suggestions based on the initial seed network architecture, where each suggestion is independently configured as a modification to the network architecture, and the latency reduction amount after each modification is not less than 10ms; Execute the suggestions sequentially, and fill in the newly proposed architecture after each modification to the network architecture; Judge whether the existing architecture meets the preset resource requirement limit. If not, prune the filters based on the gradient information, and perform short-term fine-tuning on the currently pruned layer to restore its accuracy; After all layers are pruned, select the one that meets the resource requirement limit and has the best network performance after fine-tuning as the network architecture scheme with the highest accuracy; Iteratively execute the above steps until the network architecture reaches the target latency, and at this time, perform long-term fine-tuning on the network architecture as the adaptive network architecture output.
3. A face localization and detection method according to claim 1, characterized in that, The images of the face dataset are segmented and adjusted according to 10 different sizes, and the scale features of the grid cells include 13*13 and 26*26.
4. A face location detection method according to claim 3, wherein, For each grid cell, output the bounding box, the corresponding confidence, and the conditional probability of the face, and use the non-maximum suppression algorithm to suppress redundant bounding boxes, where the confidence formula is: where P T (Object) is the conditional probability of a human face. If a human face is included, then P T (Object) = 1; otherwise P T (Object) = 0, is the intersection over union between the bounding box and the actual box.
5. A face location detection method according to claim 4, characterized in that, In the face localization and detection algorithm model, the algorithm loss function consists of the center error term of the bounding box, the width and height error terms of the bounding box, the error term of the predicted confidence, and the error term of the predicted category.
6. A fatigue driving detection method, including extracting the facial area of the driver, extracting the coordinates of the eyes and mouth according to the facial area coordinates, calculating the eye feature vector and the mouth feature vector in real time, determining the states of the eyes and mouth, classifying the verified identity status, calculating the blinking frequency and yawning frequency within a certain period of time, and making a fatigue judgment, characterized in that, The face localization and detection method according to any one of claims 1-5 is used to extract and locate the facial area of the driver from the complex background.
7. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the face localization and detection method according to any one of claims 1-5.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the face localization and detection method according to any one of claims 1-5.
Citation Information
Patent Citations
Neural network compression method, apparatus and device, and storage medium
CN111967594A
Mask detection method based on yolov4
CN113762201A