Machine room external invasion personnel early warning method based on face recognition
By employing a face recognition-based early warning method, and utilizing the YOLOV5-Face+ECANet network for image enhancement, feature extraction, and occlusion processing, the accuracy and timeliness issues of traditional data center intrusion early warning systems have been resolved. This approach achieves efficient real-time early warning and identity recognition, thereby enhancing the security management level of bank data centers.
Patent Information
- Application Number
- CN202510894023.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional methods for early warning of unauthorized personnel entering and exiting data centers suffer from several drawbacks, including fixed infrared detection range, numerous blind spots, insufficient accuracy and timeliness of warnings, low efficiency of manual patrols, susceptibility to human factors, and difficulty in recognizing facial obstructions. These issues contribute to the low accuracy of early warnings.
A face recognition-based early warning method is adopted, which uses face image enhancement, feature extraction, real-time detection, key point detection and occlusion detection. The YOLOV5-Face+ECANet network is used for real-time face recognition and early warning, and the attention mechanism is combined to handle occluded areas to improve recognition accuracy.
It improves the accuracy of early warning for unauthorized personnel entering the data center, reduces false alarms and missed alarms, enhances the robustness and applicability of the system, effectively addresses situations where faces are obscured, achieves real-time early warning, and ensures data center security.
Smart Images

Figure FT_1 
Figure QLYQS_1 
Figure QLYQS_2
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of face recognition, and particularly relates to a method for warning intruders outside a computer room based on face recognition. BACKGROUND
[0002] In the modern financial system, the bank data center, as the core hub to ensure the normal operation of the bank, is self-evident in importance. The data center is composed of servers, storage, networks, UPS and other types of computer rooms, which carry massive key data and core business systems of the bank. Since the data stored and processed in the computer room has high sensitivity and importance, ensuring the safe and stable operation of the computer room becomes the top priority of bank security management.
[0003] The traditional method for warning intruders outside the computer room mainly relies on infrared technology. By setting infrared detection devices around the computer room, the infrared emission and reception principle is used. When a person enters the infrared detection range, an alarm signal is triggered to achieve preliminary detection of personnel intrusion. In addition, some computer rooms also combine video monitoring systems to monitor personnel activities in the monitoring area through manual patrol of the monitoring screen or simple motion detection algorithms. When abnormal personnel activities are found, appropriate warning measures are taken. Although it can achieve warning of intruders to some extent, there are still many problems.
[0004] Limitations of the traditional method for warning intruders outside the computer room: First, the traditional method for warning intruders outside the computer room mainly relies on infrared technology. The infrared detection range is fixed and cannot be flexibly adjusted according to the actual scene and demand. For some complex and variable computer room environments, it is difficult to cover all possible intrusion paths comprehensively, which may lead to monitoring blind spots, causing intruders to bypass the infrared detection area and failing to issue timely warnings. Second, the traditional method for warning intruders outside the computer room has deficiencies in accuracy and timeliness. Manual patrol of the monitoring screen is inefficient and easily affected by human factors, which may result in missed detection, false detection and other situations, leading to low accuracy of the warning. In addition, the traditional method performs poorly in complex situations such as face occlusion. When intruders intentionally block their faces, the traditional monitoring system is difficult to accurately identify their identities, thus failing to effectively issue warnings. SUMMARY
[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a method for warning intruders outside a computer room based on face recognition to solve the problems raised in the background art.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solution: a computer room intruder early warning method based on face recognition, comprising the following steps: S1: face image enhancement, S2: face feature extraction, S3: real-time face detection, S4: face key point detection, S5: face occlusion detection, and S6: real-time face recognition;
[0007] S1: Face image enhancement: Monitor the target computer room in real time, obtain the face image in the target computer room monitoring screen, and automatically enhance the face image through image enhancement technology;
[0008] S2: Facial feature extraction: Input the enhanced face image into the YOLOV5-Face+ECANet network and extract facial features based on the YOLOV5-Face network;
[0009] S3: Real-time face detection: Based on the faces appearing in the target room surveillance footage, the detection algorithm in the YOLOV5-Face+ECANet network is used to detect the face position and frame the face in real time;
[0010] S4: Face key point detection: Based on the YOLOV5-Face+ECANet network, the key point regression branch is used to detect the key points of the facial features of the face image entering the frame, output the key point detection results, and use the Wing loss function to calculate the loss of the key point detection;
[0011] S5: Face occlusion detection: Identify occluded areas based on key point detection results, use the attention mechanism ECANet to learn the unoccluded face parts, and increase the training weight of the unoccluded parts;
[0012] S6: Real-time face recognition: The extracted facial features are compared with the existing faces in the face database to obtain the face recognition accuracy predicted by the network. The predicted face accuracy is compared with the set accuracy threshold. If it is greater than the set accuracy threshold, the identity corresponding to the face is returned. If it is less than the set accuracy threshold, a warning text message is issued to the page.
[0013] Preferably, the execution content of the facial image enhancement is as follows:
[0014] The target computer room is monitored in real time using monitoring equipment to obtain facial images in the monitoring screen of the target computer room. The facial images are enhanced using the MSRCR algorithm based on Retinex theory. The enhanced facial images are input into the YOLOV5-Face+ECANet network for the next step of face detection and face recognition.
[0015] Preferably, the facial feature extraction is performed as follows:
[0016] The enhanced face image is input into the YOLOV5-Face+ECANet network, and the face features are extracted based on the YOLOV5-Face network. A lightweight attention module ECANet is added in the C3 module of the YOLOV5-Face backbone network. The implementation principle of the lightweight attention module ECANet is as follows: the feature map with an input dimension of H*W*C is subjected to spatial feature compression through global average pooling to obtain a feature map with a dimension of 1*1*C. Then, the compressed feature map is subjected to 1*1 convolution to learn the importance between different channel features, and an attention feature map with a size of 1*1*C is obtained. Finally, the attention feature map with a size of 1*1*C is multiplied with the original input feature H*W*C channel by channel to obtain a final feature map with channel attention. Wherein, H represents the height of the feature map, that is, the number of pixels corresponding to the vertical direction; W represents the width of the feature map, that is, the number of pixels corresponding to the horizontal direction; and C represents the number of channels of the feature map, representing the feature dimension of each spatial position. Figure 1
[0017] Preferably, the execution steps of the face key point detection are specifically as follows:
[0018] S41: On the basis of the YOLOV5-Face network, a key point regression branch is added, which runs in parallel with the target detection branch. The detected face region is input into the YOLOV5-Face+ECANet network, and multi-scale feature maps are extracted through convolution layers;
[0019] S42: In the feature extraction process, the ECANet channel attention module is introduced to enhance the weight of the key feature channel. Specifically, the feature map is subjected to global average pooling to obtain a channel descriptor. A one-dimensional convolution is used to capture the inter-channel dependency to generate channel weights and apply them to the original feature map;
[0020] S43: A key point prediction layer is added at the end of the YOLOV5-Face+ECANet network. A fully connected layer or a convolution layer is used to map the features to 2k values. The predicted coordinates are normalized, the normalized predicted coordinates are converted into actual positions in the original image coordinate system, and the key point detection result is output. The detected key points are marked on the face image, and the Wing loss loss function is used to calculate the loss of key point detection.
[0021] Preferably, the execution steps of the face key point detection are specifically as follows:
[0022] S51: Obtain the face key point detection result in S4, including the position coordinates and confidence values of the facial feature points;
[0023] S52: Occlusion judgment is performed on each key point: if the confidence value of the key point is lower than a preset confidence threshold T1, the key point is marked as being occluded; if the confidence value of the key point is higher than or equal to the preset confidence threshold T1, the key point is marked as not being occluded;
[0024] S53: Based on the occlusion state of the key points, an occlusion mask graph of the face region is generated: the face region is divided into a plurality of sub-regions, each sub-region being associated with a specific key point; for a sub-region associated with a key point marked as being occluded, the pixel value at the corresponding position in the occlusion mask graph is set to M1; for a sub-region associated with a key point marked as not being occluded, the pixel value at the corresponding position in the occlusion mask graph is set to M2, and M1 and M2 are not equal.
[0025] S54: The occlusion mask graph and the face feature graph are fused, and the attention mechanism ECANet is used to learn the unoccluded face part and increase the training weight of the unoccluded part.
[0026] Preferably, the specific content of the fusion of the occlusion mask graph and the face feature graph is: the ECANet attention mechanism is applied to the face feature graph to generate channel attention weights; the channel attention weights are adjusted according to the occlusion mask graph: the feature channel weight corresponding to the unoccluded region is multiplied by an enhancement coefficient a; the feature channel weight corresponding to the occluded region is multiplied by an inhibition coefficient b.
[0027] Preferably, the execution content of the real-time face recognition is as follows:
[0028] Based on the YOLOV5-Face+ECANet network, the extracted face features are compared with the existing face features in the face library to obtain the network prediction face recognition accuracy Q, and the predicted face accuracy Q is compared with the set accuracy threshold T, if the predicted face accuracy Q is greater than the set accuracy threshold T, the identity corresponding to the face is returned, if the predicted face accuracy Q is less than the set accuracy threshold T, an early warning text information is sent to the page, and early warning is performed.
[0029] Preferably, the loss function of the YOLOV5-Face+ECANet network is: wherein Loss represents the total loss function of the entire network, represents the key point positioning loss, represents the sum of the positioning loss, the confidence loss and the classification loss in the network, the positioning loss is calculated using GIoU Loss, and the confidence loss and the classification loss are calculated using the cross-entropy loss function.
[0030] As described above, the machine room external invasion personnel early warning method based on face recognition provided by the present application has at least the following beneficial effects:
[0031] The face recognition-based machine room outsider warning method provided by the application improves the image quality by real-time monitoring of a target machine room, accurately capturing a face image in a monitoring picture, and automatically optimizing the face image using image enhancement technology. Then, the enhanced face image is input into a YOLOV5-Face+ECANet network, face features are extracted based on the YOLOV5-Face network to provide key information for subsequent recognition. Meanwhile, the face position in the monitoring picture is detected in real time using a detection algorithm in the network, and the face area is accurately framed. The face image in the frame is detected for facial feature points by a key point regression branch, the detection result is output, and the Wing loss function is used to calculate the loss of key point detection to improve the detection accuracy. The occlusion area is identified according to the key point detection result, the attention mechanism ECANet is used to focus on the unoccluded face part, the training weight is increased, and the learning of effective face information is strengthened. Finally, the extracted face features are compared with the existing face in the face library to obtain the face recognition accuracy predicted by the network. If the accuracy is greater than a set threshold, the identity information corresponding to the face is returned; if the accuracy is less than the set threshold, warning text information is sent to a related page. The application greatly improves the accuracy of the machine room outsider warning, reduces the false positive and false negative rates, effectively deals with complex situations such as face occlusion, enhances the robustness and applicability of the system, realizes real-time warning, and helps to prevent outsider behavior in time, thereby protecting the safety of the core equipment and data in the machine room and improving the overall safety management level of the bank data center machine room. BRIEF DESCRIPTION OF DRAWINGS
[0032] The application will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the application. Other drawings can be obtained by those of ordinary skill in the art without creative labor on the basis of the following drawings.
[0033] Figure 1 The figure is a flowchart of the face recognition-based machine room outsider warning method. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the protection scope of the application.
[0035] Please refer to Figure 1As shown, the present application provides a machine room outside personnel early warning method based on face recognition, which comprises the following steps: S1: face image enhancement, S2: face feature extraction, S3: real-time face detection, S4: face key point detection, S5: face occlusion detection and S6: real-time face recognition.
[0036] S1: face image enhancement: real-time monitoring of the target machine room, obtaining the face image in the target machine room monitoring picture, and automatically enhancing the face image through image enhancement technology;
[0037] In this embodiment, it needs to be specifically pointed out that the execution content of the face image enhancement is as follows:
[0038] The target machine room is monitored in real time by using the monitoring equipment, the face image in the target machine room monitoring picture is obtained, the MSRCR algorithm of Retinex theory is used to enhance the face image, and the enhanced face image is input into the YOLOV5-Face+ECANet network for the next step of face detection and face recognition.
[0039] Before the image is input into the network, the image is first enhanced by using the MSRCR algorithm of Retinex theory, in order to reduce the time consumption of image enhancement as much as possible on the basis of ensuring the image enhancement effect, the parameters in the MSRCR algorithm are set, specifically: the scale number is set to 3, Dynamic is set to 2, and finally the enhanced image is input into the YOLOV5-Face+ECANet network to realize face detection and face recognition.
[0040] The core idea of Retinex theory (retina cortex theory) is: by removing the interference of light component L(x, y), the reflection component R(x, y) is retained, so as to restore the real features of the image. For example, reducing the light effect on overexposed areas and enhancing the details in dark areas.
[0041] The Dynamic parameter represents the dynamic range control parameter, in the MSRCR algorithm, the Dynamic parameter directly affects the intensity of the color recovery factor C(x, y), which is used to adjust the dynamic range (light-dark contrast) and color fidelity of the enhanced image.
[0042] It needs to be specifically pointed out that the machine room monitoring picture often causes uneven environmental light (such as strong light from the window, shadow of the equipment), which leads to overexposure or dark of the face image. Such low-quality images will significantly reduce the face recognition accuracy. The core goal of image enhancement technology is to optimize the pixel value distribution through algorithm, restore the face details (such as facial features, texture), and provide high-quality input for subsequent detection and recognition.
[0043] S2: Face feature extraction: input the enhanced face image into the YOLOV5-Face+ECANet network, and extract the face features based on the YOLOV5-Face network;
[0044] In this embodiment, it needs to be specifically pointed out that the execution content of the face feature extraction is specifically as follows:
[0045] In order to meet the real-time face detection while not reducing the performance of face recognition, a lightweight attention module ECANet is added in the C3 module of the YOLOV5-Face backbone network.
[0046] The enhanced face image is input into the YOLOV5-Face+ECANet network, and the face features are extracted based on the YOLOV5-Face network. A lightweight attention module ECANet is added in the C3 module of the YOLOV5-Face backbone network. The implementation principle of the lightweight attention module ECANet is as follows: the feature map with an input dimension of H*W*C is spatially compressed by global average pooling to obtain a feature map with a dimension of 1*1*C. Then, the compressed feature map is processed by 1*1 convolution to learn the importance between different channel features, and an attention feature map with a size of 1*1*C is obtained. Finally, the attention feature map 1*1*C is multiplied with the original input feature H*W*C channel by channel to obtain a final feature map with channel attention. Wherein, H represents the height of the feature map, that is, the number of pixels corresponding to the vertical direction; W represents the width of the feature map, that is, the number of pixels corresponding to the horizontal direction; C represents the number of channels of the feature map, which represents the feature dimension of each spatial position (pixel point). Figure 1
[0047] In this embodiment, it needs to be specifically pointed out that the YOLOV5-Face+ECANet realizes fine and effective face feature information extraction on the basis of lightweight network, which not only improves the accuracy of face detection and recognition, but also meets the real-time detection requirements.
[0048] S3: Real-time face detection: according to the face appearing in the target machine room monitoring picture, the face position is detected in real time and the face is framed by using the detection algorithm in the YOLOV5-Face+ECANet network;
[0049] S4: Face key point detection: based on the YOLOV5-Face+ECANet network, the face feature key points of the framed face image are detected through the key point regression branch, and the key point detection result is output, and the Wing loss loss function is used to calculate the loss of key point detection.
[0050] In this embodiment, it needs to be specifically pointed out that the execution steps of the face key point detection are specifically as follows:
[0051] S41: On the basis of YOLOV5-Face network, a key point regression branch is added, which runs in parallel with the target detection branch, and the detected face region is input into the YOLOV5-Face+ECANet network, and multi-scale feature maps are extracted through convolutional layers;
[0052] S42: In the feature extraction process, the ECANet channel attention module is introduced to enhance the weight of the key feature channel, specifically: the channel descriptor is obtained by global average pooling of the feature map, the one-dimensional convolution is used to capture the inter-channel dependency, the channel weight is generated and applied to the original feature map;
[0053] S43: A key point prediction layer is added at the end of the YOLOV5-Face+ECANet network, which usually contains 5 key points (eyes, nose tip, and corners of the mouth), and a fully connected layer or a convolutional layer is used to map the features to 2k values (k is the number of key points, and each point has x, y coordinates), the predicted coordinates are normalized to adapt to input images of different sizes, the normalized coordinates are converted to actual positions in the original image coordinate system, and the key point detection result is output, the detected key points are marked on the face image to facilitate subsequent occlusion analysis, and the Wing loss loss function is used to calculate the loss of key point detection.
[0054] The expression of the Wing loss loss function is as follows:
[0055] , wherein Wing(x) represents the Wing Loss loss function, wln() represents the weighted w and 1n(), 1n() represents the logarithmic function with e as the base, e represents the curvature of the constraint nonlinear region, w limits the range of the nonlinear part to , C represents a constant that connects the linear part and the nonlinear part defined by the segment, .
[0056] S5: Face occlusion detection: According to the key point detection result, the occlusion area is identified, and the attention mechanism ECANet is used to learn the part of the face that is not occluded, and the training weight of the part that is not occluded is increased;
[0057] In this embodiment, it needs to be specifically pointed out that the execution steps of the face occlusion detection are as follows:
[0058] S51: Obtain the face key point detection result in S4, including the position coordinates and confidence values of the facial feature key points;
[0059] S52: Occlusion judgment is performed on each key point: if the confidence value of the key point is lower than a preset confidence threshold T1, the key point is marked as being occluded; if the confidence value of the key point is higher than or equal to the preset confidence threshold T1, the key point is marked as not being occluded;
[0060] S53: An occlusion mask map of the face region is generated based on the occlusion state of the key points: the face region is divided into a plurality of sub-regions, each of which is associated with a specific key point; for a sub-region associated with a key point marked as being occluded, the pixel value at the corresponding position in the occlusion mask map is set to M1; for a sub-region associated with a key point marked as not being occluded, the pixel value at the corresponding position in the occlusion mask map is set to M2, and M1 and M2 are not equal;
[0061] S54: The occlusion mask map is fused with the face feature map, and the attention mechanism ECANet is used to learn the part of the face that is not occluded, and the training weight of the part that is not occluded is increased;
[0062] In this embodiment, it needs to be specifically pointed out that the specific content of fusing the occlusion mask map with the face feature map is: applying the ECANet attention mechanism to the face feature map to generate channel attention weights; adjusting the channel attention weights according to the occlusion mask map: multiplying the feature channel weight corresponding to the unoccluded area by an enhancement coefficient a (a>1); multiplying the feature channel weight corresponding to the occluded area by an inhibition coefficient b (0<b<1);
[0063] In this embodiment, it needs to be specifically pointed out that based on the key point detection in the YOLOV5-Face+ECANet network, according to the face occlusion condition, the attention mechanism ECANet is used to learn the part of the face that is not occluded, and the training weight of the part that is not occluded is increased, so that the model pays more attention to the unoccluded face region, reduces the interference brought by various occlusions as much as possible, and further improves the recognition accuracy of the face under the occlusion condition.
[0064] S6: Real-time face recognition: the extracted face features are compared with the existing faces in the face library to obtain the network predicted face recognition accuracy, and the predicted face accuracy is compared with the set accuracy threshold, if greater than the set accuracy threshold, the identity corresponding to the face is returned, if less than the set accuracy threshold, a warning text information is sent to the page.
[0065] In this embodiment, it needs to be specifically pointed out that the execution content of the real-time face recognition is as follows:
[0066] Based on the YOLOV5-Face+ECANet network, the extracted face features are compared with the existing face in the face library, the network prediction face recognition accuracy Q is obtained, the predicted face accuracy Q is compared with the set accuracy threshold T, if the predicted face accuracy Q is greater than the set accuracy threshold T, the identity corresponding to the face is returned, if the predicted face accuracy Q is less than the set accuracy threshold T, the warning text information of the page is sent, and the warning is carried out, the warning text information is "the person is not in the library, please pay attention!".
[0067] The loss function of the YOLOV5-Face+ECANet network is: Wherein, Loss represents the total loss function of the whole network, represents the key point positioning loss, represents the sum of positioning loss, confidence loss and classification loss in the network, the positioning loss is calculated using GIoU Loss, the confidence loss and the classification loss are calculated using cross-entropy loss function;
[0068] The expression of the GIOU Loss is: Wherein, A, B represent any two frames, C represents the shape containing AB;
[0069] The expression of the cross-entropy loss function is: Wherein, x represents a sample, y represents a label, a represents a predicted output, and n represents the total amount of samples.
[0070] Finally: the above only for the preferred embodiments of the present application, and does not limit the present application, any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application, should be included in the protection scope of the present application.
[0071] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this, any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A computer room intrusion warning method based on face recognition is characterized by: The following steps are involved: S1: Face image enhancement: Monitor the target computer room in real time, obtain the face image in the target computer room monitoring screen, and automatically enhance the face image through image enhancement technology; S2: Facial feature extraction: Input the enhanced face image into the YOLOV5-Face+ECANet network and extract facial features based on the YOLOV5-Face network; S3: Real-time face detection: Based on the faces appearing in the target room surveillance footage, the detection algorithm in the YOLOV5-Face+ECANet network is used to detect the face position and frame the face in real time; S4: Face key point detection: Based on the YOLOV5-Face+ECANet network, the key point regression branch is used to detect the key points of the facial features of the face image entering the frame, output the key point detection results, and use the Wing loss function to calculate the loss of the key point detection; S5: Face occlusion detection: Identify occluded areas based on key point detection results, use the attention mechanism ECANet to learn the unoccluded face parts, and increase the training weight of the unoccluded parts; S6: Real-time face recognition: The extracted facial features are compared with the existing faces in the face database to obtain the face recognition accuracy predicted by the network. The predicted face accuracy is compared with the set accuracy threshold. If it is greater than the set accuracy threshold, the identity corresponding to the face is returned. If it is less than the set accuracy threshold, a warning text message is issued to the page.
2. The face recognition-based early warning method for computer room intruders according to claim 1 is characterized by: The execution content of the face image enhancement is as follows: The target computer room is monitored in real time using monitoring equipment to obtain facial images in the monitoring screen of the target computer room. The facial images are enhanced using the MSRCR algorithm based on Retinex theory. The enhanced facial images are input into the YOLOV5-Face+ECANet network for the next step of face detection and face recognition.
3. The face recognition-based early warning method for computer room intruders according to claim 1 is characterized by: The execution content of the facial feature extraction is as follows: The enhanced face image is input into the YOLOV5-Face+ECANet network, facial features are extracted based on the YOLOV5-Face network, and a lightweight attention module ECANet is added to the C3 module of the YOLOV5-Face backbone network. The implementation principle of the lightweight attention module ECANet is as follows: the feature map with an input dimension of H*W*C is spatially compressed by global average pooling to obtain a feature map with a dimension of 1*1*C. Then, the compressed feature map is subjected to 1*1 convolution to learn the importance between different channel features, and an attention feature map of size 1*1*C is obtained. Finally, the attention feature map 1*1*C is multiplied channel by channel with the original input feature H*W*C to obtain the final feature map with channel attention; where H represents the height of the feature map, that is, the number of pixels corresponding to the vertical direction; W represents the width of the feature map, that is, the number of pixels corresponding to the horizontal direction; C represents the number of channels of the feature map, representing the feature dimension of each spatial position.
4. The face recognition-based early warning method for computer room intruders according to claim 1 is characterized by: The steps for performing the facial key point detection are as follows: S41: Based on the YOLOV5-Face network, a keypoint regression branch is added. This branch runs in parallel with the object detection branch, inputs the detected face area into the YOLOV5-Face+ECANet network, and extracts multi-scale feature maps through the convolutional layer; S42: In the feature extraction process, the ECANet channel attention module is introduced to enhance the weights of key feature channels. Specifically, the channel descriptors are obtained by global average pooling of feature maps, one-dimensional convolution is used to capture the dependencies between channels, and channel weights are generated and applied to the original feature maps. S43: Add a key point prediction layer at the end of the YOLOV5-Face+ECANet network, use a fully connected layer or a convolutional layer to map the features into 2k values, normalize the predicted coordinates, convert the predicted normalized coordinates to the actual positions in the original image coordinate system, output the key point detection results, mark the detected key points on the face image, and use the Wing loss function to calculate the loss of key point detection.
5. The method for early warning of intruders in a computer room based on face recognition according to claim 1 is characterized in that: The steps for performing face occlusion detection are as follows: S51: Obtain the facial key point detection results in S4, including the position coordinates and confidence values of the facial key points; S52: Perform occlusion judgment on each key point: if the confidence value of the key point is lower than the preset confidence threshold T1, mark the key point as occluded; if the confidence value of the key point is higher than or equal to the preset confidence threshold T1, mark the key point as not occluded; S53: Based on the occlusion status of the key points, an occlusion mask map of the face region is generated: the face region is divided into a plurality of sub-regions, each sub-region is associated with a specific key point; for a sub-region where the associated key point is marked as occluded, the pixel value of the corresponding position in the occlusion mask map is set to M1; for a sub-region where the associated key point is marked as unoccluded, the pixel value of the corresponding position in the occlusion mask map is set to M2, and M1 and M2 are not equal; S54: Fuse the occlusion mask map with the facial feature map, use the attention mechanism ECANet to learn the unoccluded face parts, and increase the training weight of the unoccluded parts.
6. The method for early warning of intruders in a computer room based on face recognition according to claim 5, characterized in that: The specific content of fusing the occlusion mask image with the facial feature image is as follows: applying the ECANet attention mechanism to the facial feature image to generate channel attention weights; Adjust the channel attention weight according to the occlusion mask map: the feature channel weight corresponding to the unoccluded area is multiplied by the enhancement coefficient a; The feature channel weight corresponding to the occluded area is multiplied by the suppression coefficient b.
7. The method for early warning of intruders in a computer room based on face recognition according to claim 1 is characterized in that: The execution content of the real-time face recognition is as follows: Based on the YOLOV5-Face+ECANet network, the extracted facial features are compared with the existing faces in the face library to obtain the network-predicted face recognition accuracy Q. The predicted face accuracy Q is compared with the set accuracy threshold T. If the predicted face accuracy Q is greater than the set accuracy threshold T, the identity corresponding to the face is returned. If the predicted face accuracy Q is less than the set accuracy threshold T, a warning text message is issued to the page and an early warning is issued at the same time.
8. The method for early warning of intruders in a computer room based on face recognition according to claim 1, characterized in that: The loss function of the YOLOV5-Face+ECANet network is: , where Loss represents the total loss function of the entire network, represents the key point localization loss, It represents the sum of the positioning loss, confidence loss, and classification loss in the network. GIoU Loss is used to calculate the positioning loss, and the cross entropy loss function is used to calculate the confidence loss and classification loss.