Smart campus small target face detection method and system based on YOLOv8
By improving on the YOLOv8 network, combined with variable attention mechanisms such as deep separable convolution and self-supervision, the accuracy and speed problems of small-target face detection in complex scenarios of smart campuses are solved, and the rapid and accurate recognition and real-time detection of multiple scenarios are achieved, which improves the unity of intelligent campus management.
Patent Information
- Application Number
- CN202411940048.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-03
AI Technical Summary
The existing technology has problems of low accuracy, slow speed, poor adaptability and high model complexity in small-target face detection in complex smart campus scenarios, and it is difficult to meet the needs of real-time and universality.
Using the smart campus small-target face detection method based on YOLOv8, face image recognition is performed through the improved YOLOv8 network, combining the variable attention mechanisms such as deep separation convolution, self-supervision, and mixed loss functions, the feature extraction and detection process is optimized to achieve fast and accurate recognition in multiple scenarios and different postures.
It improves the accuracy of face detection in small targets, meets real-time requirements, realizes universal detection in multiple scenarios, reduces system maintenance costs and complexity, and improves the unity of intelligent campus management.
Smart Images

Figure CN120088821A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of face detection, and particularly relates to a small target face detection method and system for smart campuses based on YOLOv8. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] In the field of small target face detection in complex scenarios of smart campuses, the existing technologies have the following defects:
[0004] (1) In terms of accuracy, the low resolution factor in complex scenarios causes poor performance.
[0005] The large number of people, occlusion, and light changes at the school gate, the complex background and dynamic changes of people in the road scene, the occlusion, uneven light, and rapid personnel flow in the corridor all make it difficult for traditional algorithms to accurately extract the features of small target faces, resulting in frequent misjudgments and missed detections, seriously affecting the reliability of face recognition and behavior analysis, and bringing potential safety hazards to campus security management.
[0006] (2) In terms of speed, the existing algorithms cannot meet the real-time requirements.
[0007] During peak hours at the school gate, in the vast monitoring areas of the road scene, and in the classroom with multiple students, the processing speed of the algorithm is too slow, causing problems such as congestion of personnel passage, delayed discovery of abnormal behaviors, and low teaching management efficiency, and unable to respond to various situations on campus in a timely manner.
[0008] (3) In terms of adaptability, the environmental differences in different scenarios are large, and the existing technologies lack generality.
[0009] The lighting, background, and personnel activity states in various scenarios on campus are diverse. Algorithms optimized for specific scenarios are difficult to work effectively in other scenarios, and parameters need to be frequently adjusted or algorithms need to be replaced, which not only increases the maintenance cost and complexity of the smart campus system but also hinders the realization of efficient and unified personnel detection management.
[0010] (4) In terms of model complexity, some high-precision algorithm models are complex and rely on high-performance hardware resources.
[0011] When large-scale deployment is carried out in smart campuses, ordinary surveillance cameras and low-configuration servers cannot meet their operating requirements, increasing the hardware cost and restricting their wide application, which is not conducive to the comprehensive promotion of campus intelligent management. These problems urgently require more effective technical solutions to solve in order to improve the performance of small target face detection in complex scenarios of smart campuses. Summary of the Invention
[0012] To solve the above problems, the present invention proposes a method and system for small target face detection in a smart campus based on YOLOv8. Aiming at the problem of low-resolution face detection in the complex scenarios of a smart campus, face images are recognized based on the improved YOLOv8, realizing the fast and accurate recognition of faces in the smart campus in multiple scenarios and different poses.
[0013] According to some embodiments, the first solution of the present invention provides a method for small target face detection in a smart campus based on YOLOv8, adopting the following technical solutions:
[0014] A method for small target face detection in a smart campus based on YOLOv8 includes:
[0015] Obtain face images in the smart campus;
[0016] According to the obtained face images and the small target face detection model, perform small target face detection in the smart campus;
[0017] Among them, the small target face detection model uses the YOLOv8 network, performs a spatial-to-depth transformation on the obtained face images to rearrange the pixel spatial blocks, merges the obtained face images of different channel groups in the channel dimension, uses depthwise separable convolutions to extract the image features of the obtained face images, optimizes and adjusts the parameters of the self-supervised equivariant attention mechanism and the weight loss function, and completes the small target face detection in the smart campus based on YOLOv8.
[0018] As a further technical limitation, the specific process of using depthwise separable convolutions to extract the image features of the obtained face images is as follows: perform a spatial-to-depth transformation on the convolutional feature map of the P2 layer in the feature extraction network of YOLOv8, rearrange the pixel spatial blocks of the face images, increase the number of channels, reduce the spatial dimension, merge the face image features of different channel groups obtained in the channel dimension, reduce the merged face image features obtained, and replace the P4 and P5 convolutional layers in the YOLOv8 feature extraction network with depthwise separable convolutions to complete the extraction of face image features.
[0019] Furthermore, after extracting the face image features, different levels of convolutional layers are used for the fusion of face image features, and the adopted feature fusion methods are weighted fusion, splicing fusion, or attention mechanism fusion.
[0020] As a further technical limitation, the weight loss function uses the slide loss function, adjusts the parameters of the adopted slide loss function, realizes the adjustment of the sensitivity of the small target face detail features, balances the learning of the face target detail features at different scales, and identifies the face target images under different poses, expressions, and lighting conditions.
[0021] Further, the slide loss function field and other weight loss functions are weighted and combined to obtain a hybrid loss function. The sensitivity of the small target face detail features is calculated based on the obtained hybrid loss function, and the optimal hybrid loss function is determined according to the obtained sensitivity.
[0022] As a further technical limitation, the self-supervised equivariant attention mechanism is used to compensate for the loss of the occluded face image features. By adjusting the weight coefficients in the attention calculation or setting thresholds, the parameters of the supervised equivariant attention mechanism are adjusted to complete the optimization and adjustment of the self-supervised equivariant attention mechanism.
[0023] According to some embodiments, the second solution of the present invention provides a small target face detection system for a smart campus based on YOLOv8, adopting the following technical solution:
[0024] A small target face detection system for a smart campus based on YOLOv8, comprising:
[0025] An acquisition module, configured to acquire face images of the smart campus;
[0026] A detection module, configured to perform small target face detection in the smart campus according to the acquired face images and the small target face detection model;
[0027] Wherein, the small target face detection model adopts the YOLOv8 network, the obtained face images are transformed from space to depth for rearrangement of pixel space blocks, the face images of different channel groups obtained are merged in the channel dimension, depthwise separable convolution is used to extract the image features of the acquired face images, and the parameters of the self-supervised equivariant attention mechanism and the weight loss function are optimized and adjusted to complete the small target face detection in the smart campus based on YOLOv8.
[0028] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium, adopting the following technical solution:
[0029] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in a small target face detection method for a smart campus based on YOLOv8 as described in the first solution of the present invention.
[0030] According to some embodiments, the fourth solution of the present invention provides an electronic device, adopting the following technical solution:
[0031] An electronic device, comprising a memory, a processor, and a program stored on the memory and running on the processor, and when the processor executes the program, it implements the steps in a small target face detection method for a smart campus based on YOLOv8 as described in the first solution of the present invention.
[0032] According to some embodiments, the fifth aspect of the present invention provides a computer program product, adopting the following technical solution:
[0033] A computer program product includes software code, and the program in the software code executes the steps in a method for detecting small-target faces in a smart campus based on YOLOv8 as described in the first aspect of the present invention.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] Due to the fixed convolution stride and pooling layer in the YOLOv8 feature extraction layer of the present invention, the loss of fine-grained information and low-efficiency feature representation occur. Aiming at the poor detection effect of small targets, in terms of improving accuracy, by optimizing the feature extraction network structure, such as feature fusion of shallow feature maps, adding space to depth convolution modules, and seam attention mechanisms, the key features of small-target faces with low resolution are effectively captured; adopting a lightweight design strategy, using depthwise separable convolution to reduce the amount of computation to meet the real-time requirements of multiple campus scenarios; introducing an adaptive learning strategy to automatically adjust parameters according to the characteristics of different campus scenarios; under different lighting and backgrounds, the preprocessing and feature extraction can be effectively optimized to accurately detect small-target faces, achieve general detection, reduce the system maintenance cost and complexity, and improve the unity of campus intelligent management. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings forming a part of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions thereof of this embodiment are used to explain this embodiment and do not constitute an improper limitation of this embodiment.
[0037] Figure 1 It is a flowchart of the method for detecting small-target faces in a smart campus based on YOLOv8 in the first embodiment of the present invention;
[0038] Figure 2 It is a structural block diagram of the system for detecting small-target faces in a smart campus based on YOLOv8 in the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0040] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0042] In the present invention, terms such as "upper", "lower", "left", "right", "front", "rear", "vertical", "horizontal", "side", "bottom", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only relational terms determined for the convenience of describing the structural relationship of each component or element of the present invention and do not specifically refer to any component or element of the present invention and should not be construed as a limitation to the present invention.
[0043] In the present invention, terms such as "fixed connection", "connected", "connected" should be understood in a broad sense, which may mean a fixed connection, an integral connection or a detachable connection; it may be directly connected or indirectly connected through an intermediate medium. For those relevant scientific research or technical personnel in the field, the specific meanings of the above terms in the present invention can be determined according to specific circumstances and should not be construed as a limitation to the present invention.
[0044] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0045] Embodiment 1
[0046] Embodiment 1 of the present invention introduces a small target face detection method for a smart campus based on YOLOv8.
[0047] As Figure 1 shown, a small target face detection method for a smart campus based on YOLOv8 includes:
[0048] Obtain a face image of the smart campus;
[0049] According to the obtained face image and the small target face detection model, perform small target face detection in the smart campus;
[0050] Among them, the small target face detection model uses the YOLOv8 network, performs a spatial-to-depth transformation on the obtained face image to rearrange the pixel spatial blocks, merges the obtained face images of different channel groups in the channel dimension, uses depthwise separable convolution to extract the image features of the obtained face image, optimizes and adjusts the parameters of the self-supervised equivariant attention mechanism and the weight loss function, and completes the small target face detection for the smart campus based on YOLOv8.
[0051] In this embodiment, in the convolutional neural network, feature fusion is performed on the feature maps of the second convolutional layer. In the downsampling convolution, depthwise separable convolution (space to depth-Conv, abbreviated as SPDC) is used to replace the original fixed-stride convolutional layer and pooling layer. It consists of a space-to-depth space and depth layer and a non-stride convolutional layer. By downsampling the channel dimension of the feature map and retaining information, information loss can be avoided. Between feature fusion and detection, a self-supervised equivariant attention mechanism seam is added to compensate for occlusion loss and improve the recall rate. Depthwise separable convolution is used in detection to improve the detection speed, and the slideloss function is used in the classification loss to improve the sensitivity to small face targets and details.
[0052] In this embodiment, the input image is divided into grids by the YOLOv8 algorithm, and features are extracted using a deep convolutional neural network; the coordinates and confidence scores of the bounding boxes are predicted in an Anchor-Free manner, and the class probabilities are predicted at the same time; the TaskAligned Assigner is used to assign positive and negative samples, and weighted selection is performed based on the classification and regression scores; the loss function Loss calculation includes the BCE Loss for classification, the Distribution Focal Loss for regression, and the CIoU Loss; non-maximum suppression (NMS) is used in post-processing to filter overlapping boxes, and its Head becomes a decoupled head, and the regression branch adopts a specific integral form. Training augmentation introduces the YOLOX strategy, which has different variants, and the size of the backbone network affects accuracy and speed.
[0053] In the backbone convolutional neural network of YOLOv8, feature fusion operations are performed on large images during the convolution process, enabling the network to better integrate feature information at different levels, enhancing the expression ability of multi-scale face target features, and thus laying a foundation for accurately detecting face targets of different sizes in the subsequent stage. An SPDC module is designed to improve the detection effect on low-resolution small targets. The steps of the SPDC algorithm in this embodiment are as follows:
[0054] (1) Perform a space-to-depth transformation on the convolutional feature map of the P2 layer of the YOLOv8 feature extraction network, rearrange the pixel space blocks, increase the number of channels to 4, and reduce the spatial dimension by 2 times;
[0055] (2) Merge different channel groups in the channel dimension;
[0056] (3) Perform an addition calculation on the merged feature map and the feature map processed by other operations;
[0057] (4) Use a convolution with a stride of 1 on the merged feature map to reduce the channel dimension to 3, and keep the spatial resolution as 1 / 2 of the original;
[0058] (5) Replace the P4 - P5 convolutional layers of the YOLOv8 feature extraction network with the SPDC algorithm.
[0059] In this embodiment, a seam attention mechanism is introduced between feature fusion and detection; in complex campus scenarios, faces are often occluded, which causes some feature information to be lost and affects the detection effect. The seam attention mechanism can effectively compensate for the losses caused by occlusion, enabling the network to pay more attention to the complete and key face feature regions, thereby improving the detection recall rate and reducing the situation of missed detections due to occlusion.
[0060] It should be noted that the calculation of the recall rate belongs to the prior art that those skilled in the art should know, and this embodiment will not elaborate on it here.
[0061] In the detection stage of this embodiment, depthwise separable convolutions are used to replace traditional convolution operations; depthwise separable convolutions divide the convolution process into two steps: depth convolution and pointwise convolution, greatly reducing the computational amount while maintaining good feature extraction capabilities; enabling the algorithm to quickly and accurately detect a large number of face targets in complex campus scenarios and meet the real - time requirements.
[0062] In the calculation of the classification loss in this embodiment, the slideloss function is used; face targets in campus scenarios have rich detailed features, and traditional loss functions may not be able to fully capture these detailed differences, resulting in inaccurate detection of small - target faces and details. The slideloss function can improve the network's sensitivity to small - target faces and detailed features, enabling the network to pay more attention to these key features during training, thereby improving the detection accuracy and better identifying face targets under different poses, expressions, and lighting conditions.
[0063] Example analysis
[0064] In this embodiment, a YOLOv8 network model based on the above optimizations is constructed, and the network parameters are initialized. A face image dataset containing different scales, poses, lighting conditions, and occluded situations in complex campus scenarios is prepared. The dataset is divided into a training set, a validation set, and a test set and allocated according to a certain ratio, for example, 80% for training, 10% for validation, and 10% for testing.
[0065] During the training process, the training set images are input into the network. After feature extraction by the backbone network, including feature fusion in the second convolutional layer, using the space to depth convolutional module in downsampling to replace each stride convolutional layer and each pooling layer, then the features are optimized through the seam attention mechanism, and then quickly detected through depthwise separable convolutions. Finally, the classification loss is calculated according to the slideloss function, and the network parameters are updated by backpropagation to continuously optimize the network model.
[0066] In this embodiment, a validation set is used to evaluate the model during the training process, and various indicators are monitored, such as accuracy, recall rate, mAP (mean average precision), etc. The training parameters, such as learning rate, number of iterations, etc., are adjusted according to the evaluation results to improve the model performance.
[0067] In this embodiment, the trained model is deployed to a campus monitoring system or other devices that require face detection. During actual operation, when inputting campus scene images, the model can quickly and accurately detect face targets therein, including small-scale faces and partially occluded faces, and output detection results, including face position, size, category and other information, providing strong support for applications such as campus security management and personnel statistics.
[0068] The small target face detection method for smart campus based on YOLOv8 in this embodiment can be used to elaborate on the detection effect with the following detection data in face detection in complex campus scenes:
[0069] In terms of accuracy, after testing, in the campus entrance scene, the accuracy of the traditional algorithm for detecting small target faces with low resolution is about 60%, while this algorithm can be improved to more than 85%; in the campus road scene, for small target faces more than 50 meters away from the camera, the accuracy of the traditional algorithm is less than 50%, and this algorithm can reach about 85%; in the case of uneven light in the campus corridor and 30% personnel occlusion, the accuracy of the traditional algorithm is about 55%, and this algorithm can reach more than 80%. This benefits from the effective processing of face features in different scenes by feature fusion, SPDC module and seam attention mechanism.
[0070] In terms of adaptability, in the switching test of different scenes, the accuracy fluctuation of this algorithm in scenes such as campus entrance, road, corridor, classroom, etc. does not exceed 5%, while the fluctuation of the traditional algorithm reaches more than 15%, indicating that this algorithm can adapt to various scenes without frequent parameter adjustment.
[0071] In practical applications, after this algorithm is deployed in the campus monitoring system, 100 face images in different randomly selected campus scenes are tested. The overall detection accuracy reaches 96%, the recall rate is increased to 86%, and the mAP value reaches 92%, which are significantly improved compared with the traditional algorithm, providing strong support for campus security management and personnel statistics and other work, and providing a reliable face detection technology guarantee for the construction of smart campus.
[0072] In this embodiment, due to the fixed convolution stride and pooling layer in the YOLOv8 feature extraction layer, the loss of fine-grained information and inefficient feature representation occur. Aiming at the poor detection effect of small targets and in terms of accuracy improvement, by optimizing the feature extraction network structure, such as feature fusion of shallow feature maps, adding space to depth convolution modules, and seam attention mechanisms, the key features of small target faces with low resolution can be effectively captured; adopting a lightweight design strategy, using depthwise separable convolutions to reduce the computational amount to meet the real-time requirements of multiple campus scenarios; introducing an adaptive learning strategy to automatically adjust parameters according to the characteristics of different campus scenarios; under different lighting and backgrounds, it can effectively optimize preprocessing and feature extraction, accurately detect small target faces, achieve general detection, reduce the system maintenance cost and complexity, and improve the unity of campus intelligent management.
[0073] Embodiment 2
[0074] Embodiment 2 of the present invention introduces a small target face detection system for smart campuses based on YOLOv8.
[0075] As Figure 2 shown, a small target face detection system for smart campuses based on YOLOv8 includes:
[0076] An acquisition module configured to acquire face images of the smart campus;
[0077] A detection module configured to perform small target face detection in the smart campus according to the acquired face images and the small target face detection model;
[0078] Among them, the small target face detection model uses the YOLOv8 network, performs a space-to-depth transformation on the obtained face images to rearrange pixel space blocks, merges the obtained face images of different channel groups in the channel dimension, uses depthwise separable convolutions to extract the image features of the acquired face images, optimizes and adjusts the parameters of the self-supervised equivariant attention mechanism and the weight loss function, and completes the small target face detection in the smart campus based on YOLOv8.
[0079] The detailed steps are the same as those of the small target face detection method for smart campuses based on YOLOv8 provided in Embodiment 1 and will not be elaborated here.
[0080] Embodiment 3
[0081] Embodiment 3 of the present invention provides a computer-readable storage medium.
[0082] A computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the small target face detection method for smart campuses based on YOLOv8 as described in Embodiment 1 of the present invention.
[0083] The detailed steps are the same as those of the small target face detection method based on YOLOv8 provided in the first embodiment, and will not be elaborated here.
[0084] Embodiment Four
[0085] Embodiment Four of the present invention provides an electronic device.
[0086] An electronic device includes a memory, a processor, and a program stored on the memory and running on the processor. When the processor executes the program, it implements the steps in the small target face detection method based on YOLOv8 as described in the first embodiment of the present invention.
[0087] The detailed steps are the same as those of the small target face detection method based on YOLOv8 provided in the first embodiment, and will not be elaborated here.
[0088] Embodiment Five
[0089] Embodiment Five of the present invention provides a computer program product.
[0090] A computer program product includes software code, and the program in the software code implements the steps in the small target face detection method based on YOLOv8 as described in the first embodiment of the present invention.
[0091] The detailed steps are the same as those of the small target face detection method based on YOLOv8 provided in the first embodiment, and will not be elaborated here.
[0092] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0093] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0096] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0097] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
[0098] The above description is only the preferred embodiments of this embodiment and is not used to limit this embodiment. For those skilled in the art, this embodiment can have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this embodiment shall be included within the protection scope of this embodiment.
Claims
1. A small target face detection method for smart campus based on YOLOv8, characterized in that: include: Get the face image of smart campus; Based on the acquired face images and small target face detection model, small target face detection is performed on the smart campus; Among them, the small target face detection model adopts the YOLOv8 network, transforms the obtained face image from space to depth to rearrange the pixel space blocks, merges the obtained face images of different channel groups in the channel dimension, and uses deep separable convolution to extract the image features of the obtained face image, optimizes and adjusts the parameters of the self-supervised equivariant attention mechanism and the weight loss function, and completes the smart campus small target face detection based on YOLOv8.
2. A smart campus small target face detection method based on YOLOv8 as described in claim 1, characterized in that: The specific process of extracting the image features of the facial image obtained by using the depthwise separable convolution is as follows: performing a space-to-depth transformation on the P2 layer convolution feature map of the feature extraction network in YOLOv8, rearranging the face image pixel space blocks, increasing the number of channels, reducing the spatial dimension, merging the obtained face image features of different channel groups in the channel dimension, reducing the obtained merged face image features, and replacing the P4 and P5 convolution layers in the YOLOv8 feature extraction network with the depthwise separable convolution to complete the extraction of face image features.
3. A smart campus small target face detection method based on YOLOv8 as described in claim 2, characterized in that: After extracting the facial image features, different levels of convolutional layers are used to fuse the facial image features. The feature fusion methods used are weighted fusion, splicing fusion or attention mechanism fusion.
4. A smart campus small target face detection method based on YOLOv8 as described in claim 1, characterized in that: The weight loss function adopts the slideloss function, and the parameters of the adopted slideloss function are adjusted to adjust the sensitivity of the small target face detail features, balance the learning of the detail features of face targets of different scales, and recognize the face target images under different postures, expressions and lighting conditions.
5. A smart campus small target face detection method based on YOLOv8 as described in claim 4, characterized in that: The slideloss function domain and other weight loss functions are weightedly combined to obtain a mixed loss function, the sensitivity of the small target face detail features is calculated based on the obtained mixed loss function, and the optimal mixed loss function is determined according to the obtained sensitivity.
6. A smart campus small target face detection method based on YOLOv8 as described in claim 1, characterized in that: The self-supervised equivariant attention mechanism is used to compensate for the loss of occluded facial image features. The parameters of the supervised equivariant attention mechanism are adjusted by adjusting the weight coefficient in the attention calculation or setting a threshold, thereby completing the optimization adjustment of the self-supervised equivariant attention mechanism.
7. A smart campus small target face detection system based on YOLOv8, characterized in that: include: An acquisition module, configured to acquire a face image of a smart campus; A detection module, which is configured to perform small target face detection in the smart campus according to the acquired face image and the small target face detection model; Among them, the small target face detection model adopts the YOLOv8 network, transforms the obtained face image from space to depth to rearrange the pixel space blocks, merges the obtained face images of different channel groups in the channel dimension, and uses deep separable convolution to extract the image features of the obtained face image, optimizes and adjusts the parameters of the self-supervised equivariant attention mechanism and the weight loss function, and completes the smart campus small target face detection based on YOLOv8.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the smart campus small target face detection method based on YOLOv8 are implemented as described in any one of claims 1-6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the smart campus small target face detection method based on YOLOv8 as described in any one of claims 1-6 are implemented.
10. A computer program product comprising software code, characterized in that The program in the software code executes the steps of the smart campus small target face detection method based on YOLOv8 as described in any one of claims 1-6.