Face detection method and device, storage medium and electronic equipment

Through the combination of backbone network, neck Neck network and head head network, the accuracy and applicability of face detection are solved, and an accurate focus report is generated to help parents and teachers understand students' learning situation.

CN120496141APending Publication Date: 2025-08-15NEW ORIENTAL EDUCATION & TECH GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510422706.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The face detection methods in the prior art lack objectivity, low manual feature extraction efficiency, and the Anchor based and Anchor free methods have problems such as detection performance sensitivity and high computing resource consumption, resulting in insufficient accuracy and applicability of face detection.

Method used

The combination of backbone Backbone network, neck Neck network and head head network is used to extract features through residual connections and bottleneck structures, and combine expressions, eye sockets and motion recognition to generate attention information and focus reports.

Benefits of technology

It improves the accuracy and applicability of face detection and can generate accurate concentration reports to help parents and teachers understand students' learning situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496141A_ABST
    Figure CN120496141A_ABST
Patent Text Reader

Abstract

The invention relates to a face detection method and device, a storage medium and electronic equipment, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-detected image; the to-be-detected image is input into a target detection model to obtain face information, the target detection model comprises a backbone network, a neck Neck network and a head network, and the backbone network is used for feature extraction based on residual connection and a bottleneck structure; and acquiring attention information corresponding to the face information, and generating a concentration report based on the attention information. According to the invention, the face information is obtained through the target detection model, and the corresponding concentration report is generated based on the face information, so that parents and teachers can be assisted to quickly know the learning condition of students.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a face detection method, device, storage medium, and electronic device. Background Art

[0002] Classroom learning is the primary way people acquire knowledge. Understanding the individual student status and effectively responding to or teaching the lesson is an effective way to improve learning efficiency. In other words, whether students are focused directly affects learning outcomes and teaching quality. Related technologies primarily rely on manual observation to understand student focus, but this method lacks objectivity. Furthermore, when there are many students, teachers must monitor each student's focus. Summary of the Invention

[0003] The purpose of the present disclosure is to provide a face detection method, device, storage medium and electronic device to solve the technical problems existing in the related art.

[0004] According to a first aspect of the present disclosure, a face detection method is provided, the method comprising: Obtain the image to be detected; Inputting the image to be detected into a target detection model to obtain face information, wherein the target detection model includes a backbone network, a neck network, and a head network, wherein the backbone network is used for feature extraction based on residual connection and bottleneck structure; Acquire attention information corresponding to the facial information, and generate a concentration report based on the attention information.

[0005] Optionally, inputting the image to be detected into a target detection model to obtain facial information includes: Using the residual connection and the bottleneck structure of the Backbone network to extract features of the image to be detected, to obtain a multi-scale feature map; Performing feature fusion on the multi-scale feature map based on the Neck network to obtain a candidate feature map; Determine the facial information corresponding to the candidate feature map according to the Head network.

[0006] Optionally, the Backbone includes five convolution modules, four cross-order partial feature fusion CSP2F modules and a spatial pyramid pooling SPPF module. The convolution module includes a convolution layer, a batch normalization layer and a ReLU activation function layer. The convolution module is used to perform sampling operations, the CSP2F module is used to decompose and process feature maps, and the SPPF module is used to aggregate spatial information at different scales.

[0007] Optionally, the CSP2F module is configured to perform the following operations: Based on a 1x1 convolution kernel, the number of channels of the feature map transmitted by the convolution module is reduced to half of the original number; Use multiple 3x3 convolution kernels to perform convolution operation on the processed feature map to extract feature information; The number of channels of the feature information is restored based on a 1x1 convolution kernel.

[0008] Optionally, the Neck network includes a feature pyramid network (FPN), and the step of performing feature fusion on the multi-scale feature map based on the Neck network to obtain a candidate feature map includes: The multi-scale feature map is upsampled and downsampled using the FPN to obtain the candidate feature map.

[0009] Optionally, the Head network adopts a decoupling head structure, and the decoupling head structure includes a regression head and a classification head; The determining, according to the Head network, the facial information corresponding to the candidate feature map includes: Generating multiple candidate boxes corresponding to the candidate feature maps; The position of each candidate frame is predicted based on the regression head, and each candidate frame is classified using the classification head to obtain the face information.

[0010] Optionally, obtaining attention information corresponding to the facial information includes: Performing expression recognition on the facial information to obtain an expression recognition result; Performing eye socket recognition on the facial information to obtain an eye socket recognition result; Performing action recognition on a target person corresponding to the facial information to obtain an action recognition result; The attention information is comprehensively determined based on the expression recognition result, the eye socket recognition result and the action recognition result.

[0011] In a second aspect, the present disclosure provides a face detection device, comprising: An acquisition module is configured to acquire an image to be detected; An input module is configured to input the image to be detected into a target detection model to obtain face information. The target detection model includes a backbone network, a neck network, and a head network. The backbone network is used for feature extraction based on residual connections and a bottleneck structure. The generation module is configured to obtain attention information corresponding to the facial information and generate a concentration report based on the attention information.

[0012] In a third aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the methods described in the first aspect.

[0013] In a fourth aspect, the present disclosure provides an electronic device, the electronic device comprising: a memory having a computer program stored thereon; A processor is used to execute the computer program in the memory to implement the steps of any one of the methods in the first aspect.

[0014] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the first aspect when executed by a processor.

[0015] After obtaining the image to be detected, the present invention inputs the image to be detected into the target detection model to obtain facial information, wherein the target detection model includes a backbone network, a neck network and a head network. The backbone network can be used to extract features based on residual connections and bottleneck structures. On this basis, attention information corresponding to the facial information output by the target detection model is obtained, and a concentration report is generated based on the attention information. Since the acquisition of facial information is based on the output of the target detection model, the attention information obtained based on the facial information is more accurate, and the concentration report is automatically generated through the attention information, which can assist parents or teachers to know the students' learning situation more clearly.

[0016] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 The figure is a flowchart of a face detection method according to an exemplary embodiment.

[0018] Figure 2 The figure is a structural example diagram of a target detection model in a face detection method according to an exemplary embodiment.

[0019] Figure 3 The figure is an example diagram of face information in a face detection method according to an exemplary embodiment.

[0020] Figure 4 FIG. 4 is another example diagram of facial information in a face detection method according to an exemplary embodiment.

[0021] Figure 5 The figure is an example diagram showing a method for detecting a face according to an exemplary embodiment of determining motion information based on face information.

[0022] Figure 6 The figure is a flowchart of another face detection method according to an exemplary embodiment.

[0023] Figure 7 The figure is a structural example diagram of a Backbone network in another face detection method according to an exemplary embodiment.

[0024] Figure 8 FIG. 4 is a block diagram showing a face detection method according to an exemplary embodiment of the present disclosure.

[0025] Figure 9 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] The following describes the specific embodiments of the present disclosure in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present disclosure and are not intended to limit the present disclosure.

[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0033] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0034] In related technologies, face detection is primarily based on manual feature extraction. Specifically, a region of interest (ROI) is selected, i.e., an area that may contain a face. Feature extraction is then performed on this area, and finally, face detection and classification are performed on the extracted features. However, progress in implementing face detection through manual feature extraction has been slow and performance has been poor. This includes poor recognition results, low accuracy, high computational effort, and slow processing speed. These drawbacks significantly reduce the accuracy of face detection.

[0035] Although face detection can be performed using Anchor-based and Anchor-free methods, both Anchor-based and Anchor-free methods have certain defects. Specifically, when performing face detection using Anchor-based methods, the size, number, and aspect ratio of the Anchors have a greater impact on the detection performance, that is, the detection performance of Anchor-based methods is sensitive to the size, number, and aspect ratio of the Anchors. Secondly, fixed Anchors will greatly undermine the universality of the face detector, resulting in the need to reset the size and aspect ratio of the Anchors for different people. Furthermore, in order to match the real frame, a large number of Anchors usually need to be generated, but most of the Anchors will be marked as negative samples during training, which will cause sample imbalance. In addition, during the training process, the face detection model usually needs to calculate the IOU between all Anchors and the real frame, which will consume a lot of memory and time.

[0036] In addition, face detection based on anchor-free usually requires complex post-processing, and faces that are too close may be missed.

[0037] To address the aforementioned issues, the present disclosure proposes a face detection method that utilizes a backbone network, a neck network, and a head network to comprehensively acquire facial information. Because the backbone network can extract features based on residual connections and a bottleneck structure, it can improve the accuracy of face detection to a certain extent. Furthermore, after acquiring facial information, generating a concentration report can help parents and teachers better understand their students' learning progress.

[0038] Figure 1 A face detection method according to an exemplary embodiment is shown. Figure 1 As shown, the face detection method may include the following steps: In step S110 , an image to be detected is acquired.

[0039] In the embodiments of the present disclosure, the image to be detected may include multiple target objects, which may be people, for example, students. Here, the image to be detected may be a real-time image captured by an automated acquisition system, such as an image captured by a camera, or a locally pre-stored image. Furthermore, the image to be detected may be a video frame, meaning that the embodiments of the present disclosure can detect faces in videos.

[0040] After acquiring the initial image, the embodiment of the present disclosure may preprocess the initial image, such as denoising, adjusting the white balance, and increasing the contrast of the initial image to obtain the image to be detected, thereby improving the accuracy of image detection.

[0041] Optionally, the embodiment of the present disclosure may obtain multiple candidate images and extract common features of the multiple candidate images. On this basis, the multiple candidate images are fused based on the common features to obtain the image to be detected. Here, the sources of the multiple candidate images may be different, such as the source of the first candidate image is the first image acquisition device, and the source of the second candidate image may be the second image acquisition device, wherein the first image acquisition device and the second image acquisition device may be installed in the same environment. For example, the first image acquisition device and the second image acquisition device may be installed in the same classroom, but the installation positions or installation orientations of the two are different.

[0042] In step S120, the image to be detected is input into the target detection model to obtain face information.

[0043] In the embodiment of the present disclosure, the target detection model may include: Figure 2 The backbone network 210, the neck network 220 and the head network 230 are shown. Figure 2 As can be seen, after receiving the input image (the image to be detected), the object detection model 200 can first use the backbone network 210 to extract features from the image to obtain a multi-scale feature map. Here, the backbone network 210 can use a series of convolutional and deconvolutional layers to extract features, and can also use residual connections and bottleneck structures to reduce the size of the network and improve performance.

[0044] Residuals are used to solve the vanishing gradient problem in deep neural networks, that is, as the number of network layers increases, the gradient may gradually decrease, resulting in the inability to effectively train the target detection model. The residual connection in the embodiment of the present disclosure enables the network to learn "residuals" instead of directly learning the complete mapping from input to output by introducing direct jump connections between layers. In other words, the core of the residual connection is to add the input feature map directly to the output feature map through a "shortcut", thereby alleviating the vanishing gradient problem in deep network training.

[0045] The bottleneck architecture reduces computational complexity while maintaining model performance by reducing the amount of computation in intermediate layers. It first reduces the number of channels in the feature map through dimensionality reduction, then performs convolution operations, and finally restores the number of channels in the feature map. For example, the bottleneck architecture might include 1×1 convolution for channel compression, 3×3 convolution for feature processing, and then 1×1 convolution for channel expansion. This suggests that the bottleneck architecture aims to reduce computational complexity and parameter count while maintaining the network's expressive power.

[0046] As an example, the residual connection and bottleneck structure can be implemented through the CSP2F module (Cross Stage PartialTo Feature, cross-stage partial feature fusion module).

[0047] In this process, the embodiment of the present disclosure can use the CSP2F module (Cross Stage Partial ToFeature, cross-stage partial feature fusion module) as the basic component unit of the Backbone network, so that the backbone network can have fewer parameters and better feature extraction capabilities. Among them, the CSP2F module is used for staged processing of feature extraction in deep neural networks, which can effectively decompose and process feature maps to optimize the network's computational efficiency and representation capabilities. In other words, the CSP2F module can introduce richer representations while retaining important features, thereby improving the expressiveness of the target detection model.

[0048] In some implementations, after extracting features from the image to be detected using the Backbone network to generate a multi-scale feature map, the disclosed embodiments can fuse these multi-scale feature maps using the Neck network to generate candidate feature maps. Based on this, the Head network is used to regress and classify the candidate feature maps to generate a face frame.

[0049] Here, the face information (face frame) of the target object can be represented by a four-dimensional vector, which can represent the coordinates of the upper left corner and the lower right corner of the face frame respectively, as shown in detail. Figure 3 Based on Figure 3 As can be seen, when performing face detection on the image 300 to be detected, a face frame 310 can be obtained. This face frame 310 can be represented by the upper left corner coordinate A and the lower right corner coordinate B, where the upper left corner coordinate A is (x1, y1) and the lower right corner coordinate B is (x2, y2). In this case, the four-dimensional vector can be represented as (x1, y1, x2, y2).

[0050] Optionally, the face frame 310 can also be represented by the center coordinates O (x0, y0), width w and height h. In this case, the relationship between the face frame and the coordinates can be as follows: Figure 4 As shown, based on Figure 4 As we know, the four-dimensional vector can be represented as O(x0, y0, w, h). The specific representation of the face frame is not explicitly limited here and can be selected according to the actual situation. In the embodiment of the present disclosure, facial information can be represented by a four-dimensional vector.

[0051] In step S130, attention information corresponding to the facial information is obtained, and a concentration report is generated based on the attention information.

[0052] As an optional method, after obtaining facial information, the embodiment of the present disclosure can obtain attention information corresponding to the facial information, and on this basis, a concentration report can be generated based on the attention information.

[0053] Specifically, the disclosed embodiments can perform facial expression recognition on facial information to obtain an expression recognition result, perform eye socket recognition on the facial information to obtain an eye socket recognition result, and perform action recognition on the target person corresponding to the facial information to obtain an action recognition result. Based on this, the expression recognition results, eye socket recognition results, and action recognition results are integrated to obtain a target result, based on which attention information can be determined.

[0054] Here, the attention information can be represented by the degree of concentration, which can include serious, general, and not serious. Optionally, the attention information can also be represented by a percentage, such as a concentration of 95%, which means that the student's concentration is very high at this time.

[0055] In the disclosed embodiment, the expression recognition results may include anger, disgust, fear, happiness, sadness, surprise, and neutral. After obtaining the expression recognition results, the disclosed embodiment can preliminarily determine whether the student's attention is good or poor based on the expression recognition results. For example, if the student's expression is detected as disgust, it indicates that their attention is relatively poor. Conversely, if the student's expression is detected as happiness, their attention is relatively good.

[0056] As can be seen, after obtaining the expression recognition result, the embodiment of the present disclosure can determine whether it corresponds to a positive expression or a negative expression. If the expression recognition result is a positive expression (including a neutral expression), it means that the student's learning state is relatively good at this time. Conversely, if the expression recognition result is a negative expression, it means that the student's learning state is relatively poor at this time.

[0057] The eye socket recognition results include sleepy state, normal state and excited state. During the eye socket recognition process, the embodiment of the present disclosure can obtain the length of the short side of the eye socket, which is used to represent the degree of eye openness. The larger the short side length, the wider the eyes are open, and the students are usually in an excited state at this time.

[0058] Specifically, the disclosed embodiment can determine whether the short side length is less than a first threshold. If so, the corresponding student is determined to be drowsy. If the short side length is greater than or equal to the first threshold and less than or equal to a second threshold, the corresponding student is determined to be in a normal state. If the short side length is greater than the second threshold, the corresponding student is determined to be in an excited state. Based on the student's state, the corresponding attention information can be determined. Generally, when a student is drowsy, the corresponding attention information is lower, while when they are excited, the corresponding attention information is higher.

[0059] Motion recognition is used to identify and understand the behavior and actions of people in images or video sequences, that is, to identify the action information of students during class. Here, the action information may include writing, listening to the class, and raising hands, etc. In the process of detecting student actions, the embodiment of the present disclosure can collect multiple video frames, and then perform action recognition on the same person in these video frames, and determine the target person's action change rate within a specified time period based on the recognition results. If the action change rate exceeds the first preset threshold, it means that the student is not listening carefully to the class and is always moving around. In addition, if the action change rate is less than the second preset threshold, it means that the student may have fallen asleep. At this time, the eye socket recognition result and the action recognition result can be combined to comprehensively obtain attention information.

[0060] Here, action recognition can be obtained by the target person's hand changes, footstep changes, head changes and overall body changes corresponding to the face information, such as Figure 5As shown, after detecting the facial information 501 , the embodiment of the present disclosure can obtain the posture information corresponding to the facial information 501 . The posture information can be obtained by identifying and analyzing the action corresponding to the area 502 .

[0061] It should be noted that the action recognition result can be obtained by identifying the currently collected image to be detected, or by comprehensively analyzing multiple images (video frames) before the image to be detected.

[0062] It can be seen that the embodiment of the present disclosure can obtain attention information through at least one of the expression recognition results, eye socket recognition results, and motion recognition results. For example, when the expression recognition result is detected to be neutral and the eye socket recognition result is determined to be a sleepy state, if the motion change rate is determined to be less than the second preset threshold, it means that the student's attention is extremely inattentive at this time. However, in this process, the embodiment of the present disclosure can also first obtain the student's historical behavior and determine the student's personality based on the historical behavior. If the student's personality is determined to be lively / naughty, then when the expression recognition result is determined to be neutral, the eye socket recognition result is a sleepy state, and the motion change rate is less than the second preset threshold, it is determined that the student is inattentive.

[0063] On the contrary, if the student's personality is determined to be quiet, and when he listens attentively in class, he often closes his eyes to think and his movements change little, then the student's attention is determined to be focused, which can ensure the accuracy of attention information acquisition.

[0064] It should be noted that in the process of obtaining attention information corresponding to facial information, the disclosed embodiments can also determine the type of class. If the class type is offline, then when it is determined that the student's expression recognition result is neutral, the eye socket recognition result is a sleepy state, and the movement change rate is less than a second preset threshold, it can continue to detect whether the student is wearing headphones. If the student is wearing headphones, it can be further determined that the student is not paying attention. The main reason is that students in online courses usually need to wear headphones. Therefore, by detecting the class type, it is possible to avoid errors in obtaining attention information. Among them, whether the student is wearing headphones and at least one of the above three conditions are met can determine that the student is not paying attention.

[0065] Furthermore, the disclosed embodiment can also identify the eye information of the target object (student), that is, perform pupil recognition, and analyze whether the student is playing with an electronic device based on the information in the pupil. If the information in the pupil matches the preset information, it is determined that the student is playing with the electronic device, where the preset information can be the brightness information of the human eye when playing with the electronic device normally.

[0066] In summary, in the process of obtaining attention information corresponding to facial information, the disclosed embodiments can determine the target object's attention information by obtaining at least one of the following: expression recognition results, eye socket recognition results, action recognition results, ear recognition results, and pupil recognition results. For example, if the expression recognition results determine that Student A's attention information is 90 points, and the eye socket recognition results determine that Student A's attention information is 80 points, then the average of the two can be taken, i.e., Student A's attention information is 85 points.

[0067] It should also be noted that, in the process of obtaining attention information corresponding to facial information, the embodiment of the present disclosure can also obtain course information, that is, the type of course. On this basis, the attention information is comprehensively determined according to the course type and at least one of the above five conditions.

[0068] As an example, when the acquired course is "Computer Operation", pupil recognition results may not be used as a condition for obtaining attention information. As another example, when the acquired course is "English Listening", ear recognition results may not be used as a condition for obtaining attention information. As another example, when the acquired course is "Physical Education", action recognition results may not be used as a condition for obtaining attention information. This can avoid the incorrect acquisition of attention information to a certain extent.

[0069] After acquiring the image to be detected, the embodiment of the present disclosure inputs the image to be detected into the target detection model to obtain facial information, wherein the target detection model includes a backbone network, a neck network and a head network. The backbone network can be used to perform feature extraction based on residual connections and bottleneck structures. On this basis, attention information corresponding to the facial information output by the target detection model is acquired, and a concentration report is generated based on the attention information. Since the acquisition of facial information is based on the output of the target detection model, the attention information acquired based on the facial information is more accurate, and the concentration report is automatically generated through the attention information, which can assist parents or teachers in knowing the students' learning situation more clearly.

[0070] Figure 6 Another face detection method according to an exemplary embodiment is shown. Figure 6 As shown, the face detection method may include the following steps: In step S610, an image to be detected is acquired.

[0071] The specific implementation of step S610 has been described in detail in the above embodiment and will not be repeated here.

[0072] In step S620, the residual connection and bottleneck structure of the Backbone network are used to extract features of the image to be detected to obtain a multi-scale feature map.

[0073] From the above technical solution, it can be seen that the target detection model may include a backbone network, a neck network and a head network. Among them, the backbone can be used to extract features of the image to be detected to obtain a multi-scale feature map; the neck network is used to fuse the multi-scale feature maps to obtain a candidate feature map; the head network is used to regress and classify the candidate feature maps to obtain the face frame of the face.

[0074] In the embodiment of the present disclosure, Backbone may include a convolution module, a cross-order partial feature fusion CSP2F module, and a spatial pyramid pooling SPPF (Spatial Pyramid Pooling Feature) module. Specifically, the structure of the Backbone network can be as follows: Figure 7 As shown, based on Figure 7 It is known that Backbone can include five convolution modules 211 (ConvModule), four cross-order partial feature fusion CSP2F modules 212 and one spatial pyramid pooling SPPF module 213.

[0075] Among them, the convolution module 211 is used to extract low-level to high-level features from the input image to be detected, that is, through layer-by-layer convolution operations, the target detection model can gradually capture information such as edges, textures, and shapes in the image. Here, the output of the convolution module 211 can be a multi-scale feature map.

[0076] That is, the convolution module 211 can be used for downsampling, which can include a convolution layer, a batch normalization layer, and a ReLU activation function layer. Specifically, the convolution layer in each convolution module 211 in Backbone 210 can use a convolution kernel with a stride of 2 to perform a downsampling operation to reduce the size of the feature map and increase the number of channels.

[0077] In addition, the convolution module 211 can also be responsible for nonlinear representation, that is, a batch normalization layer and a ReLU activation function layer can be added after each convolution layer to enhance the nonlinear representation capability of the target detection model.

[0078] In some embodiments, the CSP2F module 212 is used to decompose and process the feature map. Specifically, the CSP2F module 212 may use a 1x1 convolution kernel to reduce the number of input channels to half, thereby reducing computational complexity and memory consumption. Based on this, multiple 3x3 convolution kernels are used to perform convolution operations to extract feature information. Subsequently, a residual connection is used to directly add the input to the output, forming a cross-layer link. Finally, a 1x1 convolution kernel is used again to restore the number of channels in the feature map.

[0079] That is, the CSP2F module 212 can be used to perform the following operations: reducing the number of channels of the feature map transmitted by the convolution module to half of the original number based on a 1x1 convolution kernel; performing a convolution operation on the processed feature map using multiple 3x3 convolution kernels to extract feature information; and restoring the number of channels of the feature information based on a 1x1 convolution kernel.

[0080] In summary, CSP2F module 212 is used to decompose and fuse features across multiple stages. Specifically, CSP2F module 212 splits the input feature map into two parts: one part is directly passed to subsequent layers, and the other part is processed through convolution. The processed features are then merged with the directly passed features, thereby preserving the original features while enhancing their expressiveness. During face detection, CSP2F module 212 can effectively process multi-scale features, enabling the object detection model to effectively detect faces of varying sizes.

[0081] In other embodiments, the SPPF module 213 is used to aggregate spatial information at different scales, thereby capturing global features at different scales, which is beneficial for detecting faces of different sizes and positions.

[0082] In other words, after the convolution module 211 and the CSP2F module 212 perform feature extraction on the image to be detected, the embodiment of the present disclosure can use the SPPF module 213 to perform multi-scale pooling to further process the features extracted by the former to obtain a multi-scale feature map.

[0083] In step S630, feature fusion is performed on the multi-scale feature maps based on the Neck network to obtain candidate feature maps.

[0084] As an optional method, after obtaining the multi-scale feature map, the embodiment of the present disclosure can perform feature fusion on the multiple feature maps extracted by Backbone based on the Neck network to obtain candidate feature maps, that is, the Neck network can be used to fuse feature maps from different stages of Backbone to enhance the feature representation capability.

[0085] The Neck network can include a Feature Pyramid Network (FPN), which can gradually upsample feature maps from deep layers to shallow layers and combine abstract features from deep layers with detailed features from shallow layers. For example, each upsampled feature map is point-by-point added or concatenated with the shallow feature map of the corresponding layer to form a new feature map.

[0086] In other words, the disclosed embodiment can use FPN to upsample and downsample the multi-scale feature maps to obtain candidate feature maps. It can be seen that the Neck network is used to process and fuse the features extracted by the Backbone network to improve the accuracy and robustness of target detection.

[0087] Specifically, the Neck network can fuse multi-scale feature maps by introducing different structures and technologies to better capture information about targets of different scales. For example, the Neck network can use the feature pyramid network structure FPN to process the multi-scale feature maps from the Backbone network. Here, FPN can enable the model to detect targets at different scales by establishing feature pyramids at different levels, that is, through upsampling and downsampling operations, it can fuse low-level detail features with high-level semantic features to obtain a more comprehensive and rich feature representation.

[0088] Feature fusion combines features from different layers, improving the accuracy of object detection models, especially for objects of varying scales. Upsampling and downsampling adjust the scale and resolution of feature maps. Upsampling scales low-resolution feature maps to the same size as high-resolution ones, preserving more detail. Downsampling reduces the size of high-resolution feature maps to reduce computational overhead and memory consumption.

[0089] By utilizing operations such as feature pyramid networks and feature fusion, the Neck network can effectively extract and fuse multi-scale features, thereby improving the performance and robustness of face detection. This in turn enables the target detection model to better adapt to targets of different scales and sizes and achieve more accurate detection results in complex scenarios.

[0090] In step S640, the facial information corresponding to the candidate feature map is determined according to the Head network.

[0091] In the embodiment of the present disclosure, the Head network can be the last few layers of the target detection model, which can be used to generate the final detection results. The Head network can adopt a decoupling head structure, wherein the decoupling head structure can include a regression head and a classification head. In the process of determining the facial information corresponding to the candidate feature map based on the Head network, the embodiment of the present disclosure can first generate multiple candidate boxes corresponding to the candidate feature map, and then, based on this, predict the position of each candidate box based on the regression head, and use the classification head to classify each candidate box to obtain facial information.

[0092] As can be seen, the Head network is used for final face detection. After obtaining the candidate feature map after the Neck network fusion, the embodiment of the present disclosure can input the candidate feature map into the decoupling head for prediction. Specifically, the embodiment of the present disclosure can calculate the position offset between the predicted box and the true box, and then input the offset into the regression head for loss calculation, outputting a four-dimensional vector. Exemplarily, the four-dimensional vector can represent the coordinates x, y of the upper left corner and the coordinates x, y of the lower right corner of the face box, respectively.

[0093] On this basis, the classification head can perform RoI Pooling (Region of Interest Pooling) and convolution operations on the candidate box to obtain an output tensor. The value at each position in the output tensor represents the probability that the candidate box belongs to each category. Finally, the embodiment of the present disclosure can filter out the final detection result through non-maximum suppression (NMS). Among them, RoI Pooling is used to extract fixed-size features of the region of interest (RoI) from the candidate feature map and map candidate regions (RoI) of different sizes to a fixed-size feature map.

[0094] It should be noted that after obtaining the candidate feature map, the embodiment of the present disclosure can first extract candidate frames, such as using Anchor Free to extract multiple candidate frames. On this basis, the classification head can perform RoI Pooling and convolution on each candidate frame extracted by Anchor Free to obtain the probability of each category. It can be seen that the classification head is used to determine whether the detected candidate area contains the target object (such as a face) and further determine the category of the target object. In other words, the output of the classification head can be the probability (or score) of each candidate area containing the target object.

[0095] In addition, in addition to the regression head and the classification head, the decoupling head can also include a detection head. The regression head can be used to predict the position of each candidate box and output the bounding box coordinates of the candidate box; the classification head can be used to predict the category of each candidate box and output the probability distribution of each candidate box belonging to each category; the detection head can be used to integrate the outputs of the regression head and the classification head to generate the final face detection result, that is, to obtain face information.

[0096] In summary, after obtaining the candidate feature map after Neck fusion, the Head network can first generate multiple candidate boxes based on the candidate feature map. Based on this, the regression head generates the location information of the candidate boxes, and the classification head generates the category probabilities of the candidate boxes. Finally, the detection head combines the outputs of the regression and classification heads to output the final detection result containing the face location and category.

[0097] Here, the structures of the detection head and the regression head are similar. For example, the structure of the detection head may include four 3×3 convolutions and two 1×1 convolutions.

[0098] The disclosed embodiments utilize an object detection model to accurately capture facial information. Because the object detection model is trained on massive amounts of data, it effectively addresses occlusions and extreme viewing angles, resulting in high face detection accuracy. Furthermore, the disclosed embodiments are applicable to a variety of complex face detection scenarios, demonstrating high applicability. Furthermore, the disclosed embodiments support simultaneous multi-user access to services and elastic scalability, demonstrating robustness.

[0099] In step S650, attention information corresponding to the facial information is obtained, and a concentration report is generated based on the attention information.

[0100] It can be seen from the above technical solution that after obtaining facial information, the embodiment of the present disclosure can obtain attention information corresponding to the facial information, and on this basis, generate a concentration report based on the attention information.

[0101] Furthermore, the disclosed embodiment can obtain attention information by combining at least one of expression recognition results, eye socket recognition results, action recognition results, ear recognition results, and pupil recognition results, that is, the attention information can be determined by at least one of the above recognition results.

[0102] As an example, the attention information can be represented by an attention level, where a higher attention level indicates better attention of the student. For example, the attention level is divided into five levels, with level one indicating the lowest attention and level five indicating the highest attention.

[0103] As another example, attention information can also be represented by an attention score, where a higher attention score indicates better attention. For example, an attention score of 100 indicates the highest level of attention, while an attention score of 0 indicates the lowest level of attention. The specific method for representing attention information is not specified here and can be selected based on actual circumstances.

[0104] In the disclosed embodiment, the attention report (concentration test report) may include an attention score, an attention comment, and attention changes, etc. Among them, the attention score may be an overall attention score; the attention comment may be generated by combining historical attention, that is, the disclosed embodiment may perform a comprehensive analysis of the target student's attention information within a first specified time period to obtain an overall comment. For example, in the process of analyzing Student B's attention within a week, it was found that Student B's attention was relatively focused from Monday to Wednesday, but his attention was extremely poor from Thursday to Friday. The corresponding attention comment may be "The student performed very well from Monday to Wednesday, but performed poorly on Thursday and Friday. Parents are advised to pay attention to the reasons for the change."

[0105] Here, the change in attention can be reflected by obtaining the target student's attention information during the second specified time period and generating an attention information change curve. In other words, the change in attention can be an attention change curve, through which the change in the target student's attention during the second specified time period can be clearly known. The second specified time period can be the same as or different from the first specified time period.

[0106] In summary, the disclosed embodiments can perform real-time face detection on students to determine their expressions, attention, etc., thereby generating a learning concentration test report, which enables parents to better understand their children's learning situation and provide advice and guidance.

[0107] After acquiring the image to be detected, the embodiment of the present disclosure inputs the image to be detected into the target detection model to obtain facial information, wherein the target detection model includes a backbone network, a neck network and a head network. The backbone network can be used to perform feature extraction based on residual connections and bottleneck structures. On this basis, attention information corresponding to the facial information output by the target detection model is acquired, and a concentration report is generated based on the attention information. Since the acquisition of facial information is based on the output of the target detection model, the attention information acquired based on the facial information is more accurate, and the concentration report is automatically generated through the attention information, which can assist parents or teachers in knowing the students' learning situation more clearly.

[0108] Figure 8 A face detection device according to an exemplary embodiment is shown. Figure 8 The face detection device 800 shown may include an acquisition module 810 , an input module 820 and a generation module 830 .

[0109] The acquisition module 810 is configured to acquire an image to be detected; The input module 820 is configured to input the image to be detected into the target detection model to obtain face information. The target detection model includes a backbone network, a neck network and a head network. The backbone network is used to extract features based on residual connections and bottleneck structures. The generation module 830 is configured to obtain attention information corresponding to the facial information and generate a concentration report based on the attention information.

[0110] In some embodiments, the input module 820 includes: An extraction submodule is configured to perform feature extraction on the image to be detected using the residual connection and the bottleneck structure of the Backbone network to obtain a multi-scale feature map; A fusion submodule is configured to perform feature fusion on the multi-scale feature map based on the Neck network to obtain a candidate feature map; The determination submodule is configured to determine the facial information corresponding to the candidate feature map according to the Head network.

[0111] In some embodiments, the Backbone includes five convolution modules, four cross-order partial feature fusion CSP2F modules and a spatial pyramid pooling SPPF module. The convolution module includes a convolution layer, a batch normalization layer and a ReLU activation function layer. The convolution module is used to perform sampling operations, the CSP2F module is used to decompose and process feature maps, and the SPPF module is used to aggregate spatial information at different scales.

[0112] Optionally, the CSP2F module is configured to perform the following operations: Based on a 1x1 convolution kernel, the number of channels of the feature map transmitted by the convolution module is reduced to half of the original number; Use multiple 3x3 convolution kernels to perform convolution operation on the processed feature map to extract feature information; The number of channels of the feature information is restored based on a 1x1 convolution kernel.

[0113] In some embodiments, the Neck network includes a feature pyramid network (FPN), and the step of performing feature fusion on the multi-scale feature map based on the Neck network to obtain a candidate feature map includes: The multi-scale feature map is upsampled and downsampled using the FPN to obtain the candidate feature map.

[0114] In some embodiments, the Head network adopts a decoupling head structure, which includes a regression head and a classification head; the determination submodule is also configured to: generate multiple candidate boxes corresponding to the candidate feature map; predict the position of each candidate box based on the regression head, and use the classification head to classify each candidate box to obtain the face information.

[0115] In some embodiments, the generation module 830 is configured to perform expression recognition on the facial information to obtain an expression recognition result; perform eye socket recognition on the facial information to obtain an eye socket recognition result; perform action recognition on the target person corresponding to the facial information to obtain an action recognition result; and comprehensively determine the attention information based on the expression recognition result, the eye socket recognition result, and the action recognition result.

[0116] After acquiring the image to be detected, the embodiment of the present disclosure inputs the image to be detected into the target detection model to obtain facial information, wherein the target detection model includes a backbone network, a neck network and a head network. The backbone network can be used to perform feature extraction based on residual connections and bottleneck structures. On this basis, attention information corresponding to the facial information output by the target detection model is acquired, and a concentration report is generated based on the attention information. Since the acquisition of facial information is based on the output of the target detection model, the attention information acquired based on the facial information is more accurate, and the concentration report is automatically generated through the attention information, which can assist parents or teachers in knowing the students' learning situation more clearly.

[0117] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0118] Figure 9 FIG. 1 is a block diagram of an electronic device 900 according to an exemplary embodiment. Figure 9 As shown, the electronic device 900 may include: a processor 901 , a memory 902 , and may further include one or more of a multimedia component 903 , an input / output (I / O) interface 904 , and a communication component 905 .

[0119] The processor 901 is used to control the overall operation of the electronic device 900 to complete all or part of the steps in the above-mentioned face detection method. The memory 902 is used to store various types of data to support the operation of the electronic device 900. This data may include, for example, instructions for any application or method operating on the electronic device 900, as well as application-related data such as contact information, sent and received messages, images, audio, video, etc. The memory 902 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 903 may include a screen and an audio component. The screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signal may be further stored in the memory 902 or sent through the communication component 905. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 904 provides an interface between the processor 901 and other interface modules. The above-mentioned other interface modules may be a keyboard, a mouse, buttons, etc. These buttons may be virtual buttons or physical buttons. The communication component 905 is used for wired or wireless communication between the electronic device 900 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G or 4G, or a combination of one or more of them, so the corresponding communication component 905 may include: a Wi-Fi module, a Bluetooth module, an NFC module.

[0120] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned face detection method.

[0121] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided. When executed by a processor, the program instructions implement the steps of the above-described face detection method. For example, the computer-readable storage medium may be the aforementioned memory 902 including the program instructions. The program instructions may be executed by the processor 901 of the electronic device 900 to perform the above-described face detection method.

[0122] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of the above-mentioned face detection method are implemented.

[0123] In another exemplary embodiment, a computer-readable storage medium including program instructions is further provided. When the program instructions are executed by a processor, the steps of the above-mentioned face detection method are implemented.

[0124] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a processor. When the computer program is executed by the processor, the steps of the above-mentioned face detection method are implemented.

[0125] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the scope of protection of the present disclosure.

[0126] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any appropriate manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.

[0127] In addition, the various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.

Claims

1. A face detection method, characterized in that: The method comprises: Obtain the image to be detected; Inputting the image to be detected into a target detection model to obtain face information, wherein the target detection model includes a backbone network, a neck network, and a head network, wherein the backbone network is used for feature extraction based on residual connection and bottleneck structure; Acquire attention information corresponding to the facial information, and generate a concentration report based on the attention information.

2. The method according to claim 1, characterized in that The step of inputting the image to be detected into a target detection model to obtain face information includes: Using the residual connection and the bottleneck structure of the Backbone network to extract features of the image to be detected, to obtain a multi-scale feature map; Performing feature fusion on the multi-scale feature map based on the Neck network to obtain a candidate feature map; Determine the facial information corresponding to the candidate feature map according to the Head network.

3. The method according to claim 2, characterized in that The Backbone includes five convolution modules, four cross-order partial feature fusion CSP2F modules and a spatial pyramid pooling SPPF module. The convolution module includes a convolution layer, a batch normalization layer and a ReLU activation function layer. The convolution module is used to perform sampling operations, the CSP2F module is used to decompose and process feature maps, and the SPPF module is used to aggregate spatial information at different scales.

4. The method according to claim 3, characterized in that The CSP2F module is used to perform the following operations: Based on a 1x1 convolution kernel, the number of channels of the feature map transmitted by the convolution module is reduced to half of the original number; Use multiple 3x3 convolution kernels to perform convolution operation on the processed feature map to extract feature information; The number of channels of the feature information is restored based on a 1x1 convolution kernel.

5. The method according to claim 2, characterized in that The Neck network includes a feature pyramid network FPN, and the multi-scale feature map is subjected to feature fusion based on the Neck network to obtain a candidate feature map, including: The multi-scale feature map is upsampled and downsampled using the FPN to obtain the candidate feature map.

6. The method according to claim 2, characterized in that The Head network adopts a decoupling head structure, which includes a regression head and a classification head; The determining, according to the Head network, the facial information corresponding to the candidate feature map includes: Generating multiple candidate boxes corresponding to the candidate feature maps; The position of each candidate frame is predicted based on the regression head, and each candidate frame is classified using the classification head to obtain the face information.

7. The method according to claim 1, characterized in that The obtaining of attention information corresponding to the face information includes: Performing expression recognition on the facial information to obtain an expression recognition result; Performing eye socket recognition on the facial information to obtain an eye socket recognition result; Performing action recognition on a target person corresponding to the facial information to obtain an action recognition result; The attention information is comprehensively determined based on the expression recognition result, the eye socket recognition result and the action recognition result.

8. A face detection device, characterized in that: The device comprises: An acquisition module is configured to acquire an image to be detected; An input module is configured to input the image to be detected into a target detection model to obtain face information. The target detection model includes a backbone network, a neck network, and a head network. The backbone network is used for feature extraction based on residual connections and a bottleneck structure. The generation module is configured to obtain attention information corresponding to the facial information and generate a concentration report based on the attention information.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.