Method and device for confirming attendance results

By constructing an image recognition model to identify remake images in face recognition attendance, the problem of high cost of testing the authenticity of face recognition attendance in the prior art is solved, and an efficient and economical method of confirming attendance results is achieved.

CN114332991BActive Publication Date: 2025-05-16SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111511475.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-06
Publication Date
2025-05-16
Estimated Expiration
2041-12-06

AI Technical Summary

Technical Problem

The prior art often leads to high cost problems when testing the authenticity of facial recognition attendance.

Method used

By constructing the first image recognition model and the second image recognition model, the attendance image of the target object within the preset time period is obtained, and the attendance result is determined based on the type of remake image present in the image (one or multiple uses).

Benefits of technology

This method effectively reduces the cost of testing the authenticity of face recognition attendance and provides a new method of face recognition attendance, which can be run without real-time location assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332991B_ABST
    Figure CN114332991B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technology, and provides a method and device for confirming attendance results. The method includes: inputting the attendance image into the first image recognition model, and outputting the first attendance information of the target object; inputting the attendance image into the second image recognition model, and outputting the second attendance information of the target object; according to the first attendance information and / or the second attendance information, determining the attendance result of the target object within the preset time period, and using the above technical means to solve the problem of high cost in the prior art for verifying the authenticity of face recognition attendance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method and device for confirming attendance results. Background Art

[0002] In order to ensure work efficiency, almost every unit will use attendance punching to ensure employees' attendance and working hours. Among the many attendance punching methods, facial recognition punching is the most widely used one. It is currently found that some employees will use photocopying images to cheat attendance. For example, employee A is late for work, but employee A sends his photocopying image to employee B, and employee B uses employee A's photocopying image to help employee A punch in with facial recognition. Photocopying is to make copies of originals such as photos, negatives, drawings, documents and charts. In this disclosure, photocopying images refers to using images instead of real people when employees punch in with facial recognition.

[0003] In order to solve the problem of cheating attendance by using re-photographed images in face recognition clocking in, the existing technology either detects the human body through sensors while recognizing the face, or obtains the real-time location of the clocking-in object while recognizing the face. However, whether using sensors to detect the human body or obtaining the real-time location of the clocking-in object to assist face recognition, corresponding hardware is used, which increases the cost of the attendance equipment.

[0004] In the process of realizing the concept of the present disclosure, the inventors found that there are at least the following technical problems in the related technology: in order to verify the authenticity of face recognition attendance, it often causes high costs. Summary of the invention

[0005] In view of this, the embodiments of the present disclosure provide a method, device, electronic device and computer-readable storage medium for confirming attendance results to solve the problem of high costs in the prior art for verifying the authenticity of facial recognition attendance.

[0006] According to a first aspect of an embodiment of the present disclosure, a method for confirming an attendance result is provided, comprising: constructing a first image recognition model and a second image recognition model; obtaining an attendance image of a target object within a preset time period; inputting the attendance image into the first image recognition model, and outputting first attendance information of the target object, wherein the first attendance information includes a first reprinted image in the attendance image, and the first reprinted image is a reprinted image used only once by the target object within the preset time period; inputting the attendance image into the second image recognition model, and outputting second attendance information of the target object, wherein the second attendance information includes a second reprinted image in the attendance image, and the second reprinted image is a reprinted image used multiple times by the target object within the preset time period; and determining the attendance result of the target object within the preset time period according to the first attendance information and / or the second attendance information.

[0007] According to a second aspect of an embodiment of the present disclosure, a device for confirming attendance results is provided, comprising: a model building module, configured to build a first image recognition model and a second image recognition model; an acquisition module, configured to obtain an attendance image of a target object within a preset time period; a first model module, configured to input the attendance image into the first image recognition model, and output first attendance information of the target object, wherein the first attendance information includes a first reprinted image in the attendance image, and the first reprinted image is a reprinted image used only once by the target object within the preset time period; a second model module, configured to input the attendance image into the second image recognition model, and output second attendance information of the target object, wherein the second attendance information includes a second reprinted image in the attendance image, and the second reprinted image is a reprinted image used multiple times by the target object within the preset time period; and a confirmation module, configured to determine the attendance result of the target object within the preset time period according to the first attendance information and / or the second attendance information.

[0008] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0009] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0010] Compared with the prior art, the disclosed embodiment has the following beneficial effects: because the disclosed embodiment inputs the attendance image into the first image recognition model to output the first attendance information of the target object; inputs the attendance image into the second image recognition model to output the second attendance information of the target object; and determines the attendance result of the target object within a preset time period based on the first attendance information and / or the second attendance information. Therefore, the above-mentioned technical means can solve the problem of high cost in the prior art for verifying the authenticity of face recognition attendance, thereby providing a new face recognition attendance method. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 is a schematic diagram of an application scenario of an embodiment of the present disclosure;

[0013] Figure 2 It is a flowchart of a method for confirming attendance results provided by an embodiment of the present disclosure;

[0014] Figure 3 It is a structural schematic diagram of a device for confirming attendance results provided by an embodiment of the present disclosure;

[0015] Figure 4 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0016] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.

[0017] A method and device for confirming attendance results according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0018] Figure 1 Schematic diagram of an application scenario of an embodiment of the present disclosure. The application scenario may include terminal devices 1, 2 and 3, a server 4 and a network 5.

[0019] Terminal devices 1, 2 and 3 can be hardware or software. When terminal devices 1, 2 and 3 are hardware, they can be various electronic devices with display screens and supporting communication with server 4, including but not limited to smart phones, tablet computers, laptop portable computers and desktop computers, etc.; when terminal devices 1, 2 and 3 are software, they can be installed in the above electronic devices. Terminal devices 1, 2 and 3 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present disclosure are not limited to this. Furthermore, various applications can be installed on terminal devices 1, 2 and 3, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0020] The server 4 may be a server that provides various services, for example, a background server that receives a request sent by a terminal device that establishes a communication connection with the server, and the background server may receive and analyze the request sent by the terminal device, and generate a processing result. The server 4 may be a single server, or a server cluster composed of several servers, or a cloud computing service center, which is not limited in the embodiments of the present disclosure.

[0021] It should be noted that the server 4 can be hardware or software. When the server 4 is hardware, it can be various electronic devices that provide various services for the terminal devices 1, 2, and 3. When the server 4 is software, it can be multiple software or software modules that provide various services for the terminal devices 1, 2, and 3, or it can be a single software or software module that provides various services for the terminal devices 1, 2, and 3, and the embodiments of the present disclosure are not limited to this.

[0022] The network 5 can be a wired network connected by coaxial cable, twisted pair and optical fiber, or it can be a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), infrared, etc., which is not limited in the embodiments of the present disclosure.

[0023] The user can establish a communication connection with the server 4 via the network 5 through the terminal devices 1, 2 and 3 to receive or send information, etc. It should be noted that the specific types, quantities and combinations of the terminal devices 1, 2 and 3, the server 4 and the network 5 can be adjusted according to the actual needs of the application scenario, and the embodiments of the present disclosure are not limited to this.

[0024] Figure 2 It is a flowchart of a method for confirming attendance results provided by an embodiment of the present disclosure. Figure 2 The attendance result can be confirmed by Figure 1 The terminal device or server executes. Figure 2 As shown, the confirmation method of the attendance result includes:

[0025] S201, constructing a first image recognition model and a second image recognition model;

[0026] S202, obtaining the attendance image of the target object within a preset time period;

[0027] S203, inputting the attendance image into a first image recognition model, and outputting first attendance information of the target object, wherein the first attendance information includes a first retaken image in the attendance image, and the first retaken image is a retaken image that is used only once by the target object within a preset time period;

[0028] S204, inputting the attendance image into a second image recognition model, and outputting second attendance information of the target object, wherein the second attendance information includes a second re-shot image in the attendance image, and the second re-shot image is a re-shot image used multiple times by the target object within a preset time period;

[0029] S205: Determine the attendance result of the target object within a preset time period according to the first attendance information and / or the second attendance information.

[0030] A residual module, a channel attention module, a spatial attention module and a prediction module are provided to construct a first image recognition model, and a second image recognition model is constructed by using a RetinaFace module and an Arcface module, as well as a regional clustering algorithm. Because the use of a single-shot image and a multiple-shot image in attendance will play different roles in determining the attendance result of the target object within a preset time period, the embodiment of the present disclosure processes the first shot image through the first image recognition model and processes the second shot image through the second image recognition model. Both the first image recognition model and the second image recognition model have been trained to learn and save the corresponding relationship between the attendance image and the attendance information.

[0031] It should be noted that before training the image recognition model, the attendance images need to be labeled. The disclosed embodiment proposes a method for intelligent labeling using a weak model. Specifically: first use a small batch of labeled data (such as 200 normal punch-in pictures and 200 fake pictures) to train a weak model, which can be a binary classification model ResNet18; set the classification threshold of the weak model. For example, in order to recall more suspected fake pictures, during forward reasoning, the classification threshold is set to 0.4, that is, if the probability of being a fake picture is greater than 0.4, the picture is considered to be a fake picture, and the picture with a probability of being a fake picture is less than 0.4 is marked as a non-fake picture, that is, the target object really punches in and does not cheat.

[0032] The first attendance information and the second attendance information further include information that a non-reproduced image exists in the attendance image.

[0033] According to the technical solution provided by the embodiment of the present disclosure, because the embodiment of the present disclosure inputs the attendance image into the first image recognition model to output the first attendance information of the target object; inputs the attendance image into the second image recognition model to output the second attendance information of the target object; and determines the attendance result of the target object within a preset time period according to the first attendance information and / or the second attendance information, therefore, the above-mentioned technical means can solve the problem in the prior art that high costs are often incurred in order to verify the authenticity of face recognition attendance, thereby providing a new face recognition attendance method.

[0034] Because the attendance method provided by the embodiment of the present disclosure can be run offline, it can also solve the problem of obtaining the real-time location of the clocking-in object for assisting face recognition, which requires a network connection and cannot be run in a wireless state.

[0035] In step S203, the attendance image is input into a first image recognition model, and the first attendance information of the target object is output, including: a first image recognition model, including: a residual module, a channel attention module, a spatial attention module and a prediction module; a first feature of the attendance image is extracted through the residual module, wherein the residual module is composed of a basic residual network and a residual branch; the hole rate of the channel attention module is determined through a parameter center, the first feature is input into the channel attention module, and a second feature is obtained, wherein the channel attention module includes a plurality of channel branches; the number of groupings of the spatial attention module is determined through a parameter center, the second feature is input into the spatial attention module, and a third feature is obtained, wherein the spatial attention module includes a plurality of spatial branches; the weight of the prediction module is determined through a parameter center, the third feature is input into the prediction module, and a fourth feature is obtained; and the first attendance information is determined according to the fourth feature.

[0036] The residual module is composed of a basic residual network and a residual branch. The basic residual network is ResNet50, and the residual branch is the output of the basic residual network plus the input of the basic residual network. The traditional method is to only use the basic residual network, that is, to directly map the features of the previous layer to the next layer by stacking convolutional layers. The disclosed embodiment adds a residual connection branch on the basis of the basic residual network. This branch is an identity mapping, so that the residual module only needs to learn the difference between the input and the output. The input of the basic residual network is the attendance image. Because the difference mapping is easier to optimize than the target value mapping, the gradient vanishing problem is avoided, so that the model can be designed to be deeper and more expressive.

[0037] The hole rate of the channel attention module, the number of groups of the spatial attention module, and the weight of the prediction module are determined by searching the parameters saved in the parameter center and randomly confirming them. For example, the hole rate of the channel attention module is determined by searching multiple hole rates saved in the parameter center and randomly confirming a hole rate.

[0038] Determining the first attendance information according to the fourth feature is to compare the fourth feature with the prototype feature of the target object in the prototype library, and then determine whether the attendance image is a copied image or a non-copied image, and distinguish how many times the target object has been used within a preset time period, thereby obtaining the first attendance information.

[0039] In step S203, the hole rate of the channel attention module is determined by the parameter center, and the first feature is input into the channel attention module to obtain the second feature, wherein the channel attention module includes multiple channel branches, including: in each channel branch: performing a preset convolution process on the first feature to obtain a first convolution result, using a normalized exponential function to process the convolution result to obtain a normalized result, multiplying the normalized result by the first feature to obtain a multiplication result, performing pooling process on the multiplication result to obtain a first pooling result, and passing the pooling result through a fully connected layer to obtain a first classification result; adding the classification results obtained by multiple channel branches to obtain a second feature.

[0040] Determining the hole rate of the channel attention module is to determine the hole rates of multiple channel branches in the channel attention module, and each channel branch will have a hole rate. Multiple channel branches are to give the channel attention module a larger receptive field. The preset convolution processing is to convolve the first feature with the preset convolution kernel size and the above hole rate. The normalized exponential function is softmax. The above processing of the first feature to obtain the second feature is based on different receptive fields to obtain the weights of different points in space, and then perform weighted averaging, which is different from the general practice of directly pooling the original feature map. Direct pooling is equivalent to the method of finding the channel average value, making the second feature more representative.

[0041] In step S203, the number of groups of the spatial attention module is determined by the parameter center, and the second feature is input into the spatial attention module to obtain the third feature, wherein the spatial attention module includes multiple spatial branches, including: dividing the second feature according to the number of groups to obtain a number of groups; in each spatial branch: in each group: performing a preset convolution process on the group to obtain a second convolution result, performing pooling process on the second convolution result to obtain a second pooling result, passing the second pooling result through a second fully connected layer to obtain a second classification result, using a relu activation function to process the second classification result to obtain a first activation result, passing the activation result through a third fully connected layer to obtain a third classification result, using a sigmoid activation function to process the third classification result to obtain a second activation result; performing weighted averaging process on the second activation results of all groups to obtain a weighted average result; adding the weighted average results obtained from all channel branches to obtain the third feature.

[0042] The number of groups of multiple spatial branches in the spatial attention module is the same, so you only need to confirm that they are in order. Multiple spatial branches are used to process the second feature in groups. The multiple spatial branches in the spatial attention module and the multiple channel branches in the channel attention module can be understood as channels, that is, a network path, but their functions are different. The second activation results of all groups are weighted averaged to obtain a weighted average result, where the weighted average processing is to first perform weighted summation and then average the summation results. The weighted weights are determined according to the specific situation.

[0043] In step S203, the weight of the prediction module is determined by the parameter center, and the third feature is input into the prediction module to obtain the fourth feature, including: the prediction module has four stages, including: the first stage, the second stage, the third stage and the fourth stage, and the third feature has characteristics of four stages, including: the first stage feature, the second stage feature, the third stage feature and the fourth stage feature; the second stage feature, the third stage feature and the fourth stage feature are respectively downsampled to the first stage, and the three sampling results are fused with the first stage feature to obtain the first fused feature; the first stage feature, the third stage feature and the fourth stage feature are respectively downsampled to the second stage, and the three sampling results are fused with the second stage feature to obtain the second fused feature; the first stage feature, the second stage feature and the fourth stage feature are respectively downsampled to the third stage, and the three sampling results are fused with the third stage feature to obtain the third fused feature; the first stage feature, the second stage feature and the third stage feature are respectively downsampled to the fourth stage, and the three sampling results are fused with the fourth stage feature to obtain the fourth fused feature; according to the weight, the first fused feature, the second fused feature, the third fused feature and the fourth fused feature are weighted and summed to obtain the fourth feature.

[0044] It is common knowledge in the art that a neural network model or module can be divided into four stages, which will not be elaborated here. In the fusion of the above-mentioned four features, three sampling results are involved. The three sampling results that appear in the specific fusion are the three sampling results in the specific fusion paragraph. The applicant believes that it is clear and will not be elaborated again. It should be noted that because features are often represented by matrices or vectors, the features that appear in the present disclosure can be understood as a matrix or vector. Fusion of the corresponding three sampling results with the corresponding stage features is actually a fusion of shallow features and deep features. For example, the first stage features, the second stage features, and the third stage features are shallow features, and the fourth stage features are deep features. The present disclosure obtains a more representative fourth feature through the fusion of shallow features and deep features.

[0045] In step S203, it includes: randomly determining a first number of first image recognition models according to a preset hole rate set, wherein each first image recognition model has a set of different hole rates; training each first image recognition model a preset number of times, and determining a second number of first optimal models from the trained multiple first image recognition models through testing; and storing the hole rates of the second number of first optimal models in the parameter center.

[0046] The first optimal model is the second number of first image recognition models with the best test results in the test. The second optimal model below is similar to the third optimal model.

[0047] In step S203, after the hole rates of the second number of optimal models are stored in the parameter center, the method also includes: randomly determining a first number of first image recognition models according to a preset grouping number set, wherein each first image recognition model has a different number of groupings, and the hole rate of each first image recognition model belongs to the hole rate stored in the parameter center; training each first image recognition model a preset number of times, and determining a second number of second optimal models from the trained multiple first image recognition models through testing; and storing the number of groups of the second number of second optimal models in the parameter center.

[0048] In step S203, after storing the number of groups of the second number of optimal models in the parameter center, the method also includes: randomly determining a first number of first image recognition models according to a preset weight set, wherein each first image recognition model has a different set of weights, the void rate of each first image recognition model belongs to the void rate stored in the parameter center, and the number of groups of each first image recognition model belongs to the number of groups stored in the parameter center; training each first image recognition model a preset number of times, and determining a second number of third optimal models from the multiple trained first image recognition models through testing; and storing the weights of the second number of third optimal models in the parameter center.

[0049] In step S204, the attendance image is input into the second image recognition model, and the second attendance information of the target object is output, including: inputting the attendance image into the RetinaFace module, and outputting the detection information of the target object, wherein the second image recognition model includes the RetinaFace model; extracting the fifth feature of the detection information through the Arcface module, and the second image recognition model includes the Arcface module; using the regional clustering algorithm to process the fifth feature to obtain a clustering result, and determining the second attendance information according to the clustering result.

[0050] Determining the second attendance information according to the clustering result is to compare the clustering result with the prototype features of the target object in the prototype library, and then determine whether the attendance image is a copied image or a non-copy image, and distinguish how many times the target object has been used within a preset time period, thereby obtaining the second attendance information.

[0051] In step S203, the fifth feature is processed using a regional clustering algorithm to obtain a clustering result, including: classifying multiple sample points of the fifth feature into core points, boundary points and noise points according to preset rules; wherein the preset rules are: the core points are within a preset distance and have more than a third number of sample points, the boundary points are within a preset distance and have core points and less than a third number of sample points, and the noise points are other sample points except the core points and the boundary points; connecting the core points whose distance is less than the preset distance to obtain a connected cluster; connecting each boundary point to the connected cluster with the closest distance to obtain a clustering result.

[0052] The regional clustering algorithm is a new clustering algorithm proposed by the embodiment of the present disclosure to determine the clustering results, and the above calculation process is the regional clustering algorithm. The core point has more than a third number of sample points within the preset distance, which can be understood as that the number of sample points adjacent to the core point within the range of the preset distance as the radius with the core point as the center is greater than the third number.

[0053] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.

[0054] The following are embodiments of the device disclosed herein, which can be used to execute the method embodiments disclosed herein. For details not disclosed in the device embodiments disclosed herein, please refer to the method embodiments disclosed herein.

[0055] Figure 3 Schematic diagram of a device for confirming attendance results provided by an embodiment of the present disclosure. Figure 3 As shown, the attendance result confirmation device includes:

[0056] A model building module 301 is configured to build a first image recognition model and a second image recognition model;

[0057] The acquisition module 302 is configured to acquire the attendance image of the target object within a preset time period;

[0058] The first model module 303 is configured to input the attendance image into a first image recognition model, and output first attendance information of the target object, wherein the first attendance information includes a first re-shot image in the attendance image, and the first re-shot image is a re-shot image that is only used once by the target object within a preset time period;

[0059] The second model module 304 is configured to input the attendance image into a second image recognition model and output second attendance information of the target object, wherein the second attendance information includes a second re-shot image in the attendance image, and the second re-shot image is a re-shot image used multiple times by the target object within a preset time period;

[0060] The confirmation module 305 is configured to determine the attendance result of the target object within a preset time period according to the first attendance information and / or the second attendance information.

[0061] A residual module, a channel attention module, a spatial attention module and a prediction module are provided to construct a first image recognition model, and a second image recognition model is constructed by using a RetinaFace module and an Arcface module, as well as a regional clustering algorithm. Because the use of a single-shot image and a multiple-shot image in attendance will play different roles in determining the attendance result of the target object within a preset time period, the embodiment of the present disclosure processes the first shot image through the first image recognition model and processes the second shot image through the second image recognition model. Both the first image recognition model and the second image recognition model have been trained to learn and save the corresponding relationship between the attendance image and the attendance information.

[0062] It should be noted that before training the image recognition model, the attendance images need to be labeled. The disclosed embodiment proposes a method for intelligent labeling using a weak model. Specifically: first use a small batch of labeled data (such as 200 normal punch-in pictures and 200 fake pictures) to train a weak model, which can be a binary classification model ResNet18; set the classification threshold of the weak model. For example, in order to recall more suspected fake pictures, during forward reasoning, the classification threshold is set to 0.4, that is, if the probability of being a fake picture is greater than 0.4, the picture is considered to be a fake picture, and the picture with a probability of being a fake picture is less than 0.4 is marked as a non-fake picture, that is, the target object really punches in and does not cheat.

[0063] The first attendance information and the second attendance information further include information that a non-reproduced image exists in the attendance image.

[0064] According to the technical solution provided by the embodiment of the present disclosure, because the embodiment of the present disclosure inputs the attendance image into the first image recognition model to output the first attendance information of the target object; inputs the attendance image into the second image recognition model to output the second attendance information of the target object; and determines the attendance result of the target object within a preset time period according to the first attendance information and / or the second attendance information, therefore, the above-mentioned technical means can solve the problem in the prior art that high costs are often incurred in order to verify the authenticity of face recognition attendance, thereby providing a new face recognition attendance method.

[0065] Optionally, the first model module 303 is also configured as a first image recognition model, including: a residual module, a channel attention module, a spatial attention module and a prediction module; extracting the first feature of the attendance image through the residual module, wherein the residual module is composed of a basic residual network and a residual branch; determining the void rate of the channel attention module through the parameter center, inputting the first feature into the channel attention module, and obtaining the second feature, wherein the channel attention module includes multiple channel branches; determining the number of groupings of the spatial attention module through the parameter center, inputting the second feature into the spatial attention module, and obtaining the third feature, wherein the spatial attention module includes multiple spatial branches; determining the weight of the prediction module through the parameter center, inputting the third feature into the prediction module, and obtaining the fourth feature; determining the first attendance information according to the fourth feature.

[0066] The residual module is composed of a basic residual network and a residual branch. The basic residual network is ResNet50, and the residual branch is the output of the basic residual network plus the input of the basic residual network. The traditional method is to only use the basic residual network, that is, to directly map the features of the previous layer to the next layer by stacking convolutional layers. The disclosed embodiment adds a residual connection branch on the basis of the basic residual network. This branch is an identity mapping, so that the residual module only needs to learn the difference between the input and the output. The input of the basic residual network is the attendance image. Because the difference mapping is easier to optimize than the target value mapping, the gradient vanishing problem is avoided, so that the model can be designed to be deeper and more expressive.

[0067] The hole rate of the channel attention module, the number of groups of the spatial attention module, and the weight of the prediction module are determined by searching the parameters saved in the parameter center and randomly confirming them. For example, the hole rate of the channel attention module is determined by searching multiple hole rates saved in the parameter center and randomly confirming a hole rate.

[0068] Determining the first attendance information according to the fourth feature is to compare the fourth feature with the prototype feature of the target object in the prototype library, and then determine whether the attendance image is a copied image or a non-copied image, and distinguish how many times the target object has been used within a preset time period, thereby obtaining the first attendance information.

[0069] Optionally, the first model module 303 is also configured to: in each channel branch: perform a preset convolution process on the first feature to obtain a first convolution result, use a normalized exponential function to process the convolution result to obtain a normalized result, multiply the normalized result by the first feature to obtain a multiplication result, perform pooling process on the multiplication result to obtain a first pooling result, pass the pooling result through a fully connected layer to obtain a first classification result; add the classification results obtained from multiple channel branches to obtain a second feature.

[0070] Determining the hole rate of the channel attention module is to determine the hole rates of multiple channel branches in the channel attention module, and each channel branch will have a hole rate. Multiple channel branches are to give the channel attention module a larger receptive field. The preset convolution processing is to convolve the first feature with the preset convolution kernel size and the above hole rate. The normalized exponential function is softmax. The above processing of the first feature to obtain the second feature is based on different receptive fields to obtain the weights of different points in space, and then perform weighted averaging, which is different from the general practice of directly pooling the original feature map. Direct pooling is equivalent to the method of finding the channel average value, making the second feature more representative.

[0071] Optionally, the first model module 303 is also configured to divide the second feature according to the number of groups to obtain a number of groups; in each spatial branch: in each group: perform a preset convolution process on the group to obtain a second convolution result, perform pooling process on the second convolution result to obtain a second pooling result, pass the second pooling result through a second fully connected layer to obtain a second classification result, use a relu activation function to process the second classification result to obtain a first activation result, pass the activation result through a third fully connected layer to obtain a third classification result, use a sigmoid activation function to process the third classification result to obtain a second activation result; perform weighted averaging process on the second activation results of all groups to obtain a weighted average result; add the weighted average results obtained from all channel branches to obtain the third feature.

[0072] The number of groups of multiple spatial branches in the spatial attention module is the same, so you only need to confirm that they are in order. Multiple spatial branches are used to process the second feature in groups. The multiple spatial branches in the spatial attention module and the multiple channel branches in the channel attention module can be understood as channels, that is, a network path, but their functions are different. The second activation results of all groups are weighted averaged to obtain a weighted average result, where the weighted average processing is to first perform weighted summation and then average the summation results. The weighted weights are determined according to the specific situation.

[0073] Optionally, the first model module 303 is also configured as a prediction module having four stages, including: a first stage, a second stage, a third stage and a fourth stage, and the third feature has features of four stages, including: first stage features, second stage features, third stage features and fourth stage features; the second stage features, the third stage features and the fourth stage features are respectively downsampled to the first stage, and the three sampling results are fused with the first stage features to obtain a first fused feature; the first stage features, the third stage features and the fourth stage features are respectively downsampled to the second stage, and the three sampling results are fused with the second stage features to obtain a second fused feature; the first stage features, the second stage features and the fourth stage features are respectively downsampled to the third stage, and the three sampling results are fused with the third stage features to obtain a third fused feature; the first stage features, the second stage features and the third stage features are respectively downsampled to the fourth stage, and the three sampling results are fused with the fourth stage features to obtain a fourth fused feature; according to the weights, the first fused feature, the second fused feature, the third fused feature and the fourth fused feature are weighted and summed to obtain the fourth feature.

[0074] A neural network model or module can be divided into four stages, which is common knowledge in the art and will not be elaborated here. In the fusion of the above four features, three sampling results are involved. The three sampling results that appear in the specific fusion are the three sampling results in the specific fusion paragraph. The applicant believes that it is clear and will not be elaborated again. It should be noted that because features are often represented by matrices or vectors, the features that appear in this disclosure can be understood as a matrix or vector.

[0075] Optionally, the first model module 303 is also configured to randomly determine a first number of first image recognition models based on a preset hole rate set, wherein each first image recognition model has a set of different hole rates; train each first image recognition model a preset number of times, and determine a second number of first optimal models from the trained multiple first image recognition models through testing; and store the hole rates of the second number of first optimal models in the parameter center.

[0076] The first optimal model is the second number of first image recognition models with the best test results in the test. The second optimal model below is similar to the third optimal model.

[0077] Optionally, the first model module 303 is also configured to randomly determine a first number of first image recognition models based on a preset set of grouping numbers, wherein each first image recognition model has a different number of groupings, and the hole rate of each first image recognition model belongs to the hole rate stored in the parameter center; train each first image recognition model a preset number of times, and determine a second number of second optimal models from the trained multiple first image recognition models through testing; and store the number of groups of the second number of second optimal models in the parameter center.

[0078] Optionally, the first model module 303 is also configured to randomly determine a first number of first image recognition models based on a preset weight set, wherein each first image recognition model has a different set of weights, the hole rate of each first image recognition model belongs to the hole rate stored in the parameter center, and the number of groups of each first image recognition model belongs to the number of groups stored in the parameter center; each first image recognition model is trained a preset number of times, and through testing, a second number of third optimal models are determined from the multiple trained first image recognition models; and the weights of the second number of third optimal models are stored in the parameter center.

[0079] Optionally, the second model module 304 is also configured to input the attendance image into the RetinaFace module and output the detection information of the target object, wherein the second image recognition model includes the RetinaFace model; extract the fifth feature of the detection information through the Arcface module, and the second image recognition model includes the Arcface module; use the regional clustering algorithm to process the fifth feature to obtain a clustering result, and determine the second attendance information based on the clustering result.

[0080] Determining the second attendance information according to the clustering result is to compare the clustering result with the prototype features of the target object in the prototype library, and then determine whether the attendance image is a copied image or a non-copy image, and distinguish how many times the target object has been used within a preset time period, thereby obtaining the second attendance information.

[0081] Optionally, the second model module 304 is also configured to classify multiple sample points of the fifth feature into core points, boundary points and noise points according to preset rules; wherein the preset rules are: the core points are within the preset distance and have more than a third number of sample points, the boundary points are within the preset distance and have core points and less than a third number of sample points, and the noise points are other sample points except the core points and the boundary points; the core points whose distance is less than the preset distance are connected to obtain a connected cluster; each boundary point is connected to the closest connected cluster to obtain a clustering result.

[0082] The regional clustering algorithm is a new clustering algorithm proposed in the embodiment of the present disclosure in order to determine the clustering results, and the above calculation process is the regional clustering algorithm.

[0083] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.

[0084] Figure 4 Schematic diagram of an electronic device 4 provided in an embodiment of the present disclosure. Figure 4 As shown, the electronic device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0085] Exemplarily, the computer program 403 may be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present disclosure. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 403 in the electronic device 4.

[0086] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 It is only an example of the electronic device 4 and does not constitute a limitation of the electronic device 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0087] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.

[0088] The memory 402 may be an internal storage unit of the electronic device 4, for example, a hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. Further, the memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device. The memory 402 may also be used to temporarily store data that has been output or is to be output.

[0089] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0090] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0091] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this disclosure.

[0092] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0093] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0094] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0095] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.

[0096] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A method for confirming attendance results, characterized in that: include: Constructing a first image recognition model and a second image recognition model; Obtain attendance images of target objects within a preset time period; Inputting the attendance image into the first image recognition model, and outputting first attendance information of the target object, wherein the first attendance information includes a first re-shot image in the attendance image, and the first re-shot image is a re-shot image that is only used once by the target object within the preset time period; Inputting the attendance image into the second image recognition model, and outputting second attendance information of the target object, wherein the second attendance information includes a second re-shot image in the attendance image, and the second re-shot image is a re-shot image used multiple times by the target object within the preset time period; Determining the attendance result of the target object within the preset time period according to the first attendance information and the second attendance information; Wherein, the inputting of the attendance image into the first image recognition model and the outputting of the first attendance information of the target object include: the first image recognition model includes: a residual module, a channel attention module, a spatial attention module and a prediction module; extracting the first feature of the attendance image through the residual module, wherein the residual module is composed of a basic residual network and a residual branch; determining the hole rate of the channel attention module through the parameter center, inputting the first feature into the channel attention module, and obtaining the second feature, wherein the channel attention module includes a plurality of channel branches; determining the number of groupings of the spatial attention module through the parameter center, inputting the second feature into the spatial attention module, and obtaining the third feature, wherein the spatial attention module includes a plurality of spatial branches; determining the weight of the prediction module through the parameter center, inputting the third feature into the prediction module, and obtaining the fourth feature; determining the first attendance information according to the fourth feature; The method comprises: randomly determining a first number of the first image recognition models according to a preset hole rate set, wherein each of the first image recognition models has a set of different hole rates; training each of the first image recognition models for a preset number of times, and determining a second number of first optimal models from the trained plurality of the first image recognition models through testing; and storing the hole rates of the second number of the first optimal models in the parameter center; Among them, after storing the number of groups of the second number of the optimal models in the parameter center, the method also includes: randomly determining the first number of the first image recognition models according to a preset weight set, wherein each of the first image recognition models has a different set of weights, the hole rate of each of the first image recognition models belongs to the hole rate stored in the parameter center, and the number of groups of each of the first image recognition models belongs to the number of groups stored in the parameter center; training each of the first image recognition models a preset number of times, and determining the second number of third optimal models from the multiple first image recognition models after training through testing; and storing the weights of the second number of the third optimal models in the parameter center.

2. The method according to claim 1, characterized in that The hole rate of the channel attention module is determined by the parameter center, and the first feature is input into the channel attention module to obtain the second feature, wherein the channel attention module includes multiple channel branches, including: In each of said channel branches: Performing a preset convolution process on the first feature to obtain a first convolution result, processing the convolution result using a normalized exponential function to obtain a normalized result, multiplying the normalized result by the first feature to obtain a multiplication result, performing a pooling process on the multiplication result to obtain a first pooling result, and passing the pooling result through a fully connected layer to obtain a first classification result; The classification results obtained by the multiple channel branches are added together to obtain the second feature.

3. The method according to claim 1, characterized in that The number of groups of the spatial attention module is determined by the parameter center, and the second feature is input into the spatial attention module to obtain a third feature, wherein the spatial attention module includes multiple spatial branches, including: Divide the second feature according to the number of groups to obtain the number of groups; In each of said spatial branches: In each of the groups: performing a preset convolution process on the group to obtain a second convolution result, performing a pooling process on the second convolution result to obtain a second pooling result, passing the second pooling result through a second fully connected layer to obtain a second classification result, using a relu activation function to process the second classification result to obtain a first activation result, passing the activation result through a third fully connected layer to obtain a third classification result, and using a sigmoid activation function to process the third classification result to obtain a second activation result; Performing weighted average processing on the second activation results of all the groups to obtain a weighted average result; The weighted average results obtained from all the channel branches are added together to obtain the third feature.

4. The method according to claim 1, characterized in that: The step of determining the weight of the prediction module by using the parameter center, inputting the third feature into the prediction module, and obtaining the fourth feature comprises: The prediction module has four stages, including: a first stage, a second stage, a third stage and a fourth stage, and the third feature has features of four stages, including: first stage features, second stage features, third stage features and fourth stage features; The second stage features, the third stage features, and the fourth stage features are respectively downsampled to the first stage, and the three sampling results are fused with the first stage features to obtain a first fused feature; The first stage features, the third stage features, and the fourth stage features are respectively downsampled to the second stage, and the three sampling results are fused with the second stage features to obtain second fused features; The first stage features, the second stage features, and the fourth stage features are respectively downsampled to the third stage, and the three sampling results are fused with the third stage features to obtain a third fused feature; Downsampling the first stage features, the second stage features, and the third stage features to the fourth stage respectively, and fusing the three sampling results with the fourth stage features to obtain fourth fused features; According to the weight, the first fusion feature, the second fusion feature, the third fusion feature and the fourth fusion feature are weighted and summed to obtain the fourth feature.

5. The method according to claim 1, characterized in that After storing the second number of hole rates of the optimal models in the parameter center, the method further includes: Randomly determine the first number of the first image recognition models according to a preset grouping number set, wherein each of the first image recognition models has a different grouping number, and the hole rate of each of the first image recognition models belongs to the hole rate stored in the parameter center; Training each of the first image recognition models for a preset number of times, and determining a second number of second optimal models from the trained plurality of the first image recognition models through testing; The number of groups of the second number of the second optimal models is stored in the parameter center.

6. The method according to claim 1, characterized in that The step of inputting the attendance image into the second image recognition model and outputting the second attendance information of the target object comprises: Input the attendance image into a RetinaFace module, and output detection information of the target object, wherein the second image recognition model includes the RetinaFace model; extracting a fifth feature of the detection information through an Arcface module, wherein the second image recognition model includes the Arcface module; The fifth feature is processed using a regional clustering algorithm to obtain a clustering result, and the second attendance information is determined according to the clustering result.

7. The method according to claim 6, characterized in that The using a regional clustering algorithm to process the fifth feature to obtain a clustering result includes: Classifying the plurality of sample points of the fifth feature into core points, boundary points and noise points according to a preset rule; The preset rule is: the core point has more than a third number of sample points within the preset distance, the boundary point has a core point and less than the third number of sample points within the preset distance, and the noise point is other sample points except the core point and the boundary point; Connecting the core points whose distance is less than the preset distance to obtain a connection cluster; Each of the boundary points is connected to the connection cluster with the closest distance to obtain the clustering result.

8. A device for confirming attendance results, characterized in that: include: A model building module is configured to build a first image recognition model and a second image recognition model; An acquisition module is configured to acquire the attendance image of the target object within a preset time period; A first model module is configured to input the attendance image into the first image recognition model and output first attendance information of the target object, wherein the first attendance information includes a first re-shot image in the attendance image, and the first re-shot image is a re-shot image that is only used once by the target object within the preset time period; A second model module is configured to input the attendance image into the second image recognition model, and output second attendance information of the target object, wherein the second attendance information includes a second re-shot image in the attendance image, and the second re-shot image is a re-shot image used multiple times by the target object within the preset time period; A confirmation module, configured to determine the attendance result of the target object within the preset time period according to the first attendance information and the second attendance information; The first model module is also configured as a residual module, a channel attention module, a spatial attention module and a prediction module; the first feature of the attendance image is extracted through the residual module, wherein the residual module is composed of a basic residual network and a residual branch; the hole rate of the channel attention module is determined through the parameter center, and the first feature is input into the channel attention module to obtain a second feature, wherein the channel attention module includes a plurality of channel branches; the number of groupings of the spatial attention module is determined through the parameter center, and the second feature is input into the spatial attention module to obtain a third feature, wherein the spatial attention module includes a plurality of spatial branches; the weight of the prediction module is determined through the parameter center, and the third feature is input into the prediction module to obtain a fourth feature; the first attendance information is determined according to the fourth feature; The first model module is further configured to randomly determine a first number of the first image recognition models according to a preset hole rate set, wherein each of the first image recognition models has a set of different hole rates; train each of the first image recognition models for a preset number of times, and determine a second number of first optimal models from the trained plurality of the first image recognition models through testing; and store the hole rates of the second number of the first optimal models in the parameter center; The first model module is also configured to randomly determine the first number of the first image recognition models according to a preset weight set, wherein each of the first image recognition models has a different set of weights, the hole rate of each of the first image recognition models belongs to the hole rate stored in the parameter center, and the number of groups of each of the first image recognition models belongs to the number of groups stored in the parameter center; train each of the first image recognition models a preset number of times, and determine a second number of third optimal models from the trained multiple first image recognition models through testing; and store the weights of the second number of the third optimal models in the parameter center.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Tablet personal computer-based facial speech attendance system

    CN105788018A

  • System and method for integrated learning

    US20160063873A1