Non-inductive attendance checking method and device, and storage medium

Through the collaborative work of panoramic cameras and pan-tilt cameras, high-definition facial images can be quickly identified and captured, solving the problem of misidentification caused by the insufficient clarity of panoramic cameras and improving the accuracy and efficiency of attendance.

CN120636010APending Publication Date: 2025-09-12GUANGZHOU KINDLINK INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510640259.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The panoramic camera has low facial imaging clarity for people at a distance, resulting in frequent misidentification or failure of identification of people in the back row, affecting the accuracy and efficiency of attendance.

Method used

Use the panoramic images captured by the panoramic camera to quickly identify people looking up in the activity area, control the pan-tilt camera to turn to the target area to capture close-up images, ensure high-definition face recognition, avoid invalid rotation and shooting of the pan-tilt camera, and reduce usage loss.

Benefits of technology

It improves the accuracy of identity recognition and attendance efficiency, ensures that people in the back row can also be accurately identified, and reduces the loss of PTZ cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636010A_ABST
    Figure CN120636010A_ABST
Patent Text Reader

Abstract

The invention discloses a non-inductive attendance checking method and device and a storage medium, and relates to the technical field of image processing. The method comprises the steps that a first panoramic image obtained by shooting an activity area through a panoramic camera is acquired, multiple first target objects in the first panoramic image are determined, and the first target objects are images of activity persons in the activity area in the first panoramic image; determining a head-up object in the plurality of first target objects, and determining a target area corresponding to the head-up object in the activity area; controlling the pan-tilt camera to turn to the target area, and obtaining a close-up image obtained by shooting the target area by the pan-tilt camera; and performing sign-in authentication on the moving personnel in the target area according to the close-up image. Through the technical means, the close-up image shot by the pan-tilt camera can recognize the high-definition face of the head-up object, the identity information of the head-up object is accurately recognized through the high-definition face, and the problem that in the prior art, mistaken recognition or recognition failure frequently occurs when an active person carries out identity recognition is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a contactless attendance method, device, and storage medium. Background Art

[0002] Attendance management is a crucial component of organizational management. It ensures the orderliness and efficiency of organizational activities by regulating personnel behavior and schedules. For example, in organizational scenarios such as businesses, classrooms, and conferences, attendance management is performed on participants (employees, students, and conference attendees) to ensure they attend on time and comply with regulations to maintain organizational discipline. With the rapid development of image processing technology, attendance methods have evolved from fingerprint clocking, facial recognition, and manual roll call to seamless attendance. This eliminates the need for participants to actively participate in the attendance process, improving attendance management efficiency and user experience.

[0003] In the prior art, a panoramic camera installed in the activity area captures a panoramic image of the activity area, performs face detection on the panoramic image to obtain the corresponding facial imaging, and performs identity recognition on the facial imaging to confirm the identity information of the corresponding active person, thereby confirming that the active person has successfully signed in. Since the panoramic camera can capture the facial image of each active person, the active person does not need to actively cooperate with the panoramic camera, thus achieving seamless attendance. However, due to the short focal length of the panoramic camera, the panoramic camera's ability to focus on people at a distance is insufficient, and the clarity of the facial imaging of the active people in the back row of the activity area in the panoramic image is low, resulting in frequent misidentification or recognition failure of the active people in the back row during identity recognition, affecting the accuracy and efficiency of attendance. Summary of the Invention

[0004] The present application provides a contactless attendance method, device, and storage medium, which utilize a first panoramic image captured by a panoramic camera to detect the imaging of active persons in an activity area, determine a target area within the activity area where persons with their heads raised are located based on the imaging of the active persons, control a pan-tilt camera to capture a close-up image of the target area to capture a high-definition face of the person with their heads raised, and accurately identify the identity information of the person with their heads raised through the high-definition face of the person with their heads raised, thereby solving the problem of frequent misidentification or recognition failure in the prior art when performing identity recognition of active persons, and improving the accuracy and efficiency of attendance.

[0005] In the first aspect, the present application provides a method for non-sensing attendance, including:

[0006] Acquire a first panoramic image obtained by photographing the activity area with a panoramic camera, and determine a plurality of first target objects in the first panoramic image, where the first target objects are images of people active in the activity area in the first panoramic image;

[0007] Determine a head-up object among the multiple first target objects, and determine a target area corresponding to the head-up object in the activity area;

[0008] Controlling the pan-tilt camera to turn toward the target area, and acquiring a close-up image of the target area captured by the pan-tilt camera;

[0009] The people active in the target area are checked in and authenticated based on the close-up image.

[0010] Through the above-mentioned technical means, all persons who raise their heads in the activity area can be quickly identified based on the first panoramic image taken by the panoramic camera, thereby accurately locating the target area of ​​the person who raises his head in the activity area, and controlling the pan-tilt camera to turn to the target area for shooting to collect close-up images containing the faces of the persons who raise their heads, thereby avoiding invalid rotation and invalid shooting of the pan-tilt camera, and reducing the use loss of the pan-tilt camera. Since the focal length of the pan-tilt camera is relatively long, it can capture close-up images containing high-definition faces when shooting the target area at any position in the activity area, ensuring the clarity of the facial image used for identity recognition, and improving the accuracy and success rate of identity recognition. Even if the person who raises his head is in the back row of the area, the identity information of the corresponding person who raises his head can be effectively identified using the high-definition face in the close-up image for identity recognition, which solves the problem of frequent misidentification or recognition failure when performing identity recognition for active persons in the back row in the prior art, and improves the accuracy and efficiency of attendance.

[0011] Optionally, the first target object in the first panoramic image is obtained by tracking a second target object in a historical panoramic image, the second target object and the correspondingly tracked first target object correspond to the same active person, and the historical panoramic image is obtained by photographing the activity area by the panoramic camera before photographing the first panoramic image;

[0012] Accordingly, determining the headed object from the plurality of first target objects includes:

[0013] Acquire posture information of a first target object and a second target object of the active person;

[0014] Determining the corresponding action of the active person according to the posture information of the first target object and the second target object of the active person;

[0015] In a case where the action of the activity person is raising his head, the first target object of the activity person is determined to be a raising-head object.

[0016] Through the above-mentioned technical means, multiple images of the active person in the first panoramic image and the historical panoramic image are used to predict the behavior of the active person, and when the active person is predicted to make a head-raising action, it is confirmed that he or she is a head-raising person, so that the corresponding first target object is determined as the head-raising object, to ensure that the gimbal camera is controlled to capture the high-definition face of the corresponding active person before the active person finishes the head-raising action, to avoid invalid rotation and invalid shooting of the gimbal camera, and to reduce the use loss of the gimbal camera.

[0017] Optionally, the activity area includes multiple sub-areas;

[0018] Accordingly, determining the target area corresponding to the head-up object in the activity area includes:

[0019] Obtaining location information of each head-up object;

[0020] determining a sub-region where the head-up object is located according to the position information of the head-up object;

[0021] According to the sub-areas where the head-up objects are located, counting the number of head-up objects in the sub-areas;

[0022] A target area is determined in the plurality of sub-areas according to the number of head-up objects in the plurality of sub-areas.

[0023] Through the above technical means, the number of head-up objects in each sub-area of ​​the activity area can be determined according to the position information of each head-up object. Based on the number of head-up objects in each sub-area, the sub-area with a higher attendance success rate is preferentially selected as the target area, which is conducive to improving attendance efficiency and reducing the number and frequency of turns of the gimbal camera.

[0024] Optionally, counting the number of head-up objects in a plurality of sub-areas according to the sub-areas where the head-up objects are located includes:

[0025] Count the number of people who have successfully signed in in each sub-area;

[0026] When the number of people in attendance in the sub-area does not reach the number threshold of the sub-area, the number of head-up objects in the sub-area is counted according to the head-up objects in the sub-area.

[0027] Through the above technical means, by comparing the number of attendance in the sub-area with the number threshold, it is determined whether the sub-area has completed attendance. If the attendance is not completed, the number of objects with their heads raised in the sub-area is counted to avoid the pan-tilt camera turning to the sub-area where attendance has been completed to take close-up images, reduce unnecessary rotation and shooting of the pan-tilt camera, reduce the energy consumption and wear of the pan-tilt camera, and avoid repeated attendance in the sub-area where attendance has been completed, thereby improving attendance efficiency.

[0028] Optionally, counting the number of head-up objects in a plurality of sub-areas according to the sub-areas where the head-up objects are located includes:

[0029] determining a head-up rate of the sub-area according to the number of head-up objects and the number of target objects in the sub-area;

[0030] A target area is determined in the plurality of sub-areas according to the head-up rates of the plurality of sub-areas.

[0031] Through the above technical means, the sub-areas with a higher probability of completing attendance status are preferentially selected as target areas based on the head-up rate of the sub-areas, so as to minimize the number of sub-areas that the gimbal camera can turn to, thereby reducing the number and frequency of the gimbal camera's rotation.

[0032] Optionally, determining the target area from the multiple sub-areas according to the head-up rates of the multiple sub-areas includes:

[0033] In the case that the sub-area with the highest head-up rate is the same sub-area as the sub-area continuously photographed by the gimbal camera, if the continuous photographing time of the sub-area with the highest head-up rate reaches a preset time threshold, the sub-area with the second highest head-up rate is determined as the target area; otherwise, the sub-area with the highest head-up rate is determined as the target area.

[0034] Through the above technical means, when the PTZ camera is detected to stay in a certain sub-area for a long time, it can be promptly turned to other sub-areas to capture effective close-up images, thereby improving the attendance rate of other sub-areas and thus improving the overall attendance efficiency.

[0035] Optionally, obtaining a close-up image of the target area captured by the pan-tilt camera includes:

[0036] Acquire multiple frames of close-up images of the target area captured by the pan-tilt camera at a preset frequency within a first preset time period;

[0037] Accordingly, the step of performing sign-in authentication on the active persons in the target area according to the close-up image includes:

[0038] Detecting multiple facial images in the multiple close-up images, and dividing the multiple facial images into multiple facial image sets, each facial image set corresponding to one active person;

[0039] The identity information of the corresponding activity personnel is identified based on the face image set, and the successful sign-in of the corresponding activity personnel is confirmed based on the identified identity information.

[0040] Through the above technical means, after successfully identifying the identity information of the corresponding activity personnel based on the facial image of the activity personnel, it is confirmed that the activity personnel has successfully signed in and the facial image of the activity personnel will no longer be recognized again, avoiding repeated identity recognition, saving computing resources and improving attendance efficiency.

[0041] Optionally, detecting a plurality of facial images in the plurality of close-up images and dividing the plurality of facial images into a plurality of facial image sets includes:

[0042] Performing face detection on a first frame of close-up image captured within the first preset time period to obtain at least one face image;

[0043] Face tracking is performed on the remaining close-up images within the first preset time period based on the face image, and the face image in the first frame close-up image and the corresponding tracked face images in the remaining close-up images are divided into the same face image set.

[0044] Through the above technical means, the face image detected in the first frame close-up image is used to track the remaining close-up images, so that multiple face images of active persons in multiple frames of close-up images can be quickly determined, thereby improving the detection efficiency of multiple face images of active persons.

[0045] Optionally, identifying identity information of corresponding activity personnel based on multiple facial images belonging to the same activity personnel includes:

[0046] evaluating the image quality of each face image in the face image set, and selecting the face image with the highest image quality;

[0047] The identity information of the corresponding activity person is identified according to the facial image with the highest image quality, and the successful sign-in of the corresponding activity person is confirmed based on the identified identity information.

[0048] Through the above technical means, by selecting the facial image with the highest image quality from multiple facial images corresponding to active personnel for identity recognition, identity recognition of low-quality facial images can be avoided, effectively reducing the number of times identity recognition of facial images is performed, improving identity recognition efficiency, and thus improving attendance efficiency.

[0049] Optionally, controlling the pan-tilt camera to capture multiple close-up images of the target area within a first preset time period based on a preset frequency includes:

[0050] If the activity carried out in the activity area is in a critical time period, controlling the pan-tilt camera to capture a close-up image of the target area based on a first preset frequency;

[0051] If the activity carried out in the activity area is not in the critical time period, controlling the pan-tilt camera to capture a close-up image of the target area based on a second preset frequency;

[0052] The first preset frequency is higher than the second preset frequency.

[0053] Through the above technical means, by setting key time periods, the pan-tilt camera can capture close-up images of the target area during the key time periods at a higher frequency, thereby improving the success rate of the pan-tilt camera in capturing high-definition faces, thereby improving the success rate of identity recognition and attendance efficiency.

[0054] Optionally, the method further includes:

[0055] Acquire a second panoramic image obtained by photographing the activity area with a panoramic camera, where the resolution of the second panoramic image is higher than that of the first panoramic image;

[0056] performing face detection on the second panoramic image to obtain at least one face image;

[0057] When the image quality of the facial image is greater than or equal to a preset quality threshold, the identity information of the corresponding activity person is identified according to the facial image, and the corresponding activity person is confirmed to have successfully signed in based on the identified identity information.

[0058] Through these technical means, the second panoramic image captured by the panoramic camera is used in conjunction with the close-up image captured by the pan-tilt camera for identity recognition and check-in authentication, improving attendance efficiency. Furthermore, the second panoramic image is implemented in software, without adding additional hardware costs to the panoramic camera. The introduction of the second panoramic image allows each sub-area to complete attendance more quickly, which helps reduce the number of pan-tilt camera rotations and thus reduces the wear and tear of the pan-tilt camera.

[0059] In a second aspect, the present application provides a non-sensing attendance device, comprising:

[0060] One or more processors; a memory storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the contactless attendance method as described in the first aspect.

[0061] In a third aspect, the present application provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the contactless attendance method as described in the first aspect.

[0062] In the present application, a panoramic camera is used to capture a first panoramic image of an activity area, and the images of each active person in the activity area are detected in the first panoramic image. Based on the images of each active person, the head-up object is determined, thereby determining the target area where the head-up object is located in the activity area. After the pan-tilt camera is controlled to turn to the target area, a close-up image of the target area is captured by the pan-tilt camera. The identity information of the corresponding active person is recognized based on the high-definition face in the close-up image and sign-in authentication is performed. Through the above technical means, all the people with their heads raised in the activity area can be quickly identified based on the first panoramic image captured by the panoramic camera, thereby accurately locating the target area of ​​the person with their heads raised in the activity area. The pan-tilt camera is controlled to turn to the target area to shoot to capture close-up images containing the faces of the people with their heads raised, thereby avoiding invalid rotation and invalid shooting of the pan-tilt camera and improving the utilization rate of the pan-tilt camera. Since the pan-tilt camera has a long focal length, it can capture close-up images containing high-definition faces when shooting the target area at any position in the activity area, ensuring the clarity of the facial image used for identity recognition and improving the accuracy and success rate of identity recognition. Even if the person looking up is in the back row of the area, the identity information of the corresponding person looking up can be effectively identified using the high-definition face in the close-up image for identity recognition, which solves the problem of frequent misidentification or recognition failure when performing identity recognition for people in the back row in the existing technology, and improves the accuracy and efficiency of attendance. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a flow chart of a non-sensing attendance method provided by an embodiment of the present application;

[0064] Figure 2 1 is a schematic top view of a classroom provided in an embodiment of the present application;

[0065] Figure 3 This is a flowchart of determining a header object provided by an embodiment of the present application;

[0066] Figure 4 This is a flowchart of determining a target area among multiple sub-areas provided by an embodiment of the present application;

[0067] Figure 5 Schematic diagram of the attendance status of each sub-area provided in an embodiment of the present application;

[0068] Figure 6 This is a flowchart of the check-in authentication of active personnel based on multiple close-up images provided in an embodiment of the present application;

[0069] Figure 7 This is a flowchart of face detection based on high-resolution panoramic images provided by an embodiment of the present application;

[0070] Figure 8 is a schematic diagram of the image processing process provided by an embodiment of the present application;

[0071] Figure 9 This is a structural diagram of a contactless attendance device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] In order to make the purpose, technical solutions and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. It should also be noted that, for ease of description, only some, but not all, of the contents related to the present application are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

[0073] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0074] In a common existing implementation, a panoramic camera is installed in front of the activity area, facing the activity area and capable of capturing a panoramic image encompassing the entire area. When the faces of people in the activity area face the panoramic camera, the panoramic image captured by the camera will include the facial images of the people in the activity area. During attendance checking, the panoramic image captured by the camera is acquired in real time, and face detection is performed on the panoramic image to obtain a facial image. This facial image is then used to identify the person in the activity area, thereby confirming their successful check-in. While the panoramic camera has a wide field of view and can capture images encompassing the entire activity area, its short focal length makes it difficult to focus on people at a distance. Consequently, the clarity of the facial images of people in the back row of the activity area is low in the panoramic image, resulting in frequent misidentification or failure during identification, impacting the accuracy and efficiency of attendance checking. Alternatively, a pan-tilt camera can be installed in front of the activity area. This camera has a narrow field of view but a long focal length, meaning it can only capture a portion of the people in the activity area at a time, but can capture high-definition images of people at a distance. To address this, a pan-tilt camera can be used to track and capture high-definition facial images of participants, and then perform facial recognition on these images to confirm their identities. While the use of pan-tilt cameras can improve attendance accuracy, the high frequency of rotation and capture required by these cameras undoubtedly increases their wear and tear, shortening their service life.

[0075] To solve the above problems, this embodiment provides a non-contact attendance method that quickly identifies all people with their heads raised in the active area based on the first panoramic image captured by the panoramic camera, thereby accurately locating the target area of ​​the person with their heads raised in the active area, and controlling the pan-tilt camera to rotate to the target area for shooting to capture close-up images containing the faces of the people with their heads raised, thereby avoiding invalid rotation and invalid shooting of the pan-tilt camera and reducing the loss of the pan-tilt camera. By utilizing the long focal length of the pan-tilt camera to capture close-up images containing high-definition faces, even if the person with their heads raised is in the back row of the area, the high-definition face in the close-up image can be used for identity recognition to effectively identify the identity information of the corresponding person with their heads raised, thereby avoiding the frequent misidentification or recognition failure of active people in the back row during identity recognition, thereby improving the accuracy and efficiency of attendance.

[0076] The contactless attendance method provided in this embodiment can be executed by a contactless attendance device, which can be implemented by software and / or hardware. The contactless attendance device can be composed of two or more physical entities, or it can be composed of one physical entity. For example, the contactless attendance device can be a terminal device deployed in the activity area. For example, when the activity area is a conference room, the contactless attendance device can be a conference tablet device in the conference room. When the activity area is a classroom, the contactless attendance device can be a smart blackboard device in the classroom. In addition, the contactless attendance device can also be an attendance system or a server that provides attendance services.

[0077] The contactless attendance device is installed with at least one operating system, including but not limited to Android, Linux, and Windows. The contactless attendance device can install at least one application based on the operating system. The application can be a native application of the operating system or an application downloaded from a third-party device or server. In this embodiment, the contactless attendance device has at least one application that can execute the contactless attendance method.

[0078] For ease of understanding, this embodiment is described by taking a terminal device as an example of the main body for executing the contactless attendance method.

[0079] Figure 1 This is a flow chart of a non-sensing attendance method provided by an embodiment of the present application. Figure 1 As shown, the steps of the non-sensing attendance method include:

[0080] S110 , obtaining a first panoramic image obtained by photographing the activity area with a panoramic camera, and determining a plurality of first target objects in the first panoramic image, where the first target objects are images of people active in the activity area in the first panoramic image.

[0081] Participants are individuals participating in an organized activity, such as a class or meeting, such as students in a classroom or attendees in a meeting. An activity area is an area where participants gather during an organized activity, such as a student area in a classroom or a participant area in a conference room. This embodiment uses the student area as an example and the participants as students as an example. Figure 2 Schematic diagram of a classroom provided in the embodiment of the present application. Figure 2 As shown, a panoramic camera 12 is installed at the front side of the classroom 10, and the panoramic camera 12 faces the student area 11. The entire student area is within the field of view of the panoramic camera 12, and the panoramic camera 12 can capture an image containing the entire student area 11.

[0082] The first panoramic image is an image captured by the panoramic camera at the current moment, capturing the entire activity area. After attendance checking begins, the panoramic camera can capture the activity area at a certain frequency. If there are active people sitting in the activity area, images containing the active people will be captured. Person detection is performed on the first panoramic image to obtain person detection frames within the first panoramic image. Each person detection frame corresponds to a single active person. The person detection frame is identified as the first target object, i.e., the image of the active person in the activity area.

[0083] Optionally, a pre-trained object detection model may be used to perform person detection on the first panoramic image, and output a person detection frame in the first panoramic image, wherein the object detection model is trained using sample images marked with person detection frames.

[0084] In addition, if the first panoramic image is the first frame captured by the panoramic camera for this attendance check, people are detected in the first panoramic image using a pre-trained object detection model to obtain multiple person detection frames for active persons. People are then tracked in the first panoramic image based on the person detection frames in the previous image, obtaining the person detection frames in the first panoramic image. For example, if the first panoramic image is the second frame captured by the panoramic camera for this attendance check, people are tracked in the second frame based on each person detection frame in the first frame, obtaining the person detection frames in the second frame. Person detection frames for the same active person are assigned the same tracking identifier during tracking. For example, if the tracking identifier for active person A is 001, the tracking identifier for the person detection frames for active person A in both the first and second frames is 001. If a person appears in the second frame that cannot be associated with the person detection frame in the first frame, a new person is identified in the second frame, and the corresponding person detection frame is marked and assigned a new tracking identifier.

[0085] S120: Determine a head-up object among the multiple first target objects, and determine a target area corresponding to the head-up object in the active area.

[0086] Among them, the head-up object is the image of the head-up person in the activity area facing the panoramic camera in the first panoramic image, and the target area is the area where the head-up person corresponding to the head-up object is located. It can be understood that in the classroom, students face the blackboard, that is, facing the panoramic camera, and the panoramic camera can capture all students sitting in the student area. Only when the student looks up, the panoramic camera or the pan-tilt camera captures the student's facial image, and then performs identity recognition and sign-in authentication based on the facial image. In this regard, this embodiment aims to use the panoramic image captured by the panoramic camera to identify the head-up person in the activity area, thereby controlling the pan-tilt camera to capture the target area corresponding to the head-up object to record the high-definition face of the head-up person in the close-up image, and use the high-definition face in the close-up image to accurately identify the identity of the head-up person.

[0087] For example, since the panoramic camera can only capture the facial image of the active person when the active person is in a head-up posture, that is, if a face is detected in the first target object, it indicates that the active person is in a head-up posture. Based on this, face detection can be performed on each person detection frame in the first panoramic image. If a face is detected, the corresponding person detection frame is confirmed to be a head-up object.

[0088] It should be noted that face detection is to determine whether the active person is in a head-up posture at the current moment, but the head-up posture at the current moment may be the ending posture of the head-up person. After confirming the target area corresponding to the head-up object, the pan-tilt camera will be controlled to turn to the target area for shooting. Due to the existence of steering delay and shooting delay, the head-up person will have completed the head-up action and is in a head-down posture when the pan-tilt camera turns to the target area, and the pan-tilt camera will not be able to capture the face image of the person. In this regard, based on the first panoramic image and the historical panoramic images previously taken by the panoramic camera, it can be jointly judged whether the behavior of the active person in the corresponding time period is a head-up action, and the head-up person and the area where he is located can be predicted, so as to control the pan-tilt camera to turn to the corresponding target area to shoot the head-up person before the head-up person lowers his head, so as to collect his high-definition face image. Optional, Figure 3 This is a flowchart of determining the header object provided by the embodiment of the present application. Figure 3 As shown, the step of determining the header object specifically includes S1201-S1203:

[0089] S1201: Acquire posture information of a first target object and a second target object of an active person.

[0090] In this embodiment, the first target object in the first panoramic image is obtained by tracking the second target object in the historical panoramic image. The second target object is an image of an active person in the historical panoramic image. The second target object and the corresponding tracked first target object correspond to the same active person. The historical panoramic image was captured by the panoramic camera of the active area before the first panoramic image was captured. For example, the historical panoramic image can be a sequence of images captured by the panoramic camera within a preset time period before the capture timestamp of the first panoramic image. The first target object in the first panoramic image is obtained by tracking the two target objects in the image captured by the panoramic camera at the previous moment.

[0091] As can be seen from the above, when tracking the person detection frame in the first panoramic image, the person detection frame of the same active person will be marked with the same tracking identifier during tracking. Based on the tracking identifiers of the person detection frames in the historical panoramic image and the first panoramic image, the person detection frame in the first panoramic image and the person detection frame with the same tracking identifier in the second panoramic image can be respectively determined as the first target object and the second target object of the active person, and the posture information of the first target object and the second target object of the active person can be obtained. The posture information is the head posture presented by the active person's imaging, including lowered head, half lowered head, half raised head, and raised head.

[0092] It should be noted that steps S1201-S1203 are only executed if a historical panoramic image exists. If no historical panoramic image exists, indicating that the first panoramic image is the first frame captured by the panoramic camera after attendance logging began, the head-up object can be determined based on the posture of each first target object in the first panoramic image. Furthermore, if a new active person appears in the first panoramic image, meaning that no second target object with the same tracking identifier exists for this new active person in the historical panoramic image, the new active person's posture is also used to directly determine whether they are a head-up person.

[0093] S1202: Determine the corresponding behavior of the activity person according to the posture information of the first target object and the second target object of the activity person.

[0094] Exemplarily, based on multiple person detection frames with the same tracking identifier, the postures of the people in these person detection frames are analyzed frame by frame in chronological order, and whether the posture changes meet the posture change conditions of the head-raising action, if so, it is determined that the active persons corresponding to these person detection frames have made a head-raising action, otherwise it is determined that the active persons corresponding to these person detection frames have not made a head-raising action. For example, if the postures of the active person in the historical panoramic image and the first panoramic image transition from half-lowering the head to half-raising the head, or from half-raising the head to raising the head, and remain in the head-raising state, and are in a half-raising or raising-up posture in the first panoramic image, then it can be confirmed that the active person has made a head-raising action.

[0095] S1203: When the action of the active person is raising his head, determine the first target object as the raising-head object.

[0096] For example, if the active person makes a head-up gesture in both the first panoramic image and the historical panoramic image, the first target object of the active person is determined to be the head-up object. Furthermore, if the active person is a newly added active person in the first panoramic image and is in a head-up gesture, the first target object of the active person is also determined to be the head-up object.

[0097] This embodiment uses multiple images of the active person in the first panoramic image and the historical panoramic image to predict the active person's behavior and actions, and confirms that the active person is a head-up person when it is predicted that the active person makes a head-up action, thereby determining the corresponding first target object as the head-up object, to ensure that the pan-tilt camera is controlled to capture the high-definition face of the corresponding active person before the active person finishes the head-up action, to avoid invalid rotation and invalid shooting of the pan-tilt camera, and to reduce the use loss of the pan-tilt camera.

[0098] After determining the head-up objects, the target area corresponding to the head-up person in the activity area can be located based on the pixel coordinates of the head-up objects in the first panoramic image. Optionally, the area containing all head-up objects can be determined based on the pixel coordinates of each head-up object. This area can then be mapped to candidate areas in the activity area, and the positions of each head-up object in the candidate areas can be determined. Based on the field of view of the gimbal camera, the area containing the most head-up objects in the candidate areas can be determined as the target area.

[0099] In addition, the active area can be pre-divided into multiple sub-areas, and the target area can be determined in the multiple sub-areas based on the head-up objects in each sub-area. Optionally, the target area can be determined based on the number of head-up objects in each sub-area. Figure 4 This is a flow chart of determining a target area in multiple sub-areas provided by an embodiment of the present application. Figure 4 As shown, the step of determining the target area in the multiple sub-areas specifically includes S1204-S1207:

[0100] S1204: Obtain the location information of each head-up object.

[0101] The position information is the pixel coordinates of the corresponding head-up object in the first panoramic image, and the pixel coordinates of the head-up object in the first panoramic image are obtained according to the pixel points of the head-up object.

[0102] S1205: Determine the sub-region where the head-up object is located according to the position information of the head-up object.

[0103] The sub-area is the sub-area where the head-up person corresponding to the head-up object is located. Figure 2 , the student area 11 is divided into seven sub-areas 14, and the seven sub-areas 14 are sub-area ①, sub-area ②, sub-area ③, sub-area ④, sub-area ⑤, sub-area ⑥, and sub-area ⑦. The actual coordinates of the head-up person corresponding to the head-up object in the student area can be determined based on the pixel coordinates of the head-up object in the first panoramic image and the external and internal parameters of the panoramic camera. The actual coordinates of the head-up person are compared with the coordinate ranges of each sub-area. If the actual coordinates of the head-up person fall within the coordinate range of a certain sub-area, the sub-area is determined to be the sub-area where the head-up person is located. For example, if the actual coordinates of head-up person A fall within the coordinate range of sub-area ⑦, the sub-area where head-up person A is located is determined to be sub-area ⑦.

[0104] S1206: Count the number of head-up objects in multiple sub-areas according to the sub-areas where the head-up objects are located.

[0105] Illustratively, according to the sub-regions where the head-up objects are located, the head-up objects in the same sub-region are divided into the same set, and the number of head-up objects in the set is counted as the number of head-up objects in the corresponding sub-region.

[0106] It should be noted that during the attendance process, if all active personnel in a sub-area have successfully signed in, indicating that attendance has been completed for that sub-area, the PTZ camera does not need to rotate to capture images in these sub-areas. This reduces unnecessary rotation and capture, and reduces energy consumption and wear on the camera. Accordingly, for sub-areas where attendance has been completed, the number of head-up objects corresponding to the sub-areas where attendance has been completed does not need to be counted. Instead, the number of head-up objects in sub-areas where attendance has not been completed is counted, ensuring that the PTZ camera only needs to rotate to capture close-up images of these sub-areas where attendance has not been completed.

[0107] Specifically, the number of people who have successfully signed in in each sub-area is counted; when the number of people in the sub-area does not reach the threshold number of people in the sub-area, the number of head-up objects in the sub-area is counted. The number of people in attendance is the total number of active people who have successfully signed in in the sub-area. The initial number of people in the sub-area is set to zero. Every time the identity information of the active person is successfully identified and signed in based on the face image in the close-up image corresponding to the sub-area, the number of people in attendance in the sub-area is increased by one. Get the accumulated number of people in attendance in the current sub-area, and compare the number of people in attendance in the sub-area with the threshold number of people. The threshold number of people can be the number of people seated in advance in the corresponding sub-area, for example Figure 2The figure specifies that N students are seated in subarea ②. However, some activity areas allow for random seating, and the number of seats in a subarea is not specified. In this case, the number threshold is the number of target objects in the subarea—that is, the number of first target objects detected in the subarea from the first panoramic image. If a subarea has a specified number of seats, and someone is late arriving, the attendance count for that subarea will remain below the threshold. However, all participants in that subarea have already signed in, and the PTZ camera will not be able to capture valid images during this period. This will result in additional ineffective rotations and shots. To address this issue, the number of target objects in a subarea can be set as the threshold. When the number of participants in a subarea equals the target number, it is confirmed that all participants who have arrived in the subarea have successfully signed in, thus confirming that attendance for that subarea has been completed. If the number of participants in a subarea is less than the target number, it is confirmed that some participants who have arrived in the subarea have not yet signed in, thus confirming that attendance for that subarea has not been completed. In sub-areas where attendance has not been completed, the number of objects with their heads raised is counted so that the PTZ camera can be controlled to rotate to capture close-up images of the sub-areas where attendance has not been completed. This embodiment determines whether attendance has been completed in a sub-area by comparing the number of people who have logged in in the sub-area with a threshold number of people. If attendance has not been completed, the number of objects with their heads raised in the sub-area is counted again. This prevents the PTZ camera from rotating to capture close-up images of sub-areas where attendance has been completed. This reduces unnecessary rotation and shooting of the PTZ camera, lowers its energy consumption and wear, and avoids repeated attendance verification for sub-areas where attendance has been completed, thereby improving attendance efficiency.

[0108] It should be noted that if the number threshold is set based on the number of target objects, then when the number of target objects in the sub-area identified based on the first panoramic image increases, it can be confirmed that new active personnel have entered the sub-area, thereby adjusting the number threshold and attendance status of the sub-area.

[0109] S1207: Determine a target area in the multiple sub-areas according to the number of head-up objects in the multiple sub-areas.

[0110] For example, the sub-area with the largest number of head-up objects can be identified as the target area so that the PTZ camera can capture close-up images of the most human faces. The more facial images captured in the close-up images, the more people who successfully check in, and the higher the attendance success rate. This embodiment determines the number of head-up objects in each sub-area within the activity area based on the position information of each head-up object. Based on the number of head-up objects in each sub-area, sub-areas with higher attendance success rates are preferentially selected as target areas, which helps improve attendance efficiency and reduce the number and frequency of PTZ camera turns.

[0111] Optionally, the head-up rate of a sub-area can be determined based on the number of head-up objects and the number of target objects in the sub-area; and the target area can be determined in the multiple sub-areas based on the head-up rates of multiple sub-areas. The number of target objects is the number of active persons in the sub-area, and the actual coordinates of the active persons corresponding to each first target object in the active area can be determined based on the pixel coordinates of each first target object in the first panoramic image, and the sub-area where the active persons are located can be determined based on the actual coordinates of the active persons. The number of target objects can be obtained by counting the number of active persons in each sub-area based on the sub-area where each active person is located. The ratio of the number of head-up objects to the number of target objects in the same sub-area is used as the head-up rate of the sub-area, and the sub-area with the highest head-up rate is determined as the target area. For example, Figure 5 Schematic diagram of the attendance status of each sub-area provided in the embodiment of the present application. Figure 5 As shown, the position of each sub-area in the active area can be generated Figure 5 In the attendance status table, the number of people present in sub-area ② equals the target number of objects. Therefore, sub-area ② is in the completed attendance state, and the headed object count for sub-area ② does not need to be counted. However, the number of people present in sub-areas ⑤ and ⑥ is less than the corresponding target number of objects. Therefore, sub-areas ⑤ and ⑥ are in the incomplete attendance state, and the headed object count for sub-areas ⑤ and ⑥ is counted. Since the headed object rate in sub-area ⑤ is higher than that in sub-area ⑥, sub-area ⑤ can be selected as the target area.

[0112] It can be understood that the higher the head-up rate, the more likely the PTZ camera is to capture the facial images of all active people who have not signed in in the corresponding sub-area. The probability that the sub-area will reach the completed attendance state after capturing the close-up image is greater, and the sub-area that has reached the completed attendance state will be eliminated from the candidate list of the target area. The number of sub-areas that the PTZ camera can turn to is reduced, which is conducive to reducing the number of rotations of the PTZ camera. In this regard, the sub-area with the highest head-up rate can be determined as the target area. In this embodiment, the sub-area with a higher probability of reaching the completed attendance state is preferentially selected as the target area based on the head-up rate of the sub-area, so as to minimize the number of sub-areas that the PTZ camera can turn to, thereby reducing the number and frequency of rotations of the PTZ camera.

[0113] Furthermore, when determining the target area based on the head-up rate, if the sub-area with the highest head-up rate is the same as the sub-area that the PTZ camera is continuously photographing, if the continuous photographing time of the sub-area with the highest head-up rate reaches the preset time threshold, the sub-area with the second highest head-up rate will be determined as the target area; otherwise, the sub-area with the highest head-up rate will be determined as the target area. The continuous photographing time is the time that the PTZ camera stays in the sub-area currently photographed. Figure 5If the gimbal camera is always shooting towards sub-area ⑤, but the active personnel who have not signed in successfully in sub-area ⑤ are lying down and sleeping, they do not raise their heads, while other active personnel always keep their heads up. As a result, the head-up rate in sub-area ⑤ is always the highest, but no effective close-up images can be captured. The other sub-areas cannot capture close-up images, which affects the attendance rate of other sub-areas.

[0114] If the sub-area with the highest head-up rate is the same as the sub-area where the PTZ camera is currently stationed, the length of time the PTZ camera has remained in that sub-area can be determined. If this length of time reaches a preset threshold, it indicates that the PTZ camera has not captured valid images for attendance verification for a long period of time. In this case, the sub-area with the next highest head-up rate can be designated as the target area, and the PTZ camera can be promptly redirected to capture valid close-up images of other sub-areas, thereby improving the attendance rate in other sub-areas and, therefore, overall attendance efficiency. If the sub-area with the highest head-up rate is not the sub-area where the PTZ camera is currently stationed, or if the sub-area with the highest head-up rate is the same as the sub-area where the PTZ camera is currently stationed but the duration of the stay does not reach the preset threshold, it indicates that the PTZ camera may capture valid images of the sub-area with the highest head-up rate in the future. Therefore, the sub-area with the highest head-up rate can be designated as the target area. This embodiment, by promptly redirecting the PTZ camera to capture valid close-up images of other sub-areas when it detects that the PTZ camera has remained in a particular sub-area for an extended period of time, improves the attendance rate in other sub-areas and, therefore, overall attendance efficiency.

[0115] In another embodiment, the target area can be determined in multiple sub-areas based on the number of non-signed-in objects in the head-up objects of each sub-area. Specifically, the non-signed-in objects and the signed-in objects can be distinguished in the head-up objects based on the actual coordinates corresponding to each head-up object and the actual coordinates corresponding to the signed-in objects. Among them, the non-signed-in objects are images of active persons who have not been authenticated through face recognition, and the signed-in objects are images of active persons who have been authenticated through face recognition. It should be noted that the signed-in objects can be authenticated through close-up images or the first panoramic image, and when the sign-in is successful, the actual coordinates of the corresponding active persons in the activity area can be determined based on the pixel coordinates of the signed-in objects in the close-up image or the first panoramic image. The actual coordinates corresponding to the head-up objects can be determined based on the pixel coordinates of the head-up objects in the first panoramic image.

[0116] After distinguishing between non-signed-in and signed-in subjects in the head-up objects, the number of non-signed-in subjects is counted. If the number of non-signed-in subjects is the same as the number of remaining attendance subjects in the sub-area, the sub-area is identified as the target area. The remaining attendance subject is the difference between the number of target subjects in the sub-area and the number of attendance subjects, representing the number of people in the sub-area who have not yet successfully signed in. When the number of non-signed-in subjects is the same as the number of remaining attendance subjects in the sub-area, the head-up objects in the sub-area contain all remaining active persons who have not signed in. The PTZ camera then captures the sub-area, capturing high-definition close-up images of the faces of all remaining active persons who have not signed in. This allows for simultaneous identity authentication and sign-in for all remaining active persons in the sub-area, increasing the probability that the sub-area will complete attendance after the close-up image is captured. Sub-areas that have completed attendance are removed from the candidate list of target areas, reducing the number of sub-areas the PTZ camera can turn to, thereby reducing the number of PTZ camera rotations.

[0117] If the number of non-attendants and the number of remaining attendees in each sub-area are different, the sub-area with the largest number of non-attendants is selected as the target area, or the sub-area with the largest ratio of the number of non-attendants to the number of remaining attendees is selected as the target area.

[0118] S130 , controlling the pan-tilt camera to turn toward the target area, and obtaining a close-up image of the target area captured by the pan-tilt camera.

[0119] The close-up image is captured by a gimbaled camera pointing toward the target area. A gimbaled camera has a narrow field of view but a deep focal length. By rotating the gimbal, it can capture high-definition images of various local areas within the active area. By turning the gimbaled camera toward the target area, a close-up image of the target area can be captured, capturing the high-definition face of the person looking up in the target area.

[0120] When the target area is not a sub-area, the shooting focal length and shooting angle of the gimbal camera are determined according to the positional relationship between the target area and the active area. The gimbal camera is controlled to turn toward the target area according to the shooting angle. The focal length of the gimbal camera is adjusted according to the shooting focal length, and the gimbal camera is controlled to shoot the target area to obtain a close-up image.

[0121] When the target area is a sub-area, the shooting focal length and shooting angle corresponding to each sub-area are set in advance on the gimbal camera according to the positional relationship between each sub-area and the active area. After the target area is determined, the gimbal camera is controlled to turn to the target area according to the pre-set shooting angle of the target area, the focal length of the gimbal camera is adjusted according to the shooting focal length, and the gimbal camera is controlled to shoot the target area to obtain a close-up image.

[0122] After the PTZ camera turns to the target area, it captures a single close-up image of the target area. If only a single close-up image is captured, poor image quality can affect the subsequent success rate of identity recognition. To improve this success rate, the PTZ camera can be controlled to capture multiple close-up images of the target area within a first preset duration based on a preset frequency. The first preset duration is the length of time the PTZ camera remains in the target area for each capture. Using multiple close-up images to identify individuals in the target area improves the success rate of identity recognition and, in turn, attendance efficiency.

[0123] In this embodiment, since activities in an activity area occur at different time periods and participants raise their heads at different frequencies, time periods with a higher frequency of raising their heads can be pre-defined as key time periods based on the nature of the activities. This increases the pan-tilt camera's capture frequency during these key time periods, thereby improving the camera's success rate in capturing high-definition faces of multiple participants. Specifically, if the activity in the activity area falls within the key time period, the pan-tilt camera is controlled to capture close-up images of the target area at a first preset frequency. If the activity in the activity area does not fall within the key time period, the pan-tilt camera is controlled to capture close-up images of the target area at a second preset frequency, where the first preset frequency is higher than the second preset frequency. For example, the 20 seconds before class, the 5 minutes after class, the 3 minutes before get out of class, and the 30 seconds after class are the time periods in which students frequently raise their heads, and therefore these time periods are designated as key time periods. During attendance checking, if the current time falls within one of these key time periods, the pan-tilt camera is controlled to capture close-up images of the target area at the higher first preset frequency. This allows for more close-up images within the first preset duration, increasing the camera's success rate in capturing high-definition faces, thereby improving identity recognition success and attendance checking efficiency. When the current moment is not within the aforementioned critical time period, the PTZ camera is controlled to capture close-up images of the target area at a lower second preset frequency, so that fewer close-up images are captured within the first preset duration, thereby reducing the power consumption of the PTZ camera and conserving battery power. This embodiment sets a critical time period so that the PTZ camera captures close-up images of the target area at a higher frequency during the critical time period, thereby increasing the success rate of the PTZ camera in capturing high-definition faces, thereby improving the success rate of identity recognition and attendance efficiency.

[0124] S140: Check-in and authentication of the active personnel in the target area based on the close-up image.

[0125] Exemplarily, when the gimbal camera only captures a close-up image of a target area, face detection is performed on the close-up image to obtain a face detection frame in the close-up image, identity recognition is performed based on the face detection frame to determine the identity information of the corresponding active person, and the active person's successful sign-in is confirmed based on the identity information of the active person. The process of identity recognition based on the face detection frame is as follows: extracting the corresponding first facial feature based on the face detection frame, matching the extracted first facial feature with the second facial feature in a preset facial feature library, and when the match is successful, determining the identity information associated with the corresponding matched second facial feature as the identity information of the active person corresponding to the face detection frame. The preset facial feature library stores the second facial features of all active persons, and each second facial feature is associated with the identity information of the corresponding active person and stored. The second facial feature is extracted from a reference facial image pre-photographed by the active person.

[0126] It should be noted that the reference facial image used to extract the second facial features is captured by the participant after a attendance request is made. After capturing the facial image, the participant authorizes the use of the image for attendance purposes and authorizes the capture of images of the participant in real time during the activity. Therefore, in this embodiment, all images of the participant are captured and used with the participant's authorization.

[0127] When the gimbal camera captures multiple close-up images of the target area, it can detect the face detection frames in the multiple close-up images one by one, perform feature extraction and matching on each face detection frame to confirm that the face detection frame corresponds to the identity information of the active person, and then confirm that the active person has successfully signed in based on the identity information of the active person. Because there are multiple face detection frames of an active person in the multiple close-up images, if feature extraction and matching are performed based on the face detection frames in each close-up image, the feature extraction and matching of multiple face detection frames of the active person will be repeated. Even if the active person has successfully signed in, their face detection frame will be used for repeated sign-in authentication, which not only affects attendance efficiency but also repeats many invalid calculations.

[0128] Optional, Figure 6 This is a flow chart of the check-in authentication of active personnel based on multiple close-up images provided by the embodiment of the present application. Figure 6 As shown, the step of checking in and authenticating the activity personnel based on multiple close-up images specifically includes S1401-S1402:

[0129] S1401: Detect multiple facial images in multiple close-up images, and divide the multiple facial images into multiple facial image sets, where each facial image set corresponds to one active person.

[0130] For example, a pre-trained face detection model is used to perform face detection on multiple close-up images. A face detection frame is output for each close-up image, and the image area marked by the face detection frame is used as the face image. Based on the distance and similarity between the face images in different close-up images, multiple face images corresponding to the same active person are identified and grouped into the same face image set. The face detection model is trained using multiple sample images marked with face detection frames.

[0131] Optionally, face detection is performed on the first close-up image captured within a first preset duration to obtain at least one face image; face tracking is performed on the remaining close-up images within the first preset duration based on the face image, and the face image in the first close-up image and the corresponding tracked face images in the remaining close-up images are grouped into the same face image set. Exemplarily, face detection is performed on the first close-up image using a pre-trained face detection model to obtain the face image in the first close-up image; face tracking is performed on the remaining close-up images based on the face image to find related face images. During the tracking process, the related face images are assigned the same tracking identifier as the face image used for tracking, so that the tracking identifier indicates which face images belong to the same active person. Multiple face images with the same tracking identifier can be grouped into the same face image set. This embodiment uses the face image detected in the first close-up image to track the remaining close-up images, allowing for rapid identification of multiple face images of an active person in multiple close-up images, thereby improving the efficiency of detecting multiple face images of an active person.

[0132] S1402: Identify the identity information of the corresponding activity personnel based on the face image set, and confirm that the corresponding activity personnel have successfully signed in based on the identified identity information.

[0133] Exemplarily, the facial images in the facial image set are sorted in chronological order according to the shooting time of the corresponding close-up images, and the identity information of the corresponding active person is identified starting from the first facial image in the facial image set. If the identification is successful, the corresponding active person is confirmed to have successfully signed in based on the identified identity information, and the identity identification of the facial image of the active person is no longer performed; if the identification fails, the next facial image is obtained from the facial image set, and the identity information of the active person is identified based on the next facial image until the identity identification of the active person is successful, and the corresponding active person is confirmed to have successfully signed in based on the identified identity information. This embodiment avoids repeated identity identification and improves attendance efficiency by confirming that the active person has successfully signed in and no longer identifying the facial image of the active person after successfully identifying the identity information of the corresponding active person based on the facial image of the active person.

[0134] Optionally, the image quality of each facial image in the facial image set is evaluated to select the facial image with the highest image quality. The identity information of the corresponding event attendee is identified based on the facial image with the highest image quality, and the corresponding event attendee's successful check-in is confirmed based on the identified identity information. Image quality can be considered a quality score obtained by evaluating the quality of the facial image. For example, the quality of the facial image can be comprehensively evaluated based on multiple aspects, such as facial image clarity, illumination uniformity, and facial angle, to obtain the corresponding quality score. Specifically, a corresponding clarity score can be determined based on the numerical range of the facial image clarity, a corresponding clarity score can be determined based on the numerical range of the facial image illumination uniformity, and a corresponding facial angle score can be determined based on the numerical range of the facial image angle. The scores of these three dimensions can be weighted and summed to obtain the facial image quality score. A higher quality score indicates a higher success rate for identity recognition of the corresponding facial image. The quality scores of multiple facial images in the facial image set can be compared, and the facial image with the highest quality score can be selected for identity recognition. After successful identity recognition, the corresponding event attendee's successful check-in is confirmed based on the identified identity information. This embodiment selects the facial image with the highest image quality from multiple facial images corresponding to active personnel for identity recognition, thereby avoiding identity recognition on low-quality facial images, effectively reducing the number of identity recognition times on facial images, improving identity recognition efficiency, and thus improving attendance efficiency.

[0135] It should be noted that the first panoramic image captured by the panoramic camera will also contain facial images, so face detection is also performed on the first panoramic image to obtain facial images, and identity recognition is performed on the facial images to confirm the identity information of the corresponding active persons. When the recognition is successful, the corresponding active persons are signed in according to the recognized identity information. However, due to the short focal length of the panoramic camera, the images of the active persons in the front row captured by it are clearer, while the images of the active persons in the back row are blurry. Therefore, the panoramic camera can be used to capture clear facial images of the active persons in the front row, and the gimbal camera can be used to capture clear facial images of the active persons in the middle and back rows. Accordingly, when dividing the activity area into sub-areas, the front row area in the activity area that is under the responsibility of the panoramic camera can be excluded from the sub-area that is under the responsibility of the gimbal camera. Reference Figure 2 , sub-area ① is the front area in the active area, then sub-area ① is not within the sub-area range of the gimbal camera's rotation, that is, the gimbal camera only needs to turn to sub-area ②, sub-area ③, sub-area ④, sub-area ⑤, sub-area ⑥, and sub-area ⑦, which is beneficial to reduce unnecessary rotation and shooting of the gimbal camera.

[0136] From the above, it can be seen that the panoramic images captured by the panoramic camera are used for both facial image detection and head-up object detection. If the resolution of the panoramic image is high, the computational complexity of detecting head-up objects will be high, affecting the detection efficiency of head-up objects and increasing the delay of the gimbal camera turning to the target area corresponding to the head-up object; if the resolution of the panoramic image is low, the success rate of facial image recognition will be low, affecting attendance efficiency. Therefore, different resolution images can be used for different purposes according to the purpose of the panoramic images captured by the panoramic camera. Specifically, high-resolution panoramic images can be used for face detection and identity recognition, and low-resolution panoramic images can be used for head-up object detection.

[0137] Optional, Figure 7 This is a flow chart of face detection based on high-resolution panoramic images provided by the embodiment of the present application. Figure 7 As shown, the steps of performing face detection based on the high-resolution panoramic image specifically include S210-S230:

[0138] S210: Acquire a second panoramic image obtained by photographing the activity area with a panoramic camera, where the resolution of the second panoramic image is higher than that of the first panoramic image.

[0139] The second panoramic image is a high-resolution panoramic image captured by the panoramic camera at the current moment. For example, when the panoramic camera captures the active area, software is used to simultaneously generate the high-resolution second panoramic image and the low-resolution first panoramic image. For example, the second panoramic image may be at a 4K resolution, and the first panoramic image may be at a 1080P resolution.

[0140] S220: Perform face detection on the second panoramic image to obtain at least one face image.

[0141] Exemplarily, face detection is performed on the second panoramic image using a pre-trained face detection model to obtain at least one face detection frame, and the image marked by the face detection frame is used as the face image.

[0142] Optionally, if the second panoramic image is the first high-resolution panoramic image captured after the start of attendance, face detection is performed on the second panoramic image using a pre-trained face detection model to obtain a face image in the second panoramic image. If the second panoramic image is not the first high-resolution panoramic image captured after the start of attendance, face tracking detection is performed on the second panoramic image based on the face image in the previous high-resolution panoramic image to obtain a face image in the second panoramic image. During the tracking and detection process, all facial images of active personnel will be marked with the same tracking identifier so that when confirming that the active personnel have successfully signed in, the facial image of the active personnel will no longer be identified, thus eliminating unnecessary identification steps.

[0143] S230: When the image quality of the facial image is greater than or equal to a preset quality threshold, identify the identity information of the corresponding activity person according to the facial image, and confirm that the corresponding activity person has successfully signed in based on the identified identity information.

[0144] Among them, the preset quality threshold can be regarded as the minimum quality score when the face image can be successfully identified. Exemplarily, the quality of the face image is comprehensively scored by taking into account aspects such as the clarity of the face image, the uniformity of the lighting and the angle of the face image to obtain a quality score, and the quality score of the face image is compared with the preset quality threshold. If the quality score of the face image is greater than or equal to the preset quality threshold, it indicates that the face image can be used to successfully identify the identity information of the corresponding activity person, and thus the identity information of the corresponding activity person can be obtained based on the identity identification of the face image, and the sign-in of the corresponding activity person is confirmed based on the identified identity information. If the quality score of the face image is less than the preset quality threshold, it indicates that the face image cannot be used to identify the identity information of the corresponding activity person, and therefore the face image is not used for identity identification to avoid wasting computing resources.

[0145] Optionally, after identifying the identity information of the active person and signing in, the tracking identifier of the active person can be marked as the tracking identifier of the person who has successfully signed in. Accordingly, when the facial image marked with the tracking identifier is detected after the subsequent second panoramic image, it can be confirmed based on the tracking identifier of the person who has successfully signed in that the active person in the facial image has successfully signed in, and the identity of the facial image will not be identified, thereby avoiding repeated identification of the facial images of the person who has successfully signed in, saving computing resources.

[0146] Of course, the activity area can be divided not only into the sub-area for the PTZ camera to check attendance, but also into the sub-area for the panoramic camera to check attendance. When all the active people in the sub-area for the panoramic camera to check attendance have successfully signed in, the sub-area is converted to the attendance completion state, and the panoramic camera does not need to continue to shoot the second panoramic image, so as to save the computing resources of the panoramic camera. For example, refer to Figure 2 ,Subarea ① is the subarea where the panoramic camera is responsible for attendance.,If the number of attendance persons in subarea ① is equal to the number of target objects, it can be determined that subarea ① is in the attendance completion state, and the panoramic camera is controlled to no longer generate the second panoramic image.

[0147] It should be noted that when the second panoramic image is performing face detection, it may also detect high-definition faces of people active in other sub-areas outside of sub-area ①. After the identity information of the active person is recognized based on the high-definition face and the person signs in, the attendance number of other sub-areas can be updated synchronously, so as to quickly confirm the identity information and sign in of the active people in each sub-area with the close-up image of the PTZ camera, thereby improving attendance efficiency. Moreover, if, after updating the attendance number of other sub-areas, it is confirmed that the attendance number of a certain sub-area has reached the target number of objects, then the sub-area is confirmed to be in a completed attendance state, and the PTZ camera no longer needs to turn to that sub-area. Therefore, the introduction of the second panoramic image can help reduce the number of rotations of the PTZ camera, thereby reducing the use loss of the PTZ camera.

[0148] This embodiment uses a second panoramic image captured by the panoramic camera to coordinate with the close-up image captured by the PTZ camera for identity recognition and check-in authentication, improving attendance efficiency. Furthermore, this second panoramic image is implemented in software, without adding to the hardware cost of the panoramic camera. The inclusion of this second panoramic image allows each sub-area to complete attendance more quickly, reducing the number of PTZ camera rotations and thus reducing wear and tear.

[0149] On the basis of the above embodiment, in order to more clearly understand the process of processing the first panoramic image, the second panoramic image and the close-up image to record the attendance of the active personnel, Figure 8 The image processing flow shown in FIG is described as an example. Figure 8 As shown, a panoramic camera captures a first panoramic image and a second panoramic image. Face detection is performed on the second panoramic image to obtain a facial image. Facial features are extracted from the facial image and matched with facial features stored in a preset facial feature library. If a match is successful, the identity information stored in association with the facial features is determined as the identity information of the person participating in the facial image. The person participating in the activity is then signed in based on the identity information. Simultaneously, person detection is performed on the first panoramic image to obtain first target objects. Head-up objects are identified based on the posture of each first target object. The head-up rate of each subregion is determined based on the subregion where the head-up object is located. A target region is determined within each subregion based on the head-up rate. The pan-tilt camera is controlled to rotate toward the target region based on the pre-configured shooting angle and focal length corresponding to the target region to capture a close-up image. Face detection is performed on the close-up image to obtain a facial image. Facial features are extracted from the facial image and matched with facial features stored in a preset facial feature library. If a match is successful, the identity information stored in association with the facial features is determined as the identity information of the person participating in the facial image. The person participating in the activity is then signed in based on the identity information of the person participating in the activity.

[0150] In summary, the contactless attendance method provided by the embodiment of the present application uses a panoramic camera to capture a first panoramic image of the activity area, detects the images of each active person in the activity area in the first panoramic image, determines the head-up object based on the images of each active person, and thus determines the target area where the head-up object is located in the activity area, controls the pan-tilt camera to turn to the target area, and then captures a close-up image of the target area through the pan-tilt camera. Based on the high-definition face recognition in the close-up image, the identity information of the corresponding active person is obtained and the sign-in authentication is performed. Through the above-mentioned technical means, all the people with their heads raised in the activity area can be quickly identified based on the first panoramic image captured by the panoramic camera, thereby accurately locating the target area of ​​the person with their heads raised in the activity area, and controlling the pan-tilt camera to turn to the target area to shoot to capture a close-up image containing the face of the person with their heads raised, thereby avoiding invalid rotation and invalid shooting of the pan-tilt camera and improving the utilization rate of the pan-tilt camera. Since the focal length of the pan-tilt camera is long, it can capture a close-up image containing a high-definition face when shooting the target area at any position in the activity area, ensuring the clarity of the face image used for identity recognition and improving the accuracy and success rate of identity recognition. Even if the person looking up is in the back row of the area, the identity information of the corresponding person looking up can be effectively identified by using the high-definition face in the close-up image for identity recognition, which solves the problem of frequent misidentification or recognition failure when performing identity recognition for people in the back row in the existing technology, and improves the accuracy and efficiency of attendance.

[0151] Figure 9 This is a structural diagram of a non-sensing attendance device provided in an embodiment of the present application, refer to Figure 9 The contactless attendance device includes: a processor 31, a memory 32, a communication device 33, an input device 34, and an output device 35. The number of processors 31 in the contactless attendance device can be one or more, and the number of memories 32 in the contactless attendance device can be one or more. The processor 31, memory 32, communication device 33, input device 34, and output device 35 of the contactless attendance device can be connected via a bus or other means.

[0152] The memory 32, as a computer-readable storage medium, can be used to store software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the contactless attendance method of any embodiment of the present application. The memory 32 may mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc. In addition, the memory 32 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories can be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0153] The communication device 33 is used for data transmission.

[0154] The processor 31 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 32, thereby realizing the above-mentioned contactless attendance method.

[0155] The input device 34 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 35 may include a display device such as a display screen.

[0156] The contactless attendance device provided above can be used to execute the contactless attendance method provided in the above embodiment, and has corresponding functions and beneficial effects.

[0157] An embodiment of the present application also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute a contactless attendance method. The contactless attendance method includes: obtaining a first panoramic image obtained by shooting an activity area with a panoramic camera, determining multiple first target objects in the first panoramic image, where the first target objects are images of active people in the activity area in the first panoramic image; determining a head-up object among the multiple first target objects, and determining a target area corresponding to the head-up object in the activity area; controlling the pan-tilt camera to turn to the target area, obtaining a close-up image of the target area shot by the pan-tilt camera; and performing sign-in authentication on active people in the target area based on the close-up image.

[0158] Storage medium - any of various types of memory devices or storage devices. The term "storage medium" is intended to include: installation media, such as CD-ROMs, floppy disks, or tape drives; computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memory, such as flash memory, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc. Storage media may also include other types of memory or combinations thereof. In addition, the storage medium may be located in the first computer system in which the program is executed, or it may be located in a different second computer system that is connected to the first computer system via a network (such as the Internet). The second computer system can provide program instructions to the first computer for execution. The term "storage medium" may include two or more storage media residing in different locations (e.g., in different computer systems connected via a network). The storage medium may store program instructions (e.g., embodied as a computer program) that can be executed by one or more processors.

[0159] Of course, the storage medium containing computer-executable instructions provided in the embodiment of the present application, whose computer-executable instructions are not limited to the above-mentioned contactless attendance method, can also execute related operations in the contactless attendance method provided in any embodiment of the present application.

[0160] The contactless attendance device, storage medium and contactless attendance equipment provided in the above embodiments can execute the contactless attendance method provided in any embodiment of the present application. For technical details not described in detail in the above embodiments, please refer to the contactless attendance method provided in any embodiment of the present application.

[0161] The above are only preferred embodiments of the present application and the technical principles employed. The present application is not limited to the specific embodiments described herein, and any obvious changes, readjustments, and substitutions that are apparent to those skilled in the art will not depart from the scope of protection of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of the present application. The scope of the present application is determined by the scope of the claims.

Claims

1. A non-sensing attendance method, characterized in that: include: Acquire a first panoramic image obtained by photographing the activity area with a panoramic camera, and determine a plurality of first target objects in the first panoramic image, where the first target objects are images of people active in the activity area in the first panoramic image; Determine a head-up object among the multiple first target objects, and determine a target area corresponding to the head-up object in the activity area; Controlling the pan-tilt camera to turn toward the target area, and acquiring a close-up image of the target area captured by the pan-tilt camera; The people active in the target area are checked in and authenticated based on the close-up image.

2. The non-sensing attendance method according to claim 1, characterized in that: The first target object in the first panoramic image is obtained by tracking the second target object in the historical panoramic image, the second target object and the correspondingly tracked first target object correspond to the same active person, and the historical panoramic image is obtained by photographing the activity area by the panoramic camera before photographing the first panoramic image; Accordingly, determining the headed object from the plurality of first target objects includes: Acquire posture information of a first target object and a second target object of the active person; Determining the corresponding action of the active person according to the posture information of the first target object and the second target object of the active person; In a case where the action of the activity person is raising his head, the first target object of the activity person is determined to be a raising-head object.

3. The non-sensing attendance method according to claim 1, characterized in that: The activity area includes a plurality of sub-areas; Accordingly, determining the target area corresponding to the head-up object in the activity area includes: Obtaining the location information of each head-up object; determining a sub-region where the head-up object is located according to the position information of the head-up object; According to the sub-areas where the head-up objects are located, counting the number of head-up objects in the sub-areas; A target area is determined in the plurality of sub-areas according to the number of head-up objects in the plurality of sub-areas.

4. The non-sensing attendance method according to claim 3, characterized in that: The counting of the number of head-up objects in a plurality of sub-areas according to the sub-areas where the head-up objects are located includes: Count the number of people who have successfully signed in in each sub-area; When the number of people in attendance in the sub-area does not reach the number threshold of the sub-area, the number of head-up objects in the sub-area is counted according to the head-up objects in the sub-area.

5. The non-sensing attendance method according to claim 3, characterized in that: The counting of the number of head-up objects in a plurality of sub-areas according to the sub-areas where the head-up objects are located includes: determining a head-up rate of the sub-area according to the number of head-up objects and the number of target objects in the sub-area; A target area is determined in the plurality of sub-areas according to the head-up rates of the plurality of sub-areas.

6. The non-sensing attendance method according to claim 5, characterized in that: The determining of the target area in the plurality of sub-areas according to the head-up rates of the plurality of sub-areas includes: In the case that the sub-area with the highest head-up rate is the same sub-area as the sub-area continuously photographed by the gimbal camera, if the continuous photographing time of the sub-area with the highest head-up rate reaches a preset time threshold, the sub-area with the second highest head-up rate is determined as the target area; otherwise, the sub-area with the highest head-up rate is determined as the target area.

7. The non-sensing attendance method according to claim 1, characterized in that: The acquiring of a close-up image of the target area captured by the pan-tilt camera includes: Acquire multiple frames of close-up images of the target area captured by the pan-tilt camera at a preset frequency within a first preset time period; Accordingly, the step of performing sign-in authentication on the active persons in the target area according to the close-up image includes: Detecting multiple facial images in the multiple close-up images, and dividing the multiple facial images into multiple facial image sets, each facial image set corresponding to one active person; The identity information of the corresponding activity personnel is identified based on the face image set, and the successful sign-in of the corresponding activity personnel is confirmed based on the identified identity information.

8. The non-sensing attendance method according to claim 7, characterized in that: The detecting of multiple facial images in the multiple frames of close-up images and dividing the multiple facial images into multiple facial image sets includes: Performing face detection on a first frame of close-up image captured within the first preset time period to obtain at least one face image; Face tracking is performed on the remaining close-up images within the first preset time period based on the face image, and the face image in the first frame close-up image and the corresponding tracked face images in the remaining close-up images are divided into the same face image set.

9. The non-sensing attendance method according to claim 7, characterized in that: The identifying of the identity information of the corresponding activity person based on the face image set includes: evaluating the image quality of each face image in the face image set, and selecting the face image with the highest image quality; The identity information of the corresponding active person is identified based on the facial image with the highest image quality.

10. The non-sensing attendance method according to claim 7, characterized in that: The step of controlling the pan-tilt camera to capture multiple close-up images of the target area within a first preset time period based on a preset frequency includes: If the activity carried out in the activity area is in a critical time period, controlling the pan-tilt camera to capture a close-up image of the target area based on a first preset frequency; If the activity carried out in the activity area is not in the critical time period, controlling the pan-tilt camera to capture a close-up image of the target area based on a second preset frequency; The first preset frequency is higher than the second preset frequency.

11. The non-sensing attendance method according to claim 1, characterized in that: The method further comprises: Acquire a second panoramic image obtained by photographing the activity area with a panoramic camera, where the resolution of the second panoramic image is higher than that of the first panoramic image; performing face detection on the second panoramic image to obtain at least one face image; When the image quality of the facial image is greater than or equal to a preset quality threshold, the identity information of the corresponding activity person is identified according to the facial image, and the corresponding activity person is confirmed to have successfully signed in based on the identified identity information.

12. A non-sensing attendance device, characterized in that: include: one or more processors; The memory stores one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the contactless attendance method as described in any one of claims 1 to 11.

13. A storage medium containing computer-executable instructions, characterized in that: The computer executable instructions, when executed by a computer processor, are used to execute the contactless attendance method according to any one of claims 1 to 11.