Gate channel fraud prevention processing method, device, medium and equipment

By using image sequence detection and tracking technology, the head and body features of pedestrians in the turnstile channel are identified, and motion trajectories are generated, which solves the problem of deception in the turnstile channel and improves the efficiency and accuracy of identification.

CN116704602BActive Publication Date: 2026-05-05RECONOVA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RECONOVA TECH CO LTD
Filing Date
2023-06-01
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, the authentication methods for turnstile channels are inefficient and easily deceived, leading to significant challenges in security management.

Method used

By acquiring image sequences for target detection, identifying pedestrian head features and human body detection boxes, and combining timestamp information for target tracking, a motion trajectory is generated, and impersonation to pass through the gate is identified and an alarm is triggered.

Benefits of technology

It improves the accuracy and efficiency of identifying impersonation at gates, automatically identifies deceptive behavior, and ensures the effectiveness of security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704602B_ABST
    Figure CN116704602B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, medium, and device for preventing deception at turnstile channels. The method includes: acquiring an image sequence captured from a target area, and determining the head feature information and corresponding head detection boxes of pedestrians in each image to be identified within the image sequence; predicting and determining the human body detection boxes of pedestrians based on the head feature information and head detection boxes, and extracting corresponding human appearance features; performing target tracking based on the timestamp information of the images to be identified and the head detection boxes, human body detection boxes, and human appearance features of each pedestrian in the images to be identified, generating motion trajectories corresponding to each pedestrian; and triggering a fraudulent entry alarm if, after the turnstile completes one opening / closing state, a pedestrian requesting facial recognition to pass through the turnstile fails to pass through while other pedestrians have passed through. The technical solution of this application ensures the accuracy of fraudulent entry behavior identification and improves identification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and more specifically, to a method, apparatus, medium, and equipment for preventing deception at turnstile channels. Background Technology

[0002] In modern society, to improve traffic flow efficiency, public places such as airports and subways use turnstiles for ticket checking. However, in practice, some pedestrians frequently use impersonation and other means to deceive ticket checkers and gate inspectors to evade supervision, posing significant challenges to security management. Current technical solutions often involve checking ID cards or swiping cards for identity verification, but these methods are inefficient and costly. Summary of the Invention

[0003] The embodiments of this application provide a method, apparatus, medium, and equipment for preventing fraud at turnstile channels, which can at least to a certain extent ensure the accuracy of identifying impersonation and improve identification efficiency.

[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0005] According to one aspect of the embodiments of this application, a method for preventing spoofing at turnstile channels is provided, the method comprising:

[0006] Acquire an image sequence obtained by capturing images of a target area, the image sequence comprising several consecutive images to be identified;

[0007] Target detection is performed sequentially on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes;

[0008] Based on the head feature information and head detection box of each pedestrian, a prediction is made to determine the human body detection box of the pedestrian, and the human appearance features corresponding to the position information of the human body detection box are extracted.

[0009] Based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, target tracking is performed to generate a motion trajectory corresponding to each pedestrian.

[0010] Based on the movement trajectory of each pedestrian, if a pedestrian who requested to pass through the gate by facial recognition fails to pass through the gate after the gate completes one opening and closing state, and other pedestrians have passed through the gate, an alarm for impersonation will be triggered.

[0011] According to one aspect of the embodiments of this application, a processing device for preventing spoofing at turnstile channels is provided, the device comprising:

[0012] An image acquisition module is used to acquire a sequence of images captured on a target area, the image sequence including several consecutive images to be identified;

[0013] The target detection module is used to sequentially perform target detection on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes.

[0014] The target detection module is also used to predict based on the head feature information of each pedestrian and the head detection box, determine the human body detection box of the pedestrian, and extract the human appearance features corresponding to the position information of the human body detection box.

[0015] The target tracking module is used to perform target tracking based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, and to generate a motion trajectory corresponding to each pedestrian.

[0016] The processing module is used to trigger an impersonation alarm if, after the gate completes one opening and closing state, the pedestrian who requested to pass through the gate by facial recognition fails to pass through and other pedestrians have passed through the gate, based on the movement trajectory of each pedestrian.

[0017] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the gate channel anti-fraud processing method as described in the above embodiments.

[0018] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the gate channel anti-fraud processing method as described in the above embodiments.

[0019] According to one aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the gate channel anti-spoofing processing method provided in the above embodiments.

[0020] In some embodiments of this application, the technical solutions involve acquiring an image sequence captured on a target area. This image sequence includes several consecutive images to be identified. Target detection is performed on each image in sequence to determine the head feature information and corresponding head detection boxes of pedestrians contained in the images. Based on the head feature information and head detection boxes of each pedestrian, prediction is made to determine the human body detection box of the pedestrian. Human appearance features corresponding to the position information of the human body detection box are extracted. Target tracking is performed based on the timestamp information of the images to be identified and the head detection boxes, human body detection boxes, and human appearance features of each pedestrian contained in the images to be identified. Motion trajectories corresponding to each pedestrian are generated. Based on the motion trajectories of each pedestrian, if a pedestrian requesting to pass through the gate fails to pass through after the gate completes one opening and closing state, and other pedestrians have passed through, an impersonation alarm is triggered. Therefore, based on the pedestrian's head feature information and head detection box, corresponding human body prediction and human body feature extraction are performed, which facilitates the binding of the pedestrian's head and body, thereby improving the accuracy of subsequent target tracking. This ensures the accuracy of identifying impersonation at the gate and can automatically identify impersonation at the gate, thus improving the identification efficiency.

[0021] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0023] Figure 1 A flowchart illustrating a method for preventing spoofing at turnstile channels according to an embodiment of this application is shown.

[0024] Figure 2 A block diagram of a gate channel anti-fraud processing device according to an embodiment of this application is shown.

[0025] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.

[0027] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0029] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0030] Figure 1 A flowchart illustrating a method for preventing spoofing at turnstile channels according to an embodiment of this application is shown. This method can be applied to terminal devices or servers. The terminal device may include, but is not limited to, one or more of smartphones, tablets, laptops, and desktop computers. The server may be a physical server or a cloud server; this application does not impose any special limitations on this.

[0031] Please refer to Figure 1 The method for preventing spoofing at the turnstile channel includes at least steps S110 to S150. The following description uses the application of this method to a terminal device as an example.

[0032] In step S110, an image sequence is obtained by capturing images of the target area, the image sequence including several consecutive images to be identified.

[0033] In one embodiment, an image acquisition device (such as a camera) may be installed above the turnstile. This image acquisition device can communicate with a terminal device. Those skilled in the art can pre-adjust the orientation of the image acquisition device so that it can acquire image information of a target area near the turnstile and transmit it to the terminal device for subsequent use. It should be noted that the target area may include the area before and after the turnstile passage to determine whether a pedestrian has passed through the turnstile.

[0034] The terminal device can receive in real time an image sequence captured by an image acquisition device targeting a target area. This image sequence may include several consecutive images to be identified.

[0035] In step S120, target detection is performed sequentially on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes.

[0036] In this embodiment, the terminal device can sequentially perform target detection on each image to be identified, thereby identifying the head feature information of pedestrians contained in the images and determining the corresponding head detection boxes. Specifically, the terminal device can call a pre-trained target detection model to sequentially perform target detection on each image to be identified, thereby identifying the head feature information of pedestrians contained in the images and the corresponding head detection boxes. In one example, the head feature information can be represented in the form of a vector. It should be noted that the target detection model can be trained based on existing target detection algorithms, and there are no special limitations on this.

[0037] It should be noted that after determining the head detection box corresponding to each pedestrian, an ID can be assigned to each pedestrian. It should be understood that the ID of the same pedestrian is the same in different images to be identified.

[0038] In one example, the YOLOv5 model can be used for object detection. During model training, in order to improve the diversity of training data, data augmentation can be performed on the training data, such as randomly changing the brightness, contrast, random rotation, flipping, etc.

[0039] Specifically, when performing object detection, the YOLOv5 model uses the CSPDarket53 backbone network, and the final output layer follows the YOLOv3 design with a total of 255 = (4 + 1 + 80) * 3 channels, where 4 represents the bounding box regression value, 1 represents the probability value of whether the object exists, 80 represents the number of categories that the model can predict, and 3 represents the preset number of anchor boxes. For the technical solution of this application embodiment, the number of channels in its output layer can be adjusted to 24 = (4 + 1 + 2) * 3, that is, the adjusted model only predicts two categories, namely the head of a pedestrian and the human body (explained later).

[0040] In step S130, prediction is performed based on the head feature information of each pedestrian and the head detection box to determine the human body detection box of the pedestrian, and human appearance features corresponding to the position information of the human body detection box are extracted.

[0041] In this embodiment, the target detection model can predict the human body detection box corresponding to each pedestrian based on the head feature information and head detection box corresponding to each pedestrian in the image to be identified.

[0042] In one example, during model training, based on the aforementioned YOLOv5 model, the formula used to predict head detection boxes is as follows:

[0043] b x =2σ(t) x )-0.5+c x

[0044] b y =2σ(t) y )-0.5+c y

[0045] b w =p w (2σ(t w )) 2 ,

[0046] b h =p h (2σ(t h )) 2

[0047] Among them, t x t y t w t h These are the encoded values ​​of the center point and width / height of the predicted head detection box, respectively, while c x c y These are the coordinates of the predicted grid here, p w p hThese are the width and height of the preset anchor frame, b x b y b w b h These are the center and width / height of the head detection box, respectively.

[0048] Finally, the predicted head detection boxes and the labeled ground truth boxes are used to calculate the DIOU loss function based on the following formula:

[0049]

[0050] L DIOU =1-DIOU,

[0051] Where IOU represents the intersection-union ratio, ρ represents the Euclidean distance between the centers of the ground truth bounding box and the predicted bounding box, c represents the diagonal length of the minimum bounding rectangle of the ground truth bounding box and the predicted bounding box, and L... DIOU This represents the loss function.

[0052] Then, the gradient values ​​are backpropagated to update and optimize the model.

[0053] For human body prediction, an anchor-free training method can be used, whereby the object detection model can directly predict the distance between the center point of the head detection box and the human body annotation box, as shown in the following formula:

[0054]

[0055]

[0056]

[0057]

[0058] Among them, l * t * r * b * represents the distance between the center point of the head detection box predicted by the object detection model and the left, top, right, and bottom boundaries of the human detection box, respectively, while b x b y This represents the coordinates of the center point of the head detection box. These represent the left, right, up, and down coordinates of the human body annotation box, respectively.

[0059] Then, during prediction, the human bounding box at this location is decoded using the following formula:

[0060] b x1 =b x -l *

[0061] b y1 =b y -t *

[0062] b x2 =b x +r *

[0063] b y2 =b y +b *

[0064] Finally, the loss output by the object detection model can also be based on the aforementioned L. DIOU The calculations are performed using the formula, which will not be elaborated here.

[0065] In practical use, once the human body detection box corresponding to a pedestrian is determined, the target detection model can extract the human appearance features corresponding to the position information of the human body detection box.

[0066] In one example, extracting human appearance features corresponding to the position information of the human detection box includes:

[0067] The center point of the human body detection frame is determined based on the position information of the human body detection frame;

[0068] Feature extraction is performed based on the center point of the human body detection box to determine the corresponding human appearance features.

[0069] In this embodiment, the position information of the human body detection box can include the coordinates of the left, right, up, and down of the human body detection box. The target detection model can determine the center point of the human body detection box based on the coordinates of the four corner points of the human body detection box, and then perform feature extraction based on the center point of the human body detection box to determine the human appearance features corresponding to the human body. In one example, in order to reduce the difficulty of subsequent binding between the head and the human body, the features of the head are used to predict the human body, and the center point of the human body detection box is used to learn the human appearance features. Thus, the final output channel of the model is 225 = (4 + 4 + 1 + 2 + 64) * 3, that is, 4 channels predict the head detection box, 4 channels predict the human body detection box, 1 channel predicts whether the target exists, 2 channels predict the category, and 64 channels predict the human appearance features corresponding to each human body.

[0070] In one embodiment, feature extraction is performed based on the center point of the human body detection box to determine the corresponding human appearance features, including:

[0071] Based on the center point of the human detection box, the system learns through a multi-class cross-entropy loss function during training, and extracts features directly through the network during deployment to determine the corresponding human appearance features.

[0072] In this embodiment, in the classification problem of multi-class tasks, in order to reduce the gap between the model prediction and the label, that is, the smaller the KL divergence, the better. Therefore, only the cross-entropy needs to be focused on during the optimization process. Thus, during training, optimization learning is performed based on the multi-class cross-entropy loss function. In actual use, feature extraction can be performed directly through the network, thereby improving the accuracy of the determined human appearance features and ensuring the subsequent recognition effect.

[0073] Please continue to refer to this. Figure 1 In step S140, target tracking is performed based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, and a motion trajectory corresponding to each pedestrian is generated.

[0074] In this embodiment, after determining the human appearance features, the terminal device can invoke the target tracking module to track pedestrians contained in the image to be identified. This target tracking module can be used to determine the relationships between pedestrians contained in different images to be identified, thereby tracking each pedestrian. It should be noted that the target tracking module can be based on existing target tracking algorithms, and this application does not impose any special limitations on it.

[0075] In one example, target tracking is performed based on the timestamp information of the image to be identified and the head detection boxes, body detection boxes, and human appearance features of each pedestrian contained in the image to be identified, generating motion trajectories corresponding to each pedestrian, including:

[0076] Based on the timestamp information of the image to be identified and the head detection box, human body detection box, and human appearance features of each pedestrian contained in the image to be identified, determine the head motion correlation matrix and human appearance correlation matrix between the current image to be identified and the previous image to be identified.

[0077] The corresponding cost matrix is ​​determined by calculating based on the head motion correlation matrix and the human appearance correlation matrix.

[0078] Based on the cost matrix, the Hungarian algorithm is used for matching, and Kalman filtering is performed based on the head detection boxes of the matching pedestrians in the two images to be identified to update the corresponding motion trajectory of the pedestrian.

[0079] In this embodiment, the target tracking module can determine adjacent images to be identified based on the timestamp information of the image to be identified. Thus, based on the head detection boxes, human body detection boxes, and human appearance features of each pedestrian contained in the image to be identified, the head motion correlation matrix and human appearance matrix between the current image to be identified and the previous image to be identified can be determined.

[0080] In other words, this target tracking model incorporates human appearance features f associated with the pedestrian's head. i Therefore, for two adjacent images to be identified, the head motion correlation matrix A between them can be calculated. m Correlation matrix A with human appearance e Among them, the head motion correlation matrix A m The human appearance correlation matrix A can be calculated using the Mahalanobis distance between the head features of the current frame and the previous frame. e It can be calculated by measuring the cosine similarity between the human appearance features of the current frame and the previous frame.

[0081] When the head motion correlation matrix A is determined m Correlation matrix A with human appearance e Then, the corresponding cost matrix can be calculated using the following formula:

[0082] C=λA e +(1-λ)A m ;

[0083] Here, λ takes a value between 0 and 1 to determine the emphasis on each feature; in one example, λ can be 0.5. After determining the corresponding cost matrix, the Hungarian algorithm is used for matching to identify related pedestrians in two adjacent images to be recognized, that is, matching the closest pedestrians in two images to determine that they are the same pedestrian. Furthermore, Kalman filtering is performed based on the head detection bounding boxes of the matching pedestrians in two adjacent images to update the corresponding motion trajectory of the pedestrian, ensuring the accuracy of the motion trajectory.

[0084] In one embodiment, after determining the movement trajectory of each pedestrian, a tracking record table containing the movement trajectories of all pedestrians can be generated. This tracking record table may include information such as the time of each pedestrian's first facial recognition scan and whether they passed through the gate.

[0085] In one embodiment, after performing Kalman filtering based on the matching head detection bounding boxes of a pedestrian in two images to be identified to update the motion trajectory corresponding to the pedestrian, the method further includes:

[0086] Based on the human appearance features of the same pedestrian in the previous image to be identified, the human appearance features of the pedestrian in the current image to be identified are updated.

[0087] In this embodiment, after determining the movement trajectory of the same pedestrian in two adjacent images to be identified, the terminal device can update the human appearance features corresponding to that pedestrian in the current image to be identified based on the human appearance features of the same pedestrian in the previous image to be identified. It should be understood that as a pedestrian moves, their human appearance features may change due to factors such as lighting or angle. Therefore, updating the human appearance features of the same pedestrian in the current image to be identified based on the human appearance features in the previous image to be identified ensures the accuracy of the human appearance features, thereby ensuring the accuracy of subsequent target tracking and preventing situations where the human appearance features of the same pedestrian change too much due to excessive time intervals, making matching and tracking impossible.

[0088] In one example, human appearance features are updated according to the following formula:

[0089]

[0090] Where α is momentum, typically taken as 0.9, e i f represents the physical appearance characteristics of the i-th pedestrian. i t This represents the human appearance features corresponding to the human detection box of the i-th pedestrian matched in the current image to be identified. Therefore, a moving average method is used to update the features of the same pedestrian in each frame, thus eliminating the need to save the human appearance features of a pedestrian in each frame and reducing data storage.

[0091] In step S150, based on the movement trajectory of each pedestrian, if a pedestrian who requested to pass through the gate by facial recognition fails to pass through the gate after the gate completes one opening and closing state, and other pedestrians have passed through the gate, an alarm for impersonation is triggered.

[0092] In this embodiment, based on the determined movement trajectory of each pedestrian, after the gate completes one opening and closing state, the impersonation detection strategy can be activated. At this time, the ID of the pedestrian who made the face scan request to pass through the gate can be determined first. Then, based on the determined movement trajectory of the pedestrian, it can be determined whether the pedestrian has passed through the gate. If the pedestrian has not passed through the gate, but pedestrians with other IDs have passed through the gate, it indicates that there is impersonation, thereby triggering an impersonation alarm. For example, an alarm can be triggered to the surrounding staff, or an alarm message can be generated and played on the screen or broadcast, etc.

[0093] The following describes an embodiment of the apparatus described in this application, which can be used to execute the gate access anti-spoofing processing method in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the gate access anti-spoofing processing method described above in this application.

[0094] Figure 2A block diagram of a gate channel anti-fraud processing device according to an embodiment of this application is shown.

[0095] Reference Figure 2 As shown, a gate access anti-spoofing processing device according to an embodiment of this application includes:

[0096] Image acquisition module 210 is used to acquire an image sequence obtained by capturing images of a target area, the image sequence including a number of consecutive images to be identified;

[0097] The target detection module 220 is used to sequentially perform target detection on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes.

[0098] The target detection module 220 is further configured to predict based on the head feature information of each pedestrian and the head detection box, determine the human body detection box of the pedestrian, and extract the human appearance features corresponding to the position information of the human body detection box.

[0099] The target tracking module 230 is used to perform target tracking based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, and generate a motion trajectory corresponding to each pedestrian.

[0100] The processing module 240 is used to trigger an impersonation alarm if, after the gate completes one opening and closing state, the pedestrian who made the face scan request to pass through the gate fails to pass through the gate and other pedestrians have passed through the gate, based on the movement trajectory of each pedestrian.

[0101] In one embodiment of this application, the target tracking module 230 is configured to: determine the head motion correlation matrix and the human appearance correlation matrix between the current image to be identified and the previous image to be identified based on the timestamp information of the image to be identified and the head detection boxes, human detection boxes, and human appearance features of each pedestrian contained in the image to be identified; calculate and determine the corresponding cost matrix based on the head motion correlation matrix and the human appearance correlation matrix; perform matching using the Hungarian algorithm based on the cost matrix, and perform Kalman filtering based on the matching head detection boxes of pedestrians in the two images to be identified to update the motion trajectory corresponding to the pedestrian.

[0102] In one embodiment of this application, the target tracking module 230 is further configured to: update the human appearance features corresponding to the same pedestrian in the current image to be identified based on the human appearance features of the same pedestrian in the previous image to be identified.

[0103] In one embodiment of this application, human appearance features are updated according to the following formula:

[0104]

[0105] Where α is momentum, e i f represents the physical appearance characteristics of the i-th pedestrian. i t This represents the human appearance features corresponding to the human detection box of the i-th pedestrian in the current image to be identified.

[0106] In one embodiment of this application, the target detection module 220 is used to: determine the center point of the human body detection box based on the position information of the human body detection box; and perform feature extraction based on the center point of the human body detection box to determine the corresponding human appearance features.

[0107] In one embodiment of this application, the target detection module 220 is used to: learn through a multi-class cross-entropy loss function during training based on the center point of the human body detection box, and directly extract features through the network during deployment to determine the corresponding human appearance features.

[0108] Figure 3 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.

[0109] It should be noted that, Figure 3 The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0110] like Figure 3 As shown, the computer system includes a Central Processing Unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 302 or programs loaded from storage portion 308 into Random Access Memory (RAM) 303, such as performing the methods described in the above embodiments. The RAM 303 also stores various programs and data required for system operation. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An Input / Output (I / O) interface 305 is also connected to the bus 304.

[0111] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. Removable media 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 310 as needed so that computer programs read from them can be installed into storage section 308 as needed.

[0112] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs various functions defined in the system of this application.

[0113] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0115] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0116] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0117] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0118] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0119] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0120] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for preventing deception at turnstile channels, characterized in that, include: Acquire an image sequence obtained by capturing images of a target area, the image sequence comprising several consecutive images to be identified; Target detection is performed sequentially on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes; Based on the head feature information and head detection box of each pedestrian, a prediction is made to determine the human body detection box of the pedestrian, and the human appearance features corresponding to the position information of the human body detection box are extracted. Based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, target tracking is performed to generate a motion trajectory corresponding to each pedestrian. Based on the movement trajectory of each pedestrian, if a pedestrian who requested to pass through the gate by facial recognition fails to pass through the gate after the gate completes one opening and closing state, and other pedestrians have passed through the gate, an alarm for impersonation will be triggered. The human body prediction uses an AnchorFree training method, meaning the object detection model can directly predict the distance between the center point of the head detection box and the human body annotation box, as shown in the following formula: in, These represent the distances between the center point of the head detection box predicted by the object detection model and the left, top, right, and bottom boundaries of the human detection box, respectively. This represents the coordinates of the center point of the head detection box; These represent the left, right, up, and down coordinates of the human body annotation box, respectively. Then, during prediction, the human bounding box at this location is decoded using the following formula: 。 2. The method according to claim 1, characterized in that, Based on the timestamp information of the image to be identified, as well as the head detection boxes, body detection boxes, and human appearance features of each pedestrian contained in the image to be identified, target tracking is performed to generate motion trajectories corresponding to each pedestrian, including: Based on the timestamp information of the image to be identified and the head detection box, human body detection box, and human appearance features of each pedestrian contained in the image to be identified, determine the head motion correlation matrix and human appearance correlation matrix between the current image to be identified and the previous image to be identified. The corresponding cost matrix is ​​determined by calculating based on the head motion correlation matrix and the human appearance correlation matrix. Based on the cost matrix, the Hungarian algorithm is used for matching, and Kalman filtering is performed based on the head detection boxes of the matching pedestrians in the two images to be identified to update the corresponding motion trajectory of the pedestrian.

3. The method according to claim 2, characterized in that, After performing Kalman filtering on the matching head detection bounding boxes of pedestrians in two images to be identified to update the corresponding motion trajectory of the pedestrian, the method further includes: Based on the human appearance features of the same pedestrian in the previous image to be identified, the human appearance features of the pedestrian in the current image to be identified are updated.

4. The method according to claim 3, characterized in that, Update human appearance features according to the following formula: , in, For momentum, This represents the physical appearance characteristics of the i-th pedestrian. This represents the human appearance features corresponding to the human detection box of the i-th pedestrian in the current image to be identified.

5. The method according to claim 1, characterized in that, Extracting human appearance features corresponding to the position information of the human detection box includes: The center point of the human body detection frame is determined based on the position information of the human body detection frame; Feature extraction is performed based on the center point of the human body detection box to determine the corresponding human appearance features.

6. The method according to claim 5, characterized in that, Feature extraction is performed based on the center point of the human body detection box to determine the corresponding human appearance features, including: Based on the center point of the human detection box, the system learns through a multi-class cross-entropy loss function during training, and extracts features directly through the network during deployment to determine the corresponding human appearance features.

7. A device for preventing deception at turnstile channels, characterized in that, include: An image acquisition module is used to acquire a sequence of images captured on a target area, the image sequence including several consecutive images to be identified; The target detection module is used to sequentially perform target detection on each of the images to be identified in order to determine the head feature information of pedestrians contained in the images to be identified and the corresponding head detection boxes. The target detection module is also used to predict based on the head feature information of each pedestrian and the head detection box, determine the human body detection box of the pedestrian, and extract the human appearance features corresponding to the position information of the human body detection box. The target tracking module is used to perform target tracking based on the timestamp information of the image to be identified and the head detection box, human body detection box and human appearance features of each pedestrian contained in the image to be identified, and to generate a motion trajectory corresponding to each pedestrian. The processing module is used to trigger an impersonation alarm if, after the gate completes one opening and closing state, the pedestrian who made the face scan request to pass through the gate fails to pass through the gate and other pedestrians have passed through the gate, based on the movement trajectory of each pedestrian. The human body prediction uses an AnchorFree training method, meaning the object detection model can directly predict the distance between the center point of the head detection box and the human body annotation box, as shown in the following formula: in, These represent the distances between the center point of the head detection box predicted by the object detection model and the left, top, right, and bottom boundaries of the human detection box, respectively. This represents the coordinates of the center point of the head detection box; These represent the left, right, up, and down coordinates of the human body annotation box, respectively. Then, during prediction, the human bounding box at this location is decoded using the following formula: 。 8. The apparatus according to claim 7, characterized in that, The target tracking module is used for: Based on the timestamp information of the image to be identified and the head detection box, human body detection box, and human appearance features of each pedestrian contained in the image to be identified, determine the head motion correlation matrix and human appearance correlation matrix between the current image to be identified and the previous image to be identified. The corresponding cost matrix is ​​determined by calculating based on the head motion correlation matrix and the human appearance correlation matrix. Based on the cost matrix, the Hungarian algorithm is used for matching, and Kalman filtering is performed based on the head detection boxes of the matching pedestrians in the two images to be identified to update the corresponding motion trajectory of the pedestrian.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the gate channel anti-fraud processing method as described in any one of claims 1 to 6.

10. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the gate channel anti-fraud processing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Swimming pool drowning prevention supervision method and device, computer equipment and storage medium

    CN114022910A

  • Swimming pool drowning prevention human body automatic tracking method and device, computer equipment and storage medium

    CN114359966A