Multi-target ticket evasion detection method, device, computer equipment and storage medium for gate machine

By combining camera video processing with multiple models, the problems of high cost and low accuracy in existing fare evasion detection have been solved, and efficient and accurate fare evasion behavior identification and real-time alarm have been achieved.

CN114581663BActive Publication Date: 2025-09-30SHENZHEN SUNWIN INTELLIGENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210151259.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-09-30
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

Existing subway fare evasion detection methods rely on infrared imaging or depth cameras, which are costly and affected by the environment, and cannot accurately identify fare evasion.

Method used

A camera is used to shoot video, and human instance segmentation is performed through an image model. Key point recognition is performed in combination with a posture recognition model, and target tracking is performed with a tracking model. A sequence video of posture key points and human mask information is generated, and finally, fare evasion is detected in a behavior recognition model.

Benefits of technology

It achieves high-accuracy detection of fare evasion without the need for infrared imaging technology, saving costs, and is adaptable to situations where multiple people evade fares, with real-time alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581663B_ABST
    Figure CN114581663B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a method, device, computer equipment and storage medium for detecting multi-target fare evasion at a gate. The method includes: obtaining a video captured by a camera and processing the video to obtain a picture to be detected; inputting the picture to be detected into a picture model to perform human instance segmentation to obtain a human mask prediction map; inputting the human mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map; inputting the human mask prediction map into a tracking model to perform target tracking to obtain a tracking result; generating a sequence video with posture key points and human mask information based on the tracking result combined with the human mask prediction map and the key point prediction map; inputting the sequence video into a behavior recognition model to perform fare evasion behavior detection to obtain a detection result; when the detection result is fare evasion behavior, generating a warning message and sending the warning message to the terminal. By implementing the method of the embodiment of the present invention, real-time fare evasion behavior detection can be achieved without the need for infrared imaging technology for target positioning, saving a lot of costs and having a high recognition accuracy rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fare evasion detection method, and more specifically to a gate machine multi-target fare evasion detection method, device, computer equipment and storage medium. Background Art

[0002] Currently, rail transit, such as the subway, is becoming increasingly popular as an important and convenient mode of transportation. Generally, people need to purchase a ticket before boarding the subway and pass through a gate to check their tickets when entering the station. However, many passengers evade fares every day, for example by jumping over the gate, causing significant financial losses to the subway company.

[0003] Chinese patent CN201910010440.8 discloses a method and system for detecting subway fare evasion based on infrared thermal imaging, which includes the following steps: detecting whether a pedestrian has entered the pedestrian detection range of the gate image; performing background difference of the infrared thermal imaging image using the automatically updated background to extract the infrared thermal imaging overhead view image of the pedestrian; performing morphological processing on the extracted infrared thermal imaging overhead view image of the pedestrian to obtain a binary overhead view image of the pedestrian passing through the gate based on the automatically updated appropriate threshold; performing parallel extraction of sub-regions of interest on the binary pedestrian overhead view image, setting the ROI area to be the same as the number of gates (N), and obtaining N mutually independent binary overhead views of pedestrians in the gate channels; performing N mutually independent binary overhead views of pedestrians in the gate channels Figure 2 The connected areas of the image are marked respectively to obtain the pedestrian parameters; the fare evasion behavior at N gates is determined. The invention can effectively identify and avoid fare evasion behavior. However, this method is based on infrared thermal imaging for fare evasion detection. The cost is high and a fixed camera is required, which is very unfriendly for project deployment. It has certain limitations and is based on the overhead view of pedestrians in the binary gate channel. Figure 2Judging by using a value image is often greatly affected by other environmental factors such as lighting, and the accuracy of the effect cannot be guaranteed. Chinese patent CN201911224925.3 provides a subway passenger fare evasion behavior detection system and method, which specifically includes a passenger information marking module, which identifies and marks whether the passenger is a passenger who needs to purchase a ticket to ride, whether he or she is carrying an infant, and captures facial information, and stores the passenger's dynamic riding information; a card swiping behavior recognition module, which determines whether the passenger swipes the card based on whether there is an intersection between the human skeleton motion trajectory and the card swiping area obtained by the depth camera; a card swiping information reading module, which reads the gate card swiping information and time, and determines that the card swiping behavior is successful when the card swiping behavior is performed; a fare evasion behavior judgment and warning module, which combines the passenger's ticket purchase mark information, card swiping behavior recognition information, card swiping success record and the number of people passing through to identify fare evasion behavior and issue a warning; this method is based on shooting with a depth camera, which is costly, and the system needs to combine the card swiping information reading to touch and see to obtain fare evasion behavior, and cannot directly use image information to judge, which has certain limitations. Chinese patent CN201510144081.7 discloses a gate detection system and method, which includes: a three-dimensional image information acquisition module, including at least two image data acquisition devices for acquiring two-dimensional image information of the same area to be detected from different positions, for obtaining image information of the human body in the area to be detected; a gate status acquisition module, for obtaining gate status information of the gate; a three-dimensional image information recognition and processing module, for using the image information of the human body and the status information of the gate to determine whether the human body has evaded the ticket; an alarm module, for sounding an alarm when the human body has evaded the ticket; this method requires the use of a three-dimensional image information acquisition module, which is relatively low in cost, and needs to be based on the acquisition of the gate status. It cannot be analyzed based solely on image information, and the deployment project is relatively complex. Chinese patent CN202110192793.1 relates to a method for fare evasion at subway gates based on rapid estimation of passenger posture. Identification is performed through the following steps: first, information is collected through video surveillance of subway gates, then key points of subway passenger skeletons are detected, and finally, the fare evasion behavior of passengers passing through the Forbidden City gates is identified. However, this method only judges the fare evasion behavior of passengers based on key point information. This method cannot obtain the continuous characteristics of the passenger's fare evasion behavior. The lack of identification information affects the judgment of the passenger's fare evasion behavior. Secondly, relying solely on key point information for logical judgment is often affected by the space captured by the camera, resulting in inaccurate recognition.

[0004] Therefore, it is necessary to design a new method to achieve real-time detection of fare evasion, without the need for infrared imaging technology for target positioning, saving a lot of costs and with high recognition accuracy. Summary of the Invention

[0005] The purpose of the present invention is to overcome the defects of the prior art and provide a method, device, computer equipment and storage medium for detecting multi-target fare evasion at a gate.

[0006] To achieve the above objectives, the present invention adopts the following technical solutions: a multi-target fare evasion detection method for gate machines, comprising:

[0007] Obtaining the video captured by the camera and processing the video to obtain the image to be detected;

[0008] Inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction image;

[0009] Inputting the human body mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map;

[0010] Inputting the human body mask prediction image into a tracking model to perform target tracking to obtain a tracking result;

[0011] Generate a sequence video with posture key points and human mask information according to the tracking result combined with the human mask prediction map and the key point prediction map;

[0012] Inputting the sequence video into a behavior recognition model to detect fare evasion behavior to obtain a detection result;

[0013] When the detection result is a fare evasion behavior, a warning message is generated and sent to the terminal.

[0014] Its further technical solution is: the image model is formed by adding a mask branch on the basis of the yolov5 model.

[0015] A further technical solution is: inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction image, including:

[0016] Input the image to be detected into the image model to predict the human target feature layer by the yolov5 model;

[0017] The target features are respectively intercepted from the human body target feature layer through ROI Align and the corresponding SPP mechanism to obtain a first feature and a second feature;

[0018] Performing an upsampling operation of a dilated convolution group and a deconvolution group on the first feature, and performing a downsampling operation of a variable convolution group on the second feature to obtain two processed feature vectors;

[0019] The two processed feature vectors are resized, the original features are directly spliced ​​and merged, the CBAM attention mechanism is operated, and the size is adjusted twice to obtain the human body mask prediction map.

[0020] Its further technical solution is: the posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism on the basis of the HRNet model.

[0021] Its further technical solution is: the tracking model is obtained by training the ByteTrack model with a plurality of human body coordinate information with motion trajectory labels.

[0022] A further technical solution is: generating a sequence video with posture key points and human mask information based on the tracking results in combination with the human mask prediction map and the key point prediction map, including:

[0023] Determining a human body mask prediction map having a motion trajectory according to the tracking result, and processing the determined human body mask prediction map to obtain a processing result;

[0024] Directly splicing and merging the processing result with the key point prediction map based on original features to form a fused image;

[0025] The fused images are merged in time to generate a sequence video with posture key points and human body mask information.

[0026] Its further technical solution is: the behavior recognition model is formed by modifying the 3D convolution into a variable 3D convolution based on the MoviNet model.

[0027] The present invention also provides a gate machine multi-target fare evasion detection device, comprising:

[0028] A sampling unit is used to obtain the video captured by the camera and process the video to obtain the image to be detected;

[0029] a segmentation unit, configured to input the image to be detected into an image model to perform human instance segmentation to obtain a human mask prediction image;

[0030] A key point recognition unit, configured to input the human body mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map;

[0031] A tracking unit, configured to input the human body mask prediction image into a tracking model to perform target tracking to obtain a tracking result;

[0032] A video generation unit, configured to generate a sequence video having posture key points and human mask information according to the tracking result in combination with the human mask prediction map and the key point prediction map;

[0033] a behavior detection unit, configured to input the sequence video into a behavior recognition model to detect fare evasion behavior and obtain a detection result;

[0034] The alarm unit is used to generate a warning message when the detection result is a fare evasion behavior, and send the warning message to the terminal.

[0035] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0036] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0037] The beneficial effects of the present invention compared with the existing technology are: the present invention obtains the image taken by the camera and cuts it, so as to input the generated image to be detected into the image model for human instance segmentation, and then combines it with the posture recognition model to perform key point recognition, combines it with the tracking model to perform target tracking, and generates a sequence video with posture key points and human body mask information, detects fare evasion behavior based on the sequence video, and issues an alarm when fare evasion behavior occurs, thereby realizing real-time detection of fare evasion behavior, without the need for infrared imaging technology for target positioning, saving a lot of costs, and having a high recognition accuracy rate.

[0038] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 A schematic diagram of an application scenario of the multi-target fare evasion detection method for gates provided by an embodiment of the present invention;

[0041] Figure 2 A schematic diagram of the flow chart of a multi-target fare evasion detection method for a gate provided by an embodiment of the present invention;

[0042] Figure 3 A schematic diagram of a sub-process of a multi-target fare evasion detection method for a gate provided by an embodiment of the present invention;

[0043] Figure 4 A schematic diagram of a sub-process of a multi-target fare evasion detection method for a gate provided by an embodiment of the present invention;

[0044] Figure 5 A schematic block diagram of a multi-target fare evasion detection device for a gate provided by an embodiment of the present invention;

[0045] Figure 6 A schematic block diagram of a segmentation unit of a multi-target fare evasion detection device for a gate provided by an embodiment of the present invention;

[0046] Figure 7 A schematic block diagram of a video generation unit of a multi-target fare evasion detection device for gates provided by an embodiment of the present invention;

[0047] Figure 8 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0049] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0050] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0051] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0052] See also Figure 1 and Figure 2 , Figure 1 Schematic diagram of an application scenario of the multi-target fare evasion detection method for gates provided by an embodiment of the present invention. Figure 2The present invention provides a schematic flow chart of a method for detecting fare evasion at a gate with multiple targets, provided by an embodiment of the present invention. The method is applied to a server. The server exchanges data with a terminal and a camera, obtains a video captured by the camera, processes the video to form a picture to be detected, and uses a picture model to segment human instances, a posture recognition model to identify key points, and a tracking model to track targets. A sequence video with posture key points and human mask information is then generated, and the sequence video is input into a behavior recognition model to identify fare evasion. When fare evasion occurs, a warning message is generated and sent to the terminal for prompting. The method can identify fare evasion by passengers solely by relying on an optical camera, without requiring fixed shooting restrictions on the camera. It only needs to capture the gate location to detect multiple passengers evading the fare simultaneously. The miniature model used is very friendly to real-time effects, and can promptly alert passengers of fare evasion with high recognition accuracy.

[0053] Figure 2 FIG. 1 is a flow chart of a method for detecting multi-target fare evasion at a gate provided by an embodiment of the present invention. Figure 2 As shown, the method includes the following steps S110 to S170.

[0054] S110: Obtain a video captured by a camera, and process the video to obtain a picture to be detected.

[0055] In this embodiment, the image to be detected refers to an image captured by a camera within a set gate range.

[0056] Specifically, the lens of the optical camera is aimed at the passenger card swiping gate area, and the area is monitored based on video. The optical camera adopts a fixed focal length, fixed angle and position, and the captured video is cut to generate pictures. Here, every 2 frames are cut to generate frame pictures with a certain time sequence, that is, the pictures to be detected, and the sequence order information of the frame pictures is retained. Finally, the frame pictures will be scaled to ensure that the size of the sequence pictures is consistent and meets the input requirements of the picture model. Here, the width and height are set to 640, and the input picture to be detected is scaled to this size through padding.

[0057] It only relies on optical cameras and does not require infrared imaging technology for target positioning, saving a lot of costs.

[0058] S120 , inputting the image to be detected into an image model to perform human instance segmentation to obtain a human mask prediction image.

[0059] In this embodiment, the human body mask prediction image refers to the human body mask image of each passenger in each frame.

[0060] Specifically, the image model is formed by adding a mask branch to the Yolov5 model. The image model can support both object detection and segmentation tasks, thereby achieving real-time instance segmentation. To balance efficiency and accuracy, the Yolov5x model is used as the base model.

[0061] In one embodiment, see Figure 3 , the above-mentioned step S120 may include steps S121 to S124.

[0062] S121. Input the image to be detected into the image model to predict the human target feature layer using the yolov5 model.

[0063] In this embodiment, the human target feature layer refers to the feature layer where the human target is located.

[0064] Since the target size of passenger detection is medium to large, the second layer CSPDartnet53 and the fourth layer ResBlock_body output of the yolov5x model are selected as the input feature layers PMF1 and PMF2 for predicting human masks, respectively. These two layers of features have strong spatial and semantic information.

[0065] S122 , extracting target features from the human body target feature layer through ROI Align and the corresponding SPP mechanism to obtain a first feature and a second feature.

[0066] In this embodiment, the first feature layer refers to the output of the second layer CSPDartnet53 in the yolov5x model; the second feature refers to the output of the fourth layer ResBlock_body in the yolov5x model.

[0067] Specifically, the target feature layer predicted by yolov5x is respectively intercepted from each feature layer through ROI Align and the corresponding SPP mechanism to form the first feature and the second feature.

[0068] S123. Perform an upsampling operation of a dilated convolution group and a deconvolution group on the first feature, and perform a downsampling operation of a variable convolution group on the second feature to obtain two processed feature vectors.

[0069] In this embodiment, the two processed feature vectors refer to the feature vector formed after the first feature is subjected to the upsampling operation of the hole convolution group and the deconvolution group, and the feature vector formed after the second feature is subjected to the downsampling operation of the variable convolution group.

[0070] Specifically, the convolution group includes multiple convolution operations, activation operations, batch normalization and other operations, among which the activation operation is performed using the Mish activation function.

[0071] S124. The two processed feature vectors are resized, the original features are directly spliced ​​and merged, the CBAM attention mechanism is operated, and the size is resized twice to obtain a human body mask prediction map.

[0072] Specifically, the two feature vectors are resized to ensure that they have the same feature layer size. The resized output is then merged by directly concatenating the original features to obtain a feature PE that is used to enrich both semantic and spatial information. The feature PE is then subjected to the CBAM attention mechanism to ensure that the number of channels is equal to the target category, which is set to 1 here, i.e., the human category. Finally, the feature map is resized to restore it to the original target size and the final human mask is obtained through an activation function using a sigmoid activation function for normalization.

[0073] Image model prediction and training differ in the selection of target locations for mask prediction. In prediction, the target location is selected based on the predicted box position, while in training, the target location is selected based on the information from the annotated box. The model uses the CIOU loss function for object detection. Considering the sparse density of people in the gate area, the Fast-NMS process is used to accelerate image model reasoning. Focal loss is used for object classification to reduce the impact of class imbalance. The Dice loss function is used for semantic segmentation. During training, a focus strategy is used to adjust the model input to accelerate image model reasoning.

[0074] First, the gate card swiping area is set, and detection is performed within the area. The collected samples are used to predict the human body frame and the corresponding PMP (mask information, Person Mask Picture), that is, the human body mask prediction map.

[0075] S130 , inputting the human body mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map.

[0076] In this embodiment, the key point prediction map refers to the spatial positions and categories of key points such as left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right waists, left and right knees, and left and right ankles.

[0077] The posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism to the HRNet model.

[0078] Specifically, the posture recognition model is used to identify human poses within the gate's card swipe area. Based on keypoint recognition, the model employs a top-down approach. Specifically, the model predicts keypoints within the body frame using the position predicted by the image model. This allows for better detection of partially occluded passenger torsos and facilitates behavioral prediction based on the spatial coordinates of hidden keypoints. The HRNet model is used as the foundation for keypoint detection, the foundation of the posture recognition model. The SwinTransformer self-attention feature extraction mechanism is then incorporated. The HRNet model consists of four stages, during which information is repeatedly exchanged between multi-resolution networks. High-resolution and low-resolution features are enhanced through multi-scale fusion. This embodiment concatenates the features extracted in the first stage with the SwinTransformer, resulting in a larger receptive field for shallower layers of the network and better access to global information, accelerating model convergence. The posture recognition model uses the MSE loss function. To optimize training, the model employs the OHEM strategy, overweighting keypoints that are difficult to train and underweighting those that are easy to train, ensuring better model convergence. Different channels of the model's output represent key points of different categories. The relative position of the key point in the output represents the spatial position of the key point. During the prediction process, key points with a confidence level greater than a threshold of 0.5 are selected as identified key points, while occluded key points are retained. For key points of the same category, the one with the highest probability is selected as the key point of that category.

[0079] S140: Input the human body mask prediction image into a tracking model to perform target tracking to obtain a tracking result.

[0080] In this embodiment, the tracking result is the movement trajectory of the human body, that is, the movement trajectory of the passenger.

[0081] The tracking model is obtained by training the ByteTrack model with a number of human body coordinate information with motion trajectory labels.

[0082] Specifically, in the process of detecting fare evasion at the gate, it is easy for multiple targets (multiple passengers) to appear at the same time. In order to record the movement trajectories of different passengers and facilitate subsequent behavior recognition based on time series information, the ByteTrack model is used as the tracking model. The target coordinate information predicted by the image model at different times, that is, the human body frame, is tracked to obtain the movement trajectory of each passenger and the index information of each passenger trajectory. If the distance between the center points of two passengers is lower than the threshold, the two passengers are combined and their trajectory information is retained separately. This is to determine whether there are two people who are close to each other and evade the ticket.

[0083] S150 , generating a sequence video having posture key points and human mask information according to the tracking result in combination with the human mask prediction map and the key point prediction map.

[0084] In this embodiment, the sequence video with posture key points and human body mask information refers to a time series video formed after processing the human body mask prediction map with the human body's movement trajectory, and the time series video has relevant information about the posture key points. Behavior recognition is performed not only based on image information, but also based on the sequence information between images in the video, and the recognition accuracy is greatly enhanced.

[0085] In one embodiment, see Figure 4 , the above-mentioned step S150 may include steps S151 to S153.

[0086] S151 , determining a human body mask prediction map having a motion trajectory according to the tracking result, and processing the determined human body mask prediction map to obtain a processing result.

[0087] In this embodiment, the processing result refers to an image in which the foreground of the RGB three-channel image of the human body mask prediction image with the action trajectory is set to original pixel information and the background RGB is set to (0, 0, 0).

[0088] S152: directly merge the processing result with the key point prediction image based on original features to form a fused image.

[0089] In this embodiment, the fused image refers to an image formed by merging features of the processing result and the key point prediction image.

[0090] S153: Merge the fused images by time to generate a sequence video with posture key points and human body mask information.

[0091] Specifically, a logical judgment is performed based on the motion trajectory of each passenger or group of passengers and the position of the gate to determine the passenger's process from entering the gate to leaving. During this process, the posture key points of each passenger in each frame are recorded, and multiple corresponding KPPs (Key Point Picture) are generated. Here, the key point information is set to 1, and the non-key point information is set to 0. Different key point categories are represented by different channels. At the same time, the foreground of the RGB three-channel PMP image is set to the original pixel information, and the background RGB is set to (0, 0, 0) to preserve the spatial information of different human bodies. Finally, the KPP and PMP are concatenated to obtain a multi-channel PFP (Prediction Fusion Picture), that is, a single passenger time series PFP and a combined passenger time series PFP. Finally, these time series PFPs are merged to generate a sequence video PFP-Video with posture key points and human body mask information.

[0092] S160: Input the sequence video into a behavior recognition model to detect fare evasion behavior to obtain a detection result.

[0093] In this embodiment, the detection result refers to the recognition result of whether the current image to be detected contains fare evasion behavior.

[0094] Specifically, the behavior recognition model is formed by modifying the 3D convolution into a variable 3D convolution based on the MoviNet model.

[0095] The action recognition model uses sequential PFP-Video as input for action recognition. To make the model more robust, the input frames are randomly sampled at intervals of 1-3 frames during training. The 3D convolution of the action recognition model is modified to a deformable 3D convolution to enhance its adaptability to irregular objects, especially non-grid square targets. During training, a random DropBlock operation is performed on human masks and key points. This means that any dropped mask is set to the background. If a key point exists in the area, it is also set to the background. This allows the model to adapt to partial occlusion, which can effectively predict fare evasion. The cross-entropy loss function is used as the loss function.

[0096] S170: When the detection result is a fare evasion behavior, generate a warning message and send the warning message to the terminal.

[0097] Specifically, the system categorizes passenger gate behaviors into four categories: squatting, jumping, tailgating, and passing. For single-person trajectories, these behaviors are predicted, while for multi-person trajectories, these behaviors are predicted as tailgating and passing. If squatting, jumping, or tailgating occurs at the gate, it is considered fare evasion and a warning message is generated.

[0098] The method of this embodiment relies solely on an optical camera to identify passengers' fare-dodging behavior; there is no need for fixed shooting restrictions on the camera, and it is only necessary to shoot the gate position to detect multiple people fare-dodging at the same time; the adopted micro-model is very friendly to real-time effects, and can promptly alarm passengers' fare-dodging behavior, and generates yolov5-mask model, HRNet-SwinTransformer model and DC-MoviNet model based on the existing model, among which the yolov5-mask model is an image model, the HRNet-SwinTransformer model is a posture recognition model, and the DC-MoviNet model is a behavior recognition model, which greatly improves the accuracy of fare-dodging detection.

[0099] The above-mentioned multi-target fare evasion detection method for gate machines obtains the image captured by the camera and cuts it, so as to input the generated image to be detected into the image model for human instance segmentation, and then combines it with the posture recognition model for key point recognition, combines it with the tracking model for target tracking, and generates a sequence video with posture key points and human mask information. Ticket evasion behavior is detected based on the sequence video. When ticket evasion behavior occurs, an alarm is issued to realize real-time detection of ticket evasion behavior. No infrared imaging technology is required for target positioning, which saves a lot of costs and has a high recognition accuracy.

[0100] Figure 5 : is a schematic block diagram of a gate multi-target fare evasion detection device 300 provided by an embodiment of the present invention. Figure 5 As shown, corresponding to the above gate machine multi-target fare evasion detection method, the present invention also provides a gate machine multi-target fare evasion detection device 300. The gate machine multi-target fare evasion detection device 300 includes a unit for executing the above gate machine multi-target fare evasion detection method, and the device can be configured in a server. Specifically, please refer to Figure 5 The gate multi-target fare evasion detection device 300 includes a sampling unit 301, a segmentation unit 302, a key point recognition unit 303, a tracking unit 304, a video generation unit 305, a behavior detection unit 306 and an alarm unit 307.

[0101] The sampling unit 301 is used to obtain the video captured by the camera and process the video to obtain the picture to be detected; the segmentation unit 302 is used to input the picture to be detected into the picture model for human instance segmentation to obtain a human mask prediction map; the key point recognition unit 303 is used to input the human mask prediction map into the posture recognition model for key point recognition to obtain a key point prediction map; the tracking unit 304 is used to input the human mask prediction map into the tracking model for target tracking to obtain a tracking result; the video generation unit 305 is used to generate a sequence video with posture key points and human mask information based on the tracking result in combination with the human mask prediction map and the key point prediction map; the behavior detection unit 306 is used to input the sequence video into the behavior recognition model for ticket evasion behavior detection to obtain a detection result; the alarm unit 307 is used to generate a warning message when the detection result is ticket evasion behavior, and send the warning message to the terminal.

[0102] In one embodiment, if Figure 6 As shown, the segmentation unit 302 includes an input subunit 3021 , a truncation subunit 3022 , a sampling subunit 3023 and an adjustment subunit 3024 .

[0103] The input subunit 3021 is used to input the image to be detected into the image model so that the yolov5 model predicts the human target feature layer; the interception subunit 3022 is used to intercept the target features from the human target feature layer respectively through ROI Align and the corresponding SPP mechanism to obtain the first feature and the second feature; the sampling subunit 3023 is used to perform an upsampling operation of the hole convolution group and the deconvolution group on the first feature, and perform a downsampling operation of the variable convolution group on the second feature to obtain two processed feature vectors; the adjustment subunit 3024 is used to resize the two processed feature vectors, directly splice and merge the original features, perform the CBAM attention mechanism operation and perform secondary resizing to obtain a human mask prediction map.

[0104] In one embodiment, if Figure 7 As shown, the video generation unit 305 includes a processing subunit 3051 , a fusion subunit 3052 and a merging subunit 3053 .

[0105] The processing subunit 3051 is used to determine a human body mask prediction map with a motion trajectory based on the tracking result, and process the determined human body mask prediction map to obtain a processing result; the fusion subunit 3052 is used to directly splice and merge the original features of the processing result and the key point prediction map to form a fused image; the merging subunit 3053 is used to merge the fused images by time to generate a sequence video with posture key points and human body mask information.

[0106] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned gate machine multi-target fare evasion detection device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment, and for the convenience and brevity of description, it will not be repeated here.

[0107] The above-mentioned gate multi-target fare evasion detection device 300 can be implemented in the form of a computer program. The computer program can be used in the following ways: Figure 8 Runs on the computer device shown.

[0108] See also Figure 8 , Figure 8 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.

[0109] See Figure 8 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0110] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can cause the processor 502 to execute a gate multi-target fare evasion detection method.

[0111] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0112] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a multi-target fare evasion detection method for a gate.

[0113] The network interface 505 is used to communicate with other devices through the network. Figure 8The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0114] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:

[0115] Acquire the video captured by the camera and process the video to obtain a picture to be detected; input the picture to be detected into the picture model for human instance segmentation to obtain a human mask prediction map; input the human mask prediction map into the posture recognition model for key point recognition to obtain a key point prediction map; input the human mask prediction map into the tracking model for target tracking to obtain a tracking result; generate a sequence video with posture key points and human mask information based on the tracking result and the human mask prediction map and the key point prediction map; input the sequence video into the behavior recognition model for ticket evasion behavior detection to obtain a detection result; when the detection result is ticket evasion behavior, generate a warning message and send the warning message to the terminal.

[0116] The image model is formed by adding a mask branch to the yolov5 model.

[0117] The posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism to the HRNet model.

[0118] The tracking model is obtained by training the ByteTrack model with a number of human body coordinate information with motion trajectory labels.

[0119] The behavior recognition model is formed by modifying the 3D convolution into a variable 3D convolution based on the MoviNet model.

[0120] In one embodiment, when the processor 502 implements the step of inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction image, the processor 502 specifically implements the following steps:

[0121] The image to be detected is input into the image model so that the human target feature layer is predicted by the yolov5 model; the target features are respectively extracted from the human target feature layer through ROI Align and the corresponding SPP mechanism to obtain a first feature and a second feature; the first feature is up-sampled by a hole convolution group and a deconvolution group, and the second feature is down-sampled by a variable convolution group to obtain two processed feature vectors; the two processed feature vectors are resized, the original features are directly spliced ​​and merged, the CBAM attention mechanism is operated, and the size is adjusted twice to obtain a human mask prediction map.

[0122] In one embodiment, when the processor 502 implements the step of generating a sequence video having posture key points and human mask information according to the tracking results combined with the human mask prediction map and the key point prediction map, the processor 502 specifically implements the following steps:

[0123] A human body mask prediction map with a motion trajectory is determined based on the tracking results, and the determined human body mask prediction map is processed to obtain a processing result; the processing result is directly spliced ​​and merged with the key point prediction map by original features to form a fused image; the fused image is merged by time to generate a sequence video with posture key points and human body mask information.

[0124] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0125] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0126] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0127] Acquire the video captured by the camera and process the video to obtain a picture to be detected; input the picture to be detected into the picture model for human instance segmentation to obtain a human mask prediction map; input the human mask prediction map into the posture recognition model for key point recognition to obtain a key point prediction map; input the human mask prediction map into the tracking model for target tracking to obtain a tracking result; generate a sequence video with posture key points and human mask information based on the tracking result and the human mask prediction map and the key point prediction map; input the sequence video into the behavior recognition model for ticket evasion behavior detection to obtain a detection result; when the detection result is ticket evasion behavior, generate a warning message and send the warning message to the terminal.

[0128] The image model is formed by adding a mask branch to the yolov5 model.

[0129] The posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism to the HRNet model.

[0130] The tracking model is obtained by training the ByteTrack model with a number of human body coordinate information with motion trajectory labels.

[0131] The behavior recognition model is formed by modifying the 3D convolution into a variable 3D convolution based on the MoviNet model.

[0132] In one embodiment, when the processor executes the computer program to implement the step of inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction map, the processor specifically implements the following steps:

[0133] The image to be detected is input into the image model so that the human target feature layer is predicted by the yolov5 model; the target features are respectively extracted from the human target feature layer through ROI Align and the corresponding SPP mechanism to obtain a first feature and a second feature; the first feature is up-sampled by a hole convolution group and a deconvolution group, and the second feature is down-sampled by a variable convolution group to obtain two processed feature vectors; the two processed feature vectors are resized, the original features are directly spliced ​​and merged, the CBAM attention mechanism is operated, and the size is adjusted twice to obtain a human mask prediction map.

[0134] In one embodiment, when the processor executes the computer program to implement the step of generating a sequence video having posture key points and human mask information based on the tracking results combined with the human mask prediction map and the key point prediction map, the processor specifically implements the following steps:

[0135] A human body mask prediction map with a motion trajectory is determined based on the tracking results, and the determined human body mask prediction map is processed to obtain a processing result; the processing result is directly spliced ​​and merged with the key point prediction map by original features to form a fused image; the fused image is merged by time to generate a sequence video with posture key points and human body mask information.

[0136] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0137] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0138] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0139] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0140] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A multi-target fare evasion detection method for gate machines, characterized in that: include: Obtaining the video captured by the camera and processing the video to obtain the image to be detected; Inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction image; Inputting the human body mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map; Inputting the human body mask prediction image into a tracking model to perform target tracking to obtain a tracking result; Generate a sequence video with posture key points and human mask information according to the tracking result combined with the human mask prediction map and the key point prediction map; Inputting the sequence video into a behavior recognition model to detect fare evasion behavior to obtain a detection result; When the detection result is a fare evasion behavior, a warning message is generated and sent to the terminal; The image model is formed by adding a mask branch to the yolov5 model; The posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism to the HRNet model.

2. The gate machine multi-target fare evasion detection method according to claim 1 is characterized in that: The step of inputting the image to be detected into the image model to perform human instance segmentation to obtain a human mask prediction image includes: Input the image to be detected into the image model to predict the human target feature layer by the yolov5 model; The target features are respectively intercepted from the human body target feature layer through ROI Align and the corresponding SPP mechanism to obtain a first feature and a second feature; Performing an upsampling operation of a dilated convolution group and a deconvolution group on the first feature, and performing a downsampling operation of a variable convolution group on the second feature to obtain two processed feature vectors; The two processed feature vectors are resized, the original features are directly spliced ​​and merged, the CBAM attention mechanism is operated, and the size is adjusted twice to obtain the human body mask prediction map.

3. The gate machine multi-target fare evasion detection method according to claim 1 is characterized in that: The tracking model is obtained by training the ByteTrack model with a number of human body coordinate information with motion trajectory labels.

4. The gate machine multi-target fare evasion detection method according to claim 1, characterized in that: The step of generating a sequence video having posture key points and human mask information according to the tracking result in combination with the human mask prediction map and the key point prediction map includes: Determining a human body mask prediction map having a motion trajectory according to the tracking result, and processing the determined human body mask prediction map to obtain a processing result; Directly splicing and merging the processing result with the key point prediction map based on original features to form a fused image; The fused images are merged in time to generate a sequence video with posture key points and human body mask information.

5. The gate machine multi-target fare evasion detection method according to claim 1 is characterized in that: The behavior recognition model is formed by modifying the 3D convolution into a variable 3D convolution based on the MoviNet model.

6. The multi-target ticket evasion detection device for gates is characterized by: include: A sampling unit is used to obtain the video captured by the camera and process the video to obtain the image to be detected; a segmentation unit, configured to input the image to be detected into an image model to perform human instance segmentation to obtain a human mask prediction image; A key point recognition unit, configured to input the human body mask prediction map into a posture recognition model to perform key point recognition to obtain a key point prediction map; A tracking unit, configured to input the human body mask prediction image into a tracking model to perform target tracking to obtain a tracking result; A video generation unit, configured to generate a sequence video having posture key points and human mask information according to the tracking result in combination with the human mask prediction map and the key point prediction map; a behavior detection unit, configured to input the sequence video into a behavior recognition model to detect fare evasion behavior and obtain a detection result; an alarm unit, configured to generate a warning message when the detection result is a fare evasion behavior, and send the warning message to the terminal; The image model is formed by adding a mask branch to the yolov5 model; The posture recognition model is formed by adding the SwinTransformer self-attention feature extraction mechanism to the HRNet model.

7. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 5 when executing the computer program.

8. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • System and method for detecting fare evasion at turnstiles

    CN104805784B

  • Subway fare evasion behavior detection method and system based on infrared thermal imaging

    CN110378179A

  • Methods and systems for detecting fare evasion by subway passengers

    CN111064925B

  • Subway turnstile passing fare evasion identification method based on passenger posture rapid estimation

    CN113014870A

  • Table tennis action recognition method and system based on posture segmentation and key point features

    CN110472554A