Store customer flow statistics method, device, computer equipment and storage medium

By performing upper body detection and feature vector analysis on the target image frame, the low accuracy problem of existing passenger flow counting methods is solved, and more efficient passenger flow counting and tracking is achieved.

CN114092956BActive Publication Date: 2025-09-09SF TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010744346.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-29
Publication Date
2025-09-09
Estimated Expiration
2040-07-29

AI Technical Summary

Technical Problem

Existing store customer flow counting methods have the problem of low accuracy, especially when the customer flow is large, they are prone to miscounting or missing. In addition, methods based on infrared and facial recognition are prone to serious misjudgments and omissions in complex scenarios.

Method used

By detecting the upper body of the person in the target image frame, obtaining the upper body detection frame and person category, determining the feature vector, and tracking and recording based on this information, finally performing passenger flow statistics.

Benefits of technology

It improves the accuracy of passenger flow statistics, avoids inaccurate detection caused by lower body occlusion, and achieves more efficient passenger flow tracking and statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092956B_ABST
    Figure CN114092956B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, computer device, and storage medium for counting customer flow in a store. The method comprises: acquiring a target image frame; performing upper body detection on the target image frame to obtain an upper body detection frame and person category corresponding to each target person in the target image frame; determining a feature vector of the corresponding target person based on the upper body image corresponding to each upper body detection frame in the target image frame; tracking the corresponding target person based on the upper body detection frame and the feature vector to obtain a corresponding tracking record; adding the upper body detection frame as a person tracking frame to the corresponding tracking record; adding the person category corresponding to the upper body detection frame as the person tracking frame to the corresponding tracking record; and performing customer flow counting based on the tracking record corresponding to each target person and the person category corresponding to each person tracking frame in the tracking record. This method can improve the accuracy of customer flow counting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a method, apparatus, computer equipment, and storage medium for counting customer flow in a store. Background Art

[0002] Store traffic statistics refer to the number of customers visiting a store over a period of time. They are widely used in various industries, including shopping malls, large supermarkets, banks, specialty stores, and the logistics industry. The customer flow obtained through store traffic statistics is a valuable reference for businesses conducting various commercial activities. For example, in the logistics industry, the daily conversion rate, calculated by counting the number of customers visiting a store and the number of parcels sent by customers, is crucial for evaluating store operations.

[0003] Currently, common methods for counting store traffic include manual, infrared, and facial recognition-based methods. Manual methods require significant manpower and resources and suffer from low accuracy, particularly in high-volume store environments where miscounts and omissions can further reduce accuracy. Infrared and facial recognition-based methods, on the other hand, have strict requirements for hardware installation locations and are prone to misjudgments and omissions in complex scenarios, such as those with obstacles and repeated entries and exits, reducing accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a store customer flow statistics method, device, computer equipment and storage medium that can improve the accuracy of customer flow statistics in response to the above technical problems.

[0005] A store customer flow statistics method, the method comprising:

[0006] Get the target image frame;

[0007] Performing upper body detection on the target image frame to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame;

[0008] determining a feature vector of the corresponding target person according to the upper body image corresponding to each upper body detection frame in the target image frame;

[0009] Tracking the corresponding target person according to the upper body detection frame and the feature vector to obtain a corresponding tracking record; adding the upper body detection frame as a person tracking frame to the corresponding tracking record; and adding the person category corresponding to the upper body detection frame as the person category corresponding to the person tracking frame to the corresponding tracking record;

[0010] Passenger flow statistics are performed based on the tracking record corresponding to each target person and the person category corresponding to each person tracking box in the tracking record.

[0011] A store customer flow counting device, comprising:

[0012] An acquisition module, used for acquiring a target image frame;

[0013] a detection module, configured to detect the upper body of a person in the target image frame, and obtain an upper body detection frame and a person category corresponding to each target person in the target image frame;

[0014] a determination module, configured to determine a feature vector of a corresponding target person based on an upper body image corresponding to each upper body detection frame in the target image frame;

[0015] A tracking module is configured to track a corresponding target person according to the upper body detection frame and the feature vector to obtain a corresponding tracking record; the upper body detection frame is added as a person tracking frame to the corresponding tracking record; and the person category corresponding to the upper body detection frame is added as a person category corresponding to the person tracking frame to the corresponding tracking record;

[0016] The statistics module is used to perform passenger flow statistics based on the tracking record corresponding to each target person and the person category corresponding to each person tracking frame in the tracking record.

[0017] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0018] A computer-readable storage medium stores a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0019] The above-mentioned store customer flow counting method, device, computer equipment and storage medium, by detecting the upper body of the target person in the target image frame, obtains the upper body detection frame and person category corresponding to each target person, which can avoid the problem of low detection accuracy of the target person due to a high degree of lower body occlusion. In this way, when tracking and counting customers based on the upper body detection frame and person category corresponding to the target person and with high accuracy, the accuracy of customer flow counting can be improved. Furthermore, based on the upper body detection frame with high accuracy, the upper body image of the target person can be accurately located from the target image frame. Based on this upper body image with high accuracy, the target person's feature vector can be accurately obtained. By combining the upper body detection frame and feature vector corresponding to each target person, the target person can be tracked, and then customer flow counting can be performed based on the tracking record of each target person and the person category corresponding to each person tracking frame in the tracking record, which can further improve the accuracy of customer flow counting. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 1 is a flow chart of a method for counting store customer flow in one embodiment;

[0021] Figure 2 A schematic diagram of the principle of store customer flow statistics in one embodiment;

[0022] Figure 3 A flowchart of a method for counting store customer flow in another embodiment;

[0023] Figure 4 This is a structural block diagram of a store customer flow statistics device in one embodiment;

[0024] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0026] In one embodiment, Figure 1 As shown, a method for counting customer flow in a store is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0027] Step 102: Acquire a target image frame.

[0028] The target image frame is an image frame for upper body detection of a person, and specifically may be a video frame extracted from a video.

[0029] Specifically, the terminal acquires the target image frame through an image acquisition device. The image acquisition device may be a webcam or a still camera. The image acquisition device may be built into the terminal as a component or connected externally to the terminal as a standalone device. The image acquisition device may communicate with the terminal via wired or wireless means.

[0030] In one embodiment, the terminal captures video in real time using an image acquisition device, and sequentially extracts video frames from the captured video as target image frames. Based on the sequentially extracted target image frames, store customer flow statistics are performed according to the store customer flow statistics method provided in this application. Specifically, the terminal can sequentially extract target image frames from the video at a preset sampling frequency. The preset sampling frequency can be customized according to actual needs, such as 3 frames per second.

[0031] Step 104 , performing upper body detection on the target image frame to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame.

[0032] The target person is the person to be detected in the currently acquired target image frame. The upper body detection frame is used to locate the upper body of the target person in the target image frame. The person category is the category to which the target person belongs, which can include employees and customers.

[0033] Specifically, after acquiring the target image frame, the terminal performs upper body detection on the target person in the target image frame, that is, performs upper body recognition and confirmation, determines the minimum circumscribed rectangle for the upper body of each detected target person to obtain the corresponding upper body detection frame, and uses the upper body detection frame as the upper body detection frame corresponding to the corresponding target person in the target image frame, and determines the person category corresponding to the target person based on the upper body of each detected target person.

[0034] In one embodiment, since each target person in the target image frame corresponds to an upper body detection frame and a person category, the person category corresponding to each target person can be determined as the person category corresponding to the upper body detection frame corresponding to the target person. In other words, each upper body detection frame corresponds to a person category. It is understood that the terminal may determine the person category of a target person wearing a clearly employee-style outer garment in the target image frame as an employee, and determine the person category of other target persons as customers.

[0035] In one embodiment, the terminal performs upper body detection on the target image frame using a trained target detection model to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame. Specifically, the terminal inputs the target image frame into the trained target detection model to perform upper body detection to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame. It can be understood that the upper body detection frame corresponding to the target person is the minimum circumscribed rectangle of the upper body of the target person in the target image frame, and the image area in the target image frame that is within the upper body detection frame is the upper body image corresponding to the target person. Therefore, based on the upper body detection frame, the upper body of the corresponding target person can be located from the target image frame, and the upper body image of the target person can be extracted.

[0036] In one embodiment, the training steps of the target detection model include: obtaining multiple sample image frames, labeling the upper body and person category of each person in each sample image frame, obtaining the upper body detection frame and person category corresponding to each person in each sample image frame, obtaining a first training sample set based on the multiple sample image frames and the upper body detection frame and person category corresponding to each person in each sample image frame, taking the sample image frame as the input feature, and using the upper body detection frame and person category corresponding to each person in the sample image frame as the desired output feature for model training to obtain a trained target detection model.

[0037] In one embodiment, the sample image frames in the first training sample set are collected from multiple stores. Specifically, multiple sample image frames may be collected from each store. Because different stores have different front desk locations and different image acquisition equipment acquisition angles, collecting sample image frames from multiple stores can increase the number of scenarios in which the sample image frames can be collected. Thus, the object detection model trained based on sample image frames collected in different scenarios can be applied to upper body detection of people in target image frames in different scenarios, thereby improving the scenario adaptability of the object detection model.

[0038] In one embodiment, during the model training phase, an image acquisition device captures initial sample image frames, each of which is scaled to a preset size. The scaled initial sample image frames are then subjected to one or more image enhancement techniques to obtain sample image frames for training the object detection model. The preset size can be customized based on actual conditions, such as 512 x 416 pixels. Image enhancement techniques include, but are not limited to, random flipping, color transformation, and geometric transformation.

[0039] In one embodiment, during the model training stage, the terminal may obtain a first test sample set in a similar manner to obtaining the first training sample set, so as to test the target detection model trained by the first training sample set through the first test sample set, and when the test passes, the trained target detection model is determined as a trained target detection model, otherwise, the target detection model continues to be iteratively trained.

[0040] In one embodiment, the machine learning algorithms involved in the training process of the target detection model include, but are not limited to, yolov3 (a machine learning algorithm that uses multi-scale features for object detection).

[0041] Step 106 : Determine a feature vector of the corresponding target person based on the upper body image corresponding to each upper body detection frame in the target image frame.

[0042] Among them, the upper body image refers to the image or image area corresponding to the upper body of the target person in the target image frame. The feature vector is a vector used to characterize the upper body features of the target person. The feature distance between multiple feature vectors corresponding to the same target person is small, and the feature distance between feature vectors corresponding to different target persons is large. Therefore, when the feature distance between two feature vectors is small enough, the two feature vectors are determined to be feature vectors corresponding to the same target person. The feature distance between two feature vectors is small enough, which specifically refers to the feature distance being less than or equal to a preset feature distance threshold. The preset feature distance threshold can be dynamically determined based on the subsequent first feature distance threshold and the second feature distance threshold, and is not specifically limited here. It can be understood that since the present application realizes the detection and positioning of the target person by detecting the upper body of the target person, the feature vector corresponding to the upper body image of the target person can be determined as the feature vector corresponding to the target person.

[0043] Specifically, after the terminal detects the upper body detection frame corresponding to each target person from the target image frame, it extracts the upper body image corresponding to the target person from the target image frame according to the upper body detection frame corresponding to each target person, and determines the feature vector corresponding to the corresponding target person based on each extracted upper body image.

[0044] In one embodiment, the terminal extracts the image region where each upper body detection frame is located from the target image frame as the upper body image corresponding to the upper body detection frame.

[0045] In one embodiment, the terminal uses a trained pedestrian re-identification model to identify each upper body image extracted from the target image frame, obtains the feature vector corresponding to each upper body image, and determines the feature vector corresponding to each upper body image as the feature vector corresponding to the corresponding target person.

[0046] In one embodiment, the training steps of the pedestrian re-identification model include: obtaining a sample image frame set, the sample image frame set includes multiple sample image frame subsets, the multiple sample image frames in each sample image frame subset are collected from the same store, and the multiple sample image frames are multiple video frames continuously extracted from the same video, or the multiple sample image frames are multiple video frames extracted sequentially from the same video according to a preset sampling frequency, and the multiple sample image frames in each sample image frame subset include the same person; determining the upper body detection frame corresponding to each person in each sample image frame in the sample image frame through a trained target detection model or manual labeling method, and extracting the corresponding upper body sample image from the corresponding sample image frame according to the upper body detection frame.

[0047] Furthermore, for each subset of sample image frames, the upper body sample images corresponding to each person are clustered according to the intersection-over-union ratio between the upper body detection frames in two adjacent sample image frames, and the multiple upper body sample images corresponding to each person in the sample image frame subset are clustered into an upper body sample image sequence; the upper body sample image sequence corresponding to each person in the sample image frame set is determined in a similar manner, and based on the upper body sample image sequence corresponding to each person, an upper body sample image set is obtained as the second training sample set; the model is iteratively trained based on the upper body sample image set to obtain a trained pedestrian re-identification model.

[0048] In the training process of a pedestrian re-identification model based on an upper body sample image set, the upper body sample image set is divided into multiple upper body sample image subsets, each upper body sample image subset includes multiple upper body sample images corresponding to multiple persons. In each iterative training process, each upper body sample image in any upper body sample image subset is used as an input feature and input into the pedestrian re-identification model to be trained for recognition, and a predicted feature vector corresponding to each upper body sample image is obtained. The feature distance between multiple predicted feature vectors corresponding to the same person is smaller, and the feature distance between the predicted feature vectors corresponding to different persons is larger. The target loss function is determined, and the model parameters of the pedestrian re-identification model to be trained are reversely adjusted based on the target loss function. Then, the pedestrian re-identification model with adjusted model parameters is iteratively trained in the above manner until the iteration stops, and a trained pedestrian re-identification model is obtained.

[0049] In one embodiment, during the model training phase, each upper body sample image is resized to a target size, padded with a target number of zero pixels, and randomly cropped to the target size. These images are then used as upper body sample images for training the person re-identification model. For example, the target size is 256*128 pixels, and the target number is 10 pixels.

[0050] In one embodiment, the machine learning algorithms involved in the training of the person re-identification model include, but are not limited to, a ReID network (Person Re-identification). In one embodiment, the ReID network may specifically include a ResNet50 network for extracting feature maps from upper body images and inputting the feature maps into subsequent networks for processing.

[0051] Step 108: Track the corresponding target person according to the upper body detection frame and the feature vector to obtain a corresponding tracking record; the upper body detection frame is added to the corresponding tracking record as a person tracking frame; the person category corresponding to the upper body detection frame is added to the corresponding tracking record as a person category corresponding to the person tracking frame.

[0052] Specifically, the terminal matches and tracks each target person based on the upper body detection frame and feature vector corresponding to each target person in the target image frame to obtain a tracking record corresponding to each target person. The terminal uses the upper body detection frame corresponding to each target person as the person tracking frame, and the person category corresponding to the upper body detection frame as the person category corresponding to the person tracking frame. The terminal then adds the person tracking frame and the corresponding person category to the tracking record corresponding to the target person.

[0053] In one embodiment, the terminal also uses the feature vector corresponding to each upper body detection frame as the feature vector corresponding to the corresponding person tracking frame, and adds the feature vector together with the corresponding person tracking frame and person category to the corresponding tracking record.

[0054] In one embodiment, if the terminal has one or more tracking records stored locally, each tracking record includes one or more person tracking frames corresponding to the same person, and a feature vector corresponding to each person tracking frame. The terminal matches the upper body detection frame and feature vector corresponding to each target person in the target image frame with the person tracking frame and feature vector in each tracking record stored locally, so as to determine the tracking record that matches each target person in the target image frame locally based on the matching results. If a tracking record that matches the target person exists locally, the terminal updates the tracking record corresponding to the target person based on the upper body detection frame and feature vector corresponding to the target person, that is, the upper body detection frame corresponding to the target person is added as a person tracking frame to the tracking record corresponding to the target person, and the feature vector corresponding to the target person is added as a feature vector corresponding to the person tracking frame to the corresponding tracking record. If there is no tracking record that matches the target person, it indicates that the target person is a new person, and the terminal initializes the tracking record corresponding to the target person based on the upper body detection frame and feature vector corresponding to the target person.

[0055] In one embodiment, the terminal determines the person identification that matches each target person in the target image frame from the matcher based on the upper body detection frame and feature vector corresponding to each target person in the target image frame, and the person tracking frame and feature vector corresponding to each person identification in the matcher, and updates the tracking record corresponding to the person identification that matches the target person in the matcher based on the upper body detection frame and feature vector corresponding to each target person, that is, updates the tracking record corresponding to the target person, so as to track the target person based on the tracking record. It can be understood that for multiple target image frames extracted sequentially from the video, the terminal determines the upper body detection frame and feature vector corresponding to each target person in the currently extracted target image frame in accordance with the above method, and updates the upper body detection frame and feature vector corresponding to each target person to the tracking record corresponding to the matching person identification in the matcher according to the above tracking method, so as to store the upper body detection frame and feature vector corresponding to the same target person in the multiple target image frames corresponding to the same person identification, and obtain the tracking record corresponding to the person identification.

[0056] Step 110 , performing passenger flow statistics based on the tracking record corresponding to each target person and the person category corresponding to each person tracking frame in the tracking record.

[0057] Specifically, based on the upper body detection frame and feature vector corresponding to each target person in the currently acquired target image frame, after updating the tracking records corresponding to each target person in accordance with the record update method provided in one or more embodiments of this application, the terminal performs passenger flow statistics based on the number of personnel tracking frames in the tracking record corresponding to each target person, and the personnel category corresponding to each personnel tracking frame to obtain the current corresponding passenger flow.

[0058] In one embodiment, the terminal performs passenger flow statistics based on the number of tracking frames within the counting area in the tracking record corresponding to each target person, as well as the person category corresponding to each tracking frame within the counting area. For each target person, if the number of first tracking frames within the counting area in the tracking record corresponding to the target person is greater than or equal to a threshold number, and the number of second tracking frames within the counting area with the person category of customer is greater than or equal to a second threshold number, then passenger flow is counted for that target person, i.e., the passenger flow is incremented by 1. Thus, based on the passenger flow obtained based on the previous target image frame, passenger flow is counted for each target person in the currently acquired target image frame according to the passenger flow counting method to obtain the current corresponding passenger flow. This current corresponding passenger flow is then determined as the passenger flow obtained based on the currently acquired target image frame, which is the passenger flow counted as of the acquisition time of the target image frame. The counting area refers to the valid area in which target persons can be counted, and specifically may refer to the image area within the target image frame that includes the foreground position.

[0059] In one embodiment, if the tracking record corresponding to the target person does not meet the passenger flow counting conditions, passenger flow counting will not be performed for the target person. If the tracking record meets the passenger flow counting conditions for the target person, after passenger flow counting is performed according to the passenger flow counting method, a counted flag will be stored corresponding to the tracking record of the target person so that the target person will not be counted again in the subsequent passenger flow counting process, which can improve the accuracy of passenger flow counting. It is understood that if the target person has been counted, the tracking record of the target person will still be updated according to the upper body detection frame and feature vector corresponding to the target person in the subsequent passenger flow counting process, but passenger flow counting will not be performed for the target person again.

[0060] In one embodiment, for each target person, the terminal counts the length of time the target person is in the counting area based on the various tracking frames of the people in the counting area in the tracking record corresponding to the target person. Specifically, the terminal determines the length of time the target person is in the counting area based on the number of the first tracking frames of the people tracking frames in the counting area in the tracking record corresponding to each target person, and the preset sampling frequency of the target image frame. For example, if the preset sampling frequency is 3 frames / second and the number of the first tracking frames is 180, it indicates that the target person is in the counting area for 1 minute. It can be understood that the length of time may refer to the duration of time the target person is continuously in the counting area, or it may refer to the cumulative length of time the target person is intermittently in the counting area.

[0061] In one embodiment, the terminal determines the matching person identifier of each target person in the target image frame, the passenger flow counted as of the acquisition time point of the target image frame, and the length of time each target person is in the counting area as of the acquisition time point of the target image frame according to the matching method and passenger flow counting method provided in one or more embodiments of the present application, and marks the upper body of the corresponding target person in the target image frame through the upper body detection frame corresponding to each target person, and marks the matching person identifier of each target person and the length of time each target person is in the counting area according to the marked position of the upper body detection frame. The passenger flow counted as of the acquisition time point of the target image frame can also be marked at a preset position in the target image frame, such as the upper left corner of the target image frame, which is not specifically limited here. In this way, visual passenger flow counting and tracking can be achieved. Furthermore, based on displaying the marked target image frames in the order of extraction, dynamic and visual passenger flow counting and tracking can be achieved.

[0062] The store customer flow counting method detects the upper body of a target person in a target image frame to obtain an upper body detection frame and person category corresponding to each target person. This method can avoid the problem of low target person detection accuracy due to a high degree of lower body occlusion, so that when tracking and counting customers based on the upper body detection frame and person category corresponding to the target person with high accuracy are performed, the accuracy of customer flow counting can be improved. Furthermore, based on the high-accuracy upper body detection frame, the upper body image of the target person can be accurately located from the target image frame. Based on this high-accuracy upper body image, the target person's feature vector can be accurately obtained. By combining the upper body detection frame and feature vector corresponding to each target person, the target person can be tracked. Then, customer flow counting is performed based on the tracking record of each target person and the person category corresponding to each person tracking frame in the tracking record, which can further improve the accuracy of customer flow counting.

[0063] In one embodiment, step 108 includes: determining a matching relationship between the target person and a person identifier in a matcher based on an upper body detection frame and a feature vector; the matcher is used to store tracking records corresponding to the target customer; each target customer in the matcher corresponds to a person identifier; when there is a person identifier matching the target person in the matcher, updating the tracking record corresponding to the matching person identifier in the matcher based on the upper body detection frame and the feature vector corresponding to the target person to obtain an updated tracking record.

[0064] The matcher is used to store tracking records corresponding to target customers, that is, to store tracking records corresponding to the people whose customer flow needs to be counted. Target customers are people whose customer flow needs to be counted. Each target customer in the matcher is represented by a person ID.

[0065] In one embodiment, the matcher can be understood as a tracking queue group, including one or more tracking queues, each personnel identifier corresponds to a tracking queue, and the tracking queue corresponding to each personnel identifier is used to store the personnel tracking frame and the corresponding feature vector obtained when tracking the target person matching the personnel identifier. The personnel tracking frames stored in sequence in each tracking queue constitute the tracking record corresponding to the corresponding target person, so that the corresponding target person can be tracked based on the personnel tracking frames stored in sequence in the tracking queue, that is, the movement trajectory of the corresponding target person can be determined.

[0066] Specifically, the terminal matches the upper body detection frame corresponding to each target person in the target image frame with the person tracking frame corresponding to each person identifier in the matcher, and matches the feature vector corresponding to each target person in the target image frame with the feature vector corresponding to each person identifier in the matcher, and confirms the matching relationship between each target person in the target image frame and each person identifier in the matcher based on the matching results, that is, it determines whether there is a person identifier that matches each target person in the matcher. When it is determined that there is a person identifier that matches the target person in the target image frame in the matcher, the terminal updates the tracking record corresponding to the person identifier that matches the target person in the matcher based on the upper body detection frame and feature vector corresponding to each target person with the matching person identifier, that is, it updates the tracking record corresponding to the target person in the matcher to obtain the updated tracking record corresponding to the target person.

[0067] In one embodiment, each person tracking frame in the tracking record corresponding to each person identification is an upper body detection frame detected in sequence from multiple target image frames and corresponding to the same target person. Therefore, based on the upper body detection frame and feature vector corresponding to each target person in the current target image frame, when it is determined that the target person matches the person identification in the matcher, it indicates that the target person and the person corresponding to the matched person identification in the matcher are the same person. Therefore, the upper body detection frame and feature vector corresponding to the target person are updated to the tracking record corresponding to the matched person identification.

[0068] In the above embodiment, the upper body detection frame and feature vector corresponding to each target person are updated to the tracking record corresponding to the matched person identification in the matcher to update the tracking record corresponding to the target person, so as to accurately track the target person based on the updated tracking record.

[0069] In one embodiment, the matching relationship between the target person and the person identification in the matcher is determined based on the upper body detection frame and the feature vector, including: determining the matching relationship between the person identification in the matcher and the target person based on the intersection-over-union ratio and the center distance between the upper body detection frame corresponding to each target person and the person tracking frame corresponding to each person identification in the matcher, and the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identification.

[0070] The intersection-over-union (IOU) ratio between the upper body detection frame and the person tracking frame, also known as the IOU, specifically refers to the intersection area between the upper body detection frame and the person tracking frame divided by the combined area between the upper body detection frame and the person tracking frame. It can be understood that since each target image frame extracted sequentially has the same size, and the upper body detection frame and the person tracking frame each have corresponding coordinate information in the corresponding target image frame, the positions of the upper body detection frame and the person tracking frame can be determined in the same target image frame based on their respective coordinate information. Based on their respective positions in the same image frame, the intersection area and combined area of ​​the upper body detection frame and the person tracking frame can be calculated, and the corresponding IOU ratio can be derived based on the intersection area and combined area. The center distance between the upper body detection frame and the person tracking frame refers to the distance between the center point of the upper body detection frame and the center point of the person tracking frame. The feature distance between feature vectors is used to characterize the similarity between the two feature vectors, and may include, but is not limited to, Euclidean distance, cosine similarity, and correlation coefficient.

[0071] Specifically, the terminal calculates the intersection-over-union ratio and center distance between the upper body detection frame corresponding to each target person in the target image frame and the person tracking frame corresponding to each person identifier in the matcher, and calculates the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identifier in the matcher, and determines the matching relationship between each target person in the target image frame and each person identifier in the matcher based on the calculated intersection-over-union ratio, center distance and feature distance.

[0072] In one embodiment, the terminal combines each target person in the target image frame with each person identifier in the matcher to obtain a plurality of combination pairs, each combination pair comprising a target person and a person identifier. The terminal selects a first combination pair from the plurality of combination pairs, wherein the upper body detection frame corresponding to the target person is located in a first predetermined area of ​​the target image frame. The terminal calculates an intersection-over-union (IoU) between the upper body detection frame corresponding to the target person in each first combination pair and the most recently recorded person tracking frame corresponding to the person identifier, to obtain an IoU corresponding to the first combination pair. When one and only one IoU corresponding to each first combination pair is greater than or equal to an IoU threshold, the target person in the first combination pair having an IoU greater than or equal to the IoU threshold is determined to match the person identifier. In other words, the upper body detection frame corresponding to the target person is determined to match the most recently recorded person tracking frame corresponding to the person identifier, and the first combination pair having an IoU greater than or equal to the IoU threshold is determined to be a matched pair. The matching device then updates the tracking record corresponding to the matched person identifier in the matcher based on the upper body detection frame corresponding to the target person and the feature vector.

[0073] It is understood that the first preset area refers to a pre-defined image area in the target image frame, such as a counting area, or an image area determined by a first preset numerical range. The counting area refers to an effective area in which the target person can be counted. Specifically, it can refer to the image area in the target image frame that includes the front desk position. When the target person is within the counting area in the target image frame, it indicates that the target person may want to send a package, indicating that passenger counting may be required for the target person. The upper body detection frame being within the counting area means that the center point of the upper body detection frame is within the counting area. The first preset numerical range is determined by a first value and a second value. The first value is less than the second value. The first and second values ​​are customizable. For example, if the first value is 0.05 and the second value is 0.95, the first preset numerical range is 0.05 to 0.95. The upper body detection frame is determined to be within the first preset area of ​​the target image frame when the horizontal coordinate of the right border of the upper body detection frame is greater than or equal to the product of the width of the target image frame and the first value, or when the horizontal coordinate of the left border of the upper body detection frame is less than or equal to the product of the width of the target image frame and the second value. The IoU threshold can be customized, for example, 0.85.

[0074] Furthermore, the terminal removes the first combination pair whose intersection-and-union ratio is greater than or equal to the intersection-and-union ratio threshold value, and the combination pair including the target person or person identification in the first combination pair whose intersection-and-union ratio is greater than or equal to the intersection-and-union ratio threshold value from the multiple combination pairs initially constructed, to obtain the second combination pair. The terminal calculates the feature distance based on the feature vectors corresponding to the target person and the person identification in each second combination pair, to obtain the feature distance corresponding to each second combination pair. For each second combination pair, the terminal may calculate the feature distance between the feature vector corresponding to the target person in the second combination pair and the feature vector of the latest record corresponding to the person identification, to obtain the feature distance corresponding to the second combination pair. The terminal may also calculate the feature distance between the feature vector corresponding to the target person in the second combination pair and the feature vector of the latest record corresponding to the person identification, to obtain the feature distance of the preset number of feature distances, and use the average feature distance obtained from the preset number of feature distances as the feature distance corresponding to the second combination pair. When the smallest characteristic distance among the characteristic distances corresponding to each second combination pair is less than or equal to the first characteristic distance threshold, it is determined that the target person in the second combination pair with the smallest characteristic distance matches the person identification, and the second combination pair with the smallest characteristic distance is determined as a matching pair, and then the tracking record corresponding to the matching person identification in the matcher is updated according to the upper body detection frame and characteristic vector corresponding to the target person. Among them, the first characteristic distance threshold can be customized, such as 1.1. The preset number can be customized, such as 20. It can be understood that when the number of characteristic vectors in the tracking record corresponding to the person identification in the matcher is less than the preset number, the characteristic distance corresponding to the corresponding second combination is determined based on all the characteristic vectors in the tracking record and the characteristic vector corresponding to the corresponding target person according to the above-mentioned characteristic distance calculation method.

[0075] Furthermore, the terminal removes from the second set of pairs the second set of pairs having the smallest feature distance, where the smallest feature distance is less than or equal to the first feature distance threshold, and the second set of pairs where the upper body detection frame corresponding to the target person is not within the second preset area of ​​the target image frame, to obtain a third set of pairs. For each third set of pairs, the terminal calculates the center distance between the upper body detection frame corresponding to the target person in the third set of pairs and the most recently recorded person tracking frame corresponding to the person identifier, as well as the feature distance between the feature vector corresponding to the target person and the feature vector corresponding to the person identifier, to obtain the center distance and feature distance corresponding to the third set of pairs. If a third set of pairs exists in each third set of pairs where the center distance is less than or equal to the center distance threshold and the feature distance is less than or equal to the second feature distance threshold, the terminal selects the third set of pairs with the smallest center distance from the third set of pairs where the center distance is less than or equal to the center distance threshold and the feature distance is less than or equal to the second feature distance threshold as a matching pair, determines that the target person in the selected third set of pairs matches the person identifier, and then updates the tracking record corresponding to the matching person identifier in the matcher based on the upper body detection frame and feature vector corresponding to the target person.

[0076] It can be understood that the second preset area is, for example, a counting area, or an image area determined by the counting area and the second preset numerical range. The definition of the counting area is consistent with that of FIG, and will not be repeated here. The second preset numerical range is determined by the third numerical value and the fourth numerical value, the third numerical value is less than the fourth numerical value, and the third numerical value and the fourth numerical value can be customized. For example, if the third numerical value is 0.01 and the fourth numerical value is 0.99, then the second preset numerical range is 0.01 to 0.99. The second preset area is determined in a similar manner according to the counting area and the second preset numerical range, and will not be repeated here. The center distance threshold is dynamically determined by the upper body detection frame corresponding to the target person in the corresponding third combination pair, and the latest recorded person tracking frame corresponding to the person identification. Specifically, it can be half of the sum of the diagonals of the upper body detection frame and the person tracking frame. The second feature distance threshold can be customized, such as 1.3.

[0077] Furthermore, the terminal removes the second combination pair with the smallest feature distance, and the smallest feature distance is less than or equal to the first feature distance threshold, from the second combination pair, as well as the matching pair determined from the third combination pair in the above manner, to obtain a fourth combination pair. According to a similar process, the terminal calculates the center distance and feature distance corresponding to each fourth combination pair respectively, and selects the fourth combination pair with the smallest center distance as the matching pair from the fourth combination pairs whose center distance is less than or equal to the center distance threshold and whose feature distance is less than or equal to the second feature distance threshold, and determines that the target person in the selected fourth combination pair matches the person identification, and then updates the tracking record corresponding to the matching person identification in the matcher based on the upper body detection frame and feature vector corresponding to the target person.

[0078] It can be understood that during the sequential matching process between the target person in the target image frame and the person identifier in the matcher, if a matching relationship between each target person and the corresponding person identifier in the matcher has been determined, there is no need to continue the subsequent matching process. In this way, combining the intersection-over-union ratio, center distance, and feature distance to determine the matching relationship between the target person and the person identifier and updating the tracking record of the target person accordingly can reduce false positives and missed positives, thereby improving the accuracy of passenger flow statistics based on the tracking records.

[0079] In one embodiment, the upper body detection frame is represented by detection frame data, and based on the detection frame data, the corresponding upper body detection frame can be uniquely located in the target image frame. The detection frame data, for example, includes the center point coordinates, height, and width of the upper body detection frame, and also includes the upper left point coordinates and lower right point coordinates of the upper body detection frame, which are not specifically limited here. Similarly, the person tracking frame is represented by tracking frame data, and based on the tracking frame data, the corresponding person tracking frame can be uniquely located in the target image frame. It can be understood that in one or more embodiments, the center distance and intersection-over-union ratio between the upper body detection frame and the person tracking frame are determined based on the corresponding positions of the upper body detection frame and the person tracking frame in the target image frame, that is, based on the detection frame data corresponding to the upper body detection frame and the tracking frame data corresponding to the person tracking frame.

[0080] In the above embodiment, the matching relationship between the target person and the person identification is dynamically determined based on the intersection-over-union ratio and the center distance between the upper body detection frame corresponding to each target person and the person tracking frame corresponding to each person identification in the matcher, as well as the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identification. This can improve the accuracy of the matching relationship, so that the tracking records corresponding to each target person can be updated according to the matching relationship with higher accuracy, thereby improving the accuracy of the tracking records.

[0081] In one embodiment, step 108 further includes: when there is no person identifier matching the target person in the matcher, determining a matching person identifier from the tracker based on the upper body detection frame and feature vector corresponding to the target person; the tracker is used to store tracking records corresponding to interfering customers; each interfering customer in the tracker corresponds to a person identifier; based on the upper body detection frame and feature vector corresponding to the target person, updating the tracking record corresponding to the matching person identifier in the tracker to obtain an updated tracking record; and migrating the tracking records in the tracker that meet the record migration conditions to the matcher.

[0082] The tracker is used to store tracking records corresponding to interfering customers, that is, to store tracking records corresponding to objects that may need to be counted. Interfering customers are objects that may need to be counted. These objects may be people who will later need to be counted, people who briefly appear in the image acquisition device's field of view, or objects that are mistakenly identified as people, such as clothing. Each interfering customer in the tracker corresponds to a person identifier.

[0083] The tracker records the tracking records corresponding to interfering customers, while the matcher records the tracking records corresponding to target customers. When the tracking records for an interfering customer in the tracker meet the record migration conditions, the interfering customer is identified as the target customer, and the tracking records for the interfering customer in the tracker are migrated to the matcher as the tracking records for the target customer. The tracking records corresponding to the target customer in the matcher are then updated based on subsequently acquired target image frames. As a result, the interfering customers tracked by the tracker through the tracking records are not required for customer flow counting, while the target customers tracked by the matcher through the tracking records are required for customer flow counting.

[0084] The record migration condition is the condition or basis for determining whether to migrate tracking records from the tracker to the matcher. For example, a record migration condition requires that the person ID corresponding to the tracking record matches the target person in a specified number of sequentially extracted target image frames. This means that a matching target person exists in multiple sequentially extracted target image frames. The specified number can be customized, for example, 9.

[0085] In one embodiment, the tracker can be understood as a tracking queue group, including one or more tracking queues, each personnel identifier corresponds to a tracking queue, and the tracking queue corresponding to each personnel identifier is used to store the personnel tracking frame and the corresponding feature vector obtained when tracking the target person matching the personnel identifier. The personnel tracking frames stored in sequence in each tracking queue constitute the tracking record corresponding to the corresponding target person, so that the corresponding target person can be tracked based on the personnel tracking frames stored in sequence in the tracking queue, that is, the movement trajectory of the corresponding target person can be determined.

[0086] Specifically, when, according to the matching method provided in one or more embodiments, it is determined that there is no person identifier matching the target person in the matcher, the terminal determines, according to the matching method provided in one or more embodiments, the matching relationship between each target person and each person identifier in the tracker based on the upper body detection frame and feature vector corresponding to each target person for which there is no matching person identifier in the matcher, and the other person tracking frame and feature vector corresponding to each person identifier in the tracker, that is, the person identifier matching each target person for which there is no matching person identifier in the matcher is determined from the tracker. After the person identifier matching the target person is determined from the tracker according to the matching method, the tracking record corresponding to the matching person identifier in the tracker is updated according to the upper body detection frame and feature vector corresponding to the target person. If, according to the matching method, it is determined that there is no person identifier matching the target person in the tracker, the target person is determined to be a new person, and the tracking record corresponding to the target person is initialized in the tracker based on the upper body detection frame and feature vector corresponding to the target person, and the corresponding person identifier is configured for the tracking record as the person identifier matching the target person.

[0087] Furthermore, after updating the tracking records corresponding to the matched person identification in the tracker according to the upper body detection frame and feature vector corresponding to the target person in accordance with the above-mentioned update method, the terminal dynamically determines whether the tracking records corresponding to each person identification in the tracker meet the preset record migration conditions, and migrates the tracking records that meet the record migration conditions to the matcher, that is, adding the tracking record in the matcher and deleting the tracking record in the tracker.

[0088] In one embodiment, for a person who has newly entered the field of view of the image acquisition device, that is, for a newly added person, a tracking record corresponding to the newly added person is added in the tracker, and the tracking record of the newly added person is updated in the above-mentioned manner according to the target image frames extracted in sequence, until a target person matching the newly added person exists in a plurality of target image frames extracted in sequence, and the target person is determined to be a person who needs to be counted for passenger flow, and the tracking record corresponding to the newly added person in the tracker is migrated to the matcher, so as to facilitate further updating of the tracking record of the newly added person based on the matcher, and passenger flow counting is performed according to the passenger flow counting method provided in this application based on the updated tracking record. In this way, by tracking and filtering the newly added target person through the tracker, it is possible to avoid subsequent tracking and passenger flow counting of target persons with misjudgment, thereby improving the accuracy of passenger flow statistics.

[0089] In one embodiment, the target image frame includes multiple target persons. The terminal determines whether there is a matching person identification for each target person in the matcher according to the above matching method. For the target person whose matching person identification exists in the matcher, the terminal updates the tracking record corresponding to the matching person identification in the matcher according to the upper body detection frame and feature vector corresponding to the target person. For the target person whose matching person identification does not exist in the matcher, the terminal updates the tracking record corresponding to the matching person identification in the tracker according to the upper body detection frame and feature vector corresponding to the target person, or initializes the tracking record corresponding to the target person in the tracker according to the upper body detection frame and feature vector corresponding to the target person.

[0090] In the above embodiment, the tracking of target persons and passenger flow statistics are performed by combining the matcher with the tracker, which can reduce the occurrence of misjudgments and missed judgments, thereby improving the accuracy of passenger flow statistics.

[0091] In one embodiment, step 110 includes: counting the number of first tracking frames of the person tracking frames corresponding to the corresponding target person and located in the counting area according to the tracking records corresponding to the person identifier matched with each target person in the matcher; when the number of first tracking frames is greater than or equal to a first quantity threshold, counting the number of second tracking frames corresponding to the corresponding target person and having the person category of customer according to the person category corresponding to each person tracking frame in the counting area in the corresponding tracking record; and performing passenger flow counting for the corresponding target person when the number of second tracking frames is greater than or equal to a second quantity threshold.

[0092] The first number of tracking frames refers to the number of tracking frames of people within the counting area in the tracking records corresponding to a single person identifier in the matcher. The second number of tracking frames refers to the number of tracking frames of people within the counting area whose corresponding person category is customer, in the tracking records corresponding to a single person identifier in the matcher. The first number threshold can be customized, such as 120, or dynamically determined based on the preset sampling frequency of the target image frame and the duration threshold for passenger flow counting. The duration threshold for passenger flow counting is used to compare the duration of time the target person remains in the counting area to determine whether to count the target person. If the tracking records corresponding to the target person determine that the duration of time the target person remains in the counting area is greater than or equal to the duration threshold for passenger flow counting, the passenger flow counting process is triggered for the target person. For example, if the preset sampling frequency is 3 frames / second and the duration threshold for passenger flow counting is 2 minutes, the dynamically determined first number threshold is 3*60*2=360. The second quantity threshold can be customized, such as 70, or it can be dynamically determined based on the first quantity threshold, such as determining half of the first quantity threshold as the second quantity threshold, for example, when the first quantity threshold is 360, the second quantity threshold is 180.

[0093] Specifically, the terminal determines the person identification that matches each target person in the target image frame from the matcher according to the matching method provided in one or more embodiments of the present application, and updates the tracking records corresponding to the matched person identification in the matcher according to the upper body detection frame and feature vector corresponding to each target person identification. The tracking records corresponding to each person identification in the matcher are used as the tracking records corresponding to the target person matched with the person identification, and the person tracking frames in the counting area in the tracking records corresponding to each target person are counted to obtain the first tracking frame number of the person tracking frames corresponding to the target person and in the counting area. For each target person, the terminal compares the number of first tracking frames corresponding to the target person with a preset first number threshold. When it is determined that the number of first tracking frames is greater than or equal to the first number threshold, the terminal counts the number of personnel tracking frames in the counting area of ​​the tracking record corresponding to the target person and whose personnel category is customers according to the personnel category corresponding to each personnel tracking frame in the counting area in the tracking record, and obtains the number of second tracking frames corresponding to the target person, and compares the number of second tracking frames with the second number threshold. When it is determined that the number of second tracking frames is greater than or equal to the second number threshold, the passenger flow count is performed for the target person.

[0094] In the above embodiment, whether to perform passenger flow counting on the target person is determined based on the number of first tracking frames in the counting area corresponding to each target person, and the number of second tracking frames in the counting area whose person category is customers. Passenger flow counting is performed based on the judgment result, which can improve the accuracy of customer statistics.

[0095] In one embodiment, step 104 includes: inputting the target image frame into a trained target detection model, performing upper body detection on the target image frame through the target detection model, and obtaining an upper body detection frame and person category corresponding to each target person in the target image frame.

[0096] Among them, the target detection model is a model trained based on the first training sample set and can be used to detect the upper body detection frame and person category corresponding to each target person from the target image frame.

[0097] Specifically, the terminal inputs the target image frame into a trained target detection model, and uses the target detection model to detect the upper body of each target person in the target image frame to obtain the upper body detection frame and person category corresponding to each target person in the target image frame.

[0098] In the above embodiment, the upper body of the target person in the target image frame is detected by the trained target detection model, and the upper body detection frame and person category corresponding to each target person can be obtained quickly and accurately, so that the accuracy of passenger flow statistics can be improved when personnel tracking and passenger flow statistics are performed based on the upper body detection frame and person category with higher accuracy.

[0099] In one embodiment, step 106 includes: extracting an upper body image corresponding to each upper body detection frame from the target image frame; and identifying each upper body image using a trained person re-identification model to obtain a feature vector of the corresponding target person.

[0100] The person re-identification model is trained based on the second training sample set and can be used to identify the upper body image of each target person to obtain a corresponding feature vector.

[0101] Specifically, the terminal extracts the upper body image corresponding to each upper body detection frame of each target person in the target image frame from the target image frame as the upper body image corresponding to the corresponding target person in the target image frame. The terminal inputs each extracted upper body image into a trained person re-identification model, which then recognizes each upper body image and obtains a feature vector corresponding to the target person.

[0102] It can be understood that the image area within the upper body detection frame in the target image frame is the upper body image of the target person corresponding to the upper body detection frame in the target image frame. Therefore, according to the upper body detection frame corresponding to each target person, the upper body of the target person can be located in the target image frame, and the upper body image corresponding to the target person can be extracted.

[0103] In the above embodiment, each upper body image is recognized by using a trained person re-identification model, so that the feature vector of the corresponding target person can be obtained quickly and accurately.

[0104] Figure 2 The figure is a schematic diagram of the principle of store customer flow statistics in one embodiment. After acquiring the target image frame, the terminal inputs the target image frame into the trained target detection model to perform upper body detection frame detection, obtaining an upper body detection frame corresponding to each target person in the target image frame. Based on each upper body detection frame, the upper body image of the corresponding target person is extracted from the target image frame. Each upper body image is input into the trained pedestrian re-identification model for recognition, and the feature vector corresponding to each upper body image is obtained and used as the feature vector corresponding to the corresponding target person. Furthermore, based on the upper body detection frame and feature vector corresponding to each target person in the target image frame, as well as the person tracking frame and feature vector corresponding to each person identifier in the matcher, the terminal determines the person identifier that matches each target person from the matcher, and updates the tracking record corresponding to the matched person identifier in the matcher for each target person. Furthermore, the terminal counts the length of time that the corresponding target person is in the counting area and the passenger flow as of the acquisition time point of the target image frame based on the tracking records corresponding to each person identifier in the matcher, and marks the upper body of the corresponding target person in the target image frame through the upper body detection frame, and marks the corresponding position of the corresponding upper body detection frame in the target image frame according to the matching relationship between the target person and the person identifier in the matcher, marks the person identifier corresponding to each target person, and the length of time each target person is in the counting area, and marks the passenger flow as of the acquisition time point of the target image frame in the upper left corner of the target image frame.

[0105] like Figure 2 As shown, the target image frame includes two target persons. According to the process, one target person's category is determined to be employee, and the matching person identifier for this target person is 001. The other target person's category is customer, and the matching person identifier for this target person is 002. Furthermore, the target person's duration within the counting area is 50 seconds. It is understood that since this application primarily focuses on counting customer traffic, it is not necessary to count the duration of employee identifiers within the counting area.

[0106] In the above embodiment, by combining detection with multi-target matching, the motion trajectory of each target person can be captured simultaneously. By detecting and extracting the upper body image of the target person and performing person tracking and customer count based on the upper body image, the effects of mutual occlusion between target persons or obstruction of target persons by objects can be reduced, thereby improving the accuracy of customer counts. Distinguishing between customers and employees by person category can avoid the problem of reducing customer count accuracy due to counting employees.

[0107] It is understood that the store customer flow counting method provided in one or more embodiments of this application can capture video and target image frames based on the store's existing surveillance cameras, eliminating the need for complex equipment installation. This method can improve the accuracy of customer flow counting while reducing equipment costs. Furthermore, this store customer flow counting method can ensure the accuracy of customer flow counting in different scenarios and is universally applicable.

[0108] like Figure 3 As shown, in one embodiment, a store customer flow statistics method is provided, which specifically includes the following steps:

[0109] Step 302: Acquire a target image frame.

[0110] In step 304, the target image frame is input into the trained target detection model, and the target detection model is used to detect the upper body of the target person in the target image frame to obtain the upper body detection frame and person category corresponding to each target person in the target image frame.

[0111] Step 306: extract the upper body image corresponding to each upper body detection frame from the target image frame.

[0112] Step 308: Recognize each upper body image using the trained person re-identification model to obtain a feature vector of the corresponding target person.

[0113] Step 310, determines the matching relationship between the person identification in the matcher and the target person based on the intersection-over-union ratio and center distance between the upper body detection frame corresponding to each target person and the person tracking frame corresponding to each person identification in the matcher, and the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identification.

[0114] Step 312: When there is a person identification matching the target person in the matcher, the tracking record corresponding to the matching person identification in the matcher is updated according to the upper body detection frame and feature vector corresponding to the target person.

[0115] Step 314: When there is no person identifier matching the target person in the matcher, a matching person identifier is determined from the tracker based on the upper body detection frame and feature vector corresponding to the target person.

[0116] Step 316: Update the tracking record corresponding to the matching person identification in the tracker according to the upper body detection frame and feature vector corresponding to the target person.

[0117] Step 318: Migrate the tracking records in the tracker that meet the record migration conditions to the matcher.

[0118] Step 320 : According to the tracking records corresponding to the person identification matched with each target person in the matcher, the number of first tracking frames of the person tracking frames corresponding to the corresponding target person and located in the counting area is counted respectively.

[0119] Step 322: When the number of first tracking frames is greater than or equal to the first number threshold, the number of second tracking frames corresponding to the corresponding target person and having the person category of customer is counted according to the person category corresponding to each person tracking frame in the counting area in the corresponding tracking record.

[0120] Step 324: When the number of the second tracking frames is greater than or equal to the second number threshold, the passenger flow of the corresponding target person is counted.

[0121] In one embodiment, when the server executes the store customer flow statistics method executed by the terminal in one or more embodiments, the server can be used to execute the store customer flow statistics method for multiple stores, and can also be used to summarize and analyze the customer flow of each store, for example, the customer flow of each store in each area can be summarized and analyzed according to the region.

[0122] In one embodiment, after acquiring a target image frame, the terminal uses a trained target detection model to detect the upper body of the target person in the target image frame, obtaining an upper body detection frame and person category corresponding to each target person. Based on each target person's upper body detection frame, the upper body image of the target person is extracted from the target image frame. Each target person's upper body image is then recognized using a trained person re-identification model to obtain a feature vector for the target person. Furthermore, the terminal is locally configured with a tracker and a matcher, each of which stores tracking records corresponding to the person identifier, including the person tracking frame, person category, and feature vector. The terminal queries the person identification that matches the target person from the matcher based on the upper body detection frame and feature vector corresponding to each target person in the target image frame. When a person identification that matches the target person exists in the matcher, the terminal updates the tracking record corresponding to the matching person identification in the matcher based on the upper body detection frame, person category and feature vector corresponding to the target person. When the terminal does not match the target person identification in the matcher, the terminal further queries the person identification that matches the target person from the tracker based on the upper body detection frame and feature vector corresponding to the target person, and updates the tracking record corresponding to the matching person identification in the tracker based on the upper body detection frame, person category and feature vector corresponding to the target person. After updating the tracking records in the matcher or tracker in the above manner, the terminal dynamically detects whether each tracking record in the tracker meets the record migration condition, and migrates the tracking records that meet the record migration condition to the matcher.

[0123] It should be understood that although Figure 1 and Figure 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 and Figure 3 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0124] In one embodiment, Figure 4 As shown, a store customer flow statistics device 400 is provided, comprising: an acquisition module 401, a detection module 402, a determination module 403, a tracking module 404 and a statistics module 405, wherein:

[0125] An acquisition module 401 is used to acquire a target image frame;

[0126] Detection module 402, configured to detect the upper body of a person in a target image frame, and obtain an upper body detection frame and person category corresponding to each target person in the target image frame;

[0127] A determination module 403 is configured to determine a feature vector of a corresponding target person based on the upper body image corresponding to each upper body detection frame in the target image frame;

[0128] Tracking module 404 is configured to track the corresponding target person based on the upper body detection frame and the feature vector to obtain a corresponding tracking record; the upper body detection frame is added as the person tracking frame to the corresponding tracking record; the person category corresponding to the upper body detection frame is added as the person category corresponding to the person tracking frame to the corresponding tracking record;

[0129] The statistics module 405 is used to perform passenger flow statistics based on the tracking record corresponding to each target person and the person category corresponding to each person tracking box in the tracking record.

[0130] In one embodiment, the tracking module 404 is further used to determine the matching relationship between the target person and the person identifier in the matcher based on the upper body detection frame and the feature vector; the matcher is used to store the tracking records corresponding to the target customers; each target customer in the matcher corresponds to a person identifier; when there is a person identifier that matches the target person in the matcher, the tracking record corresponding to the matched person identifier in the matcher is updated based on the upper body detection frame and the feature vector corresponding to the target person to obtain an updated tracking record.

[0131] In one embodiment, the tracking module 404 is also used to determine the matching relationship between the person identification in the matcher and the target person based on the intersection-over-union ratio and the center distance between the upper body detection frame corresponding to each target person and the person tracking frame corresponding to each person identification in the matcher, and the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identification.

[0132] In one embodiment, the tracking module 404 is also used to determine a matching person identifier from the tracker based on the upper body detection frame and feature vector corresponding to the target person when there is no person identifier matching the target person in the matcher; the tracker is used to store tracking records corresponding to interfering customers; each interfering customer in the tracker corresponds to a person identifier; based on the upper body detection frame and feature vector corresponding to the target person, the tracking record corresponding to the matching person identifier in the tracker is updated to obtain an updated tracking record; and the tracking records in the tracker that meet the record migration conditions are migrated to the matcher.

[0133] In one embodiment, the statistics module 405 is further used to count the number of first tracking frames of the personnel tracking frames corresponding to the corresponding target person and in the counting area according to the tracking records corresponding to the personnel identification matched with each target person in the matcher; when the number of first tracking frames is greater than or equal to the first number threshold, the number of second tracking frames corresponding to the corresponding target person and whose personnel category is customer is counted according to the personnel category corresponding to each personnel tracking frame in the counting area in the corresponding tracking record; when the number of second tracking frames is greater than or equal to the second number threshold, the passenger flow count is performed for the corresponding target person.

[0134] In one embodiment, the detection module 402 is also used to input the target image frame into a trained target detection model, perform upper body detection on the target image frame through the target detection model, and obtain the upper body detection frame and person category corresponding to each target person in the target image frame.

[0135] In one embodiment, the determination module 403 is further configured to extract an upper body image corresponding to each upper body detection frame from the target image frame; identify each upper body image using a trained person re-identification model to obtain a feature vector of the corresponding target person.

[0136] The specific definition of the store customer flow counting device can be found in the definition of the store customer flow counting method above, and will not be repeated here. The various modules in the store customer flow counting device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0137] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for counting customer flow in a store is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0138] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0139] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0141] Those skilled in the art will appreciate that all or part of the processes in the embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include processes such as the embodiments of each method. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0142] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0143] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A store customer flow statistics method, characterized in that: The method comprises: Get the target image frame; Performing upper body detection on the target image frame to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame; determining a feature vector of the corresponding target person according to the upper body image corresponding to each upper body detection frame in the target image frame; Determining a matching relationship between the target person and a person identifier in a matcher based on the upper body detection frame and the feature vector; the matcher is configured to store tracking records corresponding to target customers; each target customer in the matcher corresponds to a person identifier; the upper body detection frame is added as a person tracking frame to the corresponding tracking record; and the person category corresponding to the upper body detection frame is added as the person category corresponding to the person tracking frame to the corresponding tracking record; When there is no person identifier matching the target person in the matcher, determining a matching person identifier from the tracker based on the upper body detection frame and feature vector corresponding to the target person; the tracker is used to store tracking records corresponding to the interfering customers; each interfering customer in the tracker corresponds to a person identifier; updating the tracking record corresponding to the matching person identifier in the tracker according to the upper body detection frame and feature vector corresponding to the target person to obtain an updated tracking record; Migrating the tracking records in the tracker that meet the record migration condition to the matcher; Passenger flow statistics are performed based on the tracking record corresponding to each target person and the person category corresponding to each person tracking box in the tracking record.

2. The method according to claim 1, characterized in that After determining the matching relationship between the target person and the person identifier in the matcher based on the upper body detection frame and the feature vector, the method further includes: When there is a person identification that matches the target person in the matcher, the tracking record corresponding to the matched person identification in the matcher is updated according to the upper body detection frame and feature vector corresponding to the target person to obtain an updated tracking record.

3. The method according to claim 1, characterized in that The determining, based on the upper body detection frame and the feature vector, a matching relationship between the target person and a person identifier in a matcher includes: The matching relationship between the person identification in the matcher and the target person is determined based on the intersection-over-union ratio and the center distance between the upper body detection frame corresponding to each target person and the person tracking frame corresponding to each person identification in the matcher, as well as the feature distance between the feature vector corresponding to each target person and the feature vector corresponding to each person identification.

4. The method according to claim 1, wherein The performing of passenger flow statistics according to the tracking record corresponding to each target person and the person category corresponding to each person tracking frame in the tracking record includes: According to the tracking records corresponding to the person identification matched with each target person in the matcher, respectively counting the number of first tracking frames of the person tracking frames corresponding to the corresponding target person and located in the counting area; When the number of the first tracking frames is greater than or equal to a first number threshold, counting the number of second tracking frames corresponding to the corresponding target person and having a person category of customer according to the person category corresponding to each person tracking frame in the counting area in the corresponding tracking record; When the number of the second tracking frames is greater than or equal to a second number threshold, passenger flow counting is performed on the corresponding target person.

5. The method according to any one of claims 1 to 4, characterized in that The detecting the upper body of the target person in the target image frame to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame includes: The target image frame is input into a trained target detection model, and the upper body of the person is detected in the target image frame by the target detection model to obtain the upper body detection frame and person category corresponding to each target person in the target image frame.

6. The method according to claim 5, characterized in that Determining a feature vector of a corresponding target person according to an upper body image corresponding to each upper body detection frame in the target image frame includes: Extracting an upper body image corresponding to each upper body detection frame from the target image frame; Each upper body image is identified using the trained person re-identification model to obtain the feature vector of the corresponding target person.

7. A store customer flow statistics device, characterized in that: The device comprises: An acquisition module, used for acquiring a target image frame; a detection module, configured to perform upper body detection on the target image frame to obtain an upper body detection frame and a person category corresponding to each target person in the target image frame; a determination module, configured to determine a feature vector of a corresponding target person based on an upper body image corresponding to each upper body detection frame in the target image frame; A tracking module is used to determine the matching relationship between the target person and the person identifier in the matcher based on the upper body detection frame and the feature vector; the matcher is used to store the tracking records corresponding to the target customers; each target customer in the matcher corresponds to a person identifier; the upper body detection frame is added to the corresponding tracking record as a person tracking frame; the person category corresponding to the upper body detection frame is added to the corresponding tracking record as the person category corresponding to the person tracking frame; when there is no person identifier matching the target person in the matcher, the matching person identifier is determined from the tracker based on the upper body detection frame and the feature vector corresponding to the target person; the tracker is used to store the tracking records corresponding to the interfering customers; each interfering customer in the tracker corresponds to a person identifier; based on the upper body detection frame and the feature vector corresponding to the target person, the tracking record corresponding to the matched person identifier in the tracker is updated to obtain the updated tracking record; the tracking record in the tracker that meets the record migration condition is migrated to the matcher; The statistics module is used to perform passenger flow statistics based on the tracking record corresponding to each target person and the person category corresponding to each person tracking frame in the tracking record.

8. The device according to claim 7, characterized in that The tracking module is further configured to: When there is a person identification that matches the target person in the matcher, the tracking record corresponding to the matched person identification in the matcher is updated according to the upper body detection frame and feature vector corresponding to the target person to obtain an updated tracking record.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • An indoor personnel detection method based on video image

    CN109472196A

  • Shop people flow identification and analysis method and system

    CN109871804A

  • Gesture recognition method based on skeleton point detection and tracking

    CN111368770A