People flow counting method, device, and storage medium

By calculating feature matching degree and location accuracy in pedestrian flow statistics and combining the Hungarian matching algorithm, the problems of low detection accuracy and high equipment configuration in existing technologies are solved, and high-precision lightweight pedestrian flow statistics are realized.

CN117237861BActive Publication Date: 2025-11-28HANGZHOU HUACHENG SOFTWARE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310986512.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-07
Publication Date
2025-11-28
Estimated Expiration
2043-08-07

AI Technical Summary

Technical Problem

Existing methods for counting pedestrian traffic have low detection accuracy, limited functionality, and high equipment configuration requirements for large models, resulting in poor versatility and applicability.

Method used

By acquiring the detection object information of the current video frame and historical video frames, the feature matching degree and location accuracy are calculated. The Hungarian matching algorithm is used to count the detection objects with similar features and close locations. The pedestrian flow is counted by combining feature information and trajectory information.

Benefits of technology

It improves the accuracy of people flow statistics, reduces false detections and missed detections, and achieves efficient and lightweight people flow statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237861B_ABST
    Figure CN117237861B_ABST
Patent Text Reader

Abstract

The application discloses a people flow counting method and device and a storage medium. The people flow counting method comprises the following steps: acquiring each detection object in a current video frame, position information of each detection object, and marching track information of each detection object in a historical video frame; calculating a feature matching degree between feature information of a target detection object in the current video frame and feature information of the marching track information of each detection object in the historical video frame; determining the marching track information of the target detection object from the marching track information of the detection object included in the historical video frame according to the feature matching degree; determining a position accuracy of the target detection object in the current video frame based on position information of the target detection object predicted based on the marching track information of the target detection object; and performing statistical processing on the target detection object based on the feature matching degree and the position accuracy. The above scheme can improve the accuracy of people flow counting.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, in particular to a people flow statistical method, device and storage medium. BACKGROUND

[0002] People flow statistics is commonly used to count the number of people entering and exiting a target area, so as to realize scientific management of the target area based on people flow data, facilitate optimization of business decisions, etc.

[0003] For example, in places such as stores, shopping centers and exhibition halls, manual counting, infrared sensing, gate counting and other methods are often used to count the people flow entering and exiting the target area, but these methods have low detection accuracy and single functionality, and are difficult to meet the in-depth needs of users.

[0004] With the development of computer vision technology, currently, a detection network based on a large model can be used to track multiple targets and perform pedestrian re-identification on video data to realize people flow statistics, but the large model has a large number of parameters and a complex structure, and has high configuration requirements for the device end, and has poor universality and applicability. Therefore, there is an urgent need for a lightweight and efficient people flow statistical method. SUMMARY

[0005] The present application provides at least a people flow statistical method, device, equipment and computer readable storage medium.

[0006] The first aspect of the present application provides a people flow statistical method, comprising: obtaining each detection object in a current video frame, position information of each detection object and marching trajectory information of each detection object in a historical video frame, the time sequence of the historical video frame in a video stream being earlier than the time sequence of the current video frame in the video stream; calculating a feature matching degree between feature information of a target detection object in the current video frame and feature information of the marching trajectory information of each detection object in the historical video frame; determining marching trajectory information of the target detection object from the marching trajectory information of the detection object included in the historical video frame according to the feature matching degree; determining a position accuracy of the target detection object in the current video frame based on position information of the target detection object predicted based on the marching trajectory information of the target detection object; and performing statistical processing on the target detection object based on the feature matching degree and the position accuracy.

[0007] In an embodiment, the step of performing statistical processing on the target detection object based on the feature matching degree and the position accuracy comprises: calculating a first cost matrix based on the feature matching degree and the position accuracy; performing first Hungarian matching on each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix to obtain a detection object that successfully passes the first Hungarian matching; and counting the number of the detection object that successfully passes the first Hungarian matching into the people flow statistical result.

[0008] In an embodiment, after the step of performing first Hungarian matching on each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix to obtain a detection object that successfully passes the matching, the method further comprises: obtaining a first detection region of each detection object that fails the matching in the historical video frame and a second detection region of each detection object that fails the matching in the current video frame; performing intersection over union calculation based on the first detection region and the second detection region to obtain a second cost matrix; performing second Hungarian matching on each detection object that fails the matching in the historical video frame and each detection object that fails the matching in the current video frame based on the second cost matrix to obtain a detection object that successfully passes the second Hungarian matching; and counting the number of the detection object that successfully passes the second Hungarian matching into the people flow statistical result.

[0009] In an embodiment, after the step of performing second Hungarian matching on each detection object that fails the matching in the historical video frame and each detection object that fails the matching in the current video frame based on the second cost matrix to obtain a detection object that successfully passes the second Hungarian matching, the method further comprises: obtaining a detection object that fails the second Hungarian matching to obtain an unmatched object; initializing trajectory information of the unmatched object to obtain initialized trajectory information; performing the first Hungarian matching and / or the second Hungarian matching on the unmatched object in a subsequent video frame based on the initialized trajectory information, and recording a matching time, the time sequence of the subsequent video frame in the video stream being later than the time sequence of the current video frame in the video stream; and performing deletion processing on the unmatched object whose matching time is greater than or equal to a preset time threshold.

[0010] In an embodiment, the feature information comprises a feature vector, and the step of calculating a feature matching degree between feature information of a target detection object in the current video frame and feature information of trajectory information of each detection object in the historical video frame comprises: obtaining a historical feature vector in the trajectory information of each detection object in the historical video frame and a current feature vector of the target detection object in the current video frame; and performing cosine distance calculation on the historical feature vector and the current feature vector to obtain the feature matching degree.

[0011] In an embodiment, after the step of performing cosine distance calculation on the historical feature vector and the current feature vector to obtain the feature matching degree, the method further comprises: if the feature matching degree is greater than or equal to a preset feature matching threshold, performing weighted summation calculation based on the current feature vector and the historical feature vector to obtain a calculation result; and updating the historical feature vector in the trajectory information of the detected object in the historical video frame whose feature matching degree is greater than or equal to the feature matching threshold based on the calculation result.

[0012] In an embodiment, the step of determining the position accuracy of the target detected object in the current video frame based on the trajectory information of the target detected object comprises: performing motion prediction on each detected object in the historical video frame based on the trajectory information of each detected object in the historical video frame to obtain the predicted position information of each detected object in the current video frame; and performing intersection over union calculation on the predicted position information and the current position information of each detected object in the current video frame to obtain the position accuracy between each detected object in the historical video frame and each detected object in the current video frame.

[0013] In an embodiment, the step of performing statistical processing on the target detected object based on the feature matching degree and the position accuracy comprises: determining the target detected object that matches successfully between the current video frame and the historical video frame based on the feature matching degree and / or the position accuracy; and counting the number of target detected objects that enter a preset target region among the target detected objects that match successfully to obtain a people flow statistical result.

[0014] The second aspect of the present application provides a people flow statistical device, comprising: an acquisition module configured to acquire each detected object in a current video frame, position information of each detected object, and trajectory information of each detected object in a historical video frame, wherein the time sequence of the historical video frame in a video stream is earlier than the time sequence of the current video frame in the video stream; a feature calculation module configured to calculate a feature matching degree between feature information of a target detected object in the current video frame and feature information of the trajectory information of each detected object in the historical video frame; a trajectory determination module configured to determine trajectory information of the target detected object from the trajectory information of the detected object included in the historical video frame according to the feature matching degree; a position determination module configured to determine a position accuracy of the target detected object in the current video frame based on position information of the target detected object predicted based on the trajectory information of the target detected object; and a statistical module configured to perform statistical processing on the target detected object based on the feature matching degree and the position accuracy.

[0015] The third aspect of the present application provides an electronic device, comprising a memory and a processor, the processor being configured to execute program instructions stored in the memory to implement the above-mentioned people flow counting method.

[0016] The fourth aspect of the present application provides a computer-readable storage medium, having program instructions stored thereon, the program instructions being executed by a processor to implement the above-mentioned people flow counting method.

[0017] The above-mentioned scheme, by obtaining each detection object in the current video frame, position information of each detection object and marching trajectory information of each detection object in the historical video frame, calculating the feature similarity of each detection object between the current video frame and the historical video frame based on the feature information and the position accuracy based on the trajectory information, combining the feature similarity and the position accuracy to count the detection objects with similar features and similar positions, obtaining the counting result, thereby effectively combining the feature information and the trajectory information to count the people flow and improving the accuracy of people flow counting.

[0018] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present application. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application.

[0020] Figure 1 is a flowchart of an exemplary embodiment of the people flow counting method of the present application;

[0021] Figure 2 is an effect diagram of feature matching in the people flow counting method of the present application;

[0022] Figure 3 is an effect diagram of counting the people flow entering the target area in the people flow counting method of the present application;

[0023] Figure 4 is a block diagram of a people flow counting device according to an exemplary embodiment of the present application;

[0024] Figure 5 is a structural diagram of an embodiment of the electronic device of the present application;

[0025] Figure 6 is a structural diagram of an embodiment of the computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0026] The schemes of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0027] In the following description, for the purposes of explanation and not limitation, specific details are set forth, such as particular architectures, interfaces, techniques, etc. in order to provide a thorough understanding of the present application.

[0028] The term "and / or" herein is merely used to describe associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent three cases: A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects. In addition, "multiple" herein means two or more. In addition, the term "at least one" herein means any one of multiple or any combination of at least two of multiple, for example, including at least one of A, B, and C, which can mean including any one or more elements selected from the set consisting of A, B, and C.

[0029] It should be noted that when the crowd flow in the to-be-counted area is counted, there can be one person or multiple persons in the to-be-counted area, and the present application mainly takes the scenario of multiple persons as an example for description.

[0030] Please refer to Figure 1 , Figure 1 is a flowchart of an exemplary embodiment of the crowd flow counting method of the present application. Specifically, it can include the following steps:

[0031] In step S110, each detection object in the current video frame, the position information of each detection object, and the marching trajectory information of each detection object in the historical video frame are obtained, and the time sequence of the historical video frame in the video stream is earlier than that of the current video frame in the video stream.

[0032] There can be one historical video frame or multiple historical video frames in the video stream, which is not limited herein; the historical video frame and the current video frame belong to the same video stream, and the time sequence of the historical video frame in the same video stream is earlier than that of the current video frame, preferably, the historical video frame can be the previous video frame of the current video frame, so as to better capture and analyze the information of the detection object.

[0033] Exemplarily, the obtained video stream is input into a pre-trained target detection network model to obtain each detection object in the current video frame output by the target detection network model, position information of each detection object, and path trajectory information of each detection object in the historical video frame; wherein the target detection network model usually outputs each detection object in the form of a detection region, for example, a detection box, and one detection box usually represents one detection object; the position information of each detection object corresponding to each detection box is determined based on the position information of each detection box; for a detection object detected in multiple historical video frames, its corresponding path trajectory information changing according to the sequence of the video stream frames is obtained, and similarly, there is related information of each detection object in each historical video frame on the path trajectory information of each detection object, such as feature information and position information.

[0034] Further, the obtained video stream is input into a pre-trained target detection network model to obtain the detection box result (each detection object in the current video frame) output by the target detection network model, and the detection box result of the kth video frame in the video stream is denoted as det i (i = 1, 2, 3…N), wherein N is the total number of detection boxes in the kth video frame (equivalent to the total number of detection objects detected in the kth video frame), and in addition, each detection box also carries data such as confidence, position information of the detection object, etc.

[0035] It should be noted that before step S110, there is also a model training step, exemplarily, the obtained human target detection data set is input into a lightweight target detection network YOLOv7-Tiny for model training until the model parameters of the initial neural network model converge, to obtain the target detection network model.

[0036] Step S120, calculating the feature matching degree between the feature information of the target detection object in the current video frame and the feature information of the path trajectory information of each detection object in the historical video frame.

[0037] It should be noted that the target detection object in the current video frame refers to any one of the detection objects in the current video frame, for the sake of description, the calculation process of one detection object in the current video frame is described here, and in the specific implementation process of the method of the present application, each detection object in the current video frame is a target detection object, and the operation on the target detection object is extended to each detection object, that is, the matching processing between the feature information of each detection object (target detection object) in the current video frame and the feature information included in the path trajectory information of each detection object in the historical video frame is performed, for the sake of description, the path trajectory information is also referred to as trajectory information hereinafter.

[0038] Further explanation can be made with reference to Figure 2As shown, Figure 2 is a schematic diagram of the feature matching effect in the crowd flow counting method of the present application. The historical video frame includes detection objects A, B, and C, and the current video frame includes detection objects D, E, and F. The feature matching degree between the feature information of the target detection object D and the feature information of A, B, and C is as shown in the table Figure 2 The feature matching degree between the feature information of the target detection object E and the feature information of A, B, and C is as shown in the table Figure 2 The feature matching degree between the feature information of the target detection object F and the feature information of A, B, and C is as shown in the table Figure 2 The feature similarity between the feature information of each detection object in the current video frame and the feature information of each detection object in the historical video frame is obtained according to the data in the table.

[0039] For example, the way to extract the feature information of the target detection object in the current video frame can be, but is not limited to, extracting the appearance feature of the target detection object through a re-identification network, or extracting the key feature of the target detection object that has an identification effect through other existing technologies, etc. Taking a re-identification network model based on a lightweight OSNet network structure as an example, the current video frame is input into the re-identification network model, so that the re-identification network model performs feature extraction on the target detection object det i (detection box) detected in the current video frame, and obtains a 512-dimensional appearance feature vector output by the re-identification network model. The cosine distance cost between the appearance feature vector of each detection object in the current video frame and the appearance feature vector included in the moving track information of each detection object in the historical video frame is calculated. The mathematical expression is:

[0040]

[0041] Where j represents the number of moving tracks in the historical video frame, and the moving track is a tracking track determined by the tracking module. M is the total number of moving tracks. It should be noted that for the cosine distance, the larger the cosine distance, the greater the gap between the data. Therefore, if the cosine distance is used to represent the feature matching degree, the cosine distance is less than the preset feature matching threshold value, indicating that the appearance features of the two detection objects match.

[0042] In step S130, the moving track information of the target detection object is determined from the moving track information of the detection objects included in the historical video frame according to the feature matching degree.

[0043] Exemplarily, if the feature matching degree of a detection object in a historical video frame and a detection object in a current video frame is greater than or equal to a preset feature matching threshold, it indicates that the two detection objects are similar in appearance feature matching, and the two detection objects can be the same detection object, that is, if each detection object is provided with a unique identifier (object ID), the detection object in the historical video frame and the detection object in the current video frame with the feature matching degree greater than or equal to the feature matching threshold belong to the same object ID, and a target detection object with successful appearance feature matching is obtained; further, in the existing behavior trajectory information in the trajectory library, the behavior trajectory information corresponding to the target detection object with successful appearance feature matching in the current video frame is determined.

[0044] It should be further pointed out that after the appearance feature matching based on the detection frame is successful, the appearance feature on the behavior trajectory information corresponding to the current video frame needs to be updated, exemplarily, the exponential moving average (EMA) method can be used to perform weighted calculation based on the appearance feature of the target detection object detected in the current video frame and the appearance feature on the behavior trajectory information of the historical key frame, and the appearance feature on the trajectory information corresponding to the current video frame is updated, so as to continue to track the target detection object based on the updated appearance feature on the trajectory information, and realize the extension of the trajectory information, which is mathematically expressed as:

[0045]

[0046] wherein, which represents the appearance feature vector of the t-th detection frame (target detection object) on the k-th frame (current video frame) of the behavior trajectory, which represents the appearance feature vector of the t-th detection frame on the k-1-th frame (historical video frame) of the behavior trajectory, which represents the appearance feature vector extracted from the t-th detection frame of the k-th frame, and a is a preset weight hyperparameter, which can be set to a = 0.9 in the embodiment.

[0047] In step S140, the position accuracy of the target detection object in the current video frame is determined based on the position information of the target detection object predicted based on the behavior trajectory information of the target detection object.

[0048] In combination with the above steps, after comparing the appearance features of the detection objects in the two video frames, the motion information of the detection objects in the two video frames can be further analyzed to improve the accuracy of target detection and target tracking.

[0049] Exemplarily, based on the trajectory information (such as the position, direction, pose, speed, etc. of the target detection object in the historical video frame) of the target detection object obtained from the historical video frame, motion estimation is performed in the manner of NSA Kalman filtering to obtain the predicted position information (predicted box) of the target detection object in the current video frame output by the Kalman filter; the predicted box and the actual position information (actual box) of the target detection object actually detected in the current video frame are subjected to Intersection over Union (IoU) calculation to obtain the motion distance cost of the target detection object from the historical video frame to the current video frame that is, the position accuracy, which is mathematically expressed as:

[0050]

[0051] It can be understood that if the position accuracy is greater than or equal to the preset accuracy threshold, it means that the predicted position information of the target detection object in the current video frame obtained based on the historical position information of the detection object in the historical video frame for motion estimation is close to or the same as the actual position information of the target detection object in the current video frame, and the detection object with the position accuracy greater than or equal to the accuracy threshold in the two video frames is the same detection object, and the motion information of the detection object in the two video frames is matched. It should be noted that for the motion distance cost, the greater the IoU between the two detection boxes i,j the smaller the distance between the two detection boxes, and therefore, if the position accuracy is represented by the motion distance cost, only when the position accuracy is less than the preset accuracy threshold, it means that the motion information of the two detection objects is matched.

[0052] Optionally, in addition to the above-mentioned Intersection over Union between the predicted box of each detection object in the historical video frame in the current video frame and the actual box actually detected in the current video frame to determine whether each detection object is matched in motion information, the motion estimation process can also be omitted. Since the moving speed of people is usually slow during the counting of the crowd, the detection boxes of the same person in the two video frames will not have a large difference in distance, and therefore, the Intersection over Union between the detection boxes of each detection object in the historical video frame and the detection boxes of each detection object in the current video can also be calculated to determine the detection boxes and the corresponding trajectory information that are matched in the two video frames, thereby improving the data processing efficiency.

[0053] Further, after the position accuracy is matched successfully, the noise covariance is adaptively calculated according to the confidence of the detection box of the target detection object in the current video frame and the noise covariance in the Kalman filter is updated, so as to perform trajectory prediction of the next video frame of the current video frame, which is mathematically expressed as:

[0054]

[0055] wherein, is the updated noise covariance, R k is the preset constant measurement noise covariance, c k is the detection frame confidence of the kth frame matching.

[0056] It should be noted that after the position accuracy (motion distance cost) of the detection frame of each detection object in the historical video frame and the current video frame is obtained, the cosine distance cost between the appearance features can be optimized based on the motion distance cost to improve the accuracy of the people flow counting process. Specifically, the mathematical expression of the optimized cosine distance cost is:

[0057]

[0058] wherein, θ cos is the preset cosine distance cost threshold (feature matching threshold), θ iou is the preset motion distance cost threshold (accuracy threshold). In this embodiment, θ cos = 0.25, θ iou = 0.5. After the cosine distance cost is optimized, the cosine distance between the detection frames with low appearance feature matching degree and low position accuracy will be set to 1 to reject matching association, while the detection frames with high appearance feature matching degree and high position accuracy will be given a lower cosine distance cost for matching, thereby improving the matching association ability of the feature matching method and the motion trajectory matching method for the detection frames in the video stream.

[0059] Step S150, counting and processing the target detection object based on the feature matching degree and the position accuracy.

[0060] In combination with the above steps, the optimized cosine distance cost and the motion distance cost are used as the keys to construct the first cost matrix CostMatrix1, the mathematical expression of which is:

[0061]

[0062] Hungarian matching is performed on the target detection object and each detection object in the historical video frame based on the first cost matrix until all the detection objects in the current video frame are matched, obtaining the target detection object the target detection object and the corresponding track information of the detection object in the historical video frame the target detection object The number of the target detection objects that are successfully matched is counted into the people flow statistical result.

[0063] It can be seen that, by acquiring each detection object in the current video frame, position information of each detection object, and track information of each detection object in the historical video frame, the feature similarity of each detection object between the current video frame and the historical video frame based on feature information and the position accuracy based on track information are calculated, the feature similarity and the position accuracy are combined to count the detection objects with similar features and similar positions, and the statistical result is obtained, thereby the feature information and the track information can be effectively combined to count the people flow, and the accuracy of the people flow counting is improved.

[0064] Based on the above embodiments, the embodiments of the present application explain the step of counting the target detection objects based on the feature matching degree and the position accuracy. Specifically, the method of the present embodiment includes the following steps:

[0065] The first cost matrix is calculated based on the feature matching degree and the position accuracy; the first Hungarian matching is performed between each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix, to obtain the detection objects that are successfully matched by the first Hungarian matching; and the number of the detection objects that are successfully matched by the first Hungarian matching is counted into the people flow statistical result.

[0066] In combination with the foregoing embodiments, after the first cost matrix is constructed based on the feature matching degree and the position accuracy, the first Hungarian matching is performed between the detection boxes of each detection object in the historical video frame and the detection boxes of each detection object in the current video frame based on the first cost matrix, where the principle of the Hungarian matching is to find the maximum matching by augmenting the path, to ensure that the maximum number of matching between the two ends of the data participating in the matching process can be obtained under the condition of meeting the matching condition; and then the number of the detection objects that are successfully matched by the first Hungarian matching is counted into the people flow statistical result.

[0067] Based on the above embodiments, the embodiments of the present application explain the step after the first Hungarian matching is performed between each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix, to obtain the detection objects that are successfully matched. Specifically, the method of the present embodiment includes the following steps:

[0068] obtain a first detection region of each detection object that fails to match in the historical video frame and a second detection region of each detection object that fails to match in the current video frame; perform an intersection over union calculation based on the first detection region and the second detection region to obtain a second cost matrix; perform a second Hungarian matching on each detection object that fails to match in the historical video frame and each detection object that fails to match in the current video frame based on the second cost matrix to obtain a detection object that succeeds in the second Hungarian matching; and count a number of the detection object that succeeds in the second Hungarian matching into the people flow statistical result.

[0069] In combination with the foregoing embodiments, after the first Hungarian matching is performed on the detection objects in the historical video frame and the current video frame, there can be detection objects that fail to match, which can be caused by the fact that the same detection object exhibits different feature information in the historical video frame and the current video frame, for example, when the facial features of the detection object are collected, the angles at which the facial features of the detection object are collected in the historical video frame and the current video frame are different, so that the feature matching degree in the two video frames is low, resulting in the first Hungarian matching failure, but it is still the same detection object; or the detection object appears for the first time in the historical video frame, and the trajectory information thereof has not been confirmed, which also results in the matching failure.

[0070] Therefore, for the detection objects in the historical video frame and the detection objects in the current video frame that fail to match in the first Hungarian matching, the IoU values of the detection boxes between the two video frames are calculated one by one based on the detection boxes of the detection objects in the current video frame and the detection boxes of the detection objects in the historical video frame, as the second cost matrix CostMatrix2; further, the second Hungarian matching is performed on the detection objects in the historical video frame and the detection objects in the current video frame that fail to match in the first Hungarian matching based on the second cost matrix, to obtain a detection object that succeeds in the second Hungarian matching, and the number of the detection object that succeeds in the second Hungarian matching is counted into the people flow statistical result.

[0071] Thus, the number of the detection objects that succeed in the first Hungarian matching and the number of the detection objects that succeed in the second Hungarian matching are jointly counted into the people flow statistical result, high-precision people flow statistics are achieved, and false detection and missed detection of the people flow are avoided.

[0072] On the basis of the foregoing embodiments, the steps after the second Hungarian matching is performed on each detection object that fails to match in the historical video frame and each detection object that fails to match in the current video frame based on the second cost matrix to obtain a detection object that succeeds in the second Hungarian matching are described. Specifically, the method of the embodiment includes the following steps:

[0073] obtain an un-matched object; initialize the moving track information of the un-matched object to obtain initialized moving track information; perform the first Hungarian matching and / or the second Hungarian matching on the un-matched object in subsequent video frames based on the initialized moving track information, and record a matching time, the time sequence of the subsequent video frames in the video stream being later than the time sequence of the current video frame in the video stream; and perform deletion processing on the un-matched object whose matching time is greater than or equal to a preset time threshold.

[0074] In combination with the foregoing embodiments, there can still be matched detection objects and un-matched detection objects after the second Hungarian matching. For the un-matched detection objects of the second Hungarian matching, there are false detection objects and / or detection objects that first appear in the current video frame, that is, the detection objects do not exist in the historical video frames, and there is no corresponding moving track information and feature information on the moving track information. Therefore, such detection objects do not satisfy the first Hungarian matching and the second Hungarian matching. Such detection objects are recorded as un-matched objects, and the track information of the un-matched objects is initialized so that the initialized track information can be used for the feature information, motion information, etc. of the un-matched detection objects provided in subsequent video frames. The subsequent video frames include one video frame or multiple video frames, the subsequent video frames and the current video frame belong to the same video stream, and the time sequence of the subsequent video frames in the video stream is later than the time sequence of the current video frame in the video stream. If the same un-matched object is detected in continuous multiple video frames (for example, 3 frames), it is considered that the un-matched object is not false detection, but first appears in the previous video frame. The un-matched object and its corresponding track are valid data, the track information of the un-matched object is confirmed, and conversely, the unconfirmed track information and the un-matched object can be false detection, and are deleted. The confirmed track information and the un-matched object are subjected to the first Hungarian matching and / or the second Hungarian matching according to the implementation manner in the foregoing embodiments, and the matching time of each un-matched object is recorded. For the un-matched object whose matching time is greater than or equal to a preset time threshold (for example, 30 continuous video frames), it is considered that the un-matched object still does not satisfy the matching condition, and is deleted. For the un-matched object whose matching is successful within the time threshold, the number of the un-matched object is counted into the people flow statistical result.

[0075] On the basis of the foregoing embodiments, the step of calculating the feature matching degree between the feature information of the target detection object in the current video frame and the feature information of the moving track information of each detection object in the historical video frame is described in the embodiments of the present application, wherein the feature information includes a feature vector. Specifically, the method of the present embodiment includes the following steps:

[0076] The historical feature vector in the trajectory information of each detected object in the historical video frame and the current feature vector of the target detected object in the current video frame are obtained; and the cosine distance of the historical feature vector and the current feature vector is calculated to obtain a feature matching degree.

[0077] With the foregoing embodiments, the historical feature vector of the detection frame included in the trajectory information of each detected object in the historical video frame and the current feature vector of the detection frame of each detected object in the current video frame are used to perform cosine distance calculation to obtain the feature matching degree of each detection frame between the historical video frame and the current video frame. The cosine distance can also be referred to as cosine similarity. The cosine value of the angle between the feature vectors of two detection frames is used as a measure of the difference between the two detection frames. When the cosine similarity is 1, it indicates that the directions of the two feature vectors are opposite, that is, the two feature vectors point to different directions and are irrelevant, and the feature matching degree is also 0. The features of the detection frames corresponding to the two feature vectors are not matched.

[0078] On the basis of the foregoing embodiments, the steps after the cosine distance calculation of the historical feature vector and the current feature vector to obtain the feature matching degree are described. Specifically, the method of the embodiment includes the following steps:

[0079] If the feature matching degree is greater than or equal to a preset feature matching threshold, a weighted sum calculation is performed based on the current feature vector and the historical feature vector to obtain a calculation result; and the historical feature vector in the trajectory information of the detected object in the historical video frame with the feature matching degree greater than or equal to the feature matching threshold is updated based on the calculation result.

[0080] With the foregoing embodiments, taking a detected object in the historical video frame and a detected object in the current video frame as an example, if the feature matching degree between the feature vectors of the two detected objects is greater than or equal to a preset feature matching threshold, that is, the cosine distance of the two feature vectors is less than a preset cosine distance cost threshold, it is considered that the two feature vectors are similar and the two detected objects can be identified as the same detected object in terms of appearance features. A weighted sum calculation is performed based on the current feature vector in the current video frame of the detected object and the historical feature vector in the historical video frame to obtain a calculation result, so as to update the feature information corresponding to the trajectory information of the detected object, facilitating subsequent target tracking detection and motion estimation.

[0081] On the basis of the foregoing embodiments, the steps of determining the position accuracy of the target detected object in the current video frame based on the position information of the target detected object predicted from the trajectory information of the target detected object are described. Specifically, the method of the embodiment includes the following steps:

[0082] The motion of each detection object in the historical video frame is predicted based on the trajectory information of each detection object in the historical video frame, to obtain the predicted position information of each detection object in the historical video frame in the current video frame; and the intersection-over-union calculation is performed on the predicted position information and the current position information of each detection object in the current video frame that is detected, to obtain the position accuracy between each detection object in the historical video frame and each detection object in the current video frame.

[0083] In combination with the foregoing embodiments, the motion estimation is performed in the manner of NSA Kalman filtering based on the trajectory information (such as the position, direction, pose, speed, etc. of the target detection object in the historical video frame) of the target detection object obtained from the historical video frame, to obtain the predicted position information (predicted box) of the target detection object in the current video frame output by the Kalman filter; the intersection-over-union calculation is performed on the predicted box and the actual position information (actual box) of the target detection object actually detected in the current video frame, to obtain the motion distance cost (position accuracy) of the target detection object from the historical video frame to the current video frame.

[0084] Optionally, in addition to the intersection-over-union between the predicted box of each detection object in the historical video frame in the current video frame and the actual box actually detected in the current video frame to determine whether each detection object matches in the motion information, the process of motion estimation can also be omitted; since the moving speed of people is generally slow in the process of counting the crowd, the detection boxes of the same person in two video frames will not have a large difference in distance, and thus the intersection-over-union calculation can be directly performed on the detection boxes of each detection object in the historical video frame and the detection boxes of each detection object in the current video, to determine the detection boxes and the corresponding trajectory information that match each other in the two video frames, so as to improve the data processing efficiency.

[0085] Based on the foregoing embodiments, the present embodiment describes the step of counting and processing the target detection object based on the feature matching degree and the position accuracy. Specifically, the method of the present embodiment comprises the following steps:

[0086] The target detection object that matches successfully between the current video frame and the historical video frame is determined based on the feature matching degree and / or the position accuracy; and the number of target detection objects that enter the preset target region in the target detection objects that match successfully is counted, to obtain the crowd counting result.

[0087] Reference can be made to Figure 3 , as shown in the following figure: Figure 3 is an effect diagram of counting the crowd entering the target region in the crowd counting method of the present application, wherein the target region includes but is not limited to Figure 3The rectangular region or horizontal line region shown in the middle, etc.; the human target in the video frame is detected and the motion trajectory state is continuously tracked, if the detected human target enters the target region, the statistical processing is performed, and the people flow statistical result is obtained.

[0088] Exemplarily, the moving direction of the detection object is acquired, if the detection object enters the target region, the people flow statistical result is increased, and if the detection object leaves the target region, the people flow statistical result is reduced.

[0089] It should be further explained that the execution subject of the people flow statistical method can be a people flow statistical device, for example, the people flow statistical method can be executed by a terminal device or a server or other processing device, wherein the terminal device can be a user equipment (User Equipment, UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital processing (Personal Digital Assistant, PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the people flow statistical method can be realized by a processor calling computer readable instructions stored in a memory.

[0090] Figure 4 is a block diagram of a people flow statistical device shown in an exemplary embodiment of the present application. As Figure 4 shown, the exemplary people flow statistical device 400 includes an acquisition module 410, a feature calculation module 420, a trajectory determination module 430, a position determination module 440, and a statistical module 450. Specifically:

[0091] The acquisition module 410 is configured to acquire each detection object in a current video frame, position information of each detection object, and marching trajectory information of each detection object in a historical video frame, the time sequence of the historical video frame in a video stream being earlier than the time sequence of the current video frame in the video stream.

[0092] The feature calculation module 420 is configured to calculate a feature matching degree between feature information of a target detection object in the current video frame and feature information of the marching trajectory information of each detection object in the historical video frame.

[0093] The trajectory determination module 430 is configured to determine the marching trajectory information of the target detection object from the marching trajectory information of the detection object included in the historical video frame according to the feature matching degree.

[0094] The position determination module 440 is configured to determine a position accuracy of the target detection object in the current video frame based on the position information of the target detection object predicted by the marching trajectory information of the target detection object.

[0095] The statistical module 450 is configured to statistically process the target detection object based on the feature matching degree and the position accuracy.

[0096] In the example crowd flow statistical device, by obtaining each detection object in a current video frame, position information of each detection object, and track information of each detection object in a historical video frame, the feature similarity of each detection object between the current video frame and the historical video frame based on feature information and the position accuracy based on track information are calculated, the feature similarity and the position accuracy are combined, the detection objects with similar features and similar positions are counted, and a statistical result is obtained, so that the feature information and the track information are effectively combined to count the crowd flow, and the accuracy of crowd flow counting is improved.

[0097] The functions of the modules can be referred to the embodiments of the crowd flow statistical method, which will not be repeated here.

[0098] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of an embodiment of an electronic device. The electronic device 500 includes a memory 501 and a processor 502. The processor 502 is configured to execute program instructions stored in the memory 501 to implement the steps in any of the above crowd flow statistical method embodiments. In a specific implementation scenario, the electronic device 500 can include but is not limited to a microcomputer, a server, and in addition, the electronic device 500 can also include a notebook computer, a tablet computer, and the like, without limitation.

[0099] Specifically, the processor 502 is configured to control itself and the memory 501 to implement the steps in any of the above crowd flow statistical method embodiments. The processor 502 can also be referred to as a CPU (Central Processing Unit). The processor 502 can be an integrated circuit chip with processing capability. The processor 502 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 502 can be implemented by an integrated circuit chip.

[0100] The scheme can effectively combine the feature information and the trajectory information to count the pedestrian flow, and improve the accuracy of the pedestrian flow counting.

[0101] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of an embodiment of the computer readable storage medium. The computer readable storage medium 610 stores program instructions 611 capable of being executed by a processor, and the program instructions 611 are used to implement the steps in any of the above pedestrian flow counting method embodiments.

[0102] The scheme can effectively combine the feature information and the trajectory information to count the pedestrian flow, and improve the accuracy of the pedestrian flow counting.

[0103] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0104] The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be mutually referred to. For the sake of brevity, it will not be repeated here.

[0105] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the above-described device implementation is only schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, another division mode can be used. For example, a unit or component can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0106] In addition, each of the function units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit. When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

Claims

1. A method of counting the number of people, characterized by, The method comprises: obtaining each detection object in a current video frame, position information of each detection object, and trajectory information of each detection object in a historical video frame, the time sequence of the historical video frame in a video stream being earlier than the time sequence of the current video frame in the video stream; calculating a feature matching degree between feature information of a target detection object in the current video frame and feature information of the trajectory information of each detection object in the historical video frame; determining trajectory information of the target detection object from the trajectory information of the detection object included in the historical video frame according to the feature matching degree; determining a position accuracy of the target detection object in the current video frame based on position information of the target detection object predicted by the trajectory information of the target detection object; optimizing the feature matching degree according to the position accuracy and the feature matching degree, the optimization process comprising: if the feature matching degree is less than a feature matching threshold and the position accuracy is less than an accuracy threshold, reducing the feature matching degree to obtain an optimized feature matching degree; otherwise, setting the feature matching degree to 1 to obtain an optimized feature matching degree; statistically processing the target detection object based on the optimized feature matching degree and the position accuracy; The step of statistically processing the target detection object based on the optimized feature matching degree and the position accuracy comprises: constructing a first cost matrix based on the minimum value of the optimized feature matching degree and the position accuracy; performing first Hungarian matching on each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix to obtain a detection object that successfully passes the first Hungarian matching; and counting the number of the detection object that successfully passes the first Hungarian matching into a people flow statistical result.

2. The method of claim 1, wherein, After the step of performing first Hungarian matching on each detection object in the historical video frame and each detection object in the current video frame based on the first cost matrix to obtain a detection object that successfully passes the matching, the method further comprises: obtaining a first detection area of each detection object that fails the matching in the historical video frame and a second detection area of each detection object that fails the matching in the current video frame; performing intersection over union calculation based on the first detection area and the second detection area to obtain a second cost matrix; performing second Hungarian matching on each detection object that fails the matching in the historical video frame and each detection object that fails the matching in the current video frame based on the second cost matrix to obtain a detection object that successfully passes the second Hungarian matching; counting the number of the detection object that successfully passes the second Hungarian matching into the people flow statistical result.

3. The method of claim 2, wherein, After the step of performing second Hungarian matching on each detection object that fails the matching in the historical video frame and each detection object that fails the matching in the current video frame based on the second cost matrix to obtain a detection object that successfully passes the second Hungarian matching, the method further comprises: obtaining a detection object that fails the matching in the second Hungarian matching to obtain an unmatched object; Initialize the moving track information of the unmatched object to obtain initialized moving track information; perform the first Hungarian matching and / or the second Hungarian matching on the unmatched object in a subsequent video frame based on the initialized moving track information, and record a matching time, the subsequent video frame being later in time sequence in the video stream than the current video frame; perform deletion processing on the unmatched object whose matching time is greater than or equal to a preset time threshold.

4. The method of claim 1, wherein, The feature information includes a feature vector, and the step of calculating the feature matching degree between the feature information of the target detection object in the current video frame and the feature information of the moving track information of each detection object in the historical video frame includes: obtaining a historical feature vector in the moving track information of each detection object in the historical video frame and a current feature vector of the target detection object in the current video frame; performing cosine distance calculation on the historical feature vector and the current feature vector to obtain the feature matching degree.

5. The method of claim 4, wherein, After the step of performing cosine distance calculation on the historical feature vector and the current feature vector to obtain the feature matching degree, the method further includes: if the feature matching degree is greater than or equal to a preset feature matching threshold, performing weighted summation calculation based on the current feature vector and the historical feature vector to obtain a calculation result; updating the historical feature vector in the moving track information of the detection object in the historical video frame whose feature matching degree is greater than or equal to the feature matching threshold based on the calculation result.

6. The method of claim 1, wherein, The step of determining the position accuracy of the target detection object in the current video frame based on the position information of the target detection object predicted based on the moving track information of the target detection object includes: performing motion prediction on each detection object in the historical video frame based on the moving track information of each detection object in the historical video frame to obtain predicted position information of each detection object in the historical video frame in the current video frame; performing intersection over union calculation on the predicted position information and the current position information of each detection object in the current video frame to obtain the position accuracy between each detection object in the historical video frame and each detection object in the current video frame.

7. The method of claim 1, wherein, The step of performing statistical processing on the target detection object based on the feature matching degree and the position accuracy includes: determining a target detection object that matches successfully between the current video frame and the historical video frame based on the feature matching degree and / or the position accuracy; counting the number of target detection objects that enter a preset target region in the target detection objects that match successfully to obtain a people flow statistical result.

8. An electronic device, comprising: The program instructions are executed by the processor to implement the method of any one of claims 1-7.

9. A computer-readable storage medium having stored thereon program instructions, wherein, The program instructions are executed by the processor to implement the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Pedestrian flow statistical method based on multi-target tracking

    CN115410155A