A passenger flow statistical method and device, computer equipment and storage medium
By performing cluster analysis and density clustering on the captured data, the problem of large deviations in the statistical results of dense passenger flow in existing technologies has been solved, and more accurate passenger flow statistics and recommendations have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
- Filing Date
- 2022-12-29
- Publication Date
- 2026-05-15
AI Technical Summary
Existing customer flow statistics methods suffer from significant biases and subjectivity in scenarios with high customer traffic, failing to accurately reflect irregular consumer behavior, resulting in incomplete statistics and inaccurate recommendations.
By acquiring snapshot data of the target area, cluster analysis and archiving are performed. An archived data set is formed based on image similarity. The density clustering algorithm DBSCAN is used to calculate the passenger flow popularity level, forming multiple data clusters with continuous time periods to determine the passenger flow popularity level.
It enables more accurate customer traffic statistics and recommendations, and can perform reasonable statistics based on irregular consumer behavior, thereby improving the accuracy of statistical results and the precision of recommendations.
Smart Images

Figure CN116244609B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of passenger flow statistics, and in particular to a passenger flow statistics method, apparatus, computer equipment, and storage medium. Background Technology
[0002] With the rapid development of the Internet and information technology, customer flow statistics are widely used in various scenarios such as smart retail and security monitoring. How to use customer flow statistics to help consumers more accurately understand the popularity of the store and make ranking recommendations for similar stores and their consumer products has become a research hotspot.
[0003] Existing customer flow statistics solutions rely on pre-defined categories within a specific scenario, analyzing customer traffic within predetermined timeframes (morning, noon, afternoon, etc.). However, these solutions often only provide general estimates of a limited number of parameters, are subjective, and cannot provide more accurate statistics and rankings based on irregular consumer behavior. Furthermore, when counting large, dense customer flows, conventional pre-defined categorization methods result in highly subjective and incomplete statistics, making them suitable only for locations with relatively stable foot traffic. Therefore, accurately measuring the number of dense customer flows in stores or shops, and addressing the limitations of existing technologies in small, fixed-scenario locations where daily evaluation indicators are highly sensitive to environmental and system parameters, has become a pressing technical problem for those skilled in the art. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method, apparatus, computer equipment, and storage medium for passenger flow statistics to solve the problem of deviation when conducting dense passenger flow statistics.
[0005] In a first aspect, embodiments of the present invention provide a method for passenger flow statistics, including:
[0006] Acquire capture data of the target area at various time periods, wherein the capture data includes capture image features and capture data time;
[0007] Cluster analysis is performed on the captured data to form an archive set; based on image similarity, the captured data in the archive set is continuously archived at different time periods using the same camera to form a corresponding archived data set;
[0008] The captured data for each time period in the archived data set is sorted to obtain the time density information of the captured data;
[0009] Density clustering is performed based on the density information of the captured data time to form multiple clustering results of data clusters with continuous time periods;
[0010] Based on the clustering results of the multiple data clusters with continuous time periods, the passenger flow popularity level of the target area in the corresponding time period is determined.
[0011] Secondly, embodiments of the present invention provide a passenger flow counting device, comprising:
[0012] The data acquisition module is used to acquire capture data of the target area in various time periods, wherein the capture data includes capture image features and capture data time.
[0013] The data archiving module is used to perform cluster analysis on the captured data to form an archive set; based on image similarity, the captured data in the archive set is continuously archived at different time periods of the same camera to form a corresponding archived data set;
[0014] The data sorting module is used to sort the captured data for each time period in the archived data set and obtain the time density information of the captured data.
[0015] The data density clustering processing module is used to perform density clustering processing based on the density information of the time of the captured data, forming multiple clustering results of data clusters with continuous time periods;
[0016] The data determination module is used to determine the passenger flow popularity level of the target area in the corresponding time period based on the clustering results of the multiple data clusters with continuous time periods.
[0017] Thirdly, embodiments of the present invention provide a computer device, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the passenger flow method described in the first aspect above.
[0018] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the passenger flow method described in the first aspect.
[0019] The advantages of this invention compared to the prior art are:
[0020] This invention provides a method, apparatus, computer device, and storage medium for passenger flow statistics. It acquires captured data of a target area at various time periods. The captured data includes captured image features and the capture time. Cluster analysis is performed on the captured data to form an archive set. Based on image similarity, captured data from the archive set is continuously archived at different time periods using the same camera to form corresponding archived data sets. The captured data in the archived data sets is sorted to obtain the density information of the capture data time. Density clustering is performed based on the density information of the capture data time to form multiple clusters of data with continuous time periods. Based on the clusters of data with continuous time periods, the passenger flow intensity level of the target area in the corresponding time period is determined. In this invention, after capturing passenger flow data at various time periods using a camera, the captured information is clustered and then archived to form corresponding archived data sets. This ensures high image and angle quality of the captured data in each archived data set and improves the efficiency of subsequent data processing. After determining the density information of the captured data time, the density-based DBSCAN clustering algorithm is used to calculate the data information of multiple continuous time periods. This enables more reasonable customer flow statistics based on the irregular behavior of consumers, more accurate judgment of the popularity level of stores at different times of the day, and ranking and recommendation of similar stores and their consumer products based on the actual situation. This effectively enhances the method of customer flow statistics and improves the accuracy of recommendations. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of an application environment for a passenger flow statistics method provided in an embodiment of the present invention;
[0023] Figure 2 This is a flowchart illustrating a passenger flow statistics method according to an embodiment of the present invention;
[0024] Figure 3 This is a flowchart illustrating a passenger flow density clustering algorithm provided in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of a passenger flow statistics method and apparatus provided in an embodiment of the present invention;
[0026] Figure 5This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0028] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0029] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0030] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0031] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0032] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0033] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0034] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0035] The passenger flow statistics method provided in Embodiment 1 of this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. The client includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0036] See Figure 2 This is a flowchart illustrating a passenger flow statistics method provided in Embodiment 1 of the present invention. The above-described passenger flow statistics method can be applied to… Figure 1 The server-side component is used to input the statistical data of captured passenger flow into the corresponding processing model, such as... Figure 2 As shown, this passenger flow statistics method may include the following steps:
[0037] S110: Acquire capture data of the target area in various time periods, wherein the capture data includes capture image features and capture data time;
[0038] In step S110, the target area is a store or shop. Cameras pre-installed in various corners of the store or shop can be used to capture images of the target object at different times of the day; and the same camera positions can be used to capture images of the target object entering the store or shop. Multiple cameras are installed in various corners of the store or shop according to a preset deployment method, and then the target object is captured using these cameras. The target object can be a person or other non-human objects; this embodiment does not impose specific limitations.
[0039] It should be noted that the image features of the captured target object can be positive or negative, and can be clear or blurry; no specific limitations are made in this embodiment.
[0040] Since the camera is constantly capturing images, a large amount of capture data will be obtained. At this point, it is necessary to separate the capture data containing the target object to obtain the target capture data. The target capture data includes the captured image features, the camera location, and the capture time. Specifically, the target capture data includes the capture time of the camera capturing the target object at each camera location and the correspondence between these camera locations. For example, consider camera locations C1 and C2. If the camera at location C1 captures the target object at times T1 and T2, and the camera at location C2 captures the target object at times T3 and T4, then the target capture data is: (C1, T1); (C1, T2); (C2, T3); (C2, T4). Among them, the captured image features are derived from the fact that the quality of the multiple images obtained by the camera during the capture process varies because it only captures images according to preset instructions or randomly. Therefore, the captured image features contain features from multiple images, from which the image with the best quality can be selected for subsequent analysis. The capture data time is either when the camera captures the target object at times T1 and T2, or when it captures the target object at times T3 and T4.
[0041] It should be noted that, based on the target capture data obtained through the above steps, the target object was captured 15 times at capture camera point C1. However, not all of these 15 captures may be of the target object; it's possible that the target object was captured by the camera at capture point C1 while passing by it. Therefore, in this embodiment, the core of determining the number of times each capture point is counted as a target object based on the target capture data is: the target object's dwell time at the capture point exceeds its landing time, where the landing time is a predefined duration representing the shortest dwell time of the object at a certain capture point. Furthermore, in this example, capture camera points C1 and C2 and capture data times T1, T2, T3, and T4 could also be capture camera points C3 and C4, etc., and capture data times T5, T6, T7, and T8, etc. This embodiment does not impose any limitations.
[0042] S120: Perform cluster analysis on the captured data to form an archive set; based on image similarity, continuously archive the captured data in the archive set under different time periods of the same camera to form a corresponding archived data set;
[0043] In step S120, since the target capture data containing the target object captured by different camera devices, or the target capture data containing the target object captured by the same device at different time periods, have inconsistent image quality, and for poor quality face capture images, such as face capture images that are side-view, even if the face capture image is archived, it has little facial feature information and no application value. Therefore, optionally, the image quality and angle can be compared and judged first, data clustering analysis can be performed, and the clustered data obtained after the clustering analysis can be used as the archive set.
[0044] In this embodiment of the invention, the target object is determined based on the captured data, and first captured information of the target object at different time periods is obtained. Based on the first captured information of the target object at different time periods, the first captured information is analyzed to obtain second captured information of companions of the target object at different time periods. Companions are people other than the target object appearing in captured images t seconds before and after the captured object's capture data time point. The capture data time point refers to different time points when the camera captures the target object. Cluster analysis is performed on the first and second captured information data. Based on the image quality and angle quality of the target object, the captured data in each cluster is filtered, and the filtered clusters are used as an archive set.
[0045] Therefore, by filtering the captured target objects, images with low quality scores can be effectively avoided. This considers not only image quality but also the quality of the facial angles captured within the image. This prevents situations where similarity scores are too close, leading to low accuracy in the generated archives, thus improving the accuracy of archive generation. This ensures the accuracy of subsequent operations based on the archived files.
[0046] Here, the above-mentioned determination based on image similarity is based on multiple biometric features extracted from each piece of data to be archived in the first data archive set under different time periods; wherein, the multiple biometric features include, but are not limited to, at least one of the following: facial features, gait features, body features, head and shoulder features, attribute features (such as clothing attributes, age attributes, gender attributes, etc.); the specific types and number of features referred to by the above-mentioned multiple biometric features are not limited in this application embodiment.
[0047] Specifically, when calculating the data similarity between different data to be archived within each first data archive set, for any two different data to be archived, multiple biometric features can be extracted from each data to be archived. Then, based on the weight coefficient corresponding to each biometric feature and the single feature similarity of the two different data to be archived under the same biometric feature, the similarity between the extracted data to be archived and the temporary archived data images in the first data archive set is calculated by setting a preset similarity threshold. If the similarity is greater than or equal to the preset similarity threshold, the new data to be archived is added to the first data archive set, updating the data in the first data archive set. That is, when the captured data is greater than or equal to the preset similarity threshold, the captured data is assigned to the corresponding archive set; when the captured data is less than the preset similarity threshold, a new archive set is created in the archive set, and the captured data is assigned to the new archive set, so as to continuously archive the captured data to form the corresponding archived data set.
[0048] For example, for each captured face image, the similarity between the newly captured face image and the current images in the archive can be calculated. If the similarity is greater than or equal to a similarity threshold, the newly captured face image can be added to the archive, increasing the number of images in the archive by one. If the similarity is less than the similarity threshold, the newly captured face image can be added to a new archive, increasing the number of archives. Optionally, after adding an image, it can be further determined whether to update the corresponding image.
[0049] As can be seen, in this example, after capturing images of passenger flow at various time periods using a camera, the captured information is clustered and analyzed before being archived to form corresponding archived data sets. This not only ensures high image and angle quality of the captured data in each archived data set, but also makes subsequent data processing more efficient, speeding up passenger flow statistics and improving accuracy.
[0050] It should be noted that the specific value of the above-mentioned preset similarity threshold can be set according to the user's actual archiving needs. This application embodiment does not impose any limitation on the specific value of the above-mentioned preset similarity threshold.
[0051] S130: Sort the captured data for each time period in the archived data set to obtain the density information of the captured data time.
[0052] In step S130, after determining the similarity of the captured data in each archived data set, the other captured data can be sorted according to the similarity between each data to be captured and the data to be captured, in descending order of similarity. That is, in this sorting, the higher the similarity with the data to be captured, the higher the ranking.
[0053] Specifically, for each time period of captured data in each archived dataset, neighboring data can be sorted based on their similarity to the data to be captured in each time period. Then, based on the determined sorting of neighboring data, the weight of each neighboring data for different time periods is determined. This weight is related to the ranking of the neighboring data within the determined sorting.
[0054] For example, if the number of neighboring data is 5, and the determined sorting is a sequence of neighboring data from highest to lowest similarity, the preset weights for different time periods of each neighboring data could be 0.9, 0.7, 0.5, 0.3, and 0.1. If the determined sorting is a sequence of neighboring data from lowest to highest similarity, the preset weights for different time periods of each neighboring data could be 0.1, 0.3, 0.5, 0.7, and 0.9. After determining the weights for different time periods of each neighboring data, the time density information of the data to be captured can be determined based on the similarity between the data to be captured and the data to be captured in each time period, as well as the weights of each neighboring data in different time periods.
[0055] It should be noted that the weights of the neighboring data at different time periods can be preset manually, or determined by a weighting function based on the ranking of the neighboring data at different time periods. This weighting function can be a power function, exponential function, logarithmic function, etc., and it is a monotonic function within the preset number of neighboring data (1 to k). Furthermore, when the determined sequence of neighboring data is sorted from largest to smallest similarity, it is a monotonically decreasing function; when the determined sequence is sorted from smallest to largest similarity, it is a monotonically decreasing function.
[0056] When sorting the sequences of neighboring data, a monotonically increasing function can be used. The specific content of the weight determination function can be set in various ways as needed, and this application embodiment does not impose any limitations on it.
[0057] In this embodiment, the captured data is sorted according to similarity to determine the weight of the captured data in different time periods, thus enabling the final determination of the density information of the captured data time. This ensures the accuracy of dense passenger flow information and improves the speed of passenger flow statistics.
[0058] S140: Perform density clustering processing based on the density information of the captured data time to form multiple clustering results of data clusters with continuous time periods;
[0059] In step S140, after determining the density information of the captured data time, the advantages of DBSCAN (density clustering) are utilized to perform clustering processing on the captured data sample points in the archived data set. After the captured data sample points in the archived data set are clustered by DBSCAN, a neighborhood region is set with the data as the center and a preset radius of preset length. When the number of data points in the neighborhood region of a data point is greater than or equal to a preset threshold, the data point is defined as a core point; when the number of data points in the neighborhood region of a data point is less than the preset threshold, the data point is defined as a boundary point; and the remaining data points are defined as noise points. That is, the sample points will be distinguished into three cases: core objects, edge points, or noise points.
[0060] This algorithm groups any two core object sample points whose distance is less than the neighborhood radius into the same cluster, repeating the above steps until all data in the detection dataset is selected, forming several different dense data regions. All samples with directly reachable density from captured archival data are within a circle centered on the core object. If they are not within the circle, they are not density-reachable, and the captured archival data are connected to form a density-reachable data set. That is, the set of samples with the highest density connectivity (clustering) is derived from the density-reachability relationship. Such a set contains one or more core objects. If there is only one core object, all other non-core objects in the cluster are within the neighborhood of this core object; if there are multiple core objects, then the neighborhood of any one core object must contain another core object (otherwise, they are not density-reachable). These core objects and all samples contained within their neighborhoods constitute a class.
[0061] If there is a strong connection between two core points, the data in the neighborhood of the two core points belong to the same cluster; if there is no connection between two core points, the data in the neighborhood of the two core points belong to different clusters.
[0062] Within each cluster consisting of core objects and edge points, random undersampling is performed to remove samples.
[0063] The processing focus here is on core objects; edge points are not processed. This is because core objects have a high density of sample points and significant sample overlap. Removing some only reduces information redundancy, not a loss of information.
[0064] As an optional implementation of this invention, density clustering is performed on the dataset of multiple time-period captured data samples. Based on the clustering of the multiple time-period captured data samples, a sample point removal operation is performed on the multiple time-period captured data samples to obtain processed multiple time-period captured data sample points. This includes: performing density clustering on the dataset of multiple time-period captured data samples to obtain multiple clusters of data with continuous time periods; determining whether the multiple clusters of data with continuous time periods are noise point sets; if they are not noise point sets, then two sample points are randomly selected, and it is determined whether one of the two sample points is a core object sample point and the other is an edge point; if one is a core object sample point and the other is an edge point, then the core object sample point is removed. In this way, duplicate sample points of multiple time-period captured data sample points can be removed.
[0065] For example, in this embodiment, by using the neighborhood parameter ( , This is used to describe the density of data distribution in a neighborhood, where... This describes the neighborhood distance threshold of a given set of data. The distance to a certain sample is described as The threshold for the number of samples in the neighborhood. Different archive datasets captured from the same camera location are input to the server. and neighborhood parameters ( , Then, the process is divided into different clusters C, where, This represents a collection of different file data captured from the same camera location. This represents snapshot data from different time periods within the archive dataset;
[0066] (1) Initialize the number of clusters k=0
[0067] (2) For To find all core objects, follow these steps:
[0068] The core object contains the captured archive data.
[0069] By using distance metrics, After sorting the data according to the capture time, starting from the first data point and counting downwards, check if the first and second data points meet the following conditions: If the condition is met, continue calculating whether the condition is met using the first and third data points. If the conditions are not met, then use the second and third data to calculate. Check if the neighborhood satisfies the condition, and repeat the analogy downwards until the calculation ends.
[0070] The above calculations satisfy The results of the neighborhood are retained, and the results are counted sequentially based on the consecutive calculation results that satisfy the condition of being less than or equal to the nearest neighbor. , sample Add to core object sample collection
[0071] (3) Perform intersection calculation on the data in the core object set calculated in (2) above to determine whether core1 contains the core point information of core2. If it does, merge the two core objects into a large set C1. If it does not, end the formation of set C2 and perform subsequent merging calculations in sequence.
[0072] (4) The final clustering output is: cluster division C={C1,C2,...,Ck}, then each cluster represents a time period with continuous customer flow. Within this continuous time period, a corresponding popularity level is set for it, and the same type of stores and their consumer products are ranked and recommended according to their popularity level.
[0073] It should be noted that the letters in the above series of parameters can be replaced by other letters, and no specific limitation is made in this embodiment.
[0074] As an optional implementation of this invention, the passenger flow statistics method provided in this embodiment only removes duplicate information, rather than weakening useful information. This reduces the workload of passenger flow statistics, making it faster and more efficient to obtain clustering results for dense passenger flow clustering. It also makes passenger flow information statistics faster and more convenient, and improves the accuracy of recommendations.
[0075] S150: Based on the clustering results of the multiple data clusters with continuous time periods, determine the passenger flow popularity level of the target area in the corresponding time period.
[0076] After performing density-based DBSCAN clustering algorithm calculations on the time density information of the captured data, multiple clustering results of continuous time period data will be generated. Each cluster represents a time period with continuous customer flow. Based on the time period with continuous customer flow, the heat level corresponding to the target area is set. Finally, based on the heat level corresponding to the target area, the ranking and recommendation of similar stores and their consumer products are made.
[0077] This embodiment considers various potential factors related to customer flow information. Specifically, it performs density-based DBSCAN clustering calculations based on customer flow information to obtain clustering results for multiple time periods with continuous customer flow. Based on the time periods with continuous customer flow, it obtains corresponding popularity levels. Based on these popularity levels, it identifies similar stores and studies the customer flow characteristics of similar stores. This not only helps consumers understand the popularity of stores more accurately and makes ranking recommendations for similar stores and their products, but also provides decision-making basis for store sales, operation, and management.
[0078] As can be seen, in this embodiment of the invention, the present invention provides a method, device, computer equipment, and storage medium for intensive store traffic statistics. It acquires snapshot data of a target area at various time periods. The snapshot data includes snapshot image features and snapshot time. Cluster analysis is performed on the snapshot data to form an archive set. Based on image similarity, snapshot data from the archive set is continuously archived at different time periods using the same camera to form corresponding archived data sets. The snapshot data in the archived data sets is sorted to obtain the density information of the snapshot data time. Density clustering is performed based on the density information of the snapshot data time to form multiple clustering results of data clusters with continuous time periods. Based on the clustering results of multiple clusters of data clusters with continuous time periods, the traffic intensity level of the target area in the corresponding time period is determined. In this invention, after capturing traffic data at various time periods using a camera, the snapshot information is clustered and then archived to form corresponding archived data sets. This ensures high image quality and angle quality of the snapshot data in each archived data set and improves the efficiency of subsequent data processing. After determining the density information of the captured data time, a density-based DBSCAN clustering algorithm is applied to calculate the data across multiple continuous time periods. This enables more reasonable customer traffic statistics based on irregular consumer behavior, more accurately determining the popularity level of a store at different times of the day, and providing ranking recommendations for similar stores and their associated products. This effectively enhances the method of customer traffic statistics and improves the accuracy of recommendations. The DBSCAN-based clustering method does not require pre-specifying the number of clusters; the number of clusters is automatically determined through the density-based clustering process, which is of great significance in real-world scenarios with complex user distributions.
[0079] It is understood that in the specific implementation of this application, data related to passenger flow statistics methods are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0080] As one implementation method of this embodiment, such as Figure 2As shown, in step S120, density clustering is performed based on the density information of the captured data time to form multiple clusters of data with continuous time periods. The results include:
[0081] Step S121: For each capture data to be clustered, based on the density information of the capture data time, determine whether there is a density information of the capture data to be clustered in different time periods that is higher than the density information of the capture data time; if so, perform clustering processing on the density information of the capture data time and the density information of the capture data in different time periods.
[0082] Step S122: Pre-set the neighborhood radius and minimum threshold, and randomly select one of multiple time period capture data from the archived data set, wherein the selected capture data is considered as the initial data cluster;
[0083] Step S123: Determine whether the density information of the captured data in different time periods within the search range expanded from the initial data cluster of the neighborhood radius exceeds the minimum threshold; if so, use the density information of the captured data in different time periods within the search range as the core data.
[0084] Step S124: If the distance between core data is less than the neighborhood radius, classify the two core data that are less than the neighborhood radius into the same cluster and regard them as a new data cluster;
[0085] Step S125: Repeat the clustering steps for the capture data density information of other time periods until the capture data of all time periods in the archived data set are clustered, forming multiple clustering results of data clusters with continuous time periods.
[0086] In steps S121 to S125 above, a data object point p is randomly selected from the archived data set. A neighborhood region is defined with this data point as the center and a preset radius as the preset length. When the number of data points within the neighborhood region is greater than a preset threshold, the data point is defined as a core point; when the number of data points within the neighborhood region is less than the preset threshold, the data point is defined as a boundary point; the remaining data points are defined as noise points.
[0087] The selected data object point p is the center point, the neighborhood radius is Eps, and the minimum number of points MinPts is the point count threshold; the neighborhood radius is preset, and the preset value of the neighborhood radius is used as the neighborhood region; the minimum number of points is preset, and in this embodiment, the preset value of the minimum number of points is used as the point count threshold; proceed to the next step of calculation.
[0088] Assuming the total number of data points in the archived dataset is 100, the neighborhood radius is preset to 3, and the minimum number of points is preset to 5. When the number of data points within a neighborhood of radius 3 centered on data object point p is greater than 5, then data object point p is a core point of data object p; otherwise, it is a non-core point of data object p. All data in the detection dataset are processed sequentially to complete the judgment of all data in the detection dataset. If the distance between core points is less than the neighborhood radius, two core points with a distance less than the neighborhood radius are grouped into the same dense region. All core points obtained through the above steps are calculated sequentially; the relationship between the distance between different core points and the neighborhood radius is calculated. When the distance between core points is less than the neighborhood radius, it can be determined that the core points are in the same data cluster; when the distance between core points is greater than or equal to the neighborhood radius, ...
[0089] It can be determined that the core points are not in the same data cluster.
[0090] In this embodiment, it is assumed that there are different core points m and n. When the distance between core point m and core point n is less than 3, they are considered to be in the same data cluster. When the distance between core point m and core point n is greater than or equal to 3, they are considered to be in different data clusters. The above steps are repeated until all data in the detection dataset is selected, forming several different data clusters.
[0091] As can be seen from the above steps, after processing all the core points in sequence, several data clusters with continuous time periods are finally obtained. Each cluster represents a time period with continuous customer flow. Based on the time period with continuous customer flow, the corresponding popularity level is set, and the stores of the same type and consumer products are ranked and recommended.
[0092] Assuming the total number of data points in the archived dataset is 100, and after the above steps, nine different data clusters are obtained, it can be determined that a complete and continuous time-period data cluster for nine individuals within the target area has been obtained. The total number of data points in the archived dataset can also be other numbers; this embodiment does not impose a specific limitation.
[0093] In this embodiment, DBSCAN inputs multiple pedestrian features fused from 2048-dimensional means. The parameters of DBSCAN, such as neighborhood radius Eps and minimum number of points threshold MinPts, are flexibly selected according to different sites, and finally outputs densely connected regions.
[0094] In this embodiment, the DBSCAN clustering algorithm was used to achieve cross-probe clustering. Other different clustering methods, such as rerank and k-means, can also be used to achieve cross-probe clustering.
[0095] In this embodiment, the DBSCAN clustering algorithm is used to calculate dense passenger flow. This method defines two parameters: the maximum radius of the cluster and the minimum number of points a cluster should contain. Clustering continues as long as the density (number of objects or data points) of neighboring areas exceeds a certain threshold. Finally, clusters within the same cluster are considered a single group. Since passenger flow data is dense and the dataset is not convex, the DBSCAN clustering algorithm can better statistically analyze dense passenger flow data of arbitrary shapes. It can identify outliers during clustering, achieving more accurate data clusters, and the clustering results are unbiased. Coordinate information is used to describe the density of the captured data distribution in the neighborhood. Different archive datasets captured from the same camera location are input to divide the data into different clusters. The intersection of the captured archive datasets is calculated to determine if one captured archive dataset contains the core point information of another. If so, the two captured archive datasets are merged into a larger set; otherwise, the set formation process ends, and subsequent merging calculations are performed sequentially. By setting up this step, you can achieve more reasonable customer traffic statistics based on consumers' irregular behavior, making the data more accurate, intuitive, and straightforward. This greatly improves the efficiency of customer traffic statistics while achieving good recommendation results.
[0096] Please see Figure 4 , Figure 4 This is a schematic diagram of a passenger flow counting device provided in an embodiment of the present invention. In this embodiment, the terminal includes units used for performing... Figure 2 The steps in the corresponding embodiments. Please refer to the details. Figure 2 The relevant descriptions in the corresponding embodiments are shown below. For ease of explanation, only the parts relevant to this embodiment are shown. See also... Figure 4 The passenger flow statistics device 30 includes: a data acquisition module 31, a data archiving module 32, a data sorting module 33, a data density clustering processing module 34, and a data determination module 35.
[0097] The data acquisition module 31 is used to acquire capture data of the target area in various time periods, wherein the capture data includes capture image features and capture data time.
[0098] The data archiving module 32 is used to perform cluster analysis on the captured data to form an archive set; based on image similarity, the captured data in the archive set is continuously archived to form a corresponding archived data set under different time periods of the same camera;
[0099] The data sorting module 33 is used to sort the captured data for each time period in the archived data set and obtain the density information of the captured data time.
[0100] The data density clustering processing module 34 is used to perform density clustering processing based on the density information of the time of the captured data to form multiple clustering results of data clusters with continuous time periods.
[0101] The data determination module 35 is used to determine the passenger flow popularity level of the target area in the corresponding time period based on the clustering results of the multiple data clusters with continuous time periods.
[0102] Optionally, the data acquisition module 31 mentioned above includes:
[0103] The first capture unit is used to determine the capture object based on the capture data and to obtain the first capture information of the capture object in different time periods.
[0104] The second capture unit is used to analyze the first capture information of the subject at different time periods and obtain the second capture information of the subject's companions at different time periods.
[0105] Optionally, the aforementioned second capture unit includes:
[0106] The accompanying personnel capture unit is used to capture images of people other than the captured object that appear in the captured images t seconds before and after the capture data time point of the captured object.
[0107] The time determination subunit is used to determine the time points at different stages when the target area captures the captured object;
[0108] The clustering analysis and filtering unit is used to perform clustering analysis on the first and second capture information. Based on the image quality and angle quality of the captured object, the capture data of each cluster data pile is filtered, and the filtered cluster data piles are used as an archive set.
[0109] Optionally, the data archiving module 32 mentioned above includes:
[0110] The archive collection unit is used to classify the captured data into the corresponding archive collection when the captured data is greater than or equal to the preset similarity threshold;
[0111] A new set unit is created to create a new file set in the file set when the captured data is less than the preset similarity threshold, and the captured data is then included in the new file set.
[0112] The final archive collection unit is used to continuously archive the captured data to form a corresponding archive data collection.
[0113] Optionally, the data sorting module 33 mentioned above includes:
[0114] The sorting unit is used to sort the captured data according to the similarity of the captured data in each time period in the archived data set;
[0115] The weight determination unit is used to determine the weight of the captured data in different time periods after sorting the captured data.
[0116] The density information determination unit is used to determine the time density information of the captured data based on the similarity of the captured data in each time period in the archived data set and the weight of the captured data in different time periods.
[0117] Optionally, the aforementioned data density clustering processing module 34 includes:
[0118] The calculation unit is used to perform density clustering processing on the density information of the time of the captured data based on the density clustering algorithm, and obtain the clustering results of multiple data clusters with continuous time periods.
[0119] Optionally, the above-mentioned computing unit includes:
[0120] The clustering processing subunit is used to determine, for each data point to be clustered, whether there is a density information of the data point to be clustered in different time periods that is higher than the density information of the data point to be clustered in the time period, based on the density information of the time of the data point to be clustered in the time period; if so, then the density information of the time of the data point to be clustered and the density information of the data point to be clustered in different time periods are clustered.
[0121] The preset sub-unit is used to pre-set the neighborhood radius and minimum threshold, and randomly select one of the multiple time period capture data from the archived data set, wherein the selected capture data is considered as the initial data cluster;
[0122] The core data determination sub-unit is used to determine whether the density information of the captured data in different time periods within the search range extended from the initial data cluster of the neighborhood radius exceeds the minimum threshold; if so, the density information of the captured data in different time periods within the search range is used as the core data.
[0123] A data cluster subunit is formed, which is used to classify two core data that are less than the neighborhood radius into the same cluster and regard them as a new data cluster when the distance between the core data is less than the neighborhood radius.
[0124] The clustering result subunit is used to repeat the clustering steps of the capture data for the density information of capture data in other time periods until the capture data of all time periods in the archived data set are clustered, forming multiple clustering results of data clusters with continuous time periods.
[0125] Optionally, the data determination module 35 mentioned above includes:
[0126] The time period is represented by multiple consecutive time period data clusters, meaning that each cluster represents a time period with continuous passenger flow.
[0127] A popularity rating unit is used to set the popularity rating of the target area according to the time period with continuous passenger flow.
[0128] The recommendation unit ranks and recommends similar stores and consumer products based on the popularity level corresponding to the target area.
[0129] It should be noted that the information interaction and execution process between the above modules, units, and sub-units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0130] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described embodiments of the passenger flow statistics method.
[0131] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0132] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0133] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0134] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of the computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0135] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electrical carrier signal or a telecommunication signal.
[0136] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0137] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0138] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0139] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0140] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0141] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0142] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for counting passenger flow, characterized in that, include: Acquire capture data of the target area at various time periods, wherein the capture data includes capture image features and capture data time; Cluster analysis is performed on the captured data to form an archive set; based on image similarity, the captured data in the archive set is continuously archived at different time periods using the same camera to form a corresponding archived data set; The captured data for each time period in the archived data set is sorted to obtain the time density information of the captured data; Density clustering is performed based on the density information of the captured data time to form multiple clustering results of data clusters with continuous time periods; Based on the clustering results of the multiple data clusters with continuous time periods, the visitor traffic popularity level of the target area in the corresponding time period is determined; The step of sorting the captured data for each time period in the archived data set to obtain the time density information of the captured data includes: The captured data is sorted according to the similarity of the captured data in each time period in the archived data set; After sorting the captured data, the weight of the captured data in different time periods is determined. Based on the similarity of the captured data in each time period in the archived data set, and the weight of the captured data in different time periods, the time density information of the captured data is determined.
2. The passenger flow statistics method as described in claim 1, characterized in that, The process of clustering the captured data to form an archive set includes: Based on the captured data, the capture object is determined, and the first capture information of the capture object in different time periods is obtained; Based on the first capture information of the captured object at different time periods, the first capture information is analyzed to obtain the second capture information of the companions of the captured object at different time periods. The companions are people other than the captured object who appear in the capture images t seconds before and after the capture data time point of the captured object. The capture data time point is the time point at different stages when the captured object is captured in the target area. Cluster analysis is performed on the first and second capture information. Based on the image quality and angle quality of the captured object, the capture data of each cluster is filtered, and the filtered clusters are used as an archive set.
3. The passenger flow statistics method as described in claim 1, characterized in that, The process of continuously archiving the captured data in the archive set based on image similarity at different time periods using the same camera to form a corresponding archived data set includes: When the captured data is greater than or equal to a preset similarity threshold, the captured data is classified into the corresponding file set; When the captured data is less than the preset similarity threshold, a new archive set is created in the archive set, and the captured data is included in the new archive set, so as to continuously archive the captured data to form a corresponding archive data set.
4. The passenger flow statistics method as described in claim 1, characterized in that, Based on the density information of the captured data time, density clustering is then performed to form multiple clustering results of data clusters with continuous time periods, including: Density clustering algorithm is used to perform density clustering on the time density information of the captured data to obtain multiple clustering results of data clusters with continuous time periods.
5. The passenger flow statistics method as described in claim 4, characterized in that, The density clustering algorithm is used to perform density clustering processing on the time density information of the captured data to obtain multiple clustering results of data clusters with continuous time periods, including: For each set of capture data to be clustered, based on the density information of the time of the capture data to be clustered, determine whether there is a density information of the capture data to be clustered in different time periods that is higher than the density information of the time of the capture data to be clustered; if so, then perform clustering processing on the density information of the time of the capture data to be clustered and the density information of the capture data in different time periods. A neighborhood radius and a minimum threshold are preset, and one of the multiple time period snapshot data is randomly selected from the archived data set, wherein the selected snapshot data is considered as the initial data cluster; Determine whether the density information of the captured data in different time periods within the search range expanded from the initial data cluster of the neighborhood radius exceeds the minimum threshold; if so, use the density information of the captured data in different time periods within the search range as the core data. If the distance between core data is less than the neighborhood radius, the two core data that are less than the neighborhood radius are classified into the same cluster and regarded as a new data cluster; Repeat the clustering steps for the capture data density information of other time periods until the capture data of all time periods in the archived data set are clustered, forming multiple clustering results of data clusters with continuous time periods.
6. The passenger flow statistics method as described in claim 5, characterized in that, In the multiple data clusters with continuous time periods, each cluster represents a time period with continuous passenger flow. The process of forming the clustering results of multiple data clusters with continuous time periods includes: Based on the time period with continuous passenger flow, set the popularity level corresponding to the target area; Based on the popularity level corresponding to the target area, similar stores and consumer products are ranked and recommended.
7. A passenger flow counting device, characterized in that, include: The data acquisition module is used to acquire capture data of the target area in various time periods, wherein the capture data includes capture image features and capture data time. The data archiving module is used to perform cluster analysis on the captured data to form an archive set; based on image similarity, the captured data in the archive set is continuously archived at different time periods of the same camera to form a corresponding archived data set; The data sorting module is used to sort the captured data for each time period in the archived data set and obtain the time density information of the captured data. The data density clustering processing module is used to perform density clustering processing based on the density information of the time of the captured data, forming multiple clustering results of data clusters with continuous time periods; The data determination module is used to determine the passenger flow popularity level of the target area in the corresponding time period based on the clustering results of the multiple data clusters with continuous time periods; The step of sorting the captured data for each time period in the archived data set to obtain the time density information of the captured data includes: The captured data is sorted according to the similarity of the captured data in each time period in the archived data set; After sorting the captured data, the weight of the captured data in different time periods is determined. Based on the similarity of the captured data in each time period in the archived data set, and the weight of the captured data in different time periods, the time density information of the captured data is determined.
8. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the passenger flow statistics method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the passenger flow statistics method as described in any one of claims 1 to 6.