Reid-Assisted Passenger Flow Statistics Analysis System and Method
Through the Reid-based passenger flow statistical analysis system, using cameras and artificial intelligence technology to detect and identify human objects, the problem of insufficient behavioral distinction in traditional passenger flow statistics methods is solved, and accurate passenger flow data is provided to support business decisions.
Patent Information
- Application Number
- CN202510246469.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-04
AI Technical Summary
Traditional customer flow statistics methods cannot distinguish different types of customer behavior, resulting in insufficient fine-grained customer flow analysis, difficulty in accurately judging real customers, and ineffective non-customer personnel, resulting in inaccurate data.
The passenger flow statistical analysis system based on reid assist is adopted, and the door scene monitoring images are collected through the camera, and the human object detection and orientation recognition are carried out. The trained reid model is used to distinguish people entering the store, wandering and leaving the store, and combining object re-identification technology to avoid repeated counting.
It realizes accurate identification of customer behavior patterns, provides more accurate customer flow statistics, supports merchants to optimize store layout and service configuration, and improves operational efficiency.
Smart Images

Figure CN119810752B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of passenger flow statistics and analysis, and more specifically, to a passenger flow statistics and analysis system and method assisted by reid. Background Art
[0002] With the continuous development of the retail and service industries, merchants are increasingly emphasizing the improvement of customer experience and store operation efficiency. To better understand customer behavior patterns, optimize store layouts, and allocate personnel, accurate passenger flow statistics and analysis have become particularly important. Traditional passenger flow statistics methods usually rely on simple sensors (such as infrared or access control counters), which can only provide a rough count of the number of people entering and leaving, and cannot distinguish different types of customer behaviors, such as wandering, entering, or leaving. In addition, there is insufficient experience in the counting application of person re-identification, resulting in insufficient fine-grainedness in passenger flow analysis and difficulty in obtaining accurate results. In particular, the method of accurately judging real customers is ignored, making it difficult to effectively exclude non-customer personnel (such as delivery workers, staff, etc.).
[0003] Therefore, an optimized passenger flow statistics and analysis solution is desired. Summary of the Invention
[0004] This application provides a passenger flow statistics and analysis system and method assisted by reid, which can identify the behavior patterns of each human target object at the store entrance, thereby performing orientation recognition and detection on it, which helps subsequent passenger flow statistics and analysis and provides strong support for business decisions.
[0005] In the first aspect, a passenger flow statistics and analysis method assisted by reid is provided, including:
[0006] Obtain the entrance scene monitoring image collected by the camera deployed at the store entrance;
[0007] Perform human target detection on the entrance scene monitoring image to obtain a set of human target data;
[0008] Use the trained reid-assisted model to perform human orientation recognition on each piece of human target data in the set of human target data to obtain a set of human orientation recognition results;
[0009] If the human orientation recognition result is positive, regard the human target data as a person entering the store and increment the passenger flow statistics count by one;
[0010] If the human orientation recognition result is side, regard the human target data as a wandering person;
[0011] If the human orientation recognition result is back, regard the human target data as a person leaving the store and decrement the passenger flow statistics count by one.
[0012] In a second aspect, a passenger flow statistical analysis system assisted by ReID is provided, including:
[0013] A door scene monitoring image acquisition module for acquiring door scene monitoring images collected by a camera deployed at the entrance of a store;
[0014] A human target detection module for performing human target detection on the door scene monitoring images to obtain a set of human target data;
[0015] A human orientation recognition module for using a trained ReID-assisted model to perform human orientation recognition on each piece of human target data in the set of human target data to obtain a set of human orientation recognition results;
[0016] An entering-store personnel statistics module for, if the human orientation recognition result is positive, regarding the human target data as entering-store personnel and incrementing the passenger flow statistics count by one;
[0017] A wandering personnel determination module for, if the human orientation recognition result is side, regarding the human target data as wandering personnel;
[0018] A leaving-store personnel statistics module for, if the human orientation recognition result is back, regarding the human target data as leaving-store personnel and decrementing the passenger flow statistics count by one.
[0019] A passenger flow statistical analysis system and method assisted by ReID provided by the present application, after performing human target detection on the collected door scene monitoring images, uses image analysis and recognition algorithms based on artificial intelligence and machine vision to analyze each piece of human target data, so as to respectively capture the significant feature representations of the human target in each piece of human target data, which helps to perform human orientation discrimination to obtain human orientation recognition results, including positive, back, and side. In this way, it is possible to perform behavior pattern recognition on each human target object at the store entrance, thereby performing orientation recognition and detection on it, which helps subsequent passenger flow statistical analysis and provides strong support for business decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present application and do not limit the present application.
[0021] Figure 1 It is a schematic flowchart of a passenger flow statistical analysis method assisted by ReID according to an embodiment of the present application.
[0022] Figure 2Schematic diagram of data flow of the passenger flow statistics and analysis method assisted by reid according to the embodiment of the present application.
[0023] Figure 3 Schematic flowchart of S3 in the passenger flow statistics and analysis method assisted by reid according to the embodiment of the present application.
[0024] Figure 4 Schematic flowchart of S31 in the passenger flow statistics and analysis method assisted by reid according to the embodiment of the present application.
[0025] Figure 5 Schematic flowchart of S32 in the passenger flow statistics and analysis method assisted by reid according to the embodiment of the present application.
[0026] Figure 6 Schematic flowchart of S323 in the passenger flow statistics and analysis method assisted by reid according to the embodiment of the present application.
[0027] Figure 7 Schematic block diagram of the passenger flow statistics and analysis system assisted by reid according to the embodiment of the present application. Detailed implementation mode
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts also belong to the scope of protection of the present application.
[0029] In recent years, with the development of computer vision technology and deep learning algorithms, especially the progress of pedestrian re-identification (Re-Identification, reid) technology, it has become possible to provide more accurate passenger flow statistics. The reid technology focuses on identifying the same pedestrian under different camera perspectives, even if the appearance of the pedestrian changes due to factors such as angle, lighting, and occlusion. This technology can be used for tracking and identifying specific individuals in a monitoring environment, thereby achieving more detailed and accurate passenger flow analysis.
[0030] Based on this, in the technical solution of the present application, a passenger flow statistics and analysis method assisted by reid is proposed, which can combine object detection and improved person re-identification technology to perform accurate crowd counting to support passenger flow statistics. Specifically, as Figure 1 and Figure 2As shown, in the above-mentioned ReID-assisted passenger flow statistics analysis method, the specific implementation steps are as follows: S1, obtain the door scene monitoring images collected by the cameras deployed at the store entrance; S2, perform human target detection on the door scene monitoring images to obtain a set of human target data; S3, use the trained ReID-assisted model to perform human orientation recognition on each piece of human target data in the set of human target data to obtain a set of human orientation recognition results; S4, if the human orientation recognition result is positive, regard the human target data as the people entering the store and increment the passenger flow statistics count by one; S5, if the human orientation recognition result is side, regard the human target data as the wandering people; S6, if the human orientation recognition result is back, regard the human target data as the people leaving the store, and decrement the passenger flow statistics count by one. In addition, the ReID-assisted passenger flow statistics analysis method further includes the steps: S7, perform object re-identification on each piece of human target data in the set of human target data to obtain a set of object re-identification results; S8, in response to the object re-identification result being the same person entering the store, do not perform duplicate counting on the human target data.
[0031] Exemplarily, in step S1, obtain the door scene monitoring images collected by the cameras deployed at the store entrance. It should be understood that obtaining the door scene monitoring images collected by the cameras deployed at the store entrance is for accurate passenger flow statistics analysis. By doing so, all human target objects entering and leaving the store can be captured, and their behavior patterns can be recognized, so as to accurately judge whether the customers are entering the store, wandering, or leaving. This helps merchants better understand the behavior intentions and traffic changes of customers, optimize the store layout and service configuration, and improve the operation efficiency. By analyzing the monitoring images of the door scene, the specific behavior patterns of customers can be judged according to the human orientation (front, side, or back), such as whether they really enter the store or just wander at the door. Traditional people counting methods are easily affected by environmental factors, resulting in inaccurate data. Through image analysis technology, the problem of duplicate counting can be effectively avoided, and non-customer personnel (such as deliverymen and staff) can be filtered out, providing more accurate passenger flow information. Based on the accurate passenger flow statistics results, merchants can obtain valuable information about customer visit time, frequency, etc., so as to make more reasonable business decisions, such as adjusting business hours, promotional activity arrangements, etc.
[0032] In one embodiment, in order to obtain the monitoring image of the doorway scene, it is first necessary to install a high-definition camera at the entrance of the store to ensure that the entire entrance and exit area is covered, so as to completely record all human targets entering and leaving. The camera should have features such as night vision function and wide-angle lens to meet the monitoring requirements under different lighting conditions. Then connect the camera to a local server or a cloud platform to ensure that the image data can be transmitted and stored in real time, and at the same time ensure that the network bandwidth is stable enough to prevent image loss or delay due to network problems.
[0033] Exemplarily, in the step S2, human target detection is performed on the monitoring image of the doorway scene to obtain a set of human target data. It should be understood that performing human target detection on the monitoring image of the doorway scene to obtain a set of human target data is to accurately identify all human objects entering and leaving the store from a complex background environment. This process is a key step in passenger flow statistical analysis. It can not only distinguish customers from non-customers (such as deliverymen, staff), but also provide necessary basic information for subsequent behavior pattern recognition and orientation discrimination. Through human target detection, specific features of each individual, such as position, posture, etc., can be captured, which helps to more accurately track the behavior trajectory of individuals and make more detailed business decisions based on this, such as optimizing the store layout, improving service quality, arranging promotional activities, etc.
[0034] In one embodiment, the present application uses an anchor box-based target detection network to perform human target detection. Specifically, the entire process is implemented by a specially trained personnel target detection network, which can effectively filter out irrelevant features and focus on identifying human targets. Specifically, the door scene monitoring image is subjected to human target detection to obtain a set of human target data, including the following steps: First, the acquired door scene monitoring image is input into the personnel target detection network. This network is an anchor window-based target detection model, such as Fast R-CNN, Faster R-CNN or RetinaNet, etc., which all have powerful target positioning capabilities. Then, the target anchoring layer in these models is used, that is, a series of anchor boxes B of predefined sizes and proportions are used to perform sliding window scanning on the image. Each anchor box B will try to match the human targets that may exist in the image according to its location and generate the corresponding candidate area. When the anchor box B slides on the image, it not only frames the potential human target interest area, but also evaluates the possibility that these areas contain actual human bodies. For each candidate region that is considered to be a human target, the network will output a set of coordinate information to define the position of the human body in the region, as well as a confidence score that indicates the reliability of this prediction. In order to improve the quality of the detection results, the non-maximum suppression (NMS) algorithm is usually implemented to eliminate redundant overlapping detection boxes to ensure that the final result is the most accurate and non-repetitive human target data set. After this processing, all candidate regions that are considered to contain human bodies constitute a set of human target data, providing a reliable basis for subsequent analysis.
[0035] Exemplarily, in step S3, the trained ReID auxiliary model is used to perform human orientation recognition on each human target data in the set of human target data to obtain a set of human orientation recognition results. In particular, the process of performing human orientation recognition on each human target data in the set of human target data using the trained ReID auxiliary model is crucial, because different behavior patterns of human targets can be distinguished through human orientation recognition. For example, a person facing forward may be more likely to be a customer entering the store, while a person facing backward may be a customer leaving the store. A person facing sideways may be wandering at the door and has not yet decided whether to enter the store. This distinction helps merchants understand the behavioral intentions of customers more accurately. Simply relying on the number of people entering and leaving the store cannot provide enough information to determine whether a person has truly become a customer. Orientation recognition can help filter out pedestrians who only pass by the store but do not enter, thereby improving the accuracy of passenger flow statistics.
[0036] Based on this, in the process of performing human body orientation recognition on each piece of human body target data in the set of the human body target data, the technical concept of the present application is to perform human body target detection on the collected monitoring images of the doorway scene, and then use image analysis and recognition algorithms based on artificial intelligence and machine vision to analyze each piece of human body target data, so as to respectively capture the significant feature representations of the human body target in each piece of human body target data, which helps to perform human body orientation discrimination to obtain human body orientation recognition results, including front, back, and side. In this way, the behavior patterns of each human body target object at the store entrance can be recognized, so as to perform orientation recognition and detection on it, which helps subsequent passenger flow statistical analysis and provides strong support for business decisions.
[0037] In one embodiment, as Figure 3 shown, use the trained reid auxiliary model to perform human body orientation recognition on each piece of human body target data in the set of the human body target data to obtain a set of human body orientation recognition results, including: S31, input the human body target data into a human multi-scale feature extractor based on a dynamic convolution labeling model to obtain a human body target multi-scale image feature encoding map; S32, perform feature selection on the human body target multi-scale image feature encoding map to obtain a human body target image significant feature encoding map; S33, perform human body orientation discrimination based on the human body target image significant feature encoding map to obtain the human body orientation recognition result.
[0038] Exemplarily, in step S31, input the human body target data into a human multi-scale feature extractor based on a dynamic convolution labeling model to obtain a human body target multi-scale image feature encoding map. It should be understood that in actual passenger flow statistics, the camera may capture human body targets at different distances and angles. For example, the camera at the doorway may capture pedestrians at both near and far distances at the same time. Therefore, in order to be able to analyze the human body target data more comprehensively and meticulously, and thus better capture the behavior patterns of the human body target data, in the technical solution of the present application, the human multi-scale feature extractor based on the dynamic convolution labeling model can combine the advantages of CNN and Transformer, which can not only capture the detailed features of local image patches of the human body target data, but also learn global semantic information, which enables the system to better adapt to various environmental changes, improve the generalization ability, and thus provide more accurate data support in complex passenger flow statistical scenarios.
[0039] In one embodiment, as Figure 4As shown, inputting the human target data into a human multi-scale feature extractor based on a dynamic convolution tagging model to obtain a human target multi-scale image feature encoding map includes: S311, performing local image block segmentation on the human target data to obtain a set of human target local image blocks; S312, inputting each human target local image block in the set of human target local image blocks into a human target local semantic feature extractor based on CNN to obtain a set of human target local feature maps; S313, splicing the set of human target local feature maps into a human target global representation feature map and then inputting it into a human target feature global correlation encoder based on a Transformer model to obtain the human target multi-scale image feature encoding map.
[0040] Specifically, the process of inputting human target data into a human multi-scale feature extractor based on a dynamic convolution tagging model to obtain a human target multi-scale image feature encoding map is a continuous and comprehensive operation. This process begins with local image block segmentation of the detected human target data, that is, dividing the entire human target area into several small blocks, each of which represents a part of the human body, such as the head, shoulders, or arms. This segmentation method helps capture detailed information of different parts and provides the basic material for subsequent feature extraction. Subsequently, each of these segmented local image blocks is separately fed into a human target local semantic feature extractor based on a convolutional neural network (CNN). Through training, this extractor can learn rich semantic features from each local image block, such as texture, color, and shape. The output result is a series of local feature maps, each corresponding to a specific local area, enabling the system to more carefully understand the uniqueness of each local area and enhancing the adaptability to complex scenarios. Next, all local feature maps are spliced together to form a human target global representation feature map. This is not just a simple splicing, but rather integrating the relationships between local features through a certain strategy (such as max pooling, average pooling, or self-attention mechanism) to construct a global feature representation containing overall structural information. This global representation feature map not only retains local details but also introduces spatial correlation, making the features more rich and representative. To further enhance the feature expression ability, a human target feature global correlation encoder based on a Transformer model is used to process the spliced global representation feature map. At this stage, dynamic convolution technology is used to adjust the size and shape of the convolution kernel according to different inputs to achieve feature capture of the human target at different scales. At the same time, combined with multi-scale feature fusion methods, information from different levels is combined to generate the final human target multi-scale image feature encoding map. Dynamic convolution and multi-scale feature fusion ensure that the model can work stably under various distance and angle conditions, improving the robustness and generalization performance of feature extraction.
[0041] Exemplarily, in step S32, feature selection is performed on the multi-scale image feature encoding map of the human target to obtain the significant feature encoding map of the human target image. It should be understood that the multi-scale image feature encoding map of the human target contains multi-scale encoded features regarding the human target. However, not all of the extracted multi-scale features of the human target are helpful for the final tasks of human orientation discrimination and passenger flow statistical analysis. Some features may carry redundant information or be irrelevant to the tasks. Removing these irrelevant features can reduce noise interference, thereby improving the recognition accuracy of the model. Based on this, in order to more effectively identify and highlight the key feature semantics and significant features of the human target, reduce the amount of data for subsequent processing, save computing resources, and provide a basis for subsequent real-time passenger flow statistics, in the technical solution of this application, further feature selection is performed on the multi-scale image feature encoding map of the human target to obtain the significant feature encoding map of the human target image. In particular, through the process of feature selection, it is possible to optimize the feature representation in the deep learning model by evaluating the interaction between the local features of the human target and their neighboring local features, so as to effectively compress and sparsify the original multi-scale image feature encoding map of the human target.
[0042] In one embodiment, as Figure 5 shown, performing feature selection on the multi-scale image feature encoding map of the human target to obtain the significant feature encoding map of the human target image includes: S321, performing feature decomposition on the multi-scale image feature encoding map of the human target to obtain a set of local feature vectors of the multi-scale image semantic encoding of the human target; S322, inputting each local feature vector of the multi-scale image semantic encoding of the human target in the set of local feature vectors of the multi-scale image semantic encoding of the human target into an importance measurement module to obtain a set of importance score values of the local feature vectors of the multi-scale image semantic encoding of the human target; S323, based on the set of importance score values of the local feature vectors of the multi-scale image semantic encoding of the human target, performing feature selection and feature shape reshaping on the set of local feature vectors of the multi-scale image semantic encoding of the human target to obtain the significant feature encoding map of the human target image.
[0043] In one embodiment, in step S321, performing feature decomposition on the multi-scale image feature encoding map of the human target to obtain a set of local feature vectors of the multi-scale image semantic encoding of the human target includes: performing feature decoupling and feature flattening on the multi-scale image feature encoding map of the human target to obtain the set of local feature vectors of the multi-scale image semantic encoding of the human target. Specifically, this process can be represented by the formula:
[0044]
[0045]
[0046] Among them, is the multi-scale image feature encoding map of the human target, is the feature decoupling operation, respectively represent the 1st, 2nd, th, and the th multi-scale image semantic encoding feature matrices of the human target in the set of multi-scale image semantic encoding feature matrices of the human target, is the feature flattening process, respectively represent the 1st, 2nd, th, and the th multi-scale image semantic encoding local feature vectors of the human target in the set of multi-scale image semantic encoding local feature vectors of the human target.
[0047] It should be understood that when the multi-scale image feature encoding map of the human target extracted from the surveillance image is preliminarily processed, it contains rich information, but not all information is useful for subsequent recognition tasks. Therefore, through feature decomposition, the complex global feature map can be decoupled into multiple local feature vectors, and each vector represents the detailed features of a certain part of the human target (such as the head, shoulders, arms, etc.). This decomposition enables the system to capture the changes of different parts of the human body more carefully. Especially for human targets with various postures, local features often can reflect the uniqueness of individuals better than the overall shape.
[0048] In the step S322, each multi-scale image semantic encoding local feature vector in the set of multi-scale image semantic encoding local feature vectors of the human target is input into the importance measurement module to obtain a set of multi-scale image semantic encoding local feature importance score values. Specifically, this process can be expressed by the formula as follows:
[0049]
[0050] Among them, and are respectively the corresponding weight matrix and bias vector, is the matrix multiplication, is the modulation vector, is the corresponding multi-scale image semantic encoding local feature importance score value of
[0051] Next, the set of human target multi-scale image semantic encoded local feature vectors obtained above is fed into the importance measurement module. The role of this module is to assign a numerical "importance score" to each local feature vector to measure the weight of the feature in describing the entire individual. Specifically, features that carry more identity information (such as the face) are usually considered more important; while relatively less important features have lower scores. In this way, the system can select the most discriminative parts from among many possible feature combinations, reducing the interference caused by redundant information. After introducing the importance measurement mechanism, not only can the data dimension be significantly reduced and the calculation speed be accelerated, but also the reliability of the final result is improved. Because in practical applications, not all detected features are helpful for the recognition task, some may be noise or common features similar to other individuals. By quantifying the importance of each feature, the system can focus more on the unique markers that are truly helpful for distinguishing different people, thus improving the accuracy of re-identification. At the same time, this also means that even in the case of low quality or partial occlusion, as long as there are enough important features remaining, a relatively reliable match can still be achieved.
[0052] In one embodiment, as Figure 6 shown, in the step S323, based on the set of importance score values of the human target multi-scale image semantic encoded local features, feature selection and feature shape reshaping are performed on the set of human target multi-scale image semantic encoded local feature vectors to obtain the human target image significant feature encoded map, including: S3231, based on the set of importance score values of the human target multi-scale image semantic encoded local features, the set of human target multi-scale image semantic encoded local feature vectors is sorted in descending order to obtain a descending sequence of human target multi-scale image semantic encoded local feature vectors. Specifically, this process can be represented by the formula:
[0053]
[0054] where represents the descending order operation, respectively represent the 1st, 2nd, th, and th human target multi-scale image semantic encoded local feature vectors in the descending sequence of human target multi-scale image semantic encoded local feature vectors.
[0055] It should be understood that according to the importance score value of each local feature vector, it is sorted in descending order to form a descending sequence of human target multi-scale image semantic encoded local feature vectors sorted by importance. This operation ensures that the features that best represent the individual's identity are considered first, while those relatively unimportant features are ranked behind.
[0056] S3232. Calculate the feature neighborhood activity of each multi-scale image semantic encoding local feature vector of the human target based on the neighborhood features of each multi-scale image semantic encoding local feature vector of the human target in the descending sequence, so as to obtain a sequence of feature neighborhood activities. Specifically, this process can be expressed by the formula:
[0057]
[0058] where is the -th position feature value of the -th multi-scale image semantic encoding local feature vector of the human target in the descending sequence of multi-scale image semantic encoding local feature vectors of the human target, is the -th multi-scale image semantic encoding local feature vector of the human target, is the -th neighborhood feature factor of the multi-scale image semantic encoding local feature vector of the human target, is the -th neighborhood feature factor of the multi-scale image semantic encoding local feature vector of the human target, is the -th feature neighborhood activity of the multi-scale image semantic encoding local feature vector of the human target.
[0059] That is, for the feature vectors that have been sorted by importance, further analyze the correlation and distribution between them, that is, calculate the "feature neighborhood activity" of each local feature vector. Specifically, this refers to evaluating the activity degree of other relevant features around a certain local feature, so as to measure the uniqueness and representativeness of this local feature in the whole image.
[0060] S3233. Based on the sequence of the feature neighborhood activities, perform feature selection on the descending sequence of the multi-scale image semantic encoding local feature vectors of the human target to obtain a descending sequence of the selected multi-scale image semantic encoding local feature vectors of the human target. Specifically, this process can be expressed by the formula:
[0061]
[0062]
[0063] where is a predetermined threshold, is the feature selection process, is the 1st, 2nd, the first and the local feature vectors of the multi-scale image semantic encoding of the human target after the selection of the second
[0064] That is, based on the sequence of feature neighborhood activities, the feature vectors sorted in descending order are selected to remove redundant information and retain those feature vectors that are both important and have high activity. This step checks whether each feature vector meets the preset threshold condition, and only those that meet the criteria are retained in the final selection list.
[0065] S3234. Reshape the descending sequence of the selected local feature vectors of the multi-scale image semantic encoding of the human target to obtain the significant feature encoding map of the human target image. Specifically, this process can be expressed by the formula:
[0066]
[0067] where is the feature shape reshaping process, is the significant feature encoding map of the human target image.
[0068] That is, the selected feature vectors are reorganized into a new feature encoding map, and this process is called feature shape reshaping. In this way, the originally dense feature representation is converted into a sparse form, retaining only the most core part. The sparsified feature encoding map not only has a smaller volume but is also easier to store and transmit, and can better adapt to the deployment requirements on different hardware platforms.
[0069] That is, in one embodiment, based on the sequence of the feature neighborhood activities, feature selection is performed on the descending sequence of the local feature vectors of the multi-scale image semantic encoding of the human target to obtain the descending sequence of the selected local feature vectors of the multi-scale image semantic encoding of the human target, including: extracting the feature neighborhood activity corresponding to the first local feature vector of the multi-scale image semantic encoding of the human target from the sequence of the feature neighborhood activities to obtain the first feature neighborhood activity; comparing the first feature neighborhood activity with a predetermined threshold, and in response to the first feature neighborhood activity being greater than or equal to the predetermined threshold, retaining the first local feature vector of the multi-scale image semantic encoding of the human target, otherwise deleting the first local feature vector of the multi-scale image semantic encoding of the human target.
[0070] Specifically, the process of feature selection for the multi-scale image feature encoding map of the human target can achieve effective compression and optimization of the input multi-scale image semantic encoding feature map of the human target through multi-level analysis of human target features. This not only improves the efficiency of feature extraction but also enhances the understanding of the internal structure of the human target, promoting deeper learning of human target semantics and better generalization ability. Compared with traditional feature selection algorithms, the feature selection process can identify and remove features that, although important for local human target features, are highly correlated with their neighbors, thus effectively reducing redundant information and improving model efficiency. At the same time, by comprehensively considering the relationship between local human target features and their neighborhoods, it can better capture complex patterns in the data and the interaction between multi-scale image local features of the human target. In addition, since the feature selection method based on importance measurement emphasizes the relationship between features rather than simply relying on the importance of individual features, the selected features often better reflect the true characteristics of human target semantics, thereby enhancing the transparency and interpretability of model decisions, facilitating the understanding of the impact of each local human target feature on the final feature representation and subsequent human orientation recognition tasks, and thus improving the interpretability and credibility of the model. In this way, it is possible to more accurately distinguish real customers from other non-customers (such as deliverymen, staff), provide more in-depth customer behavior analysis, and support merchants in making better business decisions.
[0071] Exemplarily, in step S33, human orientation discrimination is performed based on the significant feature encoding map of the human target image to obtain the human orientation recognition result. In one embodiment, performing human orientation discrimination based on the significant feature encoding map of the human target image to obtain the human orientation recognition result includes: inputting the significant feature encoding map of the human target image into a classifier-based human orientation discriminator to obtain the human orientation recognition result. That is to say, the classification process is carried out using the human target image features strengthened by feature selection significance, so as to perform human orientation discrimination to obtain the human orientation recognition result, including front, back, and side. It is worth mentioning that here, the front represents an in-store customer, the side represents a passerby, and the back represents a leaving customer. In this way, it is possible to perform behavior pattern recognition on each human target object at the store entrance, and thus perform orientation recognition and detection on it, which helps subsequent passenger flow statistical analysis and provides strong support for business decisions.
[0072] In one embodiment, the classifier-based human body orientation discriminator uses a support vector machine (SVM). SVM is a supervised learning algorithm that is particularly good at handling binary or multi-classification problems in high-dimensional space. It maximizes the interval between different categories by finding the optimal hyperplane, thereby achieving better generalization performance. In this application, three categories of SVM can be set, corresponding to the three orientations of front, side and back.
[0073] Here, since the multi-scale image feature coding map of the human target represents the multi-scale image semantic coding features of the human posture features of the human target data, when performing feature selection based on the feature neighborhood activity, the feature population attributes of different local features of human posture will have feature selection fairness differences based on the difference in feature neighborhood activity measurement, thereby affecting the micro-macro representation bias of the significant feature coding map of the human target image, and reducing the accuracy of the human orientation recognition results obtained by the classifier-based human orientation discriminator.
[0074] In a preferred example, when the human target image salient feature coding map is input into a human orientation discriminator based on a classifier to obtain a human orientation recognition result, the human target image salient feature coding map is optimized, and the optimization process includes:
[0075] Expanding the human target image salient feature coding map into a human target image salient feature coding vector;
[0076] Based on the feature value of the significant feature coding vector of the human target image Distance and Distance matrix and the distance matrix of the salient feature encoding of human target image ;
[0077] Determine the number of super-distributed eigenvalues in the significant feature coding vector of the human target image whose difference from the eigenvalue mean is greater than the eigenvalue variance, and respectively calculate the reciprocal of the logarithm of the number of super-distributed eigenvalues with base two and the exponential value of the reciprocal of the number of super-distributed eigenvalues with base natural constants to obtain the first human target image significant feature coding full network interactive representation value and the second human target image salient feature encoding full network interactive representation value :
[0078]
[0079]
[0080]
[0081] Among them, represents the significant feature encoding vector of the human target image, represents the th eigenvalue of the significant feature encoding vector of the human target image, represents the calculation of number, is the number of hyper-distribution eigenvalues in the significant feature encoding vector of the human target image , and and are the mean and variance of all eigenvalues of the significant feature encoding vector of the human target image ;
[0082] Calculate the first human target image significant feature encoding hidden basic feature vector and the second human target image significant feature encoding hidden basic feature vector with the reciprocal of the first human target image significant feature encoding whole network interaction representation value and the reciprocal of the second human target image significant feature encoding whole network interaction representation value as exponents respectively for the significant feature encoding vector of the human target image:
[0083]
[0084]
[0085] Among them, represents the first human target image significant feature encoding hidden basic feature vector, represents the second human target image significant feature encoding hidden basic feature vector;
[0086] Multiply the first human target image significant feature encoding hidden basic feature vector as a row vector by the significant feature encoding one distance matrix of the human target image to obtain a significant feature encoding one distance query vector of the human target image , where represents matrix multiplication;
[0087] Multiply the significant feature encoding two distance matrix of the human target image by the second human target image significant feature encoding hidden basic feature vector as a column vector to obtain a significant feature encoding two distance response vector of the human target image ;
[0088] Calculate the weighted sum of the significant feature encoding one distance query vector of the human target image and the significant feature encoding two distance response vector of the human target image to obtain an optimized significant feature encoding vector of the human target image. Finally, pass the optimized significant feature encoding vector of the human target image through a human orientation discriminator based on a classifier to obtain a human orientation recognition result.
[0089] That is, for the problem of insufficient overall active query coding performance caused by long distances exceeding the local spacing limit in the fixed numerical sequence layout of the human target image significant feature coding vector, a multi-dimensional attribute-based form integrating the self-distance matrix of the human target image significant feature coding vector is adopted to capture the complex architecture of the whole-network interaction of its value range, and the performance of the live sales time-series semantic query response coding vector is reconstructed by simulating the large-scale multi-dimensional attribute hiding basis of the human target image significant feature coding vector, so as to realize the reconstruction of the active search performance connection of the human target image significant feature coding vector, and further reflect the predictable reconstruction under the actual series arrangement actions of the human target image significant feature coding vector, and improve the accuracy of the human body orientation recognition result obtained by the human body orientation discriminator based on the classifier for the human target image significant feature coding vector.
[0090] Exemplarily, in the step S4, if the human body orientation recognition result is positive, the human target data is regarded as a person entering the store and the passenger flow statistical count is incremented by one. It should be understood that in an actual scenario, a person facing the store entrance usually means that they have a relatively high probability of entering the store. People tend to face the direction they are about to go, so a person facing the store entrance frontally is more likely to be a potential customer preparing to enter the store. By using the human body orientation recognition technology, the real customers and other pedestrians can be more accurately distinguished, thereby improving the accuracy of the passenger flow statistical data. Accurate passenger flow statistics are crucial for merchants, as it can help optimize the store layout, personnel allocation, and service quality. By counting the people facing forward, more accurate customer flow information can be obtained, which can further support better business decision-making.
[0091] Exemplarily, in the step S5, if the human body orientation recognition result is side, the human target data is regarded as a wandering person. It should be understood that in an actual scenario, if a person appears near the store entrance facing sideways, it usually means that the person is not clearly moving towards the direction of entering the store, but may be considering whether to enter the store, looking at the window display, or just passing by. This behavior pattern indicates that they have not yet made a decision to become a real customer, so they are classified as "wandering people".
[0092] Exemplarily, in the step S6, if the human body orientation recognition result is back, the human target data is regarded as a person leaving the store and the passenger flow statistical count is decremented by one. It should be understood that in an actual scenario, if a person appears near the store entrance facing backwards, it usually means that the person is leaving or about to leave the store. People naturally turn their backs to the entrance direction when leaving the store, so the back orientation is a strong signal indicating that this person is no longer a customer inside the store but is walking out of the store.
[0093] Exemplarily, in the steps S7 and S8, object re-identification is performed on each piece of human target data in the set of human target data to obtain a set of object re-identification results, and in response to the object re-identification results being the same entering-store personnel, the human target data is not double-counted. It should be understood that performing object re-identification on each piece of human target data in the set of human target data to obtain a set of object re-identification results is to ensure that the same person is not double-counted during the passenger flow statistics process and to accurately distinguish different individuals. In actual application scenarios, customers may enter and exit the store multiple times, or non-customer personnel such as staff and deliverymen may also appear within the monitoring range. Without discrimination, these situations will cause the passenger flow statistics data to be distorted, affecting the decisions made by merchants based on this data. Through object re-identification technology, the system can identify the same individual at different time points and camera perspectives, thereby avoiding the problem of double-counting, while also effectively eliminating the data interference of non-customer personnel and improving the accuracy and reliability of passenger flow statistics.
[0094] In one embodiment, performing object re-identification on each piece of human target data in the set of human target data to obtain a set of object re-identification results includes: First, it is necessary to extract the features of each detected human target from the monitoring image, including but not limited to significant features such as appearance, posture, and clothing. Then, a powerful feature encoder is trained using a deep learning model, and this encoder can convert complex and variable human targets into compact and discriminative feature vectors. Next, the system will compare all the human target feature vectors detected in the current frame with the previously recorded historical feature library to find the matching item with the highest similarity. If a match is found, it is considered that the same person appears again; if no sufficiently similar result is found, it is regarded as a newly appeared individual and added to the historical feature library.
[0095] In summary, the passenger flow statistics analysis method based on reid assistance according to the embodiments of the present application is elucidated. After performing human target detection on the collected monitoring images of the door scene, it uses image analysis and recognition algorithms based on artificial intelligence and machine vision to analyze each piece of human target data, so as to respectively capture the significant feature representations of the human target in each piece of human target data, which helps to perform human orientation discrimination to obtain human orientation recognition results, including front, back, and side. In this way, it is possible to perform behavior pattern recognition on each human target object at the store door, thereby performing orientation recognition and detection on it, which helps subsequent passenger flow statistics analysis and provides strong support for business decisions.
[0096] Figure 7Schematic block diagram of the passenger flow statistics analysis system based on reid assistance according to an embodiment of the present application. As Figure 7 shown, the passenger flow statistics analysis system 100 based on reid assistance includes: a door scene monitoring image acquisition module 110 for acquiring door scene monitoring images collected by a camera deployed at the entrance of a store; a human target detection module 120 for performing human target detection on the door scene monitoring images to obtain a set of human target data; a human orientation recognition module 130 for using a trained reid assistance model to perform human orientation recognition on each piece of human target data in the set of human target data to obtain a set of human orientation recognition results; an in-store personnel statistics module 140 for, if the human orientation recognition result is positive, regarding the human target data as in-store personnel and incrementing the passenger flow statistics count by one; a wandering personnel determination module 150 for, if the human orientation recognition result is side-facing, regarding the human target data as wandering personnel; and an out-store personnel statistics module 160 for, if the human orientation recognition result is back-facing, regarding the human target data as out-store personnel and decrementing the passenger flow statistics count by one.
[0097] Here, those skilled in the art can understand that the specific operations of each module and unit in the above passenger flow statistics analysis system based on reid assistance have been introduced in detail in the description of the reid assistance-based passenger flow statistics analysis method above, and therefore, the repeated description thereof will be omitted. Figures 1 to 6 of the reid assistance-based passenger flow statistics analysis method, and thus, the repeated description thereof will be omitted.
[0098] An embodiment of the present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it causes the computer to implement the methods in the above embodiments of the present application.
[0099] An embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions run on a computer, it causes the computer to implement the methods in the above embodiments of the present application.
[0100] An embodiment of the present application also provides a chip, which includes a circuit for executing the methods in the above embodiments of the present application.
[0101] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0102] In the embodiments of the present application, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no restrictive effect on the position, order, priority, quantity, content, etc. of the described objects. The use of prefix words such as ordinal numbers for distinguishing described objects in the embodiments of the present application does not constitute a restriction on the described objects. The statements of the described objects shall refer to the descriptions in the context of the claims or embodiments, and no redundant restrictions shall be constituted due to the use of such prefix words.
[0103] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0104] In each embodiment of the present application, if there is no special description and logical conflict, the terms and / or descriptions among the various embodiments are consistent and can be mutually referred to. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0105] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0107] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A passenger flow statistical analysis method based on reid assistance, characterized in that: include: Obtaining door scene monitoring images collected by cameras deployed at the store entrance; Performing human target detection on the door scene monitoring image to obtain a set of human target data; Using the trained ReID auxiliary model, human body orientation recognition is performed on each human body target data in the set of human body target data to obtain a set of human body orientation recognition results; If the human body orientation recognition result is positive, the human body target data is regarded as a person entering the store and the customer flow statistics count is increased by one; If the human body orientation recognition result is sideways, the human body target data is regarded as a wandering person; If the human body is facing the back, the human body target data is regarded as a person leaving, and the passenger flow statistics count is reduced by one; Performing object re-identification on each human target data in the set of human target data to obtain a set of object re-identification results; In response to the object re-identification result being the same person entering the store, not repeatedly counting the human target data; Wherein, using the trained reid auxiliary model to perform human body orientation recognition on each human body target data in the set of human body target data to obtain a set of human body orientation recognition results, including: Inputting the human target data into a human multi-scale feature extractor based on a dynamic convolutional labeling model to obtain a human target multi-scale image feature coding map; Performing feature selection on the multi-scale image feature coding map of the human target to obtain a significant feature coding map of the human target image; Performing human body orientation discrimination based on the human body target image salient feature coding map to obtain the human body orientation recognition result; Among them, performing body orientation discrimination based on the significant feature coding map of the human target image to obtain the body orientation recognition result includes: inputting the significant feature coding map of the human target image into a body orientation discriminator based on a classifier to obtain the body orientation recognition result.
2. The passenger flow statistical analysis method based on reid assistance according to claim 1 is characterized in that: Inputting the human target data into a human multi-scale feature extractor based on a dynamic convolutional labeling model to obtain a human target multi-scale image feature coding map, including: Performing local image block segmentation on the human target data to obtain a set of human target local image blocks; Input each human target local image block in the set of human target local image blocks into a human target local semantic feature extractor based on CNN to obtain a set of human target local feature maps; The set of local feature maps of the human target is spliced into a global representation feature map of the human target and then input into a global association encoder of human target features based on a Transformer model to obtain a multi-scale image feature coding map of the human target.
3. The passenger flow statistical analysis method based on reid assistance according to claim 2 is characterized in that: Performing feature selection on the multi-scale image feature coding map of the human target to obtain a significant feature coding map of the human target image, including: Performing feature decomposition on the human target multi-scale image feature coding map to obtain a set of local feature vectors of the human target multi-scale image semantic coding; Inputting each human target multi-scale image semantic coding local feature vector in the set of human target multi-scale image semantic coding local feature vectors into an importance measurement module to obtain a set of human target multi-scale image semantic coding local feature importance score values; Based on the set of importance score values of local features of the multi-scale image semantic coding of the human target, feature selection and feature shape reshaping are performed on the set of local feature vectors of the multi-scale image semantic coding of the human target to obtain a significant feature coding map of the human target image.
4. The passenger flow statistical analysis method based on reid assistance according to claim 3 is characterized in that: The method comprises performing feature decomposition on the multi-scale image feature coding map of the human target to obtain a set of local feature vectors of the multi-scale image semantic coding of the human target, including: performing feature decoupling and feature flattening on the multi-scale image feature coding map of the human target to obtain a set of local feature vectors of the multi-scale image semantic coding of the human target.
5. The passenger flow statistical analysis method based on reid assistance according to claim 4 is characterized in that: Based on the set of importance score values of the local features of the multi-scale image semantic coding of the human target, feature selection and feature shape reshaping are performed on the set of local feature vectors of the multi-scale image semantic coding of the human target to obtain a significant feature coding map of the human target image, including: Based on the set of importance score values of the local features of the multi-scale image semantic coding of the human target, the set of the local feature vectors of the multi-scale image semantic coding of the human target is arranged in descending order to obtain a descending sequence of the local feature vectors of the multi-scale image semantic coding of the human target; Based on the neighborhood features of each human target multi-scale image semantic coding local feature vector in the descending sequence of the human target multi-scale image semantic coding local feature vector, calculating the feature neighborhood activity of each human target multi-scale image semantic coding local feature vector to obtain a sequence of feature neighborhood activity; Based on the sequence of the feature neighborhood activity, feature selection is performed on the descending sequence of the local feature vectors of the multi-scale image semantic coding of the human target to obtain a descending sequence of the local feature vectors of the multi-scale image semantic coding of the human target after selection; The descending sequence of the semantically encoded local feature vectors of the selected human target multi-scale image is reshaped to obtain a significant feature encoding map of the human target image.
6. The passenger flow statistical analysis method based on reid assistance according to claim 5 is characterized in that: Based on the sequence of the feature neighborhood activity, feature selection is performed on the descending sequence of the local feature vectors of the multi-scale image semantic coding of the human target to obtain a descending sequence of the local feature vectors of the multi-scale image semantic coding of the human target after selection, including: Extracting the feature neighborhood activity corresponding to the local feature vector of the multi-scale image semantic encoding of the first human target from the sequence of feature neighborhood activity to obtain a first feature neighborhood activity; The first feature neighborhood activity is compared with a predetermined threshold, and in response to the first feature neighborhood activity being greater than or equal to the predetermined threshold, the first human target multi-scale image semantically encoded local feature vector is retained, otherwise the first human target multi-scale image semantically encoded local feature vector is deleted.
7. A passenger flow statistics analysis system based on reid assistance, used to execute the passenger flow statistics analysis method based on reid assistance according to claim 1, characterized in that: include: The door scene monitoring image acquisition module is used to acquire the door scene monitoring image collected by the camera deployed at the door of the store; A human target detection module, used for performing human target detection on the door scene monitoring image to obtain a set of human target data; A human body orientation recognition module, used for performing human body orientation recognition on each human body target data in the set of human body target data by using the trained ReID auxiliary model to obtain a set of human body orientation recognition results; A store entry personnel statistics module, configured to, if the human body orientation recognition result is positive, regard the human body target data as a store entry person and increase the customer flow statistics count by one; a wandering person determination module, configured to regard the human target data as a wandering person if the human body orientation recognition result is sideways; The leaving person statistics module is used to regard the human target data as a leaving person if the human body direction recognition result is the back, and the passenger flow statistics count is reduced by one.
Citation Information
Patent Citations
Passenger flow volume statistics method and device
CN106295788A
Improved 4S store passenger flow statistical algorithm based on pedestrian re-identification
CN119359359A