Pedestrian data processing method and device, and storage medium

By combining visual and terminal signals, and utilizing the trained target detection model and wireless terminal location information, the accuracy problem of pedestrian counting was solved, achieving more accurate people counting.

CN120953357BActive Publication Date: 2026-03-27BEIJING BIG DATA CENT
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, pedestrian counting methods based on vision and terminal signals each have their own accuracy issues. Vision methods are easily affected by changes in lighting and occlusion, while terminal signal methods are affected by multipath effects, resulting in inaccurate pedestrian counting results.

Method used

By combining methods based on target pedestrian images and terminal signals, the position information of pedestrians in the world coordinate system is determined through a trained target detection model, and the population statistics are calculated by combining the position information of wireless terminals.

Benefits of technology

It improves the accuracy of pedestrian counting, reduces the impact of occlusion and multipath effects on the statistical results, and provides more reliable population statistics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953357B_ABST
    Figure CN120953357B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of pedestrian data processing, in particular, the present application provides a kind of pedestrian data processing method, device and storage medium.The processing method comprises: based on the at least two target pedestrian images of target world area successively acquired in target period, determine the first world position information corresponding to each target pedestrian in each target pedestrian image in world coordinate system.Based on the at least two terminal signal sets of target world area successively acquired in target period, determine the second world position information corresponding to each wireless terminal in each terminal signal set in world coordinate system.Based on the first world position information corresponding to each target pedestrian respectively and the second world position information corresponding to each wireless terminal respectively, determine the number of people statistical result corresponding to target world area in target period.The above processing method can improve the accuracy of the number of people statistical result.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pedestrian data processing, and more particularly to a pedestrian data processing method, a pedestrian data processing device and a storage medium. BACKGROUND

[0002] In the fields of intelligent security, smart business and city management, accurate statistics of the number of pedestrians is the basis for realizing applications such as passenger flow analysis, density monitoring and abnormal early warning. Accurate crowd counting results can not only reflect the personnel density and flow trend in the region, but also provide important data support for decision-making such as resource scheduling, emergency response and space utilization optimization, and have significant social value and commercial significance. Therefore, how to accurately count the number of pedestrians is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0003] The present application is proposed in consideration of the above problems. The present application provides a pedestrian data processing method, a pedestrian data processing device and a storage medium.

[0004] According to an aspect of the present application, a pedestrian data processing method is provided, which comprises:

[0005] Based on at least two target pedestrian images of a target world region acquired in sequence within a target period, first world position information corresponding to each target pedestrian in the world coordinate system is determined; based on at least two terminal signal sets of the target world region acquired in sequence within the target period, second world position information corresponding to each wireless terminal in the world coordinate system is determined, wherein for each target pedestrian image, the acquisition time of the target pedestrian image is the same as the acquisition time of the terminal signal set corresponding to the target pedestrian image; based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal, a number of people statistical result of the target world region within the target period is determined.

[0006] Exemplarily, based on at least two target pedestrian images of a target world region acquired in sequence within a target period, first world position information corresponding to each target pedestrian in the world coordinate system is determined, comprising:

[0007] For each target pedestrian image in the at least two target pedestrian images,

[0008] the target pedestrian image is input into the trained target detection model to obtain first image position information corresponding to each target pedestrian in the target pedestrian image, wherein the first image position information is used to represent the position information of the target pedestrian in the target pedestrian image;

[0009] For each target pedestrian in the target pedestrian image, first world position information corresponding to the target pedestrian is determined based on the first image position information corresponding to the target pedestrian and a first preset relationship, wherein the first preset relationship is used to represent a corresponding relationship between an image coordinate system of the target pedestrian image and a world coordinate system.

[0010] Exemplarily, the trained target detection model is trained based on a training pedestrian image, and annotation information of at least one target region corresponding to the training pedestrian image, wherein the at least one target region includes a head-shoulder region and / or a full-body region of a training pedestrian in the training pedestrian image, the annotation information of the head-shoulder region is used to represent position information of a region from a head to a shoulder of the training pedestrian in the training pedestrian image, and the annotation information of the full-body region is used to represent position information of a region from a head to a foot of the training pedestrian in the training pedestrian image.

[0011] Exemplarily, based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal, a people count result of the target world region in the target period is determined, including:

[0012] For each target pedestrian, a first motion trajectory of the target pedestrian in the world coordinate system in the target period is determined based on the first world position information corresponding to the target pedestrian;

[0013] For each wireless terminal, a second motion trajectory of the wireless terminal in the world coordinate system in the target period is determined based on the second world position information corresponding to the wireless terminal;

[0014] Based on the first motion trajectory of each target pedestrian and the second motion trajectory of each wireless terminal, a wireless terminal corresponding to each target pedestrian is determined, wherein for a target pedestrian corresponding to a wireless terminal, a coincidence degree of the first motion trajectory of the target pedestrian and the second motion trajectory of the wireless terminal is higher than a preset threshold;

[0015] Based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal or the second world position information corresponding to the wireless terminal corresponding to the target pedestrian, the people count result of the target world region in the target period is determined.

[0016] Exemplarily, the target world region includes an entrance region and an exit region, and the people count result includes a pedestrian inflow amount and a pedestrian outflow amount, and the people count result of the target world region in the target period is determined based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal, including:

[0017] determining, based on the first world position information corresponding to the target pedestrian with the wireless terminal, the entrance world position information corresponding to the entrance area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, a first inflow amount of the target period in the entrance area, wherein the head orientation information is used to represent the orientation of the head of the target pedestrian;

[0018] determining, based on the first world position information corresponding to the target pedestrian with the wireless terminal, the exit world position information corresponding to the exit area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, a first outflow amount of the target period in the exit area.

[0019] Exemplarily, based on the first world position information corresponding to each target pedestrian, and the second world position information corresponding to each wireless terminal, determining a people counting result of the target world area in the target period, comprising:

[0020] determining, based on the first world position information corresponding to each target pedestrian, a visual counting result in the target period, wherein the visual counting result is used to represent the total number of target pedestrians in the target world area;

[0021] determining, based on the second world position information corresponding to each wireless terminal, a wireless counting result in the target period, wherein the wireless counting result is used to represent the total number of wireless terminals in the target world area;

[0022] determining, based on the visual counting result and the wireless counting result, the people counting result of the target world area in the target period.

[0023] Exemplarily, the target world area includes the entrance area and the exit area, and the visual counting result includes the first inflow amount and the first outflow amount, wherein the visual counting result in the target period is determined based on the first world position information corresponding to each target pedestrian, comprising:

[0024] determining, based on the first world position information corresponding to each target pedestrian, the entrance world position information corresponding to the entrance area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, the first inflow amount of the target period in the entrance area, wherein the head orientation information is used to represent the orientation of the head of the target pedestrian;

[0025] determining, based on the first world position information corresponding to each target pedestrian, the exit world position information corresponding to the exit area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, the first outflow amount of the target period in the exit area.

[0026] Exemplarily, the above processing method further comprises:

[0027] determine a pedestrian flow directed graph based on the respective number of people statistics of each target world region, wherein the pedestrian flow directed graph comprises a plurality of target points connected via directed connection lines, the plurality of target points respectively correspond to the plurality of target world regions one by one, for a directed connection line from a target point corresponding to a first world region in the plurality of target world regions to a target point corresponding to a second world region in the plurality of target world regions, the directed connection line is used to represent that a target pedestrian unidirectionally passes from the first world region to the second world region, and the directed connection line corresponds to a pedestrian flow amount, the pedestrian flow amount is used to represent a total amount of pedestrians moving from the first world region to the second world region within a target time period.

[0028] According to still another aspect of the present application, there is also provided a pedestrian data processing apparatus, comprising:

[0029] a first world position information determination module configured to determine, based on at least two target pedestrian images of a target world region acquired in sequence within a target time period, first world position information corresponding to each target pedestrian in each target pedestrian image in a world coordinate system;

[0030] a second world position information determination module configured to determine, based on at least two terminal signal sets of the target world region acquired in sequence within the target time period, second world position information corresponding to each wireless terminal in each terminal signal set in the world coordinate system, wherein the at least two target pedestrian images respectively correspond to the at least two terminal signal sets one by one, and for each target pedestrian image, an acquisition time of the target pedestrian image is the same as an acquisition time of the terminal signal set corresponding to the target pedestrian image;

[0031] a number of people statistics determination module configured to determine, based on the respective first world position information of each target pedestrian and the respective second world position information of each wireless terminal, number of people statistics corresponding to the target world region within the target time period.

[0032] According to still another aspect of the present application, there is also provided a storage medium. The storage medium stores program instructions, which when executed, are used to execute the above-mentioned pedestrian data processing method.

[0033] According to the above scheme of the embodiment of the present application, the first world position information corresponding to each pedestrian in each target pedestrian image can be generated based on the at least two target pedestrian images of the target world region acquired in sequence within the target period. Then, the second world position information corresponding to each wireless terminal in each terminal signal set can be generated based on the at least two terminal signal sets of the target world region acquired in sequence within the target period. Finally, the number of people statistical result of the target world region within the target period can be determined based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal. The above scheme can determine the position information of each pedestrian in the target world region through the target pedestrian image and the terminal signal set, which can reduce the influence of the problems such as occlusion and multipath effect of the pedestrian on the number of people statistical result, and further improve the accuracy of the number of people statistical result. BRIEF DESCRIPTION OF DRAWINGS

[0034] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0035] Figure 1 a schematic flow chart of a pedestrian data processing method according to an embodiment of the present application is shown;

[0036] Figure 2 a schematic diagram of a pedestrian flow directed graph according to an embodiment of the present application is shown;

[0037] Figure 3 a schematic block diagram of a pedestrian data processing apparatus according to an embodiment of the present application is shown;

[0038] Figure 4 a schematic block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0039] In order to make the objects, technical solutions and advantages of the present application more apparent, the following will describe the example embodiments according to the present application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present application.

[0040] In the prior art, the method for counting the number of pedestrians mainly relies on single-mode perception technology, and common solutions include a visual analysis method based on video monitoring and a device detection method based on terminal signals (such as Wi-Fi signals, Bluetooth signals, etc.).

[0041] However, the single-mode perception technology has the following significant limitations. On the one hand, the visual-based pedestrian counting method is susceptible to factors such as light changes, occlusions, and visual angle blind areas, and especially in crowded or dynamic scenes, it often has problems such as missed detection and repeated detection, resulting in low accuracy of the counting result. On the other hand, the counting method based on terminal signals counts by determining the physical address of a terminal device (such as a mobile phone), which is not affected by visual occlusions, but the positioning accuracy is affected by multipath effects, making it difficult to accurately reflect the real number of people, which also leads to low accuracy of the counting result.

[0042] To at least partially solve the above problems, an embodiment of the present application provides a pedestrian data processing method. Figure 1 A schematic flowchart of a pedestrian data processing method according to an embodiment of the present application is shown. As shown in Figure 1 The processing method can include the following steps S110-S130.

[0043] In step S110, based on at least two target pedestrian images of a target world region acquired in sequence within a target time period, the first world position information corresponding to each target pedestrian in a world coordinate system is determined for each target pedestrian in each target pedestrian image.

[0044] The target world region can be a physical space region (for example, an entrance area of a station, an exit area of a station, a waiting area of a station, a platform area of a station, an entrance gate area of a subway, an exit area of a shopping mall, etc.) that needs to be counted, and the target time period can be a time interval in the target world region that needs to be counted. For example, the number of people in the entrance area of a train station from 10:00 to 10:10 on a certain day needs to be counted, and the time period from 10:00 to 10:10 on that day can be set as the target time period, and the entrance area of the train station can be set as the target world region. It can be understood that the above-mentioned region that needs to be counted and the above-mentioned time period that needs to be counted can be flexibly set by the user according to the actual situation.

[0045] The target pedestrian can be a pedestrian located in the target world region. Each target pedestrian image can include all visible pedestrians (i.e., the above-mentioned target pedestrians) in the target world region. Each target pedestrian image can be acquired by an image acquisition device (such as a network camera, a panoramic camera, etc.) deployed in the target world region or capable of capturing the complete target world region.

[0046] The origin of the world coordinate system, the extension direction of the horizontal coordinate axis (i.e., the x-axis), the extension direction of the vertical coordinate axis (i.e., the y-axis), and the extension direction of the vertical coordinate axis (i.e., the z-axis) can be determined by the user according to the actual situation. For example, still taking the above-mentioned target world area as the entrance area of a train station, a certain fixed point (for example, the corner of the floor tile, the column of the entrance area, the center point of the entrance door of the entrance area, etc.) in the entrance area can be taken as the origin of the world coordinate system. The horizontal width direction of the entrance area hall (for example, from the left security check to the right security check) is taken as the extension direction of the horizontal coordinate axis. The direction along the passenger's direction of entering the station (for example, from the entrance of the entrance area to the entrance gate, and then to the platform) is taken as the extension direction of the vertical coordinate axis, and the direction perpendicular to the plane of the horizontal coordinate axis and the vertical coordinate axis is taken as the extension direction of the z-axis.

[0047] The first world position information can be used to represent the position of the target pedestrian in the world coordinate system. For example, the pixel coordinates of the target pedestrian in the corresponding target pedestrian image (for example, the pixel coordinates of the feet of the target pedestrian in the corresponding target pedestrian image), and the geometric mapping relationship between the image coordinate system of the target pedestrian image and the world coordinate system (which can be determined by camera calibration) can be used to calculate the coordinate position of the corresponding pixel of the target pedestrian in the world coordinate system (i.e., the first world position information). It can be understood that the first world position information can be used to represent the position of the feet or head of the target pedestrian in the world coordinate system.

[0048] The present application provides an example for reference. Taking the target world area as the entrance area of train station A and the target time period as from 10:00 to 10:10 on a certain day, the target pedestrian images T1 and T2 of the target world area can be obtained in sequence according to the monitoring camera deployed in the entrance area of train station A. The first world position information corresponding to each target pedestrian in the target pedestrian image T1 can be determined through the above process. Based on the same process, the first world position information corresponding to each target pedestrian in the target pedestrian image T2 can be determined. It can be understood that there can be images located outside the target world area in the images obtained by the image acquisition device deployed in the target world area. The images corresponding to the target world area in the images can be retained through image cropping, and the cropped images can be taken as the above-mentioned target pedestrian images. It can also be understood that the image acquisition device for obtaining the target pedestrian images can also be an image acquisition device located outside the target world area and facing the target world area.

[0049] In step S120, based on the at least two terminal signal sets of the target world area obtained in sequence within the target time period, the second world position information corresponding to each wireless terminal in each terminal signal set in the world coordinate system is determined.

[0050] The terminal signal set can include terminal signals emitted by all wireless terminals in the target world region and received by the wireless access point (AP) at the time of acquisition. The wireless access point can support communication protocols such as Wi-Fi, Bluetooth, etc. Each terminal signal emitted by the wireless terminal can include a Wi-Fi probe request, a Bluetooth beacon broadcast packet, etc. Taking one of the terminal signal sets as an example, the terminal signal set includes a terminal signal emitted by the wireless terminal Z1. The wireless access point or the positioning system connected thereto can extract the following feature information from the received terminal signal: received signal strength indicator (RSSI), media access control address (MAC address) of the wireless terminal sending the signal, signal reception timestamp, etc. Then, according to the above feature information and a preset positioning algorithm (for example, a triangular positioning algorithm, a fingerprint positioning algorithm, etc.), the position of the wireless terminal Z1 in the coordinate system preset by the wireless access point can be determined. According to the geometric mapping relationship between the coordinate system preset by the wireless access point and the world coordinate system, the position of the wireless terminal Z1 in the coordinate system preset by the wireless access point is converted into the position of the wireless terminal Z1 in the world coordinate system (i.e., the second world position information). The second world position information can be used to represent the absolute spatial position of the wireless terminal Z1 in the target world region.

[0051] The at least two target pedestrian images and the at least two terminal signal sets are one-to-one corresponding. For each target pedestrian image, the acquisition time of the target pedestrian image is the same as the acquisition time of the terminal signal set corresponding to the target pedestrian image. For example, to ensure the accurate association between the target pedestrian image and the terminal signal set, the time reference of the image acquisition device and the wireless access point can be calibrated by using high-precision clock synchronization technology (for example, using network time protocol (NTP) or global positioning system (GPS) time synchronization service). In actual operation, considering factors such as data processing delay and network transmission delay, a certain range of time deviation can be allowed, and the deviation can be adjusted according to specific circumstances. For example, assuming that at a certain time t, a target pedestrian image T1 is acquired from a monitoring camera located in the target world region, and at the same time t, a terminal signal set S1 including terminal signals emitted by several wireless terminals is received by the wireless access point AP1. Through the above steps, the first world position information of each target pedestrian in the target pedestrian image T1 and the second world position information of each wireless terminal in the terminal signal set S1 can be obtained.

[0052] At step S130, based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal, a number of people in the target world region in the target period is determined.

[0053] The number of people in the target world region in the target period can include a peak of people flow, a distribution of target pedestrians, a change of target pedestrians, etc. For example, the number of people in the target world region in the target period is the change of target pedestrians, the total number of target pedestrians in the first target pedestrian image obtained in the target period (hereinafter referred to as the total number of target pedestrians) is 10, the total number of target pedestrians in the last target pedestrian image obtained in the target period is 20, the total number of wireless terminals corresponding to the first terminal signal set obtained (hereinafter referred to as the total number of wireless terminals) is 15, and the total number of wireless terminals corresponding to the last terminal signal set obtained is 22. One of the total number of target pedestrians in the first target pedestrian image obtained in the target period and the total number of wireless terminals corresponding to the first terminal signal set obtained can be determined as an initial total number of target pedestrians, and one of the total number of target pedestrians in the last target pedestrian image obtained in the target period and the total number of wireless terminals corresponding to the last terminal signal set obtained can be determined as an end total number of target pedestrians. Then, the absolute value of the difference between the end total number of target pedestrians and the initial total number of target pedestrians is taken as the change of target pedestrians.

[0054] It can be understood that the determination rule of the initial total number of target pedestrians and the end total number of target pedestrians can be affected by the environment of the target world region. For example, there are more metal products in the target world region, so that the influence of the multipath effect on the number of people statistics results can be reduced. The total number of target pedestrians in the first target pedestrian image obtained in the target period can be taken as the initial total number of target pedestrians, and the total number of target pedestrians in the last target pedestrian image obtained in the target period can be taken as the end total number of target pedestrians. In the case that the light in the target world region is relatively dark, the total number of wireless terminals corresponding to the first terminal signal set obtained in the target period can be taken as the initial total number of target pedestrians, and the total number of wireless terminals corresponding to the last terminal signal set obtained in the target period can be taken as the end total number of target pedestrians. It can also be understood that the determination rule of the initial total number of target pedestrians and the end total number of target pedestrians is not limited to the influence of the environment of the target world region.

[0055] In some possible embodiments, in the case that the number of people in the target world region in the target period corresponds to a change in the number of target pedestrians, if there is a device (for example, an entry gate, an exit gate, etc.) that can record the number of people entering and leaving in the target world region, the change in the number of target pedestrians can be determined according to the recording data of the device that records the number of people entering and leaving, the total number of wireless terminals corresponding to each terminal signal set, and the total number of target pedestrians corresponding to each target pedestrian image.

[0056] In one embodiment, if the image acquisition device used to acquire the target pedestrian image is damaged or blocked, so that the image of part of the target world region cannot be acquired, the total number of target pedestrians in the part of the target world region can be determined only according to the terminal signal set.

[0057] In yet another embodiment, a particle filter technique can be employed to fuse the position of each target pedestrian in the target world region based on the first world position information of each target pedestrian and the second world position information of each wireless terminal, and to obtain more accurate third world position information of each target pedestrian. It can be understood that if the third world position information of at least one target pedestrian is located in the alarm region in the target world region and the stay time is greater than the preset stay time, an alarm can be sent to the staff of the place where the target world region is located. It can be understood that when it is determined that the third world position information of at least one target pedestrian is located in the alarm region in the target world region and the stay time is greater than the preset stay time, a camera capable of obtaining the image of the alarm region can be automatically called to determine whether there is really a target pedestrian located in the alarm region in the target world region and the stay time is greater than the preset stay time. For example, if the camera capable of obtaining the image of the alarm region is a fixed camera, the latest video clip obtained by the camera can be automatically called. If the camera capable of obtaining the image of the alarm region is a PTZ (Pan / Tilt / Zoom, i.e. a pan / tilt / zoom camera) camera, a control instruction can be automatically generated according to the position information of the alarm region in the world coordinate system, so that the PTZ camera can quickly aim at the alarm region. The dispatching instruction of the PTZ camera can be reasonably set to reduce the occurrence of a PTZ camera processing multiple alarm notifications. It can also be understood that the dispatched camera can perform high-precision analysis. For example, whether the pedestrian triggering the alarm exists (for example, whether there is really a pedestrian entering the alarm region), detailed identification of the pedestrian triggering the alarm (for example, face recognition, behavior recognition, clothing recognition, and carrying object recognition), continuous tracking of the pedestrian triggering the alarm (for example, if necessary, the pedestrian can be tracked across regions), and the like. The dispatched camera can determine whether to send an alarm notification to the staff of the place where the target world region is located according to the analysis result of the high-precision analysis.

[0058] In one embodiment, whether a wireless terminal is located in the alarm region in the target world region can be determined according to the second world position information of each wireless terminal. If a wireless terminal is located in the alarm region in the target world region, whether the wireless terminal is an authorized wireless terminal can be determined according to the MAC address of the wireless terminal. If the wireless terminal is an unauthorized wireless terminal, an alarm can be sent to the staff of the place where the target world region is located.

[0059] In yet another embodiment, whether there is a wireless terminal staying in the alarm area of the target world region for a time greater than a preset staying time can be determined according to the second world location information corresponding to each wireless terminal respectively. If there is, an alarm can be sent to the staff of the place where the target world region is located.

[0060] In still another embodiment, the total number of target pedestrians in the target world region at each target pedestrian image acquisition time can be determined according to the total number of target pedestrians corresponding to each target pedestrian image of the target world region in the target period, and the total number of wireless terminals corresponding to each terminal signal set. Then, the area density value of target pedestrians in the target world region at each target pedestrian image acquisition time can be determined according to the total number of target pedestrians in the target world region at each target pedestrian image acquisition time, and the area of the target world region. In the case where the area density value is greater than a preset area density threshold, an alarm can be sent to the staff of the place where the target world region is located.

[0061] In yet another embodiment, the peak period of people flow in the target world region in a future certain time can be predicted by a time series model. It can be understood that the time series model can be a Long Short-Term Memory Network (LSTM) or a Prophet model.

[0062] According to the above scheme of the embodiments of the present application, the first world location information corresponding to each pedestrian in each target pedestrian image can be determined based on at least two target pedestrian images of the target world region acquired in sequence in the target period. Then, the second world location information corresponding to each wireless terminal in each terminal signal set can be determined based on at least two terminal signal sets of the target world region acquired in sequence in the target period. Finally, the number of people statistics result corresponding to the target world region in the target period can be determined based on the first world location information corresponding to each target pedestrian respectively, and the second world location information corresponding to each wireless terminal respectively. The above scheme can determine the location information of each pedestrian in the target world region through the target pedestrian image and the terminal signal set, which can reduce the influence of the problems such as occlusion and multipath effect on the number of people statistics result, and further improve the accuracy of the number of people statistics result. In addition, the target pedestrian in the target pedestrian image is represented by the first world location information in the world coordinate system, and the wireless terminal is represented by the second world location information in the world coordinate system, which is also conducive to aligning the target pedestrian with the wireless terminal, so as to improve the accuracy of the number of people statistics result determined based thereon.

[0063] Exemplarily, for each of the at least two target pedestrian images, the step S110 of determining, based on the at least two target pedestrian images of the target world region acquired in sequence within the target period, the first world position information corresponding to each target pedestrian in each target pedestrian image in the world coordinate system, comprises a step S111 and a step S112.

[0064] In the step S111, the target pedestrian image is input into the trained target detection model to obtain the first image position information corresponding to each target pedestrian in the target pedestrian image.

[0065] The first image position information is used to represent the position information of the target pedestrian in the target pedestrian image. For example, the first image position information is in the form of pixel coordinates. For example, the first image position information can be the pixel coordinates of the feet of the target pedestrian in the corresponding target pedestrian image in the step S110. For another example, the first image position information can be a set of pixel coordinates corresponding to the boundary of the image region where the target pedestrian is located in the target pedestrian image. For another example, the first image position information can be in the form of a bounding box. The bounding box can be used to frame the image region of the target pedestrian in the target pedestrian image. The shape of the bounding box can include a rectangle, an ellipse or other geometric shapes, and is usually a rectangle. Each rectangular box is represented by the pixel coordinates of the first pixel in the upper left corner and the pixel coordinates of the last pixel in the lower right corner, and can also be represented by the pixel coordinates of the center point and the width and height of the bounding box.

[0066] The embodiment of the present application provides a training method of a target detection model for reference. A first training data set is obtained. The target detection model is iteratively trained through the training data set, and the model parameters of the target detection model are adjusted through a loss function until the training is completed. The first training data set can include target pedestrian images for training. The label of each target pedestrian image is the first image position information corresponding to each target pedestrian in the target pedestrian image for training. The condition for completing the training can be that the loss value calculated by the loss function tends to be stable, the number of iterations reaches a preset number, etc. The trained target detection model can output the first image position information of each target pedestrian in the target pedestrian image according to the input target pedestrian image.

[0067] In an embodiment, in order to improve the generalization ability and detection robustness of the target detection model in complex real-world scenarios, a data augmentation technique can be used to diversify the original target pedestrian images used for training, generating representative simulated training samples. Each training target pedestrian image can be accurately labeled to form an initial training dataset. On this basis, the training target pedestrian image can be enhanced by one or more of the following image transformation methods: light change simulation, motion blur, perspective transformation, small target replication and pasting, random occlusion, etc. The above light change simulation is to adjust the random brightness and contrast of the training target pedestrian image, generating overexposed (i.e. loss of detail in highlight areas of the image) and underexposed (i.e. insufficient information in the shadow area of the image) image samples to simulate the lighting conditions at different times (e.g. backlight, night, etc.). Motion blur simulation is to apply directional Gaussian blur to the pedestrian target area in the image to simulate the imaging blur phenomenon caused by the rapid movement of the target pedestrian, improving the model's ability to recognize dynamic targets. Perspective transformation simulation is to perform affine transformation or perspective transformation on the target pedestrian to simulate the deformation characteristics of images taken by cameras with different installation angles (e.g. high-angle downward view, low-angle upward view, etc.), enhancing the adaptability of the target detection model to multi-view targets. Small target replication can be to crop small size target pedestrians (e.g. distant target pedestrians, etc.) in the training target pedestrian image and paste them at different positions, scales and rotation angles in the background area of the same training target pedestrian image, increasing the number of small target samples and alleviating the problem of missed detection caused by the scarcity of small targets in the target detection model. Random occlusion simulation can add random shape occlusion blocks (e.g. rectangles, irregular polygons, etc.) around the target pedestrian target or in the local area of the body to simulate the visual effect of mutual occlusion in dense crowds, improving the detection stability of the target detection model in crowded scenes. The first training dataset generated by the above enhancement processing covers more interference factors and edge cases in the real world. The target detection model trained using this enhanced dataset can more accurately identify pedestrian targets under complex lighting, fast movement, different angles, partial occlusion, and long distance, significantly improving the detection performance and reliability of the system in actual deployment environments.

[0068] In an embodiment, in order to improve the detection accuracy of the target detection model on pedestrian targets in a specific scene, an anchor box clustering method based on the K-means++ clustering algorithm is used to determine the anchor box aspect ratio that is most suitable for the current application scenario. Specifically, a large number of training images containing target pedestrians (for example, the images in the first training data set described above) are first collected. When the first image position information is in the form of a bounding box, the bounding box corresponding to the target pedestrian in each training image can be extracted to obtain the width and height information of each bounding box. The width and height values of each bounding box are taken as a two-dimensional data point (width, height) to form a size data set composed of the width and height values of the bounding box. Then, the K-means++ clustering algorithm is used to perform clustering analysis on the size data set. Through iterative calculation, the above clustering algorithm divides the width and height values of all bounding boxes into several categories, and the clustering center of each category represents an optimal anchor box size pattern. According to actual detection requirements (for example, more small targets or concentrated width-height ratio, etc.), a suitable number of clusters K can be selected, and the most representative clustering center is selected as the anchor box aspect ratio of the target detection model. For example, in an embodiment, the width-height ratio of the optimal anchor box obtained after clustering analysis is [0.34, 0.62], indicating that the detection box of the target pedestrian in this scene is on average relatively slender, and it is suitable to use anchor boxes with such proportions for matching and prediction. The anchor box size determined in the above manner is more consistent with the distribution characteristics of the actual target, and compared with the use of fixed or empirical anchor box size, it can significantly improve the positive sample matching rate in the target detection process, and thus improve the detection accuracy and recall rate of the target detection model.

[0069] In step S112, for each target pedestrian in the target pedestrian image, the first world position information corresponding to the target pedestrian is determined based on the first image position information corresponding to the target pedestrian and the first preset relationship.

[0070] The first preset relationship is used to represent the corresponding relationship between the image coordinate system of the target pedestrian image and the world coordinate system. For example, the first preset relationship can be the geometric mapping relationship between the image coordinate system and the world coordinate system of the target pedestrian image in step S110 described above.

[0071] The present application provides an example for reference. In the form of a detection frame, the first image position information corresponding to the target pedestrian is defined by the pixel coordinates of the first pixel in the upper left corner and the pixel coordinates of the last pixel in the lower right corner. In the case where the first image position information corresponding to the target pedestrian is (x1, y1) and (x2, y2), the first image position information (x1, y1) in the world coordinate system can be determined as (x3, y3, z3) according to the first preset relationship, and the first image position information (x2, y2) in the world coordinate system can be determined as (x4, y4, z4). The position information of the feet of the target pedestrian in the world coordinate system (i.e., the first world position information) can be determined according to the coordinates (x3, y3, z3) in the world coordinate system and the coordinates (x4, y4, z4) in the world coordinate system.

[0072] In one embodiment, if the target world region corresponding to the target pedestrian image contains an alarm region, it can be determined whether there is a target pedestrian located in the alarm region according to the relative positional relationship between the detection frame corresponding to each target pedestrian in the target pedestrian image and the alarm region in the target pedestrian image. Further, it can be determined whether to send an alarm notification to the staff of the place where the target world region is located according to the stay time of the target pedestrian in the alarm region. When it is determined that there is a target pedestrian located in the alarm region and the stay time is greater than a preset stay time, the camera capable of obtaining the image of the above-mentioned alarm region is automatically called to obtain an analysis result of high-precision analysis (for details, please refer to the content in step S110 above, which will not be repeated here). According to the analysis result, it is determined whether to send an alarm notification to the staff of the place where the target world region is located.

[0073] According to the above-mentioned scheme of the embodiments of the present application, for each target pedestrian image in the at least two target pedestrian images, the target pedestrian image can be input into the trained target detection model to obtain the first image position information corresponding to each target pedestrian in the target pedestrian image. Then, for each target pedestrian in the target pedestrian image, the first world position information corresponding to the target pedestrian is determined based on the first image position information corresponding to the target pedestrian and the first preset relationship. The above-mentioned scheme can detect the target pedestrian in the target pedestrian image through the trained target detection model, can accurately identify each target pedestrian in the target pedestrian image, and output the accurate first image position information thereof. Using the first preset relationship calibrated in advance, the first image position information corresponding to each target pedestrian can be quickly and accurately converted into the first world position information in the real world coordinate system. This provides a reliable data basis for subsequent calculation of the number of people, and is further conducive to improving the accuracy of the number of people.

[0074] Exemplarily, the trained target detection model is trained based on the training pedestrian images and the annotation information of the at least one target region corresponding to the training pedestrian images.

[0075] The at least one target region includes a head-shoulder region and / or a full-body region of the training pedestrian in the training pedestrian image, the annotation information of the head-shoulder region is used to represent the position information of the region from the head to the shoulder of the training pedestrian in the training pedestrian image, and the annotation information of the full-body region is used to represent the position information of the complete body region from the head to the foot of the training pedestrian in the training pedestrian image.

[0076] Each training pedestrian image can be annotated with at least one target region corresponding to the training pedestrian. The target region can be the region corresponding to the training pedestrian in the training pedestrian image. It can include the image region corresponding to part of the body or the whole body of the training pedestrian in the training pedestrian image. For example, the target region can include the region corresponding to the head to the shoulder of the training pedestrian in the training pedestrian image, and / or the region corresponding to the head to the foot of the training pedestrian in the training pedestrian image.

[0077] Each target region can correspond to annotation information. The annotation information of each target region can be used to represent the spatial position of the training pedestrian in the target region in the training pedestrian image, which corresponds to the above-mentioned first image position information, and can be used to supervise the learning of the target detection model. It can be understood that the form of the annotation information can be consistent with the form of the first image position information, that is, the form of the annotation information can also be expressed in the form of pixel coordinates. For details, please refer to the content of the first image position information in the above-mentioned step S111, which will not be repeated here. It can also be understood that the annotation information can be different according to the different parts of the training pedestrian included in the target region. For example, in the case where the target region includes the region corresponding to the head to the shoulder of the training pedestrian in the training pedestrian image, the annotation information of the target region can be used to represent the position of the local region from the head to the shoulder of the training pedestrian in the training pedestrian image (for example, the pixel coordinates of the head and the shoulder of the training pedestrian in the training pedestrian image, or the pixel coordinates of the first pixel in the upper left corner of the detection box and the pixel coordinates of the last pixel in the lower right corner of the detection box for framing the head to the shoulder of the training pedestrian). In the case where the target region includes the region corresponding to the head to the foot of the training pedestrian in the training pedestrian image, the annotation information of the target region can be used to represent the position of the complete body region from the head to the foot of the training pedestrian in the training pedestrian image (for example, the pixel coordinates of the head and the foot of the training pedestrian in the training pedestrian image, or the pixel coordinates of the first pixel in the upper left corner of the detection box and the pixel coordinates of the last pixel in the lower right corner of the detection box for framing the head to the foot of the training pedestrian).

[0078] The embodiment of the present application provides a training method of a target detection model for reference. A second training data set is obtained. The target detection model is iteratively trained through the second training data set, and the model parameters of the target detection model are adjusted through a loss function until the training is completed. The second training data set can include a plurality of training pedestrian images. The label corresponding to each training pedestrian image can be the annotation information of the target region corresponding to each training pedestrian in the training pedestrian image. The condition of the training completion can be that the loss value calculated through the loss function tends to be stable, the iteration number reaches a preset number, etc. The trained target detection model can output the first image position information of each target pedestrian in the target pedestrian image according to the input target pedestrian image (equivalent to the annotation information).

[0079] According to the above scheme of the embodiment of the present application, the trained target detection model can be trained based on the training pedestrian image, the annotation information of at least one target region corresponding to the training pedestrian image. The at least one target region includes the head-shoulder region and / or the whole body region of the training pedestrian in the training pedestrian image. The above scheme can make the target detection model learn the target regions of two different scales and features, that is, the head-shoulder region and the whole body region, so that the target detection model can more comprehensively understand the multi-dimensional features of the target pedestrian, and then the target pedestrian in the target pedestrian image can be detected by the target detection model as long as the target pedestrian exposes at least the upper body. This is beneficial to improve the accuracy of the target detection model in detecting the target pedestrian in the target pedestrian image, ensures the usability and reliability of the above processing method in real complex environment, and also provides a reliable data basis for subsequent calculation of the number of people, which is beneficial to improve the accuracy of the number of people.

[0080] Exemplarily, the step S130 of determining the number of people corresponding to the target world region in the target period based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal includes steps S131a to S134a.

[0081] In step S131a, for each target pedestrian, the first motion trajectory of the target pedestrian in the world coordinate system in the target period is determined based on the first world position information corresponding to the target pedestrian.

[0082] The present application provides an example for reference. After obtaining the first world position information of each target pedestrian in each frame of target pedestrian image in the target period, the first world position information of the same target pedestrian in different target pedestrian images can be matched and associated based on the spatio-temporal correlation between the continuous target pedestrian images to determine the first motion trajectory corresponding to the target pedestrian. For example, the first prediction world position information of each target pedestrian can be obtained by predicting the first world position information of each target pedestrian at the acquisition time of the current frame of target pedestrian image according to the first world position information, motion trend and appearance characteristics of each target pedestrian in the previous frame of target pedestrian image. Then, the actual first world position information of each target pedestrian at the acquisition time of the current frame of target pedestrian image is matched with the first prediction world position information of each target pedestrian (for example, the comprehensive position distance and appearance similarity between the two are calculated, and the Hungarian algorithm or other assignment algorithm is used to match the two), to determine the same target pedestrian in the previous frame of target pedestrian image and the current frame of target pedestrian image. A unique identity identifier is assigned to the same target pedestrian. The above process can be recursively performed frame by frame until the last frame of target pedestrian image in the target period is processed, so as to construct the continuous motion path of each target pedestrian in the target period as the first motion trajectory of each target pedestrian. It can be understood that the above process can be performed by selecting a suitable tracker, for example, a lightweight tracker such as Deep Learning-based Simple Online and Realtime Tracking (DeepSORT), ByteTrack, etc.

[0083] In step S132a, for each wireless terminal, the second motion trajectory of the wireless terminal in the world coordinate system in the target period is determined based on the second world position information corresponding to the wireless terminal.

[0084] The present application provides an example for reference. After obtaining the second world position information of each wireless terminal in each terminal signal set in the target period, the second world position information of the same wireless terminal at the acquisition time of different terminal signal sets can be associated across time according to the identification information (for example, MAC address) of each wireless terminal in each terminal signal set. Then, the multiple second world position information corresponding to each wireless terminal is arranged in ascending order according to the corresponding time stamp to form an ordered position sequence. Based on the position sequence, the position change process of each wireless terminal in the target period is determined, and the continuous motion path of each wireless terminal in the world coordinate system is constructed as the second motion trajectory of the wireless terminal.

[0085] In step S133a, each wireless terminal corresponding to each target pedestrian is determined based on the first motion trajectory of each target pedestrian and the second motion trajectory of each wireless terminal.

[0086] For a target pedestrian corresponding to a wireless terminal, the first motion trajectory of the target pedestrian and the second motion trajectory of the wireless terminal have a degree of coincidence higher than a preset threshold. For example, the similarity between each first motion trajectory and each second motion trajectory can be determined according to a preset trajectory similarity algorithm. The second motion trajectory with a similarity to the first motion trajectory greater than the preset threshold is determined as the second motion trajectory corresponding to the first motion trajectory. Then, the wireless terminal corresponding to the second motion trajectory is determined as the wireless terminal corresponding to the target pedestrian corresponding to the first motion trajectory. It can be understood that the preset trajectory similarity algorithm can be an algorithm for measuring the similarity between two spatiotemporal trajectories (i.e., the first motion trajectory and the second motion trajectory). For example, the preset trajectory similarity algorithm can be a dynamic time warping algorithm, a longest common subsequence algorithm, a Hausdorff distance algorithm, a spatial distance average algorithm, etc. It can also be understood that the preset threshold can be determined according to actual conditions. It can also be understood that there can be at least two second motion trajectories with a degree of coincidence higher than the preset threshold with a first motion trajectory. Taking the number of the at least two second motion trajectories as 2 as an example, in this case, the second motion trajectory with the highest degree of coincidence with the first motion trajectory among the two second motion trajectories is determined as the second motion trajectory corresponding to the first motion trajectory. It can also be understood that if there are at least two second motion trajectories with a degree of coincidence higher than the preset threshold with a first motion trajectory, the wireless terminals corresponding to the two motion trajectories are both determined as the wireless terminal corresponding to the target pedestrian corresponding to the first motion trajectory, which can be regarded as the target pedestrian carrying two wireless terminals.

[0087] In step S134a, the number of people in the target world area in the target period is determined based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal or the second world position information corresponding to the wireless terminal corresponding to the target pedestrian.

[0088] In one example, the number of people in the target world region in the target period is determined based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal. For example, in the case that the number of people in the target world region is the change of the target pedestrian in step S130 above, the difference between the total number of the first world position information corresponding to the target pedestrian corresponding to the wireless terminal in the first target pedestrian image obtained in the target period and the total number of the first world position information corresponding to the target pedestrian corresponding to the wireless terminal in the last target pedestrian image obtained in the target period can be taken as the change of the target pedestrian in the target world region in the target period, i.e. the number of people in the target world region in the target period.

[0089] In another example, the number of people in the target world region in the target period is determined based on the second world position information corresponding to the wireless terminal corresponding to the target pedestrian. For example, in the case that the number of people in the target world region is the change of the target pedestrian in step S130 above, the difference between the total number of the second world position information corresponding to the wireless terminal corresponding to the target pedestrian in the first terminal signal set obtained in the target period and the total number of the second world position information corresponding to the wireless terminal corresponding to the target pedestrian in the last terminal signal set obtained in the target period can be taken as the change of the target pedestrian in the target world region in the target period, i.e. the number of people in the target world region in the target period.

[0090] In one embodiment, the displacement change of each target pedestrian in each period can be determined according to the first motion trajectory corresponding to each target pedestrian in the target period and the acquisition time interval of adjacent target pedestrian images, and the motion speed and flow direction of each target pedestrian can be determined to determine whether there is a situation such as reverse movement, stagnation, abnormal gathering of target pedestrians in the target world region in the target period. The motion speed and flow direction of each wireless terminal can also be determined according to the second motion trajectory corresponding to each wireless terminal in the target period and the timestamp information of each terminal signal set. For the wireless terminal corresponding to the target pedestrian, the motion state of the wireless terminal can be used as supplementary information to correct or verify the motion speed and flow direction of the target pedestrian determined based on the target pedestrian image.

[0091] According to the scheme, for each target pedestrian, a first motion trajectory of the target pedestrian in the world coordinate system in the target period is determined based on the first world position information corresponding to the target pedestrian. Then, for each wireless terminal, a second motion trajectory of the wireless terminal in the world coordinate system in the target period is determined based on the second world position information corresponding to the wireless terminal. Then, each target pedestrian corresponding to a wireless terminal is determined based on the first motion trajectory of each target pedestrian and the second motion trajectory of each wireless terminal. Finally, the number of people in the target world region in the target period is determined based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal or the second world position information corresponding to the wireless terminal corresponding to the target pedestrian. The scheme can determine the target pedestrian corresponding to the wireless terminal by combining the first motion trajectory and the second motion trajectory, so as to accurately determine and count the target pedestrian, and improve the accuracy of the number of people.

[0092] For example, the target world region includes an entrance region and an exit region, and the number of people includes the number of people entering and the number of people leaving.

[0093] Each target world region can include at least one entrance region and at least one exit region. The entrance region can be used for the target pedestrian to enter the entrance region in one direction, and the exit region can be used for the target pedestrian to leave the entrance region in one direction. For example, the target world region can be a part of a train station, and the entrance of the part of the train station can be used as the entrance region of the part of the train station. The exit of the part of the train station can be used as the exit region of the part of the train station.

[0094] The step S134a of determining the number of people in the target world region in the target period based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal can include a step S134a1 and a step S134a2.

[0095] In the step S134a1, the number of people entering the entrance region in the target period is determined based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal, the entrance world position information corresponding to the entrance region in the world coordinate system, and the head direction information corresponding to each target pedestrian in the target pedestrian image.

[0096] The head orientation information is used to indicate the orientation of the head of the target pedestrian. For example, the head orientation information can be obtained according to the target detection model in the foregoing, which can be represented as the included angle between the head orientation of the target pedestrian and a preset direction. Embodiments of the present application provide a training manner of a target detection model capable of outputting head orientation information of a target pedestrian for reference. A third training data set is obtained. The target detection model is iteratively trained through the third training data set, and the model parameters of the target detection model are adjusted through a loss function until the training is completed. The third training data set can include a plurality of training pedestrian images. The label corresponding to each training pedestrian image can be the annotation information of the target region corresponding to each training pedestrian in the training pedestrian image, and the head orientation of each training pedestrian. The condition for completing the training can be that the loss value calculated by the loss function tends to be stable, the number of iterations reaches a preset number, etc. The trained target detection model can output the first image position information (equivalent to the annotation information) of each target pedestrian and the head orientation of each target pedestrian according to the input target pedestrian image.

[0097] The present application provides an example for reference. Based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal, the entrance world position information corresponding to the entrance region in the world coordinate system, it is determined whether the first world position information corresponding to the target pedestrian corresponding to the wireless terminal in each target pedestrian image is located in the entrance region. If it is determined that the target pedestrian M1 and the target pedestrian M2 in the first target pedestrian image acquired in the target period are located in the entrance region, the included angle between the head orientation of the target pedestrian M1 and the passing direction (in this example, the preset direction in the foregoing) vector of the entrance region is less than a preset angle (for example, 60°), the target pedestrian M1 is determined as the target pedestrian flowing into the target world region. If the included angle between the head orientation of the target pedestrian M2 and the passing direction is greater than the preset angle, it indicates that the target pedestrian M2 does not belong to the target pedestrian flowing into the target world region. Through the above process, the target pedestrian flowing into the target world region in each target pedestrian image can be determined. Further, the total number of target pedestrians flowing into the target world region in at least two target pedestrian images of the target world region acquired in the target period can be taken as the pedestrian inflow of the entrance region in the target period. It can be understood that the preset angle and the entrance world position information corresponding to the entrance region in the world coordinate system can be determined according to actual conditions.

[0098] In step S134a2, based on the first world position information corresponding to the target pedestrian corresponding to the wireless terminal, the exit world position information corresponding to the exit region in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image, the pedestrian outflow of the exit region in the target period is determined.

[0099] The present application provides an example for reference. Based on the first world position information corresponding to the target pedestrian with the wireless terminal, and the exit world position information corresponding to the exit area in the world coordinate system, it can be determined whether the first world position information corresponding to the target pedestrian with the wireless terminal in each target pedestrian image is located in the exit area. If it is determined that the target pedestrian M3 and the target pedestrian M4 in the first acquired target pedestrian image in the target period are located in the exit area, the angle between the head direction of the target pedestrian M3 and the passing direction vector of the exit area is less than a preset angle (for example, 60°), the target pedestrian M3 is determined as the target pedestrian flowing out of the target world area. If the angle between the head direction of the target pedestrian M4 and the passing direction is greater than the above-mentioned preset angle, it indicates that the target pedestrian M4 does not belong to the target pedestrian flowing out of the target world area. Through the above process, the target pedestrian flowing out of the target world area in each target pedestrian image can be determined. Further, the total number of target pedestrians flowing out of the target world area in at least two target pedestrian images of the target world area acquired in sequence in the target period can be taken as the pedestrian outflow of the exit area in the target period. It can be understood that the above-mentioned preset angle and the exit world position information corresponding to the exit area in the world coordinate system can be determined according to actual conditions.

[0100] According to the above-mentioned scheme of the embodiment of the present application, the pedestrian inflow of the entrance area in the target period can be determined based on the first world position information corresponding to the target pedestrian with the wireless terminal, the entrance world position information corresponding to the entrance area in the world coordinate system, and the head direction information corresponding to each target pedestrian in the target pedestrian image. Then, the pedestrian outflow of the exit area in the target period can be determined based on the first world position information corresponding to the target pedestrian with the wireless terminal, the exit world position information corresponding to the exit area in the world coordinate system, and the head direction information corresponding to each target pedestrian in the target pedestrian image. Compared with the method of determining the pedestrian inflow and the pedestrian outflow only by relying on the first world position information corresponding to the target pedestrian with the wireless terminal, the entrance world position information, and the exit world position information, the above-mentioned scheme is beneficial to improve the accuracy of determining the pedestrian inflow and the pedestrian outflow, and further can ensure the reliability of the statistical result.

[0101] Exemplarily, the step S130 of determining the person statistical result of the target world area in the target period based on the first world position information corresponding to each target pedestrian and the second world position information corresponding to each wireless terminal comprises the step S131b to the step S133b.

[0102] In the step S131b, the visual statistical result in the target period is determined based on the first world position information corresponding to each target pedestrian.

[0103] The visual statistical result is used to represent the total number of target pedestrians in the target world region. For example, the total number of target pedestrians in the target world region at the time when the target pedestrian image is acquired can be taken as the visual statistical result according to the total number of the first world position information corresponding to all target pedestrians in each target pedestrian image.

[0104] In step S132b, the wireless statistical result in the target period is determined based on the second world position information corresponding to each wireless terminal.

[0105] The wireless statistical result is used to represent the total number of wireless terminals in the target world region. For example, the total number of wireless terminals in the target world region at the time when the terminal signal set is acquired can be taken as the wireless statistical result according to the total number of the second world position information corresponding to all wireless terminals in each terminal signal set.

[0106] In step S133b, the people statistical result corresponding to the target world region in the target period is determined based on the visual statistical result and the wireless statistical result.

[0107] The present application provides an example for reference. Taking the change amount of target pedestrians as an example of the people statistical result corresponding to the target world region in the target period. The people statistical result corresponding to the target world region in the target period can be determined according to the visual statistical result corresponding to the first target pedestrian image acquired in the target period (in this example, the visual statistical result is equivalent to the total number of target pedestrians in step S130 above), the visual statistical result corresponding to the last target pedestrian image acquired in the target period, the wireless statistical result corresponding to the first terminal signal set acquired, and the wireless statistical result corresponding to the last terminal signal set acquired. The specific process can be referred to the specific content in step S130 above, and the present application does not repeat it. It can be understood that the people statistical result corresponding to the target world region in the target period can also be determined according to the visual statistical result corresponding to the first target pedestrian image acquired in the target period, the visual statistical result corresponding to the last target pedestrian image acquired in the target period, the wireless statistical result corresponding to the first terminal signal set acquired, the wireless statistical result corresponding to the last terminal signal set acquired, the total number of target pedestrians corresponding to the wireless terminal in the first target pedestrian image acquired, and the total number of target pedestrians corresponding to the wireless terminal in the last target pedestrian image acquired.

[0108] According to the above scheme of the embodiment of the present application, the visual statistical result in the target period can be determined based on the first world position information corresponding to each target pedestrian. Then, the wireless statistical result in the target period can be determined based on the second world position information corresponding to each wireless terminal. Finally, the number statistical result of the target world region in the target period can be determined based on the visual statistical result and the wireless statistical result. The above method can generate the visual statistical result and the wireless statistical result in the target period based on the first world position information of the target pedestrian and the second world position information of the wireless terminal respectively, thereby realizing the bimodal observation of the target pedestrian flow in the same physical space. The visual statistical result can accurately reflect the instantaneous distribution of the target pedestrian in the target world region by relying on image detection and space mapping. The wireless statistical result is obtained based on terminal signal detection, and the terminal signal has the advantages of penetrating through the barrier and continuously tracking. The above scheme can effectively overcome the limitations of a single mode by fusing the two kinds of visual statistical results and the wireless statistical result, and the number statistical result generated thereby has higher accuracy.

[0109] Exemplarily, the target world region includes an entrance region and an exit region, and the visual statistical result includes a first inflow and a first outflow.

[0110] The step S133b includes a step S133b1 and a step S133b2.

[0111] In the step S133b1, the first inflow of the entrance region in the target period is determined based on the first world position information corresponding to each target pedestrian, the entrance world position information of the entrance region in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image.

[0112] The head orientation information is used to indicate the orientation of the head of the target pedestrian. The content corresponding to the head orientation information can refer to the above step S1341a. The present application provides an example for reference. Based on the first world position information corresponding to each target pedestrian, the exit world position information corresponding to the exit area in the world coordinate system, it can be determined whether the first world position information corresponding to each target pedestrian in each target pedestrian image is located in the exit area. If it is determined that the target pedestrian M5 and the target pedestrian M6 in the first acquired target pedestrian image in the target period are located in the exit area, the angle between the head orientation of the target pedestrian M5 and the passing direction vector of the exit area is less than a preset angle (for example, 60°), the target pedestrian M5 is determined as the target pedestrian flowing into the target world area. If the angle between the head orientation of the target pedestrian M6 and the passing direction is greater than the above-mentioned preset angle, it indicates that the target pedestrian M6 does not belong to the target pedestrian flowing into the target world area. Through the above process, the target pedestrian flowing into the target world area in each target pedestrian image can be determined. Further, the total number of target pedestrians flowing into the target world area in at least two target pedestrian images of the target world area acquired in sequence in the target period can be taken as the first outflow amount of the exit area in the target period. It can be understood that the above-mentioned preset angle and the exit world position information corresponding to the exit area in the world coordinate system can be determined according to actual conditions.

[0113] In step S133b2, based on the first world position information corresponding to each target pedestrian, the exit world position information corresponding to the exit area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, the first outflow amount of the exit area in the target period is determined.

[0114] The present application provides an example for reference. Based on the first world position information corresponding to each target pedestrian, and the exit world position information corresponding to the exit area in the world coordinate system, it can be determined whether the first world position information corresponding to each target pedestrian in each target pedestrian image is located in the exit area. It is determined that the target pedestrian M7 and the target pedestrian M8 in the first acquired target pedestrian image in the target period are located in the exit area. If the angle between the head orientation of the target pedestrian M7 and the passing direction vector of the exit area is less than a preset angle (for example, 60°), the target pedestrian M7 is determined as a target pedestrian flowing out of the target world area. If the angle between the head orientation of the target pedestrian M8 and the passing direction is greater than the above-mentioned preset angle, it indicates that the target pedestrian M8 does not belong to the target pedestrian flowing out of the target world area. Through the above process, the target pedestrian flowing out of the target world area in each target pedestrian image can be determined. The total number of target pedestrians flowing out of the target world area in at least two target pedestrian images of the target world area acquired in sequence in the target period can be used as the first outflow of the exit area in the target period. It can be understood that the above-mentioned preset angle and the exit world position information corresponding to the exit area in the world coordinate system can be determined according to actual conditions.

[0115] According to the above-mentioned scheme of the embodiments of the present application, the first inflow of the entrance area in the target period can be determined based on the first world position information corresponding to each target pedestrian, the entrance world position information corresponding to the entrance area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image. Then, the first outflow of the exit area in the target period can be determined based on the first world position information corresponding to each target pedestrian, the exit world position information corresponding to the exit area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image. The above-mentioned method can determine the first inflow and the first outflow according to the first world position information corresponding to each target pedestrian and the head orientation information corresponding to each target pedestrian, which is beneficial to improve the accuracy of determining the number of pedestrians flowing into the target world area and the number of pedestrians flowing out of the target world area, and further can ensure the reliability of the statistical results.

[0116] Exemplarily, the above-mentioned processing method further comprises: determining a pedestrian flow directed graph based on the number of pedestrians corresponding to each target world area.

[0117] The directed connection line corresponds to a pedestrian flow amount, and the pedestrian flow amount is used to represent the total number of pedestrians moving from the first world area to the second world area in the target period.

[0118] The pedestrian flow directed graph comprises a plurality of target points connected by directed connection lines. The plurality of target points respectively correspond one-to-one to the plurality of target world areas. For example, referring to Figure 2As shown, Figure 2 A schematic diagram of a pedestrian flow directed graph according to an embodiment of the present application is shown. Figure 2 Each of the target points P1 to P6 corresponds to a target world region, and different target points correspond to different target world regions. For example, if the target point P1 corresponds to the entrance area of a train station, the target point P2 corresponds to the first floor waiting area of the train station, the target point P3 corresponds to the second floor waiting area of the train station, the target point P4 corresponds to the VIP waiting area of the train station, the target point P5 corresponds to the ticket checking area of the train station, and the target point P6 corresponds to the platform area of the train station. It can be understood that the place where the target world region is located can be reasonably divided into multiple target world regions according to the physical structure (for example, a passageway, a gate, a shop, etc.) of the place.

[0119] For a directed connection line from a target point corresponding to a first world region in the multiple target world regions to a target point corresponding to a second world region in the multiple target world regions, the directed connection line is used to represent that a target pedestrian unidirectionally passes from the first world region to the second world region. For example, the above-mentioned first world region can be a target world region from which a target pedestrian flows out, and the target world region into which the target pedestrian flowing out of the first world region flows is the second world region corresponding to the first world region. For reference Figure 2 As shown, the connection lines L1 to L6 can be used to connect from a target point corresponding to a first world region in the multiple target world regions to a target point corresponding to a second world region in the multiple target world regions. For example, Figure 2 Taking the connection line L1 in the above-mentioned connection lines as an example, the connection line L1 represents that the target point P1 (i.e., the first world region in the example) flows out a target pedestrian to the target point P2 (i.e., the second world region in the example). It can be understood that the above-mentioned connection line is only used to connect target points that are physically adjacent and in which passenger flow can directly flow (for example, a target pedestrian can directly reach).

[0120] The present application provides an example for reference. For reference Figure 2 As shown, the number of target pedestrians corresponding to each connection line (for example, the number of inflow or the number of outflow) can be determined according to the number of people counted by the steps S110 to S130. The number of target pedestrians corresponding to each connection line can be marked on the connection line in the pedestrian flow directed graph. For example, according to the calculation of the steps S110 to S130, 6 target pedestrians flow out of the target world region P1 to the target world region P2 in the target time period, and the number of target pedestrians corresponding to the connection line L1 can be marked on the connection line L1 in the pedestrian flow directed graph. Figure 2The connection line L1 is marked 6. In this way, the pedestrian outflow and the pedestrian inflow of each of the target world area P2 to the target world area P6 can be calculated, and the determined pedestrian outflow and the pedestrian inflow can be marked on the corresponding connection line, so as to determine a more detailed pedestrian flow directed graph.

[0121] According to the above scheme of the embodiment of the present application, the pedestrian flow directed graph can be determined based on the number of people in each target world area. The above method can intuitively depict the migration path and flow trend of the target pedestrian in the multiple target world areas. Compared with only providing the number of people of the target pedestrian in each target world area independently, the above method can reveal the correlation and passenger flow guiding relationship between different target world areas, such as identifying the main passage path, the path with congestion, etc. Especially in the multi-target world area correlation analysis scene, the pedestrian flow directed graph can intuitively reflect the behaviors such as reverse flow, abnormal aggregation, path deviation, etc., and then the obtained pedestrian flow directed graph can be used to optimize the entrance and exit, place guide signs in the target world area prone to path deviation, plan emergency evacuation paths for the path with congestion, etc.

[0122] The embodiment of the present application also provides a pedestrian data processing device. Figure 3 A schematic block diagram of a pedestrian data processing device 200 according to an embodiment of the present application is shown. In combination with the above description of the method, the pedestrian data processing device 200 can include a first world position information determination module 210, a second world position information determination module 220, and a number of people statistical result determination module 230. Figure 3 As shown, the processing device 200 can include a first world position information determination module 210, a second world position information determination module 220, and a number of people statistical result determination module 230.

[0123] The first world position information determination module 210 is configured to determine the first world position information of each target pedestrian in a world coordinate system based on at least two target pedestrian images of the target world area acquired in sequence within a target period.

[0124] The second world position information determination module 220 is configured to determine the second world position information of each wireless terminal in a world coordinate system based on at least two terminal signal sets of the target world area acquired in sequence within the target period, wherein the at least two target pedestrian images and the at least two terminal signal sets correspond to each other in one-to-one manner, and for each target pedestrian image, the acquisition time of the target pedestrian image is the same as the acquisition time of the terminal signal set corresponding to the target pedestrian image.

[0125] The number of people statistical result determination module 230 is configured to determine the number of people statistical result of the target world area within the target period based on the first world position information of each target pedestrian and the second world position information of each wireless terminal.

[0126] Exemplarily, for each target pedestrian image in the at least two target pedestrian images, the first world position information determining module 210 comprises: a first image position information determining module and a first world position information determining submodule.

[0127] The first image position information determining module is configured to input the target pedestrian image into the trained target detection model to obtain first image position information corresponding to each target pedestrian in the target pedestrian image, wherein the first image position information is used to represent the position information of the target pedestrian in the target pedestrian image.

[0128] The first world position information determining submodule is configured to determine, for each target pedestrian in the target pedestrian image, first world position information corresponding to the target pedestrian based on the first image position information corresponding to the target pedestrian and a first preset relationship, wherein the first preset relationship is used to represent the corresponding relationship between the image coordinate system of the target pedestrian image and the world coordinate system.

[0129] Exemplarily, the trained target detection model is trained based on a training pedestrian image and annotation information of at least one target region corresponding to the training pedestrian image, wherein the at least one target region comprises a head-shoulder region and / or a full-body region of a training pedestrian in the training pedestrian image, the annotation information of the head-shoulder region is used to represent the position information of the region from the head to the shoulder of the training pedestrian in the training pedestrian image, and the annotation information of the full-body region is used to represent the position information of the region from the head to the foot of the training pedestrian in the training pedestrian image.

[0130] Exemplarily, the number of people counting result determining module 230 comprises: a first motion trajectory determining module, a second motion trajectory determining module, a first determining module and a second determining module.

[0131] The first motion trajectory determining module is configured to determine, for each target pedestrian, a first motion trajectory of the target pedestrian in the world coordinate system in the target time period based on the first world position information corresponding to the target pedestrian.

[0132] The second motion trajectory determining module is configured to determine, for each wireless terminal, a second motion trajectory of the wireless terminal in the world coordinate system in the target time period based on the second world position information corresponding to the wireless terminal.

[0133] The first determining module is configured to determine, based on the first motion trajectory of each target pedestrian and the second motion trajectory of each wireless terminal, a wireless terminal corresponding to each target pedestrian, wherein for a target pedestrian corresponding to a wireless terminal, the degree of coincidence between the first motion trajectory of the target pedestrian and the second motion trajectory of the wireless terminal is higher than a preset threshold.

[0134] The second determining module is configured to determine a people counting result of the target world region in the target period based on the first world position information corresponding to the target pedestrian of the wireless terminal or the second world position information corresponding to the wireless terminal of the target pedestrian.

[0135] The target world region includes an entrance region and an exit region, and the people counting result includes a pedestrian inflow amount and a pedestrian outflow amount. The second determining module includes a pedestrian inflow amount determining module and a pedestrian outflow amount determining module.

[0136] The pedestrian inflow amount determining module is configured to determine the pedestrian inflow amount of the entrance region in the target period based on the first world position information corresponding to the target pedestrian of the wireless terminal, entrance world position information corresponding to the entrance region in a world coordinate system, and head orientation information corresponding to each target pedestrian in the target pedestrian image, wherein the head orientation information is used to represent the orientation of the head of the target pedestrian.

[0137] The pedestrian outflow amount determining module is configured to determine the pedestrian outflow amount of the exit region in the target period based on the first world position information corresponding to the target pedestrian of the wireless terminal, exit world position information corresponding to the exit region in the world coordinate system, and head orientation information corresponding to each target pedestrian in the target pedestrian image.

[0138] The people counting result determining module 230 includes a visual counting result determining module, a wireless counting result determining module, and a third determining module.

[0139] The visual counting result determining module is configured to determine a visual counting result in the target period based on the first world position information corresponding to each target pedestrian, wherein the visual counting result is used to represent the total number of target pedestrians in the target world region.

[0140] The wireless counting result determining module is configured to determine a wireless counting result in the target period based on the second world position information corresponding to each wireless terminal, wherein the wireless counting result is used to represent the total number of wireless terminals in the target world region.

[0141] The third determining module is configured to determine the people counting result of the target world region in the target period based on the visual counting result and the wireless counting result.

[0142] The target world region includes an entrance region and an exit region, the visual counting result includes a first inflow amount and a first outflow amount, and the visual counting result determining module includes a first inflow amount determining module and a first outflow amount determining module.

[0143] The first inflow amount determining module is configured to determine a first inflow amount of the entrance area in the target period based on the first world position information corresponding to each target pedestrian, the entrance world position information corresponding to the entrance area in the world coordinate system, and head orientation information corresponding to each target pedestrian in the target pedestrian image, wherein the head orientation information is used to indicate the orientation of the head of the target pedestrian.

[0144] The first outflow amount determining module is configured to determine a first outflow amount of the exit area in the target period based on the first world position information corresponding to each target pedestrian, the exit world position information corresponding to the exit area in the world coordinate system, and head orientation information corresponding to each target pedestrian in the target pedestrian image, wherein the head orientation information is used to indicate the orientation of the head of the target pedestrian.

[0145] Exemplarily, the processing apparatus 200 further comprises a pedestrian flow directed graph determining module.

[0146] The pedestrian flow directed graph determining module is configured to determine a pedestrian flow directed graph based on the number of people corresponding to each target world area, wherein the pedestrian flow directed graph comprises a plurality of target points connected via directed connection lines, the plurality of target points correspond to the plurality of target world areas one by one, for a directed connection line from a target point corresponding to a first world area in the plurality of target world areas to a target point corresponding to a second world area in the plurality of target world areas, the directed connection line is used to indicate that the target pedestrian unidirectionally passes from the first world area to the second world area, and the directed connection line corresponds to a pedestrian flow amount, wherein the pedestrian flow amount is used to indicate the total amount of pedestrians moving from the first world area to the second world area in the target period.

[0147] According to another aspect of the present application, an electronic device is also provided. Figure 4 A schematic block diagram of an electronic device 300 according to an embodiment of the present application is shown. As shown, the electronic device 300 comprises a processor 310 and a memory 320, and the memory 320 stores a computer program, and the computer program is used to execute the above-mentioned pedestrian data processing method when the computer program is executed by the processor 310. Figure 4

[0148] ​Further, according to another aspect of the present application, there is provided a storage medium having stored thereon program instructions, which when executed by a computer or processor, cause the computer or processor to perform the corresponding steps of the above pedestrian data processing method and to implement the corresponding modules in the above pedestrian data processing apparatus or the corresponding modules in the above electronic device according to embodiments of the present application. The storage medium may, for example, include a memory card of a smart phone, a memory component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0149] According to still another aspect of the present application, there is provided a computer program product comprising computer program instructions, which when executed by a computer or processor, cause the computer or processor to perform the corresponding steps of the above pedestrian data processing method.

[0150] Those skilled in the art can understand the specific implementation of the above electronic device and storage medium by reading the above description of the pedestrian data processing method. For brevity, no further elaboration is given here.

[0151] Although the example embodiments have been described herein with reference to the accompanying drawings, it is to be understood that the example embodiments are merely exemplary and are not intended to limit the scope of the present application. Those skilled in the art can make various changes and modifications without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as claimed in the appended claims.

[0152] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0153] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the above-described device embodiments are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be omitted or not executed.

[0154] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0155] Similarly, it is to be understood that the phraseology or terminology employed herein, and not otherwise specified, is for the purpose of description only and not of limitation. Rather, the use of such terms is merely for providing one skilled in the art with a technical conception of the inventive subject matter. The foregoing detailed description has been presented for purposes of clarity of understanding only. It is most

[0156] Those skilled in the art recognize that all features disclosed in this specification, including the claims, abstract, and drawings, can be combined in any combination, except combinations that would be technically infeasible. Each individual feature disclosed in this specification, including the claims, abstract, and drawings, can be replaced by alternative features that are both equivalent in individual function and result.

[0157] Furthermore, those skilled in the art will recognize that references in the specification to "one embodiment", "an embodiment", "an example embodiment", etc., mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.

[0158] The various component embodiments of the present application can be implemented in hardware, or as software modules running in one or more processors, or combinations thereof. Those skilled in the art will appreciate that some or all of the functionality of some of the modules in the pedestrian data processing apparatus according to embodiments of the present application can be implemented in practice using a microprocessor or a digital signal processor (DSP). The present application can also be implemented as a program for executing, in part or in whole, the methods described herein, for example, a computer program and a computer program product. Such program implementing the present application can be stored on a computer readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0159] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that one skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps other than those listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In a unitary claim, several of the devices, apparatuses or means, if any, can be implemented by one and the same item of hardware. The use of the words 'first','second' and 'third', etc. do not imply any order but rather are used for naming purposes only. Further, the word'step' can not imply any order and can not necessarily be construed to refer to a sequence of steps.

[0160] Any discussion of documents, acts, materials, devices, articles or the like that has been included in the present summary is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present application, as it exists worldwide at the priority date of each claim of this application.

Claims

1. A method for processing pedestrian data, characterized in that, The method includes: Based on at least two target pedestrian images of the target world region acquired sequentially within the target time period, the first world position information of each target pedestrian in each target pedestrian image in the world coordinate system is determined, wherein the target world region includes an entrance region and an exit region; Based on at least two terminal signal sets of the target world region acquired sequentially within the target time period, the second world location information corresponding to each wireless terminal in the world coordinate system is determined. For each target pedestrian image, the acquisition time of the target pedestrian image is the same as the acquisition time of the terminal signal set corresponding to the target pedestrian image. Based on the first-world location information corresponding to each target pedestrian and the second-world location information corresponding to each wireless terminal, the population statistics of the target world area within the target time period are determined, wherein the population statistics include pedestrian inflow and pedestrian outflow. The determination of the population statistics for the target world region within the target time period, based on the first-world location information corresponding to each target pedestrian and the second-world location information corresponding to each wireless terminal, includes: For each target pedestrian, based on the first world location information corresponding to the target pedestrian, the first motion trajectory of the target pedestrian in the world coordinate system within the target time period is determined; For each wireless terminal, based on the second world location information corresponding to the wireless terminal, the second motion trajectory of the wireless terminal in the world coordinate system within the target time period is determined; Based on the first motion trajectory of each target pedestrian and the second motion trajectory of each wireless terminal, the wireless terminal corresponding to each target pedestrian is determined. For a target pedestrian with a corresponding wireless terminal, the degree of overlap between the first motion trajectory of the target pedestrian and the second motion trajectory of the wireless terminal is higher than a preset threshold. Based on the first-world location information of the target pedestrian with a corresponding wireless terminal, determine the population statistics of the target world region within the target time period; The determination of the population statistics for the target world region within the target time period based on the first-world location information corresponding to the target pedestrian with a wireless terminal includes: Based on the first world location information of the target pedestrian with a wireless terminal, the entrance world location information of the entrance area in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image, the pedestrian inflow volume corresponding to the entrance area during the target time period is determined, wherein the head orientation information is used to indicate the orientation of the target pedestrian's head. Based on the first world location information of the target pedestrian with a corresponding wireless terminal, the exit world location information of the exit area in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image, the pedestrian outflow volume corresponding to the exit area within the target time period is determined.

2. The method as described in claim 1, characterized in that, The determination of the first-world position information of each target pedestrian in each target pedestrian image in the world coordinate system, based on at least two target pedestrian images of the target world region acquired sequentially within the target time period, includes: For each of the at least two target pedestrian images, The target pedestrian image is input into the trained target detection model to obtain the first image position information corresponding to each target pedestrian in the target pedestrian image, wherein the first image position information is used to represent the position information of the target pedestrian in the target pedestrian image; For each target pedestrian in the target pedestrian image, based on the first image location information corresponding to the target pedestrian and the first preset relationship, the first world location information corresponding to the target pedestrian is determined, wherein the first preset relationship is used to represent the correspondence between the image coordinate system of the target pedestrian image and the world coordinate system.

3. The method as described in claim 2, characterized in that, The trained target detection model is trained based on a training pedestrian image and the annotation information of at least one target region corresponding to the training pedestrian image. The at least one target region includes: the head and shoulder region and / or the whole body region of the training pedestrian in the training pedestrian image. The annotation information of the head and shoulder region is used to indicate the position information of the region from the head to the shoulders of the training pedestrian in the training pedestrian image. The annotation information of the whole body region is used to indicate the position information of the region from the head to the feet of the training pedestrian in the training pedestrian image.

4. The method as described in claim 1, characterized in that, The determination of the population statistics for the target world region within the target time period, based on the first-world location information corresponding to each target pedestrian and the second-world location information corresponding to each wireless terminal, includes: Based on the first-world location information corresponding to each target pedestrian, the visual statistical results within the target time period are determined, wherein the visual statistical results are used to represent the total number of target pedestrians in the target world region; Based on the second-world location information corresponding to each wireless terminal, the wireless statistics results within the target time period are determined, wherein the wireless statistics results are used to represent the total number of wireless terminals in the target world region; Based on the visual statistics and the wireless statistics, the population statistics corresponding to the target world region within the target time period are determined.

5. The method as described in claim 4, characterized in that, The target world region includes an entrance region and an exit region, and the visual statistics results include a first inflow volume and a first outflow volume. Determining the visual statistics results within the target time period based on the first world location information corresponding to each target pedestrian includes: Based on the first world location information corresponding to each target pedestrian, the entrance world location information corresponding to the entrance area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, the first inflow of the entrance area within the target time period is determined, wherein the head orientation information is used to indicate the orientation of the target pedestrian's head; Based on the first world location information corresponding to each target pedestrian, the exit world location information corresponding to the exit area in the world coordinate system, and the head orientation information corresponding to each target pedestrian in the target pedestrian image, the first outflow of the exit area within the target time period is determined.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Based on the population statistics corresponding to each target world region, a directed pedestrian flow graph is determined. The directed pedestrian flow graph includes multiple target points connected by directed connecting lines. Each target point corresponds to one of multiple target world regions. For a directed connecting line from a target point corresponding to a first world region among the multiple target world regions to a target point corresponding to a second world region among the multiple target world regions, the directed connecting line is used to represent the one-way passage of the target pedestrians from the first world region to the second world region. The directed connecting line corresponds to the pedestrian flow volume, which is used to represent the total number of pedestrians moving from the first world region to the second world region during the target time period.

7. A pedestrian data processing device, characterized in that, The device includes: The first-world location information determination module is used to determine the first-world location information of each target pedestrian in each target pedestrian image in the world coordinate system based on at least two target pedestrian images of the target world region acquired sequentially within the target time period. The target world region includes an entrance region and an exit region. The second world location information determination module is used to determine the second world location information of each wireless terminal in the world coordinate system based on at least two terminal signal sets acquired sequentially within the target time period and in the target world region. The at least two target pedestrian images correspond one-to-one with the at least two terminal signal sets. For each target pedestrian image, the acquisition time of the target pedestrian image is the same as the acquisition time of the terminal signal set corresponding to the target pedestrian image. The people counting result determination module is used to determine the people counting result corresponding to the target world area within the target time period based on the first world location information corresponding to each target pedestrian and the second world location information corresponding to each wireless terminal. The people counting result includes pedestrian inflow and pedestrian outflow. The people counting result determination module includes: a first motion trajectory determination module, a second motion trajectory determination module, a first determination module, and a second determination module; The first motion trajectory determination module is used to determine, for each target pedestrian, the first motion trajectory of the target pedestrian in the world coordinate system within the target time period, based on the first world location information corresponding to the target pedestrian; The second motion trajectory determination module is used to determine, for each wireless terminal, the second motion trajectory of the wireless terminal in the world coordinate system within the target time period based on the second world location information corresponding to the wireless terminal. The first determining module is used to determine the wireless terminal corresponding to each target pedestrian based on the first movement trajectory of each target pedestrian and the second movement trajectory of each wireless terminal. For a target pedestrian with a corresponding wireless terminal, the degree of overlap between the first movement trajectory of the target pedestrian and the second movement trajectory of the wireless terminal is higher than a preset threshold. The second determining module is used to determine the population statistics of the target world region within the target time period based on the first world location information corresponding to the target pedestrian with a wireless terminal. The second determining module includes: a pedestrian inflow determining module and a pedestrian outflow determining module; The pedestrian inflow determination module is used to determine the pedestrian inflow corresponding to the entrance area within the target time period based on the first world location information of the target pedestrian with a wireless terminal, the entrance world location information of the entrance area in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image. The head orientation information is used to indicate the orientation of the target pedestrian's head. The pedestrian outflow determination module is used to determine the pedestrian outflow corresponding to the exit area within the target time period based on the first world location information of the target pedestrian with a corresponding wireless terminal, the exit world location information of the exit area in the world coordinate system, and the head orientation information of each target pedestrian in the target pedestrian image.

8. A storage medium storing computer program instructions, characterized in that, The computer program instructions, when executed, are used to perform the pedestrian data processing method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method and system for estimation of population density and mobility based on multi-data fusion

    CN107801203A