Device and method for detecting, tracking and identifying people using wireless signals and images

By combining wireless transceivers and cameras, utilizing CSI information and image data, and integrating people's motion trajectory estimation, the problem of accurate identification of people in large-scale crowd environments is solved, fine-grained tracking and identification is achieved, and privacy protection requirements are met.

CN112381853BActive Publication Date: 2025-09-23ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010737062.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-29
Filing Date
2020-07-28
Publication Date
2025-09-23
Estimated Expiration
2040-07-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately detect, track, and identify people using facial recognition algorithms in places such as retail stores and airports, especially in large crowd environments. Facial recognition algorithms are subject to privacy and regulatory restrictions, and the location information provided by RSSI technology of wireless signals is coarse-grained and not accurate enough.

Method used

Combining wireless transceivers and cameras, using CSI information and image data, through the fusion of camera motion and wireless packet motion, multi-mode sensor fusion and neural network algorithms, the movement trajectory of people is estimated and identified.

Benefits of technology

It achieves fine-grained identification and tracking of people in large-scale crowd environments, improves location accuracy, reduces dependence on multiple infrastructures, and adapts to privacy protection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112381853B_ABST
    Figure CN112381853B_ABST
Patent Text Reader

Abstract

An apparatus comprises: a wireless transceiver configured to transmit packet data to a mobile device associated with one or more persons in the vicinity of the wireless transceiver; and a controller in communication with the wireless transceiver and a camera, the controller being configured to: receive a plurality of packet data from the one or more persons, wherein the packet data includes at least amplitude information associated with a wireless channel in communication with the wireless transceiver; and receive an image containing a motion trajectory of an individual from the camera; and perform detection, tracking, and pseudo-identification of the individual by fusing the motion trajectory from the wireless signal and the camera image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a wireless, camera-based monitoring system. Background Art

[0002] Retail stores, airports, convention centers, and smart districts / communities can monitor nearby people. Detecting, tracking, and pseudo-identifying people can have various use cases across different applications. In many applications, cameras can be used to track people. For example, cameras in retail stores may be mounted in the ceiling, looking downward, and lack the ability to accurately identify people using facial recognition algorithms. Furthermore, facial recognition algorithms may perform poorly in locations where thousands of people may be present, such as airports or large retail stores. Summary of the Invention

[0003] According to one embodiment, an apparatus includes a wireless transceiver configured to transmit packet data to a mobile device associated with one or more persons in a vicinity of the wireless transceiver. The apparatus further includes a camera configured to capture image data of the one or more persons in the vicinity. The apparatus further includes a controller in communication with the wireless transceiver and the camera, the controller configured to: receive a plurality of packet data from the mobile device, wherein the packet data includes at least amplitude information associated with a wireless channel in communication with the wireless transceiver; determine camera motion representing motion of the one or more persons using the image data, and determine group motion representing motion of the one or more persons using the packet data; identify each of the one or more persons in response to the camera motion and the group motion; and output information associated with each of the one or more persons in response to the identification of the one or more persons.

[0004] According to another embodiment, a system includes a wireless transceiver configured to transmit packet data to a mobile device associated with one or more persons in a vicinity of the wireless transceiver. The system also includes a camera configured to identify the one or more persons and identify estimated camera motion using at least image data. The system also includes a controller in communication with the wireless transceiver and the camera. The controller is configured to: receive a plurality of packet data from the mobile device, wherein the packet data includes at least amplitude information associated with a wireless channel in communication with the wireless transceiver; estimate a packet motion of each of the one or more persons in response to the plurality of packet data; and identify each of the one or more persons in response to the estimated camera motion and the estimated packet motion.

[0005] According to yet another embodiment, a method for identifying a person using a camera and a wireless transceiver includes: receiving packet data from a mobile device associated with one or more persons in proximity to the wireless transceiver; obtaining image data associated with a camera associated with the wireless transceiver; determining an estimated camera motion of the person using the image data; determining an estimated motion of the person using the packet data from the mobile device; comparing the estimated camera motion to the estimated motion; and identifying the one or more persons in response to the comparison. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Figure 1 is an overview system diagram of a wireless system according to an embodiment of the present disclosure.

[0007] Figure 2 is an exemplary image of image data collected in a camera according to an embodiment of the present disclosure.

[0008] Figure 3 is an exemplary flow chart of an algorithm according to an embodiment of the present disclosure.

[0009] Figure 4 is an exemplary flow chart of an embodiment for comparing a camera's motion signature with a motion signature from a wireless transceiver.

[0010] Figure 5 is an exemplary flow chart of a second embodiment of comparing a camera's motion signature with a motion signature from a wireless transceiver.

[0011] Figure 6 is an exemplary visualization of the motion sign in the case of using a camera and the motion sign in the case of using Wi-Fi. DETAILED DESCRIPTION

[0012] Embodiments of the present disclosure are described herein. However, it should be understood that the disclosed embodiments are merely examples, and other embodiments may take various and alternative forms. The figures are not necessarily drawn to scale; some features may be enlarged or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but should merely be interpreted as a representative basis for teaching those skilled in the art to apply the embodiments in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any one of the figures may be combined with features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. The combination of illustrated features provides representative embodiments for typical applications. However, for specific applications or implementations, various combinations and modifications of features consistent with the teachings of the present disclosure may be desired.

[0013] Detecting, tracking, and identifying people can be important for a wide range of applications, including retail stores, airports, convention centers, and smart cities. In many applications, cameras are used to track people. However, in many environments, such as retail stores, cameras are mounted in the ceiling, looking downward, and lack the ability to accurately identify people using facial recognition algorithms. Furthermore, facial recognition algorithms may not scale across thousands of people, such as in airports or large retail stores. Furthermore, facial recognition may be regulated in some locations due to privacy concerns. However, for many analytical applications, fine-grained person identification is not necessary, as the focus is on individual data rather than specific individuals. However, obtaining such data requires distinguishing between different people and re-identifying the same person when they are nearby. In one embodiment, camera data, along with data obtained from a wireless device carried by the person (such as unique device details—MAC address, list of access points being sought, etc.), is used to capture the person's movements or itinerary within a given environment (retail store, office, shopping mall, hospital, etc.).

[0014] To pseudo-identify and track people, wireless technologies can be used to track wireless devices (such as phones and wearables) carried by users. For example, Bluetooth and Wi-Fi packets can be sniffed to identify and locate nearby people. However, most current solutions use the RSSI function of wireless signals and obtain a coarse-grained location, such as if a given wireless device (and person) is within a certain radius (e.g., 50 meters). Similarly, to locate wireless devices, some technologies require the deployment of multiple infrastructure anchor points to simultaneously receive packets from the device and then use RSSI values ​​to perform triangulation. The accuracy of these solutions is affected by fluctuations and the lack of information provided by RSSI. Compared to RSSI, CSI (channel state information) provides much richer information about how the signal propagates from transmitter to receiver and captures the combined effects of signal scattering, fading, and power loss over distance. Our proposed solution uses CSI and a single system unit (with multiple antennas), thus reducing the effort required to deploy multiple units. However, for better performance, multiple system units can be deployed.

[0015] Figure 11 is an overview system diagram of a wireless system 100 according to an embodiment of the present disclosure. The wireless system 100 may include a wireless unit 101 for generating and transmitting CSI data. The wireless unit 101 may communicate with a mobile device (e.g., a cellular phone, a wearable device, a tablet) of an employee 115 or a customer 107. For example, the mobile device of the employee 115 may transmit a wireless signal 119 to the wireless unit 101. Upon receiving the wireless packet, the system unit 101 obtains a CSI value associated with the packet reception. Furthermore, the wireless packet may contain identifiable information regarding a device ID, such as a MAC address for identifying the employee 115. Therefore, the system 100 and the wireless unit 101 may determine various hotspots without utilizing data exchanged from the employee 115's device.

[0016] Although Wi-Fi may be used as the wireless communication technology, any other type of wireless technology may be utilized. For example, if the system can obtain CSI from a wireless chipset, Bluetooth may be utilized. The system unit may be able to include a Wi-Fi chipset attached to up to three antennas, as shown by wireless unit 101 and wireless unit 103. In one embodiment, the system unit may include a receiving station that includes a Wi-Fi chipset that includes up to three antennas. The system unit can be mounted at any height or mounted on a ceiling. In another embodiment, a chipset that utilizes CSI information may be used. Wireless unit 101 may include a camera to monitor various people walking around the POI. In another example, wireless unit 103 may not include a camera and may simply communicate with a mobile device.

[0017] System 100 can cover various aisles, such as 109, 111, 113, and 114. Aisles can be defined as walking paths between shelves 105 or store walls. Data collected between aisles 109, 111, 113, and 114 can be used to generate heat maps and focus on store traffic. The system can analyze data from all aisles and use this data to identify traffic in other areas of the store. For example, data collected from various customers' 107 mobile devices can identify areas of the store that receive high traffic. This data can be used to position certain products. By leveraging this data, store managers can determine where high-traffic real estate is located compared to low-traffic real estate. Additionally, by using WiFi to fuse pseudo-identification information with camera-based analytics (e.g., gender, age range, ethnicity), the system can build profiles of individual customers and provide customer-specific analytics for individual aisles. Moreover, by capturing a single customer's entire journey, the system can provide store-wide customer-specific analytics.

[0018] CSI data can be transmitted in packets found in wireless signals. In one example, wireless signal 121 can be generated by customer 107 and its associated mobile device. System 100 can utilize various information found in wireless signal 121 to determine whether customer 107 is an employee or other characteristics, such as the angle of arrival (AoA) of the signal. Customer 107 can also communicate with wireless unit 103 via signal 122. In addition, packet data found in wireless signal 121 can be communicated with both wireless unit 101 or unit 103. The packet data in wireless signals 121, 119, and 117 can be used to provide information related to movement trajectory and traffic data related to the employee / customer's mobile device.

[0019] Figure 2 is an exemplary image of image data collected in a camera according to an embodiment of the present disclosure. As shown in the image data, Figure 2 The camera in FIG2 may be mounted in the wireless unit 101 in the ceiling. In other embodiments, the wireless unit 101 may be mounted anywhere else, such as on a shelf or wall. A motion trajectory 201 of a person 203 is shown, and the motion trajectory of the person 203 can be determined according to various embodiments disclosed below. The image data captured by the camera can be used to gather information about people moving around a space (e.g., customer or employee, gender, age range, ethnicity). As further described below, this image data can also be overlaid with a heat map or other information. The camera can add a bounding box 205 around the person. The camera can use object detection techniques (such as YOLO, SSD, Faster RCNN, etc.) to detect humans. The bounding box can identify a boundary around a person or object, which can be displayed on the graphical image to identify the person or object. The bounding box can be tracked using optical flow, mean-shift tracking, a Kalman filter, a particle filter, or other types of mechanisms. The tracking can be estimated by analyzing the person's position over any given period of time. Furthermore, identification numbers 207 can be assigned to people identified using various techniques explained further below.

[0020] Figure 3is an exemplary flow chart of an algorithm according to an embodiment of the present disclosure. System 100 illustrates an exemplary embodiment of an algorithm for identifying camera images using motion tracking by a wireless transceiver (e.g., a Wi-Fi transceiver, a Bluetooth transceiver, an RFID transceiver, etc.). In one embodiment, the system may utilize a camera 301 and a wireless transceiver 302. At a high level, camera 301 may utilize various data (e.g., image data), and wireless transceiver 302 may utilize its own data (e.g., wireless packet data) to estimate the motion of various people or objects identified by camera 301 and wireless transceiver 302, and then utilize these two data sources to determine whether the person can be matched.

[0021] At step 303, the camera 301 may detect a person in one or more frames. i The person may be detected using object identification tracking for each frame, which utilizes algorithms to identify various objects. The camera 301 may use object detection techniques (such as YOLO, SSD, Faster RCNN, etc.) to detect humans. Upon detecting the person, the camera 301 may then estimate the person's i A bounding box around a person or object. A bounding box may identify a boundary around a person or object that may be displayed on a graphic image to identify the person or object. At step 307, the camera may identify the person by i The position of the body is used to estimate the person i For example, the camera 301 may estimate the position by viewing a foot or another object associated with the person. In another example, the bounding box may be tracked using optical flow, mean shift tracking, a Kalman filter, a particle filter, or other types of mechanisms. At step 309, the system 300 may estimate the position of the person. i Tracking. It is possible to analyze people at any given time i For example, the system may track a person at different intervals (such as 1 second, 5 seconds, 10 seconds, etc.). In one example, it may be optional for the camera unit 301 to estimate the position of the person and track the person. At step 311, the camera 301 may track the person. i The motion markers are estimated as MS i cam . Figure 6 An exemplary motion marker MS is shown in i cam The motion marker MS may be a track within a specific time period (eg, TW1). For example, the time period may be 5 seconds or a longer or shorter amount of time. The motion marker MS associated with the camera i camcan be associated with a person and can be used in conjunction with an estimate established by wireless transceiver 302 to identify the person i .

[0022] In one embodiment, the wireless transceiver 302 can work simultaneously with the camera 301 to identify people. j The wireless transceiver receives the signal received by the personnel at step 313. j The packets generated by the smartphone or other mobile device of the user. Thus, the mobile device may generate wireless traffic (e.g., Wi-Fi traffic) that is received at the central system unit. At step 315, the wireless transceiver (e.g., EyeFI unit) may receive the wireless packets (e.g., Wi-Fi packets). The CSI value may be extracted from the received packets. When a wireless data packet is received, the corresponding MAC address may be utilized to identify the person. The MAC address may then be assigned to the associated person. Thus, at step 317, the wireless transceiver 302 may extract the pseudo identification information as the ID j .

[0023] At step 319, the system can measure the angle of arrival (AoA) of the packet by utilizing CSI values ​​from multiple antennas using an algorithm (e.g., SpotFi algorithm, or neural network based AoA estimation algorithm). Using AoA and / or raw CSI values, the system can estimate the angle of arrival (AoA) of the person. j The motion trajectory of the wireless marker can be within a predetermined time window (e.g., TW2) when using a wireless signal (e.g., Wi-Fi). If the wireless transceiver and the camera are time-synchronized, TW2 may be less than TW1 because the mobile device may not generate traffic during the entire period when the person is visible through the camera. At step 321, the wireless transceiver 302 may then estimate the person's j Sports logo MS j Wi-Fi . Figure 6 An exemplary motion marker MS is shown in j Wi-Fi .

[0024] Then, the system 300 can i cam With MS j Wi-Fi Compare them to determine if they are similar. If they are similar, EyeFi uses the ID j To identify personnel iAt step 323, the system may determine whether the estimated motion signatures from the camera 301 and the wireless transceiver 302 are similar. The camera data and the wireless data may be fused to determine the identity of the person. An algorithm may be used to perform multi-modal sensor fusion. The comparison of the two motion signatures may be used to determine the identity of the person. i and personnel j If the deviation between the two comparisons determines that they are similar, the system 300 can identify the person as ID j or hashed ID j If the comparison between the motion signatures shows a further deviation from the threshold, the system may try the next pair of motion signatures (MS i cam With MS j Wi-Fi ). The system can then determine that there is no match between the two estimated motion signatures and restart the process of identifying another subset of the camera data and wireless data to identify the person. Otherwise, it will use the next pair of people detected by the camera and wireless transceiver to check whether they are similar.

[0025] Figure 4 is an exemplary flow diagram of a system 400 that compares a camera's motion signature with a motion signature from a wireless transceiver. Figure 4 The camera's motion signature is compared with the motion signature from the wireless transceiver ( Figure 3 323) of the second embodiment of the exemplary flow chart. In this embodiment, Figure 4In steps 401 and 413, data from each modality is mapped to "semantic concepts" that are sensor-invariant. For example, in steps 403 and 415, the following "concepts" are extracted for each person using a neural network-based or other method, independent of other sensing modalities: (a) location, (b) change in location, (c) angle of arrival (AoA), (d) change in AoA, (e) standing versus moving, (f) direction of movement, (g) person orientation, (i) footsteps / gait, (j) phone in hand or pocket, (i) scale of nearby obstacles, and (k) motion trajectory from vision and WiFi. A cascade of neural networks may be required to refine these concepts. In another embodiment, instead of estimating semantic concepts from each modality (e.g., wireless transceiver, camera), an IMU (inertial measurement unit) can be used as a middle-ground to estimate some concepts, as in steps 405 and 417. This is because estimating gait features from either a wireless transceiver or a camera can be difficult. However, if a person has the app installed and carries a mobile phone (which uses the phone's accelerometer or gyroscope to generate IMU data), it can be relatively easy to translate CSI data into IMU data, and image data into IMU data. The translated IMU data from each modality can provide a middle ground to estimate the similarity between the same concept from both modalities. IMU data can also be used in the training phase. After training is complete and the translation function has been learned for each concept from each modality, IMU data and app installation may no longer be required for product use.

[0026] At step 405, concepts may be extracted using inertial mobile unit (IMU) data from the smartphone. j Wi-Fi To extract similar semantic concepts. At step 407, the semantic concepts can be refined and updated by fusing the camera and Wi-Fi data. A single neural network for each modality may be required, or a cascade of neural networks may be required to jointly refine the semantic concepts. Finally, SC i cam and SC j Wi-Fi The system then estimates the normalized weighted distance D to determine SC i cam With SC j Wi-FiAre they similar? For this purpose, it can also use cosine similarity, Euclidean distance or cluster analysis or other similarity metrics, or a combination thereof. If D is less than the threshold T1, the system returns "yes", meaning that SC i cam With SC j Wi-Fi Otherwise, it returns "No".

[0027] In one embodiment, the data from each modality is mapped to "semantic concepts" that are sensor invariant. For example, the following "concepts" are extracted for each person, either independently or by fusing the two sensing modalities: (a) location, (b) change in location, (c) angle of arrival (AoA), (d) change in AoA, (e) standing vs. moving, (f) direction of movement, (g) person orientation, (i) footsteps / gait, (j) phone in hand vs. in pocket, (i) scale of nearby obstacles, and (k) motion trajectory from WiFi and vision. Once we have these semantic concepts for each person from each sensing modality, the system unit performs similarity estimation between the concepts from each modality for each pair of people. As an example, let SC i Cam Become a person who uses a camera i The semantic concept of SC j Wi-Fi Become a person using Wi-Fi j The semantic concept of MAC j Is with staff j The MAC address associated with the received packet. In order to determine the SC i Cam With SC j Wi-Fi If the two semantic concepts are similar, the system uses cosine similarity, normalized weighted distance, Euclidean distance, cluster analysis, or other similarity metrics or a combination thereof. Once the two semantic concepts are found to be similar, MAC is used. j To identify personnel i , and for personnel i The actual MAC address (MAC j ) or a hash of the MAC address is tagged for future tracking, which provides a consistent identification marker.

[0028] Figure 5 is an exemplary flow chart 500 of a second embodiment for comparing a camera's motion signature with a motion signature from a wireless transceiver (as Figure 3Part of component 323). An LSTM (Long Short-Term Memory) network can be used to capture the motion signature of each person using image data, and another LSTM network can be used to capture the motion pattern using wireless CSI data. A similarity measure is applied to determine the similarity between the two motion signatures. LTSM can be a variant of RNN (Recurrent Neural Network). Other RNN networks can also be used for this purpose, such as GRU (Gated Recurrent Unit). Additionally, instead of using LSTM, an attention-based network can also be used to capture the motion signature from each sensing modality. In order to estimate the similarity between two motion signatures, the system can use cosine similarity, or normalized weighted distance, or Euclidean distance, or cluster analysis, or other similarity measures or combinations thereof.

[0029] In this embodiment, at step 505, the two input features can be fused to refine and update the input features. At steps 507 and 509, the input features of the camera data (e.g., motion trajectory, orientation information) are fed to the LSTM network N1, and the input features of the wireless packet data are fed to the LSTM network N2. Step 507 indicates that the input features are fed from the image data, and at step 509, the input features are fed from the wireless data. At step 511, the N1 network generates the motion trajectory MT of person i. i cam At step 513, the N2 network generates a latent representation of j The motion trajectory of MT j Wi-Fi At step 515, in MT i cam With MT j Wi-Fi At decision 517, the system determines whether the cosine similarity S is less than a predetermined threshold T. At step 519, if S is less than threshold T2, it reports "yes", thereby indicating that the person i and personnel j Otherwise, at step 521, the system determines that S is greater than the threshold value T2 and returns "No", indicating that the person i and personnel j Not the same person. Therefore, the system 500 can restart the process to another subset of people identified by the image data and wireless transceivers.

[0030] In another embodiment, instead of three antennas, a different number of antennas is used, such as one, two, four, and others. In another embodiment, instead of providing all CSI values ​​for all subcarriers to the LSTM model, PCA (Principal Component Analysis) is applied, and the first few principal components are used, thereby discarding CSI values ​​from noise subcarriers. In another embodiment, RSSI is used independently or in addition to CSI. In yet another embodiment, instead of people, robots or other objects carry wireless chipsets, and the aforementioned methods are used to detect, track, and identify them. In another embodiment, instead of smartphones, a small fob or device containing a wireless chipset can be carried. In another embodiment, instead of using a single system unit, multiple system units are deployed throughout an area to capture mobility patterns throughout the space.

[0031] Figure 6 is an exemplary visualization of a motion sign using a camera and a motion sign using a wireless signal (eg, Wi-Fi signal). i cam 605 is the person using the camera i The camera motion mark 605 may be a matrix that provides the motion trajectory of the person in the time window TW1. i Wi-Fi Example 603, and personnel using a wireless transceiver (such as a Wi-Fi transceiver) i Sports logo. Assumption personnel i If a person is carrying a smartphone or other mobile device that may be generating wireless (e.g., Wi-Fi) traffic, the packet is received in our system unit. The CSI value can be extracted from the received packet. The system unit can also measure the angle of arrival (AoA) of the packet by utilizing the CSI values ​​from multiple antennas and using the SpotFi algorithm (or any other type of algorithm). Using the AoA and / or the raw CSI value, the system can estimate the number of people i The trajectory of motion, such as Figure 6 Use MS i Wi-Fi As shown in 603.

[0032] The processes, methods, or algorithms disclosed herein can be delivered to, or implemented by, a processing device, controller, or computer, including any existing programmable electronic control unit or dedicated electronic control unit. Similarly, these processes, methods, or algorithms can be stored in many forms as data and instructions executable by a controller or computer, including but not limited to information permanently stored on non-writable storage media (such as ROM devices) and information revocably stored on writable storage media (such as floppy disks, magnetic tapes, CDs, RAM devices, and other magnetic and optical media). These processes, methods, or algorithms can also be implemented in software executable objects. Alternatively, these processes, methods, or algorithms can be embodied in whole or in part using suitable hardware components (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), state machines, controllers, or other hardware components or devices), or a combination of hardware, software, and firmware components.

[0033] Although exemplary embodiments have been described above, it is not intended that these embodiments describe all possible forms encompassed by the claims. The words used in the specification are descriptive rather than restrictive, and it should be understood that various changes may be made without departing from the spirit and scope of the present disclosure. As previously described, features of various embodiments may be combined to form further embodiments of the present invention, which may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desired characteristics, one of ordinary skill in the art will recognize that one or more features or characteristics may be compromised to achieve desired overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, lifecycle cost, marketability, appearance, packaging, size, repairability, weight, manufacturability, ease of assembly, etc. Thus, to the extent that any embodiment is described as less desirable with respect to one or more characteristics than other embodiments or prior art implementations, such embodiments do not exceed the scope of the present disclosure and may be desirable for a particular application.

Claims

1. A device for identifying a person using a camera and a wireless transceiver, comprising: a wireless transceiver configured to communicate packet data with a mobile device associated with one or more persons in proximity to the wireless transceiver; a camera configured to capture image data of one or more persons in the vicinity; and A controller that communicates with the wireless transceiver and the camera, the controller being configured to: receiving a plurality of packet data from a mobile device, wherein the packet data includes at least amplitude information associated with a wireless channel in communication with a wireless transceiver; determining, using the image data, a camera-related motion signature representing motion of the one or more persons, and determining, using the grouping data, a grouping-related motion signature representing motion of the one or more persons; identifying each of the one or more persons responsive to the camera-related motion indicia and the group-related motion indicia; as well as Information associated with each of the one or more persons is output responsive to the identification of the one or more persons. 2 . The apparatus of claim 1 , wherein the controller is further configured to identify the one or more persons in response to a comparison of the camera-related motion signature and the group-related motion signature. The apparatus of claim 1 , wherein the wireless transceiver comprises three or more antennas. The apparatus of claim 1 , wherein the controller is configured to communicate with one or more applications stored on a mobile device. 5 . The apparatus of claim 1 , wherein the wireless transceiver is configured to receive a media access control (MAC) address associated with the mobile device, and the controller is configured to hash the MAC address. The apparatus of claim 1 , wherein the wireless transceiver is configured to receive inertial movement data from a mobile device. 7 . The apparatus of claim 1 , wherein the controller is further configured to estimate arrival angles of the one or more persons in response to the plurality of packet data.

8. The apparatus of claim 1, wherein the controller is configured to determine the group-related motion signature using at least a long short-term memory model.

9. The apparatus of claim 1, wherein the wireless transceiver is a Wi-Fi transceiver or a Bluetooth transceiver.

10. The apparatus of claim 1, wherein the controller is further configured to determine the packet-dependent motion flag using at least the estimated angle of arrival.

11. The apparatus of claim 1, wherein the controller is further configured to determine the group-related motion signature using at least inertial movement data from a mobile device.

12. A system for identifying a person using a camera and a wireless transceiver, comprising: a wireless transceiver configured to communicate packet data with a mobile device associated with one or more persons in proximity to the wireless transceiver; a camera configured to identify the one or more persons and to identify an estimated camera-related motion signature using at least the image data; as well as A controller that communicates with the wireless transceiver and the camera, the controller being configured to: receiving a plurality of packet data from a mobile device, wherein the packet data includes at least amplitude information associated with a wireless channel in communication with a wireless transceiver; estimating a group-related motion signature for each of the one or more persons in response to the plurality of group data; as well as Each of the one or more persons is identified responsive to the estimated camera-related motion signature and the estimated group-related motion signature.

13. The system of claim 12, wherein the controller is configured to output a graphical image responsive to the identification of each of the one or more persons, the graphical image comprising a bounding box around an image of each of the one or more persons.

14. The system of claim 12, wherein the controller is further configured to determine the estimated packet-related motion flag using at least the estimated angle of arrival.

15. The system of claim 12, wherein the camera is configured to identify an estimated camera-related motion signature using at least the estimated position and the estimated tracking of the image data of the one or more persons.

16. The system of claim 12, wherein the wireless transceiver is a Wi-Fi transceiver or a Bluetooth transceiver.

17. The system of claim 12, wherein the controller is configured to determine the estimated packet-dependent motion flag in response to channel state information of the packet data.

18. A method for identifying a person using a camera and a wireless transceiver, comprising: receiving packet data from a mobile device, the packet data representing one or more persons in proximity to the wireless transceiver; obtaining image data representing one or more persons associated with a camera; determining an estimated group-related motion signature of the one or more persons using the group data from the mobile device; determining an estimated camera-related motion signature of the one or more persons using the image data; comparing the estimated group-dependent motion flag with the estimated camera-dependent motion flag; as well as The one or more persons are identified responsive to the comparing.

19. The method of claim 18, wherein the packet data includes channel state information.

20. The method of claim 18, wherein the packet data comprises wireless channel state information data.

Citation Information

Patent Citations

  • Personnel trajectory tracking method based on CSI (Channel State Information)

    CN108882171A

  • Methods, devices, servers, apparatus, and systems for wireless internet of things applications

    WO2017155634A1