System and method for tracking physical location

By using a distributed tracking system and a specially designed arrangement of cameras and weight sensors, the problems of computing power and synchronization were solved, enabling efficient and accurate position tracking in large spaces.

CN114830187BActive Publication Date: 2026-01-277-ELEVEN INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080083466.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2020-10-23
Publication Date
2026-01-27
Estimated Expiration
2040-10-23

AI Technical Summary

Technical Problem

Existing location tracking systems are limited by computer computing power and sensor synchronization issues when expanded in large physical spaces, making it difficult to accurately track the positions of people and objects.

Method used

A distributed tracking system is adopted, using a camera array, multiple camera clients, camera servers, weight sensors, weight servers, and a central server. Information is coordinated through timestamp and time window processing, combined with a specific arrangement of cameras and weight sensors to improve system scalability and accuracy.

Benefits of technology

It enables efficient and accurate tracking of people and objects in a larger space, reduces the impact of asynchrony, and improves the system's resilience and coverage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114830187B_ABST
    Figure CN114830187B_ABST
Patent Text Reader

Abstract

An extensible tracking system processes video of a space to track the location of a person within the space. The tracking system determines local coordinates of the person within a frame of video, and then assigns those coordinates to a time window based on when the frame was received. The tracking system then combines or clusters certain local coordinates that have been assigned to the same time window to determine the combined coordinates of the person during that time window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to distributed systems for tracking the physical location of people and objects. Background Technology

[0002] Location tracking systems are used to track the physical location of people and / or objects. Summary of the Invention

[0003] Location tracking systems are used to track the physical location of people and / or objects in a physical space, such as a store. These systems typically use sensors (e.g., cameras) to detect the presence of people and / or objects and use computers to determine their physical location based on signals from the sensors. In a store environment, other types of sensors can be installed to track the movement of inventory within the store. For example, weight sensors can be installed on shelves and partitions to determine when items are removed from those shelves and partitions. By tracking the location of people in the store and when items are removed from partitions, a computer can potentially determine which user in the store removed an item and charge that user for it without having to pay at a checkout. In other words, a person can walk into the store, take an item, and leave without stopping for the regular checkout process.

[0004] For larger physical spaces (e.g., convenience stores and grocery stores), additional sensors can be installed throughout the space to track the positions of people and / or objects as they move within the space. For example, additional cameras can be added to track positions in a larger space, and additional weight sensors can be added to track additional items and partitions. However, the number of sensors that can be added is limited until the computing power of computers is realized. Therefore, the computing power of computers limits the coverage of tracking systems.

[0005] One way to expand these systems to handle larger spaces is to add additional computers and distribute the sensors among these computers, so that each computer processes a subset of the signals from the sensors. However, distributing sensors among multiple computers introduces synchronization problems. For example, sensors will not send signals to their corresponding computers at the same time or simultaneously. As another example, sensors and their corresponding computers can have different time delays, so a signal from one sensor may take longer to reach the computer than a signal from another sensor. Consequently, the sensors and computers become out of sync, and it becomes more difficult for the computers to determine the location of people or objects in the space and when items were removed in a coherent manner.

[0006] This disclosure envisions an unconventional distributed tracking system that can be scaled to handle larger spaces. The system uses an array of cameras, multiple camera clients, a camera server, weight sensors, a weight server, and a central server to determine which person in the space took an item and should be charged for it. Each camera client processes video frames from a different subset of cameras in the camera array. Each camera client determines the coordinates of a person detected in the frame and then timestamps these coordinates based on the time the camera client received the frame. The camera client then transmits the coordinates and timestamps to the camera server, which coordinates the information from the camera clients. The camera server determines the person's location in the space based on the coordinates and timestamps from the camera clients. The weight server processes signals from the weight sensors to determine when an item was removed from a partition in the space. The central server uses the location of the person in the space from the camera server and the determination from the weight server regarding when the item was removed from the partition to determine which person in the space took which items and therefore should be charged.

[0007] Generally, camera servers prevent desynchronization by assigning coordinates from multiple camera clients to time windows based on timestamps. The camera server then processes the coordinates assigned to a specific time window to determine the overall coordinates of a person in space during that time window. The duration of the time window can be set longer than the expected desynchronization to mitigate its effects. For example, if cameras and camera clients are expected to be asynchronous for a few milliseconds, the time window can be set to last 100 milliseconds to combat desynchronization. In this way, the number of cameras and camera clients can be increased to scale the system to handle any suitable space.

[0008] This disclosure also envisions an unconventional method for connecting cameras in a camera array to camera clients. The cameras are arranged in a rectangular grid above space. Each camera in the grid is connected to a specific camera client according to certain rules. For example, no two directly adjacent cameras in the same row or column of the grid are connected to the same camera client. As another example, cameras arranged diagonally along the grid are connected to the same camera client. In this way, a small area of ​​the grid should include the cameras connected to each individual camera client in the system. Therefore, even if one camera client goes offline (e.g., for maintenance, error, or crash), the remaining camera clients still have sufficient coverage to track the location of a person in that small area. Thus, this arrangement of cameras improves the resilience of the system.

[0009] This disclosure also envisions an unconventional shelf and partition design that integrates weight sensors to track when items are removed from the shelves and partitions. Generally, the shelf includes a base, vertical panels, and shelves. The base forms an enclosed space in which a printed circuit board is positioned, and the base includes a drawer that opens to provide access to the enclosed space and the circuit board. The vertical panels are attached to the base, and the shelves are attached to the vertical panels. Weight sensors are positioned within the shelves. Each of the base, panel, and shelf has a custom cavity. The cavities in the shelves and the cavities in the panels are at least partially aligned. Each weight sensor transmits a signal to the printed circuit board via a cable that passes from the weight sensor through the cavities of the shelf, the panel, and the base to the circuit board.

[0010] Some embodiments include an unconventional tracking system comprising separate components (e.g., a camera client, a camera server, a weight server, and a central server) that perform different functions to track the position of people and / or objects in a space. By extending the system's functionality across these various components, the system is able to handle signals from more sensors (e.g., cameras and weight sensors). Due to the increased number of sensors, the system can track people and / or objects in a larger space. Therefore, the system can be expanded to handle larger spaces (e.g., by adding additional camera clients). Some embodiments of the tracking system are described below.

[0011] According to an embodiment, a system includes an array of cameras, a first camera client, a second camera client, a camera server, a plurality of weight sensors, a weight server, and a central server. The camera array is positioned above a space. Each camera in the camera array captures video of a portion of the space. The space contains a person. The first camera client receives a first plurality of frames of a first video from the first camera of the camera array. Each of the first plurality of frames shows a person within the space. For the first frame of the first plurality of frames, the first camera client determines a first defined region around the person shown in the first frame and generates a first timestamp of when the first camera client received the first frame. For a second frame of the first plurality of frames, the first camera client determines a second defined region around the person shown in the second frame and generates a second timestamp of when the first camera client received the second frame. The second camera client is separate from the first camera client. The second camera client receives a second plurality of frames of a second video from the second camera of the camera array. Each of the second plurality of frames shows a person within the space. For a third frame of the second plurality of frames, the second camera client determines a third defined region around the person shown in the third frame and generates a third timestamp of when the second camera client received the third frame. For the fourth frame in the second plurality of frames, the second camera client determines the fourth defined area around the person shown in the fourth frame and generates a fourth timestamp of when the second camera client received the fourth frame.

[0012] The camera server is separate from the first and second camera clients. The camera server determines that a first timestamp falls within a first time window and, in response, assigns the coordinates defining a first bounding region to the first time window. The camera server also determines that a second timestamp falls within the first time window and, in response, assigns the coordinates defining a second bounding region to the first time window. The camera server further determines that a third timestamp falls within the first time window and, in response, assigns the coordinates defining a third bounding region to the first time window. The camera server determines that a fourth timestamp falls within a second time window following the first time window and, in response, assigns the coordinates defining a fourth bounding region to the second time window.

[0013] The camera server also determines that the coordinates assigned to the first time window should be processed, and in response to determining that the coordinates assigned to the first time window should be processed, the camera server calculates the combined coordinates of a person during the first time window for the first video from the first camera, based at least on the coordinates defining a first bounding region and the coordinates defining a second bounding region, and calculates the combined coordinates of a person during the first time window for the second video from the second camera, based at least on the coordinates defining a third bounding region. The camera server also determines the person's spatial position during the first time window based at least on the combined coordinates of the person during the first time window for the first video from the first camera and the combined coordinates of the person during the first time window for the second video from the second camera.

[0014] Multiple weight sensors are positioned within the space. Each of the multiple weight sensors generates a signal indicating the weight borne by that weight sensor. A weight server is separate from the first and second camera clients and the camera server. The weight server determines that an item positioned above the first weight sensor has been removed, based at least on the signal generated by the first weight sensor. A central server is separate from the first and second camera clients, the camera server, and the weight server. The central server determines that an item has been removed, based at least on the person's position within the space during a first time window. Based at least on the determination that the first person removed an item, a fee is charged to that person when that person leaves the space for that item.

[0015] According to another embodiment, a system includes an array of cameras, a first camera client, a second camera client, a camera server, multiple weight sensors, a weight server, and a central server. The camera array is positioned above a space. Each camera in the camera array captures video of a portion of the space. The space contains a person. For each frame of a first video received from the first camera of the camera array, the first camera client determines a defined area around the person shown in that frame of the first video and generates a timestamp of when that frame of the first video was received by the first camera client. For each frame of a second video received from the second camera of the camera array, the second camera client determines a defined area around the person shown in that frame of the second video and generates a timestamp of when that frame of the second video was received by the second camera client.

[0016] The camera server is separate from the first and second camera clients. For each frame of the first video, the camera server assigns the coordinates of the bounding region around the person shown in that frame to one of a plurality of time windows, based at least on the timestamp of when that frame was received by the first camera client. For each frame of the second plurality of frames, the camera server assigns the coordinates of the bounding region around the person shown in that frame to one of a plurality of time windows, based at least on the timestamp of when that frame was received by the second camera client. For the first time window of the plurality of time windows, the camera server calculates the combined coordinates of the person during the first time window for the first video from the first camera, based at least on (1) the bounding region around the person shown in the first plurality of frames and (2) the coordinates assigned to the first time window, and calculates the combined coordinates of the person during the first time window for the second video from the second camera, based at least on (1) the bounding region around the person shown in the second plurality of frames and (2) the coordinates assigned to the first time window. The camera server determines the person's spatial position during the first time window based at least on the combined coordinates of the person during the first time window for the first video from the first camera and the combined coordinates of the person during the first time window for the second video from the second camera.

[0017] Multiple weight sensors are located within the space. The weight server is separate from the first and second camera clients and the camera server. The weight server determines that an object positioned above the first weight sensor has been removed, based at least on the signal generated by the first weight sensor among the multiple weight sensors. The central server is separate from the first and second camera clients, the camera server, and the weight server. The central server determines that an object has been removed, based at least on the person's position within the space during a first time window.

[0018] Some embodiments of the tracking system perform an unconventional tracking process that allows for some degree of asynchrony between system components (e.g., camera client and camera server). Generally, the system processes information according to time windows. These time windows can be set to be larger than the expected asynchrony in the system. Information assigned to a time window is processed together. Therefore, even if there is some asynchrony between the information, it is still processed together within the same time window. In this way, the tracking system can handle increased amounts of asynchrony, especially as the system is scaled up to include more components, allowing the system to handle a larger space. Thus, the system can scale up to handle a larger space while maintaining reliability and accuracy. Some embodiments of the tracking process are described below.

[0019] According to an embodiment, a system includes an array of cameras, a first camera client, a second camera client, and a camera server. The array of cameras is positioned above a space. Each camera in the array of cameras captures video of a portion of the space. The space contains a person. The first camera client receives first plurality of frames of the first video from the first cameras of the camera array. Each of the first plurality of frames shows a person within the space. For the first frame of the first plurality of frames, the first camera client determines a first defined region around the person shown in the first frame and generates a first timestamp of when the first camera client received the first frame. For a second frame of the first plurality of frames, the first camera client determines a second defined region around the person shown in the second frame and generates a second timestamp of when the first camera client received the second frame. For a third frame of the first plurality of frames, the first camera client determines a third defined region around the person shown in the third frame and generates a third timestamp of when the first camera client received the third frame.

[0020] The second camera client receives a second plurality of frames of a second video from the second camera in the camera array. Each of the second plurality of frames shows a person within space. For the fourth frame of the second plurality of frames, the second camera client determines a fourth defined region around the person shown in the fourth frame and generates a fourth timestamp indicating when the second camera client received the fourth frame. For the fifth frame of the second plurality of frames, the second camera client determines a fifth defined region around the person shown in the fifth frame and generates a fifth timestamp indicating when the second camera client received the fifth frame.

[0021] The camera server is separate from the first and second camera clients. The camera server determines that a first timestamp falls within a first time window and, in response, assigns the coordinates defining a first bounding region to the first time window. The camera server also determines that a second timestamp falls within the first time window and, in response, assigns the coordinates defining a second bounding region to the first time window. The camera server further determines that a third timestamp falls within a second time window following the first time window and, in response, assigns the coordinates defining a third bounding region to the second time window. The camera server also determines that a fourth timestamp falls within the first time window and, in response, assigns the coordinates defining a fourth bounding region to the first time window. The camera server further determines that a fifth timestamp falls within the second time window and, in response, assigns the coordinates defining a fifth bounding region to the second time window.

[0022] The camera server further determines that the coordinates assigned to the first time window should be processed, and in response to determining that the coordinates assigned to the first time window should be processed, the camera server processes the combined coordinates of the first video computer from the first camera during the first time window based at least on the coordinates of a defined first bounding region and a defined second bounding region, and processes the combined coordinates of the second video computer from the second camera during the first time window based at least on the coordinates of a defined fourth bounding region. After determining that the coordinates assigned to the first time window should be processed, the camera server determines that the coordinates assigned to the second time window should be processed, and in response to determining that the coordinates assigned to the second time window should be processed, the camera server processes the combined coordinates of the first video computer from the first camera during the second time window based at least on the coordinates of a defined third bounding region, and processes the combined coordinates of the second video computer from the second camera during the second time window based at least on the coordinates of a defined fifth bounding region.

[0023] According to another embodiment, a system includes an array of cameras, a first camera client, a second camera client, and a camera server. The camera array is positioned above a space. Each camera in the camera array captures video of a portion of the space. The space contains people. The first camera client receives a first plurality of frames of the first video from the first cameras of the camera array. Each frame of the first plurality of frames shows a person within the space. For each frame of the first plurality of frames, the first camera client determines a defined area around the person shown in that frame and generates a timestamp of when that frame was received by the first camera client. The second camera client receives a second plurality of frames of the second video from the second cameras of the camera array. Each frame of the second plurality of frames shows a person within the space. For each frame of the second plurality of frames, the second camera client determines a defined area around the person shown in that frame and generates a timestamp of when that frame was received by the second camera client.

[0024] The camera server is separate from the first and second camera clients. For each of the first plurality of frames, the camera server assigns the coordinates of the defined area around the person shown in that frame to one of a plurality of time windows, based at least on the timestamp of that frame received by the first camera client, and for each of the second plurality of frames, the coordinates of the defined area around the person shown in that frame are assigned to one of a plurality of time windows, based at least on the timestamp of that frame received by the second camera client.

[0025] The camera server also determines that the coordinates of a first time window assigned to multiple time windows should be processed, and in response to determining that the coordinates assigned to the first time window should be processed, calculates the combined coordinates of the person during the first time window for a first video from a first camera based at least on (1) defining the bounding area around the person shown in the first multiple frames and (2) assigning the coordinates to the first time window, and calculates the combined coordinates of the person during the first time window for a second video from a second camera based at least on (1) defining the bounding area around the person shown in the second multiple frames and (2) assigning the coordinates to the first time window.

[0026] Some embodiments include an unconventional arrangement of cameras and camera clients that enhances the resilience of the camera system. Generally, cameras are arranged in a rectangular grid providing coverage of the physical space, and each camera is communicatively coupled to a single camera client. No camera in the same row or column of the grid is directly adjacent to another camera communicatively coupled to the same camera client. Cameras arranged diagonally along the grid are communicatively coupled to the same camera client. In this way, even if one camera client in the system goes offline, the grid still provides sufficient coverage of the physical space. Therefore, the camera arrangement improves the resilience of the system. Some embodiments of the camera arrangement are described below.

[0027] According to an embodiment, a system includes a first camera client, a second camera client, a third camera client, and an array of cameras. The second camera client is separate from the first camera client. The third camera client is separate from the first and second camera clients. The camera array is positioned above space. The cameras in the camera array are arranged in a rectangular grid, including a first row, a second row, a third row, a first column, a second column, and a third column. The array includes a first, second, third, fourth, fifth, and sixth camera.

[0028] The first camera is positioned in the first row and first column of the grid. The first camera is communicatively coupled to a first camera client. The first camera transmits video of the first portion of the space to the first camera client. The second camera is positioned in the first row and second column of the grid, making it directly adjacent to the first camera within the grid. The second camera is communicatively coupled to a second camera client. The second camera transmits video of the second portion of the space to the second camera client. The third camera is positioned in the first row and third column of the grid, making it directly adjacent to the second camera within the grid. The third camera is communicatively coupled to a third camera client. The third camera transmits video of the third portion of the space to the third camera client. The fourth camera is positioned in the second row and first column of the grid, making it directly adjacent to the first camera within the grid. The fourth camera is communicatively coupled to a second camera client. The fourth camera transmits video of the fourth portion of the space to the second camera client. The fifth camera is positioned in the second row and second column of the grid, making it directly adjacent to both the fourth and second cameras within the grid. The fifth camera is communicatively coupled to a third camera client. The fifth camera transmits video of the fifth portion of the space to the third camera client. The sixth camera is positioned in the third row and first column of the grid, making it directly adjacent to the fourth camera within the grid. The sixth camera's communication is coupled to the third camera client. The sixth camera transmits video from the sixth segment of the space to the third camera client.

[0029] According to another embodiment, a system includes a plurality of camera clients and an array of cameras. The plurality of camera clients includes several camera clients. The array of cameras is positioned above space. Each camera in the camera array transmits video of a portion of the space to only one of the plurality of camera clients. The cameras in the camera array are arranged such that each of the plurality of camera clients is communicatively coupled to at least one camera in an N×N portion of the array. N is the number of camera clients minus one.

[0030] Some embodiments include unconventional shelving for holding items. The shelving includes a base and a panel for holding shelves and weight sensors. The weight sensors are wired to a circuit board located in a drawer within the base. Cables run from the weight sensors through cavities and spaces defined by the shelves, panel, and base. Some embodiments of the shelving are described below.

[0031] According to an embodiment, a system includes a circuit board and a shelf. The shelf includes a base, a panel, a shelf, a first weight sensor, a second weight sensor, a first cable, and a second cable. The base includes a bottom surface, a first side surface, a second side surface, a third side surface, a top surface, and a drawer. The first side surface is coupled to the bottom surface of the base. The first side surface of the base extends upward from the bottom surface of the base. The second side surface is coupled to the bottom of the base and the first side surface. The second side surface of the base extends upward from the bottom surface of the base. The third side surface is coupled to the bottom of the base and the second side surface. The third side surface of the base extends upward from the bottom surface of the base. The top surface is coupled to the first, second, and third side surfaces of the base such that the bottom surface and the top surface of the base define a space with the first, second, and third side surfaces of the base. The top surface of the base defines a first opening for entering the space. The drawer is positioned within the space. The circuit board is positioned within the drawer.

[0032] A panel is coupled to and extends upward from a base. The panel defines a second opening extending along its width. A shelf is coupled to the panel such that the shelf is vertically positioned above the base and extends away from the panel. The shelf includes a bottom surface, a front surface extending upward from the bottom surface, and a rear surface extending upward from the bottom surface. The rear surface of the shelf is coupled to the panel. The rear surface of the shelf defines a third opening. A portion of the third opening is aligned with a portion of the second opening.

[0033] A first weight sensor is coupled to the bottom surface of the shelf and positioned between the front and rear surfaces of the shelf. A second weight sensor is coupled to the bottom surface of the shelf and positioned between the front and rear surfaces of the shelf. A first cable is coupled to the first weight sensor and the circuit board. The first cable extends from the first weight sensor through second and third openings and downwards into the space through the first opening. A second cable is coupled to the second weight sensor and the circuit board. The second cable extends from the second weight sensor through second and third openings and downwards into the space through the first opening.

[0034] Some embodiments may exclude, include, or incorporate some or all of the technical advantages discussed above. From the accompanying drawings, description, and claims included herein, those skilled in the art will readily understand one or more other technical advantages. Attached Figure Description

[0035] To gain a more complete understanding of this disclosure, reference is now made to the following description in conjunction with the accompanying drawings, wherein:

[0036] Figure 1A-1C The illustration shows an example store that defines a physical space;

[0037] Figure 2 The diagram illustrates a sample tracking system used in a physical store;

[0038] Figure 3A-3T The illustration shows an example camera subsystem and its operation in a tracking system;

[0039] Figures 4A-4D The illustration shows an example light detection and ranging subsystem and its operation in a tracking system;

[0040] Figure 5A-5J The illustration shows an example weight subsystem and its operation in a tracking system;

[0041] Figures 6A-6C The diagram illustrates the operation of a sample central server used in conjunction with a tracking system; and

[0042] Figure 7 The example computer is illustrated. Detailed Implementation

[0043] By referring to the attached figures Figures 1A to 7 To best understand the embodiments and advantages of this disclosure, the same reference numerals are used for the same and corresponding parts of the various figures. Additional information is disclosed in U.S. Patent Application No. ___ entitled “Customer-Based Video Feed” (Attorney General’s File No. 090278.0187) and U.S. Patent Application No. ___ entitled “Topview Object Tracking Using a Sensor Array” (Attorney General’s File No. 090278.0180), both of which are incorporated herein by reference as if their entire contents were reproduced.

[0044] Location tracking systems are used to track the physical location of people and / or objects in a physical space, such as a store. These systems typically use sensors (e.g., cameras) to detect the presence of people and / or objects and use computers to determine their physical location based on signals from the sensors. In a store environment, other types of sensors can be installed to track the movement of inventory within the store. For example, weight sensors can be installed on shelves and partitions to determine when items are removed from those shelves and partitions. By tracking the location of people in the store and when items are removed from partitions, a computer can potentially determine which user in the store removed an item and charge that user for it without having to pay for it at a checkout. In other words, a person can walk into the store, take an item, and leave without stopping for the regular checkout process.

[0045] For larger physical spaces (e.g., convenience stores and grocery stores), additional sensors can be installed throughout the space to track the positions of people and / or objects as they move within it. For example, additional cameras can be added to track positions in a larger space, and additional weight sensors can be added to track additional items and partitions. However, the number of sensors that can be added is limited until the computing power of computers is realized. Therefore, the computing power of computers limits the coverage of tracking systems.

[0046] One way to expand these systems to handle larger spaces is to add additional computers and distribute the sensors among these computers, so that each computer processes a subset of the signals from the sensors. However, distributing the sensors among multiple computers introduces synchronization problems. For example, sensors may not send signals to their respective computers at the same time or simultaneously. As another example, sensors and their corresponding computers may have different time delays, so a signal from one sensor may take longer to reach the computer than a signal from another sensor. Consequently, the sensors and computers become out of sync, and it becomes more difficult for the computers to determine the location of people or objects in the space and when items were removed in a coherent manner.

[0047] This disclosure envisions an unconventional distributed tracking system that can be scaled to handle larger spaces. The system uses an array of cameras, multiple camera clients, a camera server, weight sensors, a weight server, and a central server to determine which person in the space took an item and should be charged for it. Each camera client processes video frames from a different subset of cameras in the camera array. Each camera client determines the coordinates of a person detected in the frame and then timestamps these coordinates based on the time the camera client received the frame. The camera client then transmits the coordinates and timestamps to the camera server, which coordinates the information from the camera clients. The camera server determines the person's location in the space based on the coordinates and timestamps from the camera clients. The weight server processes signals from the weight sensors to determine when an item was removed from a partition in the space. The central server uses the person's location in the space from the camera servers and the determination from the weight server regarding when the item was removed from the partition to determine which person in the space took which items and therefore should be charged. Figure 1A-7 Describe the system in more detail.

[0048] Generally, camera servers prevent desynchronization by assigning coordinates from multiple camera clients to time windows based on timestamps. The camera server then processes the coordinates assigned to a specific time window to determine the overall coordinates of a person in space during that time window. The duration of the time window can be set longer than the expected desynchronization to mitigate its effects. For example, if cameras and camera clients are expected to be asynchronous for a few milliseconds, the time window can be set to last 100 milliseconds to combat desynchronization. In this way, the number of cameras and camera clients can be increased to expand the system to handle any suitable space. [The last sentence appears to be incomplete and possibly refers to a different topic.] Figure 1A-3Q A more detailed description of the camera, camera client, and camera server.

[0049] This disclosure also envisions an unconventional method for connecting cameras in a camera array to camera clients. The cameras are arranged in a rectangular grid above space. Each camera in the grid is connected to a specific camera client according to certain rules. For example, no two directly adjacent cameras in the same row or column of the grid are connected to the same camera client. As another example, cameras arranged diagonally along the grid are connected to the same camera client. In this way, a small area of ​​the grid should include the cameras connected to each individual and each camera client in the system. Therefore, even if one camera client goes offline (e.g., maintenance, error, or crash), the remaining camera clients still have sufficient coverage to track the location of a person in that small area. Thus, this arrangement of cameras improves the resilience of the system. [The following text will be used...] Figures 3A-3E Describe the camera array in more detail.

[0050] This disclosure also envisions an unconventional shelf and partition design that integrates weight sensors for tracking when items are removed from the shelves and partitions. Generally, the shelf includes a base, vertical panels, and shelves. The base forms an enclosed space in which a printed circuit board is positioned, and the base includes a drawer that opens to provide access to the enclosed space and the circuit board. The vertical panels are attached to the base, and the shelves are attached to the vertical panels. Weight sensors are positioned within the shelves. Each of the base, panel, and shelf has a custom cavity. The cavities in the shelves and panels are at least partially aligned. Each weight sensor transmits a signal to the printed circuit board via a cable that passes from the weight sensor through the cavities of the shelf, the panel, and the base to the circuit board. Figures 5A-5J Describe the shelving and shelf design in more detail.

[0051] The system may also include the use of Figures 4A-4D A more detailed description of the light detection and ranging (LiDAR) subsystem. This system also includes a central server that binds the camera subsystem, weight subsystem, and LiDAR subsystem together. [The system will use...] Figures 6A-6CA more detailed description of the central server.

[0052] I. System Overview

[0053] Figures 1A-1C A tracking system installed in an example store space is illustrated. As discussed above, a tracking system can be installed in a store space so that shoppers do not need to participate in a traditional checkout process. Although the example of a store space is used in this disclosure, this disclosure envisions that tracking systems can be installed and used in any type of physical space, such as warehouses, storage centers, amusement parks, airports, office buildings, etc. In general, tracking systems (or components thereof) are used to track the location of people and / or objects within these spaces for any suitable purpose. For example, in an airport, a tracking system can track the location of passengers and employees for security purposes. As another example, in an amusement park, a tracking system can track the location of park visitors to measure the popularity of attractions. As yet another example, in an office building, a tracking system can track the location of employees and staff to monitor their productivity levels.

[0054] Figure 1A Example store 100 is shown. Store 100 is a physical space where shoppers can purchase items for sale. Figure 1A As seen in the diagram, store 100 is a physical building including an entrance passageway through which shoppers enter and exit. A tracking system can be installed in store 100, allowing shoppers to purchase items without participating in a traditional checkout process. This disclosure envisions store 100 as any suitable physical space. For example, store 100 could be a convenience store or grocery store. This disclosure also envisions store 100 not being a physical building, but rather a physical space or environment where shoppers can shop. For example, store 100 could be an airport food storage room, an office building kiosk, a park outdoor market, etc.

[0055] Figure 1B A portion of the interior of store 100 is shown. (For example...) Figure 1B As seen in the image, store 100 includes shopper 105, shelves 115, and door 125. Shopper 105 may have entered store 100 through one of the doors 125 that allow entry and exit from store 100. Door 125 prevents shopper 105 from entering and / or leaving the store unless door 125 is opened.

[0056] Door 125 may include scanners 110 and 120. Scanners 110 and 120 may include QR code scanners, barcode scanners, or any other suitable type of scanner capable of receiving electronic codes embedded with information such as information uniquely identifying shopper 105. Shopper 105 may scan a personal device (e.g., a smartphone) on scanner 110 to enter store 100. When shopper 105 scans their personal device on scanner 110, the personal device may provide scanner 110 with an electronic code uniquely identifying shopper 105. When shopper 105 is identified and / or authenticated, door 125, including scanner 110, opens to allow shopper 105 to enter store 100. Each shopper 105 may have already registered an account with store 100 to receive an identification code for their personal device.

[0057] After entering store 100, shopper 105 can move around inside store 100. As shopper 105 moves through the space, shopper 105 can purchase item 130 by removing item 130 from shelf 115. Figure 1B As seen in the image, store 100 includes shelves 115 for holding items 130. When a shopper 105 wishes to purchase a specific item 130, the shopper 105 can remove the item 130 from the shelves 115. The shopper 105 can remove multiple items 130 from store 100 to purchase those items 130.

[0058] When shopper 105 finishes shopping for item 130, shopper 105 approaches door 125. In some embodiments, door 125 will open automatically so that shopper 105 can leave store 100. In other embodiments, shopper 105 scans a personal device on scanner 120 before door 125 opens to allow shopper 105 to leave store 100. When shopper 105 scans their personal device on scanner 120, the personal device can provide an electronic code that uniquely identifies shopper 105 to indicate that shopper 105 is leaving store 100. When shopper 105 leaves store 100, a charge is applied to shopper 105's account for item 130 removed from store 100.

[0059] Figure 1C The interior of store 100 is shown, along with a tracking system 132 that allows shoppers 105 to simply leave store 100 with their items 130 without participating in the traditional checkout process. Figure 1C As seen in the image, the tracking system 132 includes an array 135 of sensors positioned on the ceiling of store 100. The sensor array 135 can provide coverage of the interior space of store 100. (See also: Regarding...) Figures 3A-3ETo explain further in detail, sensors 135 are arranged in a grid pattern across the ceiling of store 100. Sensors 135 can be used to track the position of shopper 105 within the space of store 100. This disclosure envisions sensors 135 as any suitable sensor (e.g., camera, light detection and ranging sensor, millimeter-wave sensor, etc.).

[0060] The tracking system 132 also includes a weight sensor 140 positioned on the shelf 115. The weight sensor 140 can detect the weight of the item 130 positioned on the shelf 115. When the item 130 is removed from the shelf 115, the weight sensor 140 can detect a decrease in weight. The tracking system 132 can use that information to determine that a specific item 130 has been removed from the shelf 115.

[0061] Tracking system 132 includes computer system 145. Computer system 145 may include multiple computers operating together to determine which shopper 105 took which items 130 from shelf 115. This will be used... Figures 2 to 7 The components of computer system 145 and their operation are described in more detail. Generally, computer system 145 uses information from sensor 135 and weight sensor 140 to determine which shopper 105 removed which items 130 from store 100. In this way, when shopper 105 leaves store 100 through door 125, shopper 105 can be automatically charged for the items 130.

[0062] Figure 2 A block diagram of an example tracking system 132 used in store 100 is shown. Figure 2 As seen in the diagram, the tracking system 132 includes a camera subsystem 202, a light detection and ranging (LiDAR) subsystem 204, and a weight subsystem 206. The tracking system 132 includes various sensors 135, such as a camera 205, a light detection and ranging (LiDAR) sensor 210, and a weight sensor 215. These sensors 135 are communicatively coupled to various computers in the computer system 145. For example, the camera subsystem 202 includes a camera 205, which is communicatively coupled to one or more camera clients 220. These camera clients 220 are communicatively coupled to a camera server 225. The LiDAR subsystem 204 includes a LiDAR sensor 210 communicatively coupled to a LiDAR server 230. The weight subsystem 206 includes a weight sensor 215 communicatively coupled to a weight server 235. The camera server 225, LiDAR server 230, and weight server 235 are communicatively coupled to a central server 240.

[0063] Generally, camera 205 generates video of portions of the space. This video may include frames or images of shoppers 105 within the space. Camera client 220 processes frames from camera 205 to detect shoppers 105 within the frames and assigns frame coordinates to those shoppers 105. Camera server 225 generally processes frame data from camera client 220 to determine the physical location of shoppers 105 within the space. LiDAR sensor 210 generally generates coordinates of shoppers 105 within the space. LiDAR server 230 processes these coordinates to determine the location of shoppers 105 within the space. Weight sensor 215 detects the weight of items 130 on shelves 115 within the space. Weight server 235 processes these weights to determine when certain items 130 have been removed from shelves 115.

[0064] Central server 240 processes location information of shopper 105 from camera server 225 and LiDAR server 230, and weight information from weight server 235, to determine which shopper 105 removed which items 130 from shelf 115. When shopper 105 leaves the space, they can then be charged for those items 130. Figures 3A to 6C Describe the operation of these components in more detail.

[0065] In one embodiment, each component of the tracking system 132 (e.g., camera client 220, camera server 225, LiDAR server 230, weight server 235, and central server 240) is a separate computing device from the other components of the tracking system 132. For example, each of these components may include its own processor, memory, and physical housing. In this way, the components of the tracking system 132 are distributed to provide additional computing power compared to a tracking system that includes only one computer.

[0066] II. Camera Subsystem

[0067] Figure 3A-3R An example camera subsystem 202 and its operation within the tracking system 132 are illustrated. As discussed above, the camera subsystem 202 includes a camera 205, a camera client 220, and a camera server 225. Generally, the camera 205 captures video of the space and sends the video to the camera client 220 for processing. These videos are a series of frames or images of the space. The camera client 220 detects the presence of people (e.g., shoppers 105) in the frames and determines the coordinates of these people in the frames (also referred to as "frame coordinates"). The camera server 225 analyzes the frame coordinates from each camera client 220 to determine the physical location of the people in the space.

[0068] 1. Camera array

[0069] Figure 3A The illustration shows an example camera array 300. (Example:) Figure 3A As shown, camera array 300 includes a plurality of cameras 305. While this disclosure shows a camera array 300 including twelve cameras 305, camera array 300 may include any suitable number of cameras 305. Generally, camera array 300 is positioned above a space such that the cameras 305 can capture overhead video of portions of the space. This video can then be processed by other components of camera subsystem 202 to determine the physical location of people (e.g., shopper 105) within the space. Figure 3A In the example, camera array 300 includes cameras 305A, 305B, 305C, 305D, 305E, 305F, 305G, 305H, 305I, 305J, 305K and 305L.

[0070] Generally, the cameras 305 in the camera array 300 are arranged to form a rectangular array. Figure 3A In the example, camera array 300 is a 3×4 array of cameras 305 (e.g., cameras 305 arranged in three rows and four columns). Camera array 300 may include any suitable number of cameras 305 arranged in an array of any suitable size.

[0071] Each camera 305 of the camera array 300 is communicatively coupled to a camera client 220. Figure 3A In the example, each camera 305 of the camera array 300 is communicatively coupled to one of camera client 1 220A, camera client 2 220B, or camera client 3220C. Each camera 305 transmits captured video to the camera client 220 to which it is communicatively coupled. The cameras 305 are communicatively coupled to the camera client 220 according to certain rules to improve the resilience of the tracking system 132. Generally, the cameras 305 are communicatively coupled to the camera client 220 such that even if one camera client 220 is offline, the physical space coverage provided by the cameras 305 communicatively coupled to the remaining camera clients 220 is sufficient to allow the tracking system 132 to continue tracking the position of a person in the space.

[0072] Camera 305 is communicatively coupled to camera client 220 using any suitable medium. For example, camera 305 can be hardwired to camera client 220. As another example, camera 305 can be wirelessly coupled to camera client 220 using any suitable wireless protocol (e.g., WiFi). Camera 305 transmits captured video to camera client 220 via the communication medium.

[0073] Camera 305 can be any suitable device for capturing video of space. For example, camera 305 can be a 3D camera capable of capturing two-dimensional video of space (e.g., the xy plane) and also detecting the height of people and / or objects in the video (e.g., the z plane). As another example, camera 305 can be a 2D camera capturing two-dimensional video of space. Camera array 300 can include a mixture of different types of cameras 305.

[0074] Figure 3B The illustration shows the coverage provided by camera 305 of camera array 300. (See diagram for reference.) Figure 3B As shown, the ground space is covered by different fields of view 310. Each field of view 310 is provided by a camera 305 of the camera array 300. For example, field of view 310A is provided by camera 305A. Field of view 310B is provided by camera 305B. Field of view 310C is provided by camera 305C, and so on. Each field of view 310 is generally rectangular in shape and covers a portion of the ground space. Each camera 305 captures video of the portion of the ground space covered by that camera's field of view 310. For example, camera 305A captures video of the portion of the ground space covered by field of view 310A. Camera 305B captures video of the portion of the ground space covered by field of view 310B. Camera 305C captures video of the portion of the ground space covered by field of view 310C, and so on.

[0075] The shading of each field of view 310 differs from that of its neighbors to distinguish fields of view 310. Fields of view 310A, 310C, 310I, and 310K are shaded using lines sloping downwards and to the right. Fields of view 310B, 310D, 310J, and 310L are shaded using lines sloping upwards and to the right. Fields of view 310E and 310G are shaded using horizontal lines, and fields of view 310F and 310H are shaded using vertical lines. The shading of each field of view 310 is intended to distinguish that field of view 310 from other directly adjacent fields of view 310. The shading is not intended to indicate a specific characteristic of a field of view 310. In other words, even if some fields of view 310 share the same shading, similar shading does not indicate that these fields of view 310 share certain characteristics (e.g., size, coverage, duration, and / or shape). Fields of view 310 may share one or more of these characteristics regardless of their respective shading.

[0076] like Figure 3BAs seen in the diagram, each field of view 310 overlaps with other fields of view 310. For example, field of view 310A overlaps with fields of view 310B, 310E, and 310F. As another example, field of view 310F overlaps with fields of view 310A, 310B, 310C, 310E, 310G, 310I, 310J, and 310K. Just like fields of view 310A and 310F, other fields of view 310 (e.g., fields of view 310B, 310C, 310D, 310E, 310G, 310H, 310I, 310J, 310K, and 310L) also overlap with adjacent fields of view 310. The shadows in the overlapping region are combinations of shadows in the individual fields of view that form the overlapping region. For example, the overlapping region formed by fields of view 310A and 310B includes diagonal lines extending in opposite directions. As another example, the overlapping region formed by fields of view 310A, 310B, 310E and 310F includes diagonal lines, horizontal lines and vertical lines extending in opposite directions.

[0077] The overlapping field of view 310 can be a result of cameras 305 in camera array 300 being close to each other. Generally, with the overlapping field of view 310, certain portions of the ground space can be captured by multiple cameras 305 of camera array 300. Therefore, even if some cameras 305 are offline, the remaining cameras 305 can still provide sufficient coverage for the tracking system 132 to operate. Furthermore, the overlapping field of view 310 can improve the tracking of people's (e.g., shoppers 105) positions as they move through the space.

[0078] Figure 3C The illustration shows an example camera grid 315. (Example:) Figure 3C As seen in the diagram, camera grid 315 includes the number of rows and columns corresponding to the number of rows and columns in camera array 300. Each box in camera grid 315 represents a camera 305 of camera array 300. Camera grid 315 illustrates how cameras 305 of camera array 300 are communicatively coupled to camera client 220. Figure 3A As in the previous example, camera grid 315 shows cameras 305A, 305D, 305G, and 305J communicatively coupled to camera client 1 220A. Camera grid 315 also shows cameras 305B, 305E, 305H, and 305K communicatively coupled to camera client 2 220B. Camera grid 315 also shows cameras 305C, 305F, 305I, and 305L communicatively coupled to camera client 3 220C.

[0079] Camera grid 315 shows cameras 305 communicatively coupled to camera client 220 according to certain rules. For example, a camera 305 communicatively coupled to a particular camera client 220 is not directly adjacent to another camera 305 communicatively coupled to the same camera client 220 in the same row or column of camera grid 315. Figure 3CAs seen in the diagram, for example, camera 305A is directly adjacent to cameras 305B and 305E in the same row or column of camera grid 315. Camera 305A is communicatively coupled to camera client 1220A, while cameras 305B and 305E are communicatively coupled to camera client 2220B. Camera 305F is directly adjacent to cameras 305B, 305E, 305G, and 305J in the same row or column of camera grid 315. Camera 305F is communicatively coupled to camera client 3220C, while cameras 305B, 305E, 305G, and 305J are communicatively coupled to either camera client 1220A or camera client 2220B.

[0080] As another example, a camera 305 that is communicatively coupled to a particular camera client 220 is diagonally opposite another camera 305 that is communicatively coupled to the same camera client 220 in the camera grid 315. Figure 3C As seen in the diagram, cameras 305D, 305G, and 305J are diagonally positioned relative to each other and are communicatively coupled to camera client 1 220A. Cameras 305C, 305F, and 305I are diagonally positioned relative to each other and are all communicatively coupled to camera client 3 220C.

[0081] The result of arranging the cameras 305 in this manner is that each camera client 220 is communicatively coupled to at least one camera 305 within a portion of the camera grid 315. For example... Figure 3C As seen in the example, each of camera client 1 220A, camera client 2220B, and camera client 3 220C is communicatively coupled to at least one camera in any 2×2 section of the camera grid 315. Therefore, even if one camera client 220 is offline, the other cameras in the 2×2 section can still provide sufficient coverage of that 2×2 section to allow the tracking system 132 to operate. This improves the resilience of the tracking system 132.

[0082] While the previous example used a certain number of cameras 305 and a certain number of camera clients 220, the tracking system 132 can use any suitable number of cameras 305 and any suitable number of camera clients 220 to provide the desired level of overlap, scalability, and resilience. Figure 3D An example camera array 300 including an additional camera 305 is shown. Figure 3D Examples also include additional camera clients 220: camera client 1 220A to camera client N 220D. Cameras 305 in camera array 300 can be configured according to... Figures 3A to 3C The same rules or principles described herein are used to communicate with camera client 220.

[0083] Figure 3EThis illustrates how camera 305 can be communicatively coupled to camera client 220. For example... Figure 3E As shown, the camera grid 315 comprises several rows and several columns. Across a row, cameras 305 are communicatively coupled to camera clients 220 in a sequential manner. After a camera 305 is communicatively coupled to camera client N 220d, this sequence repeats until the end of the row is reached. Similarly, cameras 305 in a column are sequentially coupled to camera clients 220. After a camera 305 is communicatively coupled to camera client N 220d, this pattern repeats.

[0084] like Figure 3D and 3E As shown, the tracking system 132 can be expanded to include any number of cameras 305 and any number of camera clients 220. Generally, cameras 305 communicatively coupled to a particular camera client 220 are not directly adjacent to another camera 305 communicatively coupled to the same camera client 220 in the same row or column of the camera grid 315. Furthermore, cameras 305 along the diagonal of the camera grid 315 are communicatively coupled to the same camera client 220. Additionally, each camera client 220 is communicatively coupled to at least one camera 305 in a portion of the camera grid 315. The size of this portion can depend on the number of camera clients 220 in the tracking system 132. Generally, the size of this portion is less than the number of camera clients 220 in the tracking system 132. Therefore, in Figure 3D and 3E In the example, the size of this part is (N-1) x (N-1).

[0085] 2. Initialization

[0086] Figure 3F The initialization of the camera subsystem 202 is shown. (As follows) Figure 3F As seen in the diagram, camera subsystem 202 includes camera array 300, camera client 1 220A, camera client 2 220B, camera client 3 220C, and camera server 225. Camera subsystem 202 may include any suitable number of camera array 300, camera client 220, and camera server 225. Generally, during initialization, cameras 305 of camera array 300 are started and begin sending video 302 to camera client 220. Furthermore, camera client 220 and camera server 225 synchronize their internal clocks 304. After cameras 305 in camera array 300 have been started and the internal clocks 304 have been synchronized, camera client 220 can begin processing video 302 and transmitting information to camera server 225 to perform tracking operations of camera subsystem 202.

[0087] During initialization, cameras 305 of camera array 300 can be powered on and execute a startup sequence. For example, components of camera 305 can be powered on and / or warmed up. Cameras 305 can then begin capturing video segments and transmitting video 302 to their respective camera clients 220. Cameras 305 of camera array 300 can take varying amounts of time to initialize. For example, some cameras 305 may take less or more time to initialize compared to other cameras 305 of camera array 300. Because cameras 305 of camera array 300 do not wait for other cameras 305 of camera array 300 to complete initialization before sending video 302 to camera client 220, cameras 305 of camera array 300 can each begin sending video 302 to camera client 220 at different times. Therefore, video 302, and particularly frames of video 302, may be out of sync with frames of other videos 302. In other words, frames of these videos 302 are not captured and transmitted simultaneously or at the same time by their respective cameras 305. Therefore, the frames of these videos 302 arrive at the camera client 220 at different times or at the same time.

[0088] During initialization, camera client 220 and camera server 225 are powered on and / or execute a startup sequence. After startup, camera client 220 and camera server 225 synchronize their internal clocks 304. Figure 3F In the example, camera client 1 220A has an internal clock 1 304A. Camera client 2 220B has an internal clock 2 304B. Camera client 3 220C has an internal clock 3 304C. Camera server 225 has an internal clock 4 304D. Camera client 220 and camera server 225 can synchronize their internal clocks 304 in any suitable manner. For example, camera client 220 and camera server 225 can use a synchronization protocol such as Network Time Protocol (NTP) or Precision Time Protocol (PTP) to synchronize their internal clocks 304. Although a synchronization protocol can be used to synchronize the internal clocks 304 of camera client 220 and camera server 225, this does not mean that these internal clocks 304 show exactly the same time or are perfectly synchronized with each other. Therefore, there may still be some degree of asynchrony between camera client 220 and camera server 225.

[0089] Camera client 220 can track which cameras 305 in camera array 300 have completed initialization by tracking which cameras 305 have transmitted video 302 to camera client 220. When camera client 220 determines that each camera 305 in camera array 300 has started transmitting video 302 to camera client 220, camera client 220 can determine that camera array 300 has completed initialization. In response to that determination, camera client 220 can begin processing frames of video 302 and transmit information from those frames to camera server 225. Camera server 225 can then analyze the information from camera client 220 to determine the physical locations of people and / or objects in the space.

[0090] 3. Camera App

[0091] Figure 3G-3I The operation of camera client 220 in camera subsystem 202 is illustrated. Generally, camera client 320 processes video 302 from camera 305. Camera client 320 can identify people or objects within frames 320 of these video 302 and determine the coordinates 322 of these people or objects. Camera client 320 can also generate timestamps 324 indicating when camera client 320 received a specific frame 320 (e.g., by using an internal clock 304). Camera client 320 transmits these timestamps 324 and coordinates 322 to camera server 225 for further processing.

[0092] Figure 3G-3I The operation of camera client 210 is illustrated when an event unfolds in store 100. During this event, for example, a first shopper 105 (e.g., a man) removes item 130 from a shelf in store 100 and a second shopper 105 (e.g., a woman) moves toward the shelf. Camera client 320 analyzes frames 320 of video 302 to determine the coordinates 322 of the man and woman in frame 320.

[0093] like Figure 3G As seen in the image, the man is standing near the shelf, while the woman is standing further away. Two cameras, 305A and 305B, are positioned above the space and capture video 302 of the man, woman, and the shelf. These cameras 305A and 305B send their video 302 to two different camera clients 220A and 220B. Camera 305A sends video 305 to camera client 220A. Camera 305B sends video 305 to camera client 220B.

[0094] Camera client 220A receives video 305 from camera 305A, and specifically frame 320A of that video 305. Camera client 220A processes frame 320A. As seen in frame 320A, the man is standing near the shelf, while the woman is standing further away from the shelf. Camera client 220A processes frame 320A to determine the defined areas 325A and 325B around the man and woman. Figure 3G In the example, defining regions 325A and 325B are rectangular regions that enclose the man and woman, respectively. Defining regions 325A and 325B approximate the positions of the man and woman in the frame. This disclosure envisions the camera client 220 determining defining regions 325 with any suitable shape and any suitable size. For example, defining regions 325 may be circular or may be irregularly shaped (e.g., to follow the outline of shopper 105 in frame 320).

[0095] Camera client 220A determines the coordinates 322 (also known as "frame coordinates") of the defined regions 325A and 325B within frames 320A and 320B. Figure 3G In the example, camera client 228 determines the coordinates 322 (x1, y1) and (x2, y2) of the delimited region 325A and the coordinates 322 (x3, y3) and (x4, y4) of the delimited region 325B. These coordinates 322 do not represent absolute coordinates in physical space, but rather coordinates within frame 320A. Camera client 220 can determine any suitable number of coordinates 322 for delimiting region 325.

[0096] The camera client 220A then generates frame data 330A containing information about frame 320A. For example... Figure 3G As seen in the image, frame data 330A includes an identifier for camera 305A (e.g., "camera=1"). Camera client 220A can also generate a timestamp 324 indicating when camera client 220A received frame 320A (e.g., using internal clock 304). Figure 3G In the example, timestamp 324 is t1. Frame data 320A also includes information about people or objects within frame 320A. Figure 3G In the example, frame data 330A includes information about object 1 and object 2. Object 1 corresponds to the man, and object 2 corresponds to the woman. Frame data 330A indicates the man's coordinates 322 (x1, y1) and (x2, y2) and his height z1. As discussed earlier, camera 305 can be a 3D camera capable of detecting the height of objects and / or people. Camera 305 may have already provided the heights of the man and woman to camera client 320. Figure 3GIn the example, camera 305A may have already detected the heights of the man and woman as z1 and z2, respectively. Frame data 330A also includes information about the woman, including coordinates 322 (x3, y3) and (x4, y4) and height z2. When frame data 330A is ready, camera client 220A can transmit frame data 330A to camera server 225.

[0097] In a corresponding manner, camera client 220B can process video 302 from camera 305B. For example... Figure 3G As seen in the image, camera client 220B receives frame 320B from camera 305B. Because camera 305B is located in a different position than camera 305A, frame 320B will show a perspective view of events in store 100 that is slightly different from frame 320A. Camera client 220B determines the defining regions 325C and 325D around the man and woman, respectively. Camera client 220B determines the frame coordinates 322 (x1, y1) and (x2, y2) of defining region 325C, and the frame coordinates 322 (x3, y3) and (x4, y4) of defining region 325D. Camera client 220B also determines and generates a timestamp 324 t2 indicating when camera client 220B received frame 320B (e.g., using internal clock 304). Camera client 220B then generates frame data 330B for frame 320B. Frame data 330B indicates that frame 320B was generated by camera 305B and received by camera client 220B at t2. Frame data 330B also indicates that a man and a woman were detected in frame 320B. The man corresponds to coordinates 322 (x1, y1) and (x2, y2) with a height of z1. The woman corresponds to coordinates 322 (x3, y3) and (x4, y4) with a height of z2. When frame data 320B is ready, camera client 220B transmits frame data 320B to camera server 225.

[0098] The coordinates 322 generated by camera clients 220A and 220B for frame data 330A and 330B can be coordinates within a specific frame 320 rather than coordinates in physical space. Furthermore, although the same subscript is used for the coordinates 322 in frame data 330A and 330B, this does not mean that these coordinates 322 are the same. Rather, because cameras 305A and 305B are in different positions, the coordinates 322 in frame 330A are likely to be different from the coordinates 322 in frame data 330B. Camera clients 220A and 220B are determining the coordinates 322 of a defined region 325 within frame 320, not in physical space. Camera clients 220A and 220B determine these local coordinates 322 independently of each other. The subscript indicates the sequence of coordinates 322 generated by the respective camera client 220. For example, (x1, y1) indicates the first coordinate 322 generated by camera client 220A and the first coordinate 322 generated by camera client 220B, which can be different values.

[0099] exist Figure 3H In shop 100, the events have already taken place. The man is still standing next to the shelf, and the woman has moved closer to it. Camera clients 220A and 220B receive additional frames 320C and 320D from cameras 305A and 305B. Camera client 220A again determines the defining regions 325C and 325D for the man and woman, respectively, and the coordinates 322 of these defining regions 325. Camera client 220A determines the coordinates 322 (x5, y5) and (x6, y6) of defining region 325C and the coordinates 322 (x7, y7) and (x8, y8) of defining region 325D. Camera client 220A also generates a timestamp 324 indicating that frame 320C was received at time t3. Camera client 220A generates frame data 330C indicating that frame 320C was generated by camera 305A and received by camera client 220A at t3. Frame data 330C also indicates that the man corresponds to coordinates 322 (x5, y5) and (x6, y6) in frame 320C and has a height of z3, and the woman corresponds to coordinates 322 (x7, y7) and (x8, y8) in frame 320C and has a height of z4.

[0100] Similarly, camera client 220B receives frame 320D from camera 305B. Camera client 220B determines defining regions 325E and 325F for the man and woman respectively. Camera client 220B then determines the coordinates 322(x5, y5) and (x6, y6) of defining region 325E and the coordinates 322(x7, y7) and (x8, y8) of defining region 325F. Camera client 220B generates a timestamp 324 indicating that frame 320D was received at time t4. Camera client 220B generates frame data 330D indicating that frame 320D was generated by camera 305B and received by camera client 220B at t4. Frame data 330D indicates that the man corresponds to coordinates 322(x5, y5) and (x6, y6) in frame 320D and has a height z3. Frame data 330D also indicates that the woman corresponds to coordinates 322 (x7, y7) and (x8, y8) within frame 320D and has a height z4. When frame data 330C and 330D are ready, camera clients 220A and 220B transmit frame data 330C and 330D to camera server 225.

[0101] exist Figure 3I In the middle, the events in store 100 further develop, and the man has removed item 130 from the shelf. Camera client 220A receives frame 320E from camera 305A. Camera client 220A determines the boundary areas 325G and 325H around the man and woman, respectively. Camera client 220A determines the coordinates 322(x9, y9) and (x...) of the boundary area 325G. 10 y 10 ) and the coordinates 322 (x) of the region 325H. 11 y 11 ) and (x 12 y 12 Camera client 220A generates a timestamp 324 indicating when frame 320E was received by camera client 220A (e.g., by using internal clock 304). Camera client 220A generates frame data 330E indicating that frame 320E was generated by camera 305A and received by camera client 220A at t5. Frame data 330E indicates that the man is within frame 320E with coordinates 322(x9, y9) and (x... 10 y 10 ) corresponds to and has a height of z5. Frame data 330E also indicates that the woman in frame 320E corresponds to coordinates 322 (x 11 y 11 ) and (x 12 y 12 It corresponds to and has a height of z6.

[0102] Camera client 220B receives frame 320F from camera 305B. Camera client 220B determines the boundary regions 325I and 325J around the man and woman, respectively. Camera client 220BA determines the coordinates 322(x9, y9) and (x...) of the boundary region 325I. 10 y 10 ) and the coordinates 322 (x) of the region 325J. 11 y 11 ) and (x 12 y 12 Camera client 220B generates a timestamp 324 indicating when frame 320F was received by camera client 220B (e.g., by using internal clock 304). Camera client 220B then generates frame data 330F indicating that frame 320F was generated by camera 305B and received by camera client 220B at t6. Frame data 330F indicates the man in frame 320F with coordinates 322(x9, y9) and (x... 10 y 10 ) corresponds to and has a height of z5. Frame data 330F also indicates that the woman in frame 320F corresponds to coordinates 322 (x 11 y 11 ) and (x 12 y 12 (Corresponding to and having a height of z6.) Camera clients 220A and 220B transmit frame data 330E and 330F to camera server 225 when ready.

[0103] 4. Camera server

[0104] Figure 3J-3P The operation of camera server 225 in camera subsystem 202 is illustrated. Generally, camera server 225 receives frame data 330 (e.g., 330A-330F) from camera client 220 in camera subsystem 202. Camera server 225 synchronizes and / or assigns the frame data 330 to specific time windows 332 based on timestamp 324 in the frame data 330. Camera server 225 then processes the information assigned to the specific time windows to determine the physical locations of people and / or objects in space during those time windows 332.

[0105] exist Figure 3JIn this process, camera server 225 receives frame data 330 from camera client 220 in camera subsystem 202. Camera server 225 assigns frame data 330 to time window 332 based on timestamps 324 within the frame data 330. Using the previous example, camera server 225 can determine that timestamps 324 t1, t2, and t3 fall within a first time window 322A (e.g., between times T0 and T1) and timestamps 324 t4, t5, and t6 fall within subsequent time windows 332B (e.g., between times T1 and T2). Therefore, camera server 225 assigns frame data 330 of frames 320A, 320B, and 320C to time window 1 332A and assigns frame data 330 of frames 320D, 320E, and 320F to time window 2 332B.

[0106] By assigning frame data 330 to time window 332, camera server 225 can handle asynchrony occurring between camera 305, camera client 220, and camera server 225 within camera subsystem 202. The duration of time window 332 can be set longer than the expected asynchrony to mitigate its effects. For example, if camera 305 and camera client 220 are expected to be asynchronous for a few milliseconds, time window 332 can be set to last 100 milliseconds to counteract the asynchrony. In this way, when camera subsystem 202 is expanded to handle more space by including more cameras 305 and camera clients 220, camera server 225 can mitigate the effects of asynchrony. Figure 3J In the example, camera server 225 sets the duration of time window 1 332A between T0 and T1 and sets the duration of time window 2 332B between T1 and T2. Camera server 225 can set the duration of time window 332 to any suitable amount to mitigate the effects of desynchronization. In some embodiments, T0 may be the time when camera 305 in camera subsystem 202 has completed initialization.

[0107] Figure 3K An embodiment is shown where camera server 225 uses cursor 335 to assign frame data 330 to time window 332. Each cursor 335 may correspond to a specific camera client 220 in camera subsystem 202. Figure 3KIn the example, cursor 335A corresponds to camera client 1 220A, cursor 335B corresponds to camera client 3 220C, and cursor 335C corresponds to camera client 2 220B. Each cursor 335 points to a specific time window 332. When frame data 330 is received from camera client 220, that frame data 330 is generally assigned to the time window 332 pointed to by the cursor 335 of that camera client 220. For example, if frame data 330 is received from camera client 1 220A, then that frame data 330 is generally assigned to time window 1 332A because cursor 335A points to time window 1 332A.

[0108] When frame data 330 is received from camera client 220 corresponding to cursor 335, camera server 225 can determine whether to advance cursor 335A. If that frame data 330 has a timestamp 324 belonging to a subsequent time window 332, then camera server 225 can advance cursor 335 to that time window 332, thereby instructing camera server 225 not to expect to receive any more frame data 330 from the camera client 220 belonging to the previous time window 332. In this way, camera server 225 can quickly and efficiently assign frame data 330 to time windows 332 without checking each time window 332 upon receiving frame data 330. For example, if camera client 220B is faster in sending information than camera client 1220A and camera client 320C, then cursor 335C may be far ahead of cursors 335A and 335B. When camera server 225 receives frame data 330 from camera client 2220B, camera server 225 does not need to check each time window 332 starting from time window 1 332A to determine which time window 332 the frame data 330 should be assigned to. Instead, camera server 225 can start at the time window 332 pointed to by cursor 335C. In other words, camera server 225 does not need to first check whether the timestamp 324 in the frame data 330 from camera client 2220B indicates a time falling within time window 1 332A and then check whether that time falls within time window 2 332B. Instead, camera server 225 can first check whether that time falls within time window 3 332C and ignore checking whether that time falls within time window 1 332A and time window 2 332B. Therefore, frame data 330 is quickly and efficiently assigned to the correct time window 332.

[0109] Figure 3LThe illustration shows a camera server 225 removing frame data 330 that has been assigned to a specific time window 332 for processing. Generally, the camera server 225 can determine that the frame data 330 assigned to the specific time window 332 is ready for processing. In response to that determination, the camera server 225 can move the frame data 330 from the specific time window 332 to a task queue 336. The information in the task queue 336 is then processed to determine the physical location of a person or object in space during the specific time window 332.

[0110] Camera server 225 determines that frame data 330 assigned to a specific time window 332 is ready to be processed in any suitable manner. For example, when a specific time window 332 has frame data 330 from frames 320 of a sufficient number of cameras 305, camera server 225 can determine that time window 332 is ready for processing. Camera server 225 can use a threshold 338 to make this determination. When a specific time window 332 has been assigned frame data 330 from frames 320 of more than the threshold 338, camera server 225 can determine that time window 332 is ready for processing and move the information for that time window 332 to the task queue 336. For example, suppose threshold 338 indicates that frame data 330 from frames 320 of ten cameras 305 in an array 300 of twelve cameras 305 needs to be received before time window 332 is ready for processing. If time window 332 contains frame data 330 from frames 320 of only eight cameras 305, then camera server 225 determines that time window 332 is not yet ready for processing. Therefore, time window 332 waits for frame data 330 to be assigned from additional cameras 305 for frames 320. When time window 332 receives frame data 330 of frames 320 from ten or more cameras 305, camera server 225 determines that time window 332 is ready for processing and moves the frame data 330 in time window 332 to task queue 336.

[0111] When a subsequent time window 332 has received frame data 330 for frame 320 from multiple cameras 305 exceeding a threshold 338, the camera server 225 can still determine that a specific time window 332 is ready for processing. Using the previous example, even if time window 1 332A has been assigned frame data 330 for frame 320 from eight cameras, the camera server 225 can still determine that time window 1 332A is ready for processing when time window 2 332B has been assigned frame data 330 for frame 320 from ten or more cameras 305 (e.g., from each camera 305 in the camera array 300). In this scenario, the camera server 225 can assume that no additional frame data 330 will be assigned to time window 1 332A because frame data 330 for frame 320 from a sufficient number of cameras 305 has already been assigned to the subsequent time window 2332B. In response, camera server 225 moves frame data 330 from time window 1 332A to task queue 336.

[0112] Camera server 225 can also determine when a specific time window 332 is ready for processing if it has been waiting for a specific period of time. For example, if an error or glitch occurs in the system and frames 320 from multiple cameras 305 are not sent or are lost, then time window 332 cannot receive enough frame data 330 for frames 320 from enough cameras 305. Therefore, processing for that time window 332 will stop or be delayed. Camera server 225 can use timeouts or expirations, after which time window 332 does not wait for processing. Therefore, when time window 332 has not been processed for a period of time after timeout or expiration, camera server 225 can still send frame data 330 from that time window 332 to task queue 336. Using the previous example, assume the timeout is 200 milliseconds. If time window 1 332A has been stuck for more than 200 milliseconds with frame data 330 from frames 320 of the eight cameras 305, then camera server 225 can determine that time window 1 332A has waited long enough for additional frame data 330 and is ready for processing. In response, camera server 225 moves the frame data 330 in time window 1 332A to task queue 336.

[0113] In some embodiments, when time window 332 times out or ages out, camera server 225 can adjust threshold 338 to make future time windows 332 less likely to time out or age out. For example, camera server 225 can lower threshold 338 when time window 332 times out or ages out. Similarly, when subsequent time windows 332 do not time out or age out, camera server 225 can increase threshold 338. Camera server 225 can adjust threshold 338 based on the number of cameras 305 that have sent information for a particular time window 332. For example, if a particular time window 332 times out or ages out when it has frame data 330 for frame 320 from eight cameras 305, and threshold 338 is ten cameras 305, then camera server 225 can reduce threshold 338 to a value close to eight cameras. Therefore, that time window 332 can then have frame data 330 for frame 320 from a sufficient number of cameras 305 and be moved to task queue 336. When the subsequent time window 332 has not timed out because it has received frame data 330 for frame 320 from the nine cameras 305, the camera server 225 can adjust the threshold 338 toward the nine cameras 305. In this way, the camera server 225 can dynamically adjust the threshold 338 to prevent delays in the camera subsystem 202 caused by bugs, errors and / or latency.

[0114] In some embodiments, camera server 225 processes time windows 332 sequentially. In other words, camera server 225 does not process subsequent time windows 332 until a previous time window 332 is ready for processing. Figure 3L In the example, camera server 225 may not place time window 2 332B into task queue 336 until time window 1 332A has already been placed into task queue 336. In this way, the progress of events in store 100 is evaluated sequentially (e.g., as events unfold), which allows for accurate tracking of people's locations in store 100. If time window 332 is not evaluated sequentially, then tracking system 132 may assume that events in store 100 are occurring in a different and incorrect order.

[0115] Figure 3M The diagram illustrates task queue 336 of camera server 225. (For example...) Figure 3MAs shown, task queue 336 includes frame data 330 from two time windows 332. At the beginning of task queue 336 are frame data 330 for frames 320A, 320B, and 320C. Following in task queue 336 are frame data 330 for frames 320D, 320E, and 320F. Camera server 225 can process entries in task queue 336 sequentially. Therefore, camera server 225 can first process the first entry in task queue 336 and then process the frame data 330 for frames 320A, 320B, and 320C. Camera server 225 processes the entries in task queue 336 and then moves that entry to the results queue.

[0116] To process entries in task queue 336, camera server 225 can combine or cluster the coordinates 322 of the same object detected by the same camera 320 to calculate the combined coordinates 332 of that object. As a result of this processing, each time window 332 should include only one set of coordinates 322 for each object per camera 305. After this processing, the combined coordinates 322 are placed in the results queue. Figure 3N The diagram illustrates the result queue 340 of camera server 225. (For example...) Figure 3N As shown, the result queue 340 includes combined coordinates 332 for two time windows 332.

[0117] As an example, camera server 225 first processes the first entry in task queue 336, which includes frame data 330 for frames 320A, 320B, and 320C. Frames 320A and 320C originate from the same camera 320A. Therefore, camera server 225 can use the frame data 330A and 330C for frames 320A and 320C to calculate the combined coordinates 322 of a person or object detected by camera 320A. Figure 3N As seen in the image, camera server 225 has determined the combined coordinates 322 (x, y) of object 1 detected by camera 1 305A. 13 y 13 ) and (x 14 y 14 ) and combined height z7, and combined coordinates 322 (x) of object 2 detected by camera 1 305A. 15 y 15 ) and (x 16 y 16The combined coordinates 322 and combined height z8 are the combined coordinates 322 and combined height z8 of the man and woman in the video frame 302 received by camera 305A during the first time window 332A. Similarly, camera server 225 can determine the combined coordinates 322 and combined height z8 of the object detected by camera 2 305B during the first time window 332A. For example, camera server 225 can use frame data 330B for frame 320B (and frame data 330 received by camera 2 305B during the first time window 332A for any other frame 320) to determine the combined coordinates 322 (x, y, z8) of object 1 detected by camera 2 305B. 13 y 13 ) and (x 14 y 14 ) and the combined height z7 and the combined coordinates 322 (x) of object 2 detected by camera 2 305B. 15 y 15 ) and (x 16 y 16 And the combined height z8. The camera server 225 can determine the combined coordinates 322 of each object detected by the camera 305 in the first time window 332A in this way.

[0118] Camera server 225 then determines the combined coordinates 322 of the object detected by camera 305 during the second time window 332B in a similar manner. For example, camera server 225 can use frame data 330E for frame 320E (and frame data 330 received by camera 1305A during the second time window 332B for any other frame 320) to determine the combined coordinates 322 (x, y, x) of object 1 detected by camera 1305A. 17 y 17 ) and (x 18 y 18 ) and the combined height z9 and the combined coordinates 322 (x) of object 2 detected by camera 1 305A. 19 y 19 ) and (x 20 y 20 ) and combined height z 10 Camera server 225 can also use frame data 330D and 330F for frames 320D and 320F to determine the combined coordinates 322 (x, y, y) of object 1 detected by camera 2 305B. 17 y 17 ) and (x 18 y 18 ) and the combined height z9 and the object 2 detected by camera 2 305B and the combined coordinates 322 (x 19 y 19) and (x 20 y 20 ) and combined height z 10 .

[0119] Camera server 225 calculates the combined coordinates 322 and combined height in any suitable manner. For example, camera server 225 can calculate the combined coordinates 322 and combined height by taking the average of the coordinates 322 and height of a specific object detected by the same camera 305 within a specific time window 332. Figure 3M In the example, camera server 225 can calculate the combined coordinates 322(x1, y1) and (x5, y5) for camera 1 305A by taking the average of the coordinates 322(x1, y1) and (x5, y5) from frame data 330A and 330C. 13 y 13 Similarly, camera server 225 can determine the combined coordinates 322(x2, y2) and (x6, y6) for camera 1 305A by taking the average of the coordinates 322(x2, y2) and (x6, y6) from frame data 330A and 330C. 14 y 14 Camera server 225 can determine the combined height z7 of camera 1 305A by taking the average of the heights z1 and z3 from frame data 330A and 330C. Similarly, camera server 225 can determine the combined coordinates 322(x5, y5) and (x9, y9) of camera 2 305B by taking the average of the coordinates 322(x5, y5) and (x9, y9) from frame data 330D and 330F. 17 y 17 Similarly, camera server 225 can retrieve the coordinates 322(x6, y6) and (x...) from frame data 330D and 330F. 10 y 10 The average value of ) is used to determine the combined coordinates of camera 2305B 322 (x 18 y 18 Camera server 225 can determine the combined height z9 of camera 2 305B by taking the average of the heights z3 and z5 from frame data 330D and 330F. Camera server 225 takes these averages because these are the coordinates 322 and height of the same object determined by the same camera 305 during the same time window 332.

[0120] Camera server 225 can follow a similar process to determine or calculate the combined coordinates of object 2 detected by cameras 1 305A and 2 305B. Camera server 225 can calculate the combined coordinates 322(x, y3) of camera 1 305A by taking the average of the coordinates 322(x, y3) and (x, y7) from frame data 330A and 330C. 15 y15 Similarly, camera server 225 can determine the combined coordinates 322(x4, y4) and (x8, y8) for camera 1305A by taking the average of the coordinates 322(x4, y4) and (x8, y8) from frame data 330A and 330C. 16 y 16 Camera server 225 can determine the combined height z8 for camera 1 305A by taking the average of the heights z2 and z4 from frame data 330A and 330C. Similarly, camera server 225 can determine the combined height z8 for camera 1 305A by taking the coordinates 322(x7, y7) and (x...) from frame data 330D and 330F. 11 y 11 The average value of ) is used to determine the combined coordinates 322 (x) for camera 2 305B. 19 y 19 Similarly, camera server 225 can retrieve the coordinates 322(x8, y8) and (x...) from frame data 330D and 330F. 12 y 12 The average value of ) is used to determine the combined coordinates 322 (x) for camera 2 305B. 20 y 20 Camera server 225 can determine the combined height z for camera 2305B by taking the average of heights z4 and z6 from frame data 330D and 330F. 10 .

[0121] Camera server 225 may use any other suitable computation to calculate the combined coordinates and combined height. For example, camera server 225 may take the median coordinates 322 and height for objects detected by the same camera 305 during a time window 332. Camera server 225 may also use clustering processes to calculate the combined coordinates 322 and combined height. For example, camera server 225 may use K-means clustering, density-based spatial clustering (DBSCAN) for noisy applications, k-medoids, Gaussian mixture models, and hierarchical clustering to calculate the combined coordinates 322 and combined height.

[0122] After camera server 225 has calculated the combined coordinates 322 and combined height, camera server 225 has determined the coordinates 322 of each object detected by each camera 305 during time window 332. However, camera server 225 may perform additional processing to determine whether objects detected by different cameras 305 are the same object. Camera server 225 may use linking and homography to determine which objects detected by which cameras 305 are actually the same person or object in space. Camera server 225 may then take the combined coordinates 322 for those objects from different cameras 305 and use homography to determine the physical location of that person or object in physical space during time window 332. An embodiment of this process is described in U.S. Patent Application No. ___ (Attorney General's File No. 090278.0180), entitled "Topview Object Tracking Using a Sensor Array," the contents of which are incorporated herein by reference in their entirety. In this way, camera server 225 determines the physical location of a person and / or object in space during a specific time window 332.

[0123] In a particular embodiment, camera client 220 may also use the same time window 332 as camera server 225 to batch transmit frame data 330 to camera server 225. For example... Figure 3O As seen in the diagram, camera client 220 assigns frame date 330 to time window 332 based on timestamp 324 within frame data 330. Camera client 220 can determine a specific time window 332 is ready to be transmitted to camera server 225 in a similar manner to how camera server 225 determines that time window 332 is ready for processing. When camera client 220 determines that a specific time window 332 is ready (e.g., when each camera 305 to which communication is coupled has already transmitted frames within that time window 332), camera client 220 transmits frame data 330 to camera server 225 as a batch assigned to that time window 332. In this way, camera server 225 can assign frame data 330 to time window 332 even faster and more efficiently because camera server 225 receives frame data 330 as a batch for time window 332 from camera client 220.

[0124] In some embodiments, even if the camera server 225 and the camera client 220 are out of sync, the camera server 225 can respond to the asynchrony by adjusting the timestamp 324 in the frame data 330 (e.g., by the asynchrony of the internal clock 302, by the latency difference between the camera client 220 and the camera server 225, by processing the speed difference between the camera clients 220, etc.). Figure 3PThe image shows camera server 225 adjusting timestamp 324. As discussed earlier, frame data 330 includes timestamp 324 generated by camera client 220, indicating when camera client 220 received frame 320. Figure 3P In the example, frame data 330 indicates that camera client 220 received frame 320 at time t1. If camera client 220 and camera server 225 are out of sync, then timestamp 324 t1 is relatively meaningless to camera server 225, because camera server 225 cannot guarantee that timestamps 324 from different camera clients 220 are accurate relative to each other. Therefore, it is difficult, if not impossible, to accurately analyze frame data 330 from different and / or multiple camera clients 220.

[0125] Camera server 225 can adjust the timestamp 324 of a specific camera 305 to cope with asynchrony. Generally, camera server 225 determines the latency of that camera 305 by tracking the latency of the previous frame 320 from each camera 305. Camera server 225 then adjusts the timestamp 324 of the frame data 330 for frame 320 from that camera 305 based on the determined latency. Figure 3P In the example, camera server 225 determines the timestamp 324 (labeled t) indicated in the frame data 330 for each frame 320(x) from camera 1. x The time (marked as T) when the time camera server 225 receives frame data 330 x The time difference between (marked as Δ) x The camera server 225 determines the latency of camera 1 305A by the time difference (Δ) between multiple previous frames 320. x The average delay (denoted as Δ) is calculated by averaging. Figure 3P In the example, camera server 225 averages the time difference of the previous thirty frames 320 to determine the average latency. The camera server then adds the average latency (Δ) to the timestamp 324 used for frame data 330 to adjust the timestamp 324 to cope with asynchrony. In this way, even if camera client 220 and camera server 225 are not synchronized (e.g., according to a clock synchronization protocol), camera server 225 and tracking system 132 can still function normally.

[0126] 5. Example Method

[0127] Figure 3Q and 3R This is a flowchart illustrating an example method 342 for operating the camera subsystem 202. In a particular embodiment, various components of the camera subsystem 202 perform the steps of method 342. Generally, by performing method 342, the camera subsystem 202 determines the physical location of a person or object in space.

[0128] like Figure 3Q As seen in the diagram, method 342A begins with cameras 305A and 305B generating frames 320A and 320D respectively and transmitting them to camera clients 220A and 220B. Camera clients 220A and 220B then determine the coordinates 322 of two people detected in frames 320A and 320B. These coordinates can define a bounded area 325 around these people.

[0129] Camera 305A then generates frame 320B and transmits it to camera client 220A. Camera client 220A generates coordinates 322 for the two people shown in frame 320B. During that process, camera 305B generates frame 320E and transmits it to camera client 220B. Camera client 220B then determines the coordinates 322 of the two people detected in frame 320E. Camera 305A then generates frame 320C and transmits it to camera client 220A. Camera client 220A determines the coordinates 322 of the two people detected in frame 320C. Importantly, Figure 3Q The diagram shows that frames from cameras 305A and 305B can be generated and transmitted at different times or synchronously. Furthermore, the coordinates of a person detected in frame 320 may not be generated simultaneously or synchronously in camera clients 220A and 220B.

[0130] Figure 3R It shows from Figure 3Q Method 342A continues with method 342B. For example... Figure 3RAs seen in the diagram, camera client 220A generates frame data 330 based on the coordinates 322 of the two people detected in frame 320A. Similarly, camera client 220B generates frame data 330 using the coordinates 322 of the two people detected in frame 320D. Camera clients 220A and 220B transmit frame data 330 to camera server 225. Camera client 220A generates additional frame data 330 using the coordinates 322 of the two people detected in frame 320B. Camera client 220A then transmits that frame data 330 to camera server 225. Camera server 225 can assign frame data 330 to time window 332. Camera server 225 can determine in step 344 that time window 332 is ready for processing, and in response, in step 346, put the frame data 330 in that time window 332 into task queue 336. Camera server 225 can then combine or cluster coordinates 322 within that time window 322 to determine combined coordinates 322 in step 348. For example, camera server 225 can average the coordinates 322 within that time window to determine combined coordinates 322 of people detected by different cameras 305 during that time window 332. In step 350, camera server 225 can then map the people detected by different cameras 305 to people in space. Camera server 225 can then determine the location of the people during that time window 332 in step 352. Camera server 225 transmits these determined locations to central server 240.

[0131] Can be Figure 3Q and 3R The method 342 described herein may be modified, added to, or omitted. Method 342 may include more, fewer, or other steps. For example, the steps may be performed in parallel or in any suitable order. Although discussed as a specific component of the execution steps of the camera subsystem 202, any suitable component of the camera subsystem 202 may perform one or more steps of the method.

[0132] 6. Other characteristics

[0133] In a particular embodiment, the camera subsystem 202 may include a second camera array that operates in series with the first camera array 300 of the camera subsystem 202. Figure 3S An embodiment including two camera arrays 300 and 354 is shown. Camera array 300 includes camera 305M. Camera array 354 includes camera 305N. Camera 305N operates in the same manner as camera 305M and can be used for... Figure 3A-3R The same techniques described are used to determine the location of objects and / or people in space.

[0134] Each camera 305N is positioned slightly offset from camera 305M of camera array 300. In this way, camera 305M captures video similar to that captured by camera 305N. In some embodiments, camera 305M may use a different version of software, or a different version of software may be used relative to camera 305N to process video from camera 305M. In this way, updated software can be run for camera 305N to test the effectiveness of that software. Testing that software does not interrupt the operation of camera subsystem 202, as camera 305M may still be using the previous software, which also serves as a baseline for comparison with the operation of the new software running on camera 305N. For example, the accuracy of position tracking provided by the new software can be determined and compared with the accuracy provided by the old software. If the new software is less accurate than the old software, then the old software should continue to be used.

[0135] In some embodiments, if camera server 225 cannot determine the location of a person based on frame data 330 from camera client 220, then camera server 225 can retrieve video clips from camera client 220 or shared memory. Figure 3T A camera server 225 is shown retrieving video 302 from camera client 220 and / or shared memory 356. Generally, camera client 220 stores video received from the camera locally or in shared memory 356. If camera server 225 cannot determine the location of a person based on frame data 330, then that video 302 becomes available to camera server 225. Camera server 225 can analyze video 302 to determine the location of a person in space. Compared to camera client 220, camera server 225 can perform better and more accurate analysis of raw video footage; therefore, camera server 225 can generate more accurate frame data 330 than camera client 220. In some embodiments, camera server 225 may have frame data 330 from one camera client 220 that conflicts with or is misaligned with frame data 330 from another camera client 220. Camera server 225 can retrieve raw video footage to determine which frame data 330 should be accepted and used.

[0136] exist Figure 3TIn the example, camera client 220A stores video 302A in local or shared memory 356. Camera client 220B stores video 302B in local or shared memory 356. When camera server 225 cannot determine the location of a person based on frame data 330, camera server 225 sends a request 358 to camera client 220A and / or shared memory 356. In response, camera client 220A and / or shared memory 356 send video 302A to camera server 225. Camera server 225 can then analyze video 302A to determine the location of the person in the space.

[0137] III. LiDAR Subsystem

[0138] Some embodiments of the tracking system 132 include a LiDAR subsystem 204. Figures 4A-4D The LiDAR subsystem 204 and its operation within the tracking system 132 are illustrated. Generally, the LiDAR subsystem 204 uses LiDAR sensors and a LiDAR server to track the positions of people and / or objects in physical space. The LiDAR subsystem 204 can be used alone or in conjunction with other subsystems (e.g., camera subsystem 202) to track the positions of people and / or objects in space.

[0139] Figure 4A An example LiDAR subsystem 204 is shown. (e.g.) Figure 4A As seen in the diagram, the LiDAR subsystem 204 includes a LiDAR array 400 and a LiDAR server 230. Generally, the LiDAR sensors 405 in the LiDAR array 400 detect the presence of people and / or objects in space and determine the coordinates of these people and / or objects. The LiDAR server 230 processes these coordinates to determine the physical location of people and / or objects in space.

[0140] LiDAR array 400 is an array of LiDAR sensors 405. LiDAR array 400 can be positioned above physical space to detect the presence and location of people and / or objects within that space. Figure 4A In the example, LiDAR array 400 is a 3×4 array of LiDAR sensors 405. LiDAR array 400 includes any suitable number of LiDAR sensors 405 arranged in an array of any suitable size.

[0141] Each LiDAR sensor 405 detects the presence of people and / or objects within a portion of physical space. Generally, the LiDAR sensor 405 emits light pulses into space. When these light pulses come into contact with people and / or objects in the space, they are reflected back towards the LiDAR sensor 405. The LiDAR sensor 405 tracks characteristics of the reflected light pulses, such as the return time and wavelength of the returning light pulses, to detect the presence of people and / or objects within the physical space. The LiDAR sensor 405 can also determine the coordinates of the detected people and / or objects. The LiDAR sensor 405 transmits the coordinates of the detected people and / or objects to the LiDAR server 230.

[0142] LiDAR sensor 405 can be communicatively coupled to LiDAR server 230 in any suitable manner. For example, LiDAR sensor 405 can be hardwired to LiDAR server 230. As another example, LiDAR sensor 405 can be wirelessly coupled to LiDAR server 230 using any suitable wireless standard (e.g., WiFi). LiDAR sensor 405 transmits the coordinates of detected people and / or objects to LiDAR server 230 via a communication medium.

[0143] Figure 4B A LiDAR sensor 405 is shown transmitting coordinates 410 to a LiDAR server 230. Generally, the LiDAR sensor 405 analyzes the characteristics of reflected light pulses to determine the coordinates 410 of people and / or objects in space. The LiDAR sensor 405 transmits these coordinates 410 to the LiDAR server 230 for further processing. Figure 4B In the example, LiDAR sensor 405 detects the coordinates 410 of at least two people and / or objects in space. The coordinates 410 of these people and / or objects are (x1, y1) and (x2, y2). LiDAR sensor 405 transmits these coordinates 410 to LiDAR server 230 for further processing.

[0144] Figure 4C The diagram illustrates the general operation of the LiDAR server 230. (For example...) Figure 4C As seen in the diagram, LiDAR server 230 processes coordinates 410 received from LiDAR sensor 405. LiDAR server 230 assigns coordinates 410 to time window 332 in a similar manner to how camera server 225 assigns frame data 330 to time window 332. For example, LiDAR server 230 may assign coordinates 410 to a specific time window 332 based on the time when LiDAR server 230 receives coordinates 410 from LiDAR sensor 405.

[0145] The LiDAR server 230 can process coordinates 410 assigned to a time window 332 to determine the physical location of people and / or objects in space. Figure 4C In the example, LiDAR server 230 receives coordinates 410 of two people from two different LiDAR sensors 405. One LiDAR sensor 405 provides coordinates 410 (x1, y1) and (x2, y2) for the two people respectively. The other LiDAR sensor 405 provides coordinates 410 (x1, y1) and (x2, y2) for the same two people respectively. Similar to camera client 220 and camera server 225, the subscripts of these coordinates 410 do not indicate that these coordinates 410 have the same value, but rather indicate that these are the first and second coordinates 410 provided by a specific LiDAR sensor 405.

[0146] LiDAR server 230 uses these coordinates 410 to determine the physical location of a person within space. Similar to camera server 225, LiDAR server 230 can determine that coordinates 410 provided by two different LiDAR sensors 405 correspond to the same person in physical space. In response, LiDAR server 230 can take these coordinates 410 and use homography to determine the person's position in physical space within a specific time window 332. Figure 4C In the example, LiDAR server 230 uses coordinates 410 to determine the location of the first person during time window 332 as (x3, y3). LiDAR server 230 also uses coordinates 410 to determine the physical location of the second person during time window 332 as (x4, y4). LiDAR server 230 transmits these physical locations to central server 240 for further processing.

[0147] Figure 4D A method 415 for operating the LiDAR subsystem 204 in the tracking system 132 is illustrated. Generally, the LiDAR subsystem 204 performs method 415 to determine the location of people and / or objects in physical space.

[0148] LiDAR sensor 405 determines the coordinates 410 of a detected person and transmits these coordinates 410 to LiDAR server 230. LiDAR sensor 405 can determine these coordinates 410 by emitting light pulses and analyzing the characteristics of those light pulses when they are reflected back to LiDAR sensor 405. For example, LiDAR sensor 405 can analyze the return time and / or wavelength of the reflected light pulses to determine whether a person exists in physical space and that person's coordinates 410.

[0149] In step 416, LiDAR server 230 analyzes the coordinates 410 from LiDAR sensor 405 to determine the location of a person in physical space during a first time window 332. LiDAR server 230 then transmits these locations to central server 240. LiDAR sensor 405 can subsequently determine the coordinates 410 of detected persons and transmit these coordinates 410 to LiDAR server 230. In step 418, LiDAR server 230 can again determine the locations of these persons in subsequent time windows 332 and transmit these locations to central server 240.

[0150] Similar to the camera subsystem 202, the central server 240 can use these locations to determine which person removed item 130 from the space during a specific time window 332. This will be used... Figures 6A to 6C A more detailed description of the operation of the central server 240.

[0151] Can be Figure 4D The method 415 described herein may be modified, added to, or omitted. Method 415 may include more, fewer, or other steps. For example, the steps may be performed in parallel or in any suitable order. Although discussed as a component of the LiDAR subsystem 204 that performs these steps, any suitable component of the tracing system 132 (e.g., such as the central server 240) may perform one or more steps of the method.

[0152] IV. Weight Subsystem

[0153] Tracking system 132 includes a weight subsystem 206, which includes a weight sensor 215 and a weight server 235. Generally, the weight sensor 215 detects the weight of an item positioned above or near it. The weight sensor 215 may be positioned on a non-standard shelf 115 used to hold the item. The weight server 235 tracks the weight detected by the weight sensor 215 to determine whether and when the item 130 is removed from the shelf 115. Figure 5A-5J The weight sensor 215, the shelf 115, and the weight server 235 are described in more detail.

[0154] Figure 5A The illustration shows an example weight sensor 500 of the weight subsystem 206. (See diagram for example.) Figure 5A As seen in the diagram, the weight sensor 500 includes plates 510A and 510B, weighing units 505A, 505B, 505C, and 505D, and cables 515A, 515B, 515C, 515D, and 520. Generally, the components of the weight sensor 500 are assembled such that the weight sensor 500 can detect the weight of an object 130 positioned above or near the weight sensor 500.

[0155] Plate 510 forms a surface that distributes the weight of article 130 across the surface. Plate 510 can be made of any suitable material, such as, for example, metal and / or plastic. Articles 130 can be positioned above or near plate 510 and the weight of these articles 130 can be distributed across plate 510.

[0156] Weighing unit 505 is positioned between plates 510A and 510B. Weighing unit 505 generates an electrical signal based on the weight it bears. For example, weighing unit 505 can be a sensor that converts an input mechanical force (e.g., weight, tension, compression, pressure, or torque) into an output electrical signal (e.g., current or voltage). As the input force increases, the output electrical signal can increase proportionally. Weighing unit 505 can be any suitable type of weighing unit (e.g., hydraulic, pneumatic, and strain gauge). Although weighing units 1310 are shown as cylindrical, they can be any suitable size and shape appropriate for the particular embodiment conceived.

[0157] Signals from the weighing unit 505 can be analyzed to determine the total weight of the item 130 positioned above or near the weight sensor 500. The weighing unit 505 can be positioned such that the weight of the item 130 positioned above or near the weight sensor 500 is evenly distributed across each weighing unit 505. Figure 5A In the example, weighing unit 505 is positioned substantially equidistant from the corners of plates 510A and 510B. For example, weighing unit 505A is positioned at a corner distance d1 from plates 510A and 510B. Weighing unit 505B is positioned at a corner distance d2 from plates 510A and 510B. Weighing unit 505C is positioned at a corner distance d3 from plates 510A and 510B. Weighing unit 505D is positioned at a corner distance d4 from plates 510A and 510B. Distances d1, d2, d3, and d4 can be substantially equal to each other. This disclosure envisions distances differing by 5 to 10 millimeters and still considered substantially equal to each other. By positioning the weighing unit 505 at substantially equal distances from the corners of plates 510A and 510B, the weight of an object positioned above or near the weight sensor 500 is uniformly distributed across the weighing unit 505. Therefore, the total weight of an object positioned above or near the weight sensor 500 can be determined by summing the weights borne by each weighing unit 505.

[0158] Weighing unit 505 transmits an electrical signal indicating the weight borne by the weighing unit 505. For example, weighing unit 505 may generate a current that varies according to the weight or force borne by the weighing unit 505. Each weighing unit 505 is coupled to a cable 515 carrying the electrical signal. Figure 5AIn the example, weighing unit 505A is coupled to cable 515A; weighing unit 505B is coupled to cable 515B; weighing unit 505C is coupled to cable 515C; and weighing unit 505D is coupled to cable 515D. Cables 515 are grouped together to form cable 520 extending from weight sensor 500. Cable 520 carries the electrical signals generated by weighing unit 505 to a circuit board, which then transmits the signals to weight server 235.

[0159] Weight sensors 500 can be deployed in unconventional shelves 115 designed to hold items. Figure 5B Example shelf 525 is shown. (As...) Figure 5B As seen in the diagram, shelf 525 includes a base 530, one or more panels 535, and one or more shelves 540. Generally, the base 530 is located at the bottom of shelf 525 and forms the foundation for other components of shelf 525. Panels 535 extend vertically upward from the base 530. Shelves 540 are coupled to panels 535 and / or the base 530. For example, two shelves 540 may be coupled to and extend away from panels 535. Generally, panels 535 and the base 530 allow shelf 540 to hold the weight of items placed on shelf 540. A weight sensor 500 may be deployed within shelf 540 to detect the weight of items placed on shelf 540.

[0160] Figure 5C An exploded view of shelf 525 is shown. (As shown) Figure 5CAs seen in the diagram, the base 530 is formed using several surfaces 532. Surface 532A forms the bottom surface of the base 530. Surfaces 532B and 532D form the sides of the base 530. Surface 532C forms the back surface of the base 530. Surface 532E forms the top surface of the base 530. This disclosure contemplates using any suitable material (such as, for example, wood, metal, glass, and / or plastic) to form the base 530. Surface 532A may be coupled to surfaces 532B, 532C, and 532D. Surface 532B may be coupled to surfaces 532A, 532E, and 532C. Surface 532C may be coupled to surfaces 532A, 532B, 532D, and 532E. Surface 532D may be coupled to surfaces 532A, 532C, and 532E. Surface 532E may be coupled to surfaces 532B, 532C, and 532D. Surfaces 532B, 532C, and 532D extend upward from surface 532A. Generally, surfaces 532A, 532B, 532C, 532D, and 532E form a box-like structure around space 542. Base 530 includes a drawer 545 that can be opened to allow access to space 542. Drawer 545 is positioned within space 542. When drawer 545 is closed, base 530 can form a housing around space 542. When drawer 545 is open, access to space 542 is provided through the open drawer 545. In some embodiments, a door may be used to provide access to space 542 instead of drawer 545.

[0161] The surface 532E definition also allows access to the cavity 534 of the space 542. Generally, the cavity 534 allows the cable 520 from the weight sensor 500 to extend into the space 542.

[0162] Panel 535 extends upward from base 530. Panel 535 can be formed using any suitable material, such as, for example, wood, metal, and / or plastic. Figure 5C As seen in the diagram, panel 535 defines one or more cavities 550 extending along the width of panel 535. Cavities 550 allow cables 520 from weight sensor 500 to extend into spaces 552 defined by panel 535. Generally, space 552 is the hollow interior of panel 535. Cable 520 extends through cavities 550 and downwards toward cavity 534 into space 552. In this way, cable 520 can extend downwards from weight sensor 500 into space 542 in base 530. Each cavity 550 may correspond to a shelf 540 coupled to panel 535.

[0163] Each shelf 540 is coupled to a panel 535 and / or a base 530. A weight sensor 500 is deployed in the shelf 540. The shelf 540 may be coupled to the panel 535 such that cables 520 of the weight sensor 500 deployed in the shelf 540 extend from the weight sensor 500 through a cavity 550 into a space 552. These cables 520 then extend downwards along the space 552 and enter a space 542 through a cavity 534.

[0164] Figure 5D and 5E The illustration shows example shelf 540. Figure 5D A front view of shelf 540 is shown. Figure 5D As shown, shelf 540 includes a bottom surface 560A, a front surface 560B, and a rear surface 560C. Bottom surface 560A is coupled to front surface 560B and rear surface 560C. Front surface 560B and rear surface 560C extend upward from bottom surface 560A. A plurality of weight sensors 500 are positioned on bottom surface 560A between front surface 560B and rear surface 560C. Each weight sensor 500 is positioned to detect the weight of an item 130 positioned within a certain area 555 of shelf 540. A divider 558 can be used to designate each area 555. Items placed within a specific area 555 will be detected and weighed by the weight sensor 500 for that area 555. This disclosure contemplates shelf 540 made of any suitable material, such as, for example, wood, metal, glass, and / or plastic. Cables 515 and 520 are not included. Figure 5D As shown in the diagram, the structure of shelf 540 can be clearly displayed, but they are from... Figure 5D The omissions should not be interpreted as their removal. This disclosure contemplates the presence of cables 515 and 520 and... Figure 5D In the example, it is connected to a weight sensor 500.

[0165] Figure 5E A rear view of shelf 540 is shown. (As shown) Figure 5E As seen in the diagram, the rear surface 560C defines cavity 562. Cable 520 of weight sensor 500 extends from weight sensor 500 through cavity 562. Generally, the rear surface 560C of shelf 540 is coupled to panel 535 such that cavity 562 is at least partially aligned with cavity 550 in panel 535. In this way, cable 520 can extend from weight sensor 500 through cavity 562 and through cavity 550.

[0166] In some embodiments, a weight sensor 500 is positioned within a shelf 540 such that the weight sensor 500 detects the weight of items positioned within a specific area 555 of the shelf 540. For example, in Figure 5D and 5EAs seen in the example, shelf 540 includes four regions 555 positioned above four weight sensors 500. Each weight sensor 500 detects the weight of an item positioned within its corresponding region 555. Due to the positioning of the weight sensors 500, the weight sensors 500 can be unaffected by the weight of an item 130 positioned in a region 555 that does not correspond to that weight sensor 500.

[0167] Figure 5F Example base 530 is shown. Figure 5F As seen in the diagram, the base 530 can also accommodate a weight sensor 500. For example, the weight sensor 500 can be positioned on the top surface 532E of the base 530. Cables 520 for these weight sensors 500 can extend from the weight sensor 500 through the cavity 534 into the space 542. Thus, items can be positioned on the base 530 and their weight can be detected by the weight sensor 500.

[0168] Circuit board 565 is positioned within space 542. Circuit board 565 includes a port to which cables 520 from weight sensor 500, located on shelf 525, are connected. In other words, circuit board 565 is connected to cables 520 from weight sensor 500 positioned on base 530 and shelf 540. These cables 520 enter space 542 through cavity 534 and connect to circuit board 565. Circuit board 565 receives electrical signals generated by weighing unit 505 of weight sensor 500. Circuit board 565 then transmits a signal indicating the weight detected by weight sensor 500 to weight server 235. Drawer 545 can be opened to allow access to space 542 and circuit board 565. For example, drawer 545 can be opened to allow inspection and / or repair of circuit board 565.

[0169] Figure 5G Example circuit board 565 is shown. (Example...) Figure 5G As seen in the diagram, circuit board 565 includes processor 566 and multiple ports 568. Generally, ports 568 are coupled to cables 520 from weight sensor 500. This disclosure envisions circuit board 565 including any suitable number of ports 568 for connection to cables 520 from weight sensor 500 on shelf 525. Processor 566 receives and processes signals from ports 568.

[0170] Circuit board 565 can transmit signals to weight server 235 via any suitable medium. For example, circuit board 565 can transmit signals to weight server 235 via Ethernet connection, wireless connection (e.g., WiFi), Universal Serial Bus connection, and / or Bluetooth connection. Circuit board 565 can automatically select a connection through which to transmit signals to weight server 235. Circuit board 565 can select a connection based on priority. For example, if the Ethernet connection is active, circuit board 565 can select the Ethernet connection to communicate with weight server 235. If the Ethernet connection is disconnected but the wireless connection is active, circuit board 565 can select the wireless connection to communicate with weight server 235. If both the Ethernet and wireless connections are disconnected but the Universal Serial Bus connection is active, circuit board 565 can select the Universal Serial Bus connection to communicate with weighing server 235. If the Ethernet, wireless, and Universal Serial Bus connections are disconnected but the Bluetooth connection is active, circuit board 565 can select the Bluetooth connection to communicate with weight server 235. In this way, the circuit board 565 has improved resilience because it can continue to communicate with the weight server 235 even if some communication connections are lost.

[0171] Circuit board 565 can receive power through various connections. For example, circuit board 565 may include a power port 570 for supplying power to circuit board 565. A cable plugged into a power outlet can be coupled to power port 570 to supply power to circuit board 565. Circuit board 565 can also receive power via Ethernet connection and / or Universal Serial Bus connection.

[0172] Figure 5H The signal 572 generated by the weight sensor 500 is shown. (Example) Figure 5H As seen in the diagram, signal 572 begins by indicating a weight detected by weight sensor 500. Near time t1, an item positioned above weight sensor 500 is removed. Therefore, weight sensor 500 detects a decrease in weight, and signal 572 experiences a corresponding decrease. Beyond time t1, signal 572 continues to hover near a lower weight because item 130 has been removed. This disclosure contemplates that signal 572 may include noise introduced by the environment, making signal 572 not a perfectly straight or smooth signal.

[0173] Figure 5I An example operation of the weight server 235 is shown. Figure 5IAs seen in the image, weight server 235 receives a signal 572 indicating weight w0 from weight sensor 500 at time t0. Similar to camera server 225, weight server 235 can assign this information to a specific time window 332A based on the time indicated by t0. Later, weight server 235 can receive a signal 572 from weight sensor 500 indicating that a new weight w1 has been detected at time t1. Weight w1 may be less than weight w0, thus indicating that item 130 may have been removed. Weight server 235 assigns the information to a subsequent time window 332C based on the time indicated by t1.

[0174] The weight server 235 can implement an internal clock 304E that is synchronized with the internal clock 304 of other components of the tracking system 132, such as camera client 220, camera server 225, and central server 240. The weight server 235 can use a clock synchronization protocol (e.g., network time protocol and / or precision time protocol) to synchronize the internal clock 304E. The weight server 235 can use clock 304E to determine the time to receive signals 572 from the weight sensor 500 and assign these signals 572 to their appropriate time windows 332.

[0175] In some embodiments, the time window 332 in the weight server 235 is aligned with the time window 332 in the camera client 220, the camera server 225, and / or the central server 240. For example, the time window 332A in the weight server 235 may have a time window 332A aligned with the time window 332 in the camera client 220, the camera server 225, and / or the time window 332 in the central server 240. Figure 3J In the example, the time window 332A in camera server 225 has the same start time (T0) and end time (T1). In this way, information from different subsystems of tracking system 132 can be grouped according to the same time window 332, which allows this information to be correlated with each other in time.

[0176] Similar to camera server 225, weight server 235 can sequentially process information within time window 332 when it is ready for processing. Weight server 235 can process information within each time window 332 to determine whether item 130 was moved during that specific time window 332. Figure 5IIn the example, when weight server 235 processes the third time window 332C, weight server 235 can determine that sensor 1500 detected that two items were removed during time window 332C; thus, the weight drops from w0 to w1. Weight server 235 can make this determination by determining the difference between w0 and w1. Weight server 235 can also know (e.g., through a lookup table) the weight of item 130 located above or near weight sensor 500. Weight server 235 can divide by the difference between w0 and w1 to determine the number of items 130 that were removed. Weight server 235 can transmit this information to central server 240 for further processing. Central server 240 can use this information, along with the tracked location of people in the space, to determine which person in the space removed item 130.

[0177] Figure 5J An example method 580 for operating the weight subsystem 206 is shown. Generally, the various components of the weight subsystem 206 perform method 580 to determine when certain items 130 have been taken.

[0178] Weight sensor 215 detects the weight 582 borne above or around it and transmits this detected weight 582 to weight server 235 via electrical signal 572. Weight server 235 can analyze the signal 572 from weight sensor 215 to determine the number 584 of items 130 removed during a first time window 332. Weight server 235 can then transmit this determination to central server 240. Weight sensor 215 can subsequently detect the weight 586 borne by it and transmit that weight 586 to weight server 235. Weight server 235 can analyze that weight 586 to determine the number 588 of items 130 removed during a second time window 332. Weight server 235 can then transmit that determination to central server 240. Central server 240 can track whether items 130 were removed during a specific time window 332. If so, central server 240 can determine which person in the space took those items 130.

[0179] Can be Figure 5J The method 580 described herein may be modified, added to, or omitted. Method 580 may include more, fewer, or other steps. For example, the steps may be performed in parallel or in any suitable order. While various components of the execution steps of the weighted subsystem 206 have been discussed, any suitable component of the tracking system 132 (e.g., such as the central server 240) may perform one or more steps of the method.

[0180] V. Central Server

[0181] Figures 6A-6C The operation of the central server 240 is illustrated. Generally, the central server 240 analyzes information from various subsystems (e.g., camera subsystem 202, LiDAR subsystem 204, weight subsystem 206, etc.) and determines which person removed which items from the space. As discussed earlier, these subsystems group the information into time windows 332 aligned across subsystems. By grouping the information into aligned time windows 332, the central server 240 can find relationships between information from different subsystems and collect additional information (e.g., which person removed which item 130). In some embodiments, when people leave store 100, the central server 240 also charges them for the items they removed from the space.

[0182] Figure 6A and 6B Example operation of central server 240 is shown. Figure 6A As seen in the diagram, central server 240 receives information from various servers during specific time windows. Figure 6A In the example, the central server 240 receives the physical positions of two people in space from the camera server 225 during the first time window 332A. This disclosure uses uppercase "X" and uppercase "Y" to represent the physical coordinates 602 of a person or object in space and distinguishes the physical coordinates 602 of a person or object in space determined by the camera server 225 and the LiDAR server 230 from local coordinates determined by other components (e.g., coordinates 322 determined by the camera client 220 and coordinates 410 determined by the LiDAR sensor 405).

[0183] According to camera server 225, the first person is at physical coordinates 602 (X1, Y1), and the second person is at physical coordinates 602 (X2, Y2). Furthermore, central server 240 receives the physical locations of these two people from LiDAR server 230. According to LiDAR server 230, the first person is at coordinates 602 (X7, Y7), and the second person is at coordinates 602 (X8, Y8). Additionally, central server 240 also receives information from weight server 235 during the first time window 332A. According to weight server 235, no item 130 was taken during the first time window 332A.

[0184] This disclosure envisions that the central server 240 uses any suitable process to analyze the physical location of a person from the camera server 225 and the LiDAR server 230. Although the coordinates 602 provided by the camera server 225 and the LiDAR server 230 may differ from each other, the central server 240 can use any suitable process to reconcile these differences. For example, if the coordinates 602 provided by the LiDAR server 230 differ from the coordinates 602 provided by the camera server 225 by an amount exceeding a threshold, then the central server 240 can use the coordinates 602 provided by the camera server 225. In this way, the coordinates 602 provided by the LiDAR server 230 are used as a check against the coordinates 602 provided by the camera server 225.

[0185] During the second time window 332B, the central server 240 receives the physical coordinates 602 of the two individuals from the camera server 225. According to the camera server 225, during the second time window 332B, the first person is at coordinates 602 (X3, Y3), and the second person is at coordinates 602 (X4, Y4). During the second time window 332B, the camera server 240 also receives the physical coordinates 602 of the two individuals from the LiDAR server 230. According to the LiDAR server 230, during the second time window 332B, the first person is at coordinates 602 (X9, Y9), and the second person is at coordinates 602 (X4, Y4). 10 Y 10 In addition, the central server 240 learned from the weight server 235 that no item 130 was taken during the second time window 332B.

[0186] During the third time window 332C, camera server 240 receives the physical coordinates 602 of two people from camera server 225. According to camera server 225, the first person is at coordinates 602 (X5, Y5), and the second person is at coordinates 602 (X6, Y6). Central server 240 also receives the physical coordinates 602 of two people from LiDAR server 230 during the third time window 332C. According to LiDAR server 230, during the third time window 332C, the first person is at coordinates 602 (X5, Y5), and the second person is at coordinates 602 (X6, Y6). 11 Y 11 The second person is at coordinate 602 (X). 12 Y 12 In addition, the central server 240 learns from the weight server 235 that a specific weight sensor 500 detected that two items 130 were taken during the third time window 332C.

[0187] In response to the detection by weight sensor 500 that two items 130 have been taken, central server 240 may perform additional analysis to determine who took the two items 130. Central server 240 executes any appropriate procedures to determine who took the items 130. Some of these procedures are disclosed in U.S. Patent Application No.____ (Attorney General's File No. 090278.0180) entitled "Topview Object Tracking Using a Sensor Array," the contents of which are incorporated herein by reference.

[0188] Figure 6B The example analysis performed by central server 240 is shown to determine who took item 130. Figure 6B As seen in the diagram, the central server 240 first determines the physical coordinates 602 of the two individuals during the third time window 332C. The central server 240 determines that the first person was at coordinates 602 (X5, Y5) and the second person was at coordinates 602 (X6, Y6) during the third time window 332C. The central server 240 also determines the physical location of the weight sensor 500 that detected the removed item. Figure 6B In the example, the central server 240 determines that the weight sensor 500 is located at coordinate 602 (X). 13 Y 13 ) place.

[0189] The central server 240 then determines the distance from each person to the weight sensor 500. The central server 240 determines that the distance from the first person to the weight sensor 500 is 1 and the distance from the second person to the weight sensor 500 is 2. The central server 240 then determines which person is closer to the weight sensor 500. Figure 4B In the example, the central server 240 determines that distance 1 is less than distance 2, therefore the first person is closer to the weight sensor 500 than the second person. Therefore, the central server 240 determines that the first person took both items 130 during the third time window 332C and that the first person should be charged for both items 130.

[0190] Figure 6C An example method 600 for operating a central server 240 is illustrated. In a particular embodiment, the central server 240 performs the steps of method 600 to determine which person in the space took item 130.

[0191] At step 605, the central server 240 begins receiving the coordinates 602 of the first person in the space during time window 332. At step 610, the central server 240 receives the coordinates 602 of the second person in the space during time window 332. In step 615, the central server 240 receives an indication that item 130 was taken during time window 332. In response to receiving that indication, the central server 240 analyzes the information to determine which person took item 130.

[0192] In step 620, the central server 240 determines that during time window 332, the first person was closer to item 130 than the second person. The central server 240 may make this determination based on a predetermined distance between the person and the weight sensor 500 that detected the removal of item 130. In step 625, in response to determining that the first person was closer to item 130 than the second person, the central server 240 determines that the first person took item 130 during time window 332. Then, when the first person leaves store 100, the first person can be charged for item 130.

[0193] Can be Figure 6C The method 600 described herein may be modified, added to, or omitted. Method 600 may include more, fewer, or other steps. For example, the steps may be performed in parallel or in any suitable order. Although the execution of steps is discussed as being performed by the central server 240, any suitable component of the tracing system 132 may execute one or more steps of the method.

[0194] VI. Hardware

[0195] Figure 7 An example computer 700 used in tracking system 132 is illustrated. Generally, computer 700 can be used to implement components of tracking system 132. For example, computer 700 can be used to implement camera client 220, camera server 225, LiDAR server 230, weight server 235, and / or central server 240. Figure 7 As seen herein, computer 700 includes various hardware components such as processor 705, memory 710, graphics processor 715, input / output ports 720, communication interface 725, and bus 730. This disclosure contemplates that the components of computer 700 are configured to perform any of the functions discussed herein as camera client 220, camera server 225, LiDAR server 230, weight server 235, and / or central server 240. Circuit board 565 may also include some components of computer 700.

[0196] Processor 705 is any electronic circuit system, including but not limited to microprocessors, application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), and / or state machines, which is communicatively coupled to memory 710 and controls the operation of computer 700. Processor 705 may be 8-bit, 16-bit, 32-bit, 64-bit, or any other suitable architecture. Processor 705 may include an arithmetic logic unit (ALU) for performing arithmetic and logical operations, processor registers for supplying operands to the ALU and storing the results of ALU operations, and a control unit for fetching instructions from memory and executing them by coordinating the operation of the ALU, registers, and other components. Processor 705 may include other hardware with operating software to control and process information. Processor 705 executes software stored in memory to perform any of the functions described herein. Processor 705 controls the operation and management of computer 700 by processing information received from memory 710 and / or other computer 700. Processor 705 may be a programmable logic device, a microcontroller, a microprocessor, any suitable processing device, or any suitable combination thereof. Processor 705 is not limited to a single processing device and may encompass multiple processing devices.

[0197] Memory 710 may store data, operating software, or other information for processor 705, either permanently or temporarily. Memory 710 may include any one or a combination of volatile or non-volatile local or remote devices suitable for storing information. For example, memory 710 may include random access memory (RAM), read-only memory (ROM), magnetic storage devices, optical storage devices, or any other suitable information storage devices or combinations thereof. Software refers to any suitable set of instructions, logic, or code implemented in a computer-readable storage medium. For example, software may be implemented in memory 710, a disk, CD, or flash drive. In a particular embodiment, software may include an application that can be run by processor 705 to perform one or more of the functions described herein.

[0198] The graphics processor 715 can be any electronic circuit system that receives and analyzes video data, including but not limited to microprocessors, application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), and / or state machines. For example, the graphics processor 715 can process video data to determine the correct signals to send to the display so that the display can show an appropriate image. The graphics processor 715 can also process video data to identify certain features (e.g., people or objects) within the video. The graphics processor 715 can be a component of a video card installed in computer 700.

[0199] Input / output port 720 allows peripheral devices to be connected to computer 700. Port 720 can be any suitable port, such as a parallel port, serial port, optical port, video port, network port, etc. Peripherals such as keyboards, mice, joysticks, optical tracking devices, touchpads, etc., can be connected to computer 700 through port 720. Input and output signals are transmitted between computer 700 and peripheral devices through port 720.

[0200] The communication interface 725 includes any suitable hardware and / or software for communication over a network. For example, the communication interface 725 may include a modem, network interface card (NIC), Ethernet port / controller, wireless radio / controller, cellular radio / controller, and / or Universal Serial Bus port / controller. The computer 700 can use the communication interface 725 to communicate with other devices over a communication network.

[0201] Bus 730 allows components of computer 700 to communicate with each other. Computer 700 may include a bus controller 730 that manages communication via bus 730.

[0202] While this disclosure includes several embodiments, numerous changes, variations, alterations, modifications and alterations may be made to those skilled in the art, and this disclosure is intended to cover such changes, variations, alterations, modifications and alterations that fall within the scope of the appended claims.

[0203] Terms and Conditions

[0204] 1. A system comprising:

[0205] An array of cameras, positioned above the space, with each camera in the array configured to capture video of a portion of the space containing people;

[0206] The first camera client is configured as follows:

[0207] The first multiple frames of a first video are received from the first camera in the array of cameras, each of the first multiple frames showing a person in space;

[0208] For the first frame in the first plurality of frames:

[0209] Determine the first defined area around the person shown in the first frame; and

[0210] Generate the first timestamp of when the first frame is received by the first camera client;

[0211] For the second frame in the first set of multiple frames:

[0212] Determine the second defined area around the person shown in the second frame; and

[0213] Generate a second timestamp indicating when the second frame is received by the first camera client;

[0214] The second camera client is separate from the first camera client, and the second camera client is configured as follows:

[0215] The second camera in the array of cameras receives a second plurality of frames of the second video, each of which shows a person in the space;

[0216] For the third frame in the second set of frames:

[0217] Determine the third defined region surrounding the person shown in the third frame; and

[0218] Generate a third timestamp indicating when the third frame is received by the second camera client;

[0219] For the fourth frame in the second set of frames:

[0220] Determine the fourth defined region around the person shown in the fourth frame; and

[0221] Generate a fourth timestamp indicating when the fourth frame is received by the second camera client;

[0222] The camera server, separate from the first camera client and the second camera client, is configured as follows:

[0223] Ensure the first timestamp falls within the first time window;

[0224] In response to determining that the first timestamp falls within the first time window, the coordinates of the first defined region are assigned to the first time window;

[0225] Ensure the second timestamp falls within the first time window;

[0226] In response to determining that the second timestamp falls within the first time window, the coordinates defining the second bounding region are assigned to the first time window;

[0227] Ensure that the third timestamp falls within the first time window;

[0228] In response to determining that the third timestamp falls within the first time window, the coordinates defining the third bounding region are assigned to the first time window;

[0229] Determine that the fourth timestamp falls within the second time window following the first time window;

[0230] In response to determining that the fourth timestamp falls within the second time window, the coordinates defining the fourth bounding region are assigned to the second time window;

[0231] Determine the coordinates that should be processed and assigned to the first time window;

[0232] In response to determining the coordinates that should be processed and assigned to the first time window:

[0233] Based at least on the coordinates of the first defined region and the coordinates of the second defined region, the combined coordinates of the person during the first time window are calculated for the first video from the first camera; and

[0234] At least based on the coordinates of the third defined region, the combined coordinates of the person during the first time window are calculated for the second video from the second camera; and

[0235] The location of a person in space during the first time window is determined based at least on the combined coordinates of the person during the first time window for the first video from the first camera and the combined coordinates of the person during the first time window for the second video from the second camera.

[0236] Multiple weight sensors are positioned in space, each of the multiple weight sensors being configured to generate a signal indicating the weight borne by that weight sensor;

[0237] A weight server, separate from the first camera client, the second camera client, and the camera server, is configured to determine, at least based on a signal generated by a first weight sensor among the plurality of weight sensors, that an item positioned above the first weight sensor has been removed; and

[0238] A central server, separate from the first camera client, the second camera client, the camera server, and the weight server, is configured to determine that a person has removed an item based at least on the person's location within the space during a first time window, wherein, based at least on the determination that the first person has removed an item, the person is charged for the item when leaving the space.

[0239] 2. The system as described in Clause 1 further includes:

[0240] An array of LiDAR (Light Detection and Ranging) sensors is positioned above space; and

[0241] The LiDAR server, separate from the first camera client, the second camera client, the camera server, the weight server, and the central server, is configured to determine the location of a person in space during a first time window based at least on coordinates received from LiDAR sensors in an array of LiDAR sensors.

[0242] 3. The system as described in Clause 1, wherein the determination of the first person removing the item is based at least on (1) the distance between the person's position in space and the position of the first weight sensor in space, and (2) the time when the weight server receives the signal generated by the first weight sensor within the first time window.

[0243] 4. The system as described in Clause 1, wherein:

[0244] The first camera client is also configured to receive the height of the person shown in the first frame and the height of the person shown in the second frame.

[0245] The second camera client is also configured to receive the height of the person shown in the third frame and the height of the person shown in the fourth frame; and

[0246] The camera server is also configured as follows:

[0247] Based at least on the height of the person shown in the first frame and the height of the person shown in the second frame, calculate the combined height of the person during the first time window for the first video from the first camera; and

[0248] Based at least on the height of the person shown in the third frame, the combined coordinates of the person during the first time window are calculated for the second video from the second camera.

[0249] 5. The system as described in Clause 1, wherein:

[0250] The first camera client implements a first clock for generating the first and second timestamps;

[0251] The second camera client implements a second clock for generating the third and fourth timestamps; and

[0252] The camera server implements a third clock, and the first, second, and third clocks are synchronized using a clock synchronization protocol.

[0253] 6. The system as described in Clause 5, wherein the weight server implements a fourth clock, which is synchronized with the first, second, and third clocks using a clock synchronization protocol.

[0254] 7. The system as described in Clause 1, wherein the combined coordinates of a person during a first time window for the first video from the first camera include the average of coordinates defining a first defined region and coordinates defining a second defined region.

[0255] 8. The system as described in Clause 1, wherein the array of cameras is arranged in a grid such that a camera that is communicatively coupled to a first camera client is not directly adjacent to another camera that is communicatively coupled to the first camera client in the same row or column of the grid.

[0256] 9. The system as described in Clause 1 further includes:

[0257] A shelf within a space, the shelf comprising a shelf panel and a base including drawers, the base being vertically positioned below the shelf panel. The shelf panel is divided into a first area and a second area. A first weight sensor of a plurality of weight sensors is positioned within the first area and configured to generate a first signal based at least on the weight borne by the first weight sensor within the first area. A second weight sensor of the plurality of weight sensors is positioned within the second area and configured to generate a second signal based at least on the weight borne by the second weight sensor within the second area. Each weight sensor of the plurality of weight sensors includes a plurality of weighing units.

[0258] A circuit board, positioned inside the drawer, is communicatively coupled to a first weight sensor and a second weight sensor. The circuit board is configured as follows:

[0259] Receive the first signal and the second signal;

[0260] Transmit a signal to the weight server indicating the weight borne by the first weight sensor; and

[0261] A signal indicating the weight borne by the second weight sensor is transmitted to the weight server.

[0262] 10. The system as described in Clause 1, wherein the camera server is further configured to request a first frame from a first camera client and a third frame from a second camera client, and the determination of the person's position is further based on at least the first and third frames.

[0263] 11. A method comprising:

[0264] The first camera client receives the first plurality of frames of the first video from the first camera in the camera array, the camera array being positioned above space, each camera in the camera array being configured to capture video of a portion of the space containing people, and each of the first plurality of frames showing people in the space.

[0265] For the first frame in the first plurality of frames:

[0266] The first camera client determines a first defined area around the person shown in the first frame; and

[0267] The first timestamp of when the first frame was received by the first camera client is generated by the first camera client;

[0268] For the second frame in the first set of multiple frames:

[0269] The second defined area around the person shown in the second frame is determined by the first camera client; and

[0270] A second timestamp is generated by the first camera client to indicate when the second frame was received by the first camera client.

[0271] A second camera client, separate from the first camera client, receives a second plurality of frames of the second video from the second camera in the array of cameras, each of the second plurality of frames showing a person in the space;

[0272] For the third frame in the second set of frames:

[0273] The third defined area around the person shown in the third frame is determined by the second camera client; and

[0274] A third timestamp generated by the second camera client to indicate when the third frame was received by the second camera client;

[0275] For the fourth frame in the second set of frames:

[0276] The fourth defined region around the person shown in the fourth frame is determined by the second camera client; and

[0277] A fourth timestamp generated by the second camera client to indicate when the fourth frame was received by the second camera client;

[0278] The camera server, which is separate from the first camera client and the second camera client, determines whether the first timestamp falls within the first time window;

[0279] In response to determining that the first timestamp falls within the first time window, the camera server assigns the coordinates of the first defined region to the first time window.

[0280] The camera server determines whether the second timestamp falls within the first time window;

[0281] In response to the determination that the second timestamp falls within the first time window, the camera server assigns the coordinates of the defined second bounding region to the first time window.

[0282] The camera server determines whether the third timestamp falls within the first time window;

[0283] In response to the determination that the third timestamp falls within the first time window, the camera server assigns the coordinates of the defined third bounding region to the first time window.

[0284] The camera server determines that the fourth timestamp falls within the second time window following the first time window.

[0285] In response to the determination that the fourth timestamp falls within the second time window, the camera server assigns the coordinates of the fourth defined region to the second time window.

[0286] The camera server determines the coordinates that should be processed and assigned to the first time window;

[0287] In response to determining the coordinates that should be processed and assigned to the first time window:

[0288] The camera server calculates, based at least on the coordinates of a person during a first time window, for the first video from the first camera, a combination of coordinates of a person within a first time window, using coordinates defining a first bounding region and coordinates defining a second bounding region; and

[0289] The camera server calculates the combined coordinates of the person during the first time window based on the coordinates of the second video from the second camera, at least based on the coordinates of the third defined region.

[0290] The camera server determines the location of a person in space during the first time window based at least on the combined coordinates of the person during the first time window for the first video from the first camera and the combined coordinates of the person during the first time window for the second video from the second camera.

[0291] Multiple weight sensors positioned in space generate signals indicating the weight borne by the multiple weight sensors;

[0292] A weight server, separate from the first camera client, the second camera client, and the camera server, determines, at least based on a signal generated by the first weight sensor among the plurality of weight sensors, that an item positioned above the first weight sensor has been removed.

[0293] A central server, separate from the first camera client and the second camera client, the camera server, and the weight server, determines that a person has removed an item based at least on the person's location within the space during a first time window. The person is charged for the item when they leave the space, based at least on the determination that the first person has removed the item.

[0294] 12. The method of Clause 11 further includes determining the location of a person in space during a first time window by a light detection and ranging (LiDAR) server separate from the first camera client and the second camera client, the camera server, the weight server and the central server, based at least on coordinates received from LiDAR sensors in an array of LiDAR sensors positioned above space.

[0295] 13. The method as described in Clause 11, wherein the determination of the first person removing the item is based at least on (1) the distance between the person's position in space and the position of the first weight sensor in space, and (2) the time when the weight server receives the signal generated by the first weight sensor within the first time window.

[0296] 14. The system as described in Clause 11 further includes:

[0297] The height of the person shown in the first frame and the height of the person shown in the second frame are received by the first camera client;

[0298] The height of the person shown in the third frame and the height of the person shown in the fourth frame are received by the second camera client;

[0299] The camera server calculates a combined height of the person during a first time window, based at least on the height of the person shown in the first frame and the height of the person shown in the second frame, for the first video from the first camera; and

[0300] The camera server calculates the combined coordinates of the person during the first time window based on the height of the person shown in the third frame, for the second video from the second camera.

[0301] 15. The method described in Clause 11 further includes:

[0302] A first clock, implemented by the first camera client, is used to generate the first and second timestamps.

[0303] A second clock, implemented by the second camera client, is used to generate the third and fourth timestamps; and

[0304] The third clock is implemented by the camera server, and the first, second and third clocks are synchronized using a clock synchronization protocol.

[0305] 16. The method described in Clause 15 further includes a fourth clock implemented by a weight server, the fourth clock being synchronized with the first, second, and third clocks using a clock synchronization protocol.

[0306] 17. The method as described in Clause 11, wherein the combined coordinates of a person during a first time window for the first video from the first camera include the average of the coordinates defining a first defined region and the coordinates defining a second defined region.

[0307] 18. The method of Clause 11, wherein the array of cameras is arranged in a grid such that a camera that is communicatively coupled to a first camera client is not directly adjacent to another camera that is communicatively coupled to a first camera client in the same row or column of the grid.

[0308] 19. The system as described in Clause 11 further includes:

[0309] A circuit board receives a first signal from a first weight sensor among the plurality of weight sensors. The circuit board is positioned inside a drawer of a shelf located in the space. The shelf includes a shelf and a base. The base includes a drawer and its bottom is vertically positioned below the shelf. The shelf is divided into a first area and a second area. The first weight sensor is positioned in the first area. The first signal indicates the weight borne by the first weight sensor in the first area. Each of the plurality of weight sensors includes a plurality of weighing units.

[0310] The circuit board receives a second signal from a second weight sensor among the plurality of weight sensors. The second weight sensor is located in a second region, and the second signal indicates the weight borne by the second weight sensor in the second region.

[0311] The circuit board transmits a signal to the weight server indicating the weight borne by the first weight sensor; and

[0312] The circuit board transmits a signal to the weight server indicating the weight borne by the second weight sensor.

[0313] 20. The method as described in Clause 11 further includes a camera server requesting a first frame from a first camera client and a third frame from a second camera client, wherein the determination of the person's position is further based at least on the first and third frames.

[0314] 21. A system comprising:

[0315] First Camera Client;

[0316] The second camera client is separate from the first camera client;

[0317] The third-camera client is separate from the first-camera client and the second-camera client; and

[0318] An array of cameras, positioned above space, is arranged in a rectangular grid, comprising a first row, second row, third row, first column, second column, and third column. This array includes:

[0319] The first camera is positioned in the first row and first column of the grid. The first camera is communicatively coupled to the first camera client. The first camera is configured to transmit video of the first portion of the space covered by the first field of view of the first camera to the first camera client.

[0320] The second camera is positioned in the first row and second column of the grid, such that the second camera is directly adjacent to the first camera in the grid. The second camera is communicatively coupled to the second camera client. The second camera is configured to transmit video of the second part of the space covered by the second field of view of the second camera to the second camera client.

[0321] The third camera is positioned in the first row and third column of the grid, so that the third camera is directly adjacent to the second camera in the grid. The third camera is communicationally coupled to the third camera client. The third camera is configured to transmit video of the third part of the space covered by the third field of view of the third camera to the third camera client.

[0322] The fourth camera is positioned in the second row and first column of the grid, such that the fourth camera is directly adjacent to the first camera in the grid. The fourth camera is communicatively coupled to the second camera client. The fourth camera is configured to transmit video of the fourth part of the space covered by the fourth field of view of the fourth camera to the second camera client.

[0323] A fifth camera, positioned in the second row and second column of the grid, is directly adjacent to the fourth and second cameras within the grid. The fifth camera is communicatively coupled to a third camera client. The fifth camera is configured to transmit video of the fifth portion of the space covered by its fifth field of view to the third camera client.

[0324] The sixth camera is positioned in the third row and first column of the grid, such that the sixth camera is directly adjacent to the fourth camera in the grid. The sixth camera is communicatively coupled to the third camera client. The sixth camera is configured to transmit video of the sixth part of the space covered by the sixth field of view of the sixth camera to the third camera client.

[0325] The fifth field of view of the fifth camera partially overlaps with the first field of view of the first camera, the second field of view of the second camera, and the fourth field of view of the fourth camera, such that if the third camera client is offline, at least a portion of the fifth part of the space is still captured in the video transmitted from the first camera to the first camera client, the video transmitted from the second camera to the second camera client, and the video transmitted from the fourth camera to the second camera client.

[0326] 22. The system as described in Clause 21, wherein the array of cameras is arranged such that each camera communicating with a first camera client is not directly adjacent to another camera in the array of cameras communicating with the first camera client in the same row or column of the grid.

[0327] 23. The system as described in Clause 21, wherein the array of cameras is arranged such that a 2×2 portion of the array includes a camera communicatively coupled to a first camera client, a camera communicatively coupled to a second camera client, and a camera communicatively coupled to a third camera client.

[0328] 24. The system as described in Clause 21, wherein the camera array is arranged such that the communication of the camera array to a first camera client is coupled diagonally to another camera in the grid that is coupled to the communication of the camera array to the first camera client.

[0329] 25. The system as described in Clause 21, wherein a first portion of the space captured by the first camera partially overlaps with a second, fourth, and fifth portion of the space captured by the second, fourth, and fifth cameras, respectively.

[0330] 26. The system as described in Clause 21, wherein:

[0331] The grid also includes a fourth column; and

[0332] The camera array also includes:

[0333] The seventh camera is positioned in the first row and fourth column of the grid, making it directly adjacent to the third camera in the grid. The seventh camera is communicatively coupled to the first camera client, and the seventh camera is configured to transmit video of the seventh part of the space to the first camera client.

[0334] The eighth camera, positioned in the second row and third column of the grid, is directly adjacent to the third and fifth cameras within the grid. The eighth camera is communicatively coupled to the first camera client and is configured to transmit video from the eighth portion of the space to the first camera client.

[0335] The ninth camera is positioned in the third row and second column of the grid, making it directly adjacent to the fifth and sixth cameras in the grid. The ninth camera is communicatively coupled to the first camera client and is configured to transmit video of the ninth part of the space to the first camera client.

[0336] 27. The system as described in Clause 21, wherein:

[0337] The first camera is hardwired to the first camera client;

[0338] The second and fourth cameras are hard-wired to the second camera client; and

[0339] The third, fifth, and sixth cameras are hardwired to the third camera client.

[0340] 28. The system as described in Clause 21, wherein the first camera client is configured as follows:

[0341] Determine the user's first coordinates within the first frame of the video from the first camera;

[0342] Generate the first timestamp of when the first camera client receives the first frame;

[0343] Determine the user's second coordinates within a second frame of video from the first camera, which is received after the first frame;

[0344] Generate a second timestamp indicating when the first camera client receives the second frame;

[0345] Determine the user's third coordinates within the third frame of the video from the first camera, which is received after the second frame; and

[0346] Generate the third timestamp of when the first camera client receives the third frame.

[0347] 29. The system as described in Clause 28 further includes a camera server separate from the first camera client, the second camera client, and the third camera client, the camera server being configured as follows:

[0348] Determine if the first and second timestamps fall within the first time window;

[0349] In response to determining that the first timestamp and the second timestamp fall within the first time window, the first coordinate and the second coordinate are assigned to the first time window.

[0350] Determine that the third timestamp falls within the second time window following the first time window;

[0351] In response to determining that the third timestamp falls within the second time window, the third coordinate is assigned to the second time window;

[0352] Based on the first and second coordinates, determine the user's combined coordinates during the first time window; and

[0353] After determining the user's combined coordinates during the first time window, the user's combined coordinates during the second time window are determined based on the third coordinate.

[0354] 30. The system as described in Clause 21, wherein the first camera, second camera, third camera, fourth camera, fifth camera and sixth camera are further configured to detect the user's height.

[0355] 31. The system as described in Clause 21 further includes a second array of cameras positioned above space, the second array of cameras being arranged in a second rectangular grid, the second array of cameras being different from the array of cameras, each camera in the second array of cameras being offset from the cameras in the array of cameras.

[0356] 32. A method comprising:

[0357] The first camera in the camera array transmits video of a first portion of the space covered by the first field of view of the first camera to the first camera client. The camera array is positioned above the space, and the cameras in the camera array are arranged in a rectangular grid, which includes a first row, a second row, a third row, a first column, a second column, and a third column. The first camera is positioned in the first row and the first column of the grid, and the first camera is communicatively coupled to the first camera client.

[0358] The video of the second part of the space covered by the second field of view of the second camera in the camera array is transmitted to the second camera client, which is separate from the first camera client. The second camera is positioned in the first row and second column of the grid, such that the second camera is directly adjacent to the first camera in the grid. The second camera is communicationally coupled to the second camera client.

[0359] The video of the third part of the space covered by the third field of view of the third camera in the camera array is transmitted to the third camera client, which is separate from the first and second camera clients. The third camera is located in the first row and third column of the grid, so that the third camera is directly adjacent to the second camera in the grid. The third camera communication is coupled to the third camera client.

[0360] The video of the fourth part of the space covered by the fourth field of view of the fourth camera in the camera array is transmitted to the second camera client. The fourth camera is located in the second row and first column of the grid, so that the fourth camera is directly adjacent to the first camera in the grid. The fourth camera is communication coupled to the second camera client.

[0361] The video of the fifth portion of the space covered by the fifth field of view of the fifth camera in the camera array is transmitted to the third camera client. The fifth camera is located in the second row and second column of the grid, so that the fifth camera is directly adjacent to the fourth and second cameras in the grid. The fifth camera is communicatively coupled to the third camera client.

[0362] The sixth camera in the camera array transmits video of the sixth part of the space covered by the sixth field of view of the sixth camera to the third camera client. The sixth camera is located in the third row and first column of the grid, so that the sixth camera is directly adjacent to the fourth camera in the grid. The sixth camera is communicationally coupled to the third camera client.

[0363] The fifth field of view of the fifth camera partially overlaps with the first field of view of the first camera, the second field of view of the second camera, and the fourth field of view of the fourth camera, such that if the third camera client is offline, at least a portion of the fifth part of the space is still captured in the video transmitted from the first camera to the first camera client, the video transmitted from the second camera to the second camera client, and the video transmitted from the fourth camera to the second camera client.

[0364] 33. The method as described in Clause 32, wherein the array of cameras is arranged such that each camera in the array of cameras that is communicatively coupled to the first camera client is not directly adjacent to another camera in the array of cameras that is communicatively coupled to the first camera client.

[0365] 34. The method of claim 32, wherein the array of cameras is arranged such that a 2×2 portion of the array includes a camera communicatively coupled to a first camera client, a camera communicatively coupled to a second camera client, and a camera communicatively coupled to a third camera client.

[0366] 35. The method of claim 32, wherein the array of cameras is arranged such that a camera that is communicatively coupled to a first camera client is diagonally opposite another camera that is communicatively coupled to the first camera client in the grid.

[0367] 36. The method as described in Clause 32, wherein a first portion of the space captured by the first camera partially overlaps with a second, fourth, and fifth portion of the space captured by the second, fourth, and fifth cameras, respectively.

[0368] 37. The method described in Clause 32 further includes:

[0369] The seventh camera in the camera array transmits video of the seventh part of the space to the first camera client. The grid also includes a fourth column. The seventh camera is positioned in the first row and fourth column of the grid, so that the seventh camera is directly adjacent to the third camera in the grid. The seventh camera is communicationally coupled to the first camera client.

[0370] The eighth camera in the camera array transmits video of the eighth portion of the space to the first camera client. The eighth camera is positioned in the second row and third column of the grid, making it directly adjacent to the third and fifth cameras within the grid. The eighth camera's communication is coupled to the first camera client.

[0371] The ninth camera in the camera array transmits video of the ninth part of the space to the first camera client. The ninth camera is positioned in the third row and second column of the grid, so that the ninth camera is directly adjacent to the fifth and sixth cameras in the grid. The ninth camera is communicatively coupled to the first camera client.

[0372] 38. The method as described in Clause 32, wherein:

[0373] The first camera is hardwired to the first camera client;

[0374] The second and fourth cameras are hard-wired to the second camera client; and

[0375] The third, fifth, and sixth cameras are hardwired to the third camera client.

[0376] 39. The method described in Clause 32 further includes:

[0377] The first camera client determines the user's first coordinates within the first frame of the video from the first camera;

[0378] The first time stamp of when the first camera client received the first frame is generated by the first camera client.

[0379] The user's second coordinates within a second frame of video from the first camera are determined by the first camera client. The second frame is received after the first frame.

[0380] A second timestamp is generated by the first camera client to indicate when the first camera client received the second frame.

[0381] The user's third coordinates within the third frame of the video from the first camera are determined by the first camera client; the third frame is received after the second frame; and

[0382] The first camera client generates a third timestamp indicating when it received the third frame.

[0383] 40. The method described in Clause 39 further includes:

[0384] The camera server, which is separate from the first camera client, the second camera client, and the third camera client, determines whether the first timestamp and the second timestamp fall within the first time window;

[0385] In response to determining that the first timestamp and the second timestamp fall within the first time window, the camera server assigns the first coordinates and the second coordinates to the first time window.

[0386] The camera server determines that the third timestamp falls within the second time window following the first time window.

[0387] In response to the determination that the third timestamp falls within the second time window, the camera server assigns the third coordinates to the second time window.

[0388] The camera server determines the user's combined coordinates during the first time window based on the first and second coordinates; and

[0389] After determining the user's combined coordinates during the first time window, the camera server determines the user's combined coordinates during the second time window based on the third coordinates.

[0390] 41. The method as described in Clause 32 further includes detecting the user's height using a first camera, a second camera, a third camera, a fourth camera, a fifth camera, and a sixth camera.

[0391] 42. The method of claim 32, wherein a second array of cameras is positioned above space, the second array of cameras is arranged in a second rectangular grid, the second array of cameras is different from the array of cameras, and each camera in the second array of cameras is offset from the camera in the array of cameras.

[0392] 43. A system comprising:

[0393] Circuit boards; and

[0394] Shelves, including:

[0395] The base includes:

[0396] Bottom surface;

[0397] The first side surface is coupled to the bottom surface of the base, and the first side surface of the base extends upward from the bottom surface of the base;

[0398] The second side surface is coupled to the bottom surface of the base and the first side surface, and the second side surface of the base extends upward from the bottom surface of the base;

[0399] The third side surface is coupled to the bottom surface and the second side surface of the base, and the third side surface of the base extends upward from the bottom surface of the base.

[0400] The top surface, coupled to the first, second, and third side surfaces of the base, such that the bottom and top surfaces of the base, as well as the first, second, and third side surfaces of the base, define a space, and the top surface of the base defines a first opening into the space; and

[0401] The drawer is positioned within the space, and the circuit board is positioned inside the drawer;

[0402] A panel, coupled to and extending upward from the base, defines a second opening extending along the width of the panel;

[0403] A shelf, coupled to a panel, is vertically positioned above a base and extends away from the panel. The shelf includes a bottom surface, a front surface extending upward from the bottom surface of the shelf, and a rear surface extending upward from the bottom surface of the shelf. The rear surface of the shelf is coupled to the panel and defines a third opening, a portion of which is aligned with a portion of a second opening.

[0404] The first weight sensor is coupled to the bottom surface of the shelf and positioned between the front and rear surfaces of the shelf.

[0405] The second weight sensor is coupled to the bottom surface of the shelf and positioned between the front and rear surfaces of the shelf.

[0406] A first cable, coupled to a first weight sensor and a circuit board, extends from the first weight sensor through a second and a third opening and then downwards into the space through the first opening; and

[0407] The second cable is coupled to the second weight sensor and the circuit board. The second cable extends from the second weight sensor through the second opening and the third opening and then enters the space downward through the first opening.

[0408] 44. The system as described in Clause 43 further includes a second shelf, the second shelf including a rear surface defining a fourth opening, the panel further defining a fifth opening extending along the width of the panel, the rear surface of the second shelf being coupled to the panel such that a portion of the fourth opening is aligned with a portion of the fifth opening.

[0409] 45. The system as described in Clause 43 further includes:

[0410] The third weight sensor is coupled to the top surface of the base; and

[0411] The third cable is coupled to the third weight sensor and the circuit board, and extends from the third weight sensor into the space through the first opening.

[0412] 46. ​​The system as described in Clause 43, wherein the first weight sensor comprises:

[0413] The first weighing unit is configured to generate a first current based on the force borne by the first weighing unit;

[0414] The second weighing unit is configured to generate a second current based on the force borne by the second weighing unit;

[0415] The third weighing unit is configured to generate a third current based on the force borne by the third weighing unit; and

[0416] The fourth weighing unit is configured to generate a fourth current based on the force borne by the fourth weighing unit.

[0417] 47. The system as described in Clause 46, wherein:

[0418] The first weight sensor includes a first corner, a second corner, a third corner, and a fourth corner;

[0419] The first weighing unit is positioned at a first distance from the first corner;

[0420] The second weighing unit is located at the second distance from the second corner;

[0421] The third weighing unit is located at a third distance from the third corner; and

[0422] The fourth weighing unit is located at a distance of four from the fourth corner. The first, second, third and fourth distances are basically the same.

[0423] 48. The system as described in Clause 47, wherein the first cable comprises:

[0424] The third cable is coupled to the first weighing unit;

[0425] The fourth cable is coupled to the second weighing unit;

[0426] The fifth cable is coupled to the third weighing unit; and

[0427] The sixth cable is coupled to the fourth weighing unit.

[0428] 49. The system as described in Clause 43, wherein:

[0429] The first weight sensor is configured to transmit a first signal via a first cable to a circuit board, indicating the weight of a shelf within a first region borne by the first weight sensor, the first region being positioned above the first weight sensor; and

[0430] The second weight sensor is configured to transmit a second signal via a second cable to the circuit board, indicating the weight of a second area of ​​the shelf that the second weight sensor is borne by, the second area being located above the second weight sensor.

[0431] 50. The system as described in Clause 49 further includes a weight server, the circuit board being configured to transmit a third signal to the weight server indicating the weight borne by the first weight sensor.

[0432] 51. The system as described in Clause 50, wherein the circuit board is further configured to:

[0433] If an Ethernet connection is established, then the third signal is transmitted to the weight server via the Ethernet connection;

[0434] If no Ethernet connection is established but a wireless connection is established, then a third signal is transmitted to the weight server via the wireless connection; and

[0435] If no Ethernet or wireless connection is established but a Universal Serial Bus (USB) connection is established, then a third signal is transmitted to the weight server via the USB connection.

[0436] 52. The system as described in Clause 50, wherein the weight server is configured to determine, based on a third signal, that an item in a first area of ​​the shelf has been removed.

[0437] 53. A system comprising:

[0438] Circuit boards; and

[0439] Shelves, including:

[0440] The base includes a drawer positioned within a space defined by the base, a circuit board positioned inside the drawer, and the base also defines a first opening for entering the space.

[0441] A panel, coupled to and extending upward from the base, defines a second opening extending along the width of the panel;

[0442] A shelf, coupled to a panel, is positioned vertically above the base and extends away from the panel. The shelf defines a third opening, a portion of which is aligned with a portion of a second opening.

[0443] The first weight sensor is coupled to the shelf;

[0444] A second weight sensor is coupled to the shelf;

[0445] A first cable, coupled to a first weight sensor and a circuit board, extends from the first weight sensor through a second and a third opening and then downwards into the space through the first opening; and

[0446] The second cable is coupled to the second weight sensor and the circuit board. The second cable extends from the second weight sensor through the second opening and the third opening and then enters the space downward through the first opening.

[0447] 54. The system as described in Clause 53 further includes a second shelf defining a fourth opening, the panel further defining a fifth opening extending along the width of the panel, the second shelf being coupled to the panel such that a portion of the fourth opening is aligned with a portion of the fifth opening.

[0448] 55. The system as described in Clause 53 further includes:

[0449] A third weight sensor is coupled to the base; and

[0450] The third cable is coupled to the third weight sensor and the circuit board, and extends from the third weight sensor into the space through the first opening.

[0451] 56. The system as described in Clause 53, wherein the first weight sensor comprises:

[0452] The first weighing unit is configured to generate a first current based on the force borne by the first weighing unit;

[0453] The second weighing unit is configured to generate a second current based on the force borne by the second weighing unit;

[0454] The third weighing unit is configured to generate a third current based on the force borne by the third weighing unit; and

[0455] The fourth weighing unit is configured to generate a fourth current based on the force borne by the fourth weighing unit.

[0456] 57. The system as described in Clause 56, wherein:

[0457] The first weight sensor includes a first corner, a second corner, a third corner, and a fourth corner;

[0458] The first weighing unit is positioned at a first distance from the first corner;

[0459] The second weighing unit is located at the second distance from the second corner;

[0460] The third weighing unit is located at a third distance from the third corner; and

[0461] The fourth weighing unit is located at a distance of four from the fourth corner. The first, second, third and fourth distances are basically the same.

[0462] 58. The system as described in Clause 56, wherein the first cable comprises:

[0463] The third cable is coupled to the first weighing unit;

[0464] The fourth cable is coupled to the second weighing unit;

[0465] The fifth cable is coupled to the third weighing unit; and

[0466] The sixth cable is coupled to the fourth weighing unit.

[0467] 59. The system as described in Clause 53, wherein:

[0468] The first weight sensor is configured to transmit a first signal via a first cable to a circuit board, indicating the weight of a shelf within a first region borne by the first weight sensor, the first region being positioned above the first weight sensor; and

[0469] The second weight sensor is configured to transmit a second signal via a second cable to the circuit board, indicating the weight of a second area of ​​the shelf that the second weight sensor is borne by, the second area being located above the second weight sensor.

[0470] 60. The system as described in Clause 59 further includes a weight server, the circuit board being configured to transmit a third signal to the weight server indicating the weight borne by the first weight sensor.

[0471] 61. The system as described in Clause 60, wherein the circuit board is further configured to:

[0472] If an Ethernet connection is established, then the third signal is transmitted to the weight server via the Ethernet connection;

[0473] If no Ethernet connection is established but a wireless connection is established, then a third signal is transmitted to the weight server via the wireless connection; and

[0474] If no Ethernet or wireless connection is established but a Universal Serial Bus (USB) connection is established, then a third signal is transmitted to the weight server via the USB connection.

[0475] 62. The system as described in Clause 60, wherein the weight server is configured to determine, based on a third signal, that an item in a first area of ​​the shelf has been removed.

Claims

1. A system for tracking physical location, comprising: An array of cameras, positioned above the space, with each camera in the array configured to capture video of a portion of the space containing people; The first camera client is configured as follows: Receive first multiple frames of a first video from the first camera in the camera array, each of the first multiple frames showing a person in space; For the first frame in the first plurality of frames: Determine the first defined region around the person shown in the first frame; as well as Generate the first timestamp of when the first frame is received by the first camera client; For the second frame in the first set of multiple frames: Determine the second defined area around the person shown in the second frame; as well as Generate a second timestamp indicating when the second frame is received by the first camera client; as well as For the third frame in the first set of multiple frames: Determine the third defined region surrounding the person shown in the third frame; as well as Generate a third timestamp indicating when the third frame was received by the first camera client; The second camera client is configured as follows: The second camera in the array of cameras receives a second plurality of frames of the second video, each of which shows a person in the space; For the fourth frame in the second set of frames: Determine the fourth defined region around the person shown in the fourth frame; and Generate a fourth timestamp indicating when the fourth frame was received by the second camera client; and For the fifth frame in the second set of frames: Determine the fifth defined region surrounding the person shown in the fifth frame; as well as Generate a fifth timestamp indicating when the fifth frame was received by the second camera client; as well as The camera server, separate from the first camera client and the second camera client, is configured as follows: Ensure the first timestamp falls within the first time window; In response to determining that the first timestamp falls within the first time window, the coordinates of the first defined region are assigned to the first time window; Ensure the second timestamp falls within the first time window; In response to determining that the second timestamp falls within the first time window, the coordinates defining the second bounding region are assigned to the first time window; Determine that the third timestamp falls within the second time window following the first time window; In response to determining that the third timestamp falls within the second time window, the coordinates defining the third bounding region are assigned to the second time window; Confirm that the fourth timestamp falls within the first time window; In response to determining that the fourth timestamp falls within the first time window, the coordinates defining the fourth bounding region are assigned to the first time window; Ensure the fifth timestamp falls within the second time window; In response to determining that the fifth timestamp falls within the second time window, the coordinates defining the fifth bounding region are assigned to the second time window; The coordinates assigned to the first time window are processed using the following operations: Based at least on the coordinates of the first defined region and the coordinates of the second defined region, the combined coordinates of the person during the first time window are calculated for the first video from the first camera; as well as Based at least on the coordinates of the fourth defined region, the combined coordinates of the person during the first time window are calculated for the second video from the second camera; After processing the coordinates assigned to the first time window, the coordinates assigned to the second time window are processed using the following operations: At least based on the coordinates of the defined third bounding region, the combined coordinates of the person during the second time window are calculated for the first video from the first camera; and Based at least on the coordinates of the fifth defined region, the combined coordinates of the person during the second time window are calculated for the second video from the second camera.

2. The system for tracking physical location as claimed in claim 1, wherein the camera server is further configured to: determine the location of a person in space during the first time window based at least on a combination of the person's coordinates during the first time window for a first video from a first camera and a combination of the person's coordinates during the first time window for a second video from a second camera.

3. The system for tracking physical location as claimed in claim 1, wherein the combined coordinates of a person during a first time window from a first video from a first camera include the average of coordinates defining a first defined region and coordinates defining a second defined region.

4. The system for tracking physical locations as claimed in claim 1, wherein processing the coordinates assigned to the first time window is performed in response to: Determining the coordinates assigned to the first time window includes the coordinates of frames from multiple cameras in the camera array; and Determine if the number of cameras exceeds the threshold.

5. The system for tracking physical location as claimed in claim 1, wherein processing the coordinates assigned to the first time window is performed in response to: Determining the coordinates assigned to the second time window includes the coordinates of frames from multiple cameras in the camera array; and Determine if the number of cameras exceeds the threshold.

6. The system for tracking physical locations as claimed in claim 1, wherein processing the coordinates assigned to the first time window is performed in response to determining that the coordinates assigned to the first time window have not been processed within a timeout period.

7. The system for tracking physical location as described in claim 6, wherein the camera server is further configured to: In response to the determination that the coordinates assigned to the first time window have not been processed within the timeout period: Determining the coordinates assigned to the second time window includes the coordinates of frames from a first number of cameras in the camera array; and Reduce the threshold to the first quantity; The coordinates assigned to the third time window include the coordinates of frames from a second number of cameras in the camera array; It is determined that the second quantity exceeds the first quantity; and In response to determining that the second quantity exceeds the first quantity, the threshold is increased to the second quantity.

8. The system for tracking physical location as claimed in claim 1, wherein determining the coordinates to be processed for the first time window includes determining that frames have been received from each camera in the array of cameras during a second time window.

9. The system for tracking physical location as claimed in claim 1, wherein the first camera client is further configured to: in response to determining that during a first time window the first camera client has received frames from each camera in the array of cameras connected to the first camera client, transmit the coordinates defining a first defined region and the coordinates defining a second defined region as a batch to the camera server.

10. The system for tracking physical location as claimed in claim 1, wherein: The first camera client is also configured to receive the height of the person shown in the first frame, the height of the person shown in the second frame, and the height of the person shown in the third frame; The second camera client is also configured to receive the height of the person shown in the fourth frame and the height of the person shown in the fifth frame; and The camera server is also configured as follows: Based at least on the height of the person shown in the first frame and the height of the person shown in the second frame, calculate the combined height of the person during the first time window for the first video from the first camera; and Based at least on the height of the person shown in the fourth frame, the combined height of the person during the first time window is calculated for the second video from the second camera.

11. The system for tracking physical location as claimed in claim 1, wherein: The first camera client implements a first clock for generating the first and second timestamps; The second camera client implements a second clock for generating the fourth and fifth timestamps; as well as The camera server implements a third clock, and the first, second, and third clocks are synchronized using a clock synchronization protocol.

12. The system for tracking physical location as claimed in claim 1, wherein the array of cameras is arranged in a grid such that: Each camera in the camera array whose communication is coupled to the first camera client is not directly adjacent to another camera in the same row or column of the grid whose communication is coupled to the first camera client; and The communication in the camera array is coupled to the camera of the first camera client in the grid, and the communication in the camera array is coupled to the other camera of the first camera client diagonally.

13. The system for tracking physical location as claimed in claim 1, wherein: The space also includes a second person, which is shown in each of the first and second frames; The first camera client is also configured as follows: Determine the sixth defined region surrounding the second person shown in the first frame; Determine the seventh defined area around the second person shown in the second frame; Determine the eighth defined region surrounding the second person shown in the third frame; The second camera client is also configured as follows: Determine the ninth defined region surrounding the second person shown in the fourth frame; Determine the tenth defined area around the second person shown in the fifth frame; and The camera server is also configured as follows: In response to determining that the first timestamp falls within the first time window, the coordinates of the sixth defined region are assigned to the first time window; In response to determining that the second timestamp falls within the first time window, the coordinates of the defined seventh bounding region are assigned to the first time window; In response to determining that the third timestamp falls within the second time window, the coordinates defining the eighth bounding region are assigned to the second time window; In response to determining that the fourth timestamp falls within the first time window, the coordinates of the ninth defined region are assigned to the first time window; In response to determining that the fifth timestamp falls within the second time window, the coordinates defining the tenth bounding region are assigned to the second time window; The coordinates assigned to the first time window are further processed using the following steps: Based at least on the coordinates of the sixth and seventh defined regions, the combined coordinates of the second person during the first time window are calculated for the first video from the first camera; and Based at least on the coordinates of the ninth defined region, the combined coordinates of the second person during the first time window are calculated for the second video from the second camera; The coordinates assigned to the second time window are further processed using the following steps: Based at least on the coordinates of the defined eighth bounding region, the combined coordinates of the second person during the second time window are calculated for the first video from the first camera; and Based at least on the coordinates of the tenth defined region, the combined coordinates of the second person during the second time window are calculated for the second video from the second camera.

14. A method for tracking physical location, comprising: The first camera client receives the first plurality of frames of the first video from the first camera in the camera array, the camera array being positioned above space, each camera in the camera array being configured to capture video of a portion of the space containing people, each of the first plurality of frames showing people in the space; For the first frame in the first plurality of frames: The first camera client determines the first defined area around the person shown in the first frame; as well as The first timestamp of when the first frame was received by the first camera client is generated by the first camera client; For the second frame in the first set of multiple frames: The second defined area around the person shown in the second frame is determined by the first camera client; as well as A second timestamp is generated by the first camera client to indicate when the second frame was received by the first camera client. For the third frame in the first set of multiple frames: The third defined region around the person shown in the third frame is determined by the first camera client; as well as A third timestamp generated by the first camera client to indicate when the third frame was received by the first camera client; A second camera client, separate from the first camera client, receives a second plurality of frames of the second video from the second camera in the array of cameras, each of the second plurality of frames showing a person in the space; For the fourth frame in the second set of frames: The fourth defined area around the person shown in the fourth frame is determined by the second camera client; as well as A fourth timestamp generated by the second camera client to indicate when the fourth frame was received by the second camera client; For the fifth frame in the second set of frames: The fifth defined area around the person shown in the fifth frame is determined by the second camera client; as well as The fifth timestamp, generated by the second camera client, indicates when the fifth frame was received by the second camera client. The camera server, which is separate from the first camera client and the second camera client, determines whether the first timestamp falls within the first time window; In response to determining that the first timestamp falls within the first time window, the camera server assigns the coordinates of the first defined region to the first time window. The camera server determines whether the second timestamp falls within the first time window; In response to the determination that the second timestamp falls within the first time window, the camera server assigns the coordinates of the defined second bounding region to the first time window. The camera server determines that the third timestamp falls within the second time window following the first time window. In response to the determination that the third timestamp falls within the second time window, the camera server assigns the coordinates of the third defined region to the second time window. The camera server determines that the fourth timestamp falls within the first time window; In response to the determination that the fourth timestamp falls within the first time window, the camera server assigns the coordinates of the fourth defined region to the first time window. The camera server determines that the fifth timestamp falls within the second time window; In response to the determination that the fifth timestamp falls within the second time window, the camera server assigns the coordinates of the fifth defined region to the second time window. The coordinates assigned to the first time window are processed by the camera server through the following operations: The camera server calculates the combined coordinates of a person during a first time window for the first video from the first camera, based at least on the coordinates of a first defined region and a second defined region. as well as The camera server calculates the combined coordinates of the person during the first time window based on the coordinates of the second video from the second camera, at least based on the coordinates of the fourth defined region. After processing the coordinates assigned to the first time window, the coordinates assigned to the second time window are processed using the following operations: The camera server calculates the combined coordinates of a person during a second time window based on the coordinates of a first video from the first camera, at least based on the coordinates of a third defined region. as well as The camera server calculates the combined coordinates of the person during the second time window based on the coordinates of the second video from the second camera, at least based on the coordinates of the fifth defined region.

15. The method for tracking physical location as claimed in claim 14, further comprising determining, by a camera server, the location of a person in space during the first time window based at least on a combination of coordinates of a person during the first time window for a first video from a first camera and a combination of coordinates of a person during the first time window for a second video from a second camera.

16. The method for tracking physical location as claimed in claim 14, wherein the combined coordinates of a person during a first time window from a first video from a first camera include the average of coordinates defining a first defined region and coordinates defining a second defined region.

17. The method for tracking physical location as described in claim 14, wherein processing the coordinates assigned to the first time window is performed in response to: Determining the coordinates assigned to the first time window includes the coordinates of frames from multiple cameras in the camera array; and Determine if the number of cameras exceeds the threshold.

18. The method for tracking physical location as described in claim 14, wherein processing the coordinates assigned to the first time window is performed in response to: Determining the coordinates assigned to the second time window includes the coordinates of frames from multiple cameras in the camera array; and Determine if the number of cameras exceeds the threshold.

19. The method for tracking physical location as claimed in claim 14, wherein processing the coordinates assigned to the first time window is performed in response to determining that the coordinates assigned to the first time window have not been processed within a timeout period.

20. The method for tracking physical location as described in claim 19, further comprising: In response to the determination that the coordinates assigned to the first time window have not been processed within the timeout period: Determining the coordinates assigned to the second time window includes the coordinates of frames from a first number of cameras in the camera array; and Reduce the threshold to the first quantity; The coordinates assigned to the third time window include the coordinates of frames from a second number of cameras in the camera array; It is determined that the second quantity exceeds the first quantity; as well as In response to determining that the second quantity exceeds the first quantity, the threshold is increased to the second quantity.

21. The method for tracking physical location as claimed in claim 14, wherein determining the coordinates to be processed for the first time window includes determining that frames have been received from each camera in the array of cameras during the second time window.

22. The method for tracking physical location as claimed in claim 14, further comprising the first camera client, in response to determining that during a first time window the first camera client has received frames from each camera in the array of cameras connected to the first camera client, transmitting the coordinates defining a first defined region and the coordinates defining a second defined region as a batch to the camera server.

23. The method for tracking physical location as described in claim 14, further comprising: The height of the person shown in the first frame, the height of the person shown in the second frame, and the height of the person shown in the third frame are received by the first camera client. The height of the person shown in the fourth frame and the height of the person shown in the fifth frame are received by the second camera client; The camera server calculates a combined height of the person during a first time window based at least on the height of the person shown in the first frame and the height of the person shown in the second frame, using the first video from the first camera; and The camera server calculates the combined height of the person during the first time window based at least on the height of the person shown in the fourth frame for the second video from the second camera.

24. The method for tracking physical location as described in claim 14, further comprising: A first clock, implemented by the first camera client, is used to generate the first and second timestamps. The second clock, used to generate the fourth and fifth timestamps, is implemented by the second camera client. as well as The third clock is implemented by the camera server, and the first, second and third clocks are synchronized using a clock synchronization protocol.

25. The method for tracking physical location as described in claim 14, wherein the array of cameras is arranged in a grid such that: Each camera that is communicatively coupled to the first camera client is not directly adjacent to another camera in the grid that is also communicatively coupled to the first camera client; and The camera that is communication-coupled to the first camera client is diagonally opposite another camera that is communication-coupled to the first camera client in the grid.

26. The method for tracking physical location as described in claim 14, further comprising: The sixth defined region surrounding the second person shown in the first frame is determined by the first camera client; The seventh defined area around the second person shown in the second frame is determined by the first camera client; The eighth defined region surrounding the second person shown in the third frame is determined by the first camera client; The ninth defined area around the second person shown in the fourth frame is determined by the second camera client; The tenth defined area around the second person shown in the fifth frame is determined by the second camera client; In response to the determination that the first timestamp falls within the first time window, the camera server assigns the coordinates of the sixth defined region to the first time window. In response to the determination that the second timestamp falls within the first time window, the camera server assigns the coordinates of the seventh defined region to the first time window. In response to the determination that the third timestamp falls within the second time window, the camera server assigns the coordinates of the defined eighth bounding region to the second time window. In response to the determination that the fourth timestamp falls within the first time window, the camera server assigns the coordinates of the ninth defined region to the first time window. In response to the determination that the fifth timestamp falls within the second time window, the camera server assigns the coordinates of the tenth defined region to the second time window. The coordinates assigned to the first time window are further processed using the following steps: The camera server calculates the combined coordinates of the second person during the first time window based on at least the coordinates of the sixth and seventh defined regions of the first video from the first camera. as well as The camera server calculates the combined coordinates of the second person during the first time window based at least on the coordinates of the ninth defined region for the second video from the second camera; and The coordinates assigned to the second time window are further processed using the following steps: The camera server calculates the combined coordinates of the second person during the second time window based at least on the coordinates of the first video from the first camera, using the coordinates of the first video from the first camera, defined by the eighth bounding region. as well as The camera server calculates the combined coordinates of the second person during the second time window based on the coordinates of the defined tenth bounded area for the second video from the second camera.

Citation Information

Patent Citations

  • Self-service vending method, device and system, electronic equipment and computer readable medium

    CN109726759A

  • Crime prevention support system

    JP2005347905A