Merchandise surveillance system and method

The automated order confirmation system addresses the inefficiencies of manual monitoring by using video sensors and database-driven image detection to ensure accurate and automated verification of goods and personnel at receiving/shipping portals, enhancing inventory control and loss prevention.

JP7778247B2Active Publication Date: 2025-12-01EVERSEEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024546150
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-02
Filing Date
2023-01-31
Publication Date
2025-12-01
Estimated Expiration
2043-01-31

AI Technical Summary

Technical Problem

Manual monitoring of incoming and outgoing goods at receiving/shipping portals is tedious and inconsistent due to high traffic volumes, making it difficult to prevent unauthorized entry and ensure accurate receipt and dispatch of goods.

Method used

An automated order confirmation system using video sensors and a processing unit for event analysis, combined with a database for face and product image detection, to monitor and verify the arrival and departure of goods and personnel, ensuring accurate order confirmation.

Benefits of technology

The system provides automated and accurate monitoring of goods and personnel, reducing manual intervention and improving inventory control by detecting discrepancies and preventing loss at receiving/shipping portals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007778247000014
    Figure 0007778247000014
  • Figure 0007778247000015
    Figure 0007778247000015
  • Figure 0007778247000016
    Figure 0007778247000016
Patent Text Reader

Abstract

The order confirmation system includes a video sensor configured to capture video footage of a monitored area located in proximity to the receiving / shipping portal. The processing unit performs event analysis on the captured video footage to detect entities, detect arrival of goods from a third-party supplier from a door opening event, identify the third-party supplier, perform a check-in process for a delivery person associated therewith, detect entry and exit of goods through the receiving / shipping portal, and verify that the detected delivered products match data regarding products to be delivered by the third-party supplier. The database stores at least a dataset of face images / logos for face / brand detection and a dataset of product images for product identification. The database records the results of the order confirmation process and the check-out of the delivery person at the end of the delivery.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE DISCLOSURE The present disclosure relates generally to systems and methods for monitoring, and more particularly, to systems and methods for monitoring incoming / outgoing goods at a receiving / shipping portal. [Background technology]

[0002] Environments, such as retail environments equipped with warehouses and the like, may facilitate the entry and exit of people and goods, as well as the storage and retrieval of goods. Often, the high traffic of people and goods passing through these environments can pose challenges in tracking the individual movements of people and goods in such environments to prevent, for example, unauthorized entry, theft of goods, and proper receipt and dispatch of goods. Traditionally, these efforts may have been performed manually by deploying security personnel. However, given the high traffic of people and goods, such manual work by security personnel is tedious and inconsistent. Therefore, there is a need for a more robust system and method that can perform monitoring of people and goods entering and exiting such environments without requiring manual intervention. Summary of the Invention

[0003] In one aspect of the present disclosure, an order confirmation system is provided. The order confirmation system includes a plurality of video sensors adapted to capture video footage of a monitored area located within an order receiving or shipping area of ​​a receiving / shipping portal. The order confirmation system further includes a processing unit configured to perform event analysis on the captured video footage, detect entities within the video footage captured by the video sensors, detect arrivals from third-party suppliers from door-opening events within the captured video footage, identify the third-party suppliers, perform a check-in process for delivery personnel from the third-party suppliers, detect goods entering and leaving the receiving / shipping portal, and verify that the detected delivery products match data regarding the products to be delivered by the third-party suppliers. The order confirmation system also includes a database communicatively coupled to the processing unit. The database is configured to store at least a dataset of face images / logos for use in face / brand detection and a dataset of product images for use in product identification. The database is also configured to record results of the order confirmation process at the end of delivery and the delivery personnel's checkout for future retrieval upon request to the processing unit.

[0004] In another aspect of the present disclosure, a method for performing video surveillance is provided. The method includes using a plurality of video sensors to capture video footage of a monitored area located within an order receiving area or an order shipping area of ​​a receiving / shipping portal. The method further includes using a processing unit to perform event analysis on the captured video footage. The method further includes using the processing unit to detect entities in the video footage captured by the video sensors. The method further includes using the processing unit to detect an arrival of a package from a third-party supplier from a door-opening event in the captured video footage. The method further includes using the processing unit to identify the third-party supplier and perform a check-in process for a delivery person from the third-party supplier. The method further includes using the processing unit to detect the entry and exit of goods through the receiving / shipping portal and verify that the detected delivery product matches data regarding the product to be delivered by the third-party supplier. The method further includes using a database to store at least a dataset of face images / logos for use in face / brand detection and a dataset of product images for use in product identification. The method further includes using the database to record results of the order confirmation process and delivery person checkout at the end of the delivery. The method further includes retrieving the record by the processing unit from the database in response to a request to the processing unit.

[0005] In yet another aspect of the present disclosure, the embodiments disclosed herein are also directed to a non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed by a processing unit, cause the processing unit to perform the methods disclosed herein.

[0006] The present disclosure presents a system and method for monitoring incoming / outgoing goods at a receiving / shipping portal. The disclosure is described with reference to a retail environment. However, those skilled in the art will understand that the disclosure is not limited to use in a retail environment. Rather, the disclosure is applicable to any environment in which goods pass through a shipping portal as they leave a first order processing facility and then pass through a receiving portal as they enter a second order receiving facility, which is the goods' desired destination. Thus, the purpose of the disclosed system is to detect discrepancies between the planned inventory records of an incoming / outgoing order and the actual contents of the corresponding incoming / pre-shipped order.

[0007] Additionally, the system addresses the problem of detecting inaccurate or incomplete orders and similarly inaccurate or incomplete assembled orders before dispatch. In this manner, merchandise traffic at both ends, i.e., the shipping portal and the receiving portal of the delivery system, can be characterized and controlled to improve both the accuracy of the delivery system and the inventory control process at both the order processing facility and the order receiving facility.

[0008] Incoming merchandise orders typically pass through an entry door of an order receiving facility before being accepted and received by staff at the order receiving facility. Similarly, outgoing merchandise from an order fulfillment facility typically passes through an exit door of an order fulfillment facility before being delivered to its required destination. For simplicity, the entry door of an order receiving facility will hereafter be referred to as a receiving portal. Similarly, the exit door of an order fulfillment facility will hereafter be referred to as a shipping portal. The present disclosure addresses the problem of loss prevention at receiving / shipping portals. In particular, the present disclosure enables automated confirmation of inbound / outbound merchandise orders at a receiving / shipping portal.

[0009] In practice, order receiving facilities often include multiple receiving portals. Similarly, order fulfillment facilities often include multiple shipping portals. In fact, a given facility may undertake both order receiving and order fulfillment, in which case the facility may include a first plurality of receiving portals and a second plurality of shipping portals. The time required to manually review each order and the high volume of inbound and / or outbound order traffic typically experienced at an order receiving facility / order fulfillment facility make manual monitoring of inbound and outbound orders very difficult. The challenge is amplified when multiple deliveries occur simultaneously within the limited space of the order receiving or order fulfillment area of ​​a receiving / fulfillment portal.

[0010] Accordingly, the present disclosure discloses a system and method for automated monitoring of inbound and outbound traffic from either or both receiving and shipping portals of an order receiving facility and an order processing facility, respectively. For brevity, the system and method of the present disclosure will hereinafter be referred to as an order confirmation system and an order confirmation method, respectively.

[0011] It will be understood that features of the present disclosure are capable of being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims. [Brief explanation of the drawings]

[0012] The foregoing summary, as well as the following detailed description of exemplary embodiments, will be better understood when read in conjunction with the accompanying drawings. For the purpose of illustrating the disclosure, example structures of the disclosure are shown in the drawings. However, the disclosure is not limited to the particular methods and instrumentalities disclosed herein. Moreover, those skilled in the art will appreciate that the drawings are not to scale. Wherever possible, like elements have been designated by like numerals. [Figure 1] FIG. 1 is a perspective view of an exemplary environment in which an order confirmation system may execute, according to one embodiment of the present disclosure. [Figure 2]FIG. 2 is a schematic top-down view of a monitored area in the exemplary environment of FIG. 1, the monitored area being located proximate to a receiving / shipping portal and monitored by a video sensor of the order confirmation system of FIG. 1. [Figure 3] FIG. 3 is a schematic diagram of an order confirmation system according to one embodiment of the present disclosure. [Figure 4] FIG. 4 is a block diagram illustrating the software architecture of an order confirmation system according to one embodiment of the present disclosure. [Figure 5] FIG. 5 is a diagram of an exemplary camera configuration for generating training data and subsequent detection of doors, including people and door states, according to one embodiment of the present disclosure. [Figure 6] FIG. 6 illustrates the detection of receiving / shipping portals having various degrees of closure / openness corresponding to closed, intermediate, and fully open states according to one embodiment of the present disclosure. [Figure 7] FIG. 7 illustrates an exemplary threshold value for the height of a bounding box surrounding a receiving / shipping portal that can be used to determine whether the receiving / shipping portal is open or closed, according to one embodiment of the present disclosure. [Figure 8] FIG. 8 is a diagram illustrating an example pair of consecutive video frames from a portion of video footage, according to one embodiment of the present disclosure. [Figure 9] FIG. 9 illustrates an exemplary Yolo v5 architecture that can be used to implement a palette detector, according to one embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates a flowchart for keypoint detection according to one embodiment of the present disclosure. [Figure 11(a)] FIG. 11(a) is a virtual representation of a physical grid pattern marked on a ground area, according to one embodiment of the present disclosure. [Figure 11(b)] FIG. 11(b) is a virtual representation of the points from FIG. 11(a) projected by a camera according to one embodiment of the present disclosure. [Figure 12(a)]FIG. 12(a) is a diagram illustrating an exemplary cube with opposing corner points T' and B' according to one embodiment of the present disclosure. [Figure 12(b)] FIG. 12(b) is a virtual representation of the projections Tp and B corresponding to the corner points T' and B' taken from the view of FIG. 12(a), according to one embodiment of the present disclosure. [Figure 13] FIG. 13 is a virtual representation of an object as seen by a camera placed above the object, according to one embodiment of the present disclosure. [Figure 14] FIG. 14 is a cross-sectional view of the pallet. [Figure 15] FIG. 15 is a cross-sectional view of a pallet with two boxes stacked on top of each other, the bottom box being longer than the top box.

[0013] In the accompanying drawings, underlined numbers are employed to represent the item in which the underlined number is located or to which it is adjacent. Numbers without underlines refer to items identified by a line linking the ununderlined number to the item. When a number is accompanied by an associated arrow without an underline, the ununderlined number is used to identify the general item toward which the arrow is pointing. DETAILED DESCRIPTION OF THE INVENTION

[0014] The following detailed description sets forth embodiments of the present disclosure and how they may be practiced. While best modes of carrying out the disclosure are disclosed, those skilled in the art will recognize that there are other possible embodiments for carrying out or practicing the disclosure.

[0015] 1 and 2, order confirmation system 110 includes a plurality of video sensors 102 adapted to capture video footage of a monitored area located near receiving / shipping portal 101. In one embodiment, video sensors 102 may be attached to a frame 101a of receiving / shipping portal 101. In another embodiment, video sensors 102 may be positioned proximate receiving / shipping portal 101 such that the field of view (FOV) of video sensors 102 includes at least the monitored area having one or both of the center and sides of receiving / shipping portal 101, i.e., the area adjacent to the receiving / shipping portal.

[0016] Thus, the monitored area is located within an order receiving area or an order shipping area. The monitored area is formed from the collective field of view (FOV) of the video sensors 102 and is bounded by the receiving / shipping portal 101 and a region of interest (ROI) 204 (also known as a pallet analysis zone). In an embodiment of the present disclosure, the monitored area also comprises a first buffer zone 202 and a second buffer zone 203 that are used together as a hysteresis determination function to eliminate uncertainty in determining whether a mobile entity is located outside or inside a corresponding order receiving area or order shipping area.

[0017] In the illustrated example, the exterior monitoring zone 201 located to the left of the receiving / shipping portal 101 is considered to be outside the order receiving or order shipping area. Similarly, the area to the right of the receiving / shipping portal 101 is considered to be within the order receiving or order shipping area. Thus, in this example, the presence of an item approaching the order receiving area may be detected in the exterior monitoring zone 201. Furthermore, when an item is detected within the first buffer zone 202 and immediately thereafter, i.e., consecutively detected within the second buffer zone 203, the item is considered to have passed through the receiving / shipping portal 101 and entered the order receiving or order shipping area. Similarly, when an item is detected within the second buffer zone 203 and immediately thereafter, i.e., consecutively detected within the first buffer zone 202, the item is considered to have exited the order receiving or order shipping area.

[0018] The peripheral portion of the order receiving or order shipping area located near the receiving / shipping portal 101 is represented by the internal lingering zone 205. The internal lingering zone 205 is not within the field of view of the video sensor 102 of the order confirmation system 110. Similarly, using the naming protocol of this example, the internal lingering zone 205 is within the order receiving or order shipping area. Thus, the internal lingering zone 205 is an unmonitored area within the order receiving or order shipping area. Accordingly, the order confirmation system 110 of the present disclosure monitors the movement of entities (people, pallets) within the external monitoring zone 201, the first buffer zone 202 and the second buffer zone 203, and the region of interest (ROI) / pallet analysis zone 204; therefore, these aforementioned zones 201, 202, 203, and 204 may collectively be considered a monitoring area for the purposes of brevity in this disclosure.

[0019] Referring to FIG. 2, in a top-down view of the monitored area, the region of interest (ROI) / pallet analysis zone 204 and the first and second buffer zones 202 and 203 are defined as a set of rectangles. Those skilled in the art will understand that from a perspective view of the camera and corresponding to the configuration presented in the diagram of FIG. 1, each of the region of interest (ROI) / pallet analysis zone 204 and the first and second buffer zones 202 and 203 are formed as trapezoids. Returning to the plan view of the monitored area shown in FIG. 2, the vertical dimension of the region of interest (ROI) / pallet analysis zone 204 and the first and second buffer zones 202 and 203 is equal to the vertical dimension of the receiving / shipping portal 101. Similarly, the first buffer zone 202 and the second buffer zone 203 have horizontal dimensions equal to the horizontal dimension of the receiving / shipping portal 101. In one example, the horizontal dimension of the region of interest (ROI) / pallet analysis zone 204 is configured to be three times the horizontal dimension of a standard merchandise pallet. However, those skilled in the art will recognize that the above relationships between the horizontal dimensions of a region of interest (ROI) / pallet analysis and the horizontal dimensions of a standard merchandise pallet are exemplary in nature and are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system of the present disclosure is not limited to the above dimensional relationships. Rather, preferred embodiments are operable with any horizontal dimensions associated with an inlet / outlet channel of an order receiving facility / order fulfillment facility.

[0020] 1 and 2, in one example, an entity 103 (e.g., a pallet of goods, shown below using the same numeral "103") traverses the receiving / shipping portal 101 in a left-to-right direction (moving from the external monitoring zone 201 to the internal remaining zone 205). The entity is considered to have entered the first buffer zone 202 when two conditions are met: (a) the entity is moving through the first buffer zone 202; (b) A box substantially, eg, digitally, representing the outline of the entity detected by the video sensor 102 of the order confirmation system 110 intersects the second buffer zone 203 .

[0021] For simplicity, the box that substantially outlines the detected entity will hereafter be referred to as the bounding box around the entity.

[0022] Referring to FIG. 3, the architecture of the order confirmation system 110 of the present disclosure includes: (a) a receiving / shipping portal with a plurality of video sensors 102; (b) a processing unit 301; (c) a database 302.

[0023] If an order receiving / order fulfillment facility has multiple receiving / shipping portals 101, a separate instance of the order confirmation system 110 may be dedicated to each receiving / shipping portal 101. In such an embodiment, some components may be shared between the separate instances of the order confirmation system 110, including, among others, the video sensor 102, the software detector component associated with the processing unit 301, and the database 302 of the order confirmation system 110.

[0024] The processing unit 301 comprises one or more CPUs, main memory, and local storage. The processing unit 301 is configured to run algorithms that detect entities, such as objects or people, in video footage captured by the video sensor 102. The processing unit 301 is also configured to run algorithms that perform event analysis on the captured video footage. These algorithms are executed by a set of software detector components, described later in this specification and that form part of the order confirmation system 110. The database 302 stores information necessary to implement the algorithm execution / detector functionality, as described below. Specifically, in various embodiments, the information stored in the database 302 includes the following: A dataset of face images / logos required for face / brand detectors, A dataset of product images required for product re-identification (run through the Pallet-by-Pallet Product Classification module described below).

[0025] The order confirmation system 110 of the present disclosure facilitates automated monitoring of an order receiving or shipping area by covering various aspects such as: -Detecting arrival of goods from third-party suppliers from door-open events; - Identifying third-party suppliers and conducting a check-in process for delivery personnel from third-party suppliers; - Correlating data on products actually received with those that suppliers should deliver, e.g., advance shipping notices; -When detecting the arrival and departure of goods via the receiving / shipping portal, verify that the detected delivered products match the data that should be delivered by the third-party supplier; -Ensure the results of the order confirmation process (approval / rejection of received orders) and the delivery person's check-out at the end of the delivery.

[0026] Referring to FIG. 4, the software architecture of the order confirmation system 110 comprises three main software modules: a delivery detection module 402 , a pallet monitor module 404 , and an event (or alert) management module 406 .

[0027] The delivery detection module 402 is responsible for verifying whether the receiving / delivery process is performed correctly. The delivery detection module 402 includes a door status detector 402a, a person detector 402b, a person tracker 402c, and a quick response (QR) detector 402d. Based on an analysis of video footage captured by the video sensor 102 of the order confirmation system 110, the door status detector 402a determines whether the receiving / shipping portal 101 is in an open or closed state. The person detector 402b analyzes the video footage captured by the video sensor 102 of the order confirmation system 110 to detect whether a delivery person has arrived at the receiving / shipping portal 101. Using the same video footage, the person tracker 402c tracks the movement of the delivery person detected by the person detector 402b. The QR detector 402d detects the presence of a quick response (QR) code in the captured video footage and reads the QR code. The QR detector 402d compares the detected QR code with known pre-approved QR codes for the third-party supplier / delivery person and finds a match. If a match is found, the person presenting the QR code is classified (i.e., the person is deemed by the QR detector 402d) as an authorized entrant to the order fulfillment / order receipt facility. Thus, if the detected person's movements are tracked by the person tracker 402c and the person is deemed by the QR detector 402d to be an authorized entrant, the delivery detection module 402 grants the person access to the order fulfillment / order receipt facility and performs activities pursuant to performing the associated delivery.

[0028] The pallet monitor module 404 is dedicated to verifying the contents of goods being delivered from or received into a premises, such as an order fulfillment facility / order receiving facility. To this end, the first step is to detect pallets using a pallet detector module 404a, then track the detected pallets using a pallet tracker module 404b, and finally classify the goods on the detected pallets using a per-pallet goods classification module 404c. The pallet monitor module 404 also includes a pallet volume estimator 404d for the purpose of estimating the quantity of goods on a pallet. The final component of the pallet monitor module 404 is an in / out counter 404e, which is used to extract information regarding the total number of pallets passing through the receiving / shipping portal.

[0029] The event (or alert) management module 406 comprises an alert manager 406a and an event recorder 406b. The alert manager 406a is configured to issue alerts regarding the detection of authorized entrants and information regarding the goods being supplied / delivered, such as the goods class, the volume of the pallet, the amount of pallets received during the supply / delivery episode in question, etc. The event recorder 406b is configured to record the entire supply / delivery episode. The above software components are described in more detail below.

[0030] The person detector 402b comprises a model used to detect the presence of authorized entrants, including, but not limited to, delivery personnel from third-party suppliers and / or employees of the order fulfillment / order receiving facility, within video footage captured by the video sensor 102 of the order confirmation system 110. The output from the person detector 402b is also processed by a person re-identification model (not shown) to track persons whose presence is detected within the captured video footage. The door status detector 402a comprises a model used to detect the presence of the receiving / shipping portal 101 within the captured video footage and a door status algorithm for determining whether the receiving / shipping portal is open or closed.

[0031] In one embodiment, the person detector 402b and the door status detector 402a may be combined into a software component. In this embodiment, a neural network based on the YOLO v5 architecture is used to detect people and doors (i.e., the receiving / shipping portal 101). The selected architecture is version M, which adds feature pyramid level P6 to the neck components of the original version. CSPDarknet53 (as described in C.-Y. Wang, H.-YM Liao, Y.-H. Wu, P.-Y. Chen, J.-W. Hsieh, and I.-H. Yeh, "CSPNet: A New Backbone that can Enhance Learning Capability of CNN," 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2020, pp. 1571-1580) is the backbone of YOLO v5, used as a feature extractor. The neck is represented by a PANet (as described in S. Liu, L. Qi, H. Qin, J. Shi and J Jia, Path Aggregation for Instance Segmentation, 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2018, 8759-8768) to generate a feature pyramid to help the model generalize at different scales. The head is used for final detection by generating anchor boxes and corresponding output vectors.

[0032] However, those skilled in the art will recognize that the above-described neural networks and architectures are exemplary in nature and are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system 110 of the present disclosure is not limited to the above-described neural networks and architectures. Rather, the present disclosure can be implemented with any neural network and architecture capable of detecting people and objects, such as doors, within captured video footage. For example, the person detector 402b and the door status detector 402a may include a YOLO v5 architecture having an S-architecture or an L-architecture. Similarly, the person detector 402b and the door status detector 402a may comprise any single-shot detector (SSD), such as RetinaNet, or alternatively, may embody other types of neural networks and architectures known to those skilled in the art.

[0033] Further, in this embodiment, the person detector 402b and the door status detector 402a (or a combination of the person and door status detectors) are trained on a dataset whose labels are door, employee, and delivery person. An exemplary camera setup for generating the training data and subsequent person and door status detection is shown in FIG. 5. Those skilled in the art will understand that the order confirmation system 110 disclosed herein is not limited to the camera positions shown in FIG. 5. In particular, cameras 501, 502, and 503 can be moved 5-10 cm in any direction from the positions shown in FIG. 5. Cameras 504 and 505 should be positioned so that the bottom of the receiving / shipping portal 101 is completely within the camera's field of view and the entire pallet can be viewed, regardless of whether the pallet passes through the receiving / shipping portal 101 on its right, left, or center.

[0034] Video footage of people in the training data is labeled according to that person's clothing. The labels assume that each operator at an order fulfillment or order receiving facility has a standard uniform that all employees must wear, and that the uniform is easily distinguishable from clothing worn by non-employees. The labeling protocol also addresses the degree of uniform variation; for example, all employees may wear brown shirts but have non-standard dark pants (e.g., brown, black). Thus, using this approach, everyone wearing an operator uniform would be labeled as an "employee" and everyone else would be labeled as a "delivery person."

[0035] Exemplary details of the datasets used for training, validation, and testing (after being split into training, validation, and test sets) of the order confirmation system 110 are as follows: Image size: minimum 1920x1080 pixels Number of annotated images: 18311 Number of different cameras (at least two viewpoints as shown in Figure 5) Total number of bounding boxes surrounding objects in video frames in the dataset: 59276 The number of bounding boxes surrounding objects of a particular class in a video frame of the dataset: Receiving / Shipping Portal: 16610 Delivery Person: 19461 Employees: 23205 During inference, the people detector 402b and the door state detector 402a (or a combination of people and door state detectors) receive as input images comprising video frames from video footage captured by the video sensor 102, e.g., cameras 501-505 of the order confirmation system 110. In response to the received images, the people detector 402b (or a combination of people and door state detectors) outputs a 3D tensor comprising: (a) the coordinates of the center of a bounding box encompassing a person or door detected in the received image; (b) a width and height of the bounding box, the width and height being normalized by scaling them relative to the width and height of the received image, respectively; and (c) an objectness score, between 0 and 1, indicating the neural network's confidence that an object center exists at a given location in the received image; and (d) Two output class predictions: “Employee” and “Deliveryman.”

[0036] A non-maximum suppression algorithm is used to generate the prediction with the highest confidence score from several overlapping bounding box proposals for the same person. The term "prediction" refers to the classifications of "employee" and "deliveryman" and the corresponding locations in the received image of the people so classified.

[0037] In addition to detecting people, the delivery detection module 402 also detects the presence of a receiving / shipping portal 101 in a received video frame and determines whether the receiving / shipping portal 101 is in an open or closed state. To this end, the YOLO v5 network of the door status detector 402a (or a combination of the person and door status detector) generates an output classification of “door” upon detecting the presence of a receiving / shipping portal 101 in a received video frame. For a received video frame in which a receiving / shipping portal 101 is detected, a further output from the YOLO v5 network is a set of coordinates from which the height of a bounding box surrounding the detected receiving / shipping portal 101 can be calculated. Referring to FIG. 6 , the bounding box heights, e.g., H1, H2, and H3, are used to determine the state of the receiving / shipping portal 101. Specifically, the receiving / shipping portal 101 is determined to be either open or closed.

[0038] When the receiving / shipping portal 101 is closed, the height of the bounding box surrounding it has a maximum value. In contrast, ideally, when the receiving / shipping portal 101 is open, the height of the bounding box surrounding it is evaluated at 0, since the receiving / shipping portal 101 is no longer visible in the received video frames.

[0039] However, there may be cases where the receiving / shipping portal 101 is not fully open. In this case, to avoid classifying the receiving / shipping portal 101 as closed, the threshold variable may be preset to a threshold value by an operator. The threshold value may be empirically determined according to the environment and deployment in which the order confirmation system 110 is used. Referring to FIG. 7, if the height of the bounding box surrounding the receiving / shipping portal 101 is less than or equal to the threshold value HT, the receiving / shipping portal 101 is considered open. Otherwise, if the height value exceeds the threshold value HT, the receiving / shipping portal 101 is considered closed. Using this approach, it is recognized that the height H3 of the bounding box in FIG. 6 is less than or equal to the threshold value HT.

[0040] Returning to FIG. 4, the person tracker 402c is used to track the track path T ID, and assigns a unique ID to every person detected within an item of video footage and keeps a record of the unique IDs. In one embodiment, the person tracker 402c performs tracking using a detection algorithm based on the DeepSort algorithm (as described in Wojke N., A. Bewley A. and Paulus D., “Simple online and realtime tracking with a deep association metric,” 2017 IEEE International Conference on Image Processing (ICIP), Beijing, 2017, pp. 3645-3649). Specifically, the person tracker 402c uses the person detector 402b to establish bounding boxes around every person detected within every image of the captured video footage. A unique ID is assigned to each detected person in association with these bounding boxes. Track path T ID ={(x1,y1),(x2,y2),…} represents the vector of spatial coordinates of the center of the bounding box corresponding to the person ID, stored in the order in which the bounding boxes are established in consecutive video frames. The position is expressed in pixels in the frame coordinate system, whose origin is located at the top-upper corner of the video frame, with the OX axis running horizontally from left to right of the origin and the OY axis running vertically from top to bottom of the video frame.

[0041] 8 , five people are detected in a first video frame captured at time T0, and five bounding boxes are established around the detected people in the first video frame. The five bounding boxes are assigned unique IDs, namely, 180, 129, 159, 165, and 137, respectively. Immediately after the first video frame, in a second video frame captured at time T1, the same five people are visible, with the person with ID 137 partially obscured. Five bounding boxes are established around the people in the second video frame. Each bounding box is assigned a unique ID that corresponds to the location of the bounding box surrounding the same person appearing in the first video frame, even if the location of the bounding box in the second video frame differs from the location of the bounding box in the first video frame.

[0042] Those skilled in the art will appreciate that the unique IDs shown in the video frames of Figure 8 are provided for illustrative purposes. Notably, the people tracker of the order confirmation system is in no way limited to using these particular unique IDs or their particular values, as shown in Figure 8. Rather, the people tracker of the order confirmation system is operable with any unique ID that enables it to identify individuals between successive video frames of captured video footage and distinguish individuals appearing in a video frame from other individuals.

[0043] In one embodiment, a DeepSort algorithm is used to perform the tracking. Every new person who enters the observed scene is assigned a new ID. For a person detected in a previous video frame, the DeepSort algorithm uses a representation of the person that is sufficient to recognize the same person if that person leaves and later re-enters the observed scene. Upon detecting and recognizing the person, the DeepSort algorithm assigns the person the same ID that was assigned when the person was detected in the previous video frame.

[0044] The DeepSort algorithm comprises a Sort Tracker and a ReID module implemented with a View Knowledge Distillation (VKD) (Porrello A., Bergamini L. and Calderara S., “Robust Re-identification by Multiple View Knowledge Distillation, Computer Vision”, ECCV 2020, Springer International Publishing, European Conference on Computer Vision, Glasgow, August 2020) neural network.

[0045] The original sort tracker (as described by Bewley A, Ge Z., Ott L., Ramos F. and Upcroft B., “Simple Online and Realtime Tracking”, 2016 IEEE International Conference on Image Processing (ICIP), Phoenix, AZ, 2016, pp. 3464-3468) tracks people from one video frame to another using only position and motion cues. The original sort tracker relies on the previous track of a person detected in the current video frame. ID and a Kalman filter module that receives the bounding box of the person and the corresponding previous track T IDBased on the estimated location, the sort tracker estimates the location of previously detected people in the current video frame. The sort tracker then compares the estimated location with the details of the bounding boxes surrounding each person detected in the current video frame to find its closest match. The measurement vector of the Kalman filter is represented by the size and location of the bounding box center. Furthermore, the state vector of the Kalman filter includes motion information (i.e., the derivatives of the measurement vector components). While simple and computationally efficient, the original sort tracker suffers from frequent identity switching in crowded scenes. To overcome this limitation, the information output from the DeepSort algorithm is integrated with the sort tracker's appearance information extracted by an offline-trained deep neural network. The neural network enables ReID, i.e., the re-identification of people who have been previously detected but not seen for a while. The order confirmation system of the present disclosure uses a Views Knowledge Distillation (VKD) neural network to generate a better representation of a person's appearance for the purpose of ReID.

[0046] The VKD neural network learns numeric appearance descriptors for people. The appearance descriptors are trained so that the cosine distance between appearance descriptors obtained from different poses of the same person is small and the cosine distance between appearance descriptors from different people is large. The VKD architecture consists of a Resnet feature extractor, such as Resnet50 or Resnet101, and a classification head. The appearance descriptor represents the smoothed output of the Resnet after applying global average pooling. The VKD architecture is trained using a classification loss applied to the classification head and a triplet loss applied to the appearance descriptor. VKD achieves improved performance by learning robust representations using a teacher network and extracting knowledge into a student network. In this embodiment, the representations take the form of embedding vectors of length 2048. However, those skilled in the art will recognize that the order confirmation system of the present disclosure is not limited to embedding vectors of this length. Instead, the order confirmation system disclosed herein is operable with embedding vectors of any length that allow for recognition of a person within the settings and environmental conditions of a given order processing facility / order receiving facility.

[0047] The Deep Sort algorithm uses both motion and appearance information to assign unique IDs to people detected in received video frames. For simplicity, a person assigned a unique ID will hereafter be referred to as an enrollee. To enable tracking of an enrollee in subsequent received video frames by matching the enrollee with people detected in subsequent video frames, the Deep Sort algorithm retains the enrollee's ID, along with the enrollee's appearance descriptor, and the enrollee's corresponding position and motion information included in the corresponding Kalman filter state for a predefined number of subsequent received video frames.

[0048] If the registered person does not match the person detected in a predefined number of subsequently received video frames, the registered person's unique ID and corresponding appearance and movement information are discarded. In this embodiment, the predefined number of subsequently received video frames is 1000. However, those skilled in the art will recognize that the order confirmation system of the present disclosure is not limited to this number of subsequently received video frames. Instead, the order confirmation system disclosed herein can operate with any number of subsequently received video frames that allows for recognition of a person who may have left the field of view of the order confirmation system's video sensor and later re-entered this field of view, thereby meeting the requirements of the order receipt / order delivery process and the underlying conditions of a given order fulfillment facility / order receipt facility.

[0049] The Jonker-Volgenant algorithm estimates the track T ID is used to match the current detected person based on their location. The Jonker-Volgenant algorithm is an efficient variant of the Hungarian algorithm. In the first phase, the Jonker-Volgenant algorithm is used to match the previous track T using information about the person's appearance. ID The algorithm then compares the current detection with the previously identified ID. If there are still mismatched current detections after the first phase, the algorithm runs again in a second phase using the motion information described above. After the second phase, the previously mismatched tracks are retained in a database for use with the next received video frame, and the current mismatched detections are used to create new tracks corresponding to several newly created IDs after a specific predefined warm-up period. In this embodiment, the warm-up period is three video frames. However, those skilled in the art will recognize that the order confirmation system of the present disclosure is not so limited; rather, the number of frames used in the warm-up period can be empirically determined according to environmental conditions and the settings of the order processing facility / order receiving facility.

[0050] In another embodiment, the Jonker-Volgenant algorithm matches the person detected in the currently received video frame with the tracks of previously detected enrollees based on a weighted combination of a motion cost metric and an appearance cost metric. The motion cost metric may be calculated, for example, as the squared Mahalanobis distance between the Kalman filter measurement vector associated with the person detected in the current video frame and the measurement vector predicted by the Kalman filter for each previously detected enrollee. The appearance cost metric may be calculated, for example, as the cosine distance between the appearance descriptor of the person detected in the current video frame and the enrollee's appearance descriptor at each instance the enrollee is detected in the previously received video frame.

[0051] Those skilled in the art will recognize that the above formulations for the motion cost metric and appearance cost metric are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system of the present disclosure is not limited to these above-described formulations for the motion cost metric and appearance cost metric. Rather, the order confirmation system disclosed herein is operable with any formulation of the motion cost metric and appearance cost metric that supports matching a person detected in a currently received video frame with a previously detected enrolled person. For example, the motion cost metric and / or the appearance cost metric may instead use a formulation that includes maximum likelihood statistics.

[0052] We define a newly enrolled person as the person who was last assigned an ID, and if the newly enrolled person does not match a person detected in a predefined number of subsequent video frames, the detection leading to the newly enrolled person is considered a false positive, and therefore the newly enrolled person's unique ID and corresponding appearance and movement information are discarded.

[0053] Sorting supports short-term matching of detected persons, while ReID supports long-term matching. Sorting involves hyperparameters that need to be tuned on a validation dataset, which contains a series of video frames extracted from video footage captured by a video sensor in an order confirmation system at a constant frame rate, e.g., 4-7 frames per second. In contrast, the VKD algorithm employs a neural network trained on the ReID dataset. The ReID dataset contains cropped regions from received video frames, where a cropped region corresponds to the area occupied by a bounding box containing a person. The cropped regions in the ReID dataset are also grouped into tracklets, each representing a region extracted from a video frame belonging to the received video footage.

[0054] The neural network used in the VKD algorithm can be trained or pre-trained using open source datasets such as the Motion Analysis and Re-Identification (MARS) dataset. The ReID dataset used in the preferred embodiment has the following characteristics: Image size: variable (image of palette cropped using bounding box predicted by palette detector) Number of individuals / people: 48 Minimum number of bounding boxes per person: 30 Maximum number of bounding boxes per person: 3295 However, those skilled in the art will understand that both the values ​​associated with the above training / pre-training datasets and the above training / pre-training datasets are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system disclosed herein is not limited to using these datasets to train / pre-train neural networks with VKD. Rather, the order confirmation system disclosed herein is operable with any dataset suitable for training / pre-training neural networks used in VKD algorithms, including privately collected datasets.

[0055] Returning to Figure 4, QR detector 402d implements a quick response (QR) detection algorithm. The purpose of QR detector 402d is to enable identification of employees of a delivery person or a third-party supplier / buyer, etc., based on the presence of a QR code on a tag attached to the person's uniform. In this manner, entry to an order fulfillment / order receiving facility can be controlled such that only authorized entrants, i.e., those who present to the order fulfillment / order receiving facility a tag with a QR code that matches a known, approved QR code of the supplier / delivery person, etc., are granted access to the order fulfillment / order receiving facility.

[0056] In one embodiment, the QR detector 402d is implemented using a neural network based on the Yolo_v5 architecture, more specifically, Yolo_v5s. Yolo_v5 includes three main parts: a backbone, a neck, and a head. The backbone employs a CSP-Cross Stage Partial Network (CSP-Cross Stage Partial Network) that is used to extract features from input images / video frames. The neck is used to generate a feature pyramid. The neck contains a PANet, which helps the Yolo_v5s model generalize at different scales. The head is used for the final detection stage; specifically, the head generates anchor boxes and output vectors for the Yolo_v5s model. Those skilled in the art will recognize that the above network architecture is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system 110 of the present disclosure is not limited to the use of the above network architecture. Rather, the order confirmation system 110 disclosed herein is operable with any suitable network architecture that enables the detection and recognition of QR codes present in images. For example, the order confirmation system 110 may be operable with any other single-shot detector (SSD), such as RetinaNet, previously disclosed herein.

[0057] A detailed example of the dataset used to train the Yolo_v5 network is shown below. Number of images: 1681 Image size: 2560x1440 pixels Comments: 1681 During training, a reference frame is created, which is a video frame obtained from captured video footage of the monitored area without the presence of a QR code. In the next step, a short video is extracted from the raw video footage of the training dataset. The short video includes a series of video frames in which a QR code is displayed on a video camera. To ensure diversity in the feature distribution, video frames are extracted from the short video using an average hashing algorithm. In one embodiment, the average hashing algorithm was implemented using the open-source Python library ImageHash. However, it should be noted that the above-mentioned software tool for the average hashing algorithm is provided for illustrative purposes only. In particular, those skilled in the art will understand that the order confirmation system 110 of the present disclosure is not limited to use with the ImageHash software tool. Rather, the order confirmation system 110 disclosed herein is operable with any software implementation of the average hashing algorithm.

[0058] In the average hashing algorithm, a hash is calculated for each video frame in the short video. For simplicity, a given second or subsequent video frame in the short video is hereafter referred to as the candidate QR image, and the video frame preceding the candidate QR image in the short video is hereafter referred to as the preceding candidate QR image. In an iterative process starting with the second video frame in the short video and progressing stepwise through each remaining video frame in the short video, the hash of the candidate QR image is compared with the hash of the preceding candidate QR image and the hash of the reference frame. If the hash of the candidate QR image differs from the hash of the preceding candidate QR image by more than 5, the candidate QR image is selected, and the hash of the preceding candidate QR image is updated with the hash of the candidate QR image. Similarly, if the hash of the candidate QR image differs from the hash of the reference frame by more than 7, the candidate QR image is selected, and the hash of the reference frame is updated with the hash of the candidate QR image.

[0059] Once trained, the Yolo_v5 network of the above embodiment receives as input video frames from video footage captured by the video sensor 102 of the order confirmation system 110. In response, the Yolo_v5 network outputs three vectors as follows: (a) the coordinates of the center of a bounding box encompassing the QR code detected in the received image, along with the width and height of the bounding box, the width and height being normalized by scaling them, respectively, relative to the width and height of the received video frame; (b) an objectness score (rated between 0 and 1) indicating the neural network's confidence that an object center exists at a given location in the received video frame; and (c) Class probability of detected objects. Upon detecting a QR code in a received video frame, a corresponding region is cropped from the video frame. The cropped region corresponds to the combination of the area of ​​the video frame occupied by a bounding box surrounding the QR code plus an additional 20 pixels added to each side of the bounding box to ensure the entire QR code is contained within the cropped region. The QR code displayed in the cropped region is then decoded using a barcode reading software component. In one embodiment, the QR code reading software component is the Python library Pyzbar, which is based on the Zbar open source software suite. Those skilled in the art will appreciate that the above-described barcode reading software component is provided for illustrative purposes only. In particular, those skilled in the art will appreciate that the order confirmation system 110 of the present disclosure is not limited to use with the above-described barcode reading software component. Instead, the order confirmation system 110 disclosed herein is operable with any software component capable of reading QR codes, such as, but not limited to, PyQRCode, qrcode, and qrtools.

[0060] 4 in conjunction with FIG. 1, output from the barcode reading software component includes a string decoded from a QR code detected in a received video frame. The delivery detection module 402 associates the string with the detected person closest to the QR code in the received video frame. Thus, the ability of the person tracker 402b to re-identify a person from one frame to another (based on appearance and movement attributes) is enhanced by the combination with the identity assigned to the person based on the QR code they present to the video sensor 102 of the order confirmation system 110.

[0061] The palette detector module 404a implements a model capable of detecting, determining, and identifying the location of palettes. In one embodiment, the palette detector module 404a is implemented using a neural network based on the Yolo_v5 architecture, more specifically, Yolo_v5s, as shown in FIG. 9. Furthermore, as shown in FIG. 9, Yolo_v5 includes three main parts: a backbone, a neck, and a head. The backbone employs a Cross-Stage Partial (CSP) network, which is used to extract features from the input image. The neck of the Yolo_v5 neural network is used to generate a feature pyramid. The neck includes a PANet, which helps the Yolo_v5 neural network generalize at different scales. The head is used for the final detection stage, generating anchor boxes and output vectors from the Yolo_v5 neural network.

[0062] Those skilled in the art will recognize that the above network architectures are provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system of the present disclosure is not limited to use with the above network architectures. Rather, the order confirmation system disclosed herein can operate with any suitable network architecture that enables detection and recognition of QR codes present in images. For example, the order confirmation system disclosed herein can operate with any other single-shot detector, such as RetinaNet.

[0063] Exemplary details of the dataset used to train the Yolo_v5 network are as follows: Image size: 1920x1080 pixels Number of images (including palettes or parts of palettes taken at different angles): 1591 Number of bounding boxes enclosing palettes or parts of them in video frames of the dataset: 27873 Number of bounding boxes per class (the dataset should be balanced, i.e. each class should have the same number of bounding boxes):

[0064]

number

[0065] 1, 4, and 9, once trained, the Yolo_v5 network of the pallet detector module 404a receives as input video frames from video footage captured by the video sensor 102 of the order confirmation system 110. In response, the Yolo_v5 network outputs three vectors as follows: (a) coordinates of the center of a bounding box encompassing the detected palette in the received image, along with the width and height of the bounding box, the width and height being normalized by scaling, respectively, relative to the width and height of the received video frame; (b) An objectness score that indicates the confidence, evaluated between 0 and 1, of a neural network that a palette center exists at a given position within the received video frame, and (c) The class probability of the detected palette.

[0066] Let the time at which a first video frame of a given item of a video is captured by a video camera, for example, the video camera 502 shown in FIG. 5, be time τ. The time interval Δt between captures of consecutive video frames of the video will hereinafter be referred to as the sampling interval. Using this notation, the video can be

[0067]

Number

[0068] described as

[0069]

Number

[0070] represents an individual video frame of the video, which is captured at time τ + iΔt and will hereinafter be known as the sampling time of the video frame.

[0071] For clarity, in the following disclosure, the current sampling time t k is given by t k = τ + NΔt, where N < n. The previous sampling time t p is a sampling time prior to the current sampling time t k and is given by t p = τ + DΔt, where 0 < D < N. The current video frame Fr(t k ) is the video frame captured at the current sampling time t k . The previous video frame Fr(t p ) is the previous sampling time tp 1 and 4, the currently detected palette is the video frame captured at the current video frame Fr(t k ) detected by the palette detector module 404a. The previously detected palette is the palette detected by the palette detector module 404a in the previous video frame Fr(t p ) is the palette detected in the previous video frame Fr(t p ) by the palette detector module 404a. The current detection of the palette is k ) by the palette detector module 404a. Furthermore, the most recent previous detection of a palette is the closest previous sampling time to the current sampling time, i.e., the given current time t k The most recent previous detection of a palette is one of one or more previous detections of a given palette by the palette detector module 404a in the previous video frame. The most recent previous detection of a palette is the last detection of the palette in the previous video frame.

[0072] The pallet tracker module 404b is communicatively coupled to the pallet detector module 404a and receives therefrom a list of pallets detected in the current video frame. The pallet tracker module 404b uses the output of the pallet detector module 404a to track the movement of pallets after they are detected. To this end, the pallet tracker module 404b tracks the center of each bounding box output by the pallet detector module 404a. Specifically, the pallet tracker module 404b processes video footage from all video sensors 102 of the order confirmation system 110 to track only pallets traversing the receiving / shipping portal 101.

[0073] 2 and 4, it should be noted that the area in which pallets are tracked includes several zones: an external monitoring zone 201, a first buffer zone 202, a second buffer zone 203, and a region of interest (ROI) / pallet analysis zone 204. To describe the path taken by each pallet as it moves within the area, each pallet tracked by the pallet tracker module 404b is assigned to a "track." A track has six associated attributes: (1) a unique track identifier (Tr_ID); (2) the truck's lifetime, i.e., a variable used to count the time since the truck was last assigned to a pallet detected by the pallet detection module; and (3) A status variable indicating the status of the truck, i.e., the status variable indicates whether the truck is assigned to the pallet detected by the pallet detector module 404a, and the status variable can have one of two possible values, i.e., "assigned" and "unassigned." The default value of the status variable is "unassigned." (4) coordinates of the center of a bounding box encompassing the detected palette in the received video frame, along with the width and height of the bounding box, where the width and height are normalized by scaling them relative to the width and height of the received video frame, respectively; (5) The code of the zone shown in Figure 2 (hereinafter referred to as the zone code), namely: 201 - External Surveillance Zone 202 - First Buffer Zone 203 - Second Buffer Zone 204-Region of Interest (ROI) / Palette Analysis Zone 205-Internal Residual Zone (6) K path point vectors corresponding to each of the most recent K previous observations of the same palette.

[0074]

number

[0075] path vector containing

[0076]

number

[0077] Each such path point vector then contains four attributes derived from observation of the pallet. Specifically, the attributes of a given path point include: · The unique track identifier (Tr_ID) of the track with which the path vector is associated; and the time when the corresponding previous observation of the pallet occurred, and The coordinates of the center of the bounding box that encompasses the palette in the corresponding previous observation.

[0078] Therefore, at time t p The path point vector corresponding to the previous observation of a given pallet at

[0079]

number

[0080] where P=0, 1, . . . , k-1.

[0081] The input to the palette tracker module 404b is a list of currently detected palettes, with each element in the list having the following attributes: The coordinates of the center of the bounding box that contains the currently detected palette, The corner coordinates of the bounding box that contains the currently detected palette, a zone code representing the zone (201, 202, 203, or 204 in FIG. 2) in which the currently detected pallet (as described by the center of the bounding box encompassing the currently detected pallet) is determined to be located by the pallet detector module 404b; and A palette flag indicating whether the currently detected palette is assigned to a track. The default value of the palette flag is FALSE. However, the palette flag may be updated to TRUE by the palette tracker module 404b once a track to which the currently detected palette location may belong is identified.

[0082] Using the nomenclature above, the current video frame Fr(t k ), the pallet tracker module 404b uses the following method to match the currently detected pallet with the track maintained by the pallet tracker.

[0083] The default status of each truck from the plurality of trucks maintained by the pallet tracker is set to "unassigned."

[0084] The Euclidean distance is calculated between the center of the bounding box encompassing the currently detected pallet and the center of the bounding box encompassing each of the recently detected pallets whose track status variable has a value of "unassigned." The most recent previous detection of a pallet is indicated by the last element of the path vector of the track corresponding to the previously detected pallet in question. The currently detected pallet is assigned to the track of the previously detected pallet for which there is the smallest Euclidean distance between the most recent previous detection of that pallet and the currently detected pallet.

[0085] The palette flag of the assigned currently detected palette is set to true, and the status variable of the track in the problem is set to "assigned". Similarly, the center of the bounding box encompassing the assigned currently detected palette is added to the end of the track's path vector. The track's path vector is therefore increased in size by one path point vector, which contains the following attributes: The track's unique track identifier (Tr_ID), and The current sampling time, The coordinates of the center of the bounding box of the currently detected assigned palette.

[0086] The above sequence of processing steps is repeated for each currently detected pallet until there are no remaining tracks with the state variable "unassigned" or no currently detected pallets with the palette flag set to "false" (in other words, no potential matching pairs remaining). At the end of the process, if there are any currently detected pallets remaining that have not been assigned to a track (i.e., there are any currently detected pallets remaining with the palette flag set to "false"), a new track is created for the currently detected pallet.

[0087] The per-pallet product classification module 404c is configured to analyze the contents of the pallets. The per-pallet product classification module 404c includes two communicatively coupled modules: an instance segmentation module that performs instance segmentation, and an image search module that uses an image search algorithm to classify cropped bounding boxes of products detected by the instance segmentation module.

[0088] Because product appearances often change with the seasons and years, instead of retraining a model for each new appearance of a class, it is more scalable to have a generic model that can detect the presence of a product and a further model for recognizing the product using prior knowledge in the form of an easily updateable product database. To this end, a model for detecting the presence of a product and a model for recognizing products in the classes of "pack," "box," and "vegetable" are trained. The classes may be further expanded to include "small pack," "medium pack," and "large pack." Using this formulation, the details of a dataset that may be exemplarily used to train a model for per-pallet product classification module 404c are as follows: Image size: 1920x1080 pixels Number of images: 27050 Number of masks (annotations for different classes): 1114548 Number of masks per class: Staff: 24825 Vendor: 10565 Pack: 18400 Box: 15501 Vegetables: 101 Flowers: 45 Fruit: 22 Instance segmentation is used because products on a pallet may be stacked irregularly, and pixel-level masks improve product detection accuracy, which is the benefit of multi-task training. Bounding box-based detection and mask-based detection work synergistically to reduce errors.

[0089] In a preferred embodiment, the instance segmentation module employs a Transformer-based model inspired by Swin (described in Z. Liu, Y. Lin, Y. Cao, H. Yu, Y. Wei, Z. Zhang, S. Lin, and B. Gao, “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), 2021, pp. 10012-10022). However, those skilled in the art will recognize that the above Swin-based Transformer model is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system 110 of the preferred embodiment is not limited to the use of the Swin Transformer model. Rather, the order confirmation system 110 disclosed herein is operable with any Transformer-based or CNN-based backbone that can be used for instance segmentation.

[0090] The image retrieval module executes an algorithm for product re-identification based on a neural network that learns embeddings of each instance of a product contained in the product image database. More specifically, the image retrieval module compares the visual appearance of the pallet in the received video frames with visual appearance information of products expected to be received / shipped by the order receiving / order fulfillment facility based on, for example, an advance shipping notice.

[0091] For example, consider Vendor X, which has products "a," "b," and "c." For each of these products, the product database described above includes images representing the current appearance of these products. From these images, information about the product's appearance under various different conditions (e.g., from different viewpoints and rotation angles) can be represented as embedding vectors, which can be formed using embedding models such as VKD or Siamese Nets. Those skilled in the art will understand that these embedding models are provided for illustrative purposes only. In particular, those skilled in the art will understand that the order confirmation system 110 disclosed herein is not limited to the above-described embedding networks. Rather, the order confirmation system 110 disclosed herein can operate with any encoder model capable of forming an embedded vector representation of product appearance, such as a classical CNN trained as a classifier and then having its head removed. To train the embedding model, several images (up to 10 images) of each product are provided. Furthermore, VKD can be trained with images of the entire palette rather than images of each product. However, this approach requires providing a significantly larger number of images, e.g., at least 30 images.

[0092] The embedding vector forms a representation of the product (cola, chocolate, beer) that is robust to changes in appearance and viewpoint. This representation is used to identify the product in various images at various scales and positions, including various rotation and tilt angles between the product and the video sensor 102. The image search module of the pallet monitor module 404 compares the detected product in the received video frames with the product expected to be received / shipped by searching for the associated product image and / or embedding vector from a product database. The embedding vector is used for search and / or re-identification via a simple distance metric in the embedding space. In one embodiment, the distance metric is a cosine metric. However, those skilled in the art will recognize that the above distance metric is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the order confirmation system 110 of the present disclosure is not limited to the use of a cosine distance metric in the image search module. Rather, the order confirmation system 110 disclosed herein is operable with any suitable distance metric, such as a Euclidean distance metric.

[0093] If there are several products on a pallet, a representation of the pallet can be constructed by combining the embeddings of all the products visible on the pallet. Using this technique, it is possible to go beyond determining what products appear on the pallet; instead, it is also possible to extract information useful for estimating the number of products in a pack and the number of packs on a pallet.

[0094] "Key points" are defined as points of interest on a pallet. Each key point corresponds to a corner of the pallet. Furthermore, the term "pallet" refers to the entire configuration of the wooden body and the products placed on the wooden body. In a preferred embodiment, 16 key points are defined: 8 for the wooden body of the pallet and 8 for the products stacked on the pallet. For simplicity, the wooden body of the pallet and the products stacked on the pallet are hereafter collectively referred to as pallet subcomponents. Each of the above key points has a different class. The name of the class consists of the pallet subcomponent name and a name consisting of references to the three axes: near, far, left, right, and up and down (e.g., product_far_left_top, product_far_left_bottom).

[0095] Pallets can be one of two types: Regular Pallet: A pallet having the shape of a rectangular parallelepiped, with the products stacked on it, is considered to be a single object and keypoints are annotated accordingly. Non-regular pallet: a pallet that does not have the shape of a rectangular parallelepiped (for example, when the shape of the stack of products is not rectangular). In this case, the shape of the stack of products is divided into multiple rectangular parallelepipeds.

[0096] 4, for non-regular palettes, the palette keypoint detector (not shown) of the palette monitor module 404 is configured to detect multiple keypoints of the same class. In contrast, for regularly shaped palettes, the palette keypoint detector (not shown) detects unique keypoints (i.e., keypoints of different classes).

[0097] To this end, a palette keypoint detector (not shown) includes a convolutional neural network configured to receive a cropped region of a received video frame, the cropped region including a palette. The palette keypoint detector (not shown) is configured to process the received cropped region to generate 16 heatmaps, each heatmap including, for example, 128×128 pixels. Each heatmap determines the location of a corresponding keypoint within the cropped region.

[0098] Exemplary details of the dataset used to train the convolutional neural network are as follows: Image size: Variable, as the image for the palette is cropped using the bounding box established by the palette detector. Number of images: 2918 Number of comments: 30804 Number of annotations per class: Product_Far_Left_Top:2712 Product_Far_Right_Top:2777 Product_Far_Right_Bottom:791 Product vicinity left top: 2700 Product vicinity bottom left: 2583 Product_near_right_top:2781 Product Nearby Right Bottom: 2658 Palette_far_right_top:752 Palette_Far_Right_Bottom: 629 Palette_near_left_top:2543 Palette_near_left_bottom:2467 Palette_near_right_top:2640 Palette_near_right_bottom:2591 Palette_Far_Left_Top:800 Palette_Far_Left_Bottom: 596 Product_Far_Left_Bottom:784.

[0099] Thus, referring to FIG. 10, a palette keypoint detector (not shown) performs the following steps:

[0100] Detect palettes in received video frames 1000. Palette detection is performed by a palette detector module as disclosed earlier herein.

[0101] A region in the received video frame where the presence of a palette is detected by the palette detector module is cropped 1002. The cropped region corresponds to the combination of the area of ​​the video frame occupied by the bounding box surrounding the palette, plus an additional 20 pixels added to each side of the bounding box to ensure that the entire palette is included in the cropped region. For simplicity, this cropped region is hereafter referred to as the "cropped palette region." In practice, the cropped palette region may include, for example, 128x128 pixels, with the upper left corner of the cropped palette region located at coordinates (x1, y1) in the received video frame.

[0102] A convolutional neural network sequentially processes 1004 each of the one or more cropped palette regions from the received video frame to generate one or more heat maps. In one embodiment, the convolutional neural network is configured to generate 16 heat maps from the cropped palette regions. However, those skilled in the art will recognize that the above number of heat maps is provided for illustrative purposes only. In particular, those skilled in the art will recognize that the palette keypoint detector of the preferred embodiment is not limited to generating this number of heat maps. Rather, those skilled in the art will recognize that the palette keypoint detector is operable to generate any number of heat maps as needed to enable accurate detection of palette keypoints visible in the cropped palette regions from the cropped palette regions.

[0103] The multiple heatmaps are post-processed 1006 by a function to generate a corresponding number of lists of points, each of which corresponds to a palette keypoint.

[0104] The point is scaled 1008 to the dimensions of the crop palette area. So, for example, if the crop palette area has dimensions 128x128 and the point is defined by (x,y)=(0.46,0.76), the point is located at approximately (x',y')=(59,97) in the crop palette area's coordinate system, whose upper left corner is denoted by (0,0).

[0105] The scaled points are transformed 1010 back into the coordinate system of the received video frame, the result of the transformation being the detected palette keypoints 1012.

[0106] Returning to Figure 4, the palette volume estimation algorithm of palette volume estimator 404d calculates the volume of objects by estimating their size from a 2D image. In general, it is impossible to recover a 3D position from a 2D projection because an infinite number of points from a line in 3D space will project to the same point on the 2D projection, i.e., onto the camera plane. One possible solution to resolving the ambiguity exploited in stereo vision is to use a pair of views of the scene captured from different positions and recover depth information using triangulation principles.

[0107] The preferred embodiment assumes a flat, horizontal floor and uses a homography to calculate the real-world coordinates of floor points from camera coordinates. Figure 11(a) shows a physical grid pattern of points marked on the ground, and Figure 11(b) shows a representation of these points on the camera projection. For simplicity, these grid pattern points are hereafter referred to as the reference grid. Corner points that belong to the same vertical line (e.g., as shown in Figure 12(a), the upper-left corner T' of the front face of a rectangular parallelepiped corresponds to the lower-left corner B' of the same face) and their correspondence in projection on the floor that belongs to the basic geometry are used to calculate the elevations of the corners of the parallelepiped from the ground. From these, the volume of the object is calculated.

[0108] The first step of the algorithm is to estimate the parameters of the homography transformation mapping point (x', y') from the floor to camera pixel coordinates (x, y).

[0109]

number

[0110] Since not all parameters are independent, the homography matrix is ​​estimated up to the scale. For this purpose, the matrix is ​​normalized. For example, in equation (1), h 22 can be set to a value of 1, so that the remaining eight parameters of the H matrix can be recovered from a set of four corresponding points with known locations, taken from a known reference grid. For better accuracy, more corresponding points with known locations are used.

[0111] Returning to Figure 11, a reference grid is constructed with points marked on the floor, and homographies are estimated for several sets of four points. A least-median-of-squares robust estimation method is then used to find the solution parameters. In a possible embodiment, the spacing between individual points of the grid (i.e., the grid size parameter d) is set to 50 cm.

[0112] In many cases, a pallet can be represented by a rectangular parallelepiped object. The volume of the rectangular parallelepiped object can be calculated using the size of its edges (in pixels). This is calculated from key points representing the corners of the pallet. In particular, the width and length of the pallet are the sizes of the two adjacent edges resting on the floor. Therefore, if we define the bottom edge of the parallelepiped as its edge resting on the floor, the length of the bottom edge can be calculated from the corner of the pallet corresponding to the parallelepiped. For this purpose, the location of the pallet corner is estimated by referencing the point from the ground pattern in Figure 11(a) that is observed to be closest to the corner. From this, the length of the bottom edge of the corresponding parallelepiped is calculated using equation (1) above.

[0113] The height of a rectangular object can be calculated as the distance between one corner resting on the floor and the opposite corner located directly above it. For example, referring to Figure 13, T is located above B, so the projection T p are B and C p The camera C observing the pallet corresponding to the parallelepiped is at a known height C above the floor. height It is placed at height C height is the length of Figure 13 |CC p |

[0114]

number

[0115] and

[0116]

number

[0117] also represents the height.

[0118]

number

[0119] and

[0120]

number

[0121] is perpendicular to the floor and aligned parallel to it. Therefore, equation (2) is p C p C and triangle T p This can be established from the similarity between the BT triangle.

[0122]

number

[0123] Sometimes, pallets may exhibit shapes other than a rectangular prism. Two common cases are (a) when the number of items on different pallet rows differs from each other, and (b) when a pallet consists of a set of non-uniform packs of different items, each of which exhibits a rectangular prism shape.

[0124] For objects of different shapes (e.g., case (a) disclosed above), the line CC in FIG. p Since the surface is parallel to the CC line in Figure 13, the above method can accurately determine only the points corresponding to the line perpendicular to the floor (representing the height), so additional key points are required for volume calculation. For example, in Figure 14, point D, which is located below point A, is necessary for estimating the position of point A. Otherwise, the area of ​​surface ABCE would be estimated incorrectly if |EA| is the tilt height of the object shown in Figure 14. Specifically, point A is located at CC in Figure 13. p It is not parallel to the floor and would be estimated to be higher than the floor because the pixel distance between points A and E is greater than the pixel distance between points A and D. To accurately estimate the area of ​​surface ABCE, the areas of ABCD and ADE need to be calculated separately.

[0125] Figure 15 shows a different pallet shape representative of case (b) disclosed above, where the pallet is composed of a non-uniform set of packs of different items. In this case, the pallet shape is formed by two rectangular objects stacked on top of each other. Point G can be calculated using the above method only if point G' below it is known (given as a key point). Similarly, point E can be estimated because E' is perpendicular to the floor.

[0126] Thus, in each received video frame, keypoints representing the corners of the pallet visible therein are detected by a pallet keypoint detector (not shown), as well as other useful points (e.g., G' as shown in Figure 15). Returning to Figure 4, the pallet volume estimation algorithm (not shown) of the pallet volume estimator 404d is configured to use the above-mentioned Apache to calculate the volume of the pallet in each of the video frames if enough keypoints (pallet corners) are visible and detected, regardless of the angle of the pallet relative to the camera observing it.

[0127] Referring to FIG. 4 in conjunction with FIG. 1, the in-out counter 404e is configured to use a path determined by the pallet tracker module 404b and defined as a sequence of the form:

[0128]

number

[0129] , where k=0, 1, . . . , K, and detects which paths intersect with the receiving / shipping portal 101, thereby determining which pallets are entering or leaving the order fulfillment facility / order receiving facility. Using this information, the order confirmation system 110 records incoming and outgoing pallets (identified by their IDs) along with their time of entry / exit. Additionally, a count may be kept of the number of incoming / outgoing pallets to / from the order fulfillment facility / order receiving facility over a given period of time (as determined by the direction of pallet movement).

[0130] Additionally, a value of a pallet status variable may be recorded for the pallet. The pallet status variable characterizes the degree to which the pallet is loaded. For example, the value of the pallet status variable may be "full," "empty," "partially loaded," etc. The value of the pallet status variable may be determined using the height of the goods stacked on the pallet as determined by the pallet volume estimator 404d. Knowing the maximum allowable stack height of the pallet, the pallet may be classified as follows: - Fully loaded if the height of the stacked goods is close to the maximum permitted stack height, or - Partial load if the height of the stacked goods is approximately half the maximum permitted stack height, or - A pallet is empty if it has no products on it.

[0131] The event (or alert) management module 406 triggers specific alerts specific to individual applications based on events of interest (e.g., exceeding a maximum door open time, invalid access to a receiving / shipping portal, etc.). Events are generated using, for example, outputs from the door status detector 402a, person detector 402b, and QR detector 402d. Events of interest are also recorded as GIF files and saved to disk for later use.

[0132] The event (or alert) management module 406 executes logic for a particular alert, such as, but not limited to: The receiving / shipping portal 101 remains open for a period longer than a certain threshold; and The pallet must remain in a specific area for a period longer than a certain threshold; - The contents of pallets from certain vendors do not match the advance shipping notices; Pallets leaving the receiving / shipping portal 101 without being registered in the order confirmation system 100; The height of the pallet exceeds a certain maximum allowable height, and - No employee of the order processing / receiving facility is present when the delivery person arrives; · Delivery personnel entering order processing / receiving facilities without signing in.

[0133] These and other types of alerts are based on output from, for example, door status detector 402a, person detector 402b, and QR detector 402d. For example, when receiving / shipping portal 101 is opened, a timer is started and when the timer reaches a certain threshold, an alert is triggered and the event is logged.

[0134] When an alert is triggered, the event recorder 406b is configured to save a set of consecutive video frames on disk (the consecutive video frames can also be assembled and packed as an animated GIF file), starting from a specific period before the alert and ending a specific period after the alert (for example, the period can be extended from 10 seconds to 60 seconds based on the type of alert). The generated video frames / GIFs are saved and can be reviewed by staff at any time. The maximum number of stored video frames / GIFs and the period they are retained can be configured according to operator needs and are handled by predefined logic in the event recorder 406b based on application-specific requirements.

[0135] Modifications to the embodiments of the present disclosure described above are possible without departing from the scope of the present disclosure, which is defined by the appended claims. The terms "including," "comprising," "incorporating," "consisting of," "having," "being," and the like, as used to describe and claim the present disclosure, are intended to be interpreted in a non-exclusive manner, i.e., they construe that items, components, or elements not expressly described are also present. References to the singular should also be interpreted as relating to the plural.

Claims

1. a plurality of video sensors configured to capture video footage of a monitor area located within the order receiving or order shipping area of ​​the receiving / shipping portal; performing an event analysis on the captured video footage; Detecting an entity within the video footage captured by the video sensor; detecting an arrival of goods from a third-party supplier from a door opening event in the captured video footage; Identifying the third-party supplier and conducting a check-in process for delivery personnel from the third-party supplier; Detecting the entry and exit of goods via the receiving / shipping portal and verifying that the detected delivered products match data regarding products to be delivered by the third-party supplier; analyzing the video footage captured by the video sensor using a door status detector to determine whether the receiving / shipping portal is in an open or closed state; analyzing the video footage captured by the video sensor of the order confirmation system using a person detector to detect whether a delivery person has arrived at the receiving / shipping portal; using a person tracker to track the delivery person's movements through the captured video footage upon detecting the delivery person by the person detector; a processing unit configured to detect the presence of a Quick Response (QR) code in the captured video footage using a Quick Response (QR) detector and read the QR code; a database communicatively coupled to the processing unit, the database comprising: At least one dataset of face images / logos for use in face / brand detection; a dataset of product images for use in product identification; and the database configured to store the results of the order confirmation process and a record of the delivery person's checkout at the end of the delivery for future retrieval upon request to the processing unit; An order confirmation system comprising:

2. The QR detector compares the detected QR code with known pre-approved QR codes for third party suppliers / delivery personnel to find a match, and if a match is found, the QR detector: classifying the delivery person who presents the QR code as an authorized entrant to an order fulfillment / order receiving facility; The order confirmation system of claim 1 , wherein a delivery detection module is used to grant the delivery person access to the order processing / order receiving facility.

3. 10. The order confirmation system of claim 1, wherein the person detector and the door state detector are embodied as neural networks of a predetermined architecture configured for person and door detection.

4. The processing unit and a pallet monitor module configured to verify one or more contents of goods being delivered or received from the order fulfillment facility / order receiving facility premises, said pallet monitor module comprising: a pallet detector module configured to detect a pallet; a pallet tracker module configured to track the detected pallet; a per-pallet commodity sorting module configured to sort commodities on the detected pallets; a pallet volume estimator configured to estimate a quantity of items on the detected pallet; an in-out counter configured to extract information regarding a total number of pallets passing through the receiving / shipping portal.

5. The processing unit further comprises an event management module in communication with the delivery detection module and the pallet monitor module, the event management module comprising: an alert manager configured to issue an alert regarding the detection of an authorized entrant and information regarding the goods being supplied / delivered during a supply / delivery episode; an event recorder configured to record said supply / delivery episode; 5. The order confirmation system of claim 4, comprising:

6. The alert manager when the receiving / shipping portal remains open for a period of time that exceeds a first predefined threshold; when the pallet remains in a particular area for a period of time that exceeds a second predefined threshold; when the contents of the pallet from a particular third-party supplier do not match the corresponding details in the advance shipping notice; When the pallet leaves the receiving / shipping portal without being pre-registered, When the height of the pallet exceeds a predefined maximum allowable height, When the delivery person arrives and no employees of the order processing facility / order receiving facility are present, 6. The order confirmation system of claim 5, further configured to issue an alert when the delivery person enters the order processing / order receiving facility without first signing in.

7. 1. A method for performing video surveillance, the method comprising: using a plurality of video sensors to capture video footage of a monitor area located within an order receiving area or an order shipping area of ​​a receiving / shipping portal; using a processing unit to perform event analysis on the captured video footage; using the processing unit to detect entities within the video footage captured by the video sensor; detecting, using the processing unit, an arrival of a shipment from a third-party supplier from a door-opening event in the captured video footage; using the processing unit to identify the third-party supplier and perform a check-in process for a delivery person from the third-party supplier; using the processing unit to detect the entry and exit of goods through the receiving / shipping portal and verify that the detected delivered products match data regarding products to be delivered by the third-party supplier; analyzing the video footage captured by the video sensor to determine whether the receiving / shipping portal is in an open or closed state; analyzing the video footage captured by the video sensor of the order confirmation system to detect whether a delivery person has arrived at the receiving / shipping portal; Upon detecting a delivery person, tracking the delivery person's movements through the captured video footage; detecting the presence of a quick response (QR) code within the captured video footage; reading the detected QR code as an output; using a database to store at least a dataset of face images / logos for use in face / brand detection and a dataset of product images for use in product identification; using said database to record the results of the order confirmation process and the delivery person's checkout at the end of delivery; retrieving said record by said processing unit from said database in response to a request to said processing unit; A method comprising:

8. further comprising comparing the detected QR code with known pre-approved QR codes for third party suppliers / delivery personnel to find a match, and if a match is found in the event: classifying a delivery person who presents the QR code as an authorized entrant to an order fulfillment / order receiving facility; The method of claim 7 , further comprising granting the delivery person access to the order processing / order receiving facility.

9. The method of claim 7 , further comprising the step of: executing a neural network of a predetermined architecture configured for person and door detection.

10. The method further includes a step of verifying the contents of one or more items delivered or received from the premises of an order fulfillment facility / order receiving facility, said verifying step comprising: detecting a pallet using a pallet detector module; tracking the detected pallets using a pallet tracker module; classifying items on the detected pallets using a pallet-by-pallet item classification module; estimating the quantity of items on the detected pallet using a pallet volume estimator; and extracting information regarding the total number of pallets passing through the receiving / shipping portal using an in-out counter.

11. issuing an alert regarding the detection of an authorized entrant and information regarding the goods being supplied / delivered during a supply / delivery episode using an alert manager; The method of claim 10, further comprising: recording the supply / delivery episode using an event recorder.

12. the receiving / shipping portal remains open for a period of time that exceeds a first predefined threshold; the pallet remains in a particular area for a period of time that exceeds a second predefined threshold; The contents of the pallet from a particular third-party supplier do not match the corresponding details in the advance shipping notice; the pallet leaves the receiving / shipping portal without being pre-registered; the height of the pallet exceeds a predefined maximum allowable height; No employees of the order processing facility / order receiving facility are present when the delivery person arrives; 12. The method of claim 11, further comprising issuing an alert by the alert manager in the event the delivery person enters the order processing / order receiving facility without previously signing in.

13. When executed by a processing unit, the processing unit: capturing video footage of a monitor area located within an order receiving area or an order shipping area of ​​a receiving / shipping portal using a plurality of video sensors; performing an event analysis on the captured video footage; and detecting an entity within the video footage captured by the video sensor; detecting an arrival of a shipment from a third-party supplier using the processing unit from a door-opening event in the captured video footage; identifying the third-party supplier and performing a check-in process for a delivery person from the third-party supplier; Detecting the entry and exit of goods through the receiving / shipping portal and verifying that the detected delivered products match data regarding products to be delivered by the third-party supplier; analyzing the video footage captured by the video sensor of an order confirmation system to determine whether the receiving / shipping portal is in an open or closed state; analyzing the video footage captured by the video sensor to detect whether a delivery person has arrived at the receiving / shipping portal; Upon detecting a delivery person, tracking the delivery person's movements through the captured video footage; Detecting the presence of a Quick Response (QR) code within the captured video footage and reading the detected QR code as an output; using a database to store at least a dataset of face images / logos for use in face / brand detection and a dataset of product images for use in product identification; using said database to record the results of the order confirmation process and the delivery person's checkout at the completion of delivery; retrieving the record from the database in response to a request by the processing unit; A non-transitory computer-readable medium having stored thereon computer-executable instructions for causing the computer to execute

14. Upon execution of the executable instructions, the processing unit: configured to compare the detected QR code with known pre-approved QR codes for third-party suppliers / delivery personnel to find a match, and in the event a match is found, classifying a delivery person who presents the QR code as an authorized entrant to an order fulfillment / order receiving facility; 14. The non-transitory computer-readable medium of claim 13, wherein the delivery person is granted access to the order processing / order receiving facility.

15. Upon execution of the executable instructions, the processing unit: configured to verify one or more items of merchandise delivered or received from the premises of the order fulfillment facility / order receiving facility, said verifying comprising: detecting a pallet using a pallet detector module; tracking the detected pallet using a pallet tracker module; classifying items on the detected pallets using a pallet-by-pallet item classification module; estimating the quantity of items on the detected pallet using a pallet volume estimator; and extracting information regarding a total number of pallets passing through the receiving / shipping portal using an in-out counter.

16. Upon execution of the executable instructions, the processing unit: issuing alerts regarding the detection of authorized entrants and information regarding said goods being supplied / delivered during a supply / delivery episode using an alert manager; 16. The non-transitory computer-readable medium of claim 15, configured to record the supply / delivery episode using an event recorder.

17. Upon execution of the executable instructions, the processing unit: the receiving / shipping portal remains open for a period of time that exceeds a first predefined threshold; the pallet remains in a particular area for a period of time that exceeds a second predefined threshold; the said contents of a pallet from a particular third-party supplier do not match the corresponding details in the advance shipping notice; the pallet leaves the receiving / shipping portal without being pre-registered; the height of the pallet exceeds a predefined maximum allowable height; No employees of the order processing facility / order receiving facility are present when the delivery person arrives; 16. The non-transitory computer-readable medium of claim 15, configured to issue an alert using an alert manager in the event the delivery person enters the order processing facility / order receiving facility without first signing in.

Citation Information

Patent Citations

  • System and method for direct store distribution

    US20210233016A1