System and method for alerting incorrect placement of warehouse inventory
The warehouse management system uses machine learning to automate the detection and correction of inventory placement errors by analyzing camera images, enhancing efficiency and reducing manual effort.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2025-09-26
- Publication Date
- 2026-04-30
AI Technical Summary
Manual identification of incorrect inventory placement in warehouses is tedious and cumbersome due to human error, leading to mismatches between actual and database data.
A warehouse management system using machine learning models to analyze images captured by cameras, identify shelf and inventory identifiers, and compare them with database entries to detect and alert mismatches.
Automates the process of identifying and correcting inventory placement errors, reducing resource consumption and improving efficiency by directly alerting operators to incorrect placements.
Smart Images

Figure JP2025034193_30042026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR ALERTING INCORRECT PLACEMENT OF WAREHOUSE INVENTORY
[0001] The present disclosure relates to the field of artificial intelligence, and more particularly relates a system and method for alerting incorrect inventory placement in a warehouse.
[0002] In a warehouse, the warehouse operators place various inventory (interchangeably referred to as "packages", "parcels", "items", and / or "pallets") on specific shelves. Each shelf and inventory may be marked by a unique identifier (e.g., a barcode or alphanumeric code) for identifying the respective shelf and inventory. The warehouse operator, upon placing an inventory on a particular shelf, may input in a database about the details (e.g., identifiers) of the inventory and its associated shelf. In some instances, for example, due to human error, the inventory may be placed on a first shelf, but the database reflects that the inventory is placed on a second shelf. In such situations, where there is a mismatch between actual data and database data, rectifying this mismatch can involve manually identifying any incorrect placement of warehouse inventory, by scanning each inventory and / or shelf using handheld devices (e.g, Radio Frequency Identifier (RFID) reader or barcode reader), which can be tedious and / or cumbersome.
[0003] These and other problems are generally solved or circumvented, and technical advantages are generally achieved, by advantageous embodiments of the present disclosure.
[0004] A summary of certain embodiments disclosed herein is set forth below. It should be understood that these aspects are presented merely to provide the reader with a brief summary of these certain embodiments and that these aspects are not intended to limit the scope of this disclosure. Indeed, this disclosure may encompass a variety of aspects that may not be set forth below.
[0005] In one embodiment, the disclosure is directed towards a system. The system comprises at least one memory, and at least one processor configured to receive an image captured by a camera. The at least one processor is configured to cause the system to determine, using a first machine learning model (MLM), if the received image includes a shelf and an inventory. On determining that the received image includes the shelf, the at least one processor is configured to cause the system to determine, using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image. The at least one processor is configured to cause the system to recognize the identifier for: the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier. The at least one processor is configured to cause the system to compare the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs. The at least one processor is configured to cause the system to initiate one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs.
[0006] In another embodiment, the disclosure is directed towards a method. The method comprises receiving an image captured by a camera. The method comprises determining, using a first machine learning model (MLM), if the received image includes a shelf and an inventory. On determining that the received image includes the shelf, the method comprises determining, using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image. The method comprises recognizing the identifier for the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier. The method comprises comparing the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs. The method comprises initiating one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs.
[0007] The details of the embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the disclosure will be apparent from the description and drawings, and from the claims.
[0008] The detailed description is described with reference to the accompanying figures. The same numbers are used throughout the drawings to reference like features and components.
[0009] Fig. 1 illustrates an example image of a warehouse;Fig. 2 illustrates a block diagram depicting an environment in which the warehouse management system (WMS), according to an embodiment of the present disclosure;Fig. 3 illustrates the various components within the WMS, according to an embodiment of the present disclosure;Fig. 4 illustrates a process flow diagram for detection of incorrect placement of inventory, according to an embodiment of the present disclosure;Fig. 5 illustrates an example image captured by a camera, according to an embodiment of the present disclosure;Fig. 6 illustrates an example screenshot of table contents in a database, according to an embodiment of the present disclosure; andFig. 7 illustrates a method performed by the WMS, according to an embodiment of the present disclosure.
[0010] Exemplary embodiments now will be described with reference to the accompanying drawings. The embodiments disclosed herein may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather there are several variations and modifications which may be made without departing from the scope of the embodiments disclosed herein. The terminology used in the detailed description of the particular exemplary embodiments illustrated in the accompanying drawings is not intended to be limiting. In the drawings, like numbers refer to like elements. The term "exemplary embodiment" is meant to be interpreted as being an example embodiment and is not meant to be interpreted as a preferred embodiment.
[0011] The specification may refer to "an", "one" or "some" embodiment(s) in several locations. This does not necessarily imply that each such reference is to the same embodiment(s), or that the feature only applies to a single embodiment. Single features of different embodiments may also be combined to provide other embodiments.
[0012] As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms "includes", "comprises", "including" and / or "comprising" when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, whenever the phrase "at least one of" or "one or more of" precedes a list of elements, wherein the elements are joined by "and" or "or", it means that at least any one of the elements or at least all the elements are present. As used herein, whenever the phrase "one of" precedes a list of elements, wherein the elements are joined by "and" or "or", it means that only one of the elements are present at a given instant, unless the context permits a meaning that allows the inclusion of more than one element. The usage of the term "or" is to be understood as "inclusive or" instead of "exclusive or", unless indicated otherwise by the relevant context. Conditional language, such as among others, "can" or "may", unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments could include, while other embodiments may not include certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments. It will be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements may be present. Furthermore, "connected" or "coupled" as used herein may include wirelessly connected or coupled. As used herein, the term "and / or" includes any and all combinations and arrangements of one or more of the associated listed items.
[0013] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0014] The figures depict a simplified structure only showing some elements and functional entities, all being logical units whose implementation may differ from what is shown. The connections shown are logical connections; the actual physical connections may be different. In addition, all logical units described and depicted in the figures include the software and / or hardware components required for the unit to function. Further, each unit may comprise within itself one or more components, which are implicitly understood. These components may be operatively coupled to each other and be configured to communicate with each other to perform the function of the said unit. For the sake of illustrative purposes, the embodiments herein recite various units / modules that are associated with a particular functionality. However, this is to be construed as non-limiting as the functionality of the various units / modules may be combinable in any manner.
[0015] Referring now to the drawings, and more particularly to Figs. 1 to 7, where similar reference characters denote corresponding features used consistently throughout the figures relating to the example embodiments disclosed herein.
[0016] Fig. 1 illustrates an example image of a warehouse. The warehouse can include multiple areas dedicated to a specific purpose. Examples of the areas include inbound docks, outbound docks, a staging area, a shipping area, a storage area, and a packaging area. The inbound docks, staging area, and storage area can be used for receiving inbound inventory, organizing the inbound inventory, and storing the inbound inventory on shelves, respectively. The packaging area, shipping area, and outbound docks, can be used for packing outbound inventory, organizing the outbound inventory to be shipped, and shipping the outbound inventory, respectively.
[0017] The storage area, which can include shelves, can store one or more (inbound) packages. Each package can have a unique identifier for identifying the package. Each shelf can also have a unique identifier for identifying the shelf.
[0018] The embodiments herein disclose a warehouse management system (WMS) 102 that can receive an image of various warehouse inventory and the shelves that the inventory is placed on. The images can be captured by a camera 112 that can be located, for example, on a side of forklift. The captured image may also include the unique identifier for each of the inventory and its associated shelf, wherein the unique identifier may be affixed on the inventory and the shelf, respectively. The WMS 102 recognizes and interprets the unique identifiers of the inventory and shelf, from the received image, and compares the inventory-shelf pair identifiers with those in the database 116. In case there is a mismatch during the comparison process, the WMS 102 generates an alert to a warehouse operator about the incorrect placement of a warehouse inventory, so that corrective action can then be taken by the warehouse operator. In some embodiments, the WMS 102 updates the database 116 with the inventory-shelf pair identifier recognized from the captured image. It is to be noted that the various units / modules whose functionality have been described herein can also be described as functions performed by the WMS 102. Additionally, the WMS 102 may comprise further units / modules not explicitly described herein, but would be implicitly understood by a person skilled in the art as being present in the WMS 102 based on the functionality of the WMS 102 as described herein.
[0019] <Overview Of The Warehouse Management System (WMS)> Fig. 2 illustrates a block diagram depicting an environment 100 in which the WMS 102 can be implemented. The environment can comprise a camera 112 for capturing the images of the shelves with the inventory. The camera 112 can have a wireless communication module for transmitting the captured images to the WMS 102, for analysis, via a network 108.
[0020] The network 108 can be the Internet, or a private communication link (e.g., LAN or WAN). The network 108 may use standard communication technologies or protocols. Examples of the wireless communication modules can include, but are not limited to, WiFi module, Bluetooth module etc.
[0021] The WMS 102 can be implemented by processing circuitry ("processors") 104 and one or more storage devices ("memory") 106. By way of example, rather than limitation, the WMS 102 can operate in the capacity of a server. The server may be a physical or virtual server, and the server may be a web server, an application server, or a cloud server.
[0022] The processing circuitry 104 can include, although is not limited to, a general-purpose central processing unit (CPU), application specific integrated circuit (ASIC), microcontrollers, multiple processing units, dedicated circuitry etc. for achieving the functionality of various units / modules that will be described later herein.
[0023] The memory 104 can be any suitable processor-readable storage medium, e.g., random access memory (RAM), read-only memory (ROM), Electrical Erasable Read-only Memory (EEPROM), Flash memory, etc., suitable for storing instructions for execution by the processor, and located separate from the processing circuitry and / or integrated therewith. The memory can store various instructions that enable the processing circuitry to execute the functions of the various modules that will be described later herein.
[0024] The database 116 can be a physical device, or stored on non-volatile storage media (e.g., a hard drive), or may be cloud-based. The database 116 may also be implemented using a structured query language (SQL) datastore on a server.
[0025] The warehouse operator can have a user device 118 (e.g., a handheld device like a smartphone, tablet, or pager) that can communicate with the WMS 102 and receive alerts or notifications from the WMS 102.
[0026] Fig. 3 illustrates the various modules with the WMS 102, according to an example embodiment disclosed herein. The WMS 102 comprises the following modules: an image filtering module 1021, an identifier detection module 1022, an identifier recognition module 1023, an identifier comparison module 1024, and a mismatch handling module 1025.
[0027] The image filtering module 1021 can receive the image captured and transmitted by the camera 112. In an embodiment where the camera 112 is located on a side / both sides of a forklift and the forklift passes through an aisle, the camera 112 can capture images of the shelves and the inventory on the shelves, along with their respective unique identifiers. Fig. 5 illustrates an example image captured by the camera 112. Without limitation to the aforementioned embodiment, it may be possible that the camera 112 captures non-aisle images. In other words, the captured images may not depict the shelves, the inventory, and their respective unique identifiers.
[0028] Accordingly, the image filtering module 1021, on receiving the captured image, can determine if the image contains a depiction of an inventory and a shelf. In other words, the image filtering module 1021 determines if the captured image includes an inventory and a shelf. This determination can happen in a two-step manner, wherein in a first step, the image filtering module 1021 checks for the presence of a shelf. However, if at the first step, no shelf is present within the image, then further image analysis may not occur, If the shelf is determined to be present, then in a second step, the image filtering module 1021 checks for the presence of an inventory in the image. Regardless of whether the captured image includes or is non-inclusive of any inventory, the image may still be processed. In such an embodiment, the shelf identifier can be recognized from the image, and for the absent inventory, a special identifier can be assigned. The special identifier can correspond to any unique character or character sequence to signify that it represents an identifier for an absent inventory. This identifier pair of the shelf-absent inventory can be compared with the corresponding data in the database 116 to determine if there is a mismatch. In other words, if the image indicates that a certain shelf is free of any inventory, then this can be compared with the data in the database 116, wherein if the database 116 indicates that the shelf is supposed to contain a particular inventory, then a mismatch alert can be triggered. Since the embodiments herein refrain from further analyzing an image if a shelf is absent in the image, the embodiments herein achieve a technical effect of better utilization of resources that would have been consumed in analyzing / processing an image that lacked the necessary information for the accuracy of inventory placement in the warehouse.
[0029] In an embodiment, the image filtering module 1021 can be implemented by at least one machine learning model (referred to as "image classification model"). An example of the image classification model can be a convolutional neural network (CNN). In an example embodiment, the CNN used can be ResNet50 built using TensorFlow. The image classification model may be trained to perform binary classification of the received image from the camera 112. The image classification model may perform a first binary classification as to whether the image depicts a shelf. The image classification model may perform a second binary classification as to whether the image depicts an inventory / item / package. In an event that the image classification model renders an output that the image lacks the shelf (as part of the first binary classification), the image filtering module 1021 may discard the received image so that no further processing of the image can occur. In an embodiment, the image classification model can generate bounding boxes for the shelf and the inventory. If no bounding boxes are generated for the shelf, then it can mean that the image does not include a shelf. Similarly, if no bounding boxes are generated for an inventory, that it can mean that no inventory is present in the image.
[0030] If the image classification model determines that the image includes a shelf and, optionally, an inventory, the received image is fed to an identifier detection module 1022. The identifier detection module 1022 can detect the portion(s) in the image corresponding to the unique identifiers for the inventory (if present) and shelves. In other words, the identifier detection module 1022 can detect the presence and location of the unique identifiers in the received image. The unique identifiers can be, by way of example rather than limitation, in the form of alphabets, numbers, an alphanumeric code, or a barcode.
[0031] In an example embodiment, the identifier detection module 1022 can be implemented by a character detection model (e.g., character-region awareness for text (CRAFT) detection model (implemented by a CNN)). The CRAFT mode can be a single-stage detector, in that it can directly predict the location and a confidence score of text (e.g., alphanumeric) regions in a single forward pass. This manner of text detection is efficient and avoids a multi-stage pipeline often used in traditional text detection methods. The CRAFT model can be implemented by a pretrained CNN, like ResNet or VGG (Visual Geometry Group) to extract high level features from an input image. To enhance the discriminative powers for text regions, multiple CNN layers can be used to refine extracted features. For detecting the text region, a heatmap can be generated to localize text instances and predict the confidence score for each detected text region. The confidence score can indicate the CRAFT model's certainty of the detected region comprising actual text. The single-shot architecture of the CRAFT model can result in a faster and more efficient character detection. The CRAFT model is also relatively robust with respect variation in text appearance, such as font size and orientation.
[0032] Examples of characters (in the text region) that would constitute an identifier can include, but are not limited to alphabets, numbers, barcodes, or a combination thereof. If the character detection model determines that no identifier exists (for the shelf or inventory) in the received image (by way of failing to detect any character (e.g., alphabet, number, barcode) region in the image), the identifier detection module 1022 can indicate to a user (e.g., the warehouse operator) that no unique identifier could be detected / located through a notification / alert transmitted to the user device 118.
[0033] If the image includes the unique identifiers for the shelf and inventory, and the unique identifier regions are detected by the identifier detection module 1022, the detected characters (in the regions) are transmitted to an identifier recognition module 1023 for interpretation / recognition. In an example embodiment, the identifier recognition module 1023 can be implemented by a character recognition model (e.g., a convolutional recurrent neural network (CRNN)).
[0034] The character recognition model can be a Keras implementation of a CRNN, that was built using TensorFlow. The CRNN can combine the strengths of CNNs and recurrent neural networks (RNNs) to achieve high accuracy and robustness. The functioning of the CRNN can take place as follows. The CNN can perform feature extraction on the text region in the image, to extract high level features (e.g., edges, corners, and text patterns). The output of the CNN can be a feature map. These generated feature maps can be fed into a recurrent layer (e.g., LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit)), which can capture the sequential nature of text and learn dependencies between characters to generate a sequence of hidden states as the output. These hidden states can be fed into a Connectionist Temporal Classification (CTC) layer which transcribes hidden states into the sequence of characters. The training of the CRNN can be performed using a training dataset comprising images captured, for example by a camera on a forklift, within a warehouse.
[0035] In a warehouse, there may be different naming conventions for parcel and shelf, so that whether an identifier for an inventory can be easily distinguishable from an identifier for a shelf. Accordingly, the identifier recognition module 1023 may also comprise a rule-based engine that includes rules for identifying whether a recognized identifier belongs to a shelf or an inventory. For instance, one rule can be that shelf identifiers (e.g., as numbers) have 6 digits, whereas inventory numbers are limited to 4 digits. There can also be a rule with respect to the order of the digits in the shelf / inventory numbers. For instance, in the case of a shelf number including 6 digits, the first 2 digits among the 6 digits can represent an aisle number. The middle 2 digits can represent a horizontal position of the shelf in a rack. The last 2 digits can represent a vertical position of the shelf in a rack. In other embodiments, the identifier recognition module 1023 can include another feature extraction model (e.g., YOLO) that can extract features from an image to differentiate parcel identifiers from shelf identifiers based on, for example, the shape, size, context, and font type of the identifiers.
[0036] Further, the identifier recognition module 1023 can also associate a recognized shelf identifier with a recognized inventory identifier. In other words, if the captured image depicts 3 shelves, with each shelf having an inventory, the identifier recognition module 1023 can correctly determine which shelf is carrying which inventory. The captured image can include coordinates of the location of the shelf identifier and the inventory identifier. Therefore, once the identifier recognition module 1023 determines whether a certain identifier belongs to a shelf or an inventory, the identifier recognition module 1023 checks the coordinates of the identifiers in the image to create the identifier pairs of the inventory-shelf. As an inventory that is associated with a shelf is usually placed above a shelf, the coordinates of the inventory and shelf identifiers would be indicative of the inventory identifier being above the shelf identifier. In one embodiment, there may be a first space to the left of the shelf identifier and a second space to the right of the shelf identifier. The first and second spaces can be pre-specified. If the coordinates of the inventory identifier indicate that it is vertically above the first and second spaces, the inventory identifier can be associated with the shelf identifier. In other embodiments, there may be other manners in which the inventory may be placed in a shelf, which would then affect the way the identifier recognition module 1023 associates an inventory identifier with a shelf identifier. In an embodiment where the inventory is absent in the captured image, or a shelf is depicted to be empty, the identifier recognition module 1023 can assign a special identifier (a unique character / character sequence) for indicating that the shelf is free of any inventory. Accordingly, the special identifier of the inventory would form a part of the shelf-inventory identifier pair. The special identifier would be represented by the same character / character sequence in the database 116 as well.
[0037] Upon recognizing the identifier pairs of the inventory-shelf, the identifier comparison module 1024 can compare the identifier pairs with those that may be stored in the database 116. The content of the database 116 may be indicative of the shelf on which a specific inventory should be located. Similarly, the content of the database 116 may be indicative of the inventory that should be on a specific shelf. When the identifier comparison module 1024 engages in this comparison, the comparison can involve comparing at least one of (i) for the same inventory, if there is a match between the recognized shelf identifier (obtained from the image) and the stored shelf identifier (obtained from the database 116); or (ii) for the same shelf, if there is a match between the recognized inventory identifier (obtained from the image) and the stored inventory identifier (obtained from the database 116).
[0038] The comparison performed can be illustrated with reference to the below table.
[0039] The above table illustrates the recognized identifier pair being "500-A1", where "500" corresponds to the recognized inventory identifier and "A1" corresponds to the recognized shelf identifier, obtained from the received image. The above table also illustrates two stored identifier pairs in the database 116. The stored identifier pair 1 is "500-B1", where "500" corresponds to the inventory identifier and "B1" corresponds to the shelf identifier. The stored identifier pair 2 is "400-A1", where "400" corresponds to the inventory identifier and "A1" corresponds to the shelf identifier pair.
[0040] The identifier comparison module 1024 can compare the recognized identifier pair with the stored identifier pairs in the following manner. The identifier comparison module 1024 can check those stored identifier pairs in the database 116 that have the same inventory identifier as the recognized inventory identifier, and then compare the associated shelf identifiers to see if there is a match. With reference to the above table, the identifier comparison module 1024 can scan the database 116 and retrieve the stored identifier pair 1, as the inventor identifier "500" is the same in the recognized identifier pair and the stored identifier pair 1. Then, the identifier comparison module 1024 can compare the shelf identifier of the recognized identifier pair ("A1") with the shelf identifier of the stored identifier pair 1 ("B1"). Accordingly, the identifier comparison module 1024 can determine there is a mismatch between "A1" and "B1" for the same inventory "500" as the identifiers do not match, meaning that the inventory "500" is incorrectly placed on shelf "A1".
[0041] Similarly, the identifier comparison module 124 can scan the database 116 and retrieve the stored identifier pair 2, as the shelf identifier "A1" is the same in the recognized identifier pair and the stored identifier pair 2. Then, the identifier comparison module 1024 can compares the inventory identifier of the recognized identifier pair ("500") with the inventory identifier of the stored identifier pair 2 ("400"). Accordingly, the identifier comparison module 1024 can determine there is a mismatch between "500" and "400" as these identifiers do not match, meaning that the inventory "500" is incorrectly placed on shelf "A1".
[0042] The mismatch handling module 1025 can then undertake one or more response actions based on the inference / determination that an inventory is incorrectly placed. In one embodiment, the mismatch handling module 1025 can transmit an alert (e.g., audio or visual) to the user device 118 of the warehouse operator, wherein the alert can notify the warehouse operator of the item that is believed to be incorrectly placed. Accordingly, the warehouse operator can physically move the inventory to its associated shelf as per the database 116. In another embodiment, the mismatch handling module 1025 can update the stored identifier pair in the database 116 to match the recognized identifier pair. For example, with reference to the above table, the mismatch handling module 1025 can update the stored identifier pair 1 "500-B1" to now appear as "500-A1".
[0043] Fig. 4 illustrates a process flow diagram, according to an example embodiment disclosed herein. At step 402 in the process, the forklift operator can travel through left and right aisles on the respective side of forklift. The forklift can have cameras 112 located on the left and right sides of the forklift so as to capture images of the aisles on the respective side of the forklift. As the forklift travels between the left and right aisles, the cameras 112 continuously captures image of the respective aisles. In one or more example embodiments, the forklift can also include a memory module for storing the captured images prior to their transmission to the WMS 102. This can be useful in an instance where when the forklift is in motion, there is a drop in the connectivity of the camera 112 to the network 108, thereby rendering the camera 112 unable to transmit the captured images to the WMS 102, and the memory module can then store the captured images until the connectivity improves.
[0044] At step 404, the captured images, depicting the shelves, inventory, and their respective identifiers, are transmitted to and analyzed by the WMS 102. The WMS 102 can detect the identifier pairs for each inventory-shelf pair from the transmitted images, and also recognize the characters of the identifier pairs. Fig. 5 illustrates an example image with the identifiers of the inventory (e.g., 5101 & 5103) and shelves, respectively affixed to them.
[0045] At step 406, the WMS 102 can compare the recognized identifier pairs with the stored identifier pairs in the database 116 to see if there is a match. If there is a match, then no further action is taken. However, if there is no match or a mismatch, which is indicative of incorrect placement of an inventory, then one or more response actions are undertaken by the WMS 102 to rectify this incorrect placement.
[0046] Fig. 6 illustrates a screenshot of a table in the database 116, having two columns, with one column denoting the racks and another column denoting the parcel. Each row of the rack column represents a set of identifiers for shelves in a rack, and the same row for the parcel column represents a set of identifiers for the inventory on that rack. Fig. 5 illustrates that parcels identified as "5101" and "5103" are placed on the shelves identified as "010102" and "010402", respectively. When comparing Fig. 5 with Fig. 6, it can be seen that the aforementioned shelf identifiers "010102" and "010402", as per FIG. 6, are meant to store a different set of parcels, i.e., not parcels identified as "5101" and "5103". This can be considered as a mismatch, and thereby result in the one or more response actions being undertaken.
[0047] Fig. 7 illustrates a method 700 performed by the WMS 102, according to an example embodiment disclosed herein. At step 702, the WMS 102 receives an image captured by a camera.
[0048] At step 704, the WMS 102 determines, using a first machine learning model (MLM), if the received image includes a shelf and an inventory.
[0049] At step 706, on determining that the received image includes the shelf, the WMS 102 determines, using a second MLM, if the received image includes an identifier for the shelf and an identifier for the inventory (if determined to be present in the received image).
[0050] At step 708, the WMS 102 recognizes, using a third MLM, the identifiers for the shelf. The WMS 102 also recognizes an identifier for the inventory present in the image, using the third MLM. If the image does not include an inventory, the WMS 102 recognizes (assigns) a special identifier for this missing inventory.
[0051] At step 710, the WMS 102 compares the recognized identifier pairs of the shelf and inventory with the pre-existing (e.g., previously stored or determined) identifier pairs for a shelf and an inventory. The comparison of the identifier pairs of the shelf and inventory with the pre-existing identifier pairs can happen in at least one of the following manners. In a first manner, the recognized shelf identifier can initially be compared with the pre-existing shelf identifier (from the pre-existing identifier pair in the database) to see if there is a match. If there is a match, then the recognized inventory identifier can be compared with the inventory identifier (from the pre-existing identifier pair in the database) to see if there is a match, and if there is no match, this would be considered as a comparison mismatch. In a second manner, the recognized inventory identifier can initially be compared with the pre-existing inventory identifier (from the pre-existing identifier pair in the database) to see if there is a match. If there is a match, then the recognized shelf identifier can be compared with the shelf identifier (from the pre-existing identifier pair in the database) to see if there is a match, and if there is no match, this would be considered as a comparison mismatch.
[0052] At step 712, the WMS 102 initiates a mismatch handling response in view of a comparison mismatch.
[0053] It is to be noted that although the above-mentioned steps mention using a first MLM, a second MLM, and a third MLM, it is not to be construed as there necessarily being different / multiple MLMs. Instead, the terms "first", "second", and "third" are meant to denote the different functionalities performed by either the same MLM or a combination of MLMs.
[0054] <Technical Advantages / Effects> The following is a non-exhaustive list of technical advantages / effects achieved by the one or more embodiments disclosed herein. By filtering out images which do include a shelf, the additional processing resources for processing the image can then be omitted. In embodiments where the text detection is performed by the CRAFT detection model, the associated benefits with the CRAFT detection model are realized. In embodiments where the text recognition is performed using a combination of CNN and RNN, a higher accuracy and robustness is achieved.
[0055] In the drawings and specification, there have been disclosed exemplary embodiments of the present disclosure. Although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. It will be apparent to those having ordinary skill in this art that various modifications and variations may be made to the embodiments disclosed herein, consistent with the present disclosure, without departing from the spirit and scope of the present disclosure. Other embodiments consistent with the present disclosure will become apparent from consideration of the specification and the practice of the description disclosed herein.
[0056] <Supplementary notes> The whole or part of the example Aspects disclosed above can be described as, but not limited to, the following supplementary notes. (Supplementary note 1) A system (102), comprising: at least one memory (106); and at least one processor (104), configured to cause the system to: receive an image captured by a camera (112); determine, using a first machine learning model (MLM), if the received image includes a shelf and an inventory; on determining that the received image includes the shelf, determine, using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image; recognize the identifier for: the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier; compare the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs; and initiate one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs. (Supplementary note 2) The system according to supplementary note 1 further comprising: wherein the one or more response actions include at least one of: transmitting an alert to a user device indicating the comparison mismatch; or updating the one or more pre-existing identifier pairs to correspond to the recognized pair of identifiers. (Supplementary note 3) The system according to supplementary note 1 further comprising: wherein the camera is located on a side of a forklift. (Supplementary note 4) The system according to supplementary note 1 further comprising: wherein at least one of the first, second, or third MLMs are implemented by a convolutional neural network (CNN). (Supplementary note 5) A method (700), comprising: receiving (702) an image captured by a camera (112); determining (704), using a first machine learning model (MLM), if the received image includes a shelf and an inventory; on determining that the received image includes the shelf, determining (706), using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image; recognizing (708) the identifier for: the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier; comparing (710) the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs; and initiating (712) one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs. (Supplementary note 6) The method according to supplementary note 5 further comprising: wherein the initiating the one or more response actions include at least one of: transmitting an alert to a user device indicating the comparison mismatch; or updating the one or more pre-existing identifier pairs to correspond to the recognized pair of identifiers. (Supplementary note 7) The method according to supplementary note 5 further comprising: wherein the camera is located on a side of a forklift. (Supplementary note 8) The method according to supplementary note 5 further comprising: wherein at least one of the first, second, or third MLMs are implemented by a convolutional neural network (CNN).
[0057] This application is based upon and claims the benefit of priority from Indian Patent Application No. 202441080338, filed on October 22, 2024, the disclosure of which is incorporated herein in its entirety by reference.
[0058] 102 WAREHOUSE MANAGEMENT SYSTEM (WMS) 1021 IMAGE FILTERING MODULE 1022 IDENTIFIER DETECTION MODULE 1023 IDENTIFIER RECOGNITION MODULE 1024 IDENTIFIER COMPARISON MODULE 1025 MISMATCH HANDLING MODULE 104 PROCESSOR 106 MEMORY 108 NETWORK 112 CAMERA 116 DATABASE 118 USER DEVICE
Claims
1. A system, comprising: at least one memory; and at least one processor, configured to cause the system to: receive an image captured by a camera; determine, using a first machine learning model (MLM), if the received image includes a shelf and an inventory; on determining that the received image includes the shelf, determine, using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image; recognize the identifier for: the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier; compare the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs; and initiate one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs.
2. The system as claimed in claim 1, wherein the one or more response actions include at least one of: transmitting an alert to a user device indicating the comparison mismatch; or updating the one or more pre-existing identifier pairs to correspond to the recognized pair of identifiers.
3. The system as claimed in claim 1, wherein the camera is located on a side of a forklift.
4. The system as claimed in claim 1, wherein at least one of the first, second, or third MLMs are implemented by a convolutional neural network (CNN).
5. A method, comprising: receiving an image captured by a camera; determining, using a first machine learning model (MLM), if the received image includes a shelf and an inventory; on determining that the received image includes the shelf, determining, using a second MLM, if the received image includes an identifier for: the shelf; and the inventory, if determined to be present in the received image; recognizing the identifier for: the shelf, using a third MLM, and the inventory, wherein if the received image includes the inventory, the third MLM is used for recognizing the identifier of the inventory, and wherein if the received image does not include the inventory, the identifier of the inventory is recognized as a special identifier; comparing the recognized pair of identifiers of the shelf and the inventory with one or more pre-existing identifier pairs; and initiating one or more response actions upon there being a comparison mismatch between the recognized pair of identifiers and the one or more pre-existing identifier pairs.
6. The method as claimed in claim 5, wherein the initiating the one or more response actions include at least one of: transmitting an alert to a user device indicating the comparison mismatch; or updating the one or more pre-existing identifier pairs to correspond to the recognized pair of identifiers.
7. The method as claimed in claim 5, wherein the camera is located on a side of a forklift.
8. The method as claimed in claim 5, wherein at least one of the first, second, or third MLMs are implemented by a convolutional neural network (CNN).
Citation Information
Patent Citations
Goods shelf management method and system based on image recognition
CN118052991A
Augmented reality provision system, information processing terminal, information processor, augmented reality provision method, information processing method, and program
JP2012074014A
System, information processor, information processing method, and program
JP2018002331A
Display state determination system
JP2019185684A
Systems and methods for locating, identifying and counting items
JP2019513274A