Devices and methods that use machine learning to detect and decode a label

A machine learning-based system dynamically adjusts focus settings to enhance label detection and identification in retail environments, addressing inefficiencies in conventional imaging systems by ensuring markers are in focus before processing.

DE102025148286A1Pending Publication Date: 2026-05-28ZEBRA TECHNOLOGIES CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
ZEBRA TECHNOLOGIES CORP
Filing Date
2025-11-20
Publication Date
2026-05-28

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Devices and methods for detecting and decoding a marking are disclosed herein. The method captures, using a device and a current focus setting, a current image of an area containing one or more markings. The method determines a first distance between the device and one or more markings present in the current image, based on the current focus setting during image capture. Using a trained model, the method detects the one or more markings present in the current image and determines, for a current marking located beneath the one or more markings, a second distance between the device and the current marking, based on a known size of the current marking and / or a feature of the current marking.The procedure determines whether the current marking is in focus, based on the first distance, the second distance and at least one attribute of the device.
Need to check novelty before this filing date? Find Prior Art

Description

background

[0001] An establishment (e.g., a grocery store, a convenience store, a retail store, etc.) can include at least one support structure (e.g., a display module) with one or more support surfaces (e.g., shelves) for supporting and displaying one or more objects (e.g., products). For example, objects on a display module can be oriented so that they are positioned on a leading edge of a support surface of the display module and are oriented in such a way that they are identifiable (e.g., an employee or customer can observe an object that is associated with and oriented to a identifier on a support surface, such as a stock keeping unit (SKU) or a product code). An employee of an establishment can use a device (e.g., a display stand) to display the product.A smartphone, tablet, mobile computer, head-mounted display, scanner, portable computing device, or similar device may be used to identify each object displayed on a display module. For example, an employee can process a tag (such as scanning an SKU or product code) associated with each object. Additionally, based on this identification, an employee can perform other tasks, including but not limited to locating, selecting, and / or restocking each object displayed on a display module. Brief description of the different views of the drawings

[0002] The accompanying figures, in which the same reference numerals refer to identical or functionally similar elements in the individual views, are integrated into the specification together with the following detailed description and form a part thereof, serving to further illustrate embodiments of concepts that include the claimed invention and to explain various principles and advantages of these embodiments. Fig. Figure 1 is a representation illustrating an embodiment of a system of the present disclosure. Fig. Figures 2A-C are representations illustrating another embodiment of a system of the present disclosure. Fig. 3 is a representation showing components of the computing device of Fig. 1 and Fig. 2C illustrates. Fig. Figure 4 is a flowchart illustrating processing steps performed by an embodiment of the present disclosure. Fig. Figures 5A-B are illustrations that depict characteristics of an embodiment of the present disclosure.

[0003] Experts will recognize that elements in the figures are illustrated for the sake of simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some of the elements in the figures may be exaggerated relative to other elements to help improve the understanding of embodiments of the present invention.

[0004] Where appropriate, the apparatus and process components have been represented by conventional symbols in the drawings, which show only those specific details relevant to understanding the embodiments of the present disclosure, so as not to obscure the disclosure with details that would be obvious to persons skilled in the art referring to the present description. Detailed description

[0005] As mentioned above, a facility employee can use a device (e.g., a smartphone, tablet, mobile computer, head-mounted display, scanner, portable computing device, or similar) to identify each object displayed on a screen. For example, an employee can process an identifier (e.g., scan an SKU or product code) associated with each object. Additionally, based on this identification, an employee can perform other tasks, including, but not limited to, locating, selecting, and / or repopulating each object displayed on a screen.

[0006] Scanning a tag associated with each object is a manual process (requiring human intervention) and, as such, can be time-consuming, costly (increased associated labor costs), and prone to human error (scanning the wrong tag). For example, a display module might have hundreds, if not thousands, of tags attached to it, so manually processing each tag (scanning an SKU or product code) associated with each object to identify, locate, select, and / or determine if each object requires restocking can be time-consuming and expensive.

[0007] Conventional imaging systems for label processing can capture an image of a label and / or an associated object and process the label present in the image. However, these systems can be costly to deploy and use in a facility and / or provide insufficient performance, reducing the efficiency and overall on-timeness of label processing. For example, high-resolution imaging systems are costly to deploy and use in a facility because these systems require one or more high-resolution cameras to capture an image of sufficient quality (e.g., sufficient resolution) to detect and identify (e.g., recognize and / or decode) one or more labels present in the image that are positioned at varying distances from the cameras.

[0008] In another example, low-resolution imaging systems use one or more low-resolution cameras. However, low-resolution cameras generally capture an image of insufficient quality (e.g., insufficient resolution) to detect and identify (e.g., recognize and / or decode) one or more tags present in the image that are positioned at varying distances from the cameras. For example, low-resolution cameras may capture images containing one or more tags that are blurry and illegible (e.g., one or more identifiers of a tag cannot be recognized and / or decoded). Furthermore, these low-resolution imaging systems may still attempt to recognize and / or decode illegible tags, consuming significant amounts of performance and processing resources without providing any benefit.

[0009] Proposed techniques for mitigating shortcomings in low-resolution imaging systems include using different focus settings (e.g., autofocus, fixed focus, and multi-focus) of one or more low-resolution cameras to capture an image of sufficient quality to detect and identify (e.g., recognize and / or decode) one or more markings present in the image that are positioned at varying distances from the cameras. However, these proposed techniques may also produce markings that are blurry and illegible, and systems employing the proposed techniques may still attempt to recognize and / or decode illegible markings.For example, an autofocus setting might focus on an object and a corresponding label within the center of a low-resolution camera's field of view (FOV), rendering one or more labels positioned outside the center of the FOV blurry and illegible. In another example, cycling through multiple focus settings of a low-resolution camera might result in each focus setting failing to align with one or more labels positioned at varying distances from the camera, again rendering those labels blurry and illegible.As noted above, low-resolution imaging systems using the proposed techniques may still attempt to detect and / or decode illegible markings, consuming significant amounts of power and processing resources without providing any benefit.

[0010] Additionally, the volume of images and the multitude of labels present in each image often leads to redundant, illegible labels, reducing the processing efficiency of an imaging device and / or imaging system and the efficiency of the detection and identification (e.g., recognition and / or decoding) process of the labels present in each image.

[0011] As such, conventional systems suffer from a general lack of versatility, since they cannot automatically and dynamically determine whether a marker is in focus, process the marker when it is in focus, or modify a device's focus setting when the marker is out of focus. For example, these systems cannot automatically and dynamically determine whether a marker present in an image captured by a device is in focus according to a first distance between the device and the marker based on a current focus setting of the device, a second distance between the device and the marker based on a known size of the marker and / or a feature of the marker, and at least one attribute of the device.Additionally, these systems cannot automatically and dynamically determine a different focus setting of a device based on at least the second distance of the marking if the marking is not in focus.

[0012] Overall, this lack of versatility causes conventional systems to provide inadequate performance and reduces the efficiency and overall on-timeliness of marking processing. Therefore, an objective of the present disclosure is to eliminate these and other problems with conventional systems and methods by means of systems and methods that can detect a marking, determine whether a marking is in focus, and process the marking when the marking is in focus, or modify a focus setting of a device when the marking is out of focus.

[0013] According to the above and with the disclosure herein, the present disclosure includes improvements in computer functionality or improvements in other technologies, at least because the present disclosure describes that, for example, image processing devices and / or systems and their associated various components can be improved or extended with the disclosed dynamic system features and methods that automatically and dynamically detect a marking, determine whether a marking is in focus, and process the marking if the marking is in focus, or modify a focus setting of a device if the marking is out of focus.

[0014] That is to say, the present disclosure describes improvements to the functioning of an imaging device and / or an image processing device and / or a system and / or "any other technology or any other technical field" (e.g., the field of image processing). For example, the disclosed dynamic system features and methods improve and enhance the detection and identification of a mark by introducing the automatic and dynamic determination of whether a mark is in focus and processing the mark when the mark is in focus, or modifying a focus setting of a device when the mark is out of focus, in order to mitigate (if not eliminate) operational errors and eliminate inefficiencies typically experienced over time by systems lacking such features and methods.This improves the state of the art, at least because such previous systems are inefficient, as they lack the ability to automatically and dynamically determine whether a marking is in focus and to process the marking when the marking is in focus, or to modify a focus setting of a device when the marking is not in focus.

[0015] Furthermore, the present disclosure applies various features and functionalities, as described herein, with or using a specific machine, e.g., a processor, a device, and / or other hardware components, as described herein. In addition, the present disclosure includes other specific features beyond what is well understood as routine, conventional activity in the field, or the addition of unconventional steps, which, in various embodiments, demonstrate certain useful applications, e.g., image processing protocols of a device for automatically and dynamically determining whether a marking is in focus and processing the marking when the marking is in focus, or modifying a focus setting of a device when the marking is out of focus.

[0016] Accordingly, it would be highly advantageous to develop a system and method that can automatically and dynamically detect a marking, determine whether the marking is in focus, and process the marking if it is in focus, or modify the focus setting of a device if the marking is out of focus. The systems and methods of this disclosure address these and other needs.

[0017] In one embodiment, the present disclosure is directed to a method. The method comprises: capturing, by a device using a current focus setting, a current image of an area, wherein the area has one or more markings present therein; determining a first distance between the device and one or more markings present in the current image, based on the current focus setting of the device during the acquisition of the current image; detecting, using a trained machine learning model, one or more markings present in the current image, wherein each marking has one or more identifiers; and determining whether the one or more markings are to be evaluated.In response to determining that one or more markings are not evaluated, determine, for a current marking under the one or more markings, a second distance between the device and the current marking based on a known size of the current marking and / or a feature of the current marking; determine whether the current marking is in focus, based on the first distance, the second distance of the current marking and at least one attribute of the device; and in response to determining that the current marking is in focus, process the current marking.

[0018] In one embodiment, the present disclosure relates to a device comprising an imaging assembly; one or more processors; and a non-transitory, computer-readable memory coupled to the one or more processors. The memory stores instructions which, when executed by the one or more processors, cause the one or more processors to: receive a current image of an area, captured by the imaging assembly using a current focus setting, wherein the area has one or more markings present therein; determine a first distance between the device and one or more markings present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image;Detect, using a trained machine learning model, one or more labels present in the current image, each label having one or more identifiers; determine whether the one or more labels are evaluated; in response to determining that the one or more labels are not evaluated, determine, for a current label below the one or more labels, a second distance between the device and the current label based on a known size of the current label and / or a feature of the current label; determine whether the current label is in focus, based on the first distance, the second distance of the current label, and at least one attribute of the imaging assembly; and in response to determining that the current label is in focus, process the current label.

[0019] In one embodiment, the present disclosure is directed to a non-transitory computer-readable medium. The non-transitory computer-readable medium stores instructions on it which, when executed by one or more processors, cause the one or more processors to: receive a current image of an area, captured by an imaging assembly using a current focus setting, wherein the area has one or more markings present therein; determine a first distance between the device and one or more markings present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image;Detect, using a trained machine learning model, one or more labels present in the current image, each label having one or more identifiers; determine whether the one or more labels are evaluated; in response to determining that the one or more labels are not evaluated, determine, for a current label below the one or more labels, a second distance between the device and the current label based on a known size of the current label and / or a feature of the current label; determine whether the current label is in focus, based on the first distance, the second distance of the current label, and at least one attribute of the imaging assembly; and in response to determining that the current label is in focus, process the current label.

[0020] Referring to the drawings, Fig. Figure 100 illustrates an embodiment of a system of the present disclosure. The system can be used in an establishment (e.g., a grocery store, a convenience store, a retail store, etc.). For example, the system can be used in a section of the establishment accessible to an employee, which may be referred to as the back of the establishment (e.g., a storage room, a stockroom, an inventory room, etc.), and / or in a section of the establishment accessible to a customer, which may be referred to as the front of the establishment. Objects received at the establishment, e.g., via a receiving tray or the like, are generally placed on a support structure (e.g., a display module) with one or more support surfaces (e.g., shelves) in a back area until it is necessary to replenish the relevant objects at the front of the establishment.An employee can retrieve the items requiring refilling from the back and transport these items to the appropriate locations at the front of the facility.

[0021] As in Fig. As shown in Figure 1, the device includes at least one support structure, such as a display module 102, with one or more support surfaces 104-1, 104-2, and 104-3 (collectively referred to as support surfaces 104 and generally as support surface 104) that support and display objects 106-1, 106-2, and 106-n (collectively referred to as objects 106 and generally as object 106). The objects 106 can be of different types, so that object 106-1 differs from objects 106-2 and 106-n, object 106-2 differs from object 106-n, and so on. Furthermore, an object 106 can be grouped with one or more objects. For example, object 106-1 is grouped with eight objects 106-1, and object 106-2 is grouped with three objects 106-2.Objects 106-1, 106-2, and 106-3 can each be identified by object identifiers 108-1, 108-2, and 108-n, respectively (collectively referred to as identifiers 108 and generally as identifier 108). An identifier 108 can contain one or more identifiers (e.g., a barcode, a numeric string, an alpha string, and an alphanumeric string). For example, identifier 108 can contain an SKU and / or a product code (e.g., a Universal Product Code (UPC)) or the like.

[0022] The system may include a computing device 116 (e.g., a smartphone, tablet, mobile computer, head-mounted display, scanner, portable computing device, or the like). The computing device 116 may be operated by an associate at the facility and includes at least one imaging assembly (e.g., a camera) with a field of view (FOV) 120 and a display 124. Alternatively, the computing device 116 may be an imaging assembly (e.g., a camera). For example, the computing device 116 may be a camera mounted on a first display module 102 and have an FOV 120 of at least one section of a second display module 102 positioned opposite it.In another example, the computing device 116 can be a camera mounted in an overhead position above a display module 102 and having a field of view (FOV) 120 of at least one section of the display module 102 positioned below the computing device 116. The computing device 116 can be manipulated such that an imaging assembly can view at least one section of the display module 102 within the FOV 120 and can be configured to capture an image or a stream of images of the display module 102. From such images, the computing device 116 can detect and identify (e.g., recognize and / or decode) a marker 108 associated with an object 106.The computing device 116 can also generate and / or update a log associated with the display module 102, based on the identified identifiers 108, wherein the log provides an inventory of objects 106 positioned on the display module 102. The computing device 116 can exchange data with the server 130, for example, via a network 142, which is implemented as any suitable combination of local and wide area networks.

[0023] It goes without saying that the system can also be used in any suitable environment. For example, the system can also be used in a logistics environment, as described below in relation to Fig. 2A-C described in more detail.

[0024] The server 130 can include a processor 132 (e.g., one or more central processing units (CPUs)) connected to a non-transient, computer-readable storage medium, such as a memory 134 and an interface 140. The memory 134 includes a combination of volatile memory (e.g., random-access memory or RAM) and non-volatile memory (e.g., read-only memory or ROM, electrically erasable programmable read-only memory or EEPROM, or flash memory). The processor 132 and the memory 134 each comprise one or more integrated circuits.

[0025] Memory 134 stores computer-readable instructions for execution by processor 132. Memory 134 stores an image processing application 136 (also simply referred to as the application 136) which, when executed by processor 132, configures processor 132 to perform various functions, described in more detail below. These functions relate to automatically and dynamically detecting a marker 108, determining whether a marker 108 is in focus, and processing the marker 108 if it is in focus, or modifying a focus setting of a computing device 116 if the marker 108 is out of focus. For example, when executed by processor 132, application 136 configuresthe processor 132 to: receive a current image of an area, captured by an imaging assembly (not shown) of a computing device 116 using a current focus setting, wherein the area has one or more markings 108 present therein; determine a first distance between the device and one or more markings 108 present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image; detect, using a trained machine learning model, one or more markings 108 present in the current image, each marking 108 having one or more identifiers; determine whether the one or more markings 108 are evaluated; in response to determining that the one or more markings are not evaluated, determinefor a current marking under one or more markings, a second distance between the device 116 and the current marking 108 based on a known size of the current marking 108 and / or a feature of the current marking 108; determining whether the current marking 108 is in focus, based on the first distance, the second distance of the current marking 108 and at least one attribute of the imaging assembly; and, in response to determining that the current marking 108 is in focus, processing the current marking 108. As described below, this functionality can also be performed by the processor 202 of the device 116.

[0026] Application 136 can also be implemented as a series of different applications in other examples. Those skilled in the art will recognize that the functionality implemented by the processor 132 through the execution of application 136 can also be implemented in other embodiments by one or more specially designed hardware and firmware components, such as field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and the like.

[0027] Memory 134 also stores a database 138. Database 138 can store one or more image datasets of a variety of identifiers 108 (e.g., for training a machine learning model to detect, classify, and / or decode an identifier 108 and one or more identifiers thereof). Database 138 can also store one or more captured images (e.g., historical data) of previously detected identifiers 108, the images of which can be used to train the machine learning model to detect an identifier 108 based on its distinguishing features (e.g., size, shape, color, or the like). It is understood that database 138 can be stored in memory (not shown) of the computing device 116.

[0028] Server 130 also includes a communication interface 140, which enables Server 130 to communicate with other computing devices, including Computing Device 116, via Network 142. The communication interface 140 includes suitable hardware components (e.g., transceivers, ports, and the like) and appropriate firmware according to the communication technology used by Network 142.

[0029] Fig. 3 is a representation 200, the components of the computing device 116 of Fig. 1 and Fig. Figure 2C illustrates this. As mentioned above, the computing device 116 can be, but is not limited to, a smartphone, tablet, mobile computer, head-mounted display, scanner, portable computing device, or the like. The computing device 116 can be operated by an employee in the facility and includes at least one imaging assembly 208 (e.g., a camera) with a field of view 120 and a display 124. Alternatively, the computing device 116 can be an imaging assembly 208 (e.g., a camera). The computing device 116 can capture an image or a stream of images of an object 106 and an associated label 108. As shown in Figure 2C, the computing device 116 can be a smartphone, tablet, mobile computer, head-mounted display, scanner, portable computing device, or the like. Fig. As shown in Figure 2, the computing device 116 includes a processor 202, a memory 204, a display 124, an input / output 206, an imaging assembly 208, sensor(s) 210 and an interface 212.

[0030] The processor 202 can be one or more CPUs, a graphics processing unit (GPU), or a combination thereof, and is communicatively coupled to a memory 204 (e.g., a non-transient, computer-readable storage medium implemented as a suitable combination of volatile and non-volatile memory elements), a display 124, an input / output 206, an imaging assembly 208, sensor(s) 210, and an interface 212. The processor 202 and the memory 204 each comprise one or more integrated circuits.

[0031] Memory 204 can store a variety of computer-readable instructions, for example, in the form of an image processing application 214 (also simply referred to as the application 214), which, when executed by processor 202, configures processor 202 to perform various functions, which are described in more detail below and relate to automatically and dynamically detecting a marker 108, determining whether a marker 108 is in focus, and processing the marker 108 if it is in focus, or modifying a focus setting of a computing device 116 if the marker 108 is out of focus. For example, when executed by processor 202, application 214 configuresthe processor 202 to: receive a current image of an area, captured by an imaging assembly (not shown) of a computing device 116 using a current focus setting, wherein the area has one or more markings 108 present therein; determine a first distance between the device and one or more markings 108 present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image; detect, using a trained machine learning model, one or more markings 108 present in the current image, each marking 108 having one or more identifiers; determine whether the one or more markings 108 are evaluated; in response to determining that the one or more markings are not evaluated, determinefor a current marking under one or more markings, a second distance between the device 116 and the current marking 108 based on a known size of the current marking 108 and / or a feature of the current marking 108; determining whether the current marking 108 is in focus, based on the first distance, the second distance of the current marking 108 and at least one attribute of the imaging assembly; and in response to determining that the current marking 108 is in focus, processing the current marking 108.

[0032] Application 214 can also be implemented as a series of different applications in other examples. Those skilled in the art will recognize that the functionality implemented by the processor 202 through the execution of application 214 can, in other embodiments, also be implemented by one or more specially designed hardware and firmware components, such as FPGAs, ASICs, and the like. As noted above, in some examples, memory 204 can also store database 138, instead of database 138 being stored on server 130.

[0033] The display 124 can be any suitable display, including but not limited to a light-emitting diode (LED) display, an organic LED display, a liquid crystal display (LCD), and a touchscreen display.

[0034] The imaging assembly 208 (e.g., a camera) can include a suitable sensor (e.g., an accelerometer, a gyroscope, a magnetometer, an altimeter, or a proximity sensor) or a combination of sensors. Alternatively, the imaging assembly 208 and the sensor(s) 210 can be independent of each other. In another alternative, the device 116 can be an imaging assembly 208 (e.g., a camera) with a field of view (FOV) and one or more sensors 210 integrated therein or coupled to it.

[0035] The input / output 206 can be a device connected to the processor 202. The input device 206 is configured to receive an input (e.g., from a user of the device 116) and provide the processor 202 with data representative of the received input. The input device 206 can include any or a suitable combination of a touchscreen integrated into the display 124, a keyboard, a microphone, and the like. In addition to the display 124, the device 116 can also include an output 206. The output 206 can be a device connected to the processor 202. The output device 206 is configured to receive an output (e.g., a signal from a processor 202) and provide data representative of the received output.The output device 206 can include any or a suitable combination of a speaker, a headset, a notification LED, and the like.

[0036] The communication interface 212 enables communication between the device 116 and other computing devices (e.g., a server 130) via suitable short-range connections, networks such as network 142, and the like. The interface 212 therefore includes suitable hardware elements that execute suitable software and / or firmware to communicate via network 142 and / or other communication links.

[0037] The sensor(s) 210 can include any one or a suitable combination of sensors configured to facilitate the determination of a focus setting of the imaging assembly 208 and / or a distance between the device 116 and a marker 108 during image acquisition. For example, the sensor(s) 214 can include an inertial navigation system incorporating one or more accelerometers, gyroscopes, magnetometers, altimeters, or proximity sensors. In this way, the sensor(s) 214, in conjunction with one or more other components (e.g., the imaging assembly 208) of the device 116, provide spatial computation (e.g., Google's ARCore) to determine the position and orientation of a user-operated device 116.

[0038] Fig. Figures 2A-C are illustrations depicting another embodiment of a system of the present disclosure. In one embodiment, the system can be used in a logistics environment. In logistics operations, a wide variety of objects 106, such as packages and other cargo, can be transported in a container 150, which is implemented as a storage unit attached to or stored in a vehicle 151 from point of origin to point of destination. Each object 106 can have a respective marking 108 affixed to it to identify the object 106 and facilitate its transport. A marking 108 can include one or more identifiers (e.g., a barcode, a numeric string, an alpha string, and an alphanumeric string). An operator of the vehicle 151 can use the computing device 116 to capture an image or a stream of images of the interior of the container 150.From such images, the computing device 116 can detect and identify (e.g. recognize and / or decode) a marking 108 attached to an object 106.

[0039] Fig. 2A-B are illustrations depicting a Container 150. Fig. 2A is a representation illustrating a top view of container 150, and Fig. Figure 2B is a diagram illustrating a side view of container 150. As shown in Fig. 2A and Fig. As shown in Figure 2B, the container 150 is a storage unit attached to a vehicle 151 (e.g., a truck). In alternative embodiments, the container 150 can be a storage unit attached to or stored in a vehicle 151, including a trailer attached to a platform having one or more sets of wheels and a trailer hitch assembly for towing by the vehicle, or a unit load device (ULD) stored in an aircraft, or a storage area integrated into at least one section of a vehicle 151, including an all-terrain vehicle (SUV), van, truck, commercial vehicle, Sprinter, or step van.Container 150 can include a door opening 152, a corridor 154, and at least one supporting structure such as a shelf 156 (two shelves 156 at approximately the same height are shown) on which objects 106 can be positioned. As in . Fig. 2A and Fig. As shown in 2B, objects 106 are loaded into container 150. Additionally, as shown in Fig. Figure 2A shows that objects 106 are positioned at different depths on a shelf 156.

[0040] Fig. Figure 2C is illustration 170, which illustrates an image capture carried out by an embodiment of the present disclosure. As in Fig. As shown in Figure 2C, the computing device 116 has a display 124 and an imaging assembly 208 (not shown) with a known field of view (FOV) 120. The computing device 116 can acquire an image or a stream of images of one or more objects 106 within the FOV 120 of the imaging assembly 208, such that the image or stream of images can include one or more objects 106 and respective identifiers 108 thereof. ... Fig. As shown in Figure 2C, an FOV 120 of the imaging assembly 208 (not shown) can capture an image or a stream of images including object 106-6 with a label 108-6 and object 106-7 with a label 108-7.

[0041] Fig. Figure 4 is a flowchart illustrating processing steps performed by an embodiment of the present disclosure. The processing steps are described in connection with their execution in the system (e.g., by the device 116 or the server 130 in conjunction with the device 116). In general, by performing the processing steps, the system can automatically and dynamically detect a marker 108, determine whether the marker 108 is in focus, and process the marker 108 if it is in focus, or modify a focus setting of a device 116 if the marker 108 is not in focus.

[0042] For example, the system can, using a device and a current focus setting, receive a current image of an area, wherein the area contains one or more markings; determine, for the current image, an initial distance between the device and the one or more markings based on the device's current focus setting during image acquisition; using a trained machine learning model that detects one or more markings present in the current image, wherein each marking has one or more identifiers; determine whether the one or more markings are evaluated;In response to determining that one or more labels are being evaluated, determine a second distance between the device and the current label for a current label under the one or more labels, based on a known size of the current label and / or a feature of the current label; determine whether the current label is in focus, based on the first distance, the second distance, and at least one attribute of the device; and in response to determining that the current label is in focus, process the current label. Alternatively, in response to determining that the current label is not in focus, the current label and the second distance of the current label can be added to a list of unprocessed labels.

[0043] With reference to Fig. In step 302, the system receives a current image of an area (e.g., a display module 102, a container 150, or the like) from a device 116 using a current focus setting. The area has one or more markings 108 present therein. For example, the device 116 can capture an image or a stream of images of one or more objects 106 and associated markings 108 thereof within the FOV 120 of the device 116 or the FOV 120 of an imaging assembly 208 thereof, wherein the objects 106 are positioned on a support surface 104 of a display module 102.In another example, the device 116 can capture an image or a stream of images of one or more objects 106 and associated markings 108 thereof within the FOV 120 of the device 116 or the FOV 120 of an imaging assembly 208 thereof, wherein the objects 106 are positioned on a support surface 156 of a container 150. The device 116 can include, but is not limited to, a smartphone, a tablet, a mobile computer, a head-mounted display, a scanner, or a portable computing device. The current focus setting can be a fixed focus setting of the device 116 or of an imaging assembly 208 thereof, or an autofocus setting of the device 116 or of an imaging assembly 208 thereof.

[0044] In step 304, the system determines a first distance between the device 116 and one or more markings 108 present in the current image, based on the current focus setting of the device 116 during the acquisition of the current image. For example, the first distance can be a known distance between the device 116 and the one or more markings 108, based on the current focus setting of the device 116 or an imaging assembly 208 of the device 116. As mentioned above, the device 116 can include one or more sensor(s) 210, wherein the sensor(s) 214 can comprise an inertial navigation system that includes one or more accelerometers, gyroscopes, magnetometers, altimeters, or proximity sensors. In this way, the sensor(s) 214, in conjunction with one or more other components (e.g.,The imaging assembly 208 of the device 116 provides a spatial calculation (e.g., Google's ARCore) to determine the position and orientation of a device 116 used by a user. Optionally, the system can use a spatial calculation to confirm and / or refine a specific initial distance between the device 116 and the one or more markers 108. Additionally, and as described below, the system can optionally use a spatial calculation to confirm and / or refine a specific secondary distance between the device 116 and an actual marker 108 located beneath the one or more markers 108.

[0045] In step 306, the system, using a trained machine learning model, detects the one or more labels 108 present in the current image, where each label 108 has one or more identifiers. For example, the system can detect a current label 108 among the one or more labels 108 using a trained machine learning model by generating a bounding box corresponding to the current label 108 and determining a pixel size of the current label 108 based on a pixel length and pixel height of the bounding box. The size of a label 108 within a bounding box can change based on the initial distance and / or skewness of the view.As noted above, the database 138 can store one or more image datasets of a variety of labels 108 for training the machine learning model to detect, classify, and / or decode a label 108 and one or more of its identifiers. The database 138 can also store one or more captured images (e.g., historical data) of previously detected labels 108, the images of which can be used to train the machine learning model to detect a label 108 based on its distinguishing features (e.g., size, shape, color, or the like).As such, the system can train a machine learning model to detect one or more tags 108 based on at least one of image datasets containing images of one or more tag types, each tag type having a known size among other known and / or distinguishing features or historical data that includes one or more previously detected tags 108. The tag 108 can include one or more identifiers (e.g., a barcode, a numeric string, an alpha string, and an alphanumeric string). For example, the tag 108 can include an SKU and / or a product code (e.g., a UPC) or the like. Alternatively, the tag 108 can be an image (e.g., an image of an object, a landscape, an individual, or any suitable image).

[0046] Fig. Figures 5A-B are illustrations that depict characteristics of an embodiment of the present disclosure. Fig. 5A is a representation 400 illustrating a label 402 which contains numeric strings 404a, 404b, 404c and 404d and alphanumeric strings 406a, 406b and 406c. Fig. 5B is a representation 420 illustrating a label 422. The label 422 is a barcode consisting of parallel lines with varying widths, spacing, and sizes. As described in more detail below, the system can process a label 108 associated with an object 106.

[0047] With renewed reference to Fig.In step 308, the system determines whether each of the one or more labels 108 is evaluated. If the system determines that each of the one or more labels 108 is evaluated, the process proceeds to step 318 (described in more detail below). Alternatively, if the system determines that each of the one or more labels 108 is not evaluated, the process proceeds to step 310. As described in more detail below, some processing steps 300 may be repeated until each of the one or more labels 108 is evaluated.For example, part of the processing steps 300 may be repeated until each of the one or more tags 108 is processed in step 314, each of the one or more tags 108 and respective second spaces are added to a list of unprocessed tags 108 in step 316, or part of the one or more tags 108 is processed in step 314 and the remaining part of the one or more tags 108 and respective second spaces are added to the list of unprocessed tags 108 in step 316.

[0048] In step 310, the system determines, for a current marking 108 among one or more markings 108, a second distance between the device 116 and the current marking 108 based on a known size of the current marking 108 and / or a feature of the current marking 108. As noted above, a marking 108 can have a type, each marking type having a known size among other known and / or distinguishing features (e.g., shape, color, or the like). Additionally, a marking 108 can have a pixel size based on a pixel length and pixel height of a bounding box corresponding to the marking 108. As described above, the device 116 can include one or more sensor(s) 210, the sensor(s) 214 being able to include an inertial navigation system.In this way, the sensor(s) 214, in conjunction with one or more other components (e.g., the imaging assembly 208) of the device 116, provide a spatial calculation (e.g., Google's ARCore) to determine the position and orientation of a device 116 used by a user. Optionally, the system can use a spatial calculation to confirm and / or refine a specific second distance between the device 116 and the current marker 108.

[0049] In step 312, the system determines whether the current marker 108 is in focus, based on the first distance, the second distance of the current marker 108, and at least one attribute of the device 116. The at least one attribute of the device 116 can be a depth of field of the device 116 (e.g., a camera) or a depth of field of an imaging assembly 208 of the device 116. If the system determines that the current marker 108 is in focus, the process proceeds to step 314. Alternatively, if the system determines that the current marker 108 is not in focus, the process proceeds to step 316.

[0050] In step 314, the system processes the current identifier 108. For example, the system can decode one or more identifiers of the current identifier 108 and select a decoded identifier, or, if the current identifier 108 contains more than one identifier, select a decoded identifier that corresponds to a predetermined symbology (e.g., including, but not limited to, a Universal Product Code (UPC), European Article Number (EAN), Code 128, Code 39, and Data Matrix) and / or barcode data structure. In one embodiment, the system may only need to decode one initial identifier among the one or more identifiers if the initial identifier corresponds to the predetermined symbology and / or barcode data structure. In this way, the system can avoid processing additional identifiers.In another example, the system can use character recognition to detect one or more identifiers of the current identifier 108 and select a detected identifier, or, if the current identifier 108 contains more than one identifier, select a detected identifier that matches a predetermined string structure. The process then returns to step 308.

[0051] Referring again to step 312, if the system determines that the current label 108 is not in focus, the process proceeds to step 316. In step 316, the system indicates that the current label 108 is not in focus and adds the current label 108 and the second spacing of the current label 108 to a list of unprocessed labels 108. In this way, the system improves the processing efficiency of the device 116 by eliminating the detection, recognition, and / or decoding of illegible labels, which reduces the processing efficiency of the device 116 and / or the system and the efficiency of the detection and identification (e.g., recognition and / or decoding) process of labels 108 present in each image. The process then returns to step 308.

[0052] In step 308, the system determines whether each of the one or more tags 108 is evaluated. As mentioned above, some processing steps 300 may be repeated until each of the one or more tags 108 is evaluated. For example, some processing steps 300 may be repeated until each of the one or more tags 108 is processed in step 314, each of the one or more tags 108 and their respective second intervals are added to a list of unprocessed tags 108 in step 316, or part of the one or more tags 108 is processed in step 314 and the remaining part of the one or more tags 108 and their respective second intervals are added to the list of unprocessed tags 108 in step 316.In this way, the system can process each of the one or more labels 108 detected in the current image, even if a detected label 108 is not in focus in the current image, thereby increasing the image processing efficiency of the system.

[0053] When the system determines that each of the one or more markings 108 is being evaluated, the process proceeds to step 318. In step 318, based on the list of unprocessed markings 108, the system determines and sets a different focus setting for the device 116 for the next image acquisition. For example, the system may determine and set a different focus setting that provides for the acquisition of another image in which one or more markings 108 on the list of unprocessed markings 108 are in focus.

[0054] Specific embodiments have been described in the foregoing specification. However, those skilled in the art will recognize that various modifications and alterations can be made without departing from the scope of the invention, as set forth in the claims below. Accordingly, the specification and the figures are to be regarded in an illustrative rather than a limiting sense, and all such modifications are to be included within the scope of the present teachings.

[0055] The benefits, advantages, problem solutions, and any element(s) that may lead to or enhance a benefit, advantage, or solution shall not be construed as critical, necessary, or essential features or elements of any claim or all claims. The invention is defined exclusively by the attached claims, including all amendments made during the pendency of this application, and all equivalents of these claims as granted.

[0056] Furthermore, in this document, relational expressions such as first and second, top and bottom, and the like may be used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order of such entities or actions. The expressions "includes," "comprising," "has," "including," "containing," "incorporating," "including," or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, procedure, article, or device that includes, has, contains, or includes a list of elements may not only include those elements but may also include other elements not expressly listed or inherent in such process, procedure, article, or device. An element that "includes," "has," orThe phrase "…a", "includes…a", or "contains…a" preceding a statement does not, without further limitations, preclude the existence of additional identical elements in the process, method, article, or apparatus that includes, has, incorporates, or contains the element. The terms "a" and "a" are defined as one or more unless expressly stated otherwise herein. The terms "essentially", "generally", "approximately", "about", or any other version thereof are defined in a manner that would be closely understood by those skilled in the art, and in one non-restrictive embodiment, the term is defined as being within 10%, in another embodiment within 5%, in another embodiment within 1%, and in yet another embodiment within 0.5%.The term "coupled," as used herein, is defined as connected, although not necessarily directly and not necessarily mechanically. A device or structure that is "configured" in a particular way is configured at least in that way, but may also be configured in ways not listed.

[0057] Certain expressions may be used herein to list combinations of elements. Examples of such expressions include: "at least one of A, B, and C"; "one or more of A, B, and C"; "at least one of A, B, or C"; "one or more of A, B, or C". Unless expressly stated otherwise, the above expressions include any combination of A and / or B and / or C.

[0058] It is understood that some embodiments may include one or more specialized processors (or “processing devices”) such as microprocessors, digital signal processors, custom processors, and field-programmable gate arrays (FPGAs), and unique stored program instructions (including both software and firmware) that control the one or more processors to implement, in conjunction with certain non-processor circuitry, some, most, or all of the functions of the method and / or device described herein. Alternatively, some or all of the functions could be implemented by a state machine that does not have any stored program instructions, or in one or more application-specific integrated circuits (ASICs) in which each function, or some combinations of certain functions, are implemented as custom logic.Of course, a combination of the two approaches could be used.

[0059] Furthermore, an embodiment can be implemented as a computer-readable storage medium containing computer-readable code for programming a computer (e.g., comprising a processor) to perform a method described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, a ROM (read-only memory), a PROM (programmable read-only memory), an EPROM (erasable programmable read-only memory), an EEPROM (electrically erasable programmable read-only memory).Furthermore, it is expected that average professionals, regardless of possible considerable effort and many design decisions motivated, for example, by available time, current technology and economic considerations, guided by the concepts and principles disclosed herein, will be readily able to produce such software instructions and programs and ICs with minimal experimentation.

[0060] The summary of disclosure is provided to enable the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it is not intended to interpret or limit the scope or meaning of the claims. Furthermore, it is evident from the preceding detailed description that various features in different embodiments have been summarized for the purpose of simplifying the disclosure. This method of disclosure is not to be construed as reflecting an intention that the claimed embodiments require more features than are expressly stated in each claim. Rather, as the following claims reflect, the inventive step lies in fewer than all the features of any single disclosed embodiment.Therefore, the following claims are hereby included in the detailed description, each claim being a separate subject matter claimed on its own.

Claims

[1] Procedure, encompassing: Capturing, by a device using a current focus setting, a current image of an area, wherein the area has one or more markings present therein; Determining a first distance between the device and one or more markings present in the current image, based on the current focus setting of the device during the acquisition of the current image; Detect, using a trained machine learning model, one or more labels present in the current image, each label having one or more identifiers; Determine whether one or more labels are to be evaluated; In response to determining that one or more markings are not to be evaluated, determine, for a current marking under the one or more markings, a second distance between the device and the current marking based on a known size of the current marking and / or a feature of the current marking; Determine whether the current marking is in focus, based on the first distance, the second distance of the current marking and at least one attribute of the device; and In response to determining that the current labeling is the focus, processing the current labeling. [2] Method according to claim 1, wherein the current focus setting is one of a fixed focus setting of the device or an autofocus setting of the device. [3] Method according to claim 1, wherein the detection, using the trained machine learning model, comprises one or more markings present in the current image: Generating a boundary frame that corresponds to each label; and Determining a pixel size for each label based on a pixel length and / or a pixel height of the bounding box. [4] The method of claim 1, further comprising training a machine learning model to detect the one or more markings based on at least one of historical data containing one or more previously detected markings, or data sets containing images of one or more marking types, wherein each marking type has a known size. [5] Method according to claim 1, wherein the device of a mobile computer, head-mounted display, tablet, smartphone, camera or portable computing device; and that at least one attribute of the device is depth of field. [6] Method according to claim 1, wherein the processing of the current marking comprises: Decoding an identifier that corresponds to a predetermined symbology and / or barcode data structure; or Using character recognition to identify an identifier that matches a predetermined string structure. [7] Method according to claim 1, further comprising: in response to the determination that the current labeling is not the focus, Adding the current identifier and the second space of the current identifier to a list of unprocessed identifiers; Determine whether one or more labels are to be evaluated; In response to determining that one or more labels are being evaluated, determine and adjust, based on the list of unprocessed labels, a different focus setting of the device for the next image acquisition. [8] Device comprising: an imaging assembly; one or more processors; and a non-transitory, computer-readable memory coupled to the one or more processors, wherein the memory stores instructions which, when executed by the one or more processors, cause the one or more processors to: Receiving a current image of an area, captured by the imaging assembly using a current focus setting, wherein the area has one or more markings present therein; Determining a first distance between the device and one or more markings present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image; Detect, using a trained machine learning model, one or more labels present in the current image, each label having one or more identifiers; Determine whether one or more labels are to be evaluated; In response to determining that one or more markings are not to be evaluated, determine, for a current marking under the one or more markings, a second distance between the device and the current marking based on a known size of the current marking and / or a feature of the current marking; Determine whether the current label is in focus, based on the first distance, the second distance of the current label, and at least one attribute of the imaging assembly; and In response to determining that the current labeling is the focus, processing the current labeling. [9] Device according to claim 8, wherein the current focus setting is one of a fixed focus setting of the imaging assembly or an autofocus setting of the imaging assembly. [10] Device according to claim 8, wherein the instructions, when executed, further cause the one or more processors, using the trained machine learning model, to detect the one or more markings present in the current image by: Generating a boundary frame that corresponds to each label; and Determining a pixel size for each label based on a pixel length and / or a pixel height of the bounding box. [11] Device according to claim 8, wherein the instructions, when executed, further cause the one or more processors to train a machine learning model to detect the one or more markings based on at least one of historical data containing one or more previously detected markings, or data sets containing images of one or more marking types, each marking type having a known size. [12] Device according to claim 8, wherein the device of a mobile computer, head-mounted display, tablet, smartphone, camera or portable computing device; and that at least one attribute of the imaging assembly is depth of field. [13] Device according to claim 8, wherein the instructions, when executed, cause the one or more processors to process the current label by: Decoding an identifier that corresponds to a predetermined symbology and / or barcode data structure; or Using character recognition to identify an identifier that matches a predetermined string structure. [14] Device according to claim 8, wherein, in response to the determination that the current marking is not in focus, the instructions, when executed, further cause the one or more processors to: Add the current identifier and the second interval of the current identifier to a list of unprocessed identifiers; determine whether one or more labels are to be evaluated; In response to determining that one or more labels are being evaluated, determine and set a different focus setting of the device for the next image acquisition, based on the list of unprocessed labels. [15] Non-transitory computer-readable medium that stores instructions which, when executed by one or more processors, cause the one or more processors to: a current image of an area, acquired by an imaging assembly using a current focus setting, wherein the area has one or more markings present therein; determine an initial distance between the device and one or more markings present in the current image, based on the current focus setting of the imaging assembly during the acquisition of the current image; using a trained machine learning model that detects one or more labels present in the current image, each label having one or more identifiers; determine whether one or more labels are to be evaluated; In response to the determination that one or more markings are not to be evaluated, determine a second distance between the device and the current marking for a current marking under the one or more markings, based on a known size of the current marking and / or a feature of the current marking; determine whether the current label is in focus, based on the first distance, the second distance of the current label, and at least one attribute of the imaging assembly; and In response to the determination that the current label is the focus, process the current label. [16] Non-transitory computer-readable medium according to claim 15, wherein the instructions, when executed, further cause the one or more processors, using the trained machine learning model, to detect the one or more markings present in the current image by: Generating a boundary frame that corresponds to each label; and Determining a pixel size for each label based on a pixel length and / or a pixel height of the bounding box. [17] Non-transitory computer-readable medium according to claim 15, wherein the instructions, when executed, further cause the one or more processors to train a machine learning model to detect the one or more markings based on at least one of historical data containing one or more previously detected markings, or data sets containing images of one or more marking types, each marking type having a known size. [18] Non-transitory computer-readable medium according to claim 15, wherein the instructions, when executed, further cause the one or more processors to process the current label by: Decoding an identifier that corresponds to a predetermined symbology and / or barcode data structure; or Using character recognition to identify an identifier that matches a predetermined string structure. [19] Non-transitory computer-readable medium according to claim 15, wherein, in response to the determination that the current marking is not in focus, the instructions, when executed, further cause the one or more processors to: Add the current identifier and the second interval of the current identifier to a list of unprocessed identifiers; determine whether one or more labels are to be evaluated; In response to determining that one or more labels are being evaluated, determine and set a different focus setting of the device for the next image acquisition, based on the list of unprocessed labels.