Reading optical codes
The method combines neural networks for initial segmentation and classical image processing to dynamically adjust parameters, enhancing code reading accuracy and efficiency in dynamic scenarios by addressing challenges like varying object heights and contrasts.
Patent Information
- Application Number
- EP2024187000
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2026-01-14
AI Technical Summary
Existing code reading technologies struggle to adapt to dynamic and challenging scenarios, such as varying object heights and contrasts, leading to reduced read rates due to false positives and insufficient processing time, especially in logistics applications.
A method that employs a combination of machine learning, specifically a neural network, for initial segmentation to identify code regions, followed by classical image processing to refine these regions, dynamically adjusting parameters based on the initial candidates to enhance code detection and decoding efficiency.
Improves code reading accuracy and read rates by adapting to varying conditions without additional hardware, handling variable object heights and contrasts, and reducing manual parameter adjustments, thus optimizing processing times for more efficient code recognition.
Smart Images

Figure IMGAF001_ABST
Abstract
Description
[0001] The invention relates to a method for reading optical codes according to the preamble of claim 1 and to a corresponding optoelectronic code reader.
[0002] Code readers are commonly found at supermarket checkouts, for automatic package identification, mail sorting, baggage handling at airports, and in other logistics applications. In a code scanner, a reading beam is guided across the code using a rotating mirror or a polygonal mirror wheel. A camera-based code reader uses an image sensor to capture images of objects with their codes, and image analysis software extracts the code information from these images. Camera-based code readers can easily handle code types other than one-dimensional barcodes, such as matrix codes, which are also two-dimensional and provide more information.
[0003] In one important application group, the code-bearing objects are conveyed past the code reader. A code scanner captures the codes as they are successively guided into its reading area. Alternatively, in a camera-based code reader, a line scan camera reads the object images containing the code information successively and line by line, capturing the relative movement. A two-dimensional image sensor regularly records image data, which overlaps to a greater or lesser extent depending on the recording frequency and conveying speed. To allow the objects to be arranged in any orientation on the conveyor, several code readers are often installed on a single reading tunnel to capture objects from multiple or all sides.
[0004] As preparation for reading codes, a captured image of a code-bearing object is searched for code areas—that is, those areas in the image that could potentially contain a code. This step is called segmentation or pre-segmentation. In most current code reading applications, segmentation is performed using traditional image processing algorithms and manually created classifiers. This often allows even very small structures to be recognized effectively, thus also finding code areas with small codes, such as 2D codes with small module and symbol sizes or barcodes with a small code height or bar length. The conventional approach remains very localized in its evaluation.In challenging reading situations, such as those with many background structures, this approach can lead to the detection of numerous false-positive code areas. Under typical real-time application conditions, not all of these can be processed, thus reducing the read rate. Another segmentation approach is based on artificial neural networks, particularly deep convolutional neural networks (CNNs). Ultimately, such a neural network applies a multitude of filter kernels, trained with sample data, to the image.
[0005] Traditionally, segmentation and code reading for an application are fixed at the device level or manually adjusted to the application situation, for example, by setting processing steps or parameters in a graphical user interface. Optimization is thus only possible on average, and even in a dynamic application, the settings remain fixed from object to object or image to image. This reduces the read rate, either directly through overlooked codes or codes that are unreadable due to unsuitable decoder settings, or indirectly because the available decoder time is insufficient to evaluate all code areas and images.
[0006] As an example, consider the search contrast parameter, which determines the required contrast level in an image area for it to be recognized as a code area. A high search contrast can cause low-contrast code areas to be missed, while a low search contrast can detect a number of code areas that contain only background texture and no optical code at all (false positives). This is particularly evident in a code reader used in a stationary application on a conveyor belt with a large variation in object height. Here, increased contrast sensitivity is necessary, meaning a relatively low search contrast, to recognize codes applied to flat objects, which become low-contrast due to their considerable distance from the code reader.However, codes on larger objects have higher contrasts, so that, for example, even a cardboard texture or a pattern on it is mistakenly considered a code.
[0007] From EP 3 812 953 A1, it is known that a code reader uses a distance sensor to determine the distance to a code and, based on the measured distance, sets a parameter or incorporates an additional algorithm of the decoding process. However, this requires additional equipment for the distance sensor, and moreover, not all parameters important for segmentation and decoding can be derived from the distance.
[0008] It is also known to statistically analyze a history of decodings with regard to certain key metrics in order to adjust parameters for future code reading. However, this only works with a corresponding history and thus for gradual, continuous changes. Abrupt changes or dynamics on short timescales, such as a large variance in object height, cannot be successfully addressed in this way.
[0009] In the paper Zhao, Qijie, et al, "Deep Dual Pyramid Network for Barcode Segmentation using Barcode-30k Database", arXiv preprint arXiv:1807.11886 (2018) a large dataset is synthesized and code segmentation is performed using CNNs.
[0010] Xiao, Yunzhe, and Zhong Ming, "1D Barcode Detection via Integrated Deep-Learning and Geometric Approach", Applied Sciences 9.16 (2019): 3268 claim a performance improvement of at least 5% compared to previous approaches in locating barcodes, without the need to manually adjust parameters.
[0011] Hansen, Daniel Kold, et al, "Real-Time Barcode Detection and Classification using Deep Learning", IJCCI. 2017 detect code ranges including a rotation in real time with an Intel i5-6600 3.30 GHz and an Nvidia GeForce GTX 1080.
[0012] Zharkov, Andrey; Zagaynov, Ivan, Universal Barcode Detector via Semantic Segmentation, arXiv preprint arXiv: 1906.06281, 2019 detect barcodes and identify the code type in a CPU environment.
[0013] German patent DE 101 37 093 A1 discloses a method for recognizing a code and a code reader in which the step of locating the code within an image environment is performed using a neural network. German patent DE 10 2018 109 392 A1 proposes the use of convolutional networks for detecting optical codes. US patent DE 10 650 211 B2 discloses another code reader that uses convolutional networks to locate codes in a captured image.
[0014] EP 3 428 834 B1 uses a classical decoder that employs methods without machine learning to train a machine learning-trained classifier, or more specifically, a neural network. However, this document does not address preprocessing or the discovery of code regions in detail.
[0015] EP 3 916 633 A1 describes a camera and a method for processing image data in which segmentation is performed using a neural network in a streaming process, meaning that image data is already being processed while further image data is being read in. At least the first layers of the neural network can be implemented on an FPGA. This significantly reduces computation times and hardware requirements, but does not improve the segmentation or code reading itself.
[0016] The EP 4 231 195 A1 demonstrates a combined segmentation using classical image processing and machine learning.
[0017] Against this background, the purpose of the invention is to further improve code reading, especially in dynamic scenarios.
[0018] This problem is solved by a method for reading optical codes according to claim 1 and a corresponding optoelectronic code reader according to claim 12. The method is a computer-implemented method that runs, for example, in a processing unit of a code reader and / or an attached processing unit. An image is captured in which an object with at least one optical code affixed to it is located. The image capture can be carried out in any of the ways described above, for example, during a conveying movement or by presenting an object to a camera in its field of view. As a preparatory step for the actual code reading, code regions in the image are located, i.e., image sections (ROI, region of interest) with an optical code, a process also referred to as segmentation.The detected code areas are then fed to a decoder, which reads the code content within each area, or decodes the code, thus converting the encoded message into plaintext. It should be noted that it can only be determined afterward whether the image actually depicts an object with an optical code, or whether an optical code can be read within a given code area.
[0019] In the process of identifying code regions, an initial segmentation procedure is performed, employing machine learning. Depending on the specific implementation (to be described later), the initial candidates for code regions identified can be used as code regions themselves, prepare the way for the actual identification of code regions, represent a subset of the code regions processed further by the decoder, or serve as a criterion for what is recognized as a code region. Here, machine learning, in a broader sense, encompasses methods based on learning or training from data, and in a narrower sense, it refers to artificial intelligence methods, particularly those using a neural network. Classical image processing is used as the contrasting term.Key practical differences are that classical image processing is programmed for its task from the outset and does not require experience or training, and its performance remains more or less constant from the beginning throughout its entire runtime. In machine learning, however, all of this depends far more than just on a programmed structure; it only emerges in conjunction with the training process and the quality of the training data.
[0020] The invention is based on the fundamental idea of using the results of the first segmentation method, which is based on machine learning, to adapt the discovery of code regions and / or the decoding to the current situation. This adaptation is expressed here via parameters, which include settings, configurations, and / or the selection of the methods or process steps used, or additional algorithms. Thus, by evaluating the first candidates resulting from the first segmentation method, (control) parameters for the segmentation and / or the decoder are derived and redefined.
[0021] The invention has the advantage of enabling situation-adapted parameterization or adjustment of the segmentation or decoder. This directly increases the read rate through more codes read and indirectly through shorter processing times, thus providing more decoding time for more difficult-to-read codes under (quasi-)real-time requirements. No additional hardware, such as a distance sensor, is required. This improves the handling of variable object heights, changing contrasts, and the like. Especially in logistics, objects with specific properties, such as package patterns, are often repeated in batches, so a change in this approach promises significant improvements for an entire series of subsequent objects that would otherwise be handled conventionally like the preceding batch.The invention also simplifies handling, since at least some of the parameters no longer need to be set manually and therefore may not even be visible in a configuration tool, thus reducing its complexity.
[0022] The first segmentation method preferentially generates an initial result map, where a result map is an image with a lower resolution than the captured image. Each pixel contains information indicating whether a code area is detected at that pixel's location. This is a particularly easy-to-use representation of the initial candidates as their respective result maps (heatmaps). Due to the lower resolution compared to the original image, each pixel of the result map represents a specific region or tile of the image and provides information, either in binary form or with a scoring value, about whether this region is part of a code area or not, or how likely it is that a code (part) was detected in that tile. This information may also include classification information, such as the likely code type.
[0023] Following the identification of the code regions, a fine segmentation is preferably performed to define the code regions more precisely, particularly at the image resolution. Information about the position of a code region does not necessarily contain the exact boundaries of the code. For example, in the case of result maps, the code region is only located at their coarser resolution. Fine segmentation improves the boundaries of the code regions, preferably with pixel-level accuracy at the resolution of the captured image.
[0024] The first segmentation method preferably employs a neural network, in particular a deep neural network or a convolutional neural network (CNN). This is a particularly well-established machine learning method for image processing. The first segmentation method can therefore reliably identify initial candidates for code segments.
[0025] The neural network is preferably trained using supervised learning with example images, which are evaluated based on the results of a segmentation and / or decoding procedure, without machine learning methods. Supervised learning makes it possible to generalize from a training dataset containing examples of a predefined correct evaluation to images later presented in operation. Corresponding neural network architectures and algorithms for training and operation (inference) are well-established, so existing, well-functioning solutions can be used or built upon. Some relevant literature was mentioned in the introduction. Assigning the correct evaluation to an example image, i.e., annotating or labeling, can, in principle, be done manually, since the training takes place before runtime.Furthermore, the first segmentation method can be trained, at least partially, with example images that have been evaluated by a segmentation method using classical image processing. This does not simply reproduce the classical segmentation method using different methods, as the neural network develops its own evaluation and generalization through training. Additionally, at least one classical decoder can be used that evaluates example images and retrospectively annotates code areas as positive examples only if code could actually be read there.
[0026] The code region detection process preferably employs a second segmentation method based on classical image processing without machine learning, which identifies second candidates for code regions. In other words, an additional classical segmentation is performed. This preferably results in a second result map analogous to the first result map, i.e., an image with a lower resolution than the original captured image, whose pixels contain information indicating whether a code region is detected at the pixel's location. The second segmentation method preferably identifies second candidates within a tiled image. This allows only small image sections to be processed at a time, for example, to determine contrast or count brightness edges. The tiles can be processed iteratively sequentially or in parallel at any desired interval. The second result map preferably contains one pixel per tile.If the first and second result cards are processed together, they should preferably have the same resolution, or a resolution adjustment can be made.
[0027] From the evaluation of the first candidates, a contrast threshold for the second segmentation method is preferably determined. This second segmentation method, which employs classical image processing, uses a contrast criterion to identify further candidates. The underlying heuristic is that the light and dark areas of an optical code generate higher contrast than background structures alone. The choice of the contrast threshold thus has a significant impact on missed optical codes or, conversely, on background structures mistakenly identified as optical codes. The initial candidates, which from the perspective of the first segmentation method are precisely those image areas containing optical codes, allow for a particularly precise selection of the contrast threshold.
[0028] The contrast threshold is preferably determined locally, particularly for each area surrounding a first candidate. A local contrast threshold allows for even more precise matching. The first candidates provide clues not only for contrast values in the image as a whole, but more specifically for their immediate surroundings. In this way, it may be possible to detect low-contrast optical codes in one part of the image without simultaneously misidentifying background structures in another part of the image as optical codes. Alternatively, a global contrast threshold can be used, which is set the same for the entire image.
[0029] A segmentation mode is preferably determined from the evaluation of the first candidates. A segmentation mode defines the criteria by which image areas are considered code areas, in particular the influence of the first and second candidates on the definition of code areas. The following segmentation modes are particularly preferred: 1) Use the first candidates as code areas. In this case, the result of the first segmentation method is adopted as the overall result of the segmentation, and the second segmentation method is preferably not executed at all. 2) Use the second candidates as code areas. In this case, the first segmentation method serves only to enable the parameterization according to the invention; the first candidates themselves are not used as code areas. 3) Use only those code areas that are both first and second candidates. This is a particularly strict criterion, as both segmentation methods must agree in the sense of a logical AND or an intersection. This results in very few false-positive regions, at the cost of potentially overlooking optical codes.The ratio of these two errors can be adjusted particularly favorably by selecting this third or the following fourth segmentation mode depending on the situation. 4) Use code areas that are first candidates or second candidates. Here, all potentially detected code areas of both segmentation methods are considered in the sense of a logical OR or a union. This virtually eliminates the risk of overlooking optical codes, but it also requires the decoder to evaluate a particularly large number of code areas, so its available processing time will only be sufficient if, through a clever selection of the segmentation mode, only a few false-positive candidates remain despite the OR operation.
[0030] The ratio of first candidates to second candidates is preferably used as a criterion for determining the segmentation mode. This ratio ultimately indicates the degree of agreement between the two segmentation methods. For example, if there are significantly more second candidates than first candidates, this may suggest a pattern in the background. The second candidates are then likely to be largely false positives, having reacted solely to the contrast of the pattern and not to an optical code. Therefore, a segmentation mode can be chosen that relies solely or predominantly on the first candidates.
[0031] Based on the evaluation of the initial candidates, a filter is preferably determined that excludes initially detected code areas before decoding. This filter is therefore a false-positive filter, which subjects the initial candidates to a further test to prevent valuable decoder time from being spent on code areas that actually contain no code. Depending on the evaluation of the initial candidates, it may or may not be necessary to use at least one false-positive filter.
[0032] The filter primarily checks whether the code area has a light background and / or quiet zones of an optical code. These are two examples of a false-positive filter. Optical codes are typically found on a label or tag, and therefore in a homogeneous, light environment; this is checked in the first case. The second case involves searching for quiet zones, which an optical code should possess.
[0033] The preferred criterion for determining the filter is the ratio of the number of first candidates to the number of second candidates. This criterion is analogous to the criterion for determining the segmentation mode mentioned above as an example.
[0034] The optoelectronic code reader according to the invention comprises a light receiving element for generating image data from received light and thus for capturing an image. The light receiver can be that of a barcode scanner, for example a photodiode, and the intensity profiles of the scans are assembled line by line to form the image. Preferably, it is an image sensor of a camera-based code reader. The image sensor, in turn, can be a line sensor for capturing a code line or an area-based code image by assembling image lines, or a matrix sensor, whereby images from a matrix sensor can also be combined to form a larger image. A combination of several code readers or camera heads is also conceivable.In a control and evaluation unit, which itself can be part of a barcode scanner or a camera-based code reader or connected to it as a control device, a method according to the invention for reading optical codes according to one of the embodiments is implemented.
[0035] The invention is further explained below with regard to additional features and advantages by way of example embodiments and with reference to the accompanying drawing. The illustrations in the drawing show: Fig. 1: A schematic three-dimensional overview of an exemplary assembly of a code reader above a conveyor belt on which objects with codes to be read are conveyed; Fig. 2: An exemplary flowchart for a segmentation procedure using a neural network; Fig. 3: An example image with optical codes printed in decreasing quality; Fig. 4: The result of a process based on the image according to Figure 3applied segmentation method with neural network for 2D codes; Fig. 4b the result of a segmentation method applied to the image according to Figure 3 applied segmentation method with neural network for 1D codes; Fig. 5 an exemplary flowchart for parameterizing the segmentation and / or decoding based on an evaluation of the result of a segmentation method with neural network; Fig. 6 another representation of the example image according to Figure 3 with a close-up of a low-contrast code; Fig. 7 an exemplary flowchart for a classic segmentation method; Fig. 8 an example image of a code-bearing object with a patterned texture and a close-up of the texture; Fig. 9 the result of a process applied to the image according to Figure 8applied segmentation method with neural network; Fig. 10a another exemplary image with optical codes and noise structures, here primarily text; Fig. 10b the result of a segmentation method applied to the image according to Figure 10a applied segmentation method using a neural network; and Fig. 10c the result of a scan of the image according to Figure 10a applied classical segmentation method.
[0036] Figure 1 Figure 1 shows an optoelectronic code reader 10 in a preferred application situation mounted above a conveyor belt 12, which conveys objects 14, as indicated by arrow 16, through the detection area 18 of the code reader 10. The objects 14 bear codes 20 on their outer surfaces, which are detected and evaluated by the code reader 10. These codes 20 can only be recognized by the code reader 10 if they are located on the top surface or at least visible from above. Therefore, unlike the illustration in Figure 1, the following applies: Figure 1To read a code 22 located, for example, to the side or bottom, a plurality of code readers 10 are mounted from different directions to enable so-called omnidirectional reading from all directions. In practice, the arrangement of the multiple code readers 10 into a reading system is usually implemented as a reading tunnel. This stationary application of the code reader 10 on a conveyor belt is very common in practice. However, the invention initially relates to the code reader 10 itself or the method implemented therein for decoding codes, so this example should not be understood as limiting.
[0037] The code reader 10 uses an image sensor 24 to capture image data of the conveyed objects 14 and the codes 20, which are then further processed by a control and evaluation unit 26 using image evaluation and decoding methods. The specific imaging method is not essential for the invention, so the code reader 10 can be constructed according to any known principle. For example, only one line is captured at a time, either using a line-shaped image sensor or a scanning method, in the latter case where a simple light receiver such as a photodiode suffices as the image sensor 24. The control and evaluation unit 26 combines the lines captured during the conveying movement into the image data. A matrix-shaped image sensor allows a larger area to be captured in a single image, and here too, the merging of images is possible both in the conveying direction and perpendicular to it.The multiple images are captured sequentially and / or by multiple code readers 10, whose detection ranges 18, for example, only together cover the entire width of the conveyor belt 12, with each code reader 10 capturing only a section of the overall image and the sections being stitched together. Fragmentary decoding within individual sections followed by stitching of the code fragments is also conceivable.
[0038] The task of the code reader 10 is to recognize the codes 20 and read the codes attached to them. Recognizing the codes 20, or rather the corresponding code areas in a captured image, is also called segmentation or pre-segmentation. The code reader 10 outputs information, such as read codes or image data, via an interface 28. It is also conceivable that the control and evaluation unit 26 is not located in the actual code reader 10, i.e., the one in Figure 1The control and evaluation unit 26 is not arranged as shown in the camera, but is connected as a separate control device to one or more code readers 10. In this case, the interface 28 also serves as a connection between internal and external control and evaluation. The control and evaluation functionality can be distributed across virtually any internal and external components, with the external components also being able to be connected via a network or cloud. All of this is not further differentiated here, and the control and evaluation unit 26 is considered part of the code reader 10, regardless of the specific implementation. The control and evaluation unit 26 can comprise several components, such as an FPGA (Field Programmable Gadget Array), a microprocessor (CPU), and the like.Specialized hardware components, such as an AI processor, an NPU (Neural Processing Unit), a GPU (Graphics Processing Unit), or similar devices, can be used, particularly for the segmentation process with a neural network, which will be described later. The processing of the image data, especially the segmentation, can be performed on-the-fly during image data acquisition or streaming, particularly on an FPGA, and in the manner described for a neural network code reader in the aforementioned EP 3 916 633 A1.
[0039] Figure 2 This shows an example flowchart for a segmentation process using a neural network. A neural network, especially a deep neural network or convolutional neural network (CNN), is particularly suitable for this purpose. In step S1, the captured image is fed into the input layer of the neural network.
[0040] In step S2, the neural network generates initial candidates for code regions from the inputs in several layers S3 (inference). Three layers S3 are shown as examples; the architecture of the neural network is not restricted, and the usual tools such as feedforward, feedback, or recurrent operations, layer skipping (ResNets), and the like are possible. A characteristic feature of a convolutional network is the convolutional layers, which effectively convolve the image, or in deeper layers, the feature map, of the preceding layer with a local filter. Resolution loss (downsampling) is possible due to larger filter shift steps (strided convolution, pooling layer). Reducing the resolution, especially in early layers, is desirable to enable sufficiently fast inference with limited resources.Contrary to the representation, the neural network can also include layers without convolution or pooling.
[0041] The neural network is trained beforehand using example images with known code sections (supervised learning). These example images can be evaluated manually (labeling, annotating). Alternatively, it's possible to use a traditional decoder with classic segmentation to evaluate example images and retrospectively identify code sections based on actually readable code. Especially without time constraints in offline mode, such traditional methods are very powerful, allowing for the automatic generation of a large number of training examples.
[0042] In step S4, the neural network completes its inference, and the feature map at its output provides the first candidates for code regions. These first candidates are preferably determined in the form of a heatmap. This is an image at a resolution corresponding, for example, to the last feature map at the neural network's output, with each pixel representing a specific region of the higher-resolution image. In the case of a binary heatmap, a pixel indicates whether or not a code region is present in the represented image segment. Alternatively, numerical pixel values can indicate a probability score for a code and / or a specific code type.
[0043] Figure 3 shows a sample image with optical codes printed in decreasing quality. Figures 4a-b show the result of a calculation based on the image according to Figure 3applied segmentation method with neural network, namely a heatmap with first candidates for code areas 30 with 2D codes in Figure 4a or with 1D codes in Figure 4b .
[0044] Figure 5This diagram illustrates an example flowchart for parameterizing segmentation and / or decoding based on an evaluation of the initial candidates. The idea is to use information gained from these initial candidates to parameterize subsequent code reading steps in a context-specific manner. This could, for example, involve further segmentation using traditional methods, as described below. Traditionally, this presents a chicken-and-egg problem, as properties from the code regions are needed to locate those regions. This is resolved because the initial segmentation process, based on a neural network, is parameter-free and requires no prior information about the code regions of the current situation. Furthermore, this initial segmentation process can be implemented very early in the processing chain, for example, during the preprocessing of the newly captured image on an FPGA.Parameterization is possible as an alternative or additional step for the subsequent decoding.
[0045] As far as this still overlaps with the flowchart of the Figure 2 In step S10, the image is captured, and in step S11, the first segmentation procedure using a neural network is applied to the captured image to obtain the first candidates.
[0046] In step S12, the first candidates are evaluated to determine key parameters. These are used in step S13 to determine parameters for adapting the subsequent steps to the situation, namely finding code ranges in step S14 and / or decoding the code ranges in step S15.
[0047] There are numerous control parameters for finding code areas S14 and decoding S15 that could be set or adjusted. The following examples illustrate a search contrast for classic segmentation, the activation of false-positive filters, and the selection of a segmentation mode. The invention also relates to the adjustment of further control parameters, such as the detection of blur situations in which the captured image is entirely paused until processing time becomes available for the decoder, or special situations in which an image does not need to be processed at all or cannot reasonably be processed.
[0048] The first parameter examined in detail is the search contrast for classical segmentation. Following the first segmentation method using a neural network, this is a second segmentation method that yields second candidates for code regions. In connection with the third parameter considered, the following discussion will address how the first and second candidates can be used in different implementations to define the code regions.
[0049] The search contrast refers to a required minimum change in gray value across a certain number of pixels, i.e., a contrast threshold, which allows textureless or only weakly varying image areas to be distinguished from the light-dark structures of an optical code 20. Traditionally, a fixed search contrast would be predefined or manually parameterized once for a specific application. By evaluating the contrast of the initial candidates, a situation-specific search contrast can be determined. For example, the variance or standard deviation of the gray values in the image area of the initial candidates is determined, and a certain multiple of this is set as the search contrast. The adjusted search contrast can be set globally, i.e., across all candidates, or locally for a specific area surrounding an initial candidate.
[0050] Figure 6 shows another representation of the example image according to Figure 3Here, the contrast decreases continuously for illustrative purposes; a close-up reveals a particularly low-contrast code. Accordingly. Figure 4a The code areas 30 are known after the first segmentation procedure, so that the contrasts for each code area can be determined. Figure 6 However, it can also be understood as an illustration of a slower change, where initially high-contrast codes could be read, as in the upper part, and later low-contrast codes appear, as in the lower part.
[0051] To avoid a drop in readership with a conventionally fixed search contrast, the search contrast is adjusted based on the first candidates each time.
[0052] Figure 7The diagram also shows an example flowchart for a classical segmentation method. Classical segmentation is generally based on relatively simple computational rules that can also be performed by an FPGA. The classical segmentation shown here is purely exemplary, particularly regarding the advantageous but not mandatory use of tiles, and any image processing methods known for segmentation can be used. However, the classical segmentation method is limited by the fact that no machine learning methods, and therefore in particular no neural networks, are used.
[0053] In step S20, the image is divided into tiles, i.e., into image sections of, for example, 10x10 or another number of pixels, which can also differ in the X and Y directions. Further processing can then be carried out tile by tile, with parallel processing across multiple tiles being possible.
[0054] In step S21, the contrast is determined for each tile, since a homogeneous area with low contrast contains no code. To determine the contrast, the grayscale values of the read pixels and their squares can be summed on-the-fly, for example, on an FPGA. From these summed values, the mean and standard deviation can then be determined without re-accessing the pixels, with the latter being a measure of contrast. Incidentally, such summed values can also be used in the contrast evaluation of the first candidates described above to accelerate the calculations.
[0055] In step S22, transitions from light to dark or vice versa are counted along a line, preferably on a test cross of two perpendicular lines. These are potential edges between code elements, of which a minimum number is expected in a code area. Step S22 is an example of an optional further evaluation within the framework of classical segmentation, going beyond pure contrast evaluation.
[0056] In step S23, a contrast assessment is performed against the adjusted search contrast. Tiles with insufficient contrast are discarded; they are not suitable candidates for code ranges. For tiles with sufficient contrast, the number of brightness edges is optionally compared against an edge threshold, and tiles with an insufficient number are also discarded. In this way, the evaluation of the first candidates improves the second, classic segmentation method by adjusting the search contrast to the specific situation.
[0057] The second parameter examined in more detail concerns the situation-dependent activation of at least one false-positive filter. Figure 8The image shows an example of a code-bearing object with a patterned texture, as well as a close-up of the texture. Such textures or patterns can fulfill a contrast criterion, even with adjusted search contrast, and this can sometimes lead to the detection of numerous code areas that actually only contain the pattern and are therefore false positives. These can be at least partially eliminated by false-positive filters, such as checking for a light background corresponding to a code label or tag, or for the presence of quiet zones in an optical code. However, false-positive filters consume valuable evaluation time and are sometimes too lenient, rejecting optical codes instead of false positives, for example, if an optical code is printed directly on the edge of a label and therefore does not include a quiet zone.
[0058] In the case of a patterned texture such as in Figure 8The second, classical segmentation method will generate a large number of second candidates. In contrast, the first segmentation method using a neural network is relatively immune to such texture. Figure 9 shows a heatmap with initial candidates for the image of the Figure 8 The first segmentation method did not quite accurately identify the code areas with actual optical codes, but in any case, it certainly did not consider every area with a pattern structure as code. Therefore, there are far fewer first candidates than second candidates. A simple criterion for detecting an image with a disturbing pattern-like texture thus evaluates the ratio of the number of first candidates to second candidates, and if this ratio is, for example, below 1 / 2 or another predefined threshold, a situation with a pattern-like texture is assumed, and false-positive filters are activated.
[0059] The third parameter examined in more detail concerns the choice of a segmentation mode. The following segmentation modes, or a selection thereof, may be available: 1) The first segmentation method: the first candidates are directly considered code areas. 2) The second segmentation method: the second candidates are directly considered code areas; the first candidates contribute only indirectly by setting parameters, particularly for the second segmentation method, based on the first candidates. 3) First combination mode: code areas exist only where both first and second candidates are detected. By restricting the selection to the intersection, only the most promising code areas are fed to the decoder. The balance here lies in avoiding false positives at the cost of potential false negatives, i.e., missed optical codes. 4) Second combination mode: code areas exist where either first OR second candidates are detected. The union of these candidates feeds all candidates to the decoder.The balance here lies in avoiding false negatives at the cost of using computing time on code areas where there is no optical code at all.
[0060] The segmentation modes thus differ in the probability that a code area actually contains code, and in the computation time required by the decoder. As already discussed, false positives are frequently caused by background textures in practice. Therefore, a similar criterion can be used as in the case of false-positive filters, comparing the number of first candidates to the number of second candidates. The more second candidates there are compared to first candidates, the more false positives can be assumed. If there are a comparatively large number of second candidates, segmentation mode 1) is more likely to be used because classical segmentation is not reliable, or segmentation mode 3), which only considers particularly promising code areas.Conversely, if the ratio is close to one, all candidates in segmentation mode 4) are of interest, as this promises the highest read rate but also requires the most computation time and would be overloaded with many false positives, or one relies in this case on the long-used classic segmentation in segmentation mode 2), because it is ensured that no excessive number of false positives were found.
[0061] The Figures 10a-c This further illustrates the relationship between the first and second candidates. It shows Figure 10a Another example image with optical codes and interference patterns, primarily text. In the background, you can see a package with a white label in the foreground, which displays numerous inscriptions, including several optical codes. Figure 10b shows the result of the calculation on the image according to Figure 10aapplied first segmentation method with neural network, i.e. the first candidates, and Figure 10c The result of the second, classical segmentation method, i.e., the second set of candidates. The discrepancy is obvious, clearly indicating the presence of interference patterns. Here, for example, the first segmentation mode could be selected, especially since the first segmentation method even identified a code area that the second segmentation method overlooked due to the darker image areas and the resulting reduced contrast.
[0062] Should the first segmentation method exceptionally fail to yield any initial candidates, the second segmentation method with a fixed search contrast can be used by default. Alternatively, such an image can also be excluded from decoding, at least temporarily, because there is a high probability that no optical code was recorded at all, or that any optical code that does exist could not be read anyway due to small size, blurriness, or similar factors.
Claims
1. Method for reading optical codes (20), comprising the steps of capturing an image, finding code areas (30) in the image and decoding the optical codes (20) in the code areas (30), wherein the finding of code areas (30) includes a first segmentation procedure using machine learning to find first candidates for code areas (30). characterized by that the first candidates are evaluated to determine parameters for finding code ranges (30) and / or decoding the optical codes (20).
2. Method according to claim 1, wherein the first segmentation method generates a first result map, wherein a result map is an image of lower resolution than the recorded image, the pixels of which contain information as to whether a code area (30) is detected at the location of the pixel.
3. Method according to claim 1 or 2, wherein the first segmentation method comprises a neural network, in particular a deep convolutional network.
4. Method according to one of the preceding claims, wherein the finding of code areas (30) comprises a second segmentation method of classical image processing without machine learning, with which second candidates for code areas (30) are found.
5. Method according to claim 4, wherein a contrast threshold for the second segmentation method is determined from the evaluation of the first candidates.
6. Method according to claim 5, wherein the contrast threshold is determined locally, in particular for each environment of a first candidate.
7. Method according to one of claims 4 to 6, wherein a segmentation mode is determined from the evaluation of the first candidates, in particular one of the following segmentation modes: use the first candidates as code areas (30), use the second candidates as code areas (30), use only such code areas (30) that are both first candidates and second candidates, use code areas (30) that are either first candidates or second candidates.
8. The method according to claim 7, wherein the ratio of the number of first candidates to the number of second candidates is used as a criterion for determining the segmentation mode.
9. Method according to one of the preceding claims, wherein a filter is determined from the evaluation of the first candidates which excludes initially found code areas (30) before decoding.
10. Method according to claim 9, wherein the filter checks whether the code area (30) has a light background and / or quiet zones of an optical code (20).
11. Method according to one of claims 4 to 8 and claim 9 or 10, wherein the ratio of the number of first candidates to the number of second candidates is used as a criterion for determining the filter.
12. Optoelectronic code reader (10) with at least one light receiving element (24) for generating image data and with a control and evaluation unit (26) in which a method for reading optical codes (20) according to one of the preceding claims is implemented.
Citation Information
Patent Citations
Camera and method for processing image data
EP3916633A1
Artificial intelligence-based machine readable symbol reader
US10650211B2
Recognition of a code, particularly a two-dimensional matrix type code, within a graphical background or image whereby recognition is undertaken using a neuronal network
DE10137093A1
METHOD FOR CAPTURING OPTICAL CODES, AUTOMATION SYSTEM AND COMPUTER PROGRAM PRODUCT FOR EXECUTING THE METHOD
DE102018109392A1
Optoelectronic code reader and method for reading optical codes
EP3428834B1