Method and device for automatically generating a training data set for training a container detection model of machine learning of a real-time image analysis unit of a container handling device
The method automatically generates a training data set for container recognition models by using image data from a container handling device, enabling fully automatic labeling and achieving precise, fast, and robust container recognition in real-time.
Patent Information
- Application Number
- PCT/EP2024/081856
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-11
- Publication Date
- 2025-06-05
AI Technical Summary
Existing methods for training container recognition models in machine learning require significant manual intervention for labeling images, which is time-consuming and inefficient, especially for real-time container recognition in container handling devices.
A method and device for automatically generating a training data set by using image data from a container handling device, where edge data points are determined based on a detection area, and recognition results are derived by evaluating these edge data points against a boundary line of the recognition area, allowing for fully automatic labeling without manual intervention.
The method enables precise, fast, and robust container recognition in real-time, reducing the need for manual labeling and improving the efficiency of container handling processes.
Smart Images

Figure EP2024081856_05062025_PF_FP_ABST
Abstract
Description
[0001] Method and device for automatically generating a training data set for training a container recognition model of machine learning of a real-time image evaluation device of a container handling device
[0002] Description
[0003] The present invention relates to a method and a training data generation device for automatically generating a training data set for training, in particular for retraining, a container recognition model of machine learning of a container handling device in which containers can be guided or are transported along a predetermined transport path.
[0004] In this case, image data can be generated or is generated using an image capture device that is characteristic of a transport area along the transport path. The transport area is preferably an area within which the containers located therein are to be recognized. The image data can be and are, in particular, fed to the container recognition model as input variables for performing an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area.
[0005] The containers are preferably plastic containers (particularly PET containers), containers whose main component consists of pulp and / or glass containers and / or cans. The containers can be containers from the beverage and / or food and / or cosmetics industries. For example, they can be cans or bottles, such as glass bottles, pulp bottles and plastic bottles. The containers (to be recognized) can be fully formed containers and / or not yet fully formed containers such as preforms. They can also be empty containers or containers already filled with a product. Furthermore, the containers to be recognized can be provided with closures or not have closures.
[0006] With the help of neural networks, various objects can be recognized. The present invention particularly relates to the recognition of containers or parts of containers (lids, bottle mouths, etc.).
[0007] In order for the neural network to be able to recognize these containers, it must be trained. This requires training data. The training data consists of labeled images. The objects to be detected are marked with so-called "regions of interest" (ROIs). These can preferably be rectangles, circles, or even polygons.
[0008] If an image contains multiple relevant objects, multiple ROIs must be created. This process is called labeling. During subsequent training, the algorithm learns to recognize these ROIs in other situations. The labeling process can be done entirely manually or semi-automatically.
[0009] With manual labeling, the rectangles are defined and positioned by the user. There are various approaches to semi-automated labeling. Most approaches use a contiguous image sequence (video). The user labels a portion of the images manually, and with the help of an algorithm, these rectangles are extrapolated or interpolated to the subsequent images. It is also possible to use an existing neural network for labeling. With this variant, the user must decide whether the objects have been labeled correctly. Regardless of the variant, manual user intervention is always necessary.
[0010] Experience from prior art methods shows that a trained neural network cannot recognize every container. Retraining for each specific situation will always be necessary. Thus, labeled data is required for training. Manual labeling takes a very long time and must be performed by an operator. EP 350 302 4 A1 discloses a method for defect detection in packaging containers for liquid food products. Image data of the packaging containers are captured, and image features representing defects in the packaging containers are defined in the image data. Furthermore, time stamps are determined for the occurrence of defects during subsequent detection. Corresponding production parameters of the packaging containers in the machine are determined for the occurrence of defects based on the time stamps.
[0011] US 2019 355 128 A1 describes a method for segmenting generic foreground objects in images and videos. For this purpose, an appearance stream of an image in a video frame is processed using a first deep neural network, and a motion stream of an optical flow image in the video frame is processed using a second deep neural network. The appearance and motion streams are combined to combine complementary appearance and motion information to perform segmentation of generic objects in the video frames.
[0012] DE 10 2013 207 139 A1 discloses a method for controlling a filling system for liquid or solid products. This method comprises analyzing a dynamic state of the filling system, in which image sequences are recorded in at least one area of the filling system and evaluated by calculating an optical flow from an image sequence with a predetermined number of individual images.
[0013] US 2022 379 475 A1 discloses a system for identifying an object to be picked up by a robot from a container containing objects. The method involves obtaining a 2D color image and a 2D depth map image of the objects using a 3D camera, with pixels being assigned a value that identifies the distance from the camera to the objects.
[0014] US 2021 319 195 discloses a computer vision system for automatically identifying, tracking, and managing inventory and / or assets. The computer vision system is programmed with predetermined and configurable image data conditions that identify and capture specific types of movements.
[0015] US 2016 070950 A1 discloses a method for automatically assigning class labels to objects. This method uses object data that specify a plurality of parameters associated with an object. The method comprises identifying a plurality of cluster centers in a d-dimensional space from the object data, each cluster center corresponding to one of the class labels. For each cluster center, a surrounding region is determined based on a nearest neighboring cluster center, and the respective class label is assigned to the objects within the surrounding region.
[0016] From KR 101471519 B1 a method for detecting a moving object using cumulative difference image labeling is known.
[0017] From CN 106 331 620 A, a positioning analysis method for medicine bottles of a filling production line is known, in which the distance between a medicine bottle and the camera is calculated to ensure that the medicine bottle is further positioned.
[0018] EP 300 523 1 B1 discloses a method for counting objects on a conveyor belt. This method involves tracking the position of an extracted characteristic of each object in the image data in frames by extracting the characteristic in at least one additional subsequent frame recorded by the same camera.
[0019] US 2022 351 503 A1 discloses a method for labeling video images using an artificial neural network. A user provides a first input to label initial aspects of an object shown in a video frame.
[0020] The artificial neural network derives or predicts second aspects to be labeled for the object in a second video frame.
[0021] The present invention is based on the object of overcoming the disadvantages known from the prior art and of providing a method and a training data generation device for generating training data for training, in particular for retraining, a machine-learning container recognition model of a container handling device, which can mark or label the training images as far as possible without user intervention. A further object is to provide a precise, fast, and robust method for recognizing containers located in a transport area of a container handling device, as well as a corresponding container handling device. This object is achieved according to the invention by the subject matter of the independent claims. Advantageous embodiments and developments of the invention are the subject matter of the dependent claims.
[0022] A method according to the invention is a method, in particular a computer-implemented method, for automatically generating a training data set for training, in particular for retraining, a (trainable) container recognition model of machine learning (in particular a (processor-based) real-time image evaluation device) of a container treatment device in which (in particular during operation of the container treatment device) containers can be guided and / or transported (by means of a transport device) along a predetermined transport path, preferably in the form of a container stream.
[0023] It is also conceivable that the method according to the invention, instead of generating a training data set for training, in particular for retraining, the container recognition model described above, is (only) directed at automatic labeling of image data (wherein, on the basis of the labeled image data, a training data set is preferably generated in a subsequent step that is independent of or different from the method, for example by assigning generated annotation data to the respective image data).
[0024] In this case, image data can be generated (during the operating mode of the container treatment device) by means of (at least) one image capture device (and preferably with a plurality of image capture devices), which image data are characteristic of a (predetermined and / or predeterminable and / or in particular detectable by the at least one image capture device) transport area along the transport path (as well as containers located therein) and which can be fed to the container recognition model as input variables for carrying out an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area. The image capture device is preferably an image capture device of the container treatment device. The transport area is preferably an area within which the containers located therein are to be detected.In particular, the image data generated by means of (at least) one image capture device (and preferably with a plurality of image capture devices) depicts the containers located in the transport area. In other words, (during the working operation of the container treatment device) images are recorded with the at least one image capture device, in which the transport area (and containers located therein) are captured or depicted in the captured images. These are preferably evaluated in real time by the real-time image evaluation device using the container recognition model. For example, the real-time image evaluation device can, using the container recognition model, determine as an evaluation result or as a recognition result at least one for the number of containers detected in the transport area and / or position (or location) of the (detected) container (or containers).each container detected in the transport area) and / or orientation (or each container detected in the transport area) and / or type of container (or each container detected in the transport area) and / or condition of the container (or each container detected in the transport area) and / or the speed of the container (or each container detected in the transport area) and / or an occupancy level of the transport area or a section along the transport path with regard to containers located therein and / or distribution of the containers in a predetermined area within the transport area and / or make it available for output and / or transmission.
[0025] Preferably, the evaluation variable serves as a control variable for the container treatment device, for example, for the (automatic) execution of a treatment function on at least one and / or multiple containers. It is also conceivable for the at least one evaluation variable to be used as a control variable for controlling and / or regulating the container flow and / or a container throughput through the container treatment device and / or a transport speed and / or a container feed and / or container discharge.
[0026] The image capture device (or the plurality of image capture devices) may be an image recording device such as a camera, a CMOS sensor (CMOS abbreviation for complementary metal-oxide-semiconductor), a CCD sensor, a 3D sensor, an X-ray-based image recording device, an optical element, a thermal imaging camera, a LIDAR camera and the like, as well as combinations thereof.
[0027] In the case of the images (generated by the image capture device and / or the container recognition
[0028] The image data to be supplied to the model is preferably two-dimensional (spatially resolved) image data. The image data which can be or are supplied to the container recognition model for evaluation preferably do not include any depth information (measured immediately or directly by the image capture device), i.e. in particular no (depth and / or distance) measured values measured in the direction of the recording direction of the image capture device, which are characteristic of a distance of the image capture device to the objects imaged in the image data. In other words, the image capture device preferably does not generate a measured value which is solely characteristic of a distance and / or a (relative) position of the image capture device in relation to an imaged object.Such a distance and / or position of the image capture device relative to the objects imaged by it could be determined from image data generated by image capture device(s) at different positions from one another (e.g. via stereo recordings and / or LIDAR camera).
[0029] Preferably, the image capture device records a color image or a color video sequence or a color image sequence (for determining and / or generating the image data). However, it is also conceivable that the recorded and / or generated and / or determined image data is a grayscale image or grayscale images. In other words, it is conceivable that achromatic image data is fed to the container recognition model as input variables. This advantageously allows a larger transport area of the container handling device to be recorded or mapped and evaluated in the image data.
[0030] The image capture device (or the plurality of image capture devices) is preferably for recording (static) (individual) images and / or moving images or image sequences (or video sequences).
[0031] The image data (generated by the image capture device and / or) to be supplied to the container recognition model may be image sequences and / or individual images and / or image data recorded (essentially) at a single recording time or (essentially) simultaneously recorded or captured.
[0032] The image data (to be supplied to the container recognition model) preferably comprises image data obtained from more than one individual image. For example, two (different) image capture devices can each capture an image (preferably substantially simultaneously). Preferably, an image composed of multiple images and / or image data generated from multiple images is supplied to the container recognition model. Preferably, preprocessed image data is supplied to the container recognition model.
[0033] The machine learning container detection model is preferably based on an (artificial) neural network. The neural network is preferably designed as a deep neural network (DNN), in which the parameterizable processing chain has a plurality of processing layers, and / or a so-called convolutional neural network (CNN), and / or a recurrent neural network (RNN), and / or other DNN layer classes.
[0034] Preferably, the machine learning container recognition model is a previously trained container recognition model. In other words, the container recognition model to be trained, or more precisely, retrained, with the training data set to be generated is in a state after completion of a training process. Preferably, the training data set to be generated serves or is used for (further) fine-tuning of the container recognition model.
[0035] It is conceivable that the container recognition model has already been trained with a general training data set, preferably independent of a specific or concrete container handling device.
[0036] To generate the training data set, the method according to the invention comprises an evaluation of (specified and / or specifiable) image data as a function of a specified and / or specifiable detection area on the containers to be detected (for detecting a container to be detected). The term "detection area on the containers to be detected" is understood in particular to mean that the detection area can be a partial area of the container (e.g., the underside of the container) or an area of an element that is arranged on the container (preferably in a rotationally fixed and / or translationally fixed manner) (such as a closure). This offers the advantage that a particularly well-detectable detection area can be specified or selected for the container detection model as the ROI or as the object to be detected.
[0037] The image data specified and / or specifiable for generating the training data set preferably comprises image data generated or acquired by (at least) one image capture device of a container handling device. Preferably, the container handling device is the container handling device whose container recognition model is to be trained or retrained with the training data set to be generated. Preferably, the (at least) one image capture device is the image capture device of the container handling device whose container recognition model is to be trained or retrained with the training data set to be generated.This offers the advantage that the (post-)training process is specifically tailored to the (finely adjustable) container handling device and, for example, specific conditions of the specific container handling device, such as optical properties of the image capture device or specific lighting conditions or reflection conditions or geometric conditions (e.g. of the container guidance) in the container handling device can be directly taken into account.
[0038] However, it is also conceivable that the data (specified and / or specifiable for generating the training data set) were generated or transmitted by an identical (but different) container handling device.
[0039] Preferably, the image data provided (for generating the training data set) and / or to be evaluated (specified and / or specifiable) and / or acquired with the image acquisition device depict a transport area along a transport path of containers to be transported in a container handling device (by a transport device), wherein at least one container and preferably a plurality of containers are located in the transport area. The containers are preferably containers of the type with respect to which an evaluation is to be performed in the machine learning container recognition model to be trained and / or retrained.
[0040] Preferably, at least some of the image data (particularly preferably all of the image data) depict the same transport area (in particular, the same container handling device). Preferably, the different image data (different image recordings) have different numbers and / or distributions of the containers within the transport area.
[0041] Preferably, the image data (specified and / or specifiable for generating the training data set) are recorded and / or generated and / or determined at a different location (with respect to the location where the training data set is generated), preferably directly in the vicinity of or on the container handling device whose container recognition model is to be retrained and / or trained with the training data to be generated (preferably during the (ongoing) operation of the container handling device). The different location is therefore located in particular outside the building and / or company premises in which the container handling device is arranged and is preferably at a distance from it. Preferably, the image data generated orThe image data determined (specified and / or specifiable for generating the training data set) are transmitted to an external storage device (with respect to the container treatment device), which is accessed in particular via the Internet (and / or via a, in particular at least partially wired and / or wireless, public and / or private network).
[0042] Preferably, the method for generating a training data set comprises retrieving and / or receiving image data (which are specified and / or can be specified for generating a training data set) from the external storage device.
[0043] However, it is also conceivable that an image data set is received as image data using user data received as part of a user input.
[0044] The image data is preferably photorealistic. However, it is also conceivable that synthetic image data and / or augmented images (particularly based on photorealistic image data) are / are specified as image data. This offers the advantage that, for example, situations or conditions that do not or only rarely occur during normal operation of the container handling device (such as covered or overturned containers) can be recreated or simulated and used for training.
[0045] Preferably, the image data is processed, preferably depending on the detection area, to determine edge data points in the image data. Preferably, data points are determined or identified in the image data that are edge data points and / or that represent or depict an edge (or edge line) in the image data. The edge data points are determined or identified depending on the detection area.
[0046] In contrast to the methods known from the prior art, preferably not all of the image data are converted into an edge image independently of the detection area, but rather the detection area is taken into account when determining whether an image data point is a data point of an edge in the imaged image.
[0047] Preferably, at least one edge, and preferably several edges, are detected in the image data points using edge detection based on the detection area. This allows all edges or edge data points (of the detection area) belonging to containers depicted in the image data to be determined, so that preferably all containers depicted in the image data points can be detected.
[0048] The detection area is in particular selected such that it is suitable and determined for the (unambiguous) identification of a container based on a detection of the detection area (in particular from a viewing angle or detection direction of the at least one image capture device).
[0049] According to the invention, a recognition result of a container to be recognized is derived from the image data by evaluating determined edge data points as a function of an (outer) boundary line of the recognition area. The boundary line of the recognition area is preferably a circumferential line and / or a contour line. For example, the boundary line could be (at least in part) an (imaged) outer circumferential line of an area of an (imaged) container to be recognized, such as a circumferential line or a boundary line of a closure relative to another area of the container or an imaged background.
[0050] The boundary line can be a contour line and / or an outline line in which all (data) points or pixels have color values and / or brightness values from the same color value range and / or brightness range (e.g. all values of the same color values and / or brightness values).
[0051] In other words, based on the determined edge data points, an edge identified in the image data is compared with the boundary line of the detection area with regard to its (at least partially and / or completely and preferably fully comprehensive) course and / or size and / or scaling. In particular, the edge data points are examined or evaluated for detection of the boundary line. The proposed method offers the advantage that by evaluating the given image data with regard to a (given) detection area, a very reliable identification of a container to be detected based on the detection area is possible and, as a result, fully automatic labeling (i.e., without manual intervention by a user, for example by manually labeling or confirming a suggested label) of objects to be labeled is possible based on a short video recording or recorded images.
[0052] In a preferred method, the detection area of the container to be detected is selected from a group comprising lids, closures, can tops, can bottoms, container tops, container bottoms, container side wall areas, a mouth area of the container to be detected, a container feature (such as a label applied to the container, for example), a logo arranged on the container, a container set (entire sets), and the like, as well as combinations and sub-areas thereof. The detection area is preferably a partial area of a container depicted (in a two-dimensional image). The detection area can be characteristic of an element arranged on the container (such as a label) and / or of a (sub-)area of the container.
[0053] It is conceivable that the detection area is the closure applied to the container to be detected, in a top view or side view, or in a view taken from a detection direction obliquely from above (compared to a longitudinal direction starting from the container base along a longitudinal axis or central axis of the container toward the container mouth). For a closure with a (essentially) round cross-section, the boundary line would be a circle in this case.
[0054] It would also be conceivable to select the detection area of a container's base body (with or without a closure) in a top-down view. If the container's base body is essentially cylindrical, the boundary line would also be a circle.
[0055] In a further preferred method, a color filter is applied to the image data to determine the edge data points in the image data depending on a color range of the detection area. The color range of the detection area is understood to mean, in particular, the color spectrum represented by the detection area (preferably with a tolerance range). This offers the advantage that relevant image data points can be identified solely on the basis of the color values of the detection area. This provides a reliable and rapid pre-selection of relevant image data portions. The edge data points can be identified, for example, as data points in the image data that form an (outer) boundary of a (contiguous) area of the image data with image data points that have color values from the color range of the detection area.
[0056] By applying the color filter based on the color range (predefined and / or selected and / or predefined depending on the detection range), each area corresponding to the container detection range is advantageously identified in the image data.
[0057] In a preferred variant, the HSV color space is used for the color range. However, the use of the RGB, YCbCr, and / or L*a*b* color spaces is also conceivable.
[0058] Instead of using a color range of the detection area to determine the edge data points in the image data, it would also be conceivable to determine these as a function of a brightness profile and / or gray value range profile and / or imaged patterns and / or a contrast gradient and / or a texture feature of the detection area.
[0059] In a further preferred method, a color range of the detection area is determined and / or queried to determine the edge data points, preferably depending on user data transmitted as part of a user input. It is also conceivable that, depending on the specified detection area, the color spectrum or the color range of the color values represented in the detection area is automatically determined (and stored) once. It is also conceivable that the detection area is present and / or specified in the form of a photorealistic image of the corresponding area of or on the container to be detected.
[0060] In a further preferred method, the image data is processed into a binary image depending on a color range of the detection area by assigning exactly one of two predefined values to image data points depending on their color value depending on the color range of the detection area. The color values of the individual image data points are thus analyzed to determine whether the respective color value lies within the color range of the detection area. If this is the case for an image data point, a first predefined value (for example, a white color value) is assigned to the respective image data point. If this is not the case for an image data point, a second predefined value is assigned to the respective image data point.
[0061] The edge data points (between the areas to which a first predefined value is assigned and the areas to which the second predefined value is assigned) can then advantageously be quickly determined from the binary image.
[0062] In other words, the image data is converted into a binary image taking into account or depending on the detection area, preferably the color values of the (preferably complete) detection area. From the binary image, the edge data points are then preferably determined or identified as edges or border areas between the areas of different values.
[0063] In particular, the pixel values of the image data are converted into white or black color values depending on the specified detection range, creating a binary black-and-white image. This advantageously allows the image data to be divided or differentiated quickly and robustly into relevant image areas (which lie within the color range of the detection range, for example) and non-relevant image areas (whose image data points have color values outside the color range of the detection range).
[0064] By determining the edges between the white and black areas, the edge data points are then preferably determined and / or the image data is converted into an edge image.
[0065] In a further preferred method, the edge data points are determined from the binary image. This offers the advantage that edges are already determined in the vicinity of image data regions that have relevant color values (in particular, the color values of the detection region). These can advantageously be checked by contour comparison with the boundary line of the detection region to determine whether these edge data points form part or a section of an edge of a (imaged) detection region, or whether these edges form an object or element depicted in the image data points that is different from the detection region and (coincidentally) has identical or similar color values compared to the detection region.In a further preferred method, a Hough transformation, preferably characteristic of the boundary line of the detection area, is specified, by means of which edge data points determined in the image data are converted into a multidimensional parameter space, in particular a Hough space. The use of a Hough transformation offers the advantage of robust recognition of geometric figures or contours (here the boundary line) specified according to the Hough transformation rule in a binary gradient image (black and white image) after edge detection, or in an edge image. The use of a Hough transformation advantageously offers a robust method for recognizing contour curves even in highly noisy images.
[0066] Preferably, a standard or conventional Hough transformation is selected as the Hough transformation, in which the boundary line of the detection area can be parameterized.
[0067] Preferably, a plurality of edge data points, particularly preferably all of the determined edge data points, are transformed into the parameter space (Hough space) according to the Hough transformation. Preferably, the points transformed in the Hough space are examined for the presence of cluster points. It is conceivable that edge data points from the determined edge data points are selected—for example, randomly—and transformed into the Hough space.
[0068] Preferably, the recognition result is derived depending on (at least) one cluster point determined in the parameter space (Hough space). Preferably, a positive recognition result of a (imaged) recognition area (and thus also a positive recognition result of a (imaged) container to be recognized) is derived depending on one (preferably each) cluster point determined in the parameter space (Hough space).
[0069] If, for example, several (essentially equal or similarly pronounced) cluster points occur, it can be deduced that the boundary line of the detection area is depicted or occurs several times in the binary image or in the edge image.
[0070] If several (essentially equal or similarly pronounced) cluster points occur, it is then preferentially inferred that the detection area is depicted multiple times in the image data and / or that the container to be detected is depicted multiple times in the image data. The number of cluster points corresponds to the number of images of the container to be detected in the image data.
[0071] Additionally or alternatively, the position and / or orientation and / or scaling of the detection area depicted in the image data and / or the container to be detected is derived from the (preferably each) determined cluster point.
[0072] The Hough transform can be a generalized Hough transform. A (progressive) probabilistic Hough transform (PPHT) is also conceivable.
[0073] The generalized Hough transformation is preferably used for boundary lines that cannot be parameterized or can only be parameterized using a large number of parameters. For this purpose, a transformation rule is preferably provided depending on the boundary line, by means of which edge data points can be transformed into Hough space or the multidimensional parameter space. The transformation rule can, for example, be in the form of a look-up table. The parameters of the parameter space can characterize a position of a reference point of the boundary line (which preferably lies within the boundary line or within the detection area) and / or an orientation of the boundary line and / or (at least in sections and / or at least at certain points) the course of the boundary line.
[0074] The generalized Hough transform offers the advantage of allowing more complex perimeter lines to be selected as boundary lines. This offers the advantage, for example, that the highly flexible selection of the detection area, which results from allowing more complex perimeter lines, allows for a very narrow or (compared to the color values of, for example, the background depicted in the image data or the objects of no interest) specific color range of the detection area to be selected. This allows for even more precise detection of containers by detecting the detection areas.
[0075] It would also be conceivable to apply a generalized Hough transform depending on a texture feature of the detection area to derive a detection result based on the data points transformed into Hough space. In a further preferred method, the multidimensional parameter space is spanned by a plurality of parameters, by means of which the boundary line of the detection area and / or a position and / or orientation and / or scaling of the detection area can be parameterized. In particular, a standard or conventional Hough transform (SHT) is preferably used.
[0076] In a further preferred method, the boundary line of the detection area can be parameterized with fewer than 5, preferably fewer than 3, and preferably fewer than 2 parameters. This offers the advantage of efficient implementation of the algorithm.
[0077] Preferably, a Hough transformation is chosen such that the dimension of the Hough space is not higher than 5, preferably 3, preferably 2.
[0078] In a further preferred method, the boundary line of the detection area follows, at least in sections, a substantially circular and / or elliptical and / or rectangular line-shaped (and / or polygonal) course.
[0079] As mentioned above, for example in the case of containers with a circular cross-section and / or closures with a circular cross-section and / or with round tops and / or with mouth regions that are essentially circular in cross-section, a detection area can be selected such that a (essentially) circular and / or annular boundary line is present. In this case, the image data are preferably recorded and / or determined from a detection direction along the longitudinal axis of the containers (in particular determined from above). The detection direction of the image detection direction when recording the image data is preferably perpendicular to the transport plane (in the transport area), on which the containers depicted in the image data preferably stand or are arranged (straight) (upright).
[0080] Since the radius of the containers to be detected or the diameter of the closures of the containers to be detected is known, the radius is preferably not specified as a free parameter in the Hough transformation. In such a case, only the position of the center of the circle is preferably specified as a parameter of the Hough space.
[0081] The same applies to the use of elliptical or rectangular boundary lines.
[0082] Here, too, scaling and / or the ellipticity or ratios of the rectangle side lines can preferably be specified and / or preferably not taken into account as (free) parameters of the Hough space or the Hough transformation.
[0083] In the case of elliptical boundary lines (for example, when selecting closures or containers with a round cross-section), the detection direction of the image capture devices recording the image data can be selected obliquely to the longitudinal axis and / or central axis of the containers and / or closures.
[0084] In a further preferred method, a position and / or an orientation and / or a scaling of the boundary line and / or the detection area of the detection areas and / or containers detected in the image data is determined based on the determined edge data points, and the obtained detection areas are preferably checked for mutual intersection and / or overlap. It is also conceivable that only one position or only one position and orientation of the boundary line (and thus of the container) is determined.
[0085] If there are overlapping (detected) boundary lines, boundary lines are preferably removed, preferably repeatedly until no more boundary lines overlap. This advantageously allows containers that have been detected multiple times (possibly incorrectly) to be corrected, so that only one detected boundary line or one detected detection area remains for each container depicted in the image data.
[0086] The method preferably comprises a color filter (described above), a Hough transform (described above), and an elimination of overlapping lines. The method steps are preferably performed in the order mentioned. The method is preferably used for (fully automatic) labeling based on a (short) video recording or on recorded images of the objects to be labeled (containers, detection areas of the containers).
[0087] Preferably, the position (and / or orientation) of a (preferably each) recognized container is determined in the image data and / or the coordinates of a (preferably each) recognized container are determined in a (prescribed) coordinate system (world coordinate system) of the container handling device (and assigned to the respective image data). In a further preferred method, annotation data and preferably training data comprising an assignment of the generated annotation data to the image data are generated depending on the recognition result. The annotation data can include information on a (predefined) category, label, identification and / or marking of specific objects, localization and / or segmentation and / or video annotation (keypoints, polygons, bounding boxes for marking an object in the various frames). For example, it is conceivable that the annotation data contain coordinates orThe image data includes or contains position information of a detected container (or detection area). It is also conceivable that the annotation data includes or indicates a marking of the detected container (or detection area), for example, by means of a rectangle and / or a boundary line and / or a bounding box.
[0088] Preferably, the annotation data is characteristic of a (semantic) segmentation of the (respective) image data. In particular, (essentially) each image data point of the image data is assigned a class (for classifying an object), such as "container" or "non-container" (class annotation).
[0089] The classes can, for example, be classes for classifying a container type (e.g. can | bottle | glass bottle | PET bottle and the like) and / or for classifying a respective container size (e.g. filling volume, container height, diameter) and / or for classifying a container's features (e.g. body label, neck label, closure type, closure color, and the like).
[0090] It is conceivable that a (text) file with annotation data is generated for each image data to be labeled, which is then assigned to the image data. The assigned training dataset preferably includes the image data and the annotation data assigned to it.
[0091] Preferably, more than 100 different, preferably more than 1000 different images or image data are labeled in the manner described above, or annotation data is created for this purpose (and assigned to the respective image data). A training data set is preferably generated from this.
[0092] Preferably, the container treatment device is selected from a group comprising a transport device (such as a conveyor belt) for transporting the containers, a buffer device for temporarily buffering containers, a pasteurization device (such as a tunnel pasteurizer), a heating device for heating a preform, a forming device for forming a plastic preform into a plastic bottle, a sterilization device for sterilizing a container, in particular a plastic preform, a manufacturing device for producing a glass bottle, a filling device for filling a container with a product, an inspection device for inspecting a plastic preform or a bottle, a labeling device for labeling a container, a closing device for closing a container, in particular a filled container, a control device,a packaging device, a direct printing device for printing a container, a compilation device for assembling a plurality of containers into a compilation or a bundle, and the like.
[0093] Preferably, the container flow is a (particularly continuous) flow of successive or consecutive containers (on the transport path). The container flow can be guided or transported in a single lane or multiple lanes (by means of the transport device) in certain areas and preferably within the entire container treatment device (as a mass flow).
[0094] The transport device can also be a mass transporter for the transport of a large number of containers, preferably in multiple lanes and / or in an unordered manner.
[0095] The containers can be transported or guided standing or upright (by the transport device), preferably at least in sections and preferably along the entire transport area.
[0096] Preferably, the transport device is suitable and intended for at least partially guiding or transporting the plurality of containers, preferably along the entire transport area, of containers under dynamic pressure.
[0097] Preferably, the transport device is suitable and intended for transporting and / or guiding (at least within the transport area) at least 1 container per hour, preferably at least 5000 (recognizable) containers per hour and carries out this within the working operation of the container treatment device.
[0098] The present invention is further directed to a method for training, in particular for retraining, a container recognition model of machine learning, preferably a real-time image evaluation device, of a container treatment device in which (in a working operation of the container treatment device) containers can be guided along a predetermined transport path, in particular in the form of a container stream.
[0099] In this case, image data can be generated (during operation) using an image capture device that is characteristic of a (imaged) transport area (preferably of interest with regard to containers to be recognized) along the transport path. The image data can be fed to the container recognition model as input variables for performing an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area.
[0100] According to the invention, the container recognition model of machine learning is trained and / or retrained with training data generated according to one of the methods described above (in particular according to an embodiment described above).
[0101] It is therefore also proposed within the scope of the method according to the invention that labeled image data (as a training data set) is used for training, preferably for retraining, the container recognition model, the underlying image data of which are evaluated and labeled by determining edge data points as a function of a (predetermined and / or predeterminable) recognition area (which is on the container to be recognized) and evaluating the determined edge data points as a function of a boundary line of the recognition area.
[0102] The container recognition model and the container handling device can each be designed with all the features described above in connection with the method described above, individually or in combination with one another (and vice versa).
[0103] The container treatment device preferably records image data for generating a training data set using one or more image capture devices and transmits this data to an external storage device. The image data preferably represent (each) the transport area of the container treatment device. Image data is preferably generated (by the image capture device of the container treatment device) at different recording times and with different distributions of containers in the transport area and transmitted to the external storage device. The image data stored on the external storage device is preferably retrieved (in particular by or triggered by a training data generation device). A training data set is preferably generated (by the training data generation device) based on the retrieved image data (using the method described above according to a preferred embodiment).The training data set is preferably generated remotely with respect to the container handling device.
[0104] Preferably, the training data set generated (by the training data generation device) is retrieved (preferably by the container treatment device) from an external storage device.
[0105] Preferably, the data (to be processed), in particular the image data recorded by at least one image capture device, are fed to the container recognition model or the (artificial) neural network as input variables. Preferably, the container recognition model or the artificial neural network maps the input variables to output variables depending on a parameterizable processing chain, wherein the output variable is preferably a state of the respectively recognized container, a location of (each) recognized container, a speed of (each) recognized container, and / or a distribution and / or number of containers in the transport area and / or (at least) one variable characteristic of an occupancy level of the transport area or (at least) a section of the transport area.
[0106] It is also conceivable that the container recognition model determines as an output variable a position variable characteristic of a (current) location / position of (each) recognized container and / or a variable characteristic of a recognized container type and / or of an occupancy level of the transport area or (at least) a section of the transport area and makes it available (for transmission and / or output).
[0107] Preferably, a (semantic) segmentation of the image data (captured by the at least one image capture device) is performed using the container recognition model. Preferably, based on the obtained (semantic) segmentation of the image data, an occupancy level of a transport section of the transport area and / or the transport area or a characteristic value thereof is determined. The pixel-based or data point-based determination of the occupancy level offers the advantage of a very precise determination of the occupancy level.
[0108] Preferably, the container recognition model is suitable and determined for performing a (semantic) segmentation (of image data). Preferably, the respective occupancy size is determined using the data point-wise or pixel-wise class assignment (e.g., classes: "container" | "non-container") obtained by the (semantic) segmentation. For example, to determine the respective occupancy level of the transport area or a transport section, a ratio can be formed between the proportion of data points or pixels with the class assignment "container" and the total number of data points or pixels that represent the respective transport area or transport section.
[0109] The occupancy rate size can be determined on a container type-specific basis.
[0110] Preferably, the container handling device and / or the real-time evaluation device determines, depending on the output variables of the container recognition model, a variable characteristic of a speed of one and preferably each container recognized (in the transport area) and / or a state of one and preferably each container recognized (in the transport area) and / or a distribution and / or number of the recognized containers in the transport area and / or the occupancy level of the transport area or a section thereof.
[0111] Preferably, the container recognition model based on machine learning or the artificial neural network is (re-)trained using the generated training data, with the training parameterizing the configurable processing chain. An iterative training process is preferably selected, which is repeated until a specified recognition accuracy is achieved.
[0112] The external storage device is preferably a (non-volatile) storage device, in particular a cloud-based storage device and / or an external server (including storage device), wherein the storage device is accessed in particular via the Internet (and / or via a public and / or private network, in particular a wired and / or wireless network at least in sections). An external server is understood to mean, in particular, a server external to a container handling device and / or real-time evaluation device, in particular a backend server.
[0113] The external server is, for example, a backend, in particular of a container handling device manufacturer or a service provider, which is configured to manage image data (in particular from a plurality of image capture devices and / or a plurality of container handling devices) and / or to configure container handling devices. The functions of the backend or the external server can be performed on (external) server farms. The (external) server can be a distributed system.
[0114] The present invention is further directed to a method for detecting containers located in a transport area (a container handling device).
[0115] According to the invention, a container recognition model of machine learning trained according to the method described above for training, in particular for retraining, a container recognition model of machine learning is used for the recognition and / or tracking of containers located in the transport area.
[0116] The container treatment device preferably has a real-time image evaluation device, in particular a processor-based one, which is suitable and intended to carry out an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area by means of the container recognition model of machine learning, for which purpose the image data can be supplied to the container recognition model (and the real-time image evaluation device) as input variables.
[0117] It is therefore also proposed within the scope of the method according to the invention that a very precise container recognition model of machine learning, obtained by retraining as described above, is used to recognize the containers (and derive a recognition result) and / or determine the number and / or speed and / or distribution of the containers located in the transport area.
[0118] Preferably, the container recognition model and the container handling device can be equipped with all of the features described above, either individually or in combination with one another (and vice versa). It is also conceivable for the container recognition model to determine, as an output variable, a position variable characteristic of a (current) location of the (each) recognized container and / or a variable characteristic of a recognized container type and to provide it (for transmission and / or output).
[0119] Preferably, the container handling device and / or the real-time evaluation device determines, depending on the output variables of the container recognition model, a variable characteristic of a speed of one and preferably each container recognized (in the transport area) and / or a state of one and preferably each container recognized (in the transport area) and / or position and / or a distribution and / or number and / or container type of the recognized containers in the transport area and / or an occupancy level of the transport area or a section thereof.
[0120] Preferably, at least one detected container is tracked based on the detection result and / or the determined number and / or distribution and / or position and / or container type. For this purpose, image data recorded at different points in time are used and each evaluated (particularly in real time). Preferably, positions of the respectively detected container are determined based on image data recorded at different points in time. Preferably, one or more speeds of the container are further determined based on the determined positions and the recording times and / or on the basis of a transport speed caused by the transport device. Preferably, an expected location area of the container for at least one future point in time is determined based on the determined positions and / or speeds of the container, preferably by applying a Kalman filter.
[0121] Preferably, such tracking is applied to several and preferably all containers located in the transport area.
[0122] It is also conceivable that, based on the detected containers, their positions and / or speeds and / or their determined future expected locations and / or an occupancy level of a transport area or a section thereof, a virtual traffic jam counter is provided in which a container jam is determined or predicted with a predetermined probability (preferably occurring in the future). Preferably, in the event of a detected or predicted container jam, a warning message is provided for transmission and / or output to a user. It is also conceivable that, depending on an occupancy level or container distribution in the transport area, different warning levels are determined and corresponding messages are provided for transmission and / or output to a user.
[0123] The present invention is further directed to a, preferably processor-based, training data generation device for automatically generating a training data set for training, in particular for retraining, a container recognition model of machine learning of a container handling device in which containers can be guided along a predetermined transport path, wherein image data can be generated by means of an image capture device which are characteristic of a transport area along the transport path and which can be fed to the container recognition model as input variables for carrying out an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area.
[0124] The training data generation device is suitable and intended for generating the training data set for evaluating image data depending on a predetermined and / or predeterminable detection area on the containers to be detected.
[0125] According to the invention, the training data generation device is suitable and intended to process the image data in such a way that edge data points are determined in the image data, and to derive a recognition result of a container to be recognized in the image data by evaluating determined edge data points in dependence on a boundary line of the recognition area.
[0126] The training data generation device can be suitable, intended and / or configured to carry out one or more of the method steps described above of the method for training, in particular for retraining, a container recognition model of machine learning (of a container handling device).
[0127] The present invention is further directed to a container handling device for handling containers, comprising a transport device suitable and intended for transporting the containers along a predetermined transport path, and comprising an image capture device by means of which image data can be generated that is characteristic of a transport area along the transport path. The container handling device has a real-time image evaluation device, in particular a processor-based one, which, by means of a machine-learning container recognition model, is suitable and intended for performing an evaluation, in particular a real-time evaluation, with regard to the containers located in the transport area. For this purpose, the image data can be supplied to the container recognition model (and the real-time image evaluation device) as input variables.
[0128] According to the invention, the machine learning container recognition model is a container recognition model that has been trained and / or retrained with a training data set generated according to one of the methods described above.
[0129] Preferably, the containers are treated depending on the recognition result of the container recognition model or (at least) one output variable of the container recognition model or depending on the containers recognized by the container recognition model.
[0130] The container treatment device can have all the features described above in connection with a container treatment device alone or in combination (and vice versa) and can be suitable and intended for carrying out all the method steps described above in connection with the method for detecting containers located in a transport area of a container treatment device.
[0131] The container recognition model may comprise one or more of the features described above individually or in combination.
[0132] The present invention is further directed to a computer program or computer program product comprising program means, in particular a program code, which represents or encodes at least some of the and preferably all method steps of the respective (above-described) inventive method (for generating a training data set and / or a method for training a container recognition model of a container handling device) and preferably one of the (above-described) preferred embodiments and is designed to be executed by a processor device. The present invention is further directed to a data memory on which at least one embodiment of the inventive computer program or a preferred embodiment of the computer program is stored.
[0133] The present invention has been described with reference to containers to be recognized. The present invention is also applicable to objects to be recognized generally in machine image processing (based on machine learning image analysis models) for image recognition in the beverage and / or pharmaceutical sectors, particularly in container handling devices in the beverage and / or pharmaceutical sectors. The applicant reserves the right to also claim related subject matter, in particular the method for generating a training data set and the training data generation device.
[0134] Further advantages and embodiments can be seen from the attached drawings:
[0135] Showing:
[0136] Fig. 1 is a schematic representation of a sequence of a method according to the invention for generating a training data set in a preferred embodiment;
[0137] Fig. 2 is a schematic representation of the method steps of a training data generation device according to an embodiment;
[0138] Fig. 3 a,b an example of image data (Fig. 3a) and labeled image generated from it (Fig. 3b);
[0139] Fig. 4a, b, c an example of recorded image data of a container treatment device (Fig. 4a) with labeled detail representation (4b) and a position display resulting therefrom;
[0140] Fig. 5 ac shows an example of image data (Fig. 5a), edge image generated from it with marked detection areas (Fig. 5b) and inverted image (Fig. 5c);
[0141] Fig. 6 shows an example image of a container treatment device; Fig. 7 shows a simulation representation of the container treatment device shown in Fig. 6;
[0142] Fig. 8a-c representations for the prediction of moving containers;
[0143] Fig. 9 is a schematic representation of a container treatment device;
[0144] Fig. 10 a,b image data together with the respective annotation data;
[0145] Fig. 11 - 13 Representations for the coordinate transformation from a world coordinate system into image coordinates;
[0146] Fig. 14 shows an example representation of the composition of image data from various image capture devices;
[0147] Fig. 15 shows a further diagram illustrating the merging of image information from image data in the spatial coordinate system;
[0148] Fig. 16a, b a representation of a transport device with tracked containers.
[0149] Fig. 1 shows a schematic representation of a flow of a method according to the invention for generating a training data set in a preferred embodiment. The training data set is intended, in particular, to serve for retraining, for "refining," a pre-trained (artificial neural) network 5 (referred to above as a machine learning container recognition model). Retraining (for example, of a Yolov5s algorithm) is necessary because the algorithm has recognized some, or at least individual, containers, but not all.
[0150] For this purpose, an image capture device 2 records a video and uses the video file for automatic labeling, designated in Fig. 1 by reference numeral 14. The labeled files are retained for training to "sharpen" the network 4, thus obtaining a trained network 6. Reference numeral N designates the network via which data can be exchanged between the image capture device 2 and a device for training the network. Based on a recognition result output by the network 6, a signal can be output to indicate how many objects (in this case, containers) were detected. This can be compared with a manual count. An online connection to a sensor is also possible.
[0151] An object (here a container) recognized in image data by the container recognition model can preferably be analyzed for container properties such as geometric dimensions, such as diameter, and / or main colors.
[0152] Furthermore, a region of interest (ROI) can be determined or defined within which the container will move.
[0153] The image data (e.g. video) recorded by the image capture device 2 are preferably recorded from a vertical angle.
[0154] Fig. 2 shows a schematic representation of the system architecture and the method steps of a training data generation device 1 according to an embodiment.
[0155] With the help of several coordinated algorithms, labeling is carried out fully automatically based on a short video recording or recorded images of the objects to be labeled. The image data recorded by an image capture device 2, designated here by reference numeral 16, which can be, for example, video recordings or images, is provided to the auto-labeler (also referred to as training data generation device 1). Reference numeral 10 designates containers (here, cans) depicted in the image data 16, which are to be recognized by the container recognition model to be trained or retrained. Therefore, the training data generation device 1 or the auto-labeler is to label the input image data 16 with respect to each recognized container 10.
[0156] Fig. 2 shows the method steps preferably included in the method (in the specified order): applying a color filter F1 to the image data 16, performing a Hough transformation F2, checking a bounding box size F3, and checking an overlap F4. Preferred embodiments of these steps are explained in more detail below. Reference numeral 24 indicates the obtained result, which is label 20 (here circles) for marking the position of the detected containers 10.
[0157] A preferred embodiment of a method according to the invention for generating a training data set or a training data generation device 1 preferably comprises several algorithms that cooperate to find and identify a desired object (referred to above as a detection area). The desired objects can be partial areas of containers (lids, closures, can tops, can bottoms) or entire packages. The overall algorithm consists of a color filter, a Hough transform algorithm, and an algorithm that eliminates overlapping lines. These sub-algorithms are executed in the above-mentioned order.
[0158] In the example shown in Fig. 2, the detection area or the desired object can be the circular bottom of the can. The boundary line of this detection area, in this case, is the outer edge of the can bottom, which is a circular ring. This detection area is marked by circles 20 in the result obtained in the labeling process, i.e., the image data 24.
[0159] To apply the algorithm, individual images, image sequences, or a recorded video (as image data 16) are required. In a less preferred version, a live stream can also be used. The generated individual images, image sequences, or videos (image data 16) can, in a preferred variant, be recorded near the machine and uploaded to a cloud. The algorithm is then applied in this cloud. In a less preferred variant, the image data can also be generated near the machine or in an integrated computer system and further processed in this environment.
[0160] For automatic labeling, the individual images 16 of the information input are first processed by converting the pixel values (both color pixels and black-and-white pixels) of the image. A color filter F1 is used for this conversion. A color range is defined that preferentially contains the color (including the surrounding color spectrum) of the object to be labeled.
[0161] A preferred variant uses the HSV color space for this purpose. However, the algorithm also works in the RGB, YCbCr, and L*a*b* color spaces, among others. The result of this conversion is a binary image in which the relevant image areas are white (pixel value = 1) and the irrelevant image areas are black (pixel value = 0).
[0162] The relevant area is defined as the area containing the previously defined color range. However, if the image sequence or video contains objects of a similar color to the object to be labeled, these are also displayed as relevant areas (white areas).
[0163] To eliminate these irrelevant areas, the next algorithm follows. The Hough transform F2 is capable of detecting circles or rectangles contained in the binary image. The procedures for detecting circles and rectangles differ. In the case of circles, the edges of the white areas (border to the black pixels) are considered.
[0164] In the case of the cans 10 shown in Fig. 2, for example, if the can bottoms are selected as the detection area, the outer edge results as a circular boundary line. In the binary image above (black and white image), this appears as an edge between the white area (can bottom) and the surrounding black area. This circular edge can be detected using the Hough transform F2.
[0165] Circles with a calculated radius are drawn along these edges. This set of circles creates many intersection points. The area with the most intersection points is the center of a circle. A similar procedure is used to determine rectangles. However, here, instead of circles, straight lines are drawn with a calculated gradient. A total of four straight lines can be drawn through the most frequently occurring intersection points, which then define a rectangle. Since the container diameter is known, a circle can be drawn around the found center points. Rectangles can be drawn analogously around the found center points.
[0166] It is conceivable that in a preferred step F3, the size of the detected circle or the detected bounding box is checked. For example, the diameter of the containers (and also the bottoms of the cans) is known. Circles detected by the algorithm in the image data that lie outside the tolerance range of this circle diameter indicate the detection of another (but unrecognizable) object and can be preferentially removed. However, it is also conceivable that this step F3 is not performed. For example, a predefined radius / diameter or geometric dimensions can already be taken into account during the Hough transformation.
[0167] In the final step, the superfluous circles or rectangles must be removed. The third algorithm, F4, is preferably used for this. This checks whether the circles or the generated rectangles overlap. If this is the case, the radius or edge length is checked. If the radius or edge length is within a specified tolerance range, if it is the desired container, or has other properties such as color, then the superfluous labels are removed. The found labels (circles or rectangles) are then generated as a text file with the same name as the frame used. The file contains the coordinates of each label 20, along with the class or category of the label. This information is used to train a neural network to recognize these objects and their categories.
[0168] Advantage of the proposed embodiment:
[0169] Using the proposed embodiment, any type, size, or color of container can be labeled, and the necessary data for training the neural network can be generated. The trained network can then provide information about the condition, location, and speed of empty and / or filled containers located on a transport unit and / or on a container handling machine. The user only needs to specify the color range of the object to be labeled at the beginning of the overall algorithm.
[0170] Fig. 3a, b shows an example of a circle search in image data (Fig. 3a) captured by an image capture device 2 using edge detection. Fig. 3b shows the resulting labeled image, i.e., the image provided with labels 20 (Fig. 3b).
[0171] Containers 10 should also be recognized if there is a pixelated connection between two or more containers 10. A Hough algorithm is particularly preferred because it provides good recognition results even with noisy image data. Fig. 4a, b, c show an example of recorded image data of a container handling device 1 (Fig. 4a) with a labeled detailed representation (4b). In each case, a transport device 4 can be seen, in which the containers 10 are transported (here in a single lane). The containers are provided with closures 11, which here in a plan view (which corresponds to the perspective of the image capture device here) also have a circular cross-section. It can be seen that in both Figures 4a, b the respective closures are provided or marked with labels 20.
[0172] Fig. 4c shows a (true-scale) position display of the containers (or the closures 11 thereof) resulting from the container recognition, which can be output to the user or processed for further control of the container treatment device 1.
[0173] Fig. 5 a shows an example of image data with the labels 19 resulting from container recognition.
[0174] After applying a color filter (here, the color range of the can bottom), only image areas that contain color values from this color range are identified as relevant. In the present case, for example, the base of the containers 10 (approximately in the lower left image area) does not contain any color values from the color range of the can bottom. The guide rods shown in Fig. 5a, on the other hand, contain color values from the color range of the can bottom, and their image areas are therefore marked as relevant image areas in the binary image. An edge image is generated from this.
[0175] Fig. 5b shows the generated edge image, which represents the edges between the relevant image areas (areas with color values from the color range of the can bottom selected here as the detection area) and the non-relevant image areas (area with color values outside the color range of the can bottom).
[0176] Fig. 5c shows the inverted image of Fig. 5b.
[0177] Finally, by applying a Hough transformation to this edge image, the edges that are circular (and thus represent the circular boundary line of the can bottom) can be detected. These detected boundary lines of the detection areas are identified in Fig. 5c with reference numeral 20. Fig. 6 shows an example image of a container handling device 1 with a transport device 4, which is configured here as a mass conveyor for the random transport of containers 10 provided with closures 11.
[0178] Fig. 7 shows a simulation representation of the container treatment device 1 with transport device 4' shown in Fig. 6, in which the real-time position of the containers 10' is shown.
[0179] Fig. 8a-c show representations for predicting moving containers;
[0180] First, images / videos are captured using an image capture device 2 (e.g., a camera) (step S1). Containers are then recognized in the captured (individual) images using the (retrained) machine learning container recognition model (step S2). The coordinates of the recognized containers are converted into the spatial coordinate system (step S3).
[0181] Preferably, image data are recorded from multiple viewing angles or capture directions (multiple image capture devices 2). The image information from these multiple image data sets is combined (step S4). In step S5, it is shown that the containers are tracked (in particular by means of simulation and prediction of their future position) and preferably by means of image data recorded at different times and evaluated for container recognition. In step S6, the image data is preferably compared with a simulation or calculated prediction.
[0182] The simulation or determination of a prediction of the movement or the future position of a (each detected) container can be carried out using a Kalman filter (see, for example, Fig. 8b, 8c).
[0183] Fig. 9 shows a schematic representation of a container treatment system 3 with a transport device 4. Reference symbols E, E1, E2 indicate transport areas of different sizes arranged at different positions along the transport path of the containers, which can be captured by different image capture devices 2 (or different settings). The areas E1 and E2 overlap. In particular, when image data for the areas E1 and E2 are recorded simultaneously, these image data can be combined so that the image data depicts the entire area E.
[0184] Preferably, live information about the number and distribution of containers on the transport device, such as a main conveyor belt, is determined.
[0185] Fig. 10 a,b shows image data together with the respective annotation data. Different annotation types are selected. In Fig. 10a, the containers 10 detected on the transport device 4 are provided with a unique number so that individual containers can be distinguished. In Fig. 10b, each detected container is marked with a bounding box 34 around the container top (here, the can top). It can be seen that all annotation data include information about a (current) position of the container 10.
[0186] Figures 11 - 13 show diagrams illustrating coordinate transformation from a world coordinate system (or spatial coordinates) to image coordinates. Using mathematical transformations, image data points from the image data captured (by the image acquisition device) can be converted into spatial coordinates based on the camera's position in space.
[0187] Fig. 12 shows the coordinate transformation of spatial coordinates (here as “world coordinates” [XYZ]) into camera coordinates [Xc Yc Zc] and subsequent projection into (two-dimensional) pixel coordinates [xy].
[0188] The conversion rule given in the upper formula in Fig. 13 specifies the conversion of 3D (spatial) coordinates into 2D pixel coordinates.
[0189] The chessboard in Fig. 11 illustrates that, based on the known dimensions of a reference object, such as the chessboard here (reference numeral 50 denotes the detected points for illustration purposes, and reference numeral 52 denotes the re-projected points), the homography matrix from Fig. 13 can be determined for coordinate transformation into spatial coordinates. All containers located in the same plane as the chessboard are correctly converted using the conversion rule from Fig. 13. Objects above or below the plane are displayed distorted.
[0190] Fig. 14 shows an example representation of the composition of image data from various image acquisition devices. The image data is captured by image acquisition devices from different acquisition directions. The ROI, such as an entire transportation area, can be covered by two camera images and must be converted accordingly. Reference symbols B1 and B2 indicate the coverage of the ROI across two camera images in the spatial coordinate system and the world coordinate system, respectively.
[0191] Fig. 15 shows a further illustration of the merging of image information from image data in the spatial coordinate system (or world coordinate system) acquired by multiple image capture devices. Reference symbols MP1 indicate the positions of containers in the spatial coordinate system, which are determined by evaluating image data from an image capture device from a first perspective. MP2 indicate the positions of containers in the spatial coordinate system, which are determined by evaluating image data from an image capture device from a second perspective.
[0192] Overlapping circles represent the same can (container). An offset could be caused, for example, by a bounding box that's too large.
[0193] Fig. 16a, b shows a representation of a transport facility with tracked containers. Tracking is preferably performed using a Kalman filter. For this purpose, after initialization (determining speed vO and location CO), a theoretical location CT of the container is determined or calculated (particularly using the Kalman filter) to predict a future position of the container. The theoretical location CT is preferably compared with a measured location CM, and an adjustment is made.
[0194] The measured location is determined by recording image data with an image capture device 2 and by recognizing the containers using the (retrained) machine learning container recognition model and determining the positions of the recognized containers 10. The recognized containers can, for example, be marked with labels 62 in the corresponding image data.
[0195] The higher the frame rate, the more precise the tracking, identified in the exemplary figures by the reference symbol T. However, this also results in a longer computing time. Tracking preferably takes place after conversion to the spatial coordinate system. Calculating the container speed is also possible. The applicant reserves the right to claim all features disclosed in the application documents as essential to the invention, provided that they are new, individually or in combination, compared to the prior art. It is further pointed out that the individual figures also describe features which may be advantageous in themselves. The person skilled in the art will immediately recognize that a specific feature described in a figure may also be advantageous without the adoption of further features from this figure.Furthermore, the person skilled in the art will recognize that advantages can also arise from a combination of several features shown in individual or different figures.
Claims
Patent claims 1. A method for automatically generating a training data set for training, in particular for retraining, a machine learning container recognition model (5) of a container handling device (3) in which containers (10) can be guided along a predetermined transport path, wherein image data can be generated by means of an image capture device (2) which are characteristic of a transport area (E1, E2) along the transport path and which can be fed to the container recognition model (5) as input variables for carrying out an evaluation, in particular a real-time evaluation, with regard to the containers (10) located in the transport area (E1, E2), wherein, to generate the training data set, image data (16) are evaluated as a function of a predetermined and / or predeterminable detection area on the containers (10) to be recognized, characterized in thatthat the image data (16) are processed in dependence on the detection area in such a way that edge data points (F2) are determined in the image data (16), and that by evaluating determined edge data points (F2) in dependence on a boundary line of the detection area, a detection result of a container to be detected in the image data (16) is derived., 2. Method according to claim 1, characterized in that the recognition area of the container (10) to be recognized is selected from a group which includes lids, closures, can tops, can bottoms, container tops, container bottoms, container side wall areas, equipment of a container, logo arranged on the container, bundles of containers and the like as well as combinations and sub-areas thereof.
3. Method according to at least one of the preceding claims, characterized in that to determine the edge data points (F2) in the image data, a color filter is applied to the image data depending on a color range of the detection area.
4. Method according to at least one of the preceding claims, characterized in that, in order to determine the edge data points, a color range of the detection area is determined and / or queried, preferably as a function of user data transmitted as part of a user input.
5. Method according to at least one of the preceding claims, characterized in that the image data are processed into a binary image as a function of a color range of the recognition area by assigning image data points to exactly one of two predetermined values as a function of their color value as a function of the color range of the recognition area.
6. Method according to the preceding claim, characterized in that the edge data points are determined from the binary image.
7. Method according to at least one of the preceding claims, characterized in that a Hough transformation characteristic of the boundary line of the detection area is predetermined, by means of which edge data points determined in the image data are converted into a multidimensional parameter space, in particular a Hough space, the detection result being derived as a function of an accumulation point determined in the parameter space.
8. Method according to the preceding claim, characterized in that the multidimensional parameter space is spanned by a plurality of parameters by means of which the boundary line of the detection area and / or a position and / or orientation and / or scaling of the detection area can be parameterized.
9. Method according to the preceding claim, characterized in that the boundary line of the detection area can be parameterized with fewer than 5, preferably with fewer than 3 and preferably with fewer than 2 parameters.
10. Method according to at least one of the preceding claims, characterized in that the boundary line of the detection area follows, at least in sections, a substantially circular and / or elliptical and / or rectangular line-shaped course.
11. Method according to at least one of the preceding claims, characterized in that on the basis of the determined edge data points, a position and / or an orientation and / or a scaling of the boundary line and / or the detection area of the detection areas and / or containers detected in the image data is determined and the obtained detection areas are checked for mutual overlap.
12. Method according to at least one of the preceding claims, characterized in that annotation data (20) and preferably training data comprising an assignment of the generated annotation data (20) to the corresponding image data (16) are generated as a function of the recognition result.
13. A method for training, in particular for retraining, a machine learning container recognition model (5) of a container handling device (3) in which containers (10) can be guided along a predetermined transport path, wherein image data (16) can be generated by means of an image capture device (2), which are characteristic of a transport area (E1, E2) along the transport path and which can be fed to the container recognition model (5) as input variables for carrying out an evaluation, in particular a real-time evaluation, with regard to the containers (10) located in the transport area (E1, E2), characterized in that the machine learning container recognition model (5) is trained with training data generated according to one of the preceding claims.
14. A method for detecting containers (10) located in a transport area (E1, E2) of a container handling device, characterized in that a machine learning container detection model trained according to the preceding claim is used to detect and / or track containers located in the transport area (E1, E2).
15. Training data generation device (1) for automatically generating a training data set for training, in particular for retraining, a container recognition model of machine learning of a container handling device (3), in which containers can be guided along a predetermined transport path, wherein image data can be generated by means of an image capture device (2), which are characteristic of a transport area along the transport path and which can be fed to the container recognition model as input variables for carrying out an evaluation, in particular a real-time evaluation, with regard to the containers (10) located in the transport area, wherein the training data generation device (1) is suitable and intended for generating the training data set for evaluating image data (16) as a function of a predetermined and / or predeterminable detection area on the containers (10) to be recognized, characterized in thatthat the training data generation device (1) is suitable and intended to process the image data (16) in such a way that edge data points (F2) are determined in the image data, and to derive a recognition result of a container to be recognized in the image data by evaluating determined edge data points (F2) in dependence on a boundary line of the recognition area.
16. Container treatment device for treating containers (10) with a transport device (4) which is suitable and intended for transporting the containers (10) along a predetermined transport path, with an image capture device (2) by means of which image data (16) can be generated which are characteristic of a transport area (E1, E2) along the transport path, and with a real-time image evaluation device which is designed by means of a container recognition model of machine learning (6) to carry out an evaluation, in particular a real-time evaluation, with regard to the in the transport area located containers (10), for which the image data (16) can be supplied to the container recognition model (6) as input variables, characterized in that the container recognition model of machine learning (6) is a container recognition model which is provided with a container recognition model according to one of the claims 1-12 was trained and / or retrained.
Citation Information
Patent Citations
Medicine bottle positioning analysis method of filling production line
CN106331620A
Method for monitoring and controlling a filling plant and device for carrying out the method
DE102013207139A1
Method and device for counting objects in image data in frames
EP3005231A2
Method of defect detection in packaging containers
EP3503024A1
Apparatus and method for detecting moving objects using accumulative difference image labeling
KR101471519B1