Computer-implemented method for determining a state of a handling device, method for training or retraining a first learning algorithm and / or a second learning algorithm, computer program, and monitoring system for determining a state of a handling device
A computer-implemented method using uncalibrated cameras and dual learning algorithms automates the detection of handling device states, addressing inefficiencies in existing systems by providing accurate and timely identification of operational issues.
Patent Information
- Application Number
- DE102023134879
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-18
AI Technical Summary
Existing monitoring systems for handling devices, such as mail handling systems, face inefficiencies and inaccuracies due to the reliance on light barriers and continuous human monitoring of camera feeds, which can lead to misinterpretation and delayed detection of operational issues like jams.
A computer-implemented method using uncalibrated cameras to generate image packets, processed by a first learning algorithm to extract features and a second algorithm to determine the state of the handling device, enabling efficient and accurate detection of operational states without manual intervention.
The method provides reliable, low-error, and timely detection of handling device states, reducing human workload and improving operational efficiency by automating the detection of issues like jams with high-resolution cameras, even in uncalibrated and randomly placed setups.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
The present invention relates to a computer-implemented method for determining a state of a handling device for mailpieces, to a method for training or retraining a first learning algorithm and / or a second learning algorithm for determining a state of a handling device for mailpieces, to a computer program and to a monitoring system for determining a state of a handling device for handling mailpieces.For monitoring handling systems for handling objects, it is known to use light barriers to determine a position of the objects handled by the handling systems. For example, a missing signal at a light barrier can indicate that a problem is present upstream in a feed region.Light barriers can only provide a result at exactly one location of the belt. Dense packet streams can block a light barrier for a long time. This could be confused with a jam. The use of tracking light barriers works only at isolated currents and requires a dense network of light barriers, all of which report to the same PLC / computer for evaluation. Without an additional camera, there is no possibility of judging the situation from a control room. Light barriers are often contaminated and provide incorrect results.Cameras represent an alternative monitoring possibility. However, it is necessary here that users must continuously monitor the videos in order to determine a state of the handling device.Both the light barriers and the cameras constitute an unsatisfactory solution. More specifically, the light barriers can output a state only exactly at their position. Furthermore, the light barriers, once blocked, cannot generate any further signals. Continuous monitoring of camera images requires a high amount of manual work and is susceptible to human errors such as misinterpretation.It is therefore an object of the present invention to provide a method and an apparatus for determining a state of a handling device, in which a state of the handling device can be determined reliably on the one hand and efficiently on the other hand.The present invention solves this problem with a computer-implemented method for determining a state of a handling device having the features of claim 1, with a method for training or retraining a first learning algorithm and / or a second learning algorithm for determining a state of a handling device for mail pieces having the features of claim 10, with a computer program having the features of claim 11 and with a monitoring system having the features of claim 12.According to one aspect of the present invention, a computer-implemented method for determining a state of a handling device for mailpieces is provided. The method can comprise determining or obtaining an image data stream of the handling device when handling mailpieces. The method may include generating at least one image packet from the image data stream, wherein the image packet includes a plurality of individual images. The method may include determining features of each image of an image packet using a first learning algorithm. The method may comprise determining the state of the handling device based on the features of each image of an image packet by means of a second learning algorithm. The first learning algorithm may be different from the second learning algorithm.Compared to the known prior art, the present invention provides the advantage that a state of the handling system can be generated based on any image data stream. In other words, the image data stream need not originate from a calibrated source, such as a calibrated video camera, but may originate from arbitrarily and randomly arranged image data sources, for example. Furthermore, continuous monitoring of the handling device can be provided without misinterpreting occurring. Thus, for example, it can be automatically detected whether or not the handling device represented in the image data stream is operating without errors. Further, state detection with low error rates and a low time delay between the occurrence of an error and its detection may be provided. For example, the present invention may detect jams (as an example of a state of the handler) of mailpieces on conveyors (as an example of a handler), e.g., in parcel logistics centers. A congestion can be a blockage of a movement flow. For example, conveyor belts with high resolution cameras can be observed, with the cameras not calibrated and placed arbitrarily, but with an overview of the relevant conveyor path in 3D space. The camera can be a 1080p, RGB camera. An object of the invention is to automatically identify whether the conveying path visible in the video stream (i.e. in the image data stream) of the camera is currently blocked or a traffic jam is present. This traffic jam detection can be implemented reliably with low error rates as well as with a small time delay between the occurrence of the traffic jam and its detection.In the prior art, the problem is solved by displaying a live video stream of the camera on a computer monitor, or on multiple monitors if there are multiple cameras observed by human experts who manually recognize and report jams. The problem is difficult to solve automatically, since the video cameras used are not calibrated, i.e. their position in three-dimensional space and the focal length of the objective are not known. In addition, there may be many such video cameras within the same parcel node, each with a different viewing angle and a different conveyor belt being observed. This makes calibration and thus measurement in 3-D space more difficult. Another difficulty arises from the possibility of non-linear operation of the manipulator. For example, a chute may feed new mailpieces to a handler, and the handler may stop for a short time to allow the mailpieces from the chute to land on an empty space on the handler. Such targeted stops of the handling device and thus of all mail items located thereon should not be recognized as an erroneous state of the handling device (e.g. as a jam).The state of the handling device can be indicative, for example, of whether the handling device is operating without errors. In other words, the state may be indicative of whether or not the handler is performing its assigned task. In other words, the state may take two values, namely a fault-free function or a fault-free function. Furthermore, it is conceivable that the state can also be indicative of the probability with which state is present. It is thus possible to estimate more accurately how unambiguously the respective state is present. In addition, the state can also comprise a prediction, namely the probability of which state will occur in the future, so that it can be estimated, for example, whether or not an error has to be expected in the near future. This can provide a predictable nature and, if appropriate, a control of the handling device can be adjusted to this. The handling device can be, for example, a conveyor belt, a chute, a suction gripper, a singulator, a sorter and / or an end station. The handling device may be any device that delivers mailpieces, for example, at a parcel center. In other words, the handling device can be a device which is used in the dispatch of mail pieces. The dispatch or handling of mail items can thereby comprise a physical displacement of at least one mail item. For example, the mailpieces can be conveyed, gripped, transferred, deflected, and / or aspirated during handling. The mailpieces can be, for example, packages, pockets, shipping bags, poly bags, smalls, and the like. The image data stream may be a video. In other words, the image data stream can comprise a plurality of individual images which are output over a certain period of time. The image data stream can be provided at 24 images per second, for example. The image data stream can either be determined by itself (for example by a video camera) or can be obtained from a database or a memory. In other words, the image data stream can also be obtained after a preceding processing. The image data stream can in this case comprise a handling process of at least one mail piece by the handling device. In other words, the image data stream can represent the physical handling of one or more mailpieces by the handling device. It may be sufficient that the image data stream partially represents the handling device. An image packet can then be generated from the image data stream. An image packet may be a collection of multiple individual images from the image data stream. Thus, the image packet can represent a section of the image data stream. The image packet can directly comprise individual images of the image data stream. Preferably, the image packet comprises images following one another in time. In other words, a single image package may not include duplicate images (i.e., no images captured at the same time). Based on the image data stream, an image packet or a plurality of image packets may be generated. After this step of image acquisition, the actual processing can be carried out by means of a learning algorithm (e.g. by means of a neural network, in particular a deep neural network), which can preferably be carried out purely end-to-end without explicit intermediate steps. Thus, as an example of a learning algorithm, the deep neural network can be continuously optimized by applying backpropagation and gradient descent. The at least one image packet can then be fed to the first learning algorithm. The first learning algorithm may recognize features for each image in the image packet based on the image packet and provide them as output. The size of the image packets may be constant. The size of an image packet can be selected to be so large that one state (e.g. a traffic jam) can be distinguished from another state of the handling device (e.g. a normal operation of the handling device). Preferably, however, a size of the image packet is also small enough to not introduce an artificial delay between the occurrence of the physical scene and the processing of the associated image. For package applications, a four second image package size seems to be a good compromise. Features can be, for example, a position of a mail piece in the handling device. Furthermore, such a feature can be a relative position of two or more mail pieces with respect to one another and / or with respect to the handling device. In addition, such a feature can be a property of the handling device represented in an image. For example, a specific movement time, a movement and / or an actuation of the handling device can signal that the handling device is functioning in the specific manner. It is thus conceivable, for example, for an operating lamp of the handling device to be recognizable in the image, which indicates whether the handling direction is currently performing a handling. The first learning algorithm can determine one or more features based thereon in connection with, in particular, information about the mailpieces. Preferably, the first learning algorithm may generate a feature map for each image of an image packet. A feature map may include, for example, at least one feature or characteristic associated with one or more pixels of the image. A position of a feature in the respective image can thus be defined. The features can then be fed to a second learning algorithm, which can determine a state of the handling device based thereon. In this case, only the features can be fed to the second learning algorithm. The images, on the other hand, cannot be supplied to the second learning algorithm. In this case, the second learning algorithm can be configured to determine a state of the handling device on the basis of the features which have been determined by the first learning algorithm. In other words, the first learning algorithm and the second learning algorithm may be configured differently for their respective task, such that each algorithm may operate optimally and efficiently. Thus, the first learning algorithm and the second learning algorithm may be different from each other. The difference may be that different input data and different output data are provided. Nevertheless, the two algorithms can be connected in series. The first learning algorithm and the second learning algorithm can serve, for example, as modules of a superordinate algorithm. In other words, essential features can be extracted on an image basis by means of the first learning algorithm and then made available to the second learning algorithm. The first learning algorithm and the second learning algorithm can thus be combined in order to enable continuous determination (for example classification) during training, as a result of which the overall process can be provided more robust and cost-effective. Furthermore, the first learning algorithm and the second learning algorithm can be optimized end-to-end, whereby the overall system is very robust. The optimization may be provided by a supervised learning approach in which annotated training data may be provided cost-effectively, since this may be provided only from camera images in combination with time intervals (for example, from one time stamp to the other time stamp) that show a particular state of the handling device. The automated output of the state of the handling device can reduce a workload of a human operator and reduce the error rate of the entire detection system.Preferably, the state is determined for each image packet. In other words, the second learning algorithm can only be provided with the features of an image packet. Thus, the second learning algorithm itself does not need to determine properties or basic conditions from the images of an image packet. The output of the second learning algorithm may be a single state for the entire image packet (i.e., for all images included in the image packet). Thus, the method may have an increased efficiency.Preferably, the first learning algorithm determines a feature map for each image of an image packet. In other words, one or more pixels of an image may be assigned a property and / or features. This allows a spatial distribution of the different property and / or features to be defined in each image. Thus, the second learning algorithm can also take into account, for example, a relatively spatial arrangement of different properties and / or features with respect to one another.Preferably, the generation of the at least one image packet is parameter-free. In other words, no other parameters are used for creating the at least one image packet. As a result, no further information is necessary which would have to be supplied to the system. Thus, no information about the camera such as angle of view, focal length, aperture, shutter speed and the like need be provided. In other words, it is sufficient for the present method to use only the images from an image data stream. The method can thus be established and applied very easily. Preferably, parameters designate camera parameters exclusively (such as aperture, exposure, focus and the like). In other words, parameters cannot mean an image section, viewing angle, images per unit time and the like.Preferably, the size of several image packets is the same. In other words, if multiple image packets are provided, each image packet may comprise the same number of individual images. This ensures that the learning algorithms always operate reliably. With a varying number of images in the image packets, it may not be ensured that the learning algorithms output the desired result.Preferably, an image packet covers a period of four seconds. In other words, an image packet preferably comprises a first image and a last image, wherein the time interval between the first image and the last image is four seconds. Thus, an image packet can be indicative of an operating section of the handling device lasting four seconds. In this case, it has been found that when processing mail pieces of this time period, the best results are given in terms of accuracy and minimization of the amount of data.Preferably, the first learning algorithm generates the features of each image on a per image basis. In other words, the first algorithm can be configured to determine the features of the respective image image image by image. As a result, each image can be taken into account sufficiently and all features which are represented in the respective image can be determined. Preferably, the first learning algorithm and the second learning algorithm are different from each other by considering a spatial dimension. In other words, the second learning algorithm may omit the spatial dimension, whereas the first learning algorithm may take the spatial dimension into account. In other words, the first learning algorithm may associate the features determined in each individual image with a spatial position. In contrast, the second learning algorithm can only be configured to classify the features determined by the first learning algorithm. As a result, the second learning algorithm can be configured to be significantly more efficient compared to the case when it would likewise have to take into account a spatial dimension. Overall, the method can thereby be provided significantly more efficiently. It is also conceivable that in the second learning algorithm, the temporal dimension is also omitted in addition to the spatial dimension. This can bring about a further increase in efficiency.The first learning algorithm is preferably an artificial neural network, in particular a convolutional neural network. Artificial neural networks may serve as universal function adaptors. Data is propagated from the input layer to the output layer, an activation function providing nonlinearity. During training of the artificial neural network, an error can be determined and weights used within the artificial neural network can be adjusted layer by layer using error feedback and an optimization method. The artificial neural network is particularly useful when assigning a plurality of input data to a limited number (i.e., a significantly smaller number) of output data. Thus, for example, 100,000 to millions of pixels (for example pixels) can be converted into a small number of permitted results in comparison therewith. In the present embodiment, the input data of the first neural network may be the images of an image packet. The output data, on the other hand, may be the features of each image. The convolutional neural network (for instance convolutional neural network) can be a special case of an artificial neural network. The structure of an artificial convolutional neural network may consist of one or more convolutional layers followed by a pooling layer. For example, batch normalization is provided. This unit can in principle repeat as often as desired (for example depending on the image resolution of the individual images and / or pooling), wherein a deep convolutional neural network is obtained with sufficient repetitions. For example, as compared to a multi-layer perceptron (e.g., a multi-layer perceptron), three essential differences of the convolutional neural network may be realized, namely a 2D or 3D arrangement of the neurons, shared weights and local connectivity. As input value to the convolutional neural network, a two- or three-dimensional matrix (for example a pixel of a grayscale or color image) can be present. Accordingly, the neurons can be arranged in the convolutional layer.The activity of each neuron can be calculated via a discrete convolution. In this case, a comparatively small convolution matrix (filter kernel) can be moved stepwise over the input value. The input value of a neuron in the convolutional layer can then be calculated as the inner product of the filter kernel with the currently underlying image section. Accordingly, adjacent neurons in the convolutional layer may respond to overlapping regions. The input value of each neuron determined by means of discrete convolution can then be converted into an output value by an activation function (for example a rectified linear unit (RELU)). In a pooling layer, superfluous information can be discarded. For example, an exact position of an edge in the image is often of minor interest, since the approximate localization of a feature is sufficient. Data reduction can thereby be achieved. After some repeating convolutional layer and pooling layer units, the artificial neural network may include one or more fully connected layers. A number of neurons in the fully connected layer (which may be the last layer in the neural network, for example) may then correspond to the number of different features that the neural network is intended to determine. The output of the last layer of the convolutional neural network can be converted into a probability distribution, for example, by a softmax function, a translation-but not scale-invariant normalization over all neurons in the last layer. The softmax function is particularly advantageous at the last layer of the convolutional neural network, which outputs information (e.g. results) to a user. As a result, an output of the artificial neural network can indicate a probability with which a specific feature is present.The second learning algorithm is preferably an artificial neural network, in particular a multilayer perceptron or long short-term memory. As for the artificial neural network, the above applies analogously to the present embodiment. The multi-layer perceptron (German multilayer perceptron, hereinafter referred to as MLP) may consist of at least three layers of neurons, namely an input layer, one or more hidden layers and an output layer. The multi-layer perceptron can be considered a directed acyclic graph in which the neurons in one layer can be completely connected to the neurons of the subsequent layer. The connection between the individual neurons of two layers adjoining one another may have a weight or a variable which can influence a forwarding value from the one neuron to the other neuron. The weights may be adjusted during training. In order to calculate a weighted sum of a neuron, the input values can be multiplied by corresponding weights and subsequently summed. The weighted sums may be calculated for each neuron in the hidden layers and the output layer. After the weighted sum is calculated, it may be processed by an activation function to determine the neuron's output value. The activation function may introduce non-linear characteristics into the neuron's output and may allow the multi-layer perceptron to learn complex relationships. The sigmoid function can be used as an activation function in the case of the multi-layer perceptron. Furthermore, the rectified linear unit function can also be used in the multi-layer perceptron. The output layer of the multi-layer perceptron can then generate a prediction of the network. Each neuron in the output layer may be connected to the neurons in the last hidden layer. The output values of the neurons in the output layer may represent the prediction of the multi-layer perceptron for the given input. When a multi-layer perceptron is used, features generated by applying the first learning algorithm to the individual images can be stacked for all the images within the image packet. Thereafter, the temporal and spatial dimensions may be omitted to treat the collection of features from the image package as a "pocket of features.". This removes the explicit term of time and space from the neural network, but otherwise the neural network may learn complex dependencies between different times. The multi-layer perceptron may include a plurality of feedforward layers each having a stack normalization and a non-linear activation function. The last neural layer may not have batch normalization and have a softmax function as a non-linear activation function.Additionally or alternatively, long short-term memory may also be used as the second learning algorithm. The long short term memory (hereinafter referred to as LSTM) may also map input data to output data. In this case, a comparatively small convolution matrix (filter kernel) can be moved stepwise over the input value. An advantage of an LSTM module is that it enables a relatively constant and applicable error flow. Generally, the LSTM may have a similar input layer, hidden middle layers, and output layer construction as an MLP. In an LSTM module, it is possible to check which information item should run into and out of an inner cell. The LSTM may then be configured to design or add information on the state by providing regulated structures (e.g., gates or gates). In LSTM modules, as in the network described above, modules are connected in series in a chain-like manner, but they have a different internal structure. In particular, the provided gates offer the possibility of optionally passing information through. Instead of a neural function, there may be gates in the LSTM module, which may be configured as follows. An input gate may control an amount in which a new value flows in the cell. A fogger gate may control the extent by leaving a value in the cell and an output gate may control an extent by using the value in the cell for calculation for the next module of the string. As a result, the errors can be prevented from disappearing or becoming excessively widened overall. In the selection of an LSTM network, the features extracted from the individual images by applying the first learning algorithm may be stacked. Unlike the MLP approach, the LSTM may omit only the spatial dimensions, but the temporal dimension may be maintained. This can preserve the explicit term of time required for LSTM networks. The LSTM network may consist of multiple bidirectional LSTM layers. The last layer of the neural network may be a fully connected neural layer without stack normalization and a softmax function to facilitate the two class problem. Both neural network variants (MLP and LSTM) may process the images of an image packet in a two class problem by feature extraction by the first learning algorithm and classifying the entire image packet. The difference between MLP and LSTM networks may be the explicit term of time omitted (MLP) or maintained (LSTM).Preferably, the spatial dimension is omitted when determining the state of the handling device by means of the multilayer perceptron. In other words, after the step of determining features of each image of an image packet, a removal of the spatial and temporal dimension can be provided. As a result, only a collection of features (for example a pocket of features) can be supplied to the multi-layer perceptron. As a result, the determination of the state of the handling device can be simplified and accelerated. In other words, no sub-selection method can be provided in the temporal dimension.Preferably, in determining the state of the handling device by means of the long-short-term memory, the spatial dimension is omitted. By omitting the spatial dimension in the LSTM as the second learning algorithm, it can be increased in its efficiency. Nevertheless, a spatial dimension can nevertheless be taken into account.Preferably, the image data stream is generated or acquired by a video camera. In other words, the image data stream can be recorded by a conventional video camera, which is provided for monitoring purposes, for example. The video camera does not have to be explicitly aligned with the handling device, but it is sufficient if the handling device is at least partially recognizable in the image of the video camera. This offers the advantage that common monitoring cameras can also be used to generate the image data stream. Furthermore, the video camera can be a movable camera which images a specific region in which the handling device is provided at specified intervals.Preferably, the video camera may be an uncalibrated video camera. A calibration can be realized, for example, by recording the specific position and the angle of view of the camera. Furthermore, during a calibration, reference marks can be provided in space in order to be able to assign the image data to spatial coordinates. This is not necessary in the present invention because the first learning algorithm and the second learning algorithm are provided as the end-to-end learning algorithm. More specifically, features are also detected without calibration of the image data stream (e.g., by the first learning algorithm). This makes it possible to use the present invention without problems and in a cost-effective manner. Further, an existing video camera system may be used.Preferably, the image data stream is obtained from a plurality of video cameras. In other words, the viewing angles of each camera may be different. Thus, the image data stream can be composed of different images with different viewing angles. In such a case, it is to be provided that an image packet comprises only images that have been recorded from the same viewing angle. In such a case, it is conceivable that the image data stream from a plurality of video cameras per camera is processed in parallel to one another.The method preferably further comprises determining an area of interest in the images, wherein in particular only the features in the area of interest are used as the basis for determining the state. In other words, only a region of the images that is actually of interest for determining the state of the handling device can be taken into account. This is particularly advantageous when arbitrary image stream data is used for such determination. Thus, robust features may be learned using the first learning algorithm, as well as not viewing certain areas in the images that are not of interest. Such areas may be, for example, spatial areas in which mailpieces are temporarily stored or in which human workers walk along. Moreover, the prediction accuracy of the learning algorithm can be further improved by providing the learning algorithm with the region of interest in which the manipulator is visible at the corresponding viewpoint. This region of interest may be encoded as a black (0=n't the region of interest) or style (1=n't the region of interest) pixel mask. Such a pixel mask is cost-effective to produce, since it only has to be carried out once per camera viewpoint (i.e. once for each viewing angle). The pixel mask for the region of interest corresponding to the current viewing angle may then be used to mask the pixels that are not part of the region of interest. Preferably, the pixel mask is applied after determining the features of each image. In other words, the pixel mask may be provided between the application of the first learning algorithm and the application of the second learning algorithm. Thus, the features output by the first learning algorithm may be masked. Thus, the features that do not correspond to a region of interest are not provided to the second learning algorithm, thereby allowing it to operate more efficiently. By doing so, both robust features within the first learning algorithm can be learned and regions of no interest cannot be considered.Preferably, the region of interest may be manually specified by a user. In other words, it is possible to define once for each camera viewpoint (for each viewing angle of the camera), which region is of interest (for example, where the handling device is provided) and / or which region is less or not of interest. It is also conceivable that the region of interest is automatically defined. This can be carried out, for example, additionally by the first learning algorithm. In other words, the first learning algorithm may determine which regions in the images are of interest and which are not. For example, this can be determined on the basis of differences between images recorded one after the other in the meantime. It is thus possible to identify where nothing is moving in the images, in particular permanently, and it is defined that this region is of reduced interest.Preferably, the state indicates how likely a first state and a second state are present. In other words, the output of the second learning algorithm may be a probabilistic assignment as to whether the image packets that are supplied as input to the second learning algorithm (for example as a stack of properties) have a first state or second state. The first state and / or the second state can thus be evaluated. The first state can mean, for example, that a traffic jam is present. The second state can mean, for example, that there is no traffic jam. In other words, a clear determination of a state may often not be possible, which is why a probabilistic assignment is helpful. It is also conceivable that, in the case of further processing, the probability of a first or second state being assumed can be used as a basis for implementing further controls.The state is preferably indicative of a probability that a jam of mailpieces, vertically standing mailpieces, damaged mailpieces, double / multiple packages of mailpieces, mailpieces fallen by the handling device, stuck mailpieces, dumped liquids and / or loose mailpiece content is present in the handling device. The aforementioned states can be considered incorrect states in the handling of mail items. In other words, if such a state is present, an adaptation of the control of the handling device or a manual intervention of a user is necessary. A jam of mail items can occur when mail items are displaced in the handling device (for example a conveyor belt) such that they block one another. In this case, further transport of the mail pieces is prevented. Mail items standing vertically, on the other hand, can be mail items, for example, which rest with their smallest surface area on the handling device. Such mail items are disadvantageous since subsequent problems can occur when the mail item falls over. Damaged mail items may be mail items in which the outer packaging is applied, for example. Double / multiple packages of mailpieces may comprise, for example, a mailpiece that is contained in a bag. Such a multiple package can lead to errors or even damage to the handling device when handling the mail pieces. Mailpieces fallen by the handling device can comprise, for example, mailpieces that have fallen unintentionally by a handling device and thus are not subject to handling. For example, these may be mail pieces that have fallen from a conveyor belt. Plug-inable mailpieces can comprise mailpieces that cannot be further handled by the handling device. In this case, it is not necessarily necessary for a jam to occur; this can be, for example, a mail piece which has remained stuck to a switch of a conveying device and does not move any further. Discharged liquids can be, for example, liquids that have leaked from mail items, which can have a negative influence on operation of the handling device. Furthermore, the liquids can also be lubricating or operating substances of the handling device itself, so that a negative influence on the handling device can occur on account thereof.Preferably, the method may further comprise evaluating the state, wherein the evaluation is indicative of a quality of the state. The result of the application of the second learning algorithm is preferably transmitted to a user after its determination. The user can then take a countermeasure accordingly (for example, manually intervene). In order to simplify transmission to the user, the results can be evaluated. For example, each result can be assigned an assessment. The assessment can then be indicative of a quality of the state. Quality can be understood here to indicate how important or important the condition is. It is conceivable here to provide a color coding which, depending on a displayed color, can signal whether a critical, approximately critical or non-critical state is present. For example, a type of traffic light system can be provided here, in which non-critical results are marked with green color. Here, it is conceivable, for example, for the assigned images to be framed in green. However, if the second learning algorithm detects an error (e.g., one of the above-mentioned circumstances), the evaluation of the state may change. For example, immediately after a state change (i.e., from smooth operation to biased operation), the color marker may take on an orange one. This can help the user not have to execute corresponding countermeasures directly, but still be prepared for any error case that may occur. More precisely, for example, when a traffic jam is detected in reality, a traffic jam can actually be present, but this can also possibly be resolved again by itself. Therefore, for example, the quality of the state may be dependent on the duration of the state. In other words, the quality of the state may change when a particular state continues for the particular duration. Preferably, the quality of the state is determined during the evaluation based on a duration of the state. For example, a condition classified as susceptible to errors can be marked with orange color. If the classification for the endangered state remains for several seconds in succession, for example, the orange color can change to a red color. Furthermore, for a large number of different camera viewpoints, it can be further filtered by displaying to the user only those viewpoints which are currently marked either orange or red or have recently been marked but are now marked green again. In other words, the information provided to a user can thus be limited only to the information that is actually relevant to the user. This allows a human operator or user to be signalled when a traffic jam continues for several seconds and thus an error rate in determining the state is reduced. Furthermore, unnecessarily triggered countermeasures can be avoided. As a result, the workload of the human operators can be reduced by not having to process each item of information and, in addition, not having to initiate countermeasures directly in the case of each faulty state.Preferably, generating the at least one image packet comprises scaling and / or adjusting the resolution of the images of the image data stream, in particular scaling to an RGB value. The image data stream can be recorded, for example, using 1080 p cameras. Further, the cameras may be RGB cameras. For optimum processing of the images, the images of the image data stream can be normalized to a standard resolution and scaling of the RGB values. This provides the advantage that an image data stream can also be processed by cameras of different configuration. In other words, the image data stream can be composed of different images, which have, for example, a different resolution. Even such an image data stream can be used for the present method, since the normalization can always give the same boundary conditions.Preferably, prior to determining the state of the handling device, the features of the images of an image packet are stacked. In other words, the images of each image packet may be stacked. This can be done, for example, by the first learning algorithm. Thus, the input to the second learning algorithm may be a stack of images that exactly corresponds to the images of a packet of images. The second learning algorithm may then assign a state of the handler to each stack. In other words, the stack of images may also be a stack of the properties or features determined by the first learning algorithm. In other words, the features of each image of an image packet may be stacked. Thus, for example, a change of a feature over time (i.e. during a time period covered by the image packet) can be taken into account by the second learning algorithm. Efficient handling is thus made possible. Furthermore, the transfer between the first learning algorithm and the second learning algorithm is simplified, since both can be regarded as modules of a superordinate learning algorithm.Preferably, the image data stream comprises five images per second. More specifically, the image data stream that is used as the basis for generating at least one image packet can comprise five images per second. The image data stream recorded by the video camera can have a greater image frequency (FPS). Thus, the step of generating the at least one image packet can additionally also comprise setting the image frequency. Thus, the input data to the learning algorithm can be optimized, so that the learning algorithm can operate efficiently.Preferably, each image packet comprises 20 images. In other words, an image packet may cover a period of four seconds during the handling operation of the handling apparatus. In this regard, it has been found that this monitored time period is optimal to detect the condition of a handling device of mailpieces.Preferably, the determination of the state of the handling device is a classification of the features of the image packet. In other words, if a plurality of image packets are present, a classification can be carried out for each image packet. The classification can have two statements. On the one hand, it can be classified that troublefree operation is possible, and on the other hand, it can be classified that a problem exists. It is also conceivable that the classification is configured in a two-class classification. This allows a rapid determination of the state of the handling device to be determined.Preferably, the state of the handling device is determined in real time. In other words, during the operation of the handling device, the state of the handling device can be deduced directly. It can thus be ensured that a problem is detected in good time and can be corrected. In real time, this can mean that the state of the handling device is determined only with very little delay. The low retardation may be in the range of 1 to 10 seconds. This can ensure that a reaction can be made quickly to a state of the handling device.Preferably, at least two image packets are generated, and wherein preferably identical images are included in image packets adjoining one another in time. In other words, at least two image packets adjoining one another can overlap. Thus, the same images can be provided in two different image packets. It has been found to be particularly advantageous if an overlap of the images of two image packets adjoining one another is greater than 50%. In other words, a single image may be present in a plurality of image packets. As a result, the determination using learning algorithms can be implemented more robust. Image packets can thus be overlapping, i.e. each individual image can be comprised in a plurality of successive image packets.Preferably, a subsequent image packet comprises only a new image, compared to a previously generated image packet. For example, an image packet may comprise five individual images. The images may be taken in the order in which image packet has been taken as it has been taken in time. In a subsequent image packet, the most recently recorded image can be omitted and a new image can be recorded. In this case, the two image packets adjacent to each other overlap with four identical images. This can enhance the recognition accuracy of the learning algorithm.Preferably, the handling device is controlled based on the determined state. Thus, direct control of the handling device based on the determined state can be realized. In other words, the handling device can be controlled without a user having to actively intervene. The handling device can thus react to any errors that may occur. For example, in the case where the handler is a conveyor belt, the conveying speed may be reduced or increased to improve the state of the handler. As a result, the degree of automation can be further increased, as a result of which the efficiency can be increased.According to an embodiment of the present invention, a result of the neural network is then transmitted to a human operator to physically resolve a jam on the conveyor belt and restore normal operation. For this purpose, a "traffic light" approach is proposed: image sequences from the image packet, which are classified as "not at risk of jamming" by the neural network, are marked with a green color, e.g. as a border. The sequences classified as "stasis" by the neural network are marked with an orange color. If the classification "congestion" of the neural network remains several seconds after one another, the color changes to red. For a large number of camera viewpoints, they can be further filtered by displaying to the operator only the viewpoints currently marked either orange or red or that were in the past a short time but are now marked green again. This traffic light concept provides the advantage of signalling the human operator when a congestion continues for several seconds, thus reducing the error rate. In addition, human operator workload is reduced by showing only camera poses that are of interest and omitting camera poses that operate normally for a period of time.According to another aspect of the present invention, a method for training or retraining a first learning algorithm and / or a second learning algorithm for determining a state of a handling device for mailpieces is provided. The method comprises receiving input training data, namely annotated images of the handling device during a first time period, wherein the images are provided in ascending time sequence, in succession, receiving output training data, namely a state of the handling device for mail items during the first time period, and training or retraining the first learning algorithm and / or the second learning algorithm based on the input training data and the output training data.The input training data can be provided as image packets. During training, weights and / or variables of a learning algorithm are adjusted in order to reduce the error. Training of the learning algorithms can be performed by a collection of annotated training and evaluation data. Since the detection of a state may be modeled as a continuous classification task, no polygons or explicit object detection are necessary. The annotated training data can consist of the recorded camera images which are provided in a temporally ascending manner and without missing images. Thus, subsampling in the temporal dimension is not necessary. Omitting the subsampling can make the process efficient. Nevertheless, there may be situations where subsampling is necessary. The subsampling can be used, for example, to adapt an image rate per unit time of the individual images to a desired image rate. For example, the first and / or second learning algorithm may have been trained on image data at a specific image rate per unit time. During subsampling, this frame rate can be adjusted. In other words, if the input data is present at a frame rate of 30 frames per second and if a learning algorithm is present that has been trained at a frame rate of 5 frames per second, 25 frames of the input data can be removed during subsampling, so that the input data matches the learning algorithm. Thus, under subsampling may mean adapting the frame rate to the trained frame rate. Surplus data (e.g., images) can be removed in a substituteless manner. The annotation of the data may consist of intervals (from-to-ranges), of time stamps associated with the beginning and end of a single state. Thus, there may be a 1:1 correspondence between physical conditions on the handler and the intervals in the annotated data. The end-to-end optimization of the learning algorithms can comprise a supervised learning substitute with backpropagation and gradient descent. The input to the learning algorithm may be a stack of images (i.e., an image package) that exactly correspond to the images of an image package that was captured by the camera. The output of the learning algorithm may comprise a probabilistic mapping of whether or not this stack of images has a particular state. A cross entropy loss function may be applied to the prediction of the learning algorithm and the annotated truth intervals, completing the optimization step for the parameters of the learning algorithm. In this case, the learning algorithm can be trained for the first time or an already existing learning algorithm can be retrained. For example, it is conceivable that centrally collected data for training a learning algorithm is used for optimizing state detection. Preferably, the training, or the retraining, can comprise an end-to-end optimization by means of back propagation and / or gradient descent.According to a further aspect of the present invention there is provided a computer program comprising instructions which, when the program is executed by a computer unit, cause the computer unit to carry out methods according to any of the above embodiments. This applies to the method for determining a state of a handling device as well as to the training method for the first and / or the second learning algorithm. Alternatively, the learning algorithms can also be implemented as hardware, e.g. with fixed connections on a chip or another computer unit. The computer unit that can execute the method according to the invention can be any computer unit such as a CPU (central processing unit) or GPU (graphics processing unit). The computer unit can be part of a computer, a cloud, a server, a mobile device such as a laptop, tablet computer, mobile telephone, smartphones, etc. In particular, the computer unit can be part of a monitoring system for determining a state of a handling device. The monitoring system may include a display device, such as a computer screen.The invention also relates to a computer-readable medium comprising instructions which, when executed by a computing unit, cause the computing unit to carry out the method according to the invention, in particular the above method. Such a computer-readable medium may be any digital storage medium, for example a hard disk, a server, a cloud or a computer, an optical or a magnetic digital storage medium, a CD-ROM, an SSD card, an SD card, a DVD or a USB or other memory stick. Furthermore, the computer program can also be obtained via the Internet.According to a further aspect of the present invention, a handling device for handling mailpieces is provided. The handling device comprises at least one sub-component for handling mailpieces, and a control unit configured to control the at least one sub-component based on the state of the handling device, wherein the state of the handling device is determined by a method according to one of the above embodiments. This component may be, for example, a conveyor belt, grippers, or any other manipulation device configured to physically influence a mail piece.According to a further aspect of the present invention, a monitoring system for determining a state of a handling device for handling mailpieces is provided. The monitoring system may comprise a control unit configured to perform a method according to any of the above embodiments.The monitoring system may take an output of one for outputting the state of the handling device. According to a further aspect of the present invention, a use of the monitoring system for determining a state of a handling device is provided. The monitoring system can correspond to the above embodiments.According to one embodiment, the advantage is provided that the described neural network can learn an implicit representation of the conveyor belt, mailpieces, and their normal flows. The time is implicitly encoded or explicitly maintained, which in both cases allows recognition of jams on the conveyor belt. The advantages are that the overall system is very robust, since the neural network is optimized throughout. Optimization is a supervised learning approach, but the annotated training data is cost effective to produce because it consists only of camera images in combination with time intervals (time stamp to time stamp) that show a congestion on the conveyor belt. The traffic light approach relieves human operators and reduces the error rate of the entire detection system. The main advantages of an embodiment of the present invention is that it is an end-to-end approach without intermediate steps and that a "traffic light" approach is used to show only interesting camera perspectives to human operators.Individual features or embodiments may be combined with other features or other embodiments to form novel implementations. Advantages and configurations which are mentioned in connection with the individual features or the embodiment then apply analogously also to the novel embodiment. Advantages and configurations mentioned in connection with the method also apply analogously to the device and vice versa.Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings. FIG. 1 is a schematic view of a handling apparatus according to an embodiment of the present invention. FIG. 2 is a schematic flow diagram of a method according to an embodiment of the present invention. FIG. 3 is a schematic flow diagram of a method according to an embodiment of the present invention. FIG. 4 is a schematic illustration of an aspect of the present invention.FIG. 1 is a schematic illustration of a handling system 1 according to an embodiment of the present invention. In addition, on the right-hand side of FIG. 1, an image data stream 4 having a multiplicity of individual images 7 is schematically shown. The handling device 1 of the present embodiment is configured to handle mailpieces 2. In the present embodiment, the handling device 1 is a conveyor belt which moves mail pieces 2 in a conveying direction. Furthermore, a video camera 3 is provided which can image the handling device 1. The camera 3 has a viewing angle 31 with which the handling device 1 is imaged. The camera 3 is preferably a video camera. The camera 3 outputs an image data stream 4. The image data stream 4 has a multiplicity of individual images 7. In the present embodiment, the image data stream 4 is obtained directly from the camera. Image packets 5, 6 are then generated from the image data stream 4. Each image packet 5, 6 has a specific number of individual images 7. In the present embodiment, the first image packet 5 comprises three individual images 7 of the image data stream 4. The second image packet 6 likewise has three individual images 7. A plurality of image packets can be provided overall, which can all have the same number of individual images 7. The first image packet 5 and the second image packet 6 overlap in that they comprise two identical individual images 7. In other words, the first image packet 5 and the second image packet 6 differ in that they have a different image that the respective other image packet does not have. The image packets are then fed to a first learning algorithm 9. The first learning algorithm 9 then extracts or determines features 8 for each image of an image packet.FIG. 2 is a schematic view of an operation of an embodiment of the present invention. In FIG. 2, the individual images 7 of an image packet are supplied. Each image 7 of an image packet 5, 6 is fed to the first learning algorithm 9. Each image 7 is analyzed with the same trained parameters of the 91 of the first learning algorithm. As output of the first learning algorithm 9, features for each image 7 of an image packet 5, 6 come out as output. Optionally, an area of interest 12 may be determined in each of the images. In the present embodiment, a region of interest is determined. However, this is optional and may also be omitted. Based on the region of interest 12, a pixel mask 13 is provided. The pixel mask 13 is then applied to the features 8 determined by the first learning algorithm 9. Thus, those features that are under the pixel mask are obscured. Subsequently, the features 8 of an image packet can be stacked. The staggering 14 is carried out individually for each image packet 5, 6. The stacked features are then fed to the second learning algorithm 11. The latter determines a state 10 of the handling device based on the features of the individual images. In a further embodiment, the stacking 14 can furthermore omit the temporal and spatial dimension in the extracted features 8. This is particularly advantageous when using a multi-layer perceptron as the second learning algorithm 11.FIG. 3 is a schematic flow diagram of a method according to another embodiment of the present invention. Up to step 13, the present embodiment corresponds to the embodiment described in FIG. 2. After step 13, in the present embodiment, the spatial dimension 15 can be omitted for the extracted features of each individual image 7. Subsequently, the extracted features may be supplied to the second learning algorithm 11 configured as a long short term memory in the present embodiment. the LSTM of the present embodiment is a bidirectional LSTM 16. Subsequently, a temporal dimension may be omitted 17. Then, the extracted features may be supplied to a feed forward neural layer 18. Based thereon, a state 10 of the handling device can be output.Fig. 4 is a schematic illustration of another embodiment of the present invention. The round circles symbolise the individual image packets and the respective time at which they were recorded. The first image packet 5 was recorded, for example, at a time t=1.The second image packet 6 was recorded at a time t=2. The same applies to the further n image packets shown. If a line is shown in the image packet, then no abnormality has occurred in the determination of the state of the handling device 1. When a "J" is illustrated, there is an abnormality in operation of the handling apparatus (a congestion; J=Ja in the present embodiment). In the present embodiment, the state results are evaluated. No abnormality has been detected in the first three image packets, which is why the first quality 19 (e.g. green) can be assigned here. In the fourth and fifth image packs, it is indicated that an abnormality is present. This is defined as second grade 20 (orange). In the second quality 20, there is still no acute need to initiate countermeasures. Only in the sixth image packet is the third quality level 21 (red) displayed. In this case, some time has elapsed and it can be assumed that the anomaly does not resolve itself. The same applies to the seventh image packet. At this level of quality, a user is requested to take countermeasures. In the present case, this has taken place, so that the first quality level 19 (green) is again displayed in the eighth image packet. In the present example, the time profile runs from the left image side to the right image side (see arrow).List of reference numbers:1 Handling device 2 Mail piece 3 Camera 4 Image data stream 5 First image packet 6 Second image packet 7 Images 8 Features 9 First learning algorithm 10 State of the handling device 11 Second learning algorithm 12 Region of interest 13 Pixel mask 14 Stacks 15 Spatial dimension 16 Bidirectional LSTM 17 Temporal dimension 18 Neural layer with feed-forward 19 First quality 20 Second quality 21 Third quality
Claims
Computer-implemented method for determining a state (10) of a handling device (1) for mailpieces (2), comprising: a) determining or obtaining an image data stream (4) of the handling device (1) when handling mailpieces (2), b) generating at least one image packet (5, 6) from the image data stream (4), wherein the image packet (5) comprises a plurality of individual images (7), c) determining features (8) of each image (7) of an image packet (5, 6) by means of a first learning algorithm (9), and d) determining the state (10) of the handling device (1) based on the features (8) of each image (7) of an image packet (5, 6) by means of a second learning algorithm (11), wherein the first learning algorithm (9) is different from the second learning algorithm (10).Method according to claim 1, wherein the first learning algorithm (9) is an artificial neural network, in particular a convolutional neural network, and / or wherein the second learning algorithm (11) is an artificial neural network, in particular a multi-layer perceptron or long short-term memory.Method according to one of the preceding claims, wherein the method further comprises: determining an area of interest (12) in the images (7), wherein in particular only the features (8) in the area of interest (12) are used as a basis in the determination of the state (10).Method according to one of the preceding claims, wherein the state (10) is indicative of a probability that a jam of mailpieces (2), vertically standing mailpieces, damaged mailpieces, double / multiple packs of mailpieces, mailpieces fallen by the handling device (1), stuck mailpieces, dumped liquids and / or loose mailpiece content is present in the handling device (1).The method according to any of the preceding claims, wherein the method further comprises: evaluating the condition (10), wherein the evaluation is indicative of a quality (19, 20, 21) of the condition (10).Method according to one of the preceding claims, wherein the generation of at least one image packet (5) comprises a scaling and / or adjusting of the resolution of the images (7) of an image packet (5), in particular a scaling to an RGB value.Method according to one of the preceding claims, wherein, before the state (10) of the handling device (1) is determined, the features (8) of the images (7) of an image packet (5) are stacked (14).Method according to one of the preceding claims, wherein the determination of the state (10) of the handling device (1) is a classification of the features (8) of an image packet (5).Method according to one of the preceding claims, wherein at least two image packets (5) are generated, and wherein identical images (7) are included in image packets (5) adjoining one another in time.Method for training or retraining a first learning algorithm (9) and / or a second learning algorithm (11) for determining a state (10) of a handling device (1) for mail items (2), comprising: receiving input training data, namely annotated images (7) of the handling device (1) during a first time period, wherein the images (7) are provided in ascending time sequence, consecutively, receiving output training data, namely a state (10) of the handling device (1) for mail items (2) during the first time period, training or retraining the first learning algorithm (9) and / or the second learning algorithm (11) based on the input training data and the output training data.A computer program comprising instructions which, when the program is executed by a computing unit, cause the computing unit to carry out the method according to any one of claims 1 to 10.Monitoring system for determining a state (10) of a handling device (1) for handling mailpieces (2), comprising: a control unit which is configured to execute a method according to one of Claims 1 to 9.
Citation Information
Patent Citations
Cloth inspection method, device, terminal device, server, storage medium and system
CN109215022A
CN000109215022A