Method for processing at least one image

By inputting unmasked image data into the neural network and creating an image mask that depends on the region of interest, the problem of neural network difficulty in identifying when the data is not normally distributed is solved, achieving more accurate detection of pre-training target area and simplifying the training process.

CN120051778APending Publication Date: 2025-05-2736ZERO VISION GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072319.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-11
Filing Date
2023-10-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art neural network can provide accurate results when the input data is normally distributed, but when the mask is applied to remove the uninterested image part, the color values ​​in the remaining image data are no longer normally distributed, resulting in the neural network being unable to accurately identify the pre-trained target area.

Method used

By receiving image data and inputting unmasked image data to the artificial neural network, an image mask depends on the region of interest is detected and created, and finally the mask is applied to the image data to remove image areas that do not include the region of interest.

Benefits of technology

Ensure that the color values ​​of the input data are always normally distributed, thereby improving the recognition accuracy of neural networks, simplifying the training process and improving hardware performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051778A_ABST
    Figure CN120051778A_ABST
Patent Text Reader

Abstract

The invention relates to a method for processing at least one image (1), wherein the method comprises the following steps: receiving image data; inputting the unmasked image data into an artificial neural network to detect one or more pre-trained target regions (3); detecting one or more regions of interest (2) in the image (1) and creating an image mask dependent on the one or more regions of interest (2); the created image mask is applied to the image data to remove image regions (4) that do not include the one or more regions of interest (2).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for processing at least one image. Additionally, the present invention relates to a data processing device and an image acquisition device including such a data processing device. Furthermore, the present invention relates to a computer program product, a computer-readable medium, and a data carrier signal. Background Art

[0002] Image processing by means of an artificial neural network is known from the prior art. A neural network can be used to detect pre-trained target regions in an image. For this purpose, a mask is applied to the image to remove image portions that the application is not interested in. Then, the masked image data is input into the neural network. However, the neural network provides accurate results when the input data is normally distributed. For example, if the color values are normally distributed, the neural network works well. Applying a mask to remove one or more image portions that are not of interest results in the color values in the remaining image data input into the neural network not being normally distributed. This causes the neural network to be unable to accurately identify the pre-trained target regions. The same problem also applies to the training process when training images are applied to the neural network. This causes the neural network not to be trained to provide accurate results. Summary of the Invention

[0003] An object of the present invention is to provide a method by which pre-trained target regions can be more accurately identified.

[0004] This object is achieved by a method for processing at least one image, wherein the method comprises the following steps:

[0005] Receiving image data;

[0006] Inputting the unmasked image data into an artificial neural network to detect one or more pre-trained target regions;

[0007] Detecting one or more regions of interests in the image and creating an image mask depending on the one or more regions of interests; and

[0008] Applying the created image mask to the image data to remove image regions that do not include the one or more regions of interests.

[0009] According to the present invention, unmasked image data is input into a neural network. This ensures that the data input during normal operation or training, in particular the color values, is normally distributed. Therefore, the neural network provides accurate results. In particular, the neural network detects one or more pre-trained target regions in an accurate manner. It has been recognized that the detection of target regions and the detection of regions of interest should be performed separately from each other. In particular, the same image data is used for the detection of target regions and the detection of regions of interest. This results in better results than masking the received image data before inputting it into the neural network.

[0010] As will be described in detail later, by removing image regions that do not include one or more regions of interest, target regions of interest to the user and / or application can be identified in a simple manner. This has the advantage of simplifying training because the training can focus on identifying regions of interest and target regions. This means that all other non-interested objects do not need to be trained, and the training is independent of these non-interested regions of interest, in particular non-interested objects. The training can be performed separately for target regions and regions of interest, which also simplifies the training. Another advantage of the present invention is that the separate detection of one or more target regions and one or more regions of interest improves the hardware performance during the training phase and the operation phase. Another advantage of the present invention is that the separate detection of multiple regions of interest that can overlap or not overlap enables the parallel application and combination of different detections of one or more target regions to each region of interest.

[0011] The so-called "pre-training" means that the neural network is trained to detect target regions during the training process based on training images.

[0012] The pre-trained target regions can be flexibly selected. It can be one or more parts of one or more objects. Alternatively, the pre-trained target region can be one object or multiple objects. The pre-trained target region can depend on the application scenario of the method. If the method is used for quality inspection, the pre-trained target region can be one or more defects, such as scratches. In the method of the present invention, one or more defects of interest are arranged in the region of interest, in particular in the object. Alternatively, the pre-trained target region can also be the color of the object or object part, so that it can be checked whether the object or object part has a predetermined color. Another example is that the pre-trained target region can be multiple parts of an object, so that the method can check whether the image data includes the multiple object parts. If this is not the case, an error may have occurred during the object manufacturing process.

[0013] The expression "detecting one or more pre-trained target regions" means that the neural network determines all pre-trained target regions in the input image. Thus, the output of the neural network is whether there is one or more pre-trained target regions in the image, especially in the image data. This also includes the case where the image does not include a target region. In this case, the output of the neural network is that the image does not include a target region.

[0014] The expression "detecting one or more regions of interest" means that all target regions are determined in the input image. This also includes the case of determining that the input image does not include a region of interest. In this case, the image mask will remove the entire image because the entire image is considered non-interest. The region of interest can be part of the image or can cover the entire image.

[0015] The so-called "unmasked image data" means that the image data corresponds to the original image data acquired by the image acquisition device. By applying an image mask to the received image data, the data is changed, i.e., masked, such that it is inconsistent with the original image data. If the region of interest covers the entire image, then no image part is removed by applying the image mask.

[0016] If more than one image is obtained, each image is processed in the above manner. This means that for each image, one or more pre-trained target regions are detected, and one or more regions of interest are detected. In addition, an image mask is created for each image, and the corresponding image mask is applied to the corresponding image. Obviously, the image data received in the method of the present invention corresponds to the image, especially to the obtained image.

[0017] The method can be executed in a data processing device. The data processing device can include one or more processors or can be a processor. Alternatively, the data processing device can be a computer. The image data obtained by the image acquisition device is sent to the data processing device. Thus, the data processing device receives the image data from the image acquisition device and processes the image data.

[0018] According to one embodiment, the data processing device may process the image data such that one or more pre-trained target regions arranged in at least one (especially pre-trained) region of interest are detected. Thereby, the positions of the pre-trained target regions and the position of the (especially pre-trained) region of interest can be determined. Thus, it can be determined whether the detected pre-trained target regions are arranged in the (especially pre-trained) region of interest. One or more pre-trained target regions arranged in at least one (especially pre-trained) region of interest may be output by a neural network. This is only possible if the image data includes one or more pre-trained target regions. One or more pre-trained target regions arranged in at least one (especially pre-trained) region of interest are user-related target regions. Since the pre-trained target regions can be accurately detected (because they are detected separately from the (especially pre-trained) region of interest), all pre-trained target regions of interest can be considered. As described below, other pre-trained target regions not arranged in the region of interest may be ignored in the further processing of the image data or the pre-processing target regions.

[0019] Before the image mask is applied to the acquired image, one or more pre-trained regions of interest may be detected. As explained below, the detection of one or more pre-trained regions of interest may be independent of the detection of the pre-trained target regions and vice versa. That is, one or more pre-trained regions of interest may be detected before, during, or after the detection of the pre-trained target regions. This also applies to the training process discussed above. The training process is also independent of whether the region of interest or the target region is detected first.

[0020] The pre-trained region of interest may be at least a part of at least one object or at least one object. The object may be a discrete object. A discrete object is an object having a clear boundary and a spatial extent, and thus having spatially invariant properties. Thus, the object can be anything that is visible and / or tangible and that can be touched, such as a car, a chair, etc. Additionally or alternatively, the object may also be small enough so that its image can be acquired by an image acquisition device, especially by a camera and / or a mobile phone and / or a tablet and / or a microscope, etc.

[0021] The pre-trained target region may be different from the region of interest in terms of optical properties. For example, the pre-trained target region may have a different color, brightness, etc. Additionally or alternatively, the pre-trained target region may correspond to or be smaller than the region of interest. In other words, the pre-trained target region may have a specific shape, where the cross-section of the shape may be smaller than the cross-section of the region of interest. The pre-trained target region may correspond to a defect in the region of interest, particularly in an object. Additionally or alternatively, the pre-trained target region may also have optical properties, such as color, that are different from those of the pre-trained region of interest.

[0022] The image mask may be configured such that it removes one or more pre-trained target regions arranged outside the pre-trained region of interest. Additionally or alternatively, after one or more training target regions are detected, the image mask may be applied to the image data. This means that the image is configured such that it does not remove one or more pre-trained target regions arranged in at least one region of interest. If the image data does not include a region of interest, the image mask removes all pre-trained target regions. The masking of the image data is used to simplify and / or accelerate the processing of the pre-trained target regions, as all image regions that are not of interest and / or not relevant are removed from the image data. Thus, the image processing speed is faster after the image mask, as the image data to be processed is smaller than if the entire image data were processed. The image mask may be applied to the complete image data, not just to a part of the image data.

[0023] The process of creating the image mask is as follows. The received image is represented by a matrix. The matrix may include the values of each pixel of the image. The image mask may also be represented by a matrix. In the first step, the values of the image mask may have a value of 0. After one or more regions of interest are detected, the values of the image mask may be modified in a second step. The modification of the pixel values of the image mask occurs such that by applying the image mask with the modified values to the received image region (the image region excluding the region of interest), it is removed. This can be achieved if the pixel values corresponding to the non-interest regions are 0. Alternatively, the pixel values may have different values, based on which it can be determined that the pixel does not belong to the region of interest.

[0024] The image mask determined as described above and represented by a matrix is applied to the received image as follows. As described above, the received image may also be represented by a matrix. By applying the image mask to the received image, the output is a matrix, where the values will be the values of the received image or the values modified by the image mask. The value may be modified to 0 by the image mask, but it may also be a different value. Finally, after the image mask is applied to the image, the pixels belonging to the region of interest can be identified by their values.

[0025] One or more target regions arranged in the pre-trained region of interest can be visualized. This means that the output of the method and / or the data processing device can be visualized, for example, on a display. Additionally or alternatively, the pre-trained region of interest and / or the removed image region are not visualized.

[0026] According to an embodiment, the neural network can be a convolutional neural network. The neural network can include at least two layers. In the case where the neural network includes only two layers, the neural network includes an input layer and an output layer. The neural network can include more than two layers.

[0027] The neural network can be configured to also create an image mask and / or detect one or more regions of interest. In this case, the neural network runs all tasks, namely, detecting pre-trained targets, creating an image mask, and detecting regions of interest. These tasks can be run in parallel with each other or sequentially with each other. The output of the neural network is the detected pre-trained target region, the image mask, and the detected region of interest. In this case, only one neural network is executed in the data processing device.

[0028] Alternatively, another neural network can be executed in the data processing device. The other neural network can be configured such that it creates an image mask and / or detects regions of interest. The other neural network can output the created image mask and the detected regions of interest. The neural network and the other neural network can be run in parallel with each other or sequentially with each other. The same image data is input to both the neural network and the other neural network.

[0029] The other neural network can be a convolutional neural network. The other neural network can include at least two layers. In the case where the other neural network includes only two layers, the other neural network includes an input layer and an output layer. The other neural network can include more than two layers. In the said embodiment, the neural network outputs the detected pre-trained target region, and the other neural network outputs the detected region of interest and the created image mask.

[0030] One or more pre-trained target regions arranged in at least one region of interest are determined by the intersection of the detected one or more pre-trained target regions and the detected one or more regions of interest. The intersection processing can be run by the data processing device based on the outputs of the neural network and / or the other neural network. Alternatively, it is possible for the intersection processing to be performed by the neural network. This is possible if only one neural network is provided.

[0031] Intersection processing runs on the detected two-dimensional image and works as follows. One or more detected target regions can be assigned one or more polygons, and one or more detected regions of interest can be assigned one or more other polygons. This means that one polygon is assigned to each target region and another polygon is assigned to the region of interest. The polygons are configured such that they enclose the target regions. Similarly, the other polygon is configured such that it encloses the region of interest. When the polygon of the target region and the other polygon of the region of interest overlap or have a common surface, the target region and the region of interest intersect each other. In this case, depending on the overlapping part, a part or the whole of the target region is considered. When the polygon and the other polygon have no overlapping part or no common surface, the target region and the region of interest do not intersect, and thus the target region is not considered.

[0032] According to an embodiment, the data processing device may be configured to detect a region of interest based on a provisory region of interest. The provisory region of interest is detected by a neural network or another neural network based on the input image data. The provisory region of interest generally does not match the region of interest shown in the image. Therefore, the data processing device may determine whether the edge of the region of interest shown in the image is arranged in a predetermined region including at least a part of the edge of the provisory region of interest. The predetermined region may be a region having a predetermined number of pixels in the height direction and the width direction.

[0033] If the provisory region of interest deviates from the edge of the region of interest shown in the image in the predetermined region, the data processing device may set the region of interest to the region of interest shown in the image, particularly the edge of the region of interest. If the predetermined region does not include the edge of the region of interest deviating from the provisory region of interest, the data processing device sets the provisory region of interest as the region of interest. The above method improves the detection of the region of interest. In particular, it can ensure that the detected region of interest (particularly an object) corresponds to the region of interest (particularly an object) shown in the image.

[0034] The data processing device may determine the number of detected pre-trained target regions. Alternatively, the data processing device may determine whether the number of pre-trained target regions corresponds to a predetermined number of target regions. The number of pre-trained target regions can be used for quality control. The data processing device may determine the quality of the region of interest based on one or more detected pre-trained target regions. For example, when the number of detected pre-trained target regions does not correspond to the predetermined number of target regions, the data processing device may determine that the region of interest (particularly an object) has poor quality.

[0035] According to an embodiment, a training process is performed. During the training process, a neural network and / or another neural network are trained during an operation phase. The training objective is for the neural network to detect one or more target regions in the image data. If only one neural network is provided, the training objective is for the neural network to additionally detect one or more regions of interest and create an image mask. For the case where another neural network is executed on the data processing device, the objective of the other neural network is for the other neural network to detect one or more regions of interest and create an image mask.

[0036] The training process of the neural network can correspond to the training process of another neural network. This means that the two neural networks can be trained in the same way. This means that the same training images can be input to the two neural networks. The neural network can be trained using unmasked or masked training images, and / or another neural network can be trained using unmasked or masked training images. The so-called "unmasked image" means that no mask is applied to the training image, and the training image input to the neural network and / or another neural network corresponds to the original training image. Therefore, in the unmasked training image, no object is removed. The training image can show multiple objects that do not correspond to the regions of interest.

[0037] The training images input to the neural network and / or another neural network during the training phase can include context information. A training image with context information means that the training image includes the target region and / or region of interest to be trained. In contrast, a training image with non-context information means that the training image may or may not include the target region and may include multiple different objects, i.e., non-regions of interest.

[0038] The training process for training the neural network and / or another neural network can include two training phases. The training can be run to train the detection of the target region and / or region of interest.

[0039] In the first training phase, training images including non-background information are input into the neural network and / or another neural network. The first training phase can be run only once and can be considered as general training, in which the neural network is trained with multiple different objects. In the first training phase, the neural network and / or another neural network are trained using annotated images. By "annotation" it is meant that the image data contains information about the image height, image width, and other image information (such as color). In addition, the image signal contains information about the objects (such as screws, chairs, stairs, etc.) provided in the image signal. In this case, the training images input in the first training phase can include target regions and / or regions of interest. The specified or classified objects are at least partially surrounded by a bounding box so that the neural network can identify the position of the predetermined transport material in the image data. The bounding box can be a polygon.

[0040] During the first training process, a large number of images, especially for example millions of images, as described above, which contain information about the objects and the object positions, are input into the neural network to be trained. As described above, the images can show various different objects, where the target region and / or region of interest can be included in the image, but do not have to be.

[0041] After the first training phase, the neural network and / or another neural network are trained in the second training phase. In the second training phase, training images including the target region and / or region of interest can be input into the neural network or another neural network. The second training phase is used to train the neural network and / or another neural network for the application in which the neural network is to be used. Generally, if the neural network is to be used for other applications, retraining is required.

[0042] During training, the neural network to be trained and / or another neural network can be input with training images including the target region and / or region of interest and possibly including other objects, as well as training images that do not include the target region and / or region of interest. However, it is advantageous if 20% - 100%, especially 80% - 95%, preferably 90% - 95% of the provided training images show the target region and / or region of interest. In the second training phase, the same training images as in the first training phase can be provided. Alternatively, different training images can be provided. In the second training phase, at least one training image, especially multiple training images, can be input into the neural network and / or another neural network. Thus, some of the training images can be annotated and some can be unannotated. Alternatively, all images can be unannotated. In this case, all training images including the target region and / or region of interest will be annotated. While the training images that do not include the target region and / or region of interest are not annotated.

[0043] According to another aspect, a data processing device is provided. The data processing device includes means for performing the method of the present invention. Additionally, an image acquisition device for acquiring images is provided. The image acquisition device can be at least one of the following: a camera, a mobile phone, a microscope, and a tablet computer. The data processing device can be part of the image acquisition device. Alternatively, the data processing device can be electrically connected to the image acquisition device. The acquisition device can be configured to acquire visible light. "Visible light" means that the wavelength of the acquired light is between 380 and 750 nanometers.

[0044] According to yet another aspect of the present invention, a computer program product is provided, wherein the computer program product includes instructions that, when the program is executed by a data processing device (in particular, a computer), cause the data processing device (in particular, a computer) to perform the steps of the method of the present invention. Additionally, a computer-readable data carrier is provided, on which the computer program product is stored. A data carrier signal is also provided, which carries the computer program product.

[0045] The method of the present invention can be used in a conveying system. In particular, the method can be used to determine the quality of an object (in particular, goods) transported in the conveying system. The object (in particular, goods) can correspond to the region of interest, or form part of the region of interest, or a part of the object (in particular, goods) can be the region of interest. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In the drawings, the subject matter of the present invention is shown schematically, and elements that are the same or have similar functions mostly have the same reference numerals. The following are shown:

[0047] Figure 1 A device with an image acquisition device and a data processing device for performing the method of the present invention according to a first embodiment is shown.

[0048] Figure 2 A device with an image acquisition device and a data processing device for performing the method of the present invention according to a second embodiment is shown.

[0049] Figures 3A - 3B An image showing different processing steps of the method of the present invention.

[0050] Figure 4A A training phase for training a neural network to detect a target region is shown.

[0051] Figure 4B A training phase for detecting a region of interest is shown.

[0052] Figure 5 A process for determining a region of interest is shown.

[0053] Figure 6 Shows a transfer system including a device according to Figure 1 or Figure 2 . Detailed implementation

[0054] Figure 1 Shows a device 8 including an image acquisition device 6 and a data processing device 5, where the data processing device 5 is electrically connected to the image acquisition device 6 such that data transmission between the two devices is possible. The image acquisition device 6 may include a camera, or be a camera. Alternatively, the image acquisition device 5 may also be a camera, or a mobile phone, or a tablet computer, or a microscope. The data processing device 5 may include at least one processor, or be a processor. Alternatively, the data processing device 5 may be a computer.

[0055] The image is acquired through the optical device (such as a lens) of the image acquisition device 6. The acquired image data is output through the output part 60 of the image acquisition device 6. The output part 60 is connected to the input part 50 of the data processing device 5 such that data exchange between the image acquisition device 6 and the data processing device 5 is possible.

[0056] The data processing device 5 includes a processing part 51. The processing part 51 includes a first processing part 510 and a second processing part 511. The first processing part 510 includes a neural network, and the second processing part 511 includes another neural network. Both parts receive the image data received from the input part 50. Additionally, the data processing device 5 includes an intersection part 52. The intersection part 52 receives the outputs of the neural network and the other neural network. Additionally, the intersection part 52 serves as the output of the data processing device 5 and is connected to the display input part 90 to enable data exchange between the data processing device 5 of the device 8 and the display screen 9.

[0057] The method for processing the image acquired through the image acquisition device 6 will be introduced below.

[0058] In a first step, the image acquisition device 6 acquires an image 1. The image 1 is transmitted to the data processing device 5, in particular to the input part 50 of the data processing device 5. The received image data is processed in the processing part 50. In particular, the image data is transmitted to a neural network of the first processing part 510 in a second step. The transmitted image data is not masked. The neural network detects one or more pre-trained regions 3 in the second step. The output of the neural network is whether the image data includes one or more pre-trained target regions 3. In particular, the output of the neural network is the number and position of one or more pre-trained target regions 3 in the image 1. The output also includes the case where the image data does not include a pre-trained region 3. In such a case, the output of the neural network is that the image data does not include a pre-trained region 3. The neural network is processed in a first processing part (not shown) of the data processing device 5.

[0059] The neural network is pre-trained to detect one or more pre-trained target regions 3. The training of the neural network runs during a training process T, which is described in more detail in Figure 4A The neural network can be a convolutional neural network. Additionally or alternatively, the neural network can also include at least two layers.

[0060] In a third step, one or more regions of interest 3 in the image 1 are detected in the processing part 51. The detection can be run by the neural network discussed in the second step S2 above.

[0061] Alternatively, the detection can be run by another neural network. This case is shown in Figure 1 Thus, in the third step, the image data received through the input part 50 is transmitted to another neural network of the second processing part 511. The transmitted image data can be unmasked. In the third step, another neural network detects whether the image data includes one or more regions of interest. Additionally, in the third step, another neural network creates an image mask, where the image mask depends on one or more regions of interest 2. The image mask is created after one or more regions of interest 2 are detected. In the third step, the image mask is applied to the image data. This results in all image data that does not include one or more regions of interest 2 being removed.

[0062] The output of another neural network is whether the image data includes one or more regions of interest. In particular, another neural network determines the number and / or position of the regions of interest and the shape of the regions of interest. By shape is meant the edge of the determined region of interest. Image regions 4 that do not include one or more regions of interest 2 are not output.

[0063] Another neural network is pre-trained to detect one or more regions of interest 2. The training of the other neural network runs during the training process T, which is described in more detail in Figure 4B The other neural network may be a convolutional neural network. Additionally or alternatively, the other neural network may include at least two layers.

[0064] The received image data is input into the neural network and the other neural network. The neural network and the other neural network may process the input image data in parallel.

[0065] In the fourth step, the output of the neural network and the output of the other neural network intersect in the intersection part 52. This means that by intersecting the outputs of the two neural networks, the target region 3 arranged in the region of interest 2 is detected. The output of the fourth step is the information that the target region 3 is arranged in the region of interest 2. The said output is transmitted to the display input section 90 and is displayed in the fifth step.

[0066] The region of interest 2 in which one or more target regions 3 are arranged may be displayed. Alternatively, the region of interest 2 is not displayed such that only one or more target regions 3 are displayed.

[0067] The above method steps are executed in the data processing device 5 and specifically illustrate the computer-implemented method of the present invention for processing the image 1.

[0068] Figure 2 The device 8 with the image acquisition device 6 and the data processing device 5 for performing the method of the present invention according to the second embodiment is shown. The second embodiment is different from the Figure 1 shown first embodiment in the processing sequence. In the second embodiment, the image data received from the image acquisition device 6 is input into the neural network of the first processing component 510 only in the second step. The input image data is not masked.

[0069] Then, in the third step, the output of the neural network is input into the other neural network of the second processing part 511. The other neural network processes the output of the neural network. In particular, the other neural network processes the image data that is part of the output of the neural network in the manner described in Figure 1 The other neural network does not process the target region 3 detected by the neural network of the first processing part 510.

[0070] The data processing device 5 is configured to process the output of the other neural network in the fourth step. The output of the other processing device corresponds to Figure 1 the output of the shown neural network and the other neural network. This means that, Figure 2The output of another neural network shown is information about one or more target regions and information about one or more regions of interest 2. The information is processed in a fourth step such that an intersection process is performed in the intersection part 52. In the intersection process, one or more target regions 3 arranged in the region of interest 2 are identified. Other method steps correspond to Figure 1 the method steps discussed in

[0071] Another difference from the Figure 1 discussed method is that the data processing device 5 must include a neural network and another neural network to process the second and third steps. In contrast, as described above, in the Figure 1 shown embodiment, the second and third steps can be run by the same neural network.

[0072] Figures 3A - 3D The displayed image 1 shows different processing steps of the method of the present invention. Figure 3A The displayed image shows the image acquired by the image acquisition device 6. The image data of the acquired image 1 is transmitted to the data processing device 5. In both embodiments, the image data is input into the neural network. As Figure 3B shown, the neural network in the processing part 51 detects the target region 3 without determining the region of interest 2. This means that the output of the neural network is the position and number of the target regions 3. This information is processed by the data processing device 5 in the fourth step. The image data input into the neural network is not masked. This means that no image part in the image acquired by the acquisition device 6 is deleted, and the input image data corresponds to the original image data acquired by the image acquisition device 6.

[0073] Figure 3C The display shows the output of another neural network in the processing part 50. Figure 3A The image data of the image 1 shown in Figure 3C is input into another neural network. Another neural network creates an image mask depending on the region of interest 2. Additionally, another neural network applies the image mask to the input image data. Thus, the image region 4 that does not include the region of interest 2 is removed. The image region 4 and the region of interest 2 are shown in Figure 3C The image data of the image shown in

[0074] Figure 3D The display shows the result after the intersection process is run by the data processing device 5. During the intersection process, the image data of the image shown in Figure 3B is intersected with the image data of the image shown in Figure 3C In particular, the target regions 3 arranged in the region of interest 2 are determined. Other target regions 3 arranged outside the region of interest 2 are removed.Figure 3D Shows a target area detected by a neural network, which is arranged in the region of interest 2 detected by another neural network.

[0075] Figure 4A Shows a training process for training a neural network to detect the target area 3. The training process T includes two training phases T1 and T2.

[0076] The first training phase T1 corresponds to general training, in which a plurality of training images are input into the neural network. The training images may include the target area 3, but this is not necessary. The training images input into the neural network in the first training phase T1 (especially all training images) are labeled. Additionally, the training images show a plurality of different objects and / or information about the positions of the objects. Thus, after the first training phase T1, the neural network knows how to identify a plurality of different types of objects. It is possible to run the first training phase T1 only once to train the neural network.

[0077] After the end of the first training phase, the second training phase T2 begins. The second training phase T2 is used to train the neural network for the application using the device 8. Thus, the training images input into the neural network include the target area 3. Additionally, training images that do not include the target area 3 are input into the neural network. Contrary to the first training phase, at least some or all of the training images are unlabeled.

[0078] After the neural network runs the second training phase T2, it is possible to identify whether the input image data contains one or more target areas 3.

[0079] Figure 4B Shows a training process for detecting the region of interest 2. Depending on whether the data processing device 5 includes only a neural network, or includes a neural network and another neural network, the training process runs for the neural network or the other neural network. The following description corresponds to both cases, that is, regardless of whether there is another neural network, the training process is the same.

[0080] The first training phase T1a corresponds to general training, in which a plurality of training images are input into the neural network or the other neural network. The training images may include the region of interest 2, but this is not necessary. The training images input into the neural network in the first training phase T1a (especially all training images) are labeled. Additionally, the training images show a plurality of different objects and / or information about the positions of the objects. Thus, after the first training phase T1a, the neural network or the other neural network knows how to identify a plurality of different kinds of objects. It is possible to run the first training phase T1a only once to train the neural network.

[0081] After the end of the first training phase T1a, the second training phase T2a begins. The second training phase T2a is used to train the neural network to the application of the device 8. Therefore, the training images input to the neural network or another neural network include the target area 3. Additionally, the training images input to the neural network or another neural network do not include the target area 3. Contrary to the first training phase, at least some or all of the training images are unlabeled.

[0082] After the neural network runs the second training phase T2a, it is possible to identify whether the input image data contains one or more regions of interest 2.

[0083] Figure 5 Shows the process for determining the region of interest 2. As described above, the neural network or another neural network detects one or more regions of interest 2. For this purpose, an object detection process is run. Figure 5 Shows the preliminary region of interest 2a determined by the neural network or another neural network. Additionally, Figure 5 Shows the original region of interest 2b presented in the image 1 acquired by the image acquisition device 6.

[0084] The data processing device 5 checks whether the edge of the preliminary region of interest 2a in the predetermined region 10 deviates from the edge of the original region of interest 2b shown in the image 1. If the edge of the preliminary region of interest 2a deviates, the data processing device 5 sets the corresponding part of the region of interest 2 to the edge part corresponding to the original region of interest. This process is carried out along the circumferential direction of the original region of interest 2b. Therefore, after this process is completed, the region of interest 2 for the intersection process is determined. The region of interest 2 is as Figure 5 Shown on the right and is used for the above intersection process.

[0085] Figure 6 Shows including Figure 1 or Figure 2 The conveying system 11 of the device 8 shown. The conveying system 11 includes a conveyor belt 8 on which objects are transported. The objects correspond to the regions of interest 2 discussed above. The conveyor belt 12 is arranged to pass through the monitoring area 13 of the image acquisition device 6. This means that all objects arranged on the conveyor belt 12 will pass through the monitoring area 13. The objects can correspond to the region of interest 2 or a part of the region of interest 2.

[0086] The image acquisition device 6 acquires an image 1 of the monitoring area 13 including the objects. Then, the above method is executed in the data processing device 5 to detect one or more target areas 3 in each region of interest 2.

[0087] Reference Signs

[0088] 1 Image

[0089] 2 Region of Interest

[0090] 2a Preliminary Region of Interest

[0091] 2b Original Region of Interest

[0092] 3 Target Object

[0093] 4 Region of Non - Interest

[0094] 5 Data Processing Device

[0095] 6 Image Acquisition Device

[0096] 8 Device

[0097] 9 Display

[0098] 10 Predetermined Region

[0099] 11 Conveyor System

[0100] 12 Conveyor Belt

[0101] 13 Monitoring Region

[0102] 60 Output Portion

[0103] 50 Input Portion

[0104] 51 Processing Portion

[0105] 52 Intersection Portion

[0106] 510 First Processing Portion

[0107] 511 Second Processing Portion

[0108] 90 Display Input Portion

[0109] T Training Process

[0110] T1 First Training Phase

[0111] T2 Second Training Phase.

Claims

1. A method for processing at least one image (1), wherein the method comprises the following steps: receiving image data, inputting the unmasked image data into an artificial neural network to detect one or more pre-trained target regions (3), detecting one or more regions of interest (2) in the image (1) and creating an image mask depending on the one or more regions of interest (2), and applying the created image mask to the image data to remove image regions (4) that do not include the one or more regions of interest (2), wherein the one or more pre-trained target regions (3) arranged in at least one, in particular pre-trained, region of interest (2) are detected.

2. The method according to claim 1, characterized in that the one or more pre-trained regions of interest (2) are detected before the image mask is applied to the received image data.

3. The method according to claim 1 or 2, characterized in that the pre-trained region of interest (2) is at least a part of at least one object or is at least one object.

4. The method according to at least one of claims 1 to 3, characterized in that a. the pre-trained target region (3) is optically different from the region of interest (2), and / or b. the pre-trained target region (3) corresponds to or is smaller than the region of interest (2).

5. The method according to at least one of claims 1 to 4, characterized in that a. the image mask is configured such that it removes one or more pre-trained target regions (3) arranged outside the pre-trained region of interest (2), and / or b. the created image mask is applied to the image data after the one or more pre-trained target regions (3) are detected.

6. The method according to at least one of claims 1 to 5, characterized in that a. the one or more target regions (3) arranged in the pre-trained region of interest (2) are visualized, and / or b. the pre-trained region of interest (2) and / or the removed image region (4) are not visualized.

7. The method according to at least one of claims 1 to 6, characterized in that a. the neural network is a convolutional neural network, and / or b. the neural network comprises at least two layers.

8. The method according to at least one of claims 1 to 7, characterized in that the neural network creates the image mask and / or determines the region of interest (2).

9. The method according to at least one of claims 1 to 7, characterized in that another neural network creates the image mask and / or determines the region of interest (2).

10. The method according to claim 9, characterized in that a. the another neural network is a convolutional neural network, and / or b. the another neural network comprises at least two layers.

11. The method according to at least one of claims 1 to 10, characterized in that The one or more pre-trained target regions (3) arranged in at least one of the regions of interest (2) are determined by the intersection of the detected one or more pre-trained target regions (3) and the detected one or more regions of interest (2).

12. The method according to claim 11, wherein, in order to determine the region of interest, the edge of a preliminary region of interest (2a) is determined, and it is determined whether the edge of the region of interest (2b) shown in the image is arranged in a predetermined region (10) including at least a part of the edge of the preliminary region of interest (2a).

13. The method according to claim 12, wherein, if the preliminary region of interest (2a) deviates from the edge of the region of interest (2b) in the predetermined region (10), the region of interest (2) is set to the edge of the region of interest (2b).

14. The method according to at least one of claims 1 to 13, wherein, a. determining whether the number of the pre-trained target regions (3) corresponds to a predetermined number of target regions, and / or b. determining the quality of the region of interest (2) based on the detected one or more pre-trained target regions (3).

15. The method according to at least one of claims 9 to 14, wherein, the training of the neural network corresponds to the training of the other neural network.

16. The method according to at least one of claims 1 to 15, wherein, the neural network is trained using unmasked training images, and / or the other neural network is trained using unmasked training images.

17. The method according to at least one of claims 1 to 16, wherein, the neural network and / or the other neural network is trained using training images, and the training images include background information.

18. The method according to at least one of claims 1 to 17, wherein, the training for the neural network and / or the other neural network includes two training stages.

19. The method according to claim 18, wherein, in the first training stage, training images including non-background information are input into the neural network and / or the other neural network.

20. The method according to claim 18 or 19, wherein, in the second training stage, training images including the target regions (3) and / or the regions of interest (2) are input into the neural network or the other neural network.

21. A data processing device (5) comprising means for performing the method according to at least one of claims 1 to 20.

22. An image acquisition device (6) for acquiring an image of an object, comprising the data processing device according to claim 21.

23. A computer program product comprising instructions which, when executed by a data processing device (5), in particular a computer, cause the data processing device (5), in particular the computer, to perform the method according to at least one of claims 1 to 20.

24. A computer-readable medium having stored thereon the computer program product according to claim 23.

25. A data carrier signal carrying the computer program product according to claim 23.

26. Use of the method according to at least one of claims 1 to 21 in a conveying system, in particular for determining the quality of goods to be transported.