Training method, method and / or device for segmenting and / or classifying an object on a substrate

By preprocessing multi-channel image information from image sensors, the method enhances the accuracy of object recognition in intelligent irrigation and spraying systems, addressing the challenge of subsurface object classification and improving application efficiency.

DE102023212342A1Pending Publication Date: 2025-06-12ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102023212342
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Current intelligent irrigation and spraying systems face challenges in accurately segmenting and classifying objects on a subsurface due to variations in soil and foreign bodies, leading to inefficient herbicide application and irrigation.

Method used

A training method for a machine learning model that modifies multi-channel image information from image sensors, such as those with Bayer sensors, by preprocessing steps like calculating new image channels and replacing unused channels with preprocessed data, to enhance object recognition and classification.

Benefits of technology

The method improves the quality of object recognition on agricultural surfaces, allowing for more accurate segmentation and classification of plants and weeds, leading to optimized herbicide application and water usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for segmenting and / or classifying an object on a substrate, the method comprising the following steps: - Providing (S21) a trained machine learning model; - Providing (S22) at least one image datum comprising multi-channel, in particular three-channel, image information; - modifying (S23) at least one of the multi-channel image information by an image information preprocessing step; - providing (S24) such a modified image datum; and - Segmenting and / or classifying (S25) the object on the ground by the trained machine learning model based on the modified image data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a training method, a method and / or a device for segmenting and / or classifying an object on a substrate. State of the art

[0002] An intelligent spraying system for applying herbicides is generally known. A camera system detects plants. In a second step, the image data is used to distinguish between crops and weeds. Wherever the camera system detects weeds, appropriate nozzles are activated, thus treating or spraying the weeds with a herbicide.

[0003] Furthermore, an intelligent irrigation system is used to detect plants or plant objects on agricultural land or in a field. However, achieving the most accurate separation or segmentation of the plant objects from the respective field substrate is often challenging.

[0004] Foreign objects such as stones, straw, or even the soil composition can influence the quality and accuracy of detection. This is due to the countless variations of different soils and foreign objects on the field subsoil. Therefore, not all cases can be anticipated and intercepted.

[0005] This makes it difficult to ensure the most accurate detection possible.

[0006] Current image sensors often use so-called Bayer sensors to detect plants on agricultural land. These sensors operate in the RGB (red-green-blue) color space and are designed with color filters with, for example, 25% red, 50% green, and 25% blue. The respective color filters allow red, green, and blue color information to be recorded. A Bayer sensor of this type has twice the number of green pixels as red and blue pixels. The arrangement of the color pixels can fundamentally vary depending on the Bayer sensor. However, the green pixels, or the pixels imaged onto the image plane or photo chip by the color filter, are always diagonally opposite one another in the green color spectrum.

[0007] In a process known as demosaicing, a full-resolution and often interpolated image is calculated for each color channel, resulting in an image in the RGB color space. Demosaicing, also known as debayering or Bayer demosaicing, is a process in digital image processing used to reconstruct a color image from the raw data of an image sensor. Digital cameras often use image sensors with a so-called Bayer pattern, which consists of a matrix of pixels alternately covered with red, green, and blue color filters. This pattern allows for the capture of brightness information, but not the precise color information for each pixel.

[0008] Demosaicing is the step of reconstructing the color information for each pixel using neighboring pixel values ​​and color information. This is done by using the values ​​of neighboring pixels in the pattern to estimate the missing color information. This process results in a complete color image derived from the original monochrome image sensor raw data.

[0009] The result of demosaicing is an image that typically contains all three color channels (red, green, and blue) for each pixel, creating a full-color image suitable for human viewing. This process is critical for creating color digital images and occurs in most digital cameras and image processing applications.

[0010] Since, as already mentioned, most cameras are constructed in this way and use such filters placed in front of the pixels, it is a standard in the field of machine learning that the algorithms used process three-channel, i.e., three-dimensional, images (one channel for red, green, blue) in order to extract feature maps from the image.

[0011] Returning to smart irrigation systems, a conventional RGB image sensor is often used, but it is equipped with a bandpass filter placed in front of the image sensor, allowing only red and NIR (near infrared) light to pass through to the image sensor. The remaining wavelengths are blocked.

[0012] However, due to this camera design, an identical NIR signal is stored simultaneously in the green and blue channels. The information from the NIR signal is thus duplicated and remains unused due to unnecessary redundancy.

[0013] In common smart irrigation systems and associated monitoring systems for plant object recognition, machine learning eliminates the already unused blue channel. Due to its higher resolution (resulting from twice the number of pixels), only the green channel is used for processing the NIR signal. The blue channel is therefore no longer used in the downstream pipeline.

[0014] DE 10 2020 215 413 A1 discloses a method for generating an enhanced camera image using an image sensor. The image sensor has at least one red pixel with an R-color filter element for transmitting red light and infrared light. Further prior art can be found in the documents DE 10 2016 112 968 B4 and DE 10 2021 111 639 A1.

[0015] The invention is based on the object of providing an improved method and / or an improved device for segmenting and / or classifying an object on a substrate.

[0016] The problem is solved by a training method according to the features of patent claim 1. The problem is solved by a method for segmenting and / or classifying an object on a background according to the features of patent claim 2. The problem is solved by a device for segmenting and / or classifying an object on a background according to the features of patent claim 10. Disclosure of the invention

[0017] According to a first aspect, a training method for training a machine learning model for segmenting and / or classifying an object on a surface is provided. The training method comprises: - Providing a modified training data set comprising a plurality of image data, each with multi-channel, in particular three-channel, image information, wherein at least one of the multi-channel image information is modified by at least one image information preprocessing step; - Training the machine learning model based on the modified training data set for segmenting and / or classifying an object on a surface; and - Deploy the trained machine learning model.

[0018] According to a second aspect, an (inference) method for segmenting and / or classifying an object on a background is provided. The method comprises the following steps: - Providing a trained machine learning model; - Providing at least one image data item comprising multi-channel, in particular three-channel, image information; - modifying at least one of the multi-channel image information by an image information preprocessing step; - Providing such modified image data; and - Segmenting and / or classifying the object on the ground by the trained machine learning model based on the modified image data.

[0019] The method is preferably used in a spraying system for discharging an active agent, in particular a plant protection product or a fertilizer, or in an agricultural field sprayer. The method can also be used in an irrigation system or in a fertilizing machine or another type of machine intended to specifically detect plants and / or a subsoil. The present invention thus further relates to the application of the method in an agricultural work machine, in particular a field sprayer, and / or in an agricultural work tool, in particular a spraying device, and / or in a plant identification unit.The working tool, in particular the spraying device, is preferably configured by implementing the method to apply an active agent, in particular a plant protection agent or fertilizer, depending on the weeds identified by the method in the identified plant row. The method is preferably carried out by means of a computing unit that is also used to control the field sprayer or via which the field sprayer is controlled.

[0020] The present method can also be used to replace a number n of raw image channels, for example, n=3 (R, G, B), with, for example, nx, e.g., a single, image channel, where the nx image channels serve as input for the machine learning model. For example, a segmented plant mask (binary information), an NDVI image, and / or another type of preprocessed image channel can be used and / or replaced.

[0021] The statements made for the training procedure apply accordingly to the (inference) procedure. It is understood that linguistic modifications of procedurally formulated features can be reformulated for the system according to common linguistic practice, without such formulations having to be explicitly listed here.

[0022] It is understood that the present steps and other optional steps do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more sub-steps without thereby departing from the scope of the training method according to the invention.

[0023] The "provision of a modified training data set comprising a plurality of image data, each with multi-channel, in particular three-channel, image information, wherein at least one of the multi-channel image information items is modified by at least one image information preprocessing step" means that the machine learning model is not provided with a training data set consisting of raw image data as originally acquired by an imaging sensor, but rather at least a portion of the image information acquired by the at least one imaging sensor is preprocessed before provision. Although the raw image data may already have been interpolated before the at least one image information preprocessing step, it is subsequently altered or modified by mathematical calculations or other processing operators.

[0024] At least one additional image channel is therefore calculated from the given image information provided by the at least one imaging sensor. For example, the imaging sensor may contain image information of red values ​​in a red signal, image information of green values ​​in a green signal, image information of blue values ​​in a blue signal, and / or image information of NIR values ​​in an NIR signal. Depending on the design of the imaging sensor, the NIR signal may replace the image information of the green signal in the relevant channel. The modified signal then replaces the image information in at least one channel, so that this channel no longer provides the raw image information or, if applicable, interpolated image information, but instead provides the modified image information.

[0025] This preprocessing preferably takes place before a respective image data is provided to the machine learning model for a learning phase and / or optimization phase, and / or before an application / inference of the machine learning model is started, in which the trained ML model makes a prediction (object detection) of a new image based on the input image.

[0026] The idea is that an unused and / or duplicate image channel (e.g., NIR signal duplicated in the green and blue channels) is replaced by a new calculation from existing image information, thus offering the ML model, for example, manual, user-considered support and / or a hint as to which areas and / or types of features are important in an image to be analyzed and / or when applying the ML model (e.g., an NDVI value or a pure red channel for plant recognition). Using the specially adapted image channels defined by the user, the ML model uses these "hints" in a given input image with the remaining layers and nodes to learn more precisely and more precisely the relevant features.

[0027] Another example could be a computed derivative image, where, for example, edges (sharp and / or fast transitions of structures in the image) or corner points might be of interest in the application. The use of exact filters in preprocessing cannot preferably be specified, as this depends on the specific task and / or application of the ML model.

[0028] It is understood that, in this case, one or more image channels can be swapped out and replaced with one or more modified and / or filtered image channels, as is the case with images from standard RGB cameras, for example. Replacing or modifying the image channels can thus affect only one image channel, two, or more than two image channels of a multi-channel image (an image with multi-channel image information).

[0029] The present method and the trained ML learning model offer significant advantages, especially when using models of an artificial neural network, having a flat network architecture with a few layers and nodes. Such flat and simple neural networks are used when high computational runtime requirements and limited computing resources or hardware size and cost are involved. Smaller and flatter neural networks are limited in the number and / or variations of filters, operators, intermediate calculations, and thus in what can be learned. Thus, it can be advantageous for such network architectures to calculate suitable, targeted, and / or useful image channels in advance from “classic” image processing operations and / or filters. The ML model or the neural network can thus be trained with an extended and / or modified RGB image, and in the inference based on such preprocessed ormodified RGB images to make a classification statement and / or a segmentation statement and / or more generally a prediction.

[0030] In this case, the object recognition of various types of objects, such as plants, is improved in terms of quality (true positive). This enables a generally better recognition of objects against a background. It can also improve the classification of objects into different classes, for example, plants, into different crop and / or weed classes.

[0031] According to one embodiment, the multi-channel image information is captured by an image sensor in the form of red-green-blue image information. Alternatively, the multi-channel image information is captured by an image sensor equipped with a bandpass filter in the form of red image information and NIR image information. Other types of image information are also conceivable, as long as they can be captured by an imaging sensor.

[0032] According to one embodiment, the image sensor comprises a Bayer sensor and / or a multispectral camera and / or a hyperspectral camera, in particular an RGB-NIR camera. Other types of imaging sensors are also conceivable, so the list is not intended to be limiting.

[0033] According to one embodiment, the at least one image information preprocessing step comprises calculating at least one new image information from acquired image information and replacing and / or modifying at least one image channel for the modified image information.

[0034] According to one embodiment, the calculation comprises at least one of the following calculation steps: a difference, an addition, a subtraction, a division, a multiplication, a power, an inversion, or any arithmetic concatenation of the aforementioned calculation steps. The calculation of new image channels from the existing image information can thus consist of any calculation steps, such as the difference, addition, subtraction,

[0035] Division, multiplication, exponentiation, inversion, or any combination of these calculation steps. The calculation of new image channels in other applications can be performed using any function.

[0036] The image preprocessing to provide modified image data for the machine learning model can involve any arithmetic calculation such as addition, subtraction, multiplication, division, exponentiation, inversion, etc., or it can be performed based solely on one of the aforementioned calculation steps. Other calculation types not mentioned here are also conceivable.

[0037] According to one embodiment, the at least one image information preprocessing step comprises at least one transformation of at least one originally acquired image information and / or a brightness and / or contrast adjustment and / or a convolution or deconvolution of at least one originally acquired image information and / or a segmentation and / or a filtering based on at least one threshold value and / or an application of at least one morphological operator.

[0038] Such transformations can include, for example, rotation, distortion, scaling, reflection, and / or a Hough transform, or similar. Convolution or deconvolution can, for example, consist of smoothing, sharpening, the formation of a first and / or second derivative or nth derivative, a gradient image, and / or a combination of filters. Such filters can be, for example, LoG (Laplacian of Gaussian) filters. Segmentation can include binary or band segmentation, in particular with a "brightness" limit. Thresholding or multi-thresholding can be used. Morphological operators can include erosion, dilation, opening, closing, watersheds, and / or segmentation.

[0039] According to one embodiment, the image data is captured by an RGB-NIR image sensor and the three-channel image information comprises at least one red image information and one NIR image information, wherein the NIR image information is / is preferably mapped onto two image channels, wherein the at least one image information preprocessing step comprises calculating an NDVI value which is made available to the machine learning model as modified image information in at least one of the three image channels.

[0040] The NDVI (Normalized Difference Vegetation Index) is primarily used to assess vegetation activity and condition from satellite- or aircraft-based remote sensing data. It is used in agriculture, environmental science, and other disciplines that aim to monitor plant growth or vegetation density. The NDVI is calculated by taking the difference between the near-infrared (NIR) and red light (RED) of an image and dividing them by their sum. NDVI values ​​range from -1 to 1: values ​​close to 1 indicate dense vegetation. Values ​​around 0 indicate bare or sparsely vegetated areas. Negative values ​​can indicate bodies of water.

[0041] An NDVI threshold preferably refers to a specified NDVI value or range of values ​​used to classify or identify specific vegetation states or types. By setting thresholds, certain land cover types, such as dense vegetation, bare patches, or water bodies, can be filtered out and / or highlighted from NDVI data. An example of the use of NDVI thresholds is: An NDVI value of > 0.6 could be interpreted as dense vegetation. A value between 0.2 and 0.6 could be considered sparse or mixed vegetation. A value < 0.2 could be considered bare land or an urban area. It is important to emphasize that such thresholds are not universal and can be adapted depending on the application, region, and specific study objectives.It is preferable to set thresholds based on local knowledge and / or by combining NDVI data with other data sources.

[0042] In a further step when using the NDVI or other calculated image channels, a suitable and favorable choice of a signal level threshold (segmentation) can be used, with the prior knowledge that, for example, object pixels (e.g., plant pixels) have a high NDVI signal level and a background pixel has a low NDVI signal level. Thus, a certain modification of the image channels can be made in advance, for example, to suppress individual signal levels of pixels (or remove them entirely) while maintaining the signal level of others (or amplifying them).

[0043] Based on the knowledge of the signal information and the application of the ML model to identify plants in order to intelligently irrigate them and / or spray or treat them with an active agent, in particular a pesticide or fertilizer, the following calculations from the original RGB image information (red channel occupied by red+NIR image information, green channel occupied by NIR image information, blue channel occupied by NI R image information) are advantageous:

[0044] Calculation of a pure red signal, since the red channel has a mixed signal of red + NIR signal: Red Signalcorrected=Red Channel−Green Channel=(Red and NIR Signal)−(NIR Signal) Calculation of an NDVI value: A popular method in the state of the art to distinguish plants from a (field) background (soil). NDVI=(Green Channel−Red Channelcorrected)(Green Channel+Red Channelcorrected)=(NIR Signal−Red Signalcorrected)(NIR Signal+Red Signalcorrected) with Red Signalcorrected from above equation

[0045] The NDVI value is preferably normalized to a value range between [-1...1]. Therefore, the intensity values ​​in the image are preferably transformed to a value range for image data, e.g., the usual 8 bits from [0...255], for further calculation.

[0046] It is also possible that one of the several or more of the several image channels is replaced by the modified image information.

[0047] Possible examples instead of R(Red+NIR), G(NIR), B(NIR) in the RGB image channels are: 1. R(Red+NIR), G(NIR), B(NDVI) 2. R(Red+NIR), G(NIR), B(R_corrected) 3. R(R_corrected), G(NIR), B(NDVI)

[0048] It is also possible to swap the image channels (G <-> B) and / or (R <-> B) and / or (G <-> B), since the ML model is able to treat the individual color channels with different sensitivity. 4. R(Red+NIR), G(NDVI), B(NIR) 5. R(Red+NIR), G(NDVI), B(R_corrected) 6. R(R_corrected), G(NDVI), B(NIR)

[0049] According to one embodiment, the machine learning model comprises a neural network or another type of artificial network.

[0050] According to a third aspect, a device for segmenting and / or classifying an object on a substrate is provided. The device comprises at least one computing and / or evaluation device configured to perform the following steps: - Providing a trained machine learning model; - Providing at least one image data item comprising multi-channel, in particular three-channel, image information; - modifying at least one of the multi-channel image information by an image information preprocessing step; - Providing such modified image data; and - Segmenting and / or classifying the object on the ground by the trained machine learning model based on the modified image data.

[0051] The statements made for the method apply accordingly to the device. It is understood that linguistic modifications of procedurally formulated features can be reformulated for the system according to common linguistic practice, without such formulations having to be explicitly listed here.

[0052] The present invention also claims an intelligent irrigation system and / or an intelligent spraying system for applying an active agent, in particular a plant protection agent or a fertilizer, using a device, wherein the device is designed for object detection of plants on a field substrate. The device is preferably used for object detection of plants in an intelligent irrigation system in order to optimize object recognition and, consequently, the irrigation result. In particular, the optimized object recognition can save water, a valuable resource.

[0053] The present invention also claims a control device and / or a camera system and / or a monitoring system which may be included in an intelligent irrigation system and / or in an autonomous vehicle and / or a robotic system and / or an industrial machine, and on which a machine learning model trained according to the present method in one of its embodiments is executable.

[0054] An alternative approach to modifying an image channel of an image sensor is to expand the number of image channels, for example using 3+n-channel images instead of three. However, due to the standardization with three image channels, this appears rather complex. For example, freely available and pre-trained ML models and their architectures are already adapted and / or optimized for three-channel images. Breaking this situation would therefore be costly, as it would then be necessary to develop a completely separate ML model for 3+n image channel processing. In addition, this would mean that the more input data flows into an ML model, the longer the evaluation takes, and thus the requirements for short computing runtime might not be met. Alternatively, it is also possible to use two- or single-channel image data. For example, only one NDVI image.This allows the machine learning model to run faster.

[0055] The present invention also claims a computer program with program code for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.

[0056] The present invention also proposes a computer-readable data carrier containing program code of a computer program for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.

[0057] The described designs and further training courses can be combined as desired.

[0058] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the exemplary embodiments that are not explicitly mentioned. Short description of the drawings

[0059] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.

[0060] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.

[0061] They show: Fig. 1 a schematic flow diagram of a training procedure; Fig. 2 a schematic flow diagram of a method for segmenting and / or classifying an object on a substrate; Fig. 3 a schematic representation of a pixel distribution in a Bayer sensor; Fig. 4 a schematic representation of a flow chart for image preprocessing and for providing an input image data and / or a training image data; and Fig. 5 example images without data preprocessing and with two modifications of data preprocessing.

[0062] In the figures of the drawings, the same reference symbols designate the same or functionally identical elements, parts or components, unless otherwise stated.

[0063] Fig. 1 shows a schematic flow diagram of a training method for training a machine learning model for segmenting and / or classifying an object on a surface.

[0064] In any embodiment, the method can be carried out at least partially by a device 100, which for this purpose can comprise several components not shown in detail, for example one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the device 100 can comprise a storage device and / or an output device and / or a display device and / or an input device.

[0065] According to the invention, the computer-implemented training method comprises at least the following steps: In a step S1, a modified training data set is provided which comprises a plurality of image data, each with multi-channel, in particular three-channel, image information, wherein at least one of the multi-channel image information is modified by at least one image information preprocessing step. In a step S2, the machine learning model is trained on the basis of the modified training data set for segmenting and / or classifying an object on a surface. In step S3, the trained machine learning model is provided.

[0066] Furthermore, a schematic flow diagram of a method for segmenting and / or classifying an object on a background in Fig. 2. The method can also be carried out by the device 100.

[0067] The procedure comprises at least the following steps: In a step S21, a trained machine learning model is provided. In a step S22, at least one image data item is provided which has multi-channel, in particular three-channel, image information. In a step S23, at least one of the multi-channel image information is modified by an image information preprocessing step. In a step S24, such a modified image data is provided, ie a modified image data is provided which is based on the provided image data and the at least one modified multi-channel image information. In a step S25, the object on the background is segmented and / or classified by the trained machine learning model based on the modified image data.

[0068] Fig. Figure 3 shows a schematic representation of a pixel distribution 300 in a Bayer sensor. It can be seen that the Bayer sensor has twice as many green pixels G as blue pixels B and red pixels R. Furthermore, the green pixels G are arranged diagonally one above the other. Various pixel configurations are conceivable and depend on the type and design of the respective Bayer sensor.

[0069] Fig. 4 shows a schematic representation of a flowchart for image preprocessing and for providing input image data and / or training image data. The pixel distribution 300 shows numbers that indicate the indices of the pixels of an image matrix. The pixel distribution 300, in turn, has twice as much pixel information in the green channel of the Bayer sensor as red and blue pixel information. In this case, the Bayer sensor has a bandpass filter, so that only red color information and NIR color information are captured. The NIR color information is displayed together with the red color information in the R channel, see reference numeral 402. The NIR image information would generally be displayed twice, i.e., in both the G channel and the B channel.However, due to the plurality of pixels, the G channel has a higher resolution than the B channel, so that the NIR image information is only provided in the G channel, see reference numeral 404. The B channel 406 is therefore basically unoccupied and is available for modifying the image information by at least one image information preprocessing step, see reference numerals 408, 410. The respective red image or pixel information and NIR pixel information can each be interpolated, which is identified by the reference numeral 400.

[0070] In the image information preprocessing step 408, the following calculation is preferably carried out Red Signalcorrected=Red Channel−Green Channel=(Red and NIR)−(NIR Signal)

[0071] Corrected red image information is calculated by subtracting the image information of the R channel (red and NIR signal), specifically pixel by pixel, from the NIR image information (NIR signal). Consequently, the red image information is isolated from the NIR image information.

[0072] The isolated red image information can be used as input for the image information preprocessing step 410. In step 410, NDVI values ​​are calculated, particularly pixel by pixel. The calculation is based on the following equation: NDVI=(Green Channel−Red Channelcorrected)(Green Channel+Red Channelcorrected)=(NIR Signal−Red Signalcorrected)(NIR Signal+Red Signalcorrected)

[0073] The red image information and the NDVI image information can be provided as modified image information to the ML model for training purposes or in inference.

[0074] Fig.Figure 5 shows example images without data preprocessing and with two modifications of data preprocessing. A Bayer sensor with a bandpass filter also captures only red and NIR image information.

[0075] The top row (note the rotation of the figure) shows original image data 500 from the Bayer sensor. An input image is shown as a multi-channel composition (top left). Furthermore, the three RGB channels are shown and identified by the reference numerals 502, 504, and 506. The R channel 502 of the original image shows the red image signal together with the NIR image signal. The G channel 504 shows the NIR signal. The B channel 506 shows the NIR signal (i.e., it is present twice).

[0076] The middle row (note the rotation of the figure) shows modified image data 508 from the Bayer sensor. An input image is shown as a multi-channel composition (center left). Furthermore, the three RGB channels are shown and identified by the reference numerals 502, 504, and 506. The R channel 502 of the modified image shows the red image signal together with the NIR image signal. The G channel 504 shows the NIR signal. The B channel 506 shows an NDVI signal calculated during preprocessing.

[0077] The bottom row (note the rotation of the figure) shows modified image data 510 from the Bayer sensor. An input image is shown as a multi-channel composition (bottom left). Furthermore, the three RGB channels are shown and identified by the reference numerals 502, 504, and 506. The red image signal is shown in R channel 502 of the modified image, along with the NIR image signal. The NIR signal is shown in G channel 504. The corrected red image signal is shown in B channel 506.

[0078] The general use of NDVI offers advantages over the standard RGB image from a Bayer camera. An evaluation conducted with two identically trained neural networks (one with an RGB image, the other with RG+NDVI images as input to the network) and an evaluation on four corn image datasets, each of which the network had never seen before, showed that the overall FP (false positive) and FN (false negative) rates were significantly improved when NDVI values ​​were used as modified image information in one of the image channels. QUOTES CONTAINED IN THE DESCRIPTION

[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature

[0000] DE 10 2020 215 413 A1

[0014] DE 10 2016 112 968 B4

[0014] DE 10 2021 111 639 A1

[0014]

Claims

[1] Training method for training a machine learning model for segmenting and / or classifying an object on a surface; the training method comprising: - Providing (S1) a modified training data set comprising a plurality of image data, each with multi-channel, in particular three-channel, image information, wherein at least one of the multi-channel image information is modified by at least one image information preprocessing step; - Training (S2) the machine learning model based on the modified training data set for segmenting and / or classifying an object on a surface; and - Providing (S3) the trained machine learning model. [2] A method for segmenting and / or classifying an object on a substrate, the method comprising the following steps: - providing (S21) a machine learning model trained according to claim 1; - Providing (S22) at least one image datum comprising multi-channel, in particular three-channel, image information; - modifying (S23) at least one of the multi-channel image information by an image information preprocessing step; - providing (S24) such a modified image datum; and - Segmenting and / or classifying (S25) the object on the ground by the trained machine learning model based on the modified image data. [3] Training method and / or method according to claim 1 or 2, wherein the multi-channel image information is acquired by an image sensor in the form of red-green-blue image information; or wherein the multi-channel image information is acquired by an image sensor equipped with a bandpass filter in the form of red image information and NIR image information. [4] Training method and / or method according to claim 3, wherein the image sensor comprises a Bayer sensor and / or a multispectral camera and / or a hyperspectral camera, in particular an RGB-NIR camera. [5] Training method and / or method according to one of the preceding claims, wherein the at least one image information preprocessing step comprises calculating at least one new image information from acquired image information and replacing and / or modifying at least one image channel for the modified image information. [6] Training method and / or method according to claim 5, wherein the calculation comprises at least one of the following calculation steps: a difference, an addition, a subtraction, a division, a multiplication, an exponentiation, an inversion or any arithmetic concatenation of the aforementioned calculation steps. [7] Training method and / or method according to claim 5 or 6, wherein the at least one image information preprocessing step comprises at least one transformation of at least one originally acquired image information and / or a brightness and / or contrast adjustment and / or a convolution or deconvolution of at least one originally acquired image information and / or a segmentation and / or a filtering based on at least one threshold value and / or an application of at least one morphological operator. [8] Training method and / or method according to one of the preceding claims, wherein the image data is acquired by an RGB-NIR image sensor and the three-channel image information comprises at least one red image information and one NIR image information, wherein the NIR image information is preferably mapped to two image channels, wherein the at least one image information preprocessing step comprises calculating an NDVI value which is made available to the machine learning model as modified image information in at least one of the three image channels. [9] Training method and / or method according to one of the preceding claims, wherein the machine learning model comprises a neural network or another type of artificial network. [10] Device (100) for segmenting and / or classifying an object on a substrate; the device comprising at least one computing and / or evaluation device which is designed to carry out the following steps: - providing (S21) a machine learning model trained according to claim 1; - Providing (S22) at least one image datum comprising multi-channel, in particular three-channel, image information; - modifying (S23) at least one of the multi-channel image information by an image information preprocessing step; - providing (S24) such a modified image datum; and - Segmenting and / or classifying (S25) the object on the ground by the trained machine learning model based on the modified image data. [11] Intelligent irrigation system and / or intelligent spraying device for applying an active agent, in particular a plant protection agent or a fertilizer, with a device (100) according to claim 10, wherein the device is designed for object detection of plants on a field substrate. [12] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer. [13] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer.

Citation Information

Patent Citations

  • Method and apparatus for detecting a nuisance plant image in a camera raw image and image processing device

    DE102020215415A1

  • Method for providing a training data set quantity, method for training a classifier, method for controlling a vehicle, computer-readable storage medium and vehicle

    WO2020164841A1