Device and method for training and testing a classifier

The method leverages CycleGANs to automatically generate labeled data for classifiers, addressing the challenge of manual data labeling in critical environments, enhancing classification performance through diverse dataset creation.

JP7709857B2Active Publication Date: 2025-07-17ROBERT BOSCH GMBH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021098058
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-12
Filing Date
2021-06-11
Publication Date
2025-07-17
Estimated Expiration
2041-06-11

AI Technical Summary

Technical Problem

Supervised training of classifiers requires a large amount of labeled data, which is cumbersome and time-consuming to obtain manually, especially in critical environments like self-driving vehicles, necessitating a method for automatic and reliable labeling of signals for training and testing.

Method used

A computer-implemented method using CycleGANs to automatically generate labeled data by converting signals between classes, enabling the creation of diverse training and test datasets without human intervention, and employing CycleGANs to produce output signals and masks for classification tasks.

Benefits of technology

Enables efficient and accurate training and testing of classifiers by generating large, diverse datasets automatically, improving classification performance without human teachers, particularly in safety-critical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709857000004
    Figure 0007709857000004
  • Figure 0007709857000005
    Figure 0007709857000005
  • Figure 0007709857000006
    Figure 0007709857000006
Patent Text Reader

Abstract

To provide a method for training a classifier, a method for evaluating whether the classifier can be used in a control system, a method for operating an actuator, a computer program and a machine-readable storage medium, the classifier, the control system and a training system.SOLUTION: A training system 140 comprises the steps of: preparing a generator which provides a mask signal; providing an output signal or mask signal by the generator on the basis of an input signal; providing a difference signal corresponding to the input signal on the basis of a difference between the input signal and the output signal or the mask signal; providing the desired output signal characterizing the desired classification of the input signal corresponding to the input signal on the basis of the corresponding difference signal; and providing at least one input image and the corresponding desired output signal as a training dataset.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for training a classifier, a method for evaluating whether a classifier can be used in a control system, a method for operating an actuator, a computer program and a machine-readable storage medium, a classifier, a control system, and a training system.

Background Art

[0002] Background Art “Jun-Yan Zhu, Taesung Park, Phillip Isola and Alexei A. Efros, “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks”, https: / / arxiv.org / abs / 1703.10593v1” discloses a method for converting between unpaired images using a cycle-consistent adversarial network.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Effects of the Invention Supervised training of a classifier requires a huge amount of labeled data. Labeled data can be understood as a plurality of data points to which corresponding labels are assigned, and when each data point is provided, the classifier predicts the label. For example, a signal representing an urban traffic scene can be assigned "urban traffic" as a class label. Furthermore, it is also assumed that vehicles depicted in the scene are labeled by a bounding box predicted by the classifier.

[0005] In particular, when using a classifier in an environment where safety is important, such as in a highly automated vehicle or a self-driving vehicle, it is important to train the classifier using a large amount of labeled data. This is generally because the performance of the classifier improves when more (and diverse) data is presented during training. This reduces the risk of misclassifying the environment of the classifier during operation, and then, since the safety of the product using the classifier (e.g., a self-driving vehicle) is improved, it is very important to achieve the highest possible performance.

[0006] Similarly, to evaluate the performance of a classifier, a large amount of labeled data is required as a test data set. The true performance of the classifier with respect to unknown data during training, i.e., the data that the classifier will face during operation, is approximated by the test data set. Therefore, the larger and more diverse the test data set is, the better the approximation to the true performance of the classifier during operation.

[0007] Manually obtaining labeled data is a cumbersome and time-consuming task. Therefore, a method that enables automatic labeling of reliable signals and can be used for training or testing of a classifier is desired.

Means for Solving the Problems

[0008] The advantage of the method having the features of independent claim 1 is that the labels of signals can be obtained in an automatic and teacherless manner, i.e., without the need for a human teacher. Also, with this method, a very diverse and large dataset can be generated, and the dataset can be used as a training dataset for enhancing the performance of a classifier or as a test dataset for more reliably approximating the performance of a classifier.

[0009] Disclosure of the Invention In a first aspect, the present invention relates to a computer-implemented method for training a classifier, the classifier being configured to provide an output signal characterizing the classification of an input signal, and training the classifier comprises, in a method, based on a provided training dataset, Providing the training dataset comprises · preparing a first generator, the first generator being configured to provide an output signal having the features of a second class based on a provided input signal having the features of a first class, or the generator being configured to provide a mask signal indicating which portion of the input signal represents the features of the first class, · providing, based on the input signal, an output signal or a mask signal by the first generator, · providing a difference signal corresponding to the input signal, the difference signal being provided based on the difference between the input signal and the output signal, or the mask signal being provided as the difference signal, · providing a desired output signal corresponding to the input signal based on the corresponding difference signal, the desired output signal characterizing the desired classification of the input signal, · providing at least one input signal and the corresponding desired output signal as a training dataset, and including.

[0010] The classifier can obtain an output signal by supplying a signal to a machine learning model, particularly a convolutional neural network. In this case, the machine learning model can provide an intermediate output signal, and the intermediate output signal can be provided as the output signal of the classifier. Further, the classifier can adjust the input signal, for example, by extracting features from the input signal, before supplying the input signal to the machine learning model. Further, the classifier can also post-process the intermediate output signal and then provide this as the output signal.

[0011] One or more class labels can be assigned to the signal based on the classification characterized by the output signal. Alternatively or additionally, it can be assumed that the classification is characterized in the form of object detection by the output signal. Alternatively or additionally, it can be assumed that the semantic segmentation of the signal is characterized by the output signal.

[0012] The classifier can accept signals of various modalities, particularly images such as, for example, video images, RADAR images, LiDAR images and / or ultrasonic images, and infrared camera images. Further, it can be assumed that the signal includes data from a combination of multiple and / or different sensor modalities, for example, video, RADAR and LiDAR images.

[0013] Alternatively or additionally, the signal can include audio data, for example, raw data or pre-processed data from a microphone. For example, the signal may include MFCC features of a series of audio recordings.

[0014] Based on the content included in the signal, at least one class can be assigned to the signal from a plurality of classes. In this sense, a signal having the characteristics of a certain class can be understood as a signal including characteristics important for that class. For example, when an image includes a red octagonal symbol with the word "STOP" attached, it can be understood that the image includes the characteristics of the "STOP sign" class. In other examples, the signal may be an audio sequence including the speech of a first speaker, and a part of the sequence including the speech can be regarded as the characteristics of the class "first speaker".

[0015] The signal can be organized in a predetermined format. For example, when the signal is a grayscale camera image, etc., it may be organized as a matrix, or, for example, when the image is an RGB camera image, it may also be organized as a tensor. Therefore, the signal can be understood as having a size that characterizes the structure of the format in which the signal is organized. For example, an image signal can have a width and a height. Further, the size of the image signal may include a specific depth indicating the length of each pixel. That is, each pixel can be a numerical value or a vector of a predetermined length. For example, the depth of an RGB image is 3, that is, one channel each for red, green, and blue, and each pixel is a three-dimensional vector. In other examples, a grayscale image includes a scalar as a pixel.

[0016] An audio signal can also be organized as a matrix or a tensor. For example, a series of MFCC features can be organized as a matrix. In this case, the rows of the matrix can include MFCC features. Here, each row can represent a predetermined time point of the recording of the audio signal.

[0017] By organizing the signal in a predetermined format, specific operations on the signal can be calculated. For example, when the signal is configured as a matrix or a tensor, obtaining the difference between two signals can be realized by matrix subtraction or tensor subtraction respectively.

[0018] A signal may include individual sub - elements. For example, in the case of an image signal, a pixel is regarded as a sub - element. In the case of an audio signal, individual time steps can be regarded as sub - elements. A sub - element can be in vector form, such as a pixel of an RGB image, or a feature of an audio sequence such as MFCC features. A sub - element can have an individual position. For example, a sub - element has specific row and column positions within an image, and the features of an audio sequence have individual positions within the sequence.

[0019] The positions of two sub - elements from different signals may be considered to be equal positions. Thus, it may be possible to refer to the first sub - element of the first signal using the position of the sub - element of the second signal. For example, considering two 100x100 - pixel single - channel images where there are pixels at the position (10,12) in both signals, the number 10 indicates the row index of the pixel and the number 12 indicates the column index of the pixel.

[0020] A generator can be regarded as a device that receives an input signal and provides an output signal. The generator is preferably part of an adversarial generation network.

[0021] When obtaining a difference signal from the input signal and the output signal, the difference signal may be obtained by subtracting the input signal from the output signal, or by subtracting the output signal from the input signal.

[0022] In another aspect of the present invention, the step of preparing the first generator (G) includes training a CycleGAN including a second generator and a third generator, and training the CycleGAN includes training the second generator to provide an output signal having the characteristics of the second class based on a provided input signal having the characteristics of the first class, training the third generator to provide an input signal having the characteristics of the first class based on a provided input signal having the characteristics of the second class, and preparing the second generator as the first generator (G) after training the CycleGAN. It may further be envisioned to include.

[0023] An adversarial generation network with cycle consistency (CycleGAN) is a machine learning model that can convert a signal having the characteristics of a first class into the characteristics of a second class. Therefore, CycleGAN can be trained with signals containing the characteristics of either of the two classes. To train CycleGAN in this way, two adversarial generation networks (GANs) are used. The first GAN is trained to provide a signal having the characteristics of the second class based on a signal having the characteristics of the first class, and the second GAN is trained to provide a signal having the characteristics of the first class based on a signal having the characteristics of the second class. Each GAN includes a generator and a discriminator. The generator generates each signal, and the discriminator evaluates whether the signal is realistic with respect to the desired output. For example, when the GAN is trained to convert a signal with a stop sign into a signal with a speed limit sign, the generator first provides a signal with a speed limit sign based on a signal with a stop sign, and then the discriminator determines how realistic the signal with a speed limit sign is compared to other signals with a speed limit sign.

[0024] Compared with a normal GAN, CycleGAN imposes additional constraints during training. When the second signal provided from the first GAN based on the first signal is provided to the second GAN, the third signal provided by the second GAN needs to be as close as possible to the first signal. Here, when the respective sub - element values of the first signal and the third signal are as close as possible, the third signal can be understood to be as close as possible to the first signal.

[0025] The generator part of the first GAN can be understood as the second generator, and the generator part of the second GAN can be understood as the third generator.

[0026] The advantage of preparing the first generator based on CycleGAN is that the first generator can accurately obtain a signal with the characteristics of the second class. This is because CycleGAN is currently the optimal model for conversion between non - paired signals. That is, in the training of CycleGAN, and thus in the training of the first generator, it is not necessary to match a signal with the characteristics of the first class to a signal with the characteristics of the second class. For this reason, it becomes possible to collect a large amount of data from either of the two classes, and there is no constraint that each data point displays similar content of other classes. As a result, the generator can be trained with more data, so the performance is improved, that is, the accuracy in providing signals is improved.

[0027] In another aspect of the present invention, providing an output signal with the characteristics of the second class based on the provided input signal with the characteristics of the first class by the first generator or the second generator, · a step of obtaining an intermediate signal and a mask signal based on the input signal; · a step of obtaining the Hadamard product of the mask signal and the intermediate output signal and providing the obtained Hadamard product as a first result signal; · a step of obtaining the Hadamard product of the reciprocal of the mask signal and the input signal and providing the obtained Hadamard product as a second result signal; · providing, as an output signal, the sum of the first result signal and the second result signal; may further be assumed to be included.

[0028] During the training of the second generator, it may further be assumed that the generator is trained by optimizing a loss function, and the sum of the absolute values of the sub-elements of the mask signal is added to the loss function.

[0029] The mask signal and / or the intermediate signal are preferably of the same height and width as the input signal. Alternatively, the mask signal and / or the intermediate signal may be scaled to the same size as the input signal. The multiplication of the mask signal and the intermediate signal can be understood as an element-wise multiplication. That is, the first sub-element at a position in the first result signal can be obtained by multiplying the mask sub-element at the position in the mask signal by the intermediate sub-element at the position in the intermediate signal. The second sub-element in the second result signal can also be obtained in a corresponding and similar manner.

[0030] The reciprocal of the mask signal can be obtained by subtracting the mask signal from a signal of the same shape as the mask signal, and the sub-elements of the signal are unit vectors.

[0031] When the depth of the mask signal is 1, it may be assumed that obtaining the individual Hadamard products can be achieved by first stacking copies of the mask signal along the depth dimension to form a new mask signal that matches the depth of the input signal, and then using the new mask signal as the mask signal to calculate the Hadamard products.

[0032] The mask signal may also be assumed to be composed of scalar sub-elements in the range of 0 to 1. Therefore, the masking operation can be regarded as blending the intermediate signal with the input signal to provide the output signal.

[0033] The advantage of providing a mask signal by the second generator is that the mask signal allows obtaining an output signal by changing only a small area of the input signal. Such characteristics are further enhanced by adding the sum of the absolute values of the mask signal to the loss function during training. As a result, the generator will provide the smallest possible non-zero values for the mask signal. This has the effect that the input signal is changed only in areas containing individual features of the first class that required changes to provide an output signal with characteristics of the second class. Also, this allows finding the individual features of the first class more accurately, and as a result, the accuracy of the desired output signal is improved.

[0034] In another aspect of the present invention, it can further be assumed that providing a desired output signal includes obtaining a norm value of a sub-element of the difference signal corresponding to this sub-element and comparing the norm value with a predetermined threshold value.

[0035] The norm value of the sub-element can be obtained by calculating the L-norm of the sub-element, such as the Euclidean norm, Manhattan norm, or infinity norm. p It can be obtained by calculating the norm.

[0036] The advantage of this approach is that it is possible to compare a difference signal with a depth greater than 1 with a predetermined threshold value in order to determine where the individual features of the class are located in the input signal. As a result, the method proposed for input signals with a depth greater than 1 can be used.

[0037] In another aspect of the present invention, it can further be assumed that a bounding box is provided by determining an area within a difference signal that includes a plurality of sub-elements, where the output signal includes at least one bounding box and includes sub-elements whose corresponding norm values do not fall below a threshold value, and providing the area as the bounding box.

[0038] This region is preferably a rectangular region, and each edge of the rectangular region is parallel to one of the edges of the difference signal. Preferably, the region extends so as to exactly contain sub-elements from a plurality of sub-elements. Alternatively, it can also be assumed that the region extends a predetermined amount of sub-elements around the plurality of sub-elements, i.e., the bounding box is padded. It can further be assumed that the predetermined amount depends on the amount of the plurality of sub-elements. For example, the predetermined amount may be obtained by multiplying a constant value by the amount of the plurality of sub-elements.

[0039] It can be assumed that a plurality of regions from the signal are returned as a bounding box. The plurality of regions can be determined, for example, by connected component analysis of the difference signal, where the sub-elements of the difference signal are regarded as connected components only if the corresponding norm value does not fall below a threshold.

[0040] The advantage of this approach is that the bounding box is automatically generated without a human teacher and can then be used as a label for training a classifier to classify objects within the signal. Since no human teacher is required, this process can be applied to a large number of signals and used for training the classifier. Thereby, the classifier can learn information from more signals and achieve higher classification performance.

[0041] In another aspect of the present invention, the desired output signal includes a semantic segmentation signal of the input signal, and providing the semantic segmentation signal includes: · providing a signal having the same height and width as the difference signal; · obtaining the norm value of a sub-element in the difference signal having the position in the difference signal; · when the norm value is below a predetermined threshold, setting the sub-element at the position in the signal to belong to the first class, and otherwise setting the sub-element to the background class. · providing the signal as a semantic segmentation signal; may further be assumed to include.

[0042] Setting the sub-elements within the signal to classes can be understood as setting the values of the sub-elements to characterize their membership in a particular class. For example, the depth of the signal may well be assumed to be 1, and if the norm value is below the threshold, a value of 0 may be assigned to the sub-element to indicate membership in the first class, and otherwise a 1 may be assigned.

[0043] This approach can be understood as determining which sub-elements need to be changed in order to transform the signal from having the characteristics of the first class to having the characteristics of the second class. Thus, the identified sub-elements contain the characteristics of the first class of the first signal. All other sub-elements can be regarded as a background class that does not contain the relevant characteristics.

[0044] The advantage of this approach is that the semantic segmentation signal can be automatically obtained without a human teacher and can then be used as a label for training a classifier to perform semantic segmentation of the signal. Since no human teacher is required, this process can be applied to a large number of signals in a short time. Subsequently, the classifier can be trained using the signal. As a result, the classifier can learn information from more labeled signals than a human can discriminate in the same amount of time, and thus higher classification performance can be achieved.

[0045] In another aspect, the present invention relates to a computer-implemented method for determining whether a classifier can be used in a control system, the classifier being usable in the control system if and only if its classification performance for a test data set exhibits acceptable accuracy with respect to a threshold, and providing the test data set is · A step of preparing a first generator, wherein the first generator is configured to provide an output signal having characteristics of a second class based on a provided input signal having characteristics of a first class, or the generator is configured to provide a mask signal indicating which part of the input signal represents the characteristics of the first class; · A step of providing an output signal or a mask signal by the first generator based on the input signal; · A step of providing a difference signal corresponding to the input signal, wherein the difference signal is provided based on the difference between the input signal and the output signal, or the mask signal is provided as the difference signal; · A step of providing a desired output signal corresponding to the input signal based on the corresponding difference signal, wherein the desired output signal characterizes the desired classification that should be provided by a classifier when the input signal is provided; · A step of providing at least one input signal and the corresponding desired output signal as a test data set; including.

[0046] The method of generating a test data set is similar to the method of generating a training data set described in the first aspect of the present invention.

[0047] Depending on the task of the classifier, the classification performance can be measured by the accuracy of single-label signal classification, the (average) AP (Average Precision) of object detection, or the (average) IoU (Intersection over Union) of semantic segmentation. It is also conceivable that the negative log-likelihood is used as a classification performance metric.

[0048] The advantage of this aspect of the present invention is that the proposed approach enables obtaining insights into the internal operation of the classifier. This can function as a means for determining situations where the classification performance of the classifier is insufficient for use in a product. For example, the proposed approach can be used as part of determining whether a classifier used in a vehicle that is at least partially autonomous for detecting pedestrians has a high enough classification performance to be safely used as part of the at least partially autonomous vehicle.

[0049] It can be assumed that the method for testing the classifier may be repeatedly iterated. In each iteration, the classifier is tested with a test dataset, and the classification performance is determined. If the performance of the classifier is insufficient for use in a product, for example, preferably, the classifier can be strengthened by training the classifier with new training data generated by the method according to the first aspect of the present invention. The process here can be repeated until the classifier achieves sufficient performance with the test dataset. Alternatively, it can also be assumed that a new test dataset is required in each iteration.

[0050] Embodiments of the present invention will be discussed in more detail with reference to the following drawings.

Brief Description of the Drawings

[0051]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0052] Description of Embodiments FIG. 1 shows an embodiment of an actuator (10) in an environment (20). The actuator (10) has an interaction with a control system (40). The actuator (10) and its environment (20) are collectively referred to as an actuator system. Preferably, at equally spaced time points, a sensor (30) measures the state of the actuator system. The sensor (30) may include a plurality of sensors. Preferably, the sensor (30) is an optical sensor that captures an image of the environment (20). The output signal (S) of the sensor (30) that encodes the measured state (or, if the sensor (30) includes a plurality of sensors, the output signal (S) of each of the sensors) is transmitted to the control system (40).

[0053] Thereby, the control system (40) receives a stream of sensor signals (S). Next, the control system (40) calculates a series of actuator control commands (A) in response to the stream of sensor signals (S) and then transmits them to the actuator (10).

[0054] The control system (40) receives, in the receiving unit (50) as an optional means, a stream of sensor signals S from the sensor (30). The receiving unit (50) converts the sensor signal (S) into an input image (x). Alternatively, if there is no receiving unit (50), each sensor signal (S) may be directly used as the input image (x). The input image (x) can be provided, for example, by extracting from the sensor signal (S). Alternatively, the sensor signal (S) may be processed to generate the input image (x). The input image (x) contains image data corresponding to the signals recorded by the sensor (30). In other words, the input image (x) is provided according to the sensor signal (S).

[0055] The input image (x) is then passed to the classifier (60).

[0056] The classifier (60) is parameterized by the parameter (Φ). The parameter (Φ) is stored in the parameter storage (St1) and provided by the parameter storage (St1).

[0057] The classifier (60) obtains an output signal (y) from the input image (x). The output signal (y) contains information for assigning one or more labels to the input image (x). The output signal (y) is transmitted to the conversion unit (80) as an optional means, and the conversion unit (80) converts the output signal (y) into a control command (A). Then, the actuator control command (A) is transmitted to the actuator (10), and the actuator (10) is controlled accordingly. Alternatively, the output signal (y) may be directly used as the actuator control command (A).

[0058] The actuator (10) receives the actuator control command (A) and is controlled accordingly, performing an operation corresponding to the actuator control command (A). The actuator (10) may include control logic for converting the actuator control command (A) into a further control command, and the further control command is then used to control the actuator (10).

[0059] In a further embodiment, the control system (40) may include the sensor (30). In still other embodiments, the control system (40) may alternatively or additionally include the actuator (10).

[0060] In one embodiment, the classifier (60) can be designed to identify lanes on the road ahead, for example, by classifying the surface of the road and / or signs on the road and identifying lanes as patches of road surface between signs. Based on the output of the navigation system, a lane suitable for tracking the selected route can be selected, and depending on the current lane and the target lane, the vehicle (100) can determine whether to change lanes or stay in the current lane. In this case, the actuator control command (A) can be calculated, for example, by obtaining a predetermined operation pattern from a database corresponding to the identified operation.

[0061] Alternatively or additionally, the classifier (60) can be configured to detect road signs and / or traffic lights. When a road sign or traffic light is identified, then, depending on the type of the identified road sign or the state of the identified traffic light, constraints corresponding to possible operation patterns of the vehicle (100) can be obtained, for example, from a database, and the trajectory of the vehicle (100) can be calculated according to the constraints, and for example, the actuator control command (A) can be calculated to steer the vehicle (100) to navigate along the trajectory.

[0062] Alternatively or additionally, the classifier (60) can be configured to detect pedestrians and / or vehicles. When pedestrians and / or vehicles are identified, the predicted future actions of the pedestrians and / or vehicles can be estimated, and then, based on the estimated future actions, a trajectory can be selected to avoid collisions with the identified pedestrians and / or vehicles, and the actuator control command (A) can be calculated to steer the vehicle (100) to navigate along the trajectory.

[0063] In still other embodiments, it may be envisioned that the control system (40) controls the display (10a) instead of or in addition to the actuator (10).

[0064] Furthermore, the control system (40) may include a single processor (45) (or multiple processors) and at least one machine-readable storage medium (46). Here, the at least one machine-readable storage medium (46) stores instructions that, when executed, cause the control system (40) to perform a method according to an aspect of the present invention.

[0065] FIG. 2 shows an embodiment in which a control system (40) is used to control at least a partially autonomous robot, e.g., at least a partially autonomous vehicle (100).

[0066] The sensor (30) may include one or more video sensors and / or one or more radar sensors and / or one or more ultrasonic sensors and / or one or more LiDAR sensors and / or one or more position sensors (e.g., GPS, etc.). Some or all of these sensors are preferably incorporated into the vehicle (100), but they do not necessarily have to be. Alternatively or additionally, the sensor (30) may include an information system for obtaining the state of the actuator system. An example of such an information system is a weather information system for obtaining the current or future state of the weather in the environment (20).

[0067] For example, using the input image (x), the classifier (60) can detect an object in the vicinity of at least a partially autonomous robot. The output signal (y) may include information characterizing the location where the object is located in the vicinity of at least a partially autonomous robot. Next, the actuator control command (A) can be determined according to the information, e.g., to avoid a collision with the detected object.

[0068] Preferably, the actuator (10) incorporated into the vehicle (100) can be provided by the brakes, propulsion system, engine, drive train or steering of the vehicle 100. An actuator control command (A) can be determined to control a single actuator (or a plurality of actuators) (10) so as to avoid a collision of the vehicle (100) with a detected object. The detected object can be classified by the classifier (60) as the most likely one, for example, a pedestrian or a tree, and the actuator control command (A) can be determined according to the classification.

[0069] In a further embodiment, the at least partially autonomous robot can be provided by another mobile robot (not shown) that can move, for example, by flying, swimming, diving or walking. The mobile robot can in particular be a lawn mower that is at least partially autonomous or a cleaning robot that is at least partially autonomous. In all of the above embodiments, an actuator control command (A) can be determined to control the propulsion unit and / or steering and / or brakes of the mobile robot so that the mobile robot can avoid a collision with the identified object.

[0070] In a further embodiment, the at least partially autonomous robot can be provided by a gardening robot (not shown) that uses a sensor (30), preferably an optical sensor, to determine the state of plants in the environment (20). The actuator (10) can control a nozzle for spraying a liquid and / or a cutting device such as a blade. Depending on the identified plant species and / or the identified state of the plant, the actuator control command (A) can be determined to cause the actuator (10) to spray an appropriate amount of an appropriate liquid onto the plant and / or to cut the plant.

[0071] In a further embodiment, the at least partially autonomous robot can be provided by a household appliance (not shown), such as a washing machine, stove, oven, microwave or dishwasher. A sensor (30), such as an optical sensor, can detect the state of the object being processed by the household appliance. For example, if the household appliance is a washing machine, the sensor (30) can detect the state of the laundry inside the washing machine. Next, an actuator control command (A) can be determined according to the detected material of the laundry.

[0072] FIG. 3 shows an embodiment in which a control system (40) is used to control a manufacturing machine (11), such as a punch cutter, cutter or drill, of a manufacturing system (200) as part of a manufacturing line. The control system (40) controls an actuator (10), and the actuator (10) controls the manufacturing machine (11).

[0073] The sensor (30) can be provided, for example, by an optical sensor that captures the characteristics of the manufactured product (12). A classifier (60) can determine the state of the manufactured product (12) from these captured characteristics. Thereafter, the actuator (10) that controls the manufacturing machine (11) can be controlled according to the determined state of the manufactured product (12) for subsequent manufacturing steps of the manufactured product (12). Alternatively, it can also be envisioned that the actuator (10) is controlled during the manufacture of the subsequent manufactured product (12) according to the determined state of the manufactured product (12).

[0074] FIG. 4 shows an embodiment in which the control system (40) is used for the control of an automatic personal assistant (250). The sensor (30) may be, for example, an optical sensor for receiving video images of the gestures of the user (249). Alternatively, the sensor (30) may also be, for example, an audio sensor for receiving voice commands of the user (249).

[0075] Next, the control system (40) determines an actuator control command (A) for controlling the automatic personal assistant (250). The actuator control command (A) is determined according to the sensor signal (S) of the sensor (30). The sensor signal (S) is transmitted to the control system (40). For example, the classifier (60) can be configured to execute a gesture recognition algorithm or the like to identify the gesture made by the user (249). Next, the control system (40) can determine an actuator control command (A) for transmission to the automatic personal assistant (250). Next, the actuator control command (A) is transmitted to the automatic personal assistant (250).

[0076] For example, the actuator control command (A) may be determined according to the identified user gesture recognized by the classifier (60). The actuator control command (A) may include information for causing the automatic personal assistant (250) to obtain information from the database and output the obtained information in a form suitable for reception by the user (249).

[0077] In a further embodiment, it can be assumed that the control system (40) controls a home appliance (not shown) controlled according to the identified user gesture instead of the automatic personal assistant (250). The home appliance may be a washing machine, a stove, an oven, a microwave oven, or a dishwasher.

[0078] FIG. 5 shows an embodiment in which the control system (40) controls an access control system (300). The access control system (300) can be designed to physically control access. For example, the system may include a door (401). The sensor (30) can be configured to detect relevant scenes to determine whether to permit access. The sensor (30) may be, for example, an optical sensor that provides image or video data for detecting a human face. The classifier (60) can be configured to interpret this image or video data, for example, by comparing identification information with known people stored in a database to determine the identification information of an individual. In this case, the actuator control signal (A) can be determined according to, for example, the determined identification information, in response to the interpretation of the classifier (60). The actuator (10) can be a lock that opens and closes the door in response to the actuator control signal (A). Non-physical and logical access control is also possible.

[0079] FIG. 6 shows an embodiment in which the control system (40) controls a monitoring system (400). This embodiment is substantially the same as the embodiment shown in FIG. 5. Therefore, only the differences will be described in detail. The sensor (30) is configured to detect a scene under monitoring. The control system (40) does not necessarily need to control the actuator (10), and instead, it may control the display (10a). For example, the classifier (60) can determine whether a detected scene, for example, by the optical sensor (30), is suspicious. In this case, the actuator control signal (A) transmitted to the display (10a) can be configured to adjust the display content on the display (10a) according to the required classification, for example, to highlight an object considered suspicious by the classifier (60).

[0080] FIG. 7 shows an embodiment of a control system (40) for controlling an imaging system (500) such as, for example, an MRI apparatus, an X-ray imaging apparatus, or an ultrasonic imaging apparatus. The sensor (30) may be, for example, an imaging sensor. Next, the classifier (60) can obtain a classification of all or part of the measured image. Then, an actuator control signal (A) can be selected according to the classification, whereby the display (10a) can be controlled. For example, the classifier (60) can interpret a region of the measured image as potentially abnormal. In this case, the actuator control signal (A) can be determined to cause the display (10a) to display the image and highlight the potentially abnormal region.

[0081] FIG. 8 shows an embodiment of a training system (140) for performing training of the classifier (60). To perform the training, the training data unit (150) accesses a computer-implemented training database (St2) in which at least one set (T) of training data is stored. The set (T) includes pairs of images and corresponding desired output signals.

[0082] Next, the training data unit (150) provides a batch (x i ) of images to the classifier (60). The batch (x i ) of images can include one or more images from the set (T). Then, the classifier (60) obtains a batch of output signals i from the batch (x

Number

Number

[0083] Next, based on the batch of output signals [Number] and the batch of desired output signals (y i ), the update unit obtains, for example, a set of updated parameters (Φ’) of the classifier (60) using stochastic gradient descent. Then, the set of updated parameters (Φ’) is stored in the parameter storage (St1).

[0084] In a further embodiment, the training process is then repeated a desired number of times, and the set of updated parameters (Φ’) is provided as the set of parameters (Φ) by the parameter storage (St1) in each iteration.

[0085] Furthermore, the training system (140) can include a single processor (145) (or a plurality of processors) and at least one machine-readable storage medium (146). Here, the at least one machine-readable storage medium (146) stores instructions that, when executed, cause the training system (140) to perform the training method according to an aspect of the present invention.

[0086] FIG. 9 is a flowchart showing a method for obtaining a set (T) of images and desired output signals.

[0087] In this method, in the first step (901), a first generator is prepared. The first generator is prepared by training CycleGAN with images of a first dataset (D1) having features of a first class and images of a second dataset (D2) having features of a second class. CycleGAN is trained to convert images of the first dataset (D1) into similar images from the second dataset (D2). In this case, the generator of CycleGAN responsible for the conversion from the first dataset (D1) to the second dataset (D2) is prepared as the first generator.

[0088] Next, in the second step (902), images of a new dataset (D n ) are provided to the generator. The generator provides an output image for each image.

[0089] In the third step (903), a difference image for each image of the new dataset (D n ) is obtained by subtracting an image from the corresponding output image. In a further embodiment, it may be assumed that the difference image is obtained by subtracting the output image from the image.

[0090] In the fourth step (904), a desired output signal for each image is obtained based on the difference image corresponding to the image. For example, the desired output signal may include a bounding box indicating the presence of a specific object in the image. These bounding boxes can be obtained by first calculating the norm value of each pixel in the difference image. If the norm value does not fall below a predetermined threshold, the pixel can be considered to contain data that has changed significantly from the image to the output image. Next, connected component analysis can be performed on the pixels whose norm does not fall below the threshold. This provides a set of one or more pixels belonging to the same component. Each of these components can be considered as the associated object, and a bounding box can be obtained for each component. The bounding box can be obtained so as to have the minimum size while enclosing all the pixels of the component, that is, so as to fit tightly around the pixels of the component. Alternatively, it can also be assumed that the bounding box is selected such that its edges have a predetermined maximum distance to the pixels of the component. This approach can be considered as obtaining a bounding box padded around the pixels.

[0091] The bounding box does not necessarily have to be rectangular. Without requiring a change to the above method, the bounding box can have any shape, convex or non-convex.

[0092] In this case, the set (T) may be provided as an image from the new data set (D n ) and the bounding box corresponding to each of the images, that is, as the desired output signal.

[0093] In a further embodiment, instead of the bounding box, it can be assumed that a semantic segmentation image is obtained from the difference image. Therefore, all pixels of the signal whose norm value does not fall below the threshold can be labeled as belonging to the first class, and all other pixels of the image can be labeled as the background class. Then, the signal including the pixels labeled according to this procedure can be provided as the semantic segmentation image in the desired output signal.

[0094] In a further embodiment, it can be assumed that a mask image used as the difference image is provided by a generator used in CycleGAN. FIG. 10 shows an embodiment of a generator (G) that provides a mask image (M) to obtain an output image (O). Based on the input image (I), the generator (G) provides a mask image (M) and an intermediate image (Z).

[0095] The generator (G) is preferably configured such that the provided mask image (M) includes only values in the range of 0 to 1. This can preferably be achieved by the generator (G) providing a preliminary mask image and applying a sigmoid function to the preliminary mask image to obtain a mask image (M) having values in the range of 0 to 1.

[0096] Then, the mask image (M) is inverted by an inversion unit (Inv) to obtain an inverted mask image (M -1 ). Then, the Hadamard product of the intermediate image and the mask image (M) is provided as a first result image (R1). The Hadamard product of the input image (I) and the inverted mask image (M -1 ) is provided as a second result image (R2). Then, the sum of the first result image (R1) and the second result image (R2) is provided as the output image (O).

[0097] In a further embodiment (not shown), it may be assumed that the set (T) is used not for training the classifier (60), but for testing the classifier (60). For example, the classifier (60) is assumed to be used in a product where safety is important, such as an autonomous vehicle (10). It may further be assumed that the classifier (60) needs to achieve a predetermined classification performance before being permitted for use in the product.

[0098] The classification performance can be evaluated by determining the classification performance of the classifier (60) in the test data set, and the test data set is provided using the above steps 1 (901) to 4 (904).

[0099] Clearing of the classifier (60) may be performed, for example, by an iterative method, and as long as the classifier (60) does not meet the predetermined classification performance, it can be further trained with new training data and then tested again with the provided test data set.

Claims

A computer-implemented method for providing a training data set (T) for training a classifier (60) configured to provide an output signal (y) characterizing the classification of a first input signal (x), comprising: - A step (901) of preparing a first generator (G), wherein the first generator (G) is configured to provide an output signal (O) having characteristics of a second class based on a provided second input signal (I) having characteristics of a first class, or the first generator (G) is configured to provide a mask signal (M) indicating which portion of the second input signal (I) represents the characteristics of the first class, step (901); - A step (902) of providing an output signal (O) or a mask signal (M) by the first generator (G) based on the provided second input signal (I); - A step (903) of providing a difference signal corresponding to the second input signal (I), wherein the difference signal is provided based on the difference between the second input signal (I) and the output signal (O), or the mask signal (M) is provided as the difference signal, step (903); - A step (904) of providing a desired output signal corresponding to the second input signal (I) based on the corresponding difference signal, wherein the desired output signal characterizes the desired classification of the second input signal (I), step (904); - A step of providing at least one second input signal (I) and the corresponding desired output signal as a training data set (T); comprising: The step (901) of preparing the first generator (G) includes training a CycleGAN including a second generator and a third generator; Training the CycleGAN includes: Training the second generator to provide an output signal having characteristics of the second class based on the provided input signal having characteristics of the first class; Training the third generator to provide an input signal having characteristics of the first class based on the provided output signal having characteristics of the second class; After training the CycleGAN, preparing the second generator as the first generator (G); comprising the method. Claim 2 Providing an output signal (O) having characteristics of the second class based on a provided second input signal (I) having characteristics of the first class by the first generator (G) or the second generator is - a step of obtaining an intermediate output signal (Z) and a mask signal (M) based on the second input signal (I); - Obtain the Hadamard product of the mask signal (M) and the intermediate output signal (Z), and provide the obtained Hadamard product as a first result signal (R 1 ). - The reciprocal of the mask signal (M -1 ) and the Hadamard product of the second input signal (I) are obtained, and the obtained Hadamard product is provided as a second result signal (R 2 ); and ・Outputting the sum of the first result signal (R 1 ) and the second result signal (R 2 ) as an output signal (O); The method according to claim 1, comprising:

3. The method according to claim 2, wherein the second generator is trained using a loss function that depends on the sum of the absolute values of the sub-elements of the mask signal (M).

4. The step of providing the desired output signal is - obtaining a norm value corresponding to each sub-element of the difference signal for each sub-element of the difference signal; - comparing the norm value with a predetermined threshold; The method according to any one of claims 1 to 3, comprising:

5. The output signal includes at least one bounding box, The bounding box is provided by obtaining a region within the difference signal that includes a plurality of sub-elements whose corresponding norm values do not fall below the threshold, and providing the region as the bounding box. The method according to claim 4.

6. The desired output signal includes a semantic segmentation signal of the input signal, Providing the semantic segmentation signal includes - providing a signal having the same height and width as the difference signal; - obtaining a norm value of a sub-element within the difference signal having a position within the difference signal; - when the norm value falls below a predetermined threshold, setting the sub-element at the position within the signal to belong to the first class, and otherwise setting the sub-element to the background class; - providing the signal as a semantic segmentation signal; including The method according to claim 4.

7. A computer-implemented method for obtaining an output signal (y) characterizing the classification of an input signal (x), comprising: - training a classifier (60) by the method according to any one of claims 1 to 6; - supplying the classifier (60) to a control system (40); - A step of obtaining the output signal (y) from the control system (40), wherein the control system (40) provides the input signal (x) to the classifier (60) to obtain the output signal (y). A method comprising the above. **Claim 8** The method according to claim 7, wherein the input signal (x) is obtained based on the signal (S) of the sensor (30), and / or the actuator (10) is controlled based on the output signal (y), and / or the display device (10a) is controlled based on the output signal (y). **Claim 9** A computer-implemented method for providing a test data set for determining whether a classifier (60) is usable in a control system (40), wherein the classifier (60) is usable in the control system (40) if and only if its classification performance for the test data set exhibits acceptable accuracy with respect to a threshold value. In the method, - A step (901) of preparing a first generator (G), wherein the first generator (G) is configured to provide an output signal (O) having characteristics of a second class based on a provided input signal (I) having characteristics of a first class, or the first generator (G) is configured to provide a mask signal (M) indicating which portion of the input signal (I) represents the characteristics of the first class. Step (901). - A step (902) of providing an output signal (O) or a mask signal (M) by the first generator (G) based on the input signal (I). - A step (903) of providing a difference signal corresponding to the input signal (I), wherein the difference signal is provided based on the difference between the input signal (I) and the output signal (O), or the mask signal (M) is provided as the difference signal. Step (903). - A step (904) of providing a desired output signal corresponding to the input signal (I) based on the corresponding difference signal, wherein the desired output signal characterizes the desired classification of the input signal (I). Step (904). - A step of providing at least one input signal (I) and the corresponding desired output signal as a test data set. Including The step (901) of preparing the first generator (G) includes training a CycleGAN including a second generator and a third generator. Training the CycleGAN includes training the second generator to provide an output signal having the characteristics of the second class based on a provided input signal having the characteristics of the first class; training the third generator to provide an input signal having the characteristics of the first class based on a provided output signal having the characteristics of the second class; preparing the second generator as a first generator (G) after training the CycleGAN; A method comprising the above steps. **Claim 10** A computer program configured to cause a computer to perform all the steps of the method according to any one of Claims 1 to 9 when the computer program is executed by a processor (45, 145). **Claim 11** A machine-readable storage medium (46, 146) storing the computer program according to Claim 10. **Claim 12** A control system (40) configured to control an actuator (10) and / or a display device (10a) based on an output signal (y) of a classifier (60), wherein the classifier is trained by the method according to any one of Claims 1 to 6; A control system (40). **Claim 13** A training system (140) configured to perform the method according to any one of Claims 1 to 6.

Citation Information

Patent Citations

  • Recognition system, common feature amount extracting unit, and recognition system configuring method

    JP2018195097A

  • Area discriminator training method, area discrimination device, area discriminator training device, and program

    JP2019061658A

  • Ar compatible labeling using aligned cad models

    JP2020087440A

  • Object detection method, object detection device, and computer program

    WO2020031422A1

  • Teacher data generation device, teacher data generation method, and teacher data generation system

    WO2020049634A1