Devices and methods for classifying input signals using reversible factorization models

CN114358105BActive Publication Date: 2026-09-01ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111143407.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-29
Filing Date
2021-09-28
Publication Date
2026-09-01
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

虽然深度分类器通常能够实现良好或非常好的分类准确度,但是通常很难或甚至不可能推断出深度分类器将其决策基于输入信号的什么语义方面

Benefits of technology

在设计分类器时,通常存在折衷。分类器可以被构造成使得可以容易地从分类器推断出分类器将其决策基于输入信号的什么部分。例如,线性分类器就是这种情况。然而,也称为浅分类器的这些分类器通常不实现足够的分类准确度,尤其是当要分类的输入信号是高维的时候。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358105B_ABST
    Figure CN114358105B_ABST
Patent Text Reader

Abstract

Apparatus and method for classifying input signals using a reversible factorization model. A computer-implemented method for determining an output signal of an input signal using a classifier, wherein the output signal characterizes the classification of the input signal, the method comprising the steps of: ● determining a latent representation based on the input signal by means of a reversible factorization model included in the classifier, wherein the latent representation comprises a plurality of factors, and wherein the reversible factorization model is characterized by: - ​​a plurality of functions, wherein any one of the plurality of functions is continuous and nearly everywhere continuously differentiable, wherein the function is further configured to accept the input signal or at least one factor provided by another of the plurality of functions as input, and wherein the function is further configured to provide at least one factor; ● determining the output signal based on the latent representation by means of an internal classifier included in the classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for classifying input signals by means of a classifier, a method for training a classifier, a model for obtaining a latent representation, a classifier, a training system, a computer program, and a machine-readable storage medium. Existing technology

[0002] Jens Behrmann, Will Grathwohl, Ricky TQ Chen, David Duvenaud, and Jörn-Henrik Jacobsen's "Invertible Residual Networks" (https: / / arxiv.org / abs / 1811.00995v3, May 18, 2019) discloses a method for inverting neural networks.

[0003] Advantages of the present invention There are often trade-offs when designing classifiers. Classifiers can be constructed such that it is easy to infer from the classifier what part of the input signal it bases its decisions on. Linear classifiers are an example of this. However, these classifiers, also known as shallow classifiers, often do not achieve sufficient classification accuracy, especially when the input signal to be classified is high-dimensional.

[0004] On the other hand, so-called deep classifiers can be constructed, preferably based on deep neural networks. While deep classifiers typically achieve good or very good classification accuracy, it is often difficult or even impossible to infer what semantic aspects of the input signal the deep classifier bases its decisions on. This characteristic earns deep classifiers the concept of being a black box when it comes to classification.

[0005] Therefore, it is desirable to obtain a deep classifier that can determine which semantic aspects of the input signal to be classified cause the deep classifier to determine a specific classification.

[0006] The advantage of the method according to the invention is that the classification accuracy is as good as that of deep classifiers, while the decisive factors for classification can be clearly and easily inferred. This is achieved by first determining factors, which can be understood as the definition and semantic properties of the input signal. These factors assist the user of the classifier because the classification obtained from the classifier can be supervised by the user. With the help of these factors, the user gains insight into the internal decision-making of the classifier. If a classification is based on factors that the user knows are not allowed or are potentially error-prone, the user can identify the classifier's failures and train the classifier to correct these failures. Therefore, this method assists the user in determining the correctness of the classification during guided human-computer interaction.

[0007] Another advantage of the method proposed in this invention is that the factors on which the classification is based are at least partially unentangled through construction. This allows the classifier to base its classification decisions on more reliable features, i.e., factors, which in turn makes the classifier more robust to perturbations. This increased robustness leads to improved classification accuracy. Therefore, the proposed invention also improves the performance of the classifier. Summary of the Invention

[0008] In a first aspect, the present invention relates to a computer-implemented method for determining an output signal of an input signal using a classifier, wherein the output signal characterizes the classification of the input signal, the method comprising the following steps: ● A latent representation is determined based on the input signal using a reversible factorization model included in the classifier, wherein the latent representation comprises multiple factors, and wherein the reversible factorization model is characterized by: - A plurality of functions, wherein any one of the plurality of functions is continuous and nearly everywhere continuously differentiable, wherein the function is further configured to accept an input signal or at least one factor provided by another of the plurality of functions as input, and wherein the function is further configured to provide at least one factor, wherein the at least one factor is either provided as at least part of a potential representation or as at least part of the input of another of the plurality of functions, wherein there exists an inverse function corresponding to the function, the inverse function being continuous and nearly everywhere continuously differentiable, and being configured to determine the input of the layer based on the at least one factor provided from the function; ● The output signal is determined based on the latent representation by means of an internal classifier included in the classifier.

[0009] The input signal may include or consist of at least one image. Images may be acquired from sensors such as camera sensors, LiDAR sensors, radar sensors, ultrasonic sensors, or thermal cameras. Images may also be acquired from simulated renderings, such as virtual environments created in a computer. Digital rendering of the images is also conceivable. In particular, the input signal may also include images of different modalities or images from different sensors of the same type. For example, the input signal may include two images acquired from a stereo camera. As another example, the input signal may also include images from a camera sensor and a LiDAR sensor.

[0010] The input signal may also include audio signals, such as recordings from a microphone or audio data obtained from computer simulation or computer digital generation. Audio signals may characterize speech recordings. Alternatively or additionally, audio signals may characterize recordings of audio events, such as alarm sounds, alerts, or other notification signals.

[0011] The output signal can assign at least one label to the input signal that characterizes the classification of the input signal. For example, if the input signal includes an image, the output signal can characterize a single class to which the classifier believes the input signal belongs. Alternatively or additionally, the output signal can characterize the classification and location of objects in the image; that is, the classifier can be configured for object detection. Alternatively or additionally, the classifier can assign the class to which the classifier believes that pixel belongs to each pixel in the image; that is, the classifier can be configured for semantic segmentation.

[0012] It is also conceivable that if the input signal includes an audio signal, the output signal can assign the class to which the classifier believes the audio signal belongs. Alternatively or additionally, the output signal can also characterize audio events included in the audio signal and / or their location within the audio signal.

[0013] The classifier can preferably be understood as a neural network, wherein the neural network includes a layer for extracting factors from the input signal, i.e., a function of the invertible factorization model, and at least one layer for determining the output signal based on the obtained factors.

[0014] A reversible factorization model can be understood as a machine learning model. It takes an input signal as input and provides a latent representation, i.e., multiple factors, as output. Since each function included in a reversible factorization model is invertible, the input signal can be determined from the latent representation. Viewing a reversible factorization model as part of a neural network can be understood as the inverse propagation of factors to obtain the input signal.

[0015] A function derived from multiple functions and its corresponding inverse function can be understood as one or more layers of a reversible factorization model. The function can be used for forward propagation of the layer's input to obtain the layer's output, and the inverse function can be used for backward propagation, i.e., obtaining the input from the layer's output. By definition, the functions of a reversible model, and therefore the reversible model itself, are completely reversible. That is, if a latent representation is obtained from a reversible factorization model of an input signal, the input signal can be determined again from the latent representation. Therefore, a reversible factorization model can be considered a generative model from the field of machine learning.

[0016] In a reversible factorization model, the information flow determines the order of its functions. If the first function provides factors to the second function, then the first function is considered to precede the second function, and the second function is considered to follow the first function.

[0017] Surprisingly, the inventors discovered that by designing a reversible factorization model, at least some factors of the latent representation are disentangled. A factor can be understood as disentangled if, starting from any given input signal, a certain direction in the vector space of the input signal changes only that factor while leaving all other factors relatively unchanged. Therefore, disentanglement can be understood as a continuous measurement of the factors. Thus, a factor can be considered disentangled if a change in direction corresponding to the input signal in the vector space does not change another factor by more than a predefined threshold.

[0018] Deentangled factors are advantageous because they allow for the decomposition of the input signal, i.e., separation of the input signal into semantic aspects. For example, a first factor can change if the color of an object in an image included in the input signal changes. A second factor might change if the contrast of the image changes. A third factor might change based on the number of tires visible in the image. This is advantageous because the internal classifier bases its decisions on the factors. For example, if the internal classifier relies on the first and second factors to classify the input signal as containing a car, this indicates a strong bias in the classifier towards color and contrast, two semantic features that do not necessarily allow for the conclusion that the image contains a car. In this case, the classifier's user can identify this inappropriate relationship between the factors and the classification and train the classifier with new data to eliminate the classifier's bias towards color and contrast. Therefore, deentangled factors allow for direct insight into the internal workings of the classifier and can guide the user to train the classifier for better classification performance.

[0019] Another advantage of the unentangled factors provided by the reversible factorization model is that they allow for robust evaluation of the semantic features, i.e., the factors, of the input signal. Because the factors are robust, the classifier's classification is also more robust and therefore more accurate.

[0020] Another advantage of reversible factorization models is that the latent representation can include factors from different functions derived from the reversible factorization model. For example, one can imagine multiple functions including at least one function that provides a first factor to the latent representation and a second factor to another function derived from multiple functions.

[0021] This is advantageous because the functions of the invertible model are able to extract increasingly more abstract information from the input signal. For example, a function closer to an input signal including an image can extract factors related to the image's color or contrast, while a function closer to the latent representation can extract high-level factors related to, for example, the number of objects in the image. Similarly, when the input signal includes an audio signal, a function closer to the input signal can extract factors related to the amplitude or frequency of the audio signal, while a factor closer to the latent representation can extract high-level factors related to, for example, the number of speakers that can be heard in the audio signal.

[0022] Compared to other generative models (i.e., autoencoders) that provide their latent representation based on the output of a single function, providing factors from different functions of the invertible factorization model to the latent representation allows functions closer to the latent representation to avoid carrying information about low-level factors (e.g., color, contrast, amplitude) through the invertible factorization model if these are irrelevant to the function, i.e., if high-level factors do not need to be determined. Thus, the capacity of the invertible factorization model is not burdened by having to carry unnecessary information through the model, which in turn allows functions closer to the latent representation to extract more high-level factors, as the model has the ability to focus on high-level factors. Furthermore, high-level factors enable the internal classifier, and therefore the classifier as a whole, to base its classification on more reliable information, thereby improving the classification accuracy of the classifier.

[0023] The semantic aspects involved in determining the factors of a latent representation can be accomplished, for example, by means of an ablation test. For instance, a user can provide data where each data point differs from the others only in a single aspect. For example, a user can provide multiple images, each obtained from a single image by applying a random contrast transformation. Thus, all images show the same content and differ only in the aspect of "contrast." The user can then determine the factors of the latent representation that change with respect to the contrast, i.e., with respect to multiple images. This factor of the latent representation thus carries semantic information about the contrast in a given input image. The same approach can also be chosen for higher-level aspects, such as the number of objects visible in the image.

[0024] Alternatively, the semantic aspects of a factor can be determined by identifying the latent representation for a given input signal, the factors that change the latent representation, and the new input signal obtained from the changed latent representation.

[0025] It can be further imagined that a classifier is trained based on a first difference between a determined output signal and a desired output, wherein the determined output signal is determined for the training input signal, and the desired output signal characterizes the desired classification of the training input signal.

[0026] Because invertible models and internal classifiers can be understood as chain functions, they can be trained, for example, through backpropagation, just like neural networks. This can be achieved by: determining the output signal of the input signal, determining the determined output signal and a first difference between the determined output signal and the input signal, determining the gradients of multiple parameters of the classifier with respect to the difference, and updating the parameters based on the gradients.

[0027] Surprisingly, the authors found that when training multiple parameters of the invertible factorization model on the difference, even more factors become unentangled. This further helps users identify semantic entities in the input signal of the classifier based on its decisions, and also further improves the classification accuracy of the classifier.

[0028] It can be further envisioned that the multiple functions include at least one function configured to provide two factors according to the following formula. , in It is the input of this function. z 1 is the first factor. z 2 is the second factor. It is an encoder for an automatic encoder, and It is the decoder of the autoencoder, and further, the inverse function corresponding to this function is given by the following equation. .

[0029] The advantage of this method is that the encoder allows for the compression of information contained in its input. Therefore, the extracted factors convey a higher level of information, possessing all the advantages mentioned above.

[0030] Both encoders and / or decoders can be neural networks.

[0031] One can further imagine training the encoder based on the gradients of multiple encoder parameters relative to the first difference.

[0032] This is advantageous because the inventors found that training the encoder based on the first difference improved the deentanglement of the first and second factors.

[0033] One can further imagine training the decoder based on the gradients of multiple parameters of the decoder relative to the second difference. , in It is the input of the function, and the summation is over all the squared elements of the subtraction.

[0034] This is advantageous because the inventors found that training the decoder based on the second difference further improved the deentanglement of the first and second factors.

[0035] It can be further envisioned that the multiple functions include at least one function configured to provide three factors according to the following formula. , in It is the input of the function. z 1 is the first factor and the sum of the inputs. The result of application group standardizationz 2 is the expected value of the second factor and the group standardization. z 3 is the third factor, and the group standardization and the standard deviation of the group standardization further depend on the scaling parameter. and shift parameters The inverse of the function is given by the following equation. .

[0036] The advantage of this function is that group normalization smooths the gradients of multiple parameters in the invertible factorization model, which in turn leads to faster classifier training. Given the same amount of time, faster training allows the classifier to be trained with more input signals, thus improving its performance.

[0037] One can further imagine that multiple functions include at least one function that provides two factors according to the following formula. , in It is the input of the function. z 1 is the first factor. z 2 is the second factor, and ReLU It is the linear unit of rectification, and further, the inverse function is given by the following equation. .

[0038] The advantage of using this function in a reversible factorization model is that it transforms the model into a nonlinear one. The reversible factorization model can therefore extract factors in a nonlinear manner. Consequently, it can determine factors that have a nonlinear relationship with the input signal. This, in turn, improves the performance of the classifier.

[0039] One can further imagine that the internal classifier is a linear classifier.

[0040] The advantage of using a linear classifier as an internal classifier is that its multiple parameters, i.e., weights, determine how the factors are weighted to arrive at the classification. This provides a better insight into how the classifier arrives at the classification.

[0041] It can be further envisioned that training the classifier further includes training a reversible factorization model with the aid of a neural architecture search algorithm, wherein the objective function of the neural architecture search algorithm is based on the disentanglement of multiple factors of the latent representation.

[0042] This approach is advantageous because the disentanglement of factors is optimized during classifier training, and thus the latent representation becomes even more disentangled and even more interpretable. Therefore, users of the classifier are able to extract even more information from the latent representation. Furthermore, a more disentangled latent representation further improves the classifier's performance.

[0043] Training a reversible factorization model using neural architecture search (also known as NAS) can be understood as determining the architecture of the reversible factorization model, specifically functions from multiple functions of the reversible model, such that training a classifier leads to the best possible unentanglement of the factors.

[0044] Since the function (and its corresponding inverse function) can be understood as a layer of a neural network, the invertible factorization model can be trained according to any NAS algorithm.

[0045] The goal of applying the NAS algorithm can be understood as finding a reversible factorization model that maximizes the disentanglement of its factors. Therefore, any known disentanglement metric can be used as the objective function of the NAS algorithm, such as betaVAE, FactorVAE, mutual information difference, or DCI disentanglement.

[0046] Preferably, a NAS algorithm capable of multi-objective optimization, such as LEMONADE, is chosen. This is preferred because the architecture of the invertible factorization model can be trained to optimize both the disentanglement of the latent representation and the classification accuracy of the classifier. However, single-objective NAS algorithms are also possible. For these algorithms, the disentanglement metric can be used as the objective function. Attached Figure Description

[0047] Embodiments of the invention will be discussed in more detail with reference to the following figures. The figures illustrate: Figure 1 It is a classifier; Figure 2 It is a control system that includes a classifier that controls the actuators in its environment; Figure 3 It is a control system that controls at least partially autonomous robots; Figure 4 It is a control system for controlling manufacturing machines; Figure 5 It is a control system for controlling automated personal assistants; Figure 6 It is a control system that controls access control systems; Figure 7 It is a control system for the control and monitoring system; Figure 8 It is the control system for the imaging system; Figure 9It is a control system for the medical analysis system; Figure 10 It is a training system used to train classifiers. Detailed Implementation

[0048] Figure 1 An embodiment of a classifier (60) for determining an output signal (y) that characterizes an input signal (x) is shown, wherein the classifier includes a reversible factorization model (61) and an internal classifier (62).

[0049] The input signal (x) is provided to the reversible factorization model and processed by multiple functions (F1, F2, F3, F4). The first function (F1) accepts the input signal as input and provides a first factor (z1) to the latent representation (z) and a second factor to the second function (F2). Based on the second factor, the second function (F2) determines a third factor. The third factor is provided to the third function (F3). Based on the third factor, the third function (F3) provides a fourth factor (z3) to the latent representation (z) and a fifth factor to the fourth function (F4). Based on the fifth factor, the fourth function (F4) then provides a sixth factor (z3) and a seventh factor (z4) to the latent representation (z).

[0050] The latent representation (z) is then forwarded to an internal classifier (62), which is configured to determine an output signal (y) based on the latent representation (z). In this embodiment, a linear classifier, such as a logistic regression classifier, is used as the internal classifier (62). In other embodiments, other internal classifiers (62) may also be used.

[0051] In even other embodiments, the reversible factorization model (61) may include functions different from those described in this embodiment.

[0052] Figure 2 The diagram illustrates a control system (40) for controlling the actuator (10) based on the output signal (y) of a classifier (60). At preferably uniformly spaced time points, a sensor (30) senses the condition of the actuator system. The sensor (30) may include several sensors. Preferably, the sensor (30) includes an optical sensor, such as a camera, for capturing images of the environment (20). Alternatively or additionally, the sensor (30) may also include an audio sensor, such as a microphone. The output signal (S) of the sensor (30) (or, in the case where the sensor (30) includes multiple sensors, an output signal (S) for each sensor) is transmitted to the control system (40), which encodes the sensed condition.

[0053] Therefore, the control system (40) receives a stream of sensor signals (S). It then calculates a series of control signals (A) based on the stream of sensor signals (S) and transmits the series of control signals (A) to the actuator (10).

[0054] The control system (40) receives a stream of sensor signals (S) from the sensor (30) in an optional receiving unit (50). The receiving unit (50) transforms the sensor signals (S) into an input signal (x). Alternatively, in the absence of a receiving unit (50), each sensor signal (S) can be directly taken as the input signal (x). The input signal (x) can be given, for example, as an extract from the sensor signals (S). Alternatively, the sensor signals (S) can be processed to generate the input signal (x). In other words, the input signal (x) is provided based on the sensor signals (S).

[0055] The input signal (x) is then passed to the classifier (60).

[0056] The classifier (60) is stored in the parameter storage device ( St 1) and the parameters provided by it ( Parameterization.

[0057] The classifier (60) determines the output signal (y) based on the input signal (x). The output signal (y) is transmitted to an optional conversion unit (80), which converts the output signal (y) into a control signal (A). The control signal (A) is then transmitted to the actuator (10) to control the actuator (10) accordingly. Alternatively, the output signal (y) can be directly used as the control signal (A).

[0058] The actuator (10) receives the control signal (A), is controlled accordingly, and performs an action corresponding to the control signal (A). The actuator (10) may include control logic that transforms the control signal (A) into another control signal, which is then used to control the actuator (10).

[0059] In another embodiment, the control system (40) may include a sensor (30). In even another embodiment, the control system (40) may alternatively or additionally include an actuator (10).

[0060] In yet another embodiment, it is conceivable that a control system (40) replaces the actuator (10) or controls the display device (10a) in addition to the actuator (10).

[0061] Furthermore, the control system (40) may include at least one processor (45) and at least one machine-readable storage medium (46) thereon storing instructions which, if executed, cause the control system (40) to perform the method according to aspects of the invention.

[0062] Figure 3 An embodiment is shown in which a control system (40) is used to control a robot that is at least partially autonomous, such as a vehicle (100) that is at least partially autonomous.

[0063] The sensor (30) may include one or more video sensors and / or one or more radar sensors and / or one or more ultrasonic sensors and / or one or more LiDAR sensors. Some or all of these sensors are preferably, but not necessarily, integrated into the vehicle (100). Thus, the input signal (x) can be understood as an input image, and the classifier (60) can be understood as an image classifier.

[0064] The image classifier (60) can be configured to detect objects near at least part of the autonomous robot based on the input image (x). The output signal (y) may include information characterizing that the object is located near at least part of the autonomous robot. A control signal (A) can then be determined based on this information, for example, to avoid collisions with the detected objects.

[0065] The actuator (10), preferably integrated in the vehicle (100), can be provided by the vehicle's (100) brakes, propulsion system, engine, drivetrain, or steering. A control signal (A) can be determined to cause the actuator (10) to be controlled so that the vehicle (100) avoids collisions with detected objects. The detected objects can also be classified according to what an image classifier (60) deems most likely to be—e.g., a pedestrian or a tree—and the control signal (A) can be determined based on this classification.

[0066] Alternatively or additionally, the control signal (A) may also be used to control the display (10a), for example, to display objects detected by the image classifier (60). It is also conceivable that if the vehicle (100) approaches and collides with at least one detected object, the control signal (A) may control the display (10a) to generate a warning signal. The warning signal may be an audible warning and / or tactile signal, such as vibration of the vehicle's steering wheel.

[0067] The factors (z1, z2, z3, z4) obtained from the input image (x) can, for example, be displayed on a display (10a). The driver of the vehicle (100) will thus be given insight into which factors (z1, z2, z3, z4) cause the image classifier (60) to determine whether an object is present or absent near the vehicle (100). If the classification is based on at least one incorrect or unwise factor (z1, z2, z3, z4), the driver can, for example, take over manual control of the vehicle (100).

[0068] In even further embodiments, factors (z1, z2, z3, z4) may also be transmitted to the vehicle operator. The operator does not necessarily have to be inside the vehicle (100), but may be located outside the vehicle (100). In this case, the factors may be transmitted to the operator by means of the vehicle's (100) wireless transmission system, such as by means of a cellular network. If the classification is based on at least one erroneous or unwise factor (z1, z2, z3, z4), the operator may, for example, take over manual control of the vehicle or issue a control signal (A) that causes the vehicle (100) to be in a safe state, such as pulling over to the side of the road on which the vehicle (100) is traveling.

[0069] In even further embodiments, the sensor (30) may be a microphone or may include a microphone. The classifier (60) can then be configured, for example, to classify audio events near the vehicle (100) and control the vehicle (100) accordingly. For example, if the classifier (60) can detect the siren of a police car or fire truck. The conversion unit (80) can then determine a control signal (A) causing the vehicle (100) to open a rescue lane on its current path.

[0070] In some embodiments, the at least partially autonomous robot may be provided by another mobile robot (not shown), which may move, for example, by flying, swimming, diving, or walking. The mobile robot may, in particular, be a lawnmower that is at least partially autonomous, or a cleaning robot that is at least partially autonomous. In these additional embodiments, a control signal (A) may be determined such that the propulsion unit and / or steering and / or braking of the mobile robot are controlled, enabling the mobile robot to avoid collisions with the identified object. In these additional embodiments, the factors (z1, z2, z3, z4) may also be transmitted to an operator via wireless transmission, such as a cellular network.

[0071] In another embodiment, at least partially autonomous robotic operation may be provided by a gardening robot (not shown) that uses sensors (30), preferably optical sensors, to determine the state of the plants in the environment (20). Actuators (10) may control nozzles for spraying liquid and / or cutting devices such as blades. Depending on the type and / or state of the plant markings, control signals (A) may be determined to cause the actuators (10) to spray the appropriate amount of liquid onto the plant and / or cut the plant.

[0072] In even further embodiments, at least partially autonomous robots may be provided by household appliances (not shown), such as washing machines, stoves, ovens, microwave ovens, or dishwashers. Sensors (30), such as optical sensors, can detect the state of the object to be processed by the household appliance. For example, in the case of a washing machine, sensor (30) can detect the state of the clothes inside the washing machine. A control signal (A) can then be determined based on the detected material of the clothes.

[0073] Figure 4 An embodiment is shown, wherein a control system (40) controls a manufacturing machine (11), such as a stamping tool, cutting tool, gun drill, or jig, of a manufacturing system (200) that is, for example, part of a production line. The manufacturing machine may include transport equipment, such as a conveyor belt or assembly line, for moving manufactured products (12). The control system (40) controls an actuator (10), which in turn controls the manufacturing machine (11).

[0074] The sensor (30) can be provided by an optical sensor that captures, for example, the characteristics of the manufactured product (12). The classifier (60) can therefore be understood as an image classifier.

[0075] The image classifier (60) can determine the position of the manufactured product (12) relative to the transport equipment. The actuator (10) can then be controlled depending on the determined position of the manufactured product (12) for subsequent manufacturing steps of the manufactured product (12). For example, the actuator (10) can be controlled to cut the manufactured product at a specific location on the manufactured product itself. Alternatively, it is conceivable that the image classifier (60) classifies whether the manufactured product is broken or exhibits defects. The actuator (10) can then be controlled to remove the manufactured product from the transport equipment.

[0076] The factors (z1, z2, z3, z4) obtained from the input image (x) can, for example, be displayed on a display (10a). The operator of the manufacturing system (200) will thus be given insight into which factors (z1, z2, z3, z4) cause the image classifier (60) to determine the output signal (y). If the classification is based on at least one incorrect or unwise factor (z1, z2, z3, z4), the operator can, for example, take over the manual control of the manufacturing system (200).

[0077] Figure 5 An embodiment is shown, in which a control system (40) controls an automated personal assistant (250). The sensor (30) may be an optical sensor, such as a video image of a user's (249) gesture. Alternatively, the sensor (30) may also be an audio sensor, such as a voice command from the user (249).

[0078] The control system (40) then determines a control signal (A) for controlling the automated personal assistant (250). The control signal (A) is determined based on sensor signals (S) from the sensor (30). The sensor signals (S) are transmitted to the control system (40). For example, a classifier (60) can be configured to, for example, implement a gesture recognition algorithm to identify gestures made by the user (249). The control system (40) can then determine the control signal (A) to be transmitted to the automated personal assistant (250). It then transmits the control signal (A) to the automated personal assistant (250).

[0079] For example, a control signal (A) can be determined based on a user gesture identified by a classifier (60). This can include information that causes an automated personal assistant (250) to retrieve information from a database and output that retrieved information in a form suitable for the user (249) to receive.

[0080] In another embodiment, it is conceivable that, instead of the automated personal assistant (250), the control system (40) controls a household appliance (not shown) controlled according to an identified user gesture. The household appliance may be a washing machine, stove, oven, microwave oven, or dishwasher.

[0081] Figure 6 An embodiment is shown, in which the control system (40) controls the access control system (300). The access control system (300) can be designed to physically control access. For example, it can include a door (401). The sensor (30) can be configured to detect scenarios related to deciding whether access is permitted. For example, it can be an optical sensor for providing image or video data, such as for detecting a person's face. The classifier (60) can therefore be understood as an image classifier.

[0082] An image classifier (60) can be configured to classify a person's identity, for example, by matching the detected person's face with other faces of known people stored in a database. A control signal (A) can then be determined based on the classification by the image classifier (60), for example, according to the determined identity. An actuator (10) can be a lock that opens or closes a door depending on the control signal (A). Alternatively, the access control system (300) can be a non-physical, logical access control system. In this case, the control signal can be used to control a display (10a) to show information about the person's identity and / or whether the person has been granted access.

[0083] The factors (z1, z2, z3, z4) obtained for the input image (x) can be displayed, for example, on a display (10a). Thus, the operator of the access control system (300) will be given insight into which factors (z1, z2, z3, z4) cause the image classifier (60) to determine the output signal (y). If the classification is based on at least one incorrect or unwise factor (z1, z2, z3, z4), the operator can, for example, deny access.

[0084] Figure 7 An embodiment is shown, wherein the control system (40) controls the monitoring system (400). This embodiment is largely consistent with... Figure 6 The embodiments shown are identical. Therefore, only the differences will be described in detail. The sensor (30) is configured to detect the monitored scene. The control system (40) does not necessarily control the actuator (10), but may instead control the display (10a). For example, the image classifier (60) may determine the classification of the scene, for example, whether the scene detected by the optical sensor (30) is normal or whether the scene exhibits anomalies. The control signal (A) transmitted to the display (10a) may then be configured, for example, to cause the display (10a) to adjust the displayed content based on the determined classification, for example, highlighting objects that the image classifier (60) considers abnormal.

[0085] Figure 8 An embodiment of a medical imaging system (500) controlled by a control system (40) is shown. The imaging system may be, for example, an MRI device, an X-ray imaging device, or an ultrasound imaging device. The sensor (30) may be, for example, an imaging sensor that captures at least one image of the patient, thereby displaying, for example, different types of the patient's body tissues.

[0086] The classifier (60) can then determine the classification of at least a portion of the sensed image. Thus, at least a portion of the image is used as the input image (x) to the classifier (60). The classifier (60) can therefore be understood as an image classifier.

[0087] The control signal (A) can then be selected based on the classification, thereby controlling the display (10a). For example, the image classifier (60) can be configured to detect different types of tissue in the sensed image, for example, by classifying the tissue displayed in the image as malignant or benign tissue. This can be accomplished by semantic segmentation of the input image (x) by the image classifier (60). The control signal (A) can then be determined to cause the display (10a) to display different tissues, for example, by displaying the input image (x) and coloring different regions of the same tissue type with the same color.

[0088] The factors (z1, z2, z3, z4) obtained for the input image (x) can be displayed, for example, on a display (10a). The operator of the imaging system (500) will thus be given insight into which factors (z1, z2, z3, z4) cause the image classifier (60) to determine the output signal (y). If the classification is based on at least one incorrect or unwise factor (z1, z2, z3, z4), the operator can, for example, further examine the input image (x).

[0089] In another embodiment (not shown), the imaging system (500) can be used for non-medical purposes, such as determining the material properties of a workpiece. In these embodiments, an image classifier (60) can be configured to receive an input image (x) of at least a portion of the workpiece and perform semantic segmentation of the input image (x) to classify the material properties of the workpiece. A control signal (A) can then be determined to cause a display (10a) to show the input image (x) and information about the detected material properties.

[0090] Figure 9 An embodiment of a medical analysis system (600) controlled by a control system (40) is shown. The medical analysis system (600) is equipped with a microarray (601), wherein the microarray includes multiple points (602, also referred to as features) that have been exposed to a medical sample. The medical sample may be, for example, a human sample or an animal sample obtained from a swab.

[0091] The microarray (601) can be a DNA microarray or a protein microarray.

[0092] The sensor (30) is configured to sense a microarray (601). The sensor (30) is preferably an optical sensor, such as a video sensor. The classifier (60) can therefore be understood as an image classifier.

[0093] The image classifier (60) is configured to classify the results of the sample based on the input image (x) of the microarray supplied by the sensor (30). In particular, the image classifier (60) can be configured to determine whether the microarray (601) indicates the presence of a virus in the sample.

[0094] Then, a control signal (A) can be selected to make the display (10a) show the classification results.

[0095] The factors (z1, z2, z3, z4) obtained for the input image (x) can be displayed, for example, on a display (10a). Thus, the operator of the analysis system (600) will be given insight into which factors (z1, z2, z3, z4) cause the image classifier (60) to determine the output signal (y). If the classification is based on at least one incorrect or unwise factor (z1, z2, z3, z4), the operator can, for example, further examine the sample or issue another test or another swab.

[0096] Figure 10 An embodiment of a training system (140) for training a classifier (60) of a control system (40) using a training dataset (T) is shown. The training dataset (T) includes multiple input signals (x) used to train the classifier (60). i ), wherein the training dataset (T) further includes, for each input signal (x) i ), corresponding to the input signal (x) i ) and characterize the input signal (x) i The expected output signal (y) of the classification i ).

[0097] For training purposes, the training data unit (150) accesses a computer-implemented database (St2) that provides a training dataset (T). The training data unit (150) preferably randomly determines at least one input signal (x) from the training dataset (T). i ) and corresponding to the input signal (x) i The expected output signal (y) i ), and the input signal (x) i The signal is transmitted to the classifier (60). The classifier (60) is based on the input signal (x) and transmits it to the classifier (60). i Determine the output signal (y) i ).

[0098] Desired output signal (y) i ) and a defined output signal ( ) is transmitted to the modification unit (180).

[0099] Based on the desired output signal (y) i ) and a defined output signal ( ), modify unit (180) and then determine the new parameters for classifier (60) ( For this purpose, the modification unit (180) uses a loss function to reduce the desired output signal (y). i) and a defined output signal ( The loss function is used to determine the output signal that has been determined. ) and the desired output signal (y i The first loss value is the value that deviates from the given loss function. In the given embodiment, the negative log-likelihood function is used as the loss function. In alternative embodiments, other loss functions are also conceivable.

[0100] Furthermore, it is conceivable that a definite output signal ( ) and the desired output signal (y i Each of these components comprises multiple sub-signals, for example, in the form of a tensor, where the desired output signal (y) is... i The sub-signal of ) corresponds to a defined output signal ( The first sub-signal represents the object relative to the input signal (x). For example, it is conceivable that the classifier (60) is configured for object detection, and the first sub-signal represents the object relative to the input signal (x). i The probability of occurrence of a portion of the signal is given by the first sub-signal, and the second sub-signal characterizes the exact location of the object. If the determined output signal ( ) and the desired output signal (y i If the signal includes multiple corresponding sub-signals, then a second loss value is preferably determined for each corresponding sub-signal by means of a suitable loss function, and the determined second loss value is appropriately combined, for example by means of weighted sum, to form a first loss value.

[0101] The modification unit (180) determines new parameters based on the first loss value. In the given embodiment, this is accomplished using a gradient descent method, preferably stochastic gradient descent, Adam, or AdamW.

[0102] In other preferred embodiments, the described training iteratively repeats a predefined number of iteration steps, or iteratively repeats until a first loss value falls below a predefined threshold. Alternatively or additionally, it is conceivable that training terminates when the average first loss value relative to the test or validation dataset falls below a predefined threshold. In at least one iteration, new parameters (determined in the previous iteration) are used... ) is used as a parameter of classifier (60) ).

[0103] In even further embodiments, the classifier is trained based on a neural architecture search algorithm, wherein at least several functions (F1, F2, F3, F4) of the reversible factorization model (61) of the classifier (60) are determined based on the neural architecture search algorithm. In this embodiment, the LEMONADE algorithm is used as the neural architecture search algorithm. However, other neural architecture search algorithms are conceivable in other embodiments. The DCI deentanglement and classification accuracy of the image classifier (60) can be selected as the target of the LEMONADE algorithm.

[0104] Furthermore, the training system (140) may include at least one processor (145) and at least one machine-readable storage medium (146) containing instructions that, when executed by the processor (145), cause the training system (140) to perform a training method according to one aspect of the invention.

[0105] The term "computer" can be understood to encompass any device used to process predefined computational rules. These computational rules can be in the form of software, hardware, or a combination of both.

[0106] Generally, "multiple" can be understood as being indexed, that is, preferably by assigning consecutive integers to the elements contained in the multiple, with each element in the multiple being assigned a unique index. Preferably, if the multiple has N There are elements, among which N If the number of elements in a plurality is a given number, then the elements are assigned a number from 1 to 1. N Integers. It can also be understood that multiple elements can be accessed through their indices.

Claims

1. A computer-implemented method for determining an output signal (y) of an input signal (x) using a classifier (60), wherein the output signal (y) characterizes the classification of the input signal (x) and wherein the input signal includes an image or audio signal, the method comprising the following steps: The latent representation (z) is determined based on the input signal (x) by means of a reversible factorization model (61) included in the classifier, wherein the latent representation (z) comprises multiple factors (z1, z2, z3, z4), and wherein the reversible factorization model (61) is characterized by: Multiple functions (F1, F2, F3, F4), wherein each of the multiple functions (F1, F2, F3, F4) is continuous and almost everywhere continuously differentiable, wherein the functions (F1, F2, F3, F4) are further configured to accept an input signal (x) or at least one factor (z1, z2, z3, z4) provided by another function (F1, F2, F3, F4) as input, and wherein the functions (F1, F2, F3, F4) are further configured to provide at least one factor (z1, z2, z3, z4), wherein the at least one factor (z1, z2, z3, z4) is either provided as a potential... At least a portion of the representation (z) is provided as at least a portion of the input to another function (F1, F2, F3, F4) of the plurality of functions (F1, F2, F3, F4) of the invertible factorization model (61), wherein each function (F1, F2, F3, F4) of the invertible factorization model (61) is invertible, wherein there exists a continuous inverse function for each function (F1, F2, F3, F4), wherein each inverse function is continuous, almost everywhere continuously differentiable, and is configured to determine the input of each function (F1, F2, F3, F4) based on the at least one factor (z1, z2, z3, z4) provided from the function (F1, F2, F3, F4); The output signal (y) is determined based on the latent representation (z) by means of an internal classifier (62) included in the classifier (60); If the input signal includes an image, the output signal represents the single class to which the classifier considers the input signal to belong, and the output signal represents the classification of objects in the image and the location of the objects. as well as If the input signal includes an audio signal, the output signal will assign the audio signal to the class that the classifier believes the audio signal belongs to.

2. The method of claim 1, wherein the plurality of functions (F1, F2, F3, F4) comprises at least one function (F1, F2, F3, F4), the at least one function (F1, F2, F3, F4) providing a first factor (z1, z2, z3, z4) to the latent representation (z), and providing a second factor (z1, z2, z3, z4) to another function (F1, F2, F3, F4) from the plurality of functions (F1, F2, F3, F4).

3. The method according to claim 1 or 2, wherein the classifier (60) is based on a determined output signal ( ) and the desired output signal ( y i The first difference between ) is trained, where the determined output signal ( ) is for the training input signal ( x i The output signal is determined by and is expected to be ( ). y i ) Represents the training input signal ( x i ) expected classification.

4. The method according to any one of the preceding claims, wherein the plurality of functions (F1, F2, F3, F4) comprises at least one function (F1, F2, F3, F4), said at least one function (F1, F2, F3, F4) being configured to provide two factors according to the following formula. , in It is the input of the function (F1, F2, F3, F4). z 1 is the first factor. z 2 is the second factor. It is an encoder for an automatic encoder, and It is the decoder of the autoencoder, and further within it, The inverse function corresponding to the function is given by the following formula. 。 5. The method according to claim 4, wherein, The encoder is trained based on the gradients of multiple encoder parameters relative to a first difference.

6. The method of claim 5, wherein the decoder is trained based on the gradients of a plurality of parameters of the decoder relative to the second difference. , in It is the input of the function, and the summation is over all the squared elements of the subtraction.

7. The method according to any one of the preceding claims, wherein the plurality of functions (F1, F2, F3, F4) comprises at least one function (F1, F2, F3, F4), said at least one function (F1, F2, F3, F4) being configured to provide three factors according to the following formula. , in It is the input of the function (F1, F2, F3, F4). z 1 is the first factor and the sum of the inputs. The result of application group standardization z 2 is the expected value of the second factor and the group standardization. z 3 is the third factor, and the group standardization and the standard deviation of the group standardization further depend on the scaling parameter. and shift parameters The inverse of the function is given by the following equation. 。 8. The method according to any one of the preceding claims, wherein the plurality of functions (F1, F2, F3, F4) comprises at least one function (F1, F2, F3, F4), the at least one function (F1, F2, F3, F4) providing two factors according to the following formula. , in It is the input of the function (F1, F2, F3, F4). z 1 is the first factor. z 2 is the second factor, and ReLU It is the linear unit of rectification, and further within it, The inverse function of (F1, F2, F3, F4) is given by the following equation. 。 9. The method according to any one of the preceding claims, wherein, The internal classifier (62) is a linear classifier.

10. The method according to any one of claims 3 to 9, wherein, Training the classifier (60) further includes training a reversible factorization model (61) by means of a neural architecture search algorithm, wherein the objective function of the neural architecture search algorithm is based on the disentanglement of multiple factors (z1, z2, z3, z4) of the latent representation (z).

11. A computer program product comprising a computer program that, when executed by a processor (45, 145), implements the steps of the method according to any one of claims 1-10.

12. A machine-readable storage medium (46, 146) having a computer program stored thereon, which, when executed by a processor (45, 145), implements the steps of the method according to any one of claims 1-10.