Image classification in the sparse coding model by increasing the number of out-layer neurons

A deep hierarchical sparse-coding model with feedback connections and iterative adjustments addresses image noise and incompleteness issues, enhancing classification accuracy and reconstruction, enabling robust vehicle control through gesture recognition and real-time adaptation.

DE102025105589B3Active Publication Date: 2026-04-02MERCEDES BENZ GROUP AG
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing image classification models in vehicles suffer from decreased accuracy due to image noise and incomplete input images caused by poor lighting, obscured areas, and structural limitations, leading to missing or distorted visual information, which negatively impacts user experience in non-safety-related systems.

Method used

A deep hierarchical sparse-coding model with feedback connections and iterative sparse-coding algorithm that iteratively adjusts neuron states and applies a cost function to improve classification accuracy and reconstruct missing information, using a pre-trained model with at least three layers and feedback connections to enhance feature recognition and reconstruction quality.

Benefits of technology

The model effectively processes incomplete and noisy images, providing robust and accurate classification and reconstruction, enabling intuitive vehicle control through gesture recognition even in challenging conditions, and allowing for real-time adaptation to new gestures without distracting the driver.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for machine image classification, wherein an input image (1) generated from camera data is fed to a pre-trained sparse-coding model with layers (2) with a feedback connection (4) (S1), wherein an iterative sparse-coding algorithm is executed in the layers (2) (S3), and an inference of the model is repeatedly performed (S4), and in each of the repetitions a state value of a respective neuron (5) of the output layer (3) is artificially increased (S5) in order to iteratively simulate a result of the inference as membership in a respective class, and wherein for each artificial increase a cost function is formed based on a reconstruction quality of the input image (1) (S6), and the input image (1) is assigned to the class (S7) whose assigned neuron (5) of an output layer (3) leads to the lowest value of the cost function when artificially increased.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for machine image classification, a system for machine image classification, and a method for generating a model for machine image classification.

[0002] The problem of perceiving and classifying specific objects or actions using camera images arises in many areas. For example, in the recognition of traffic signs or other vehicles using an exterior camera, or in the recognition of an occupant's gestures using an interior camera. Robust gesture recognition for the driver and other occupants enables simple and intuitive control of infotainment systems, interior lighting, temperature, and other vehicle functions. The driver can thus adjust settings with simple gestures without taking their eyes off the road.

[0003] Many approaches, such as convolutional neural networks, have already been developed to effectively solve these problems, at least under optimal conditions. However, the classification accuracy of these models decreases rapidly when image noise occurs, for example, due to poor lighting conditions, or when obscured areas are present, for example, due to camera position, dirty lenses, or structural limitations of the vehicle. Image noise from a vehicle's interior camera can lead to an incomplete input image, but blind spots or obscured areas caused by camera position or structural limitations of the vehicle can also result in an incomplete input image for machine image recognition. This leads to important visual information being missing or distorted. See, in particular, the publication by A. Von Bernuth, G. Volk, and O.Bringmann, Augmenting Image Data Sets With Water Spray Caused by Vehicles on Wet Roads,"2021 IEEE International Intelligent Transportation Systems Conference (ITSC), Indianapolis, IN, USA, 2021, pp. 3055-3060, doi: 10.1109 / ITSC48978.2021.9564826.

[0004] Since the limitations described above can frequently occur in real-world applications, it is crucial to have an image classification algorithm that can handle these limitations. A further challenge lies in developing a method that, under these conditions, not only correctly classifies images and accurately recognizes gestures or traffic signs, but can also use this classification information to reconstruct missing information in the incomplete input image as realistically as possible, ensuring a complete and clear visual representation and thus compensating for minor errors. The visual representation of the interior can, for example, be used to display the driver's or passengers' current gestures on a screen.

[0005] This requirement is particularly relevant for non-safety-related systems in vehicles, such as entertainment or information systems, as faulty gesture recognition or poor image quality can negatively impact the user experience. Therefore, the use of image processing algorithms capable of overcoming these challenges without compromising system performance or generating incorrect content is highly desirable.

[0006] In this context, US 2024 / 135604 A1 concerns an image reconstruction system comprising: an input interface for receiving image data; a processor; and a memory in which instructions are stored that cause the processor to reconstruct an image from the image data using a self-monitoring deep learning model.

[0007] US 2023 / 0107097 A1 further relates to a method for identifying a gesture from one of several dynamic gestures, wherein each dynamic gesture comprises a unique movement performed by a user over a specific period of time within the field of view of an image-capturing device, wherein the method iteratively: captures a current image from the image-capturing device at a given time; and passes at least a portion of the current image through a bidirectionally iterating multilayer classifier. A final layer of the multilayer classifier includes an output indicating a probability that a gesture from the plurality of dynamic gestures is performed by a user during the time the image is captured.

[0008] US patent 2016 / 0335487 A1 discloses a method for detecting hand movements by localizing and tracking the hand in a video. Feature points are identified by extracting RGB and depth data, represented using a 3D mesh, and then compared with pre-trained examples to determine a hand movement category.

[0009] CN 107133361 A relates to a method and a device for gesture recognition, wherein a recognized gesture is compared with a gesture data set.

[0010] US 2025 / 0005916 A1 relates to a computer-implemented method for tuning a pre-trained machine learning network using few-shot image learning.

[0011] US 2021 / 0158168 A1 discloses a neural network for processing input data.

[0012] The object of the invention is to improve machine-based image classification when used in the operation of a vehicle.

[0013] The invention is defined by the features of the independent claims. Advantageous further developments and embodiments are the subject of the dependent claims.

[0014] A first aspect of the invention relates to a method for machine image classification, wherein an input image generated from camera data is fed to a pre-trained sparse-coding model with at least three layers, an output layer of which has a plurality of neurons, each of which is assigned a respective class, wherein at least one pair of two adjacent layers has a measure of feature recognition quality of a subsequent layer of the pair fed back to a preceding layer of the pair via a feedback connection and used for adjustment in the preceding layer, wherein an iterative sparse-coding algorithm is executed in each of the at least three layers, characterized in that an inference of the pre-trained sparse-coding model with the input image is repeatedly performed and in each of the repetitions a state value of a respective neuron of the output layer is artificially incremented.to iteratively simulate an inference result as membership in a respective class, wherein the respective artificial increase is propagated from the output layer through the layers via the feedback connections, and wherein for each artificial increase a cost function is formed based on a reconstruction quality of the input image and / or on the basis of a sparsity of the sparse-coding model, and the input image is assigned to the class whose assigned neuron of the output layer leads to the lowest value of the cost function when artificially increased.

[0015] The measure of feature recognition accuracy correlates inversely with the reconstruction error of an incomplete input image. "Incomplete" in this context means an image that, due to image noise, obscured image areas, or other image defects, does not clearly reveal its content; that is, information is lost compared to reality when generating the input image.

[0016] The artificial increase of a given neuron in the output layer corresponds to an artificial numerical increase of the state value of the respective neuron, in other words an artificial increase of the activation state of this respective neuron, for example by incrementing a current state value by one.

[0017] The deep hierarchical sparse-coding model is a multi-layered model with feedback connections. The input image to be classified serves as the input to the first layer. For each subsequent layer, the output of a neighboring preceding layer serves as the input to the next. The layers themselves can have various architectures, for example, a convolutional neural network (CNN) or another deep neural network.

[0018] The deep hierarchical sparse-coding model thus executes, in particular, a generative image optimization algorithm. Using at least three layers and their feedback connections, it allows input images to be improved by capturing deeper structures and features in the input image data. Each of the three layers extracts different features from the input image. The result is the ability to learn comprehensive abstractions of the visual data in deeper layers, while the earlier layers specialize in learning local relationships. The integration of a feedback connection within the model enables adaptive adjustment of the learning processes by dynamically adjusting the bias intensity in each layer based on the feedback signal from the deeper layers.

[0019] The feedback loop allows the algorithm to use the output of a layer to improve the input for a previous layer. This helps increase the accuracy and quality of image optimization by correcting errors and refining feature extraction. Inhibitory feedback loops exist between at least two adjacent layers, and preferably between all adjacent layers. These loops allow higher layers to influence lower layers. The strength of the inhibitory feedback loop is determined by the reconstruction error of the respective layer. With low feedback from lower layers, less is generated, resulting in more precise and context-aware image reconstructions. Furthermore, a hyperparameter can be provided to define the strength of each feedback loop.In practical application, this parameter determines how much new image content is generated or analyzed in a previous layer. Adjusting this hyperparameter allows for the desired result to be achieved for each specific application. Furthermore, the hyperparameter can be advantageously adjusted based on the measured image quality or by the user.

[0020] Machine-based image classification involves inference from a pre-trained sparse-coding model. One approach is to perform the inference normally using a current input image and to classify the image based on the highest activation state of a neuron in the output layer. This method assumes that the neuron with the highest activation best matches the actual class. However, to further improve classification accuracy, a probing procedure is employed, allowing for iterative improvement of the classification. This procedure is particularly inspired by concepts from Hopfield networks and energy-based models. In Hopfield networks, a cost function is defined, the minimization of which stabilizes the network and enables correct pattern recognition.Similarly, the sparse-coding model presented here is based on minimizing a cost function, with the feedback connection of the lower layers in the network acting as an attractor to a similar class. In the probing procedure, each possible class of the image to be classified is tested stepwise by artificially increasing the number of neurons in the initial layer. This means that the internal state of the neuron in the last layer is amplified accordingly for each class, and the activations of the layers are adjusted depending on the chosen class. For each class, a reconstruction error or an alternative metric, such as sparsity in the sparse-coding model or the maximum activation in the last layer, is calculated.The cost function describes the quality of reconstructing incomplete images and thus the classification accuracy regarding the agreement between the model's result and the respective actual class. The cost function can be defined in various ways, for example, based on a reconstruction error. In this case, the activations of the output layer are preferably projected back onto the image domain using basis functions and compared with the input image. Alternatively, the sparsity of the activations can be used as a criterion to identify the class that leads to the most economical representation.

[0021] The use of the aforementioned cost function shows parallels to other known energy-based models, as it aims to find a configuration of activations that minimizes the cost function in terms of an energy function, thereby identifying a correct class of the input image. By iteratively applying the stepwise artificial increase of neurons in the output layer for each possible class, the model's internal state is continuously optimized until the class with the lowest reconstruction error and / or the highest neuronal sparsity is found. In addition to improved classification accuracy, this artificial increase of the class deemed correct also leads to better reconstruction quality of the input image.Since the sparse-coding model is designed to learn a parsimonious and robust representation of the input image, identifying the correct class by applying the cost function leads to a finer fit of the internal representations. This further refines the reconstruction of the input image, as the model is able to minimize its internal energy toward the optimal solution, similar to Hopfield networks, which achieve stable pattern recognition by minimizing an energy function. Finally, this unsupervised procedure improves the adaptability of the sparse-coding model by implementing an unsupervised learning strategy that does not require explicitly annotated data but can identify new classes and tasks through iterative minimization of the cost function.This enables continuous adaptation and optimization of the sparse coding model based on the input images, making it a robust solution for the classification and reconstruction of image data in real-world applications.

[0022] The described sparse-coding model is advantageously able to efficiently process incomplete input images with high-frequency, noisy signals. This is achieved through its hierarchical structure and a mechanism for artificially increasing the state values ​​of the respective neurons in the output layer, enabling the recognition of robust, recurring patterns while effectively suppressing noise components. The model's ability to analyze such dynamic data formats in real time allows for the reliable detection of relevant gestures and movements of the driver or passengers when used for gesture recognition, even in situations where the signal is overlaid with significant background noise. This makes the model particularly suitable for integration into vehicle architectures that rely on neuromorphic sensors to ensure precise, fast, and interference-resistant detection of user interactions.

[0023] The deep hierarchical sparse-coding model described here integrates seamlessly into the vehicle's internal control system, providing the driver with a user-friendly interface for interacting with the system and configuring various functions. An interface is advantageously designed to offer intuitive controls without compromising driving safety. Simultaneously, it ideally allows for in-depth customization of system functions to the driver's individual preferences and enables the driver to configure various gesture control options via the interface. These controls are preferably available in an interactive menu displayed on the vehicle's infotainment screen or accessible via voice control. The driver can thus select from a variety of predefined gestures assigned to specific functions within the vehicle.Examples include: controlling the interior lighting through specific hand gestures, activating or deactivating infotainment systems through simple finger movements, adjusting the temperature control in the vehicle through swipe gestures, and opening and closing windows by means of circular hand movements.

[0024] Additionally, such a user interface offers the advantage of creating and configuring custom gestures. In this mode, a new gesture is defined and recorded by the driver. The sparse-coding model processes this gesture, and the output layer is adjusted accordingly to learn it. These custom gestures can be assigned to any vehicle function, thus providing the system with maximum flexibility.

[0025] According to an advantageous embodiment, the iterative sparse coding algorithm is a locally competitive algorithm.

[0026] In each of the at least three layers, an iterative sparse coding algorithm is executed, preferably a locally competitive algorithm (LCA). Unlike well-known artificial neural networks, each of the at least three layers is not limited to elementary linear algebraic calculations such as scalar products, but also exhibits dynamic relationships that are ideally described by differential equations. After several iterations of a layer's sparse coding algorithm, the activations of the neurons no longer change, or at least not significantly; if this is the case in all layers, the sparse coding model has converged to an equilibrium. Thus, each layer attempts to find a sparse representation of its input. However, the layers typically also include a multitude of artificial neurons, which, with their state-becoming, i.e.,Their activation state determines the output size of each layer.

[0027] According to another advantageous embodiment, the camera data has a neuromorphic data format.

[0028] The deep hierarchical sparse-coding model, with its mechanism of incrementally artificially increasing the number of neurons in an output layer, is particularly well-suited for processing neuromorphic data formats, such as those generated by the output of an event camera. These event cameras provide extremely high-resolution temporal data by capturing changes in lighting or motion as discrete events at the pixel level, rather than generating complete image frames like conventional cameras. Due to the nature of this data, which is generated at very short intervals, the output of an event camera exhibits exceptionally fine-grained temporal resolution, but is simultaneously susceptible to noise and irrelevant information.

[0029] According to a further advantageous embodiment, an input image is generated from a gesture of a vehicle occupant, whereby it is recognized when a new gesture, not yet assigned to a class, is present, and a further neuron is added to the output layer that represents a class assigned to the new gesture.

[0030] This offers the advantage of configuring the cost function mechanism directly via a user interface. The driver can activate this mechanism to learn new gestures unsupervised. The system automatically compares the gestures to all known gestures and analyzes whether a new gesture is recognized by iteratively incrementing and checking the class. If a new gesture is identified, the driver can confirm its classification and assign the corresponding vehicle functionality. This process runs fully automatically in the background, allowing the driver to configure the system without distraction while driving. In the event of a conflict where the new gesture overrides an existing one, the driver is alerted, preferably by an acoustic or visual signal, and can select an alternative assignment.

[0031] According to another advantageous embodiment, at least one of the following is adaptable by a user: - The number of iterations of the iterative sparse coding algorithm; - A degree of sparsity that is relevant for classification; - A selection of different metrics to determine the classification, in particular a maximum activation or a minimum reconstruction error; - A learning rate of the sparse coding model to set a reaction speed to input images that represent a new gesture;

[0032] The vehicle's occupant benefits from a user interface offering advanced settings. This allows for dynamic adjustments to the behavior of the sparse-coding model. In addition to these customization options, a continuous learning function is advantageously supported, where the vehicle continuously analyzes and stores new patterns and movements of the driver in the background. This learning function can be conveniently activated or deactivated via the user interface. When activated, the system continuously analyzes hand movements and gestures within the vehicle, adjusts its internal parameters, and optimizes gesture recognition without any further input from the occupant.

[0033] According to a further advantageous embodiment, the artificial enhancements of the respective neurons of the output layer are carried out in parallel, and the feedback connections are calculated in parallel for each artificial enhancement.

[0034] The application of the cost function can be efficiently parallelized by simultaneously calculating a feedback signal for each possible class. The reconstruction process for each class runs independently, allowing the iterations for each class to be parallelized. This accelerates the inference process. Therefore, an iterative update step is executed in parallel for each class independently. At the end of this process, the class with the minimum reconstruction error is selected. By parallelizing the update steps for all classes, the inference time is significantly reduced, as all classes can be processed concurrently.

[0035] Another aspect of the invention relates to a system for machine image classification, comprising a camera and a computing unit, wherein the computing unit is configured to feed an input image generated with data from the camera to a pre-trained sparse-coding model with at least three layers, an output layer of which has a plurality of neurons to which a respective class is assigned, wherein at least one pair of two adjacent layers has a feedback connection for feeding back a measure of feature recognition quality from a subsequent layer of the pair to a preceding layer of the pair, wherein the measure serves for adjustment in the preceding layer, wherein the computing unit is configured to execute an iterative sparse-coding algorithm in each of the at least three layers, characterized in that the computing unit is configured toto repeatedly perform an inference of the pretrained sparse-coding model with the input image and in each repetition artificially increment a state value of a respective neuron in the output layer in order to iteratively simulate a result of the inference as membership in a specific class, wherein the computing unit is designed to propagate the respective artificial increment from the output layer through the layers via the feedback connections, and for each increment to form a cost function based on a reconstruction quality of the input image and / or on a sparsity of the sparse-coding model, and to assign the input image to the class whose assigned neuron in the output layer leads to the lowest value of the cost function when artificially incremented.

[0036] Advantages and preferred further developments of the proposed system result from an analogous and meaningful transfer of the above statements made in connection with the proposed method for machine image classification.

[0037] Another aspect of the invention relates to a method for generating a model for machine image classification, wherein the model is a sparse-coding model having at least three layers, an initial layer having a plurality of neurons to which a respective class is assigned, and a feedback connection for feeding a measure of feature recognition quality from a subsequent layer of the pair back to a preceding layer of the pair and for adaptation in the preceding layer, wherein the model is trained by specifying an associated class for an input image, artificially increasing the value of a neuron in the initial layer belonging to the specified class, and wherein the artificial increase is propagated stepwise from the initial layer through the layers via the feedback connections.and wherein parameters in each of the at least three layers are adjusted according to the specification.

[0038] During training, artificially increasing the state value of a neuron in an output layer, representing the correct class, is used to adjust the layer parameters individually for each layer, as the layers themselves can form complex submodels. The artificial neurons of the last layer, i.e., the output layer, of the model represent specific classes, such as different images of traffic signs or user gestures. Accordingly, during training, the activation of the neuron in the last layer, which is intended to represent the class of the training image, is increased. For example, if the input image during training depicts a waving gesture, a positive value is added to the activation of the neuron that is supposed to represent the class of waving gestures.

[0039] Artificially increasing a state value, and thus the activation of a neuron in the output layer belonging to the correct class, is necessary during the training process of the sparse-coding model for the cost function to function correctly during inference. The cost function is used in inference to reliably obtain a correct classification. To this end, hypotheses about the possible class are formulated and evaluated to determine which best fits the classification pattern. Specifically, this means that the activation of a neuron in the output layer is first increased, allowing the sparse-coding model to reach a new equilibrium. Subsequently, a cost function for the sparse-coding model is determined, which can be an energy function, and, for example,The reconstruction quality of the input image or the sparsity of the sparse-coding model is considered, where sparsity indicates the number of all activations of the sparse-coding model that are zero. This is then iteratively repeated for each neuron in the output layer. The class of neuron that results in the lowest cost function is the resulting classification.

[0040] The sparse-coding model is thus trained in a partially supervised manner. During training, an input image is processed by the sparse-coding model. The iterations of the sparse-coding algorithm for each layer run concurrently. The innovation of this sparse-coding model is that, during training, the internal state of the neuron corresponding to the class of the given input image is artificially incremented. After the sparse-coding algorithm of each layer has converged, the weights of the neurons in that layer are adjusted using gradient descent. Due to the artificial increments applied to the neurons in the initial layer, each neuron in the initial layer ultimately represents a class at the end of this training process.

[0041] According to a further advantageous embodiment, parameters of a single layer are adjusted for each of the at least three layers by a separate back-propagation process.

[0042] According to a further advantageous embodiment, an extraction of features from preceding layers is used for feature classification in at least one subsequent layer further towards the initial layer.

[0043] The deep hierarchical sparse-coding model can therefore also be used as a feature extractor (for transfer learning), similar to a conventional convolutional neural network (CNN), to enable fast and efficient training of new classes. In this mode, the hierarchical structure of the model is used to extract features from preceding layers. This feature extraction serves as the basis for classification in one or more subsequent layers. Within the framework of transfer learning, only the fully connected initial layer of the hierarchical structure is retrained to perform the classification based on the extracted features. The features extracted through sparse coding in the preceding layers are generalizable representations that can be reused for various tasks.The initial layer is subsequently trained for new input images, such as new gestures, by using a small number of training examples to learn new classes. This is efficient because only the neurons of the initial layer are updated, while the preceding layers act as fixed feature extractors. After training the initial layer, the sparse-coding model can again recognize gestures through inference and the cost function by considering the new features of an input image as a subset of the classification space. This approach allows the driver to define new gestures in real time. These gestures are captured by the system, and their internal representation by the sparse-coding model is extracted. Then, only the initial layer is replaced. This results in an efficient, adaptive system that autonomously integrates and expands new gestures and functions with minimal driver interaction.

[0044] Advantages and preferred further developments of the proposed method for generating a model for machine image classification result from an analogous and meaningful transfer of the above statements made in connection with the proposed method for machine image classification.

[0045] Further advantages, features and details will become apparent from the following description, in which - possibly with reference to the drawing - at least one embodiment is described in detail.

[0046] They show: Fig. 1: A method for machine image classification according to an embodiment of the invention. Fig. 2: An architecture of a process of Fig. 1 used sparse coding model according to an embodiment of the invention.

[0047] Fig. Figure 1 shows a method for machine image classification, wherein an input image 1 generated from camera data is fed to a pre-trained sparse-coding model with at least three layers 2 S1, an output layer 3 of which has a plurality of neurons 5, each of which is assigned a respective class, wherein at least one pair of two adjacent layers 2 has a measure of feature recognition quality of a subsequent layer 2 of the pair fed back to a preceding layer 2 of the pair via a feedback connection 4 S2 and used for adaptation in the preceding layer 2, wherein an iterative sparse-coding algorithm is executed in each of the at least three layers 2 S3, characterized in that an inference of the pre-trained sparse-coding model with the input image 1 is repeatedly performed S4 and in each of the repetitions a state value of a respective neuron 5 of the output layer 3 is artificially increased S5.to iteratively simulate an inference result as membership in a respective class, wherein the respective artificial increase is propagated from the output layer 3 through the layers 2 via the feedback connections 4, and wherein for each artificial increase a cost function is formed based on a reconstruction quality of the input image 1 and / or on a sparsity of the sparse-coding model S6, and the input image 1 is assigned to the class S7 whose assigned neuron 5 of the output layer 3 leads to the lowest value of the cost function when artificially increased. The sequence of the procedural steps can vary and is not necessarily bound to the order shown; some of the steps may occur simultaneously.

[0048] Fig. Figure 2 shows a sparse-coding model with several layers, of which the outermost layer opposite an input layer is called the output layer. This output layer has a large number of neurons (of which, for clarity, the following are shown in the Fig.2 (only one of which is designated as such), to which a respective class is assigned. Furthermore, a feedback connection 4 is provided at each pair of two adjacent layers 2 for feeding back a measure of the feature recognition quality of a subsequent layer 2 of a respective pair to a preceding layer 2 of the respective pair. This measure serves to adjust the calculations of the preceding layer 2. The feedback connection to a layer 2 is determined by a reconstruction error in the next higher layer 2. The strength of the feedback connection can be adjusted with a hyperparameter. This increases the internal state of a neuron with a corresponding class of the training example. An iterative locally competitive algorithm is executed in each of the layers 2. The first of the layers 2, the input layer, is fed an input image or information based on the input image.Depending on the application, an interior or exterior camera in the vehicle continuously captures image sequences, which are transmitted via an interface to a processing unit within the vehicle. This processing unit executes the sparse-coding model and analyzes the image data to recognize, for example, traffic signs (in the case of the exterior camera) or gestures (in the case of the interior camera) made by the driver and passengers. The resulting recognitions are then fed into the vehicle's central control system, which triggers various functions based on this information. In the case of image data from the interior camera, this could include, for example, activating controls, adjusting the interior lighting, controlling infotainment systems, or updating visualizations on an infotainment display.

[0049] Although the invention has been further illustrated and explained in detail by means of preferred embodiments, the invention is not limited by the disclosed examples, and other variations can be derived from them by a person skilled in the art without departing from the scope of protection of the invention. It is therefore clear that a multitude of possible variations exist. It is also clear that the embodiments mentioned as examples are truly only examples and are not to be understood in any way as limiting, for example, the scope of protection, the possible applications, or the configuration of the invention.Rather, the preceding description and the description of the figures enable the person skilled in the art to implement the exemplary embodiments in concrete terms, whereby the person skilled in the art, with knowledge of the disclosed inventive concept, can make various changes, for example with regard to the function or the arrangement of individual elements mentioned in an exemplary embodiment, without leaving the scope of protection defined by the claims and their legal equivalents, such as further explanations in the description.

Claims

[1] A method for machine image classification, wherein an input image (1) generated from camera data is fed to a pre-trained sparse coding model with at least three layers (2) (S1), an output layer (3) of which has a plurality of neurons (5) to which a respective class is assigned, wherein at least one pair of two adjacent layers (2) a measure of feature recognition quality of a subsequent layer (2) of the pair is fed back to a preceding layer (2) of the pair via a feedback connection (4) (S2) and is used for adaptation in the preceding layer (2), wherein an iterative sparse coding algorithm is executed in each of the at least three layers (2) (S3), characterized by, that an inference of the pretrained sparse-coding model is repeatedly performed with the input image (1) (S4) and in each of the repetitions a state value of a respective neuron (5) of the output layer (3) is artificially increased (S5) in order to iteratively simulate a result of the inference as membership in a respective class, wherein the respective artificial increase is propagated from the output layer (3) through the layers (2) by means of the feedback connections (4), and wherein for each artificial increase a cost function is formed based on a reconstruction quality of the input image (1) and / or on the basis of a sparsity of the sparse-coding model (S6), and the input image (1) is assigned to the class (S7) whose assigned neuron (5) of the output layer (3) leads to the lowest value of the cost function when artificially increased. [2] Method according to claim 1, wherein the iterative sparse coding algorithm is a locally competitive algorithm. [3] Method according to any of the preceding claims, wherein the camera data has a neuromorphic data format. [4] Method according to one of the preceding claims, wherein an input image (1) is generated from a gesture of a vehicle occupant, wherein it is recognized when a new gesture not yet assigned to a class is present, wherein a further neuron (5) is added to the output layer (3) which represents a class assigned to the new gesture. [5] Method according to any of the preceding claims, wherein at least one of the following is adaptable by a user: - The number of iterations of the iterative sparse coding algorithm; - A degree of sparsity that is relevant for classification; - A selection of different metrics to determine the classification, in particular a maximum activation or a minimum reconstruction error; - A learning rate of the sparse coding model for setting a response speed to input images (1) that represent a new gesture; [6] Method according to one of the preceding claims, wherein the artificial enhancements of the respective neurons (5) of the output layer (3) are carried out in parallel and the feedback connections (4) are calculated in parallel for each artificial enhancement. [7] A machine-based image classification system comprising a camera and a computing unit, wherein the computing unit is configured to feed an input image (1) generated with data from the camera to a pre-trained sparse-coding model having at least three layers (2), an output layer (3) of which has a plurality of neurons (5) to which a respective class is assigned, wherein at least one pair of two adjacent layers (2) has a feedback connection (4) for feeding back a measure of feature recognition quality from a subsequent layer (2) of the pair to a preceding layer (2) of the pair, wherein the measure serves for adjustment in the preceding layer (2), wherein the computing unit is configured to execute an iterative sparse-coding algorithm in each of the at least three layers (2), characterized by, that the computing unit is designed to repeatedly perform an inference of the pretrained sparse-coding model with the input image (1) and in each of the repetitions artificially increase a state value of a respective neuron (5) of the output layer (3) in order to iteratively simulate a result of the inference as membership in a certain class, wherein the computing unit is designed to propagate the respective artificial increase from the output layer (3) through the layers (2) by means of the feedback connections (4), and to form a cost function for each increase based on a reconstruction quality of the input image (1) and / or on the basis of a sparsity of the sparse-coding model, and to assign the input image (1) to the class whose assigned neuron (5) of the output layer (3) leads to the lowest value of the cost function when artificially increased. [8] Method for generating a machine-based image classification model, wherein the model is a sparse-coding model having at least three layers (2), an output layer (3) having a plurality of neurons (5) to which a respective class is assigned, and a feedback connection (4) for feeding a measure of feature recognition quality from a subsequent layer (2) of the pair back to a preceding layer (2) of the pair and for adaptation in the preceding layer (2), wherein the model is trained by specifying an associated class for an input image (1), artificially increasing the value of a neuron (5) of the output layer (3) belonging to the specified class, wherein the artificial increase is propagated stepwise from the output layer (3) through the layers (2) by means of the feedback connections (4),and wherein in each of the at least three layers (2) parameters are adjusted according to the specification. [9] Method according to claim 8, wherein parameters of a single layer (2) are adjusted for each of the at least three layers (2) by a separate back-propagation process. [10] Method according to one of claims 8 to 9, wherein an extraction of features from preceding layers (2) is used for a feature classification in at least one subsequent layer (2) further towards the initial layer (3).

Citation Information

Patent Citations

  • CN000107133361A

  • Hand motion identification method and apparatus

    US20160335487A1

  • Performing Inference and Training Using Sparse Neural Network

    US20210158168A1

  • System and method for levarging multiple descriptive features for robust few-shot image learning

    US20250005916A1