Dimension reduction process

By determining task-specific feature spaces through data pair comparisons and PCA, the method aligns high-dimensional spaces with task-specific directions, enhancing training efficiency and accuracy for machine learning models, particularly in autonomous navigation.

DE102023130646B4Active Publication Date: 2025-08-21CARIAD SE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102023130646
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-06
Publication Date
2025-08-21
Estimated Expiration
2043-11-06

AI Technical Summary

Technical Problem

Existing methods for dimensionality reduction in high-dimensional feature spaces for machine learning models fail to adequately align with the task-specific features, particularly neglecting less frequent objects, leading to inefficiencies in training and application.

Method used

A method that determines task-specific feature spaces by comparing data pairs with defined feature differences, using techniques like masking or replacing data elements, followed by rotation and PCA to project relevant vectors, ensuring alignment with task-specific directions.

Benefits of technology

Enhances the alignment of feature spaces with task-specific features, improving training efficiency and accuracy, especially for less frequent objects, and enabling reliable applications in autonomous navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0001_ABST
    Figure 00000000_0001_ABST
  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for reducing the dimension of a multidimensional feature space for training a machine learning model (50) by machine learning, comprising the following steps: - Providing (101) at least one data pair (30), in which an original data element (31) and a modified data element (32) each have a feature difference (Δf) that is specific to a respective defined task for machine learning, wherein the at least one data pair (30) is specific to sensor data that results from a detection of a sensor (40), - determining (102) at least one task-specific feature space which is specific for the at least one feature difference (∆f) on the basis of a comparison of the respective data pairs (30), - carrying out (103) the dimensionality reduction on the basis of the determined task-specific feature space, characterized in that the following steps are carried out: - training the machine learning model (50) for the at least one defined task on the basis of the performed dimensionality reduction, wherein the at least one defined task comprises a recognition of the at least one feature difference (Δf), - Providing the trained machine learning model (50) for an application in which the at least one defined task is applied to the and / or further sensor data resulting from detection by the sensor (40) and / or a further sensor (40), wherein navigation of an at least partially autonomous robot and / or vehicle is carried out on the basis of the detection, wherein the image data represent a traffic scene during navigation, wherein the at least one feature difference (∆f) is provided as a difference in an image feature of the image data, which indicates a navigation-relevant difference in the traffic scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for dimensionality reduction of a feature space for training a machine learning model. Furthermore, the invention relates to a machine learning model, a computer program, a device, and a storage medium, each for this purpose. State of the art

[0002] Diversity sampling is a common facet of active learning (see, for example, Yang, Yi, et al., “Multi-class active learning by uncertainty sampling with diversity maximization.” International Journal of Computer Vision, 2015). To select a diverse set of images, the distribution within a high-dimensional feature space is often used. This feature space can be taken from the backbone of a multitasking model or from a pre-trained general-purpose model, such as “CLIP” (Ramesh, Aditya, et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv:2204.06125 [cs.CV]). Regardless of their origin, these feature spaces are often very high-dimensional. To select relevant directions, techniques such as principal component analysis (PCA) are often used to reduce the dimensionality of this space.

[0003] PCA efficiently determines and ranks the vectors of the feature space over which the data exhibit maximum variance. What generates the highest variance in an image set is likely to be of some interest, but does not necessarily correlate with the directions most important for the underlying tasks of a head—hereafter also referred to as the task head—of a multitasking model. Using earlier layers from the specific task head as the feature layer can provide some degree of task focus, but the variance in the associated directions may be very low, particularly in the case of less frequently occurring objects within the dataset. In a very simplified representation of traffic sign recognition, a feature direction correlated with the presence of yield signs may exhibit much lower variance than others, such as the direction of the turn.the presence of snow, night, buildings or vehicles, but focusing on this feature space may be desirable for selecting appropriately diverse images.

[0004] Ideally, the relevant feature space directions can be specified and preselected before applying a dimensionality reduction technique such as PCA. Using an early layer from a task header as the feature layer can somewhat restrict the directions to the target task.

[0005] The publication "JIN, Q., et al.: One-shot active learning for image segmentation via contrastive learning and diversity-based sampling. In: Knowledge-Based Systems, 2022, vol. 241, pp. 1-12. doi: 10.1016 / j.knosys.2022.108278" discloses a One-Shot Active Learning (OSAL) framework based on contrastive self-supervised learning and a diversity-based retrieval strategy.

[0006] The further publication "YAMANE, I., et al.: Multitask principal component analysis. In: Asian Conference on Machine Learning. PMLR, 2016. pp. 302-317." describes the use of principal component analysis (PCA) for dimensionality reduction in various forms.

[0007] The further publication "TIOMOKO, M., Couillet, R., Pascal, F.: PCA-based Multi Task Learning: a Random Matrix Approach. In: arXiv preprint, arXiv:2111.00924, 2021. pp. 1-11. doi: 10.48550 / arXiv.2111.0092" discloses various methods and techniques in the field of machine learning, particularly in the context of multi-task learning (MTL) and principal component analysis (PCA). Disclosure of the invention

[0008] The subject matter of the invention is a method having the features of claim 1, a machine learning model having the features of claim 8, a computer program having the features of claim 9, a device having the features of claim 10, and a computer-readable storage medium having the features of claim 11. Further features and details of the invention emerge from the respective subclaims, the description, and the drawings. Features and details described in connection with the method according to the invention naturally also apply in connection with the machine learning model according to the invention, the computer program according to the invention, the device according to the invention, and the computer-readable storage medium according to the invention, and vice versa, so that with regard to the disclosure of the individual aspects of the invention, reference is or can always be made to each other.

[0009] The invention particularly relates to a method for reducing the dimension of a multidimensional and preferably high-dimensional feature space, in particular for training a machine learning model using machine learning. The method can comprise providing at least one data pair, in which an original data element and a modified data element (as a data pair) are provided and / or have a feature difference that is specific to at least one respective defined task for machine learning. The at least one or more different defined tasks can each be provided by a task header of the machine learning model.

[0010] The at least one data pair can also be specific to sensor data resulting from detection by a sensor such as an image sensor. This means, for example, that the at least one data pair comprises the sensor data from a sensor detection and / or is based at least partially on an image recording from the image sensor and can include corresponding image information. The image information can, for example, comprise images of a vehicle's surroundings that have been detected by the image sensor. The sensor data can therefore, for example, comprise pixels or image points that represent the detected surroundings. It is also conceivable that the at least one data pair comprises the sensor data in the form of measurement data that has been determined by metrological detection by the sensor.Furthermore, it can be provided that the at least one data pair is specific to the sensor data in that it comprises at least partially constructed training data, which enables training of the machine learning model for an application on the sensor data. Various methods can also be combined here, such as augmentation or at least partial simulation of the sensor data. This can have the advantage that the machine learning model is prepared for a variety of situations and environments and thus has greater accuracy for the application.

[0011] Furthermore, the method may comprise the following steps, which are preferably carried out successively and / or repeatedly, preferably iteratively for the various defined tasks: - Determining at least one task-specific feature space, which is specific for the at least one feature difference, on the basis of a comparison of the respective data pairs, wherein the task-specific feature space is preferably designed as a subspace of the multidimensional and preferably high-dimensional feature space and / or comprises the feature space vectors relevant for the respective task, - Performing dimension reduction based on the determined task-specific feature space.

[0012] An advantage of the method is that the relevant feature space vectors can be determined using a simple method. This subspace can then be projected out, for example, by deconvolution and preferably rotation, in order to then apply a PCA to the (reduced) subspace for dimensionality reduction. In particular, the dimensionality reduction based on the determined task-specific feature space serves to specify and preselect the relevant feature space directions before applying a further dimensionality reduction technique such as PCA. Alternatively, the implementation of the dimensionality reduction can also already include the implementation of a PCA. The PCA can further provide at least one parameter, such as a transformation result, with which the dimensionality reduction can also be taken into account in a later application and in particular inference of the trained machine learning model.

[0013] In addition, the invention provides for the following steps to be carried out: - Training the machine learning model for the at least one defined task on the basis of the performed dimensionality reduction, wherein the at least one defined task comprises recognition and / or detection of the at least one feature difference, preferably in the form of classification and / or object detection, - Providing the trained machine learning model for an application and preferably inference, in which the at least one defined task is applied to the and / or further sensor data, preferably image data, which result from a detection of the sensor and / or another sensor, preferably image sensor.

[0014] It can be provided that the surroundings of a vehicle and preferably a traffic scene are represented by the values ​​of the sensor data and preferably image data. At least one of the tasks can also comprise classification and preferably image classification based on these values ​​in order to detect objects in the traffic scene, for example. The classification and image classification can also be provided in the form of semantic segmentation (i.e., pixel- or area-wise classification) and / or object detection. Furthermore, it is possible for the application to comprise traffic sign recognition and / or recognition of traffic signals of a traffic light system. Thus, an output of the machine learning model, such as a classification result, can be used to navigate an at least partially autonomous robot and / or an at least partially autonomous vehicle and / or to control it taking a traffic scene into account.

[0015] Furthermore, within the scope of the invention, it is optionally possible for the provided data elements to be specific to the image data, wherein the at least one feature difference is provided as a difference in an image feature of the image data. Preferably, the detection of the at least one feature difference (e.g., object detection) can be performed based on pixel values ​​of the image data. The image data can be, for example, images from a radar sensor, an ultrasonic sensor, a LiDAR sensor, and / or a thermal imaging camera. Accordingly, the images can also be implemented as radar images, ultrasonic images, thermal images, and / or LiDAR images.

[0016] Furthermore, within the scope of the invention, navigation of an at least partially autonomous robot and / or vehicle is carried out based on the recognition and preferably classification and / or object detection, wherein the image data represents a traffic scene during navigation, wherein the at least one feature difference is provided as a difference in an image feature of the image data, which indicates a navigation-relevant difference in the traffic scene, preferably in the form of different signals of a traffic light system and / or different traffic signs. Thus, the reliability of such navigation can be improved by the dimensionality reduction and, if appropriate, a selection of training data based thereon.

[0017] Within the scope of the invention, it can preferably be provided that a transformation result is obtained based on the performed dimensionality reduction, preferably by applying a principal component analysis (PCA) to the reduced feature space. The transformation result can preferably be specific to a weighting or loading of the PCA. The transformation result can be used in the application of the trained machine learning model to reduce the dimensionality of the sensor data. For this purpose, the transformation result can be used to form an additional layer in the machine learning model, in particular above a trained neural network of the machine learning model, in order to perform the dimensionality reduction.

[0018] Furthermore, it can be provided that the at least one defined task comprises several different tasks for which the dimensionality reduction is performed. For this purpose, a specific feature space can be determined for each of these tasks. Preferably, the different tasks are provided by different task heads of a machine learning model to enable reliable application, e.g., in an autonomous vehicle.

[0019] Furthermore, it is conceivable that the provision of the at least one data pair comprises at least one of the following steps: - Masking one of the data elements of the data pair, - replacing part of one of the data elements with part of another data element, - Performing an in-painting to modify one of the data elements of the data pair.

[0020] In this way, a simple procedure can be applied to determine the task-relevant feature space vectors: a data element is either masked, replaced by a cut-out portion of another data element, or painted over. These measures serve to generate two data elements whose difference in the feature space should be correlated with the feature difference in the two data elements.

[0021] The invention also relates to a machine learning model trained using a method according to the invention. The machine learning model can be implemented, for example, as a multitasking model that has multiple task heads to provide the various defined tasks.

[0022] The invention also relates to a computer program, in particular a computer program product, comprising instructions that, when executed by a computer, cause the computer to carry out the method according to the invention. Thus, the computer program according to the invention provides the same advantages as those described in detail with reference to a method according to the invention.

[0023] The invention also relates to a data processing device configured to carry out the method according to the invention. The device can be, for example, a computer that executes the computer program according to the invention. The computer can have at least one processor for executing the computer program. A non-volatile data memory can also be provided, in which the computer program is stored and from which the computer program can be read by the processor for execution.

[0024] The invention may also provide a computer-readable storage medium that contains the computer program according to the invention and / or includes instructions that, when executed by a computer, cause the computer to carry out the method according to the invention. The storage medium is designed, for example, as a data storage device such as a hard disk and / or a non-volatile memory and / or a memory card. The storage medium can, for example, be integrated into the computer.

[0025] Furthermore, the method according to the invention can also be implemented as a computer-implemented method.

[0026] Further advantages, features, and details of the invention will become apparent from the following description, which describes embodiments of the invention in detail with reference to the drawings. The features mentioned in the claims and in the description may be essential to the invention individually or in any combination. They show: Fig. 1 a schematic visualization of a method, a machine learning model, a device, a storage medium and a computer program according to embodiments of the invention. Fig. 2 a further illustration for visualizing a method according to embodiments of the invention.

[0027] In Fig. 1, a method 100, a device 10, a storage medium 15, a machine learning model 50 and a computer program 20 according to embodiments of the invention are schematically illustrated.

[0028] A common problem in machine learning is that the feature space of diversity selection and the task of the machine learning model 50 are poorly aligned. Embodiments of the invention can ensure that the task is not neglected during diversity selection. The proposed method has the advantage of being simple and intuitive. Furthermore, it can also be applied to multiple tasks and can be easily automated in pipelines once the target images have been created.

[0029] Embodiments of the invention can be used as an upstream part of a machine learning toolchain. A machine learning toolchain can be a collection of software tools used in a coordinated process to develop, train, and implement machine learning models. Furthermore, embodiments of the invention can also be used in the navigation of an at least partially autonomous robot and / or vehicle.

[0030] Fig. 1 illustrates, according to embodiments of the invention, a method 100 for dimensional reduction of a multi- or high-dimensional feature space for training a machine learning model 50. The training can be carried out by machine learning, in which, in particular, one or more defined tasks are learned. The dimensional reduction can precede the training, in particular to prepare the training data for training. According to a first method step 101, at least one data pair 30 can be provided for this purpose, in which an original data element 31 and a modified data element 32 can each have a feature difference Δf that is specific to the at least one defined task for the machine learning. The modified data element 32 can, if necessary, be generated on the basis of the original data element 31, e.g.by modifying the content of the original data element 31. Then, according to a second method step 102, at least one task-specific feature space can be determined, which can (in each case) be a subspace of the original multidimensional feature space of the data pairs. It may be possible for a separate task-specific feature space to be determined for each of the tasks, which may differ from one another. The respective task-specific feature space can thus be specific to a feature difference Δf. In order to determine the feature space, a comparison of the respective data pairs 30 can be provided. According to a third method step 104, the dimensionality reduction can be performed based on the (respective) determined task-specific feature space. Furthermore, a training 400 of the machine learning model 50 for the defined task can take place based on the performed dimensionality reduction.This allows a trained machine learning model 50 to be provided for an application in which the at least one defined task is applied, for example, to the and / or further sensor data, preferably image data, which result from a detection of the sensor 40 and / or another sensor 40, preferably image sensor 40.

[0031] In particular, it is an inventive idea to use a simple method to determine the (task-)relevant feature space vectors, project this subspace out by rotation, and then apply PCA to the reduced space. As a simple method for determining the relevant feature space vectors, an image can, for example, either be masked, replaced with a section of another image, or painted over to generate two images. These two images comprise the original and the modified image, the difference in the feature space of which should be correlated with the information difference in the two images. Performing a forward pass through the network and extracting the relevant feature space for both images can yield feature space vectors for both the original and the modified image. Δf⇀i=f⇀mod,i−f⇀original,i.

[0032] By repeating this procedure over a small set of several sample images, the normalized mean of this separation can be obtained, f^=〈∑iNΔf⇀i〉.

[0033] This should correspond to a direction in the feature space that is closely related to this feature, where the angle brackets here denote the normalization to a unit vector. It can then be checked whether the change is meaningfully correlated, i.e., β=1N∑iNf^⋅〈Δf⇀i〉, should not be too much smaller than 1, otherwise f̂ will not correlate meaningfully with the desired feature direction. In these cases, it is likely that the choice of feature space is insensitive to the target or the sensitivity to the target is not sufficiently localized. In some cases, it may be useful to extract more than one feature direction (e.g., a feature hypersurface) for a given target space.

[0034] For example, given an image with a green traffic light, the green traffic light could be exchanged with an image where the traffic light is yellow. The difference between these two in the feature space, i.e., the feature difference Δf i , should be highly correlated with the exchange of green light for yellow light. Repeating this with a small set of different sample images, the mean of this separation should roughly indicate the direction of the difference between green and yellow light.

[0035] In Fig. Figure 2 shows an example of this method according to embodiments of the invention. Points 201 represent the position of the difference vectors between the image pairs in the n-dimensional feature space, e.g., with and without the corresponding object in the image section. Arrow 202 is the fit line, here dominant in the directions f i and f jand with a small component in the other directions (denoted by f n ). The direction of this arrow would be f̂ for the feature highlighted by the image pairs, e.g., a yellow traffic light.

[0036] After this task is completed for a single task goal, other small image sets can be used to find other desirable features specific to the task head (or desirable features from other task heads). An orthonormal basis can be constructed iteratively by removing the projections in the previous directions, i.e., f^i=f→i / |f→i|,f→j=f→'j−∑i=1j−1f→'j⋅f^i

[0037] It is possible that the orthonormal basis encounters some redundancy in the vectors when trying to form the task-specific directions. This can be determined if |f→j|<<|f→'j|. In this case, the direction can simply be omitted, as it is already taken into account by the existing selected directions. After defining a set of vectors and extracting an orthonormal basis, the feature space can be rotated. Given p orthonormal task-specific vectors and q desired basis vectors, the feature space can be rotated as follows: f'=R1f, so that the first task-specific direction is now aligned with the first direction in n-dimensional space. This can be repeated iteratively for all p directions, resulting in a single matrix, R=∏i=1pRi,

[0038] This adapts the first p dimensions of the n >> p-dimensional space to the target subspace. This matrix can be applied to the feature vectors of all images to perform dimensionality reduction based on the determined task-specific feature space. PCA can then be applied to select the q - p dimensions with the highest variance from the np-dimensional subspace. The p orthonormal task-specific directions and the q - p directions with the highest variance from the PCA within the orthogonal subspace together result in the reduced-diversity feature space, in which task-specific directions are now encoded. In the rotated n-dimensional basis, the PCA matrix is ​​K (q - p) × (n - p) dimensional, but can be trivially extended by the p-dimensional identity and using the rotation matrix to obtain Z = (I p⊕ K) · R, the combined single q × n matrix that brings the original feature vectors into the task-specific and PCA directions.

[0039] Finally, to ensure that task-specific trait directions are sufficiently taken into account in diversity selection, post-selection scaling can be used to obtain unit variance in all directions.

[0040] The described method can be used with a selection of preprocessed images to achieve specific task-specific goals. For example, the following steps can be provided in embodiments of the invention: First, the feature direction(s) of interest can be extracted from each preselected set of image pairs. Subsequently, the resulting orthonormal basis of p vectors can be constructed. The rotation matrix R can be determined and applied to all images in the dataset. PCA can then be performed, removing the first p dimensions to obtain the q-p PCA directions. The identity can then be expanded to obtain Z, and the reduced dimensional space can be extracted. After scaling to the unit normal, a diversity sampling algorithm of choice can be applied.

[0041] This setup can also be performed online. Once the model is trained and the above steps have been performed, the modified PCA can be set as the weights of a linear neural network layer associated with the model. This allows the transformation result to be used in the application of the machine learning model. During inference, the cost of extracting these feature vectors is very low. High positive values ​​of the corresponding feature directions can be used for online triggering.

[0042] The above explanation of the embodiments describes the present invention exclusively by way of examples. Of course, individual features of the embodiments can be freely combined with one another, provided they are technically feasible, without departing from the scope of the present invention.

Claims

[1] Method (100) for reducing the dimension of a multidimensional feature space for training a machine learning model (50) by machine learning, comprising the following steps: - Providing (101) at least one data pair (30), in which an original data element (31) and a modified data element (32) each have a feature difference (Δf) that is specific to a respective defined task for machine learning, wherein the at least one data pair (30) is specific to sensor data that results from a detection of a sensor (40), - determining (102) at least one task-specific feature space which is specific for the at least one feature difference (Δf) on the basis of a comparison of the respective data pairs (30), - performing (103) the dimensionality reduction on the basis of the determined task-specific feature space, characterized by , that the following steps are carried out: - training the machine learning model (50) for the at least one defined task on the basis of the performed dimensionality reduction, wherein the at least one defined task comprises a detection of the at least one feature difference (Δf), - Providing the trained machine learning model (50) for an application in which the at least one defined task is applied to the and / or further sensor data resulting from a detection of the sensor (40) and / or a further sensor (40), wherein, based on the recognition, a navigation of an at least partially autonomous robot and / or vehicle is carried out, wherein the image data represent a traffic scene during the navigation, wherein the at least one feature difference (Δf) is provided as a difference of an image feature of the image data, which indicates a navigation-relevant difference in the traffic scene. [2] Method (100) according to claim 1, characterized by , that the at least one defined task comprises the detection of the at least one feature difference (Δf) in the form of a classification and / or object detection, and / or that the sensor data is implemented as image data and the sensor is implemented as an image sensor. [3] Method (100) according to claim 2, characterized bythat the data elements are each specific to the image data, wherein the at least one feature difference (Δf) is provided as a difference of an image feature of the image data, and the detection of the at least one feature difference (Δf) is carried out on the basis of pixel values ​​of the image data. [4] Method (100) according to claim 2 or 3, characterized by that the navigation-relevant difference in the traffic scene is in the form of different signals from a traffic light system and / or different traffic signs. [5] Method (100) according to one of claims 2 to 4, characterized bythat a transformation result is obtained on the basis of the dimensionality reduction carried out, preferably by applying a principal component analysis (PCA) of the reduced feature space, wherein the transformation result is preferably specific to a weighting or loading of the principal component analysis, wherein the transformation result is used in the application of the trained machine learning model (50) for dimensionality reduction of the sensor data. [6] Method (100) according to one of the preceding claims, characterized by that the at least one defined task comprises a plurality of different tasks for which the dimensionality reduction is carried out, wherein for this purpose a specific feature space is determined in each case, wherein preferably the different tasks are provided by different task heads of a machine learning model. [7] Method (100) according to one of the preceding claims, characterized bythat the provision (101) of the at least one data pair (30) comprises at least one of the following steps: - masking one of the data elements of the data pair (30), - replacing part of one of the data elements with part of another data element, - Performing an in-painting to modify one of the data elements of the data pair. [8] Machine learning model (50) trained by a method (100) according to any one of the preceding claims. [9] Computer program (20) comprising instructions which, when the computer program (20) is executed by a computer (10), cause the computer (10) to carry out the method (100) according to one of the preceding claims. [10] Device (10) for data processing, which is arranged to carry out the method (100) according to one of claims 1 to 7. [11] A computer-readable storage medium (15) comprising instructions which, when executed by a computer (10), cause the computer (10) to carry out the steps of the method (100) according to any one of claims 1 to 7.