Method and apparatus for identifying a data gap in a training data set

Delaunay triangulation-based data gap identification in training data sets enhances neural network reliability by systematically filling underrepresented regions, improving model performance and reducing unpredictability.

DE102024203138A1Pending Publication Date: 2025-10-09ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024203138
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-05
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

The reliability and safety of neural networks in critical applications are compromised due to limited understanding of their decision-making processes, particularly when operating outside their training dataset, necessitating an approach to identify and potentially close data gaps in training data sets to ensure robust performance.

Method used

A method using Delaunay triangulation to span the input data space with simplices, analyzing geometric parameters like volume or edge length to identify data gaps, allowing targeted data collection or generation to fill these gaps.

Benefits of technology

Enhances the reliability and accuracy of neural networks by systematically identifying and addressing underrepresented regions in the training data, improving model performance and reducing unpredictability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000007_0000
    Figure 00000007_0000
  • Figure 00000008_0000
    Figure 00000008_0000
Patent Text Reader

Abstract

The invention relates to a method for identifying a data gap in a training data set of a machine learning model, the method comprising the steps: - Providing (S1) the training data set (X, Y) with input data (X) and labels (Y) associated with the input data (X), wherein the training data set (X, Y) forms an input data space; - spanning (S2) the input data space by simplices using a triangulation method; and - Identifying (S3) the data gaps in the input data space based on a geometric parameter of the simplices depending on a predetermined limit value.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method and a device for identifying a data gap in a training data set of a machine learning model. State of the art

[0002] Neural networks have proven themselves to be one of the most powerful tools in data processing and artificial intelligence in recent years. Their ability to detect and learn complex patterns in large data sets has made them the preferred choice for a wide range of applications, from image and speech recognition to complex decision-making processes. The results achieved through the use of neural networks have pushed the boundaries of what machine learning can achieve.

[0003] Despite the advances being made in this technology, concerns remain about the safety and reliability of its application in critical areas. One problem is that our understanding of how neural networks arrive at their decisions remains limited. This uncertainty is particularly evident in safety-critical applications, where the accuracy and reliability of the results are crucial. Of particular note is the observation that neural networks can make decisions with apparent confidence even outside of their training dataset, even when they are actually unable to produce the correct answer. This tendency toward "confident" misjudgments poses a risk because it calls into question the reliability and safety of the outputs generated by these systems.

[0004] To maximize the power of neural networks while minimizing risks, it is important to ensure that the training data is as comprehensive as possible. This means identifying potential "gaps" in the training data and, where possible, addressing them to ensure robust and reliable neural network performance. The challenge is to find an approach that leverages the capabilities of neural networks while minimizing the risks associated with their current unpredictability and incomprehensibility.

[0005] Previous methods for finding gaps in training data mostly rely on the density of data points in the input space. Regions with a low density are potentially underpopulated, so additional training data is required for these regions.

[0006] It is an object of the invention to provide an improved method and / or a device for identifying and preferably closing data gaps in a training data set or an input space of a machine learning model represented by the training data set.

[0007] The object is achieved by a method according to the features of patent claim 1. The object is achieved by a device according to the features of patent claim 10. Disclosure of the invention

[0008] According to a first aspect, a method for identifying a data gap in a training dataset of a machine learning model is proposed. The method comprises the steps: - Providing the training dataset with input data and labels associated with the input data, wherein the training dataset forms an input data space; - Spanning the input data space by simplices using a triangulation method; and - Identifying the data gaps in the input data space based on a geometric parameter of the simplices depending on a predetermined threshold.

[0009] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.

[0010] According to a second aspect, a device is proposed for identifying a data gap in a training data set of a machine learning model, wherein the device comprises an evaluation and computing device which is designed to carry out the following steps: - Providing the training dataset with input data and labels associated with the input data, wherein the training dataset forms an input data space; - Spanning the input data space by simplices using a triangulation method; and - Identifying the data gaps in the input data space based on a geometric parameter of the simplices depending on a predetermined threshold.

[0011] The statements made for the method apply accordingly to the device. It is understood that linguistic modifications of procedurally formulated features can be reformulated for the device according to common linguistic practice, without such formulations having to be explicitly listed here.

[0012] The input data serves as input to the machine learning model. The labels assigned to the input data represent the model's output values, which are to be achieved by training the machine learning model. "Spanning the input data space with simplices using a triangulation method" involves structuring the input data space by spanning it with simplices. Simplices are geometric objects (e.g., lines, triangles, tetrahedra in higher dimensions) created using a triangulation method. This method serves to subdivide the space so that its geometric structure can be represented by simpler, smaller units. Identifying data gaps is done by analyzing the geometric properties of the simplices. A geometric parameter (e.g.,The analysis of the simplices' geometric parameters (e.g., size, volume, edge length) is used to identify regions that are poorly covered or not covered at all by the training dataset. The decision as to whether a data gap exists depends on a predetermined threshold. If the geometric parameter of a simplex exceeds this threshold, the region is identified as a data gap. The method provides a systematic approach for identifying regions within the input data space of a machine learning model that are insufficiently represented by the existing training dataset. This is particularly useful for improving the performance of machine learning models by specifically collecting or generating data that fills these gaps. The present approach identifies regions in the input space with sharp boundaries. The sharp demarcation allows for precise filling and, if necessary, re-measurement of underpopulated regions in the input space.The method allows the label information of the data points to be directly included in the identification of sparsely populated regions in the input space.

[0013] In other words, the method requires a training dataset to be provided, including the inputs and the corresponding labels / outputs. Using methods such as Delaunay triangulation, the input data space can be spanned using simplices. Since Delaunay triangulation triangulates the input data within its convex hull, very large simplices may be necessary in regions where data is insufficient. These regions can thus be identified geometrically, for example, by using the volume of the simplices or their edge length as an indicator of the presence of a data gap.

[0014] The method can be used to select suitable training data points and / or to expand a training dataset for training and / or testing and / or verifying and / or validating a machine learning model. The method can be used to determine in which region of the input space additional input data should be provided or samples should be measured. The boundaries of the "gaps or holes" in the input space preferably indicate opposite point settings where previous input data were provided or measurements were taken. Ideally, an interpolation of these points can target a region of interest.

[0015] The method can be used for the active selection of (training) data that a technical system transmits to a back-end computer, for example, for a high-precision simulation. This reduces data traffic. The back-end computer can use this information for training a machine learning system and / or for testing, verifying, and / or validating a machine learning system.

[0016] The method can be used to analyze data acquired by a sensor. The sensor can determine measured values ​​of an environment in the form of sensor signals, which can originate from the following sources, for example: Digital images can be processed, which can be captured by a video sensor, a radar sensor, a LiDAR sensor, an ultrasonic sensor, a motion sensor, and / or a thermal imaging sensor. The method can be used to analyze specific, particularly low-dimensional, data with which the triangulation algorithm applied in the method can function.

[0017] The method can be used for AI-based problems with medium- to low-dimensional (training) data. An example is the technical context of motor vehicles, where the method can be used to find gaps in training data used to train a machine learning model for interpretation and / or evaluation and / or to generate labels for measurement results from needle closure detection or similar virtual sensors. The method can also be used to analyze training data for data gaps in the areas of on-board fuel consumption monitoring (OBFCM) or acceleration measurement using accelerometers. The method can generally be applied to low-dimensional (intrinsic dimensionality) data that has a dimensionality suitable for calculating the convex hull.

[0018] The method can be used to detect anomalies in a technical system. Specific regions can be defined that are not sufficiently sampled in the training data. This allows outliers in production that fall into this region to be identified.

[0019] The method can be used to calculate a control signal for controlling a technical system, such as a computer-controlled machine, such as a robotic system, a vehicle, a household appliance, a power tool, a manufacturing machine, a personal assistant or an access control system.

[0020] The method can be used for measuring and / or controlling, and in particular for analyzing, data (e.g., tabular data), especially from a sensor on a test bench that measures the data. If data gaps are detected by the method, the measured technical system can be instructed to make certain settings to acquire additional data from areas where the data gap is detected and thus hardly any data has been collected by the sensor so far.

[0021] In another aspect, the triangulation method features a Delaunay triangulation.

[0022] Delaunay triangulation is a special form of triangulation that, for a given set of points in multidimensional space, generates a triangulation such that no point lies within the circumcircles of all triangulation simplices. Delaunay triangulation is known for its mathematical properties, which enable optimal triangulations under various criteria, such as minimizing the smallest angle of the triangles. Delaunay triangulation is chosen to span the input data space. This method is preferred because it provides an efficient and effective way of structuring the space while satisfying the geometric conditions favorable for identifying data gaps.

[0023] In another aspect, the geometric parameter comprises a volume of the simplices or an edge length of the simplices.

[0024] Other geometric parameters are also conceivable.

[0025] In a further aspect, identifying the data gaps comprises calculating the volumes V i of all simplices, calculating an average or expected volume V¯=1N∑i Vi over all simplices and calculating a standard deviation ΔV of the volumes, whereby a data gap is identified depending on the predetermined limit if the volume of a simplex Vs of the simplices is greater than a sum of the mean volume V and the standard deviation of the volumes ΔV, which is parameterized with a setting parameter n.

[0026] The condition can be formulated as follows: VS>V¯+n⋅ΔV

[0027] Alternatively, identifying the data gaps comprises calculating the edge length of all simplices, calculating an average edge length over all simplices, and calculating a standard deviation of the edge lengths, wherein a data gap is identified depending on the predetermined threshold value if the edge length of a simplex of the simplices is greater than a sum of the average or expected edge length and the standard deviation of the edge lengths multiplied by a setting parameter.

[0028] The setting parameter can be selected and specifies at what point a volume size is considered too large, i.e., a data gap should be identified. This setting parameter can be selected depending on the application.

[0029] In a further aspect, identifying the data gaps in the input data space comprises determining a complexity of a region of the input space based on a number of simplex vertices or based on a maximum gradient of the simplices. It should be noted that according to Delaunay triangulation, all simplices have the same number of vertices. Evaluating the gradient is particularly preferred when continuous labels are present at the vertices of the simplices.

[0030] A possible extension of the method for identifying a data gap takes into account how "complicated" a region is in the input space. In this case, the labels of the simplex vertices can be included in the evaluation.

[0031] In another aspect, determining the complexity based on the simplex vertices comprises solving a classification problem that comprises calculating expected volumes and their standard deviation depending on a number of different classes that the simplex vertices have.

[0032] In another aspect, determining the complexity involves distinguishing between discrete labels and continuous labels. Where “k” can be the number of different “discrete labels” and f(x i ) describe the labels at the corners of the simplices in the case of “continuous labels”.

[0033] For example, in the case of discrete labels, expected volumes V̅ (k) and their standard deviation ΔV (k)be calculated, depending on the number of different labels k that the simplex vertices have. Regions where all vertices have the same label may preferentially allow larger simplices. In regions with many classes, a smaller n k be defined (using the volumes from equation 1).

[0034] In another aspect, determining complexity based on the maximum gradient of the simplices involves binning into a certain number of categories.

[0035] Especially for continuous labels, simplices can be preferably classified based on their maximum gradient d max = max i,j |(y i - y j ) / (x i - x j )| clustered - where y i the labels and the vertices x irepresent - and thus data gaps can be identified analogously depending on their gradient. In order to proceed analogously to the classification, k categories can preferably be created by binning, with d l (k) < d max < d h (k) .

[0036] Complexity is defined or quantified by considering the maximum gradient present within individual simplices (e.g., triangles in a two-dimensional space, tetrahedra in a three-dimensional space, etc.). This maximum gradient is a measure of the inclination or gradient of the faces or edges of the simplices, which allows conclusions to be drawn about the variability or "unrest" within the space defined by the simplices.

[0037] To make the complexity manageable and interpretable, the concept of binning is applied. This means that the determined gradient values ​​are classified into a predefined number of categories (bins). Each category represents a specific range of gradient values. This categorization enables a simplified analysis and visualization of the complexity of the space. The number of categories is an important parameter that influences the resolution and level of detail of the complexity determination.

[0038] In a further aspect, a computer program with program code is claimed for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method(s) of the method in one of its embodiments.

[0039] In a further aspect, a computer-readable data carrier with program code of a computer program is proposed for executing at least parts of the method in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions that, when executed by a computer, cause the computer to execute the method(s) of the method in one of its embodiments.

[0040] The described designs and further training courses can be combined as desired.

[0041] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings

[0042] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.

[0043] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.

[0044] They show: Fig. 1 is a schematic flow diagram of an embodiment of the present method; and Fig. 2 a schematic representation of an entrance room divided into a plurality of simplices.

[0045] In the figures of the drawings, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.

[0046] Fig. 1 shows a schematic flow diagram of a method for identifying a data gap in a training dataset of a machine learning model.

[0047] In any embodiment, the method can be carried out at least partially by a device 100, which for this purpose can comprise several components not shown in detail, for example, one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the device 100, which can be part of a system, can comprise a storage device and / or an output device and / or a display device and / or an input device.

[0048] The computer-implemented method comprises at least the following steps: In a step S1, a training data set X, Y is provided with input data X and labels Y associated with the input data X, wherein the training data set X, Y has an input data space 200 (see Fig. 2) forms. In a step S2, the input data space 200 is spanned by simplices 202 (see Fig. 2) using a triangulation method.

[0049] In a step S3, a data gap in the input data space 200 is identified on the basis of a geometric parameter of the simplices 202 as a function of a predetermined limit value.

[0050] Fig. Figure 2 shows the input data space (X, Y) 200, which is spanned by the simplices 202 based on the Delaunay triangulation. In this case, the simplices are triangles between respective simplex vertices 204.

[0051] A data gap in the input data space 200 can now be identified based on a geometric parameter of the simplices 202 as a function of a predetermined threshold value. The geometric parameter can be a volume 206 of the simplices 202 or an edge length 208 of the simplices 202.

[0052] If the volume 206 of a simplex of the simplices 202 is greater than a sum of the mean volume and the standard deviation of the volumes, which is parameterized with a setting parameter, a data gap is identified.

[0053] Alternatively or additionally, if the edge length 208 of a simplex of the simplices 202 is greater than a sum of the mean edge length and the standard deviation of the edge lengths multiplied by a tuning parameter, a data gap is identified.

[0054] An example data gap is marked with 210. This exists because the edge length 208, 212 of a certain simplex is larger than the limit, i.e., larger than a sum of the mean edge length and the standard deviation of the edge lengths multiplied by a setting parameter.

Claims

[1] A method for identifying a data gap in a training data set of a machine learning model, the method comprising the steps: - Providing (S1) the training data set (X, Y) with input data (X) and labels (Y) associated with the input data (X), wherein the training data set (X, Y) forms an input data space; - spanning (S2) the input data space by simplices using a triangulation method; and - Identifying (S3) the data gaps in the input data space based on a geometric parameter of the simplices depending on a predetermined limit value. [2] The method of claim 1, wherein the triangulation method comprises a Delaunay triangulation. [3] The method according to claim 1 or 2, wherein the geometric parameter comprises a volume of the simplices or an edge length of the simplices. [4] Method according to claims 1 and 3, wherein identifying the data gaps comprises calculating the volumes of all simplices, calculating an average volume over all simplices and calculating a standard deviation of the volumes, wherein a data gap is identified as a function of the predetermined limit value if the volume of a simplex of the simplices is greater than a sum of the average volume and the standard deviation of the volumes, which is parameterized with a setting parameter;or wherein identifying the data gaps comprises calculating the edge length of all simplices, calculating an average edge length across all simplices, and calculating a standard deviation of the edge lengths, wherein a data gap is identified as a function of the predetermined threshold if the edge length of a simplex of the simplices is greater than a sum of the average edge length and the standard deviation of the edge lengths multiplied by a setting parameter; [5] The method of any preceding claim, wherein identifying the data gaps in the input data space comprises determining a complexity of a region of the input space based on a number of simplex vertices or based on a maximum slope of the simplices. [6] The method of claim 5, wherein determining the complexity based on the simplex vertices comprises solving a classification problem comprising calculating expected volumes and their standard deviation depending on a number of different classes having the simplex vertices. [7] Method according to one of the preceding claims, wherein determining the complexity based on the maximum gradient of the simplices comprises binning into a certain number of categories. [8] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 7 when the computer program is executed on a computer. [9] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 7 when the computer program is executed on a computer. [10] Device (100) for identifying a data gap in a training data set of a machine learning model, wherein the device (100) comprises an evaluation and computing device which is designed to carry out the following steps: - Providing the training data set (X, Y) with input data (X) and labels (Y) associated with the input data (X), wherein the training data set (X, Y) forms an input data space; - Spanning the input data space by simplices using a triangulation method; and - Identifying the data gaps in the input data space based on a geometric parameter of the simplices depending on a predetermined threshold.

Citation Information

Patent Citations

  • SECURE CONTROL OF TECHNICAL-PHYSICAL SYSTEMS

    DE102022209898A1