Computer-implemented method for quantifying the relevance of measured values ​​in time series

By quantifying relevance in industrial process data through residual and reference time series comparisons, the method reduces data volume and computational load, enhancing real-time monitoring and machine learning efficiency.

DE102024203263A1Pending Publication Date: 2025-10-16ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102024203263
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

The challenge of managing large volumes of data from industrial processes, particularly in real-time monitoring and analysis, leads to computational resource constraints and potential delays, necessitating methods to reduce data without losing important information.

Method used

A method for quantifying the relevance of measured values in time series by comparing them to a reference series, determining residual values and relevance values, and using these to identify critical data points for analysis, thereby reducing the data volume and computational load.

Benefits of technology

This approach allows for efficient real-time monitoring and analysis by focusing on critical data points, conserving computational resources and improving the effectiveness of machine learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for quantifying the relevance of measured values ​​in time series, the method comprising the following steps: - Providing a measured time series (S10), wherein the measured time series comprises a chronologically ordered sequence of measured values, wherein the measured values ​​were recorded as process parameters of an industrial process; - Providing a reference time series (S12); - Determining residual values ​​(S14) for the measured time series by comparing the measured time series with the reference time series; and - Determination of a relevance value (S16) for each measured value of the measured time series from the residual values.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to the field of monitoring and analysis of industrial processes, in particular the associated data processing. State of the art

[0002] The monitoring and analysis of industrial processes are increasingly challenged by the need to manage large amounts of data. This wealth of data results from the implementation of sensors and IoT devices in various industrial sectors. While these technologies contribute to making processes more efficient and predicting failures, they also pose a complex problem.

[0003] Industrial plants and machines produce a tremendous amount of real-time data. This data includes parameters such as temperature, pressure, flow rates, accelerations due to movement or vibration, and much more. Simply capturing and storing this data can be challenging, especially when it occurs in large quantities and at high frequency.

[0004] Furthermore, this data often needs to be analyzed in real time to identify potential problems or anomalies. This requires advanced analytics methods such as machine learning and artificial intelligence. However, processing such large amounts of data also requires significant computing resources and can lead to delays that can impair the ability to monitor and control in real time. Depending on a detected anomaly, the relevant production processes can be intervened or controlled, for example, by discarding the corresponding product, marking it as abnormal, or halting production.

[0005] To reduce the amount of data in the monitoring and analysis of industrial processes, various methods are used to extract relevant information and minimize unnecessary data. One of the most well-known methods is data aggregation. This involves summarizing raw data to reduce the number of data records that need to be stored. This aggregation can be time-based, for example, by summarizing data over specific time intervals, or spatial, by combining data from multiple sensors or sources into a single data set.

[0006] Another technique is downsampling. This involves selecting data points at a specific frequency to reduce the data volume while retaining important information. This can be particularly useful when data is collected at a higher frequency than required for analysis. However, it's difficult to distinguish which information is important and which isn't.

[0007] Additionally, techniques such as feature extraction are often used in data analysis to extract relevant information from the raw data and store or analyze only that information. This involves identifying and extracting specific characteristics or properties of the data that are of interest for analysis and monitoring. This allows for a reduction in the amount of data, but may result in the loss of important information.

[0008] The invention is therefore based on the object of proposing a method by which the amount of data in the monitoring and / or analysis of industrial processes can be reduced without losing important information.

[0009] The problem is solved by the subject matter of the independent claims. Disclosure of the invention

[0010] According to a first aspect of the invention, this object is achieved by a computer-implemented method for quantifying the relevance of measured values ​​in time series, the method comprising the following steps: - Providing a measured time series, wherein the measured time series comprises a chronologically ordered sequence of measured values, wherein the measured values ​​were recorded as process parameters of an industrial process; - Providing a reference time series; - Determining residual values ​​for the measured time series by comparing the measured time series with the reference time series; and - Determine a relevance value for each measured value of the measured time series from the residual values.

[0011] An industrial process can be any commercially applicable process for which at least one machine is used. A system, in the context of this invention, refers to a combination of at least one machine, preferably several machines. Machines can perform all kinds of industrial processes. Industrial processes include, but are not limited to: conveying, cutting, drilling, milling, grinding, punching, casting and injection molding, assembly, welding, painting, packaging, positioning, etc.

[0012] The measured values ​​for monitoring or analyzing an industrial process can therefore be any physical parameter that can be measured by the industrial process. These include, for example, the temperature of workpieces or machines and machine parts, the pressure within hydraulic or pneumatic systems, the speed of moving machine parts, the voltage or current in electronic components, etc.

[0013] The reference time series represents an ideal or average process flow, whereas the measured time series represents the actual, possibly live, process. The reference time series and the measured time series therefore relate to the same process parameter and have the same temporal division.

[0014] The measured values ​​or data points of the time series can be recorded in the microsecond, millisecond, or second range, for example. Longer intervals are also possible in principle, although these are rarely automated and monitored with the problem of data volume.

[0015] The measured time series is compared with the reference time series to identify any differences. This difference reveals where the measured time series deviates from the reference time series. These deviations are relevant for the monitored process because deviations from the norm in industrial processes often indicate process errors. The deviation is recorded as a residual value for each measured value.

[0016] For each measured value in the measured time series, a relevance value is determined from the residual value, generating a vector for the entire measured time series that indicates the relevance of each individual measured value for monitoring or analyzing the industrial process. A relevance map can therefore be generated for the entire process, which can be used for monitoring or analysis to reduce processing effort.

[0017] The measured time series may be entirely below or above the reference time series, or partially below or above it. Both deviations increase the relevance value. Therefore, the absolute value of the residual value is preferable for determining the relevance value.

[0018] Preferably, the relevance values ​​for the measured values ​​can be determined with the acquisition of the measured time series, so that the relevance value can be used directly in real-time monitoring of the industrial process.

[0019] In the manner described above, relevance values ​​can be determined for the measured values ​​of several measured time series. For example, a reference time series can be used for an entire series of industrial processes. If the reference time series does not change, this method can also detect gradual changes that could not be detected from individual time series or even when compared with a few time series.

[0020] The relevance values ​​indicate which measured values ​​are of particular interest for monitoring or analysis by reflecting the deviation of the measured values ​​from the target or normal state. For non-critical subprocesses, where generally no changes occur from one execution to the next, the relevance value is zero or almost zero, as these subprocesses hardly differ from the reference.

[0021] The industrial process could, for example, be a drilling process in which the force acting on the drill head is monitored. The industrial process begins with the drill head being positioned on the workpiece. Until the drill head is positioned, there is effectively no force acting on the drill head. This means that the measured force does not differ from the reference force, or only differs within the measurement inaccuracy of the corresponding sensor. The relevance value is therefore very small. When the drill head starts drilling, a force also acts. The force curve during drilling can, for example, differ from a reference hole if irregularities occur in the workpiece, for example if the material is brittle, the workpiece is not correctly positioned for the process, etc. If the measured force deviates from the reference force, the relevance value for these measured values ​​increases.During monitoring, a machine could then directly initiate countermeasures and, for example, correct or stop the process.

[0022] Regardless of how the measured time series is further processed, the processing of the time series can be limited to those measured values ​​of the measured time series that have a particularly high relevance value. This reduces the data to be analyzed to a relevant minimum, thus conserving computing resources and simplifying the entire processing process. The set of measured values ​​marked as relevant can be set, for example, via a threshold value. The invention thus achieves its objective.

[0023] In one embodiment, the residual values ​​are determined from the difference between the measured time series and the reference time series.

[0024] Calculating the difference between the measured time series and the reference time series is a simple way to determine a residual value. For this purpose, the monitoring system can, for example, store the reference time series in a working memory. When each measured value is recorded, the difference to the corresponding value from the reference time series can be calculated directly, and the result can be stored. This advantageously implements a particularly simple form of the invention.

[0025] In one embodiment, determining the residual values ​​comprises: - Determining spline coefficients for the reference time series; - Determine the spline coefficients for the measured time series; - Determining the difference or differences between the spline coefficients of the reference time series and the measured time series, whereby the residual values ​​are determined from the difference or differences of the spline coefficients; - Back-transforming the differences of the spline coefficients into the time domain and assigning the residual values ​​of the spline coefficients to the corresponding measured values ​​of the measured time series.

[0026] In this embodiment, determining the residual values ​​is somewhat more complex. In this embodiment, the time series can be divided into sections, each of which is assigned splines.

[0027] A spline, also known as a polynomial train, is a mathematical construct used in curve approximation and interpolation. It involves approximating a curve using a series of polynomial segments connected together to form a smooth and continuous curve. Each polynomial segment is called a segment of the spline and is usually represented by low-order polynomials, such as quadratic or cubic polynomials. Splines are used to describe complex shapes and curves. They enable a flexible and efficient representation of curves that can be easily adapted to different requirements. The coefficients of these polynomial segments are called spline coefficients.

[0028] Translating time series into splines can be used to limit the data to be monitored or analyzed to a few parameters. This reduces the amount of data to be processed, which is particularly resource-efficient, especially when using machine learning algorithms.

[0029] In one embodiment, the relevance values ​​correspond to the normalized residual values.

[0030] The relevance values ​​can be particularly simply the standardized residual values. Standardization is important for comparability. While it initially appears that a high residual value has a particularly high relevance, if the residual value is large because the values ​​of the measured time series and the reference time series are also large, then the value at this point is not necessarily meaningful for relevance.

[0031] During normalization, the residual values ​​are set in relation to the values ​​of the time series, preferably the reference time series. This can increase the accuracy of determining the relevance of the individual measured values ​​for the subsequent evaluation of the measured time series.

[0032] In one embodiment, the reference time series is an averaged time series from a plurality of measured time series.

[0033] The reference time series can be determined and provided in various ways. One possibility is to use historical data and work with average values. For example, a fixed number of measured time series can be cached, from which an average time series can then be calculated, which in turn forms the reference time series.

[0034] For example, 100 or 1,000 time series can be used for this purpose. The more time series used, the more representative the reference time series. However, with the number of time series, the storage and computing requirements for determining the reference time series also increase. Computing resources and storage requirements can be a limiting factor, especially in continuous real-time monitoring, when the reference time series is formed from the past x measurements in real time.

[0035] In one embodiment, the reference time series was determined using a model of the industrial process.

[0036] As an alternative to using historical data, synthetic time series can also be used. Synthetic time series can, for example, be generated using a model that simulates the industrial process being modeled and / or calculates the trend of the monitored process variable.

[0037] The advantage of a synthetic reference time series is that it can be generated in advance and does not require additional computing resources for real-time monitoring. Depending on the accuracy of the model, error ranges for the measured values ​​can also be taken into account when generating the reference time series, allowing errors to be estimated even when calculating the residual value.

[0038] In one embodiment, the method further comprises: - Detecting anomalies in the measured time series, where the anomalies are one or more measured values ​​with a relevance value above a defined threshold.

[0039] Anomaly detection is a data analysis method that automatically detects certain deviations in an industrial process that deviate from the norm or statistically expected behavior. The goal is to determine, through data analysis, whether a data point deviates from the normal or expected pattern or behavior. To do this, a reference must first be established to define what behavior can be considered normal.

[0040] If the relevance value for a measured value or a series of measured values ​​is particularly high, the measured time series deviates particularly strongly from the reference time series at that point. The sensitivity of the anomaly detection can be determined by the threshold. The lower the threshold, the more likely deviations are to be detected as anomalies. If the threshold is higher, the difference between the measured time series and the reference time series must also be greater to detect an anomaly.

[0041] In one embodiment, the method further comprises: - Providing another reference time series; - Providing a further measured time series, wherein the further measured time series comprises measured values ​​of a further process parameter of the industrial process; - Determination of residual values ​​for the further measured time series; - Determining a relevance value for each measured value of the further measured time series from the residual values; and - Determination of relevance times for the industrial process from the relevance values ​​of the measured time series.

[0042] When monitoring or analyzing industrial processes, several process parameters can be recorded simultaneously. For example, during a drilling operation, the borehole depth and the pressure on the drill bit can be recorded. If multiple parameters are recorded, they can also be monitored in parallel or evaluated sequentially.

[0043] When monitoring and / or analyzing multiple parameters, multiple reference time series must be available, namely at least one for each process parameter, so that the recorded data can be compared with the corresponding reference values.

[0044] The relevance values ​​of all measured time series can be summarized for the entire process into a relevance time series, in which each of the monitored parameters provides a portion of the relevance values. From the relevance time series, the most relevant points in time can be determined, which characterize the most relevant process steps, i.e., those that deviate most from the norm.

[0045] In a further aspect, the invention relates to a computer-implemented method for training a machine learning algorithm, wherein the training comprises the following steps: - Providing a training data set, wherein the training data set comprises a plurality of measured time series, wherein for each measured value of the time series a relevance value is determined using a method according to one of the preceding claims; - inputting the training dataset into the machine learning algorithm to train the machine learning algorithm; and - Deploy the trained machine learning algorithm.

[0046] The method described above can be used to train a machine learning algorithm to prepare the training data for training. During preprocessing, the measured time series can be combined with the relevance score to show the machine learning algorithm being trained which parts of the input time series are most relevant for evaluation.

[0047] By identifying the most relevant points in time in the industrial process, the training of the machine learning algorithm can be improved so that the trained machine learning algorithm can better solve its task.

[0048] Furthermore, the machine learning algorithm can be trained with time series of multiple parameters. Separate relevance values ​​are provided for each time series of each parameter, allowing the machine learning algorithm to process the time series independently and adapt to the relevance values.

[0049] In one embodiment, the machine learning algorithm comprises multiple models, each of which comprises at least one region. Each model is assigned to one or more regions and is configured to process the measured values ​​of the region(s) assigned to it.

[0050] The areas are determined using the following steps: - Determining a measured value in the measured time series whose relevance value is a local maximum or exceeds a first defined threshold, - Determine an area around the found measured value in which the relevance value is above a defined second threshold.

[0051] The reasons for deviations between the measured time curves and the respective reference time series can be diverse. Accordingly, it may be advantageous to use different models for monitoring and / or analysis for different process steps.

[0052] Based on the relevance values, the measured time series can be divided into buckets. The data from these buckets can then be input into different models, each selected according to the task at hand. This can improve the results of the monitoring or analysis.

[0053] The models may include, but are not limited to: linear models, decision trees, support vector machines, neural networks and many others.

[0054] The first and second thresholds can be used to set the areas to be selected. The lower the thresholds, the larger the areas, and vice versa.

[0055] In one embodiment, the machine learning algorithm is further trained to process range parameters, wherein the range parameters of the at least one range include the start of the range, the end of the range, the mean and / or the median of the measured time series in the range, the standard deviation of the measured time series in the range, the maximum relevance value in the range, the mean relevance value in the range, and / or other values ​​characteristic of the range.

[0056] The parameters mentioned can be easily determined for each domain. The amount of data to be processed by the machine learning algorithm is significantly reduced by selecting domain parameters, so the resources required for processing can also be reduced. Depending on the model, the model itself can even be simplified.

[0057] In one embodiment, each of the measured values ​​is assigned a weight for training the machine learning algorithm, wherein the weight correlates with the relevance of the respective measured value.

[0058] In some models, the training data can be assigned weights to reflect the relevance of a particular data point or data packet for training. Advantageously, the relevance values ​​can be used directly or indirectly as weights.

[0059] In a further aspect, the invention relates to a computer program with program code for carrying out a method as described above when the computer program is executed on a computer.

[0060] In a further aspect, the invention relates to a computer-readable data carrier with program code of a computer program for carrying out a method as described above when the computer program is executed on a computer.

[0061] In a further aspect, the invention relates to a system for quantifying the relevance of measured values ​​in time series, wherein the system is designed to carry out a method as described above.

[0062] In summary, the present invention provides a method for quantifying the relevance of measured values ​​in time series, a method for training a machine learning algorithm, a computer program, a computer-readable data carrier with program code and a system for quantifying the relevance of measured values ​​in time series.

[0063] The described designs and further training courses can be combined as desired.

[0064] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings

[0065] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate one embodiment and, in conjunction with the description, serve to explain principles and concepts of the invention.

[0066] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements shown in the drawings are not necessarily drawn to scale.

[0067] It shows: Fig. 1 schematically shows the sequence of the method according to one embodiment.

[0068] In the figure of the drawing, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.

[0069] Fig. 1 shows schematically the sequence of the method according to an embodiment.

[0070] The method begins in step S10 with the provision of a measured time series. The time series can be acquired or generated by any sensor or measuring device. Furthermore, the time series can also represent process parameters read from an industrial machine. For example, the time series can represent a voltage curve in an electric drive. The time series can therefore be provided by a sensor, measuring device, etc., or loaded from a buffer of a computing system.

[0071] In each case, the time series represents a physical quantity measured during an industrial process. The measured time series can be represented as f(t)

[0072] In embodiments, multiple time series can be provided, each reflecting the same physical quantity. This can be particularly the case with repeatable industrial processes, especially with work routines of work machines that are always performed in the same way. Even in the production of workpieces with consistently consistent properties, a measured time series can be provided for each workpiece.

[0073] In addition to the measured time series, a reference time series is provided in step S12. The reference time series specifies target values ​​for each point in time of the industrial process. It thus serves as a guide for the industrial process as it should proceed. The reference time series can be represented as fref(t).

[0074] If the measured time series represent several process parameters of the same industrial process, i.e., for example, if they were recorded in parallel or partially in parallel at the same plant, then a reference time series must be provided for each process parameter.

[0075] The time series are compared in step S14. During the comparison, residual values ​​are generated that indicate how far the measured time series deviates from the reference time series. The residual value can preferably be generated in two ways.

[0076] A first possibility is to subtract the measured time series from the reference time series and use the magnitude of the result as the residual value. The residual values ​​can be described as a function with fres(t)=|f(t)−fref(t)|.

[0077] In a second approach, the measured time series and the reference time series are approximated by splines. The difference between these spline coefficients generates another function in the time domain, from which the residual values ​​can be extracted.

[0078] The approximation can be represented, for example, by spline polynomials of n-th degree with fapprox(t)={fa(t)=antn+an−1tn−1+⋯+a1t+a0 for 0 <t≤tafb(t)=bntn+bn−1tn−1+⋯+b1t+b0 fu¨r ta<t≤tb⋮fx(t)=xntn+xn−1tn−1+⋯+x1t+x0 fu¨r tx−1<t.

[0079] For the residual values, for example, this can result in fres(t)={fres,a(t) for 0 <t≤tafres,b(t) fu¨r ta<t≤tb⋮fres,x(t) fu¨r tx−1<t with fres,a(t)=|an−ares,n|tn+|an−1−ares,n−1|tn−1+⋯+|a1−ares,1|t+|a0−ares,0|for 0 <t≤tafres,b(t)=|bn−bres,n|tn+|bn−1−bres,n−1|tn−1+⋯+|b1−bres,1|t+|b0−bres,0|fu¨r ta<t≤tb⋮fres,x(t)=|xn−xres,n|tn+|xn−1−xres,n−1|tn−1+⋯+|x1−xres,1|t+|x0−xres,0|fu¨r tx−1<t

[0080] In the final step S16, a relevance value is determined for each measured value from the associated residual value. The relevance value can preferably be determined by normalizing the residual value. This normalization advantageously makes relevance values ​​from different time series, even those with different dimensions, comparable. For example, the relevance of a temperature measurement can be compared with the relevance of a force measurement.

[0081] The relevance values ​​can be calculated, for example, with frel(t)=fres(t)fref(t)∗Max(fres(t)fref(t)).

Claims

[1] Computer-implemented method for quantifying the relevance of measured values ​​in time series, wherein the method comprises the following steps: - Providing a measured time series (S10) wherein the measured time series comprises a temporally ordered sequence of measured values, the measured values ​​being recorded as process parameters of an industrial process; - Provide a reference time series (S12); - Determining residual values ​​(S14) for the measured time series by comparing the measured time series with the reference time series; and - Determining a relevance value (S16) for each measured value of the measured time series from the residual values. [2] Computer-implemented method according to claim 1, wherein the residual values ​​are determined from the difference between the measured time series and the reference time series. [3] Computer-implemented method according to claim 1, comprising determining the residual values ​​(S14): - Determining spline coefficients for the reference time series; - Determining the spline coefficients for the measured time series; - Determining the difference or differences between the spline coefficients of the reference time series and the measured time series, whereby the residual values ​​are determined from the difference or differences of the spline coefficients; - Back-transforming the differences of the spline coefficients into the time domain and assigning the residual values ​​of the spline coefficients to the corresponding measured values ​​of the measured time series. [4] Computer-implemented method according to any of the preceding claims, wherein the relevance values ​​correspond to the normalized residual values. [5] Computer-implemented method according to any of the preceding claims, wherein the reference time series is an averaged time series from a plurality of measured time series or wherein the reference time series was determined using a model of the industrial process. [6] Computer-implemented method according to any one of the preceding claims, wherein the method further comprises: - Identifying anomalies in the measured time series, where the anomalies are one or more measured values ​​with a relevance value above a defined threshold. [7] Computer-implemented method according to any one of the preceding claims, wherein the method further comprises: - Providing an additional reference time series; - Providing an additional measured time series, wherein the additional measured time series includes measured values ​​of an additional process parameter of the industrial process; - Determining residual values ​​for the further measured time series; - Determining a relevance value for each measured value of the further measured time series from the residual values; and - Determining relevance points for the industrial process from the relevance values ​​of the measured time series. [8] Computer-implemented method for training a machine learning algorithm, wherein the training comprises the following steps: - Providing a training dataset, where the training dataset comprises a plurality of measured time series, wherein for each measured value of the time series a relevance value is determined using a method according to one of the preceding claims; - Feeding the training dataset into the machine learning algorithm to train the machine learning algorithm; and - Providing the trained machine learning algorithm. [9] Computer-implemented method according to claim 8, wherein the machine learning algorithm comprises multiple models, wherein the measured time series each comprise at least one range, wherein each model is assigned to one or more areas and is trained to process the measured values ​​of the area(s) assigned to it, the areas are determined using the following steps: - Determining a measured value in the measured time series whose relevance value is a local maximum or exceeds a first defined threshold, - Determining a range around the measured value in which the relevance value is above a defined second threshold. [10] Computer-implemented method according to claim 8 or 9, wherein the machine learning algorithm is further trained to process range parameters, the range parameters of the at least one range comprising the start of the range, the end of the range, the mean and / or median of the measured time series in the range, the standard deviation of the measured time series in the range, the maximum relevance value in the range, the mean relevance value in the range and / or other values ​​characteristic of the range. [11] Computer-implemented method according to one of claims 8 or 10, wherein each of the measured values ​​is assigned a weight for training the machine learning algorithm, the weight being correlated with the relevance of the respective measured value. [12] Computer program with program code to execute a method according to any one of claims 1 to 7 when the computer program is executed on a computer. [13] Computer-readable data carrier containing program code of a computer program for carrying out a method according to any one of claims 1 to 7 when the computer program is executed on a computer. [14] System for quantifying the relevance of measured values ​​in time series, wherein the system is configured to perform a method according to any one of claims 1 to 7.