Method for persisting data

The method addresses storage inefficiencies by clustering and selectively retaining data in technical systems, ensuring relevant data is preserved for analysis while reducing storage needs and maintaining data integrity.

WO2026092927A1PCT designated stage Publication Date: 2026-05-07SEW EURODRIVE GMBH & CO KG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SEW EURODRIVE GMBH & CO KG
Filing Date
2025-09-25
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing data persistence methods for technical systems consume excessive storage space and incur high costs due to the storage of large volumes of raw data, while historical raw data is often necessary for machine learning and anomaly identification.

Method used

A method involving data clustering, dimensional reduction, and selective data retention, where similar data sets are grouped into clusters, with only selected data sets and outliers retained, and others compressed or deleted, using techniques like autoencoders and principal component analysis.

Benefits of technology

Significantly reduces storage requirements by retaining relevant historical data while minimizing information loss, facilitating efficient data management and anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025077425_07052026_PF_FP_ABST
    Figure EP2025077425_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A method for persisting data for assessing a technical installation is proposed. This involves a plurality of data sets being recorded, each of which comprises a plurality of measured variables, each measured variable comprising a plurality of measured values. The recorded data sets are stored in a data memory. The recorded data sets are clustered, data sets that have a relatively high degree of similarity to one another being identified as belonging to a common cluster. At least one data set from each identified cluster is selected, and at least one data set from at least one identified cluster is deselected. All measured values in the selected data sets are retained in the data memory, and all measured values in the deselected data sets are erased from the data memory.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Methods for persisting data

[0002] Description:

[0003] The invention relates to a method for persisting data for the evaluation of a technical system, wherein a plurality of data sets are recorded, each comprising a plurality of measured variables, wherein each measured variable comprises a plurality of measured values, and wherein the recorded data sets are stored in a data storage device.

[0004] Technical systems, such as drive systems, comprise a multitude of technical components, including an electric motor, a converter or inverter to generate three-phase alternating current for the electric motor, and a gearbox to reduce the motor's speed. Examples of such technical systems include rotary tables, conveyor belts, and stacker cranes. After prolonged operation, the components of these systems can malfunction due to wear and tear.

[0005] It is known to monitor such technical systems by recording and evaluating specific measurements at set times. If the recorded measurements deviate significantly from predefined target values, a fault in the system is assumed, and a corresponding message is sent to the operator. For example, current generated by the converter or inverter that exceeds a defined limit may indicate stiffness in the gearbox due to wear.

[0006] Condition monitoring involves collecting a relatively large amount of data from a technical system for later analysis. This data consumes storage space and incurs costs. A proven strategy to reduce costs is to aggregate the raw data directly and store only the aggregated data. However, in certain cases, historical raw data is also relevant, for example, for training machine learning models or for visualizing the raw data to identify anomalies in the event of a fault.

[0007] ISI \ EIDOPAT 25.09.2025 Generic procedures are known from the following documents:

[0008] - PAN, Feng, et al. Finding representative set from massive data. In: Fifth IEEE International Conference on Data Mining (ICDM'05). IEEE, 2005. p. 8 pp.

[0009] - Autoencoder. In: Wikipedia, The Free Encyclopedia.

[0010] URL: https: / / de. wikipedia. org / w / index.php?title=Autoencoder&oldid=238636189.

[0011] - JOSIGER, Marcus; KIRCHNER, Kathrin. Modern clustering algorithms—a comparative analysis on two-dimensional data. In: Proc. FGML Workshop (FGML 2003), p. 2003. p. SO- 84.

[0012] - Data compression. In: Wikipedia, The Free Encyclopedia.

[0013] URL: https: / / de. wikipedia. org / w / index.php?title=Data Compression&oldid=240585389.

[0014] - LIU, Tong, et al. High-ratio lossy compression: Exploring the autoencoder to compress scientific data. IEEE Transactions on Big Data, 2021, Volume 9, No. 1, pp. 22-36.

[0015] - Envelope analysis. Last updated: February 27, 2024. URL: https: / / sensemore.io / de / hullkurvenanalyse / archived in: https: / / web.archive.Org / web / 20240227231426 / https: / / sensemore.io / de / hullkurvenanalyse / on February 27, 2024.

[0016] - K-Means algorithm. In: Wikipedia, The Free Encyclopedia. Revision as of 9.

[0017] URL: https: / / de. wikipedia. org / w / index.php?title=K-MeansAlgorithm&oldid=225216308.

[0018] - Determinantal point process. In: Wikipedia, The Free Encyclopedia. Revision as of 26 May 2023. URL: https: / / de.wikipedia.org / w / index.php?title=Determinantal_point_process&oldid=234045844.

[0019] The invention is based on the objective of further developing a method for persisting data.

[0020] The problem is solved by a method for persisting data with the features specified in claim 1. Advantageous embodiments and further developments are the subject of the dependent claims.

[0021] A method for persisting data for the evaluation of a technical system is proposed. This involves recording multiple data sets, each comprising multiple measured variables, with each measured variable comprising multiple measured values. The recorded data sets are stored in a data repository. The recorded data sets are then clustered, with data sets exhibiting a relatively high degree of similarity being identified as belonging to the same cluster. At least one data set is selected from each identified cluster, and at least one data set is deselected from each identified cluster. All measured values ​​of the selected data sets are retained in the data repository, and all measured values ​​of the deselected data sets are deleted from the data repository.

[0022] The method according to the invention is particularly suitable for persisting data for the evaluation of a technical system. The method according to the invention allows relevant historical measurement values ​​to be persisted, while redundant or less relevant data can be discarded. This results in a significant reduction in the amount of data to be stored, with virtually no loss of information.

[0023] According to an advantageous embodiment of the invention, a dimensional reduction of the recorded data sets is carried out before the clustering of the recorded data sets is performed.

[0024] Dimensional reduction is performed, for example, using principal component analysis. Principal component analysis is a mathematical method for approximating a large number of statistical variables by using a smaller number of highly informative linear combinations of those variables. Principal component analysis is described, for example, in the document "An Introduction to Statistical Learning," page 374.

[0025] According to an advantageous embodiment of the invention, data sets which do not exhibit a sufficiently high degree of similarity to any of the other data sets are selected.

[0026] In particular, outliers that do not exhibit a sufficiently high degree of similarity to any of the other datasets and therefore do not belong to any identified cluster are selected. These outliers are relevant, for example, for training machine learning models and allow, among other things, the identification of anomalies during subsequent error analysis.

[0027] According to an advantageous embodiment of the invention, the data records are recorded during a defined period of time and stored in the data storage device. The clustering of the recorded data records is then performed only after the defined period has elapsed. In order to maintain a relevant variance of the data records received during the period in the stored data records without having to store all measured values, the measured values ​​are selected based on the clusters using a specific strategy.

[0028] According to an advantageous embodiment of the invention, a compressed representation of the deselected data sets is calculated from the measured values ​​of the deselected data sets and stored in the data storage.

[0029] According to one possible embodiment, the compressed representation of the deselection data records is calculated separately for each cluster. According to an alternative embodiment, the compressed representation of the deselection data records is calculated across multiple clusters, in particular across all clusters.

[0030] According to an advantageous embodiment of the invention, a compressed representation of the selected data sets is calculated from the measured values ​​of the selected data sets and stored in the data storage.

[0031] According to one possible embodiment, the compressed representation of the selected data records is calculated separately for each cluster. According to an alternative embodiment, the compressed representation of the selected data records is calculated across multiple clusters, in particular across all clusters.

[0032] According to an advantageous embodiment of the invention, the compressed representation of the data sets is calculated by lossy compression of the measured values ​​of the data sets.

[0033] Lossy compression of the measured values ​​of the data sets allows for a later reconstruction of the data sets, resulting in only a relatively small loss of information.

[0034] According to an advantageous embodiment of the invention, the compressed representation of the data sets is calculated using a machine learning model trained on the measured values ​​of the data sets. Such a machine learning model is, for example, an autoencoder trained on the measured values ​​of the data sets. The autoencoder comprises an encoder and a decoder. The measured values ​​are converted into a latent space by the encoder. The latent space has a much lower dimensionality than the original space of the measured values. The decoder creates a sufficiently accurate representation of the data sets from the original space using the measured values ​​converted into the latent space. The autoencoder and the representation of the data sets in the latent space are stored, and the data sets from the original space are deleted. The latent representations and the autoencoder can be loaded as needed.Using the decoder of the autoencoder, a sufficiently accurate representation of the measured values ​​of the data sets can be reconstructed.

[0035] According to an advantageous embodiment of the invention, the compressed representation of the data sets is calculated by simplifying the measured values ​​of the data sets.

[0036] The line simplification of the measured values ​​in the datasets is calculated, for example, using the algorithm by Ramer, Douglas, and Peucker (https: / / utpjournals.press / doi / abs / 10.3138 / FM57-6770-U75U-7727). Such line simplification is also referred to as a "Line Simplification Algorithm".

[0037] According to an advantageous embodiment of the invention, aggregated values ​​are calculated from the measured values ​​of the deselected data records before their deletion, and the aggregated values ​​of the deselected data records are stored in the data storage.

[0038] Aggregated values ​​from datasets include, for example, a mean value of the measured values ​​and / or a maximum measured value and / or a minimum measured value.

[0039] According to an advantageous embodiment of the invention, aggregated values ​​are calculated from the measured values ​​of the selected data sets, and the aggregated values ​​of the selected data sets are stored in the data memory.

[0040] Aggregated values ​​from data sets include, for example, a mean value of the measured values ​​and / or a maximum measured value and / or a minimum measured value. According to an advantageous embodiment of the invention, the aggregated values ​​are calculated as envelopes to the data sets.

[0041] The aggregated values ​​are preferably calculated separately for individual clusters.

[0042] According to an advantageous embodiment of the invention, exactly one data record is selected from at least one detected cluster, and the number of data records belonging to the cluster is determined. The determined number of data records belonging to the cluster is then assigned to the selected data record and stored in the data storage.

[0043] According to an advantageous embodiment of the invention, exactly one data record is selected from at least one identified cluster, and the variance of the data records belonging to the cluster is determined. The determined variance of the data records belonging to the cluster is assigned to the selected data record and stored in the data storage.

[0044] According to an advantageous embodiment of the invention, a data set is selected from at least one detected cluster which is closest to a center point of the cluster.

[0045] According to an advantageous embodiment of the invention, several data sets are selected from at least one detected cluster, each of which lies at an edge of the respective cluster.

[0046] According to an advantageous embodiment of the invention, several data sets are selected from at least one detected cluster, wherein the distances between the selected data sets are maximally large.

[0047] According to an advantageous embodiment of the invention, several data sets are selected from at least one detected cluster, wherein the selection of the data sets is random.

[0048] According to an advantageous embodiment of the invention, several data records are selected from at least one identified cluster, wherein clustering of the data records within the cluster is performed, and wherein data records exhibiting a relatively high degree of similarity to one another are identified as belonging to a common subcluster. At least one data record is then selected from each identified subcluster. According to another advantageous embodiment of the invention, several data records are selected from at least one identified cluster, maximizing the diversity of the selected data records. The diversity of the data records is calculated based on a probability by using sampling via the Determinant Point Process.

[0049] The invention is not limited to the combination of features in the claims. For a person skilled in the art, further meaningful combinations of claims and / or individual claim features and / or features of the description and / or the figures will arise, in particular from the problem statement and / or the problem arising from a comparison with the prior art.

[0050] The invention will now be explained in more detail with reference to the illustrations. The invention is not limited to the embodiments shown in the illustrations. The illustrations only depict the subject matter of the invention schematically. They show:

[0051] Figure 1: a flowchart of a data set recording process,

[0052] Figure 2: a flowchart of the processing of recorded data sets and

[0053] Figure 3: a schematic representation of a clustering of the recorded data sets.

[0054] Figure 1 shows a flowchart of the acquisition of a data set 10 in a technical system. In this case, the technical system is a drive system. The drive system comprises an electric motor, a converter or inverter for generating a three-phase alternating voltage for the electric motor, and a gearbox for reducing the speed of the electric motor.

[0055] First, a data set 10 is recorded. This data set 10 comprises a number of measured variables, for example, eight different measured variables. These measured variables include, for example, DC link voltage, current, frequency, speed, torque, and possibly others. In particular, the measured variables are the speed and torque of the electric motor. These measured variables, speed and torque, are, for example, measured indirectly via the converter or inverter by measuring the frequency and current of the output current to the electric motor.

[0056] Each measurand comprises a plurality of measured values, for example, 2048 measured values. The measured values ​​of each measurand are recorded sequentially. For example, the measured values ​​of each measurand are recorded at equidistant time intervals of 5 ms each. The recorded dataset 10 thus comprises, for example, eight different measurands, each with 2048 measured values.

[0057] In a first step 101, the recorded data record 10 is stored in a data storage location 20. Several more data records 10 are recorded. For example, between ten and one hundred data records 10 are recorded daily during a month. The first step 101 is repeated, and the recorded data records 10 are each stored in the data storage location 20. Another data record 10 is recorded, for example, whenever a defined trigger condition is met. Such a trigger condition is, for example, the start of an application cycle in the technical system when the rotational speed of an electric motor in the technical system exceeds a predefined minimum speed.

[0058] Figure 2 shows a flowchart of the processing of recorded data records 10. In step 102, the previously recorded and stored data records 10 are loaded from the data storage 20.

[0059] In step 103, a dimensionality reduction of the recorded data sets 10 is first performed, for example by means of a principal component analysis. During the principal component analysis, a transformation matrix is ​​determined and applied to the data sets 10.

[0060] In step 104, the recorded data records 10 are clustered. Data records 10 that exhibit a relatively high degree of similarity to each other belong to a common cluster and are recognized as belonging to a common cluster. The clusters are preferably identified automatically using a clustering method, such as DBSCAN. Alternatively, the clusters are identified by manual assignment.

[0061] Data records 10 that do not exhibit a sufficiently high degree of similarity to any of the other data records 10 do not belong to any cluster. Such data records 10 are referred to as outliers.

[0062] If necessary, in addition to the newly recorded data sets 10, data sets 10 from previous measurement recordings and which have already been stored in the data storage 20 are also taken into account during clustering.

[0063] In step 105, a selection and deselection of data records 10 takes place. At least one data record 10, for example, exactly one data record 10, is selected from each identified cluster. The remaining data records 10 of the respective cluster are deselected. Data records 10, designated as outliers, which do not exhibit a sufficiently high degree of similarity to any of the other data records 10, are also selected.

[0064] All measured values ​​from all selected data sets 10 are retained in data storage 20. All measured values ​​from the deselection data sets 10 are deleted from data storage 20.

[0065] If necessary, aggregated values ​​are calculated from the measured values ​​of the deselected data records 10 before their deletion, and the aggregated values ​​of the deselected data records 10 calculated in this way are stored in the data storage 20.

[0066] Figure 3 shows a schematic representation of a clustering of the recorded data sets 10. Each data set 10 comprises, as already mentioned, a plurality of measured quantities with a plurality of measured values.

[0067] In step 103, as already mentioned, the dimensionality reduction of the recorded data sets 10 is carried out, and the data sets 10 are represented in a feature space.

[0068] In the representation shown here, the data records 10 are displayed only two-dimensionally for the sake of simplicity. Each of the points shown corresponds to one data record 10.

[0069] Data records 10 that exhibit at least a relatively high degree of similarity to each other belong to a common cluster and are located close to each other in the feature space. Data records 10 belonging to a common cluster are recognized as belonging to a common cluster. Recognized clusters are each marked with a circle in the representation shown here.

[0070] Those data records 10 that do not exhibit a sufficiently high degree of similarity to any of the other data records 10 do not belong to any cluster. These data records 10 lie outside the circles in the representation shown here and are marked as outliers. Therefore, in this representation, two data records 10 are marked as outliers.

[0071] As mentioned previously, one data record 10 is selected from each identified cluster. The remaining data records 10 of the respective cluster are deselected. Data records 10 designated as outliers, which do not exhibit a sufficiently high degree of similarity to any of the other data records 10, are also selected.

[0072] In the representation shown here, the selected data records (10) are marked with a cross. The unselected data records (10) are not marked.

[0073] Reference symbol list

[0074] 10 data records 20 data storage

[0075] 101-105 steps

Claims

Patent claims:

1. A method for persisting data for the evaluation of a technical system, wherein a plurality of data records (10) are recorded, each comprising a plurality of measured variables, wherein each measured variable comprises a plurality of measured values, and wherein the recorded data records (10) are stored in a data storage device (20), characterized in that a clustering of the recorded data records (10) is carried out, wherein data records (10) which exhibit a relatively high similarity to each other are recognized as belonging to a common cluster, and that at least one data record (10) is selected from each recognized cluster, and that at least one data record (10) is deselected from at least one recognized cluster, and that all measured values ​​of the selected data records (10) are retained in the data storage device (20), and that all measured values ​​of the deselected data records (10) are deleted from the data storage device (20).

2. Method according to one of the preceding claims, characterized in that a dimensionality reduction of the recorded data sets (10) is carried out before the clustering of the recorded data sets (10) is carried out.

3. Method according to one of the preceding claims, characterized in that data records (10) which do not exhibit a sufficiently high degree of similarity to any of the other data records (10) are selected.

4. Method according to one of the preceding claims, characterized in that the data records (10) are recorded during a defined period of time and stored in the data storage (20), and that the clustering of the recorded data records (10) is carried out after the defined period of time has elapsed.

5. Method according to one of the preceding claims, characterized in that a compressed representation of the deselected data sets (10) is calculated from the measured values ​​of the deselected data sets (10) and stored in the data storage (20).

6. Method according to one of the preceding claims, characterized in that a compressed representation of the selected data sets (10) is calculated from the measured values ​​of the selected data sets (10) and stored in the data storage (20).

7. Method according to one of claims 5 to 6, characterized in that the compressed representation of the data sets (10) is calculated by lossy compression of the measured values ​​of the data sets (10), or that the compressed representation of the data sets (10) is calculated using a machine learning model trained on the measured values ​​of the data sets (10), or that the compressed representation of the data sets (10) is calculated by line simplification of the measured values ​​of the data sets (10).

8. Method according to one of the preceding claims, characterized in that aggregated values ​​are calculated from the measured values ​​of the deselected data records (10) before their deletion, and that the aggregated values ​​of the deselected data records (10) are stored in the data storage (20). - 15 - 9. Method according to one of the preceding claims, characterized in that aggregated values ​​are calculated from the measured values ​​of the selected data sets (10), and that the aggregated values ​​of the selected data sets (10) are stored in the data storage (20).

10. Method according to one of the preceding claims, characterized in that exactly one data record (10) is selected from at least one detected cluster, and that a number of data records (10) belonging to the cluster is determined, and that the determined number of data records (10) belonging to the cluster is assigned to the selected data record (10) and stored in the data storage (20).

11. Method according to one of the preceding claims, characterized in that exactly one data record (10) is selected from at least one identified cluster, and that a variance of the data records (10) belonging to the cluster is determined, and that the determined variance of the data records (10) belonging to the cluster is assigned to the selected data record (10) and stored in the data storage (20).

12. Method according to one of the preceding claims, characterized in that several data sets (10) are selected from at least one detected cluster, each of which is located at an edge of the respective cluster.

13. Method according to one of the preceding claims, characterized in that several data records (10) are selected from at least one detected cluster, wherein distances between the selected data records (10) are maximally large.

14. Method according to one of the preceding claims, characterized in that several data records (10) are selected from at least one detected cluster, wherein a clustering of the data records (10) of the cluster is carried out, wherein Data records (10) which exhibit a relatively high similarity to each other are recognized as belonging to a common subcluster, and at least one data record (10) is selected from each recognized subcluster. - 16 - 15. Method according to one of the preceding claims, characterized in that several data sets (10) are selected from at least one identified cluster, wherein the diversity of the selected data sets (10) is maximal, and wherein the diversity of the data sets (10) is calculated on the basis of a probability by using sampling via Determinantal Point Process.

Citation Information

Patent Citations

  • Condition monitoring of an electric power converter

    US20230281478A1