Method and apparatus for evaluating the value of time-series data information

In industrial timing data analysis, the value benchmark template and slitting point are determined according to the business analysis goals, the timing data fragments and the benchmark template are calculated in a similarity manner, and information value is assigned, and the information value density is used to locate high information value timing data, the problem of low efficiency of timing data analysis in the existing technology is solved, and fast and accurate high information value timing data positioning is achieved.

CN116738203BActive Publication Date: 2025-05-30CERI DIGITAL TECHNOLOGY (BEIJING) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310555476.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2025-05-30
Estimated Expiration
2043-05-17

AI Technical Summary

Technical Problem

The existing technology lacks an effective and objective method for evaluating industrial timing data information, resulting in low efficiency in timing data analysis.

Method used

By determining multiple value benchmark templates based on pre-set business analysis goals, and dividing the time series data to be analyzed into multiple time series data fragments, calculating the similarity between each fragment and the reference template, giving information value, and positioning high information value timing data through information value density.

Benefits of technology

It realizes the rapid and accurate positioning of time series data with high information value, improves the efficiency of time series data analysis, and can assist personnel in quickly positioning of time series data with high information value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116738203B_ABST
    Figure CN116738203B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for evaluating the value of time-series data information, which relates to the field of information technology. The method includes: dividing the time-series data to be analyzed into multiple time-series data segments according to the segmentation point flag; assigning an information value of 0 to the first time-series data segments that are irrelevant to the business analysis objective; for the second time-series data segments that are relevant to the business analysis objective, determining the similarity between each second time-series data segment and each value benchmark template; determining the characteristic value of the second time-series data segment according to the multiple similarity values corresponding to each second time-series data segment; obtaining the information value of each second time-series data segment according to the characteristic value of each second time-series data segment; and obtaining the information value density of the time-series data to be analyzed according to the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment. The present invention can assist personnel in quickly locating time-series data with high information value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and particularly to a method and device for evaluating the value of time-series data information. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The description herein is not admitted to be prior art merely by inclusion in this section.

[0003] Industrial data acquisition systems need to collect and analyze time-series data. When analyzing, it is mainly time-interval analysis, or length conversion is performed on the basis of equal time intervals, and then equal-length analysis is carried out. However, the actual situation is that the information value provided by the data generated in each time period or length segment is not consistent. For normal situations, most do not need to be analyzed, while for abnormal situations, when analyzing, it is necessary to manually drag the time axis or length axis segment by segment to search for a large amount of data over a long time, resulting in a waste of time and low efficiency. In the actual production process, the time-series data of normal production has its inherent pattern, while that of abnormal production often has a relatively special pattern. The best way to analyze is to be able to pay attention to the entire trend process of a continuous long time and the details of abnormal time periods, so as to improve the analysis efficiency. Therefore, when the analysis system displays data, it should make a choice. As much data as possible should be displayed for time periods with high information value, and as little data as possible should be displayed for time periods with low information value. However, there is currently no unified method for defining information value. Especially for industrial systems, analyzing solely from the perspective of data will not yield an effective method. It is necessary to combine the actual business scenario to give full play to the information value of time-series data. Currently, there is a lack of an effective and objective method for evaluating the information value of industrial time-series data, resulting in low efficiency of time-series data analysis. Summary of the Invention

[0004] Embodiments of the present invention provide a method for evaluating the value of time-series data information, which is used to accurately and quickly locate time-series data with high information value and improve the efficiency of time-series data analysis. The method includes:

[0005] Determine a plurality of value benchmark templates according to a preset business analysis target; wherein, the value benchmark template is a time-series data segment generated in a normal production mode, and the business analysis target includes: the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal production curve;

[0006] Determine a time-series data segmentation point flag according to a preset business analysis target;

[0007] Divide the time-series data to be analyzed into multiple time-series data segments according to the segmentation point flag;

[0008] For a first time-series data segment that is irrelevant to the business analysis objective, assign an information value of 0;

[0009] For a second time-series data segment that is relevant to the business analysis objective, determine the similarity between each second time-series data segment and each value benchmark template, and obtain multiple similarity values corresponding to each second time-series data segment;

[0010] Determine the maximum value among the multiple similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment;

[0011] Normalize the characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective;

[0012] Based on the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment, obtain the information value density of the time-series data to be analyzed under the business analysis objective.

[0013] An embodiment of the present invention further provides a time-series data information value evaluation device for accurately and quickly locating time-series data with high information value and improving the efficiency of time-series data analysis. The device includes:

[0014] A first processing module for determining multiple value benchmark templates according to a preset business analysis objective; wherein, the value benchmark template is a time-series data segment generated in a normal production mode, and the business analysis objective includes: the deviation degree of any one or more data indicators in the actual production curve compared to the corresponding data indicators in the normal production curve;

[0015] A second processing module for determining a time-series data segmentation point flag according to a preset business analysis objective;

[0016] A third processing module for dividing the time-series data to be analyzed into multiple time-series data segments according to the segmentation point flag;

[0017] An assignment module for assigning an information value of 0 to a first time-series data segment that is irrelevant to the business analysis objective;

[0018] A similarity calculation module for determining the similarity between each second time-series data segment and each value benchmark template for a second time-series data segment that is relevant to the business analysis objective, and obtaining multiple similarity values corresponding to each second time-series data segment;

[0019] A fourth processing module for determining the maximum value among the multiple similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment;

[0020] A fifth processing module, configured to normalize the eigenvalue of each second timing data segment according to each second timing, so as to obtain the information value of each second timing data segment under the business analysis objective;

[0021] A sixth processing module, configured to obtain the information value density of the to-be-analyzed timing data under the business analysis objective according to the data length of the to-be-analyzed timing data, the information value of the first timing data segment, and the information value of the second timing data segment for each second timing.

[0022] An embodiment of the present invention further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned timing data information value evaluation method is implemented.

[0023] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned timing data information value evaluation method is implemented.

[0024] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned timing data information value evaluation method is implemented.

[0025] In an embodiment of the present invention, multiple value benchmark templates are determined according to a preset business analysis objective; wherein, the value benchmark template is a time series data segment generated in a normal production mode, and the business analysis objective includes: the degree of deviation of any one or more data indicators in the actual production curve compared to the corresponding data indicators in the normal production curve; according to the preset business analysis objective, a time series data segmentation point flag is determined; the time series data to be analyzed is divided into multiple time series data segments according to the segmentation point flag; for the first time series data segment that has nothing to do with the business analysis objective, an information value of 0 is assigned; for the second time series data segment that is related to the business analysis objective, the similarity between each second time series data segment and each value benchmark template is determined, and multiple similarity values corresponding to each second time series data segment are obtained; the maximum value among the multiple similarity values corresponding to each second time series data segment is determined as the characteristic value of each second time series data segment; according to the normalization processing of the characteristic value of each second time series data segment, the information value of each second time series data segment under the business analysis objective is obtained; according to the data length of the time series data to be analyzed, the information value of the first time series data segment, and the information value of the second time series data segment, the information value density of the time series data to be analyzed under the business analysis objective is obtained. In this way, according to the information value density of the time series data, whether it is data visualization analysis or anomaly point search, it can assist personnel in quickly locating high-information-value time series data and improving the analysis efficiency. Description of the Drawings

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. In the drawings:

[0027] Figure 1 It is a flowchart of a method for evaluating the information value of time series data provided in an embodiment of the present invention;

[0028] Figure 2 It is an example diagram of a value benchmark template provided in an embodiment of the present invention;

[0029] Figure 3 It is an example diagram of time series data to be analyzed provided in an embodiment of the present invention;

[0030] Figure 4 It is a flowchart of a method for obtaining the information value of each second time series data segment under the business analysis objective according to the normalization processing of the characteristic value of each second time series data segment provided in an embodiment of the present invention;

[0031] Figure 5 This is a flowchart of a method for obtaining the information value of each second time-series data segment under the business analysis objective by normalizing the eigenvalue of each second time-series data segment according to the present invention embodiment;

[0032] Figure 6 This is a flowchart of a method for obtaining the information value density of the time-series data to be analyzed under the business analysis objective according to the present invention embodiment;

[0033] Figure 7 This is a schematic diagram of a time-series data information value evaluation device provided by the present invention embodiment;

[0034] Figure 8 This is a schematic diagram of a computer device provided by the present invention embodiment. Detailed implementation manners

[0035] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer and more understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.

[0036] In the technical solutions of the present application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of national laws and regulations.

[0037] The term "and / or" herein merely describes an association relationship and means that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.

[0038] In the description of this specification, the terms "comprising", "including", "having", "containing", etc. are all open-ended terms, that is, they are intended to include but not limited to. The description with reference to terms such as "an embodiment", "a specific embodiment", "some embodiments", "for example", etc. means that the specific features, structures or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. The step sequences involved in each embodiment are used to schematically illustrate the implementation of the present application, and the step sequences are not limited and can be adjusted appropriately as needed.

[0039] It has been found through research that an industrial data acquisition system needs to collect and analyze time-series data. During analysis, it is mainly time-interval analysis, or length conversion is performed based on equal time intervals, followed by equal-length analysis. However, the actual situation is that the information value provided by the data generated in each time period or length segment is not consistent. For normal situations, most do not require analysis, while for abnormal situations, during analysis, it is necessary to manually drag the time axis or length axis segment by segment to search for a large amount of data over a long time, resulting in a waste of time and low efficiency. In the actual production process, the time-series data of normal production has its inherent pattern, while that of abnormal production often has a special pattern. The best way during analysis is to be able to pay attention to the entire trend process of a continuous long time and also pay attention to the detailed parts of abnormal time periods to improve the analysis efficiency. Therefore, when the analysis system displays data, it should make a trade-off. As much data as possible should be displayed for time periods with high information value, and as little data as possible should be displayed for time periods with low information value. However, there is currently no unified method for defining information value. Especially for industrial systems, analyzing solely from the perspective of data will not yield an effective method. It is necessary to combine the actual business scenario to bring the information value of time-series data into play. Currently, the lack of an effective and objective evaluation method for the information value of industrial time-series data leads to low efficiency in time-series data analysis.

[0040] In response to the above research, as Figure 1 shown, an embodiment of the present invention provides a method for evaluating the information value of time-series data, including:

[0041] S101: Determine a plurality of value benchmark templates according to a pre-set business analysis target; wherein, the value benchmark template is a time-series data segment generated in a normal production mode, and the business analysis target includes: the deviation degree of any one or more data indicators in the actual production curve compared to the corresponding data indicators in the normal production curve;

[0042] S102: Determine a time-series data segmentation point flag according to a pre-set business analysis target;

[0043] S103: Divide the time-series data to be analyzed into a plurality of time-series data segments according to the segmentation point flag;

[0044] S104: Assign an information value of 0 to the first time-series data segment that has nothing to do with the business analysis target;

[0045] S105: For the second time-series data segment related to the business analysis target, determine the similarity between each second time-series data segment and each value benchmark template, and obtain a plurality of similarity values corresponding to each second time-series data segment;

[0046] S106: Determine the maximum value among the multiple similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment;

[0047] S107: Normalize the characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective;

[0048] S108: Based on the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment, obtain the information value density of the time-series data to be analyzed under the business analysis objective.

[0049] In the embodiments of the present invention, according to a preset business analysis objective, a plurality of value benchmark templates are determined; wherein, the value benchmark template is a time-series data segment generated in the normal production mode, and the business analysis objective includes: the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal production curve; according to the preset business analysis objective, determine the time-series data segmentation point flag; divide the time-series data to be analyzed into multiple time-series data segments according to the segmentation point flag; for the first time-series data segment irrelevant to the business analysis objective, assign an information value of 0; for the second time-series data segment relevant to the business analysis objective, determine the similarity between each second time-series data segment and each value benchmark template to obtain multiple similarity values corresponding to each second time-series data segment; determine the maximum value among the multiple similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment; normalize the characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective; based on the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment, obtain the information value density of the time-series data to be analyzed under the business analysis objective. In this way, according to the information value density of the time-series data, whether it is data visualization analysis or anomaly point search, it can assist personnel in quickly locating high-information-value time-series data and improving the analysis efficiency.

[0050] The above time-series data information value evaluation method will be described in detail below.

[0051] Regarding the above S101, the business analysis objective may include, for example, the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal production curve, and the value benchmark template is a time-series data segment generated in the normal production mode.

[0052] Specifically, when a production line or equipment malfunctions, technicians or industry experts need to analyze based on the recorded process data curves to find the cause of the malfunction. The degree of deviation of the analyzed curve from the normal production curve can be regarded as the business analysis objective. For example, to locate the malfunction of a loop, one of the business analysis objectives can be: "Analyze the abnormal condition of the height curve of the bar and wire loop under load compared to the height curve during normal production". When analyzing this business objective, according to the segmentation point, the curve segments to be analyzed can be segmented. The segmentation point of the above-mentioned loop height curve under load refers to the critical point that can divide between the loop under load and no load.

[0053] Another example: Analyze the fluctuation of the torque of a certain rolling mill stand under load, and set the business analysis objective as analyzing whether the torque fluctuates too much and exceeds the designed working range. The segmentation point is the critical point between the rolling mill stand under load and no load (when under load, the torque is large, and when no load, the torque is small, and the segmentation point can be found through the torque value).

[0054] In addition, for the same analysis objective, multiple value benchmark templates can be selected. For example Figure 2 as shown, value benchmark templates T1 and T2 are selected according to the business analysis objective. The value benchmark templates can segment and select the time-series data generated during the normal production mode based on the actual scenario and historical experience.

[0055] For the above S102, according to the preset business analysis objective, determine the time-series data segmentation point flag. For example Figure 3 as shown, the business analysis objective is to analyze the abnormal condition of the height curve of the bar and wire loop under load, and use the switching point between the equipment under load and no load as the segmentation point flag, that is Figure 3 the position where the data in [ ] rapidly rises from near 0 and rapidly drops from a high position to near 0.

[0056] For the above S103, the time-series data to be analyzed is, for example, long-period time-series data X = (X 1 , X 2 ... X n ). Then, it is divided into groups according to the segmentation flag. Segment 1 is Y 1 = (X 1 , X 2 ... X 120 ), segment 2 is Y 2 = (X 121 , X 122 ... X 260 ), etc.

[0057] Here, the segment extraction is based on the segmentation flag, and the length of the time-series data segments after segmentation (the number of elements within the time-series data segments) can be equal or unequal.

[0058] Regarding the above S104, as Figure 3 shown, the time-series data to be analyzed is divided into five groups of time-series data segments: Y1 to Y5 according to the segmentation point flag. For the first time-series data segments Y2 and Y4 that are irrelevant to the business analysis target, the information value is directly assigned as 0, that is, I2 = I4 = 0. These two segments are no-load period curves and are irrelevant to the load period to be analyzed by the business target.

[0059] Regarding the above S105, for calculating the similarity, for example, the correlation coefficient, Euclidean distance, pattern distance, shape distance, DWT, etc. can be selected. For some algorithms, data preprocessing is required, such as: upsampling or downsampling algorithms to ensure that the data segment length is the same as the data length of the value benchmark template.

[0060] Therefore, as Figure 4 shown, the embodiment of the present invention provides a method for obtaining the information value of each second time-series data segment under the business analysis target by normalizing the feature value of each second time-series data segment, including:

[0061] S401: Sample each second time-series data segment to obtain the first sampled data of each second time-series data.

[0062] S402: Sample each value benchmark template to obtain the second sampled data of each value benchmark template; wherein, the data length of the first sampled data is equal to the data length of the second sampled data.

[0063] S403: Calculate the similarity between the first sampled data of each second time-series data and the second sampled data of each value benchmark template.

[0064] In addition, various interpolation and resampling algorithms can be selected for the upsampling or downsampling algorithm.

[0065] Exemplarily, as Figure 3 shown in the time-series data segment, calculate the similarity between the second time-series data segments Y1, Y3, and Y5 related to the business analysis target and the value benchmark templates T1 and T2. Taking Table 1 and the Pearson correlation coefficient calculation as an example:

[0066] Table 1 List of time-series data segments

[0067]

[0068] As shown in Table 1, the data lengths of Y1, Y3, Y5, T1, and T2 are all inconsistent. Resampling alignment needs to be performed first to ensure the same data length. Specific implementation operations include, for example: Y1’[i] = Y1[int(i * length of Y1 / length of T1)]. Where Y1’ is the resampled time series data segment of Y1, Y1’[i] represents the i-th element of Y1’, and int represents taking the integer within the parentheses. Then, Y1, Y3, and Y5 aligned with the template T1 according to resampling are Y1’(T1), Y3’(T1), and Y5’(T1), and Y1, Y3, and Y5 aligned with the template T2 are Y1’(T2), Y3’(T2), and Y5’(T2), as shown in Table 1. The aligned data is respectively calculated for correlation with the templates T1 and T2. The correlation calculation results are shown in the T1 column and T2 column of Table 2 below:

[0069] Table 2 Correlation Coefficient, Eigenvalue, and Information Value Table

[0070] T1 T2 v I Y1' 0.95731 0.96200 0.96200 0.055888 Y3' 0.94918 0.93099 0.94918 0.075247 Y5' 0.94929 0.93791 0.94929 0.075081

[0071] For the above S106, select the value with the greatest similarity to all value benchmark templates as the eigenvalue V of the second time series data segment. For example, the V column in Table 2. If there are K groups of time series data segments related to the business analysis target, then the K groups of time series segments will have K eigenvalues, which are respectively V 1 ,V 2 …V k 。

[0072] For the above S107, as Figure 5 shown, it is a method flow chart for obtaining the information value of each second time series data segment under the business analysis target by normalizing the eigenvalue of each second time series data segment provided by an embodiment of the present invention, including:

[0073] S501: Normalize the eigenvalue of each second time series data segment to obtain the normalized eigenvalue of each second time series data segment.

[0074] Specifically, normalize the eigenvalue of each second time series data segment and convert it to the 0 - 1 interval to obtain the normalized eigenvalue of each second time series data segment.

[0075] S502: Take the negative logarithm value of the normalized eigenvalue of each second time series data segment to obtain the information value of each second time series data segment under the business analysis target.

[0076] For example, the negative logarithm value can be taken using the following formula:

[0077]

[0078] Among them, I k represents the value information of the k-th second time-series data segment, V k represents the eigenvalue of the k-th second time-series data segment, V min represents the minimum eigenvalue among the k second time-series data segments, V max represents the maximum eigenvalue among the k second time-series data segments, where K is an integer.

[0079] In addition, the normalization process can adopt common algorithms in the art, and the base of the logarithm is optional, usually 2.

[0080] Exemplarily, taking the calculation of similarity using the correlation coefficient as an example, the normalization process can be: if V is greater than 0, the value remains unchanged; if V is less than 0, the value is recorded as 0. In the example of Table 2, all eigenvalues are greater than 0, so the negative logarithm operation can be directly taken, and the information value is shown in the I column of Table 2.

[0081] Regarding the above S108, as Figure 6 shown, it is a flowchart of a method for obtaining the information value density of the time-series data to be analyzed under the business analysis target provided by the embodiment of the present invention, including:

[0082] S601: Sum the information value of the first time-series data segment and the information value of the second time-series data segment to obtain the total information value of the time-series data to be analyzed.

[0083] Specifically, the information value of the entire time-series data to be analyzed for this business analysis target is the sum of the information values of all time-series data segments: Among them, I(p) represents the total information value of the time-series data to be analyzed for the p business analysis target, I m represents the information value of the m-th time-series data segment (including: the first time-series data segment, the second time-series data segment), and k is equal to the sum of the first time-series data segment and the second time-series data segment.

[0084] S602: Divide the total information value of the time-series data to be analyzed by the data length of the time-series data to be analyzed to obtain the information value density of the time-series data to be analyzed under the business analysis target.

[0085] Specifically, dividing the total information value of the time-series data to be analyzed by the data length of the time-series data to be analyzed to obtain the information value density of the time-series data to be analyzed under the business analysis target includes: Among them, I(p) represents the total information value of the time-series data to be analyzed, n is the length of the time-series data to be analyzed, and f(p) represents the information value density of the time-series data to be analyzed under the business analysis target.

[0086] Exemplarily, taking Table 2 as an example, the information value of the entire time-series data to be analyzed for this business analysis objective is the sum of the information values of all time-series data segments: I(P) = I1 + I2 + I3 + I4 + I5 = 0.055888 + 0 + 0.075247 + 0 + 0.075081 = 0.206216, then the information value density f(p) = 0.206216 / 1606 = 0.0001284.

[0087] In addition, to quickly locate the position of abnormal data in the current business, key data can be quickly located and extracted according to the information values of each time-series data and time-series data segments.

[0088] In an embodiment of the present invention, a time-series data information value evaluation device is further provided, as described in the following embodiment. Since the principle of this device for solving problems is similar to the time-series data information value evaluation method, the implementation of this device can refer to the implementation of the time-series data information value evaluation method, and the repeated parts will not be elaborated.

[0089] As Figure 7 shown, it is a schematic diagram of a time-series data information value evaluation device provided by an embodiment of the present invention, including:

[0090] A first processing module 701, configured to determine a plurality of value reference templates according to a preset business analysis objective; wherein, the value reference template is a time-series data segment generated in a normal production mode, and the business analysis objective includes: the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal production curve;

[0091] A second processing module 702, configured to determine a time-series data segmentation point flag according to a preset business analysis objective;

[0092] A third processing module 703, configured to divide the time-series data to be analyzed into a plurality of time-series data segments according to the segmentation point flag;

[0093] An assignment module 704, configured to assign an information value of 0 to a first time-series data segment that has nothing to do with the business analysis objective;

[0094] A similarity calculation module 705, configured to determine the similarity between each second time-series data segment and each value reference template for a second time-series data segment related to the business analysis objective, and obtain a plurality of similarity values corresponding to each second time-series data segment;

[0095] A fourth processing module 706, configured to determine the maximum value among the plurality of similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment;

[0096] The fifth processing module 707 is configured to normalize the eigenvalue of each second timing data segment according to each second timing, so as to obtain the information value of each second timing data segment under the business analysis objective;

[0097] The sixth processing module 708 is configured to obtain the information value density of the to-be-analyzed timing data under the business analysis objective according to the data length of the to-be-analyzed timing data, the information value of the first timing data segment, and the information value of the second timing data segment for each second timing.

[0098] In a possible implementation manner, the similarity calculation module is specifically configured to sample each second timing data segment to obtain the first sampled data of each second timing data; sample each value reference template to obtain the second sampled data of each value reference template; wherein, the data length of the first sampled data is equal to the data length of the second sampled data; calculate the similarity between the first sampled data of each second timing data and the second sampled data of each value reference template.

[0099] In a possible implementation manner, the fifth processing module is specifically configured to normalize the eigenvalue of each second timing data segment to obtain the normalized eigenvalue of each second timing data segment; take the negative logarithm value of the normalized eigenvalue of each second timing data segment to obtain the information value of each second timing data segment under the business analysis objective.

[0100] In a possible implementation manner, the sixth processing module is specifically configured to sum the information value of the first timing data segment and the information value of the second timing data segment to obtain the total information value of the to-be-analyzed timing data; divide the total information value of the to-be-analyzed timing data by the data length of the to-be-analyzed timing data to obtain the information value density of the to-be-analyzed timing data under the business analysis objective.

[0101] Based on the foregoing inventive concept, as Figure 8 shown, the present invention further provides a computer device 800, including a memory 810, a processor 820, and a computer program 830 stored in the memory 810 and executable on the processor 820. When the processor 820 executes the computer program 830, the foregoing timing data information value evaluation method is implemented.

[0102] The embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing timing data information value evaluation method is implemented.

[0103] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned method for evaluating the value of time-series data information is implemented.

[0104] In an embodiment of the present invention, according to a preset business analysis objective, a plurality of value benchmark templates are determined; wherein, the value benchmark template is a time-series data segment generated in a normal production mode, and the business analysis objective includes: the deviation degree of any one or more data indicators in the actual production curve compared to the corresponding data indicators in the normal production curve; according to the preset business analysis objective, a time-series data segmentation point flag is determined; the time-series data to be analyzed is divided into a plurality of time-series data segments according to the segmentation point flag; for the first time-series data segment that has nothing to do with the business analysis objective, an information value of 0 is assigned; for the second time-series data segment that is related to the business analysis objective, the similarity between each second time-series data segment and each value benchmark template is determined, and a plurality of similarity values corresponding to each second time-series data segment are obtained; the maximum value among the plurality of similarity values corresponding to each second time-series data segment is determined as the characteristic value of each second time-series data segment; according to the normalization processing of the characteristic value of each second time-series data segment, the information value of each second time-series data segment under the business analysis objective is obtained; according to the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment, the information value density of the time-series data to be analyzed under the business analysis objective is obtained. In this way, according to the information value density of the time-series data, whether it is data visualization analysis or anomaly point search, it can assist personnel in quickly locating high-information-value time-series data and improving the analysis efficiency.

[0105] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0106] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 in one or more flows and / or blocks Figure 1 or in one or more blocks.

[0109] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for evaluating the value of time-series data information, characterized in that, it includes: Determine multiple value benchmark templates according to the preset business analysis objectives; wherein, the value benchmark template is a time-series data segment generated in the normal production mode, and the business analysis objectives include: the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal generation curve; Determine the time-series data segmentation point flag according to the preset business analysis objective; Divide the time-series data to be analyzed into multiple time-series data segments according to the segmentation point flag; Assign an information value of 0 to the first time-series data segment that has nothing to do with the business analysis objective; For the second time-series data segment related to the business analysis objective, determine the similarity between each second time-series data segment and each value benchmark template, and obtain multiple similarity values corresponding to each second time-series data segment; Determine the maximum value among the multiple similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment; Obtain the information value of each second time-series data segment under the business analysis objective according to the normalization processing of the characteristic value of each second time-series data segment; Obtain the information value density of the time-series data to be analyzed under the business analysis objective according to the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment; Obtain the information value of each second time-series data segment under the business analysis objective according to the normalization processing of the characteristic value of each second time-series data segment, including: performing normalization processing on the characteristic value of each second time-series data segment to obtain the normalized characteristic value of each second time-series data segment; taking the negative logarithm value of the normalized characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective.

2. The method for evaluating the value of time-series data information according to claim 1, characterized in that, For the second time-series data segment related to the business analysis objective, determining the similarity between each second time-series data segment and each value benchmark template includes: Sampling each second time-series data segment to obtain the first sampling data of each second time-series data; Sampling each value benchmark template to obtain the second sampling data of each value benchmark template; wherein, the data length of the first sampling data is equal to the data length of the second sampling data; Calculate the similarity between the first sampling data of each second time-series data and the second sampling data of each value benchmark template.

3. The method for evaluating the value of time-series data information according to claim 1, characterized in that, Obtaining the information value density of the time-series data to be analyzed under the business analysis objective according to the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment includes: Adding the information value of the first time-series data segment and the information value of the second time-series data segment to obtain the total information value of the time-series data to be analyzed; Divide the total information value of the time-series data to be analyzed by the data length of the time-series data to be analyzed to obtain the information value density of the time-series data to be analyzed under the business analysis objective.

4. A time-series data information value evaluation device, characterized in that, it includes: A first processing module, configured to determine a plurality of value benchmark templates according to a preset business analysis objective; wherein, the value benchmark template is a time-series data segment generated in a normal production mode, and the business analysis objective includes: the deviation degree of any one or more data indicators in the actual production curve compared with the corresponding data indicators in the normal production curve; A second processing module, configured to determine a time-series data segmentation point flag according to a preset business analysis objective; A third processing module, configured to divide the time-series data to be analyzed into a plurality of time-series data segments according to the segmentation point flag; An assignment module, configured to assign an information value of 0 to a first time-series data segment irrelevant to the business analysis objective; A similarity calculation module, configured to determine the similarity between each second time-series data segment and each value benchmark template for a second time-series data segment related to the business analysis objective, and obtain a plurality of similarity values corresponding to each second time-series data segment; A fourth processing module, configured to determine the maximum value among the plurality of similarity values corresponding to each second time-series data segment as the characteristic value of each second time-series data segment; A fifth processing module, configured to normalize the characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective; A sixth processing module, configured to obtain the information value density of the time-series data to be analyzed under the business analysis objective according to the data length of the time-series data to be analyzed, the information value of the first time-series data segment, and the information value of the second time-series data segment; The fifth processing module is specifically configured to: normalize the characteristic value of each second time-series data segment to obtain the normalized characteristic value of each second time-series data segment; take the negative logarithm value of the normalized characteristic value of each second time-series data segment to obtain the information value of each second time-series data segment under the business analysis objective.

5. The time-series data information value evaluation device according to claim 4, characterized in that, The similarity calculation module is specifically configured to sample each second time-series data segment to obtain first sampling data of each second time-series data; sample each value benchmark template to obtain second sampling data of each value benchmark template; wherein, the data length of the first sampling data is equal to the data length of the second sampling data; calculate the similarity between the first sampling data of each second time-series data and the second sampling data of each value benchmark template.

6. The time-series data information value evaluation device according to claim 5, characterized in that, The sixth processing module is specifically configured to sum the information value of the first time-series data segment and the information value of the second time-series data segment to obtain the total information value of the time-series data to be analyzed; Dividing the total information value of the time-series data to be analyzed by the data length of the time-series data to be analyzed, the information value density of the time-series data to be analyzed under the business analysis objective is obtained.

7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 3 is implemented.

8. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

9. A computer program product, wherein, the computer program product includes a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Time series data trend feature extraction method based on dynamic grid division

    CN112765562A

  • Time sequence signal abnormal fragment detection method and system

    CN114297264A