A data intelligent analysis method and system
By decomposing the data curve into trend terms, season terms and residual terms, and dividing the data segments with the characteristics of these components, calculating the degree of abnormality of each data segment and correcting it, the problem of degradation of data abnormality detection in the prior art is solved, and higher data evaluation accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510072507.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art fails to effectively consider the characteristics of data changes and the lag of threshold judgment in data abnormality detection, resulting in a decrease in detection accuracy.
By decomposing the data curve into trend terms, season terms and residual terms, dividing the data segments with the characteristics of these components, the abnormality degree of each data segment is calculated separately, and the abnormality degree of the residual terms is corrected to meet the specific comprehensive abnormality degree relationship.
It improves the accuracy and reliability of data evaluation, can more comprehensively identify abnormal situations in the data, reduce false alarm rates, and enhance the accuracy of abnormal detection.
Smart Images

Figure CN119513549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and more specifically, to a data intelligent analysis method and system. Background Art
[0002] In the industrial production process, various sensors and detection equipment will collect a large amount of data in real time. These data reflect various aspects of the production process, such as temperature, pressure, flow, vibration, etc. Under normal circumstances, these data should fluctuate within a certain range, reflecting the stable state of the production process. When abnormalities occur in the production process, such as equipment failure, process parameter deviation, raw material quality changes, etc., these abnormalities will directly affect the changes in production data. Through intelligent analysis of the collected data, abnormalities in the production process can be detected.
[0003] Prior art, such as a patent application document with publication number CN117708636A, discloses a data analysis method based on big data, which includes: inputting original data, obtaining the distribution characteristics of the original data through a time series diagram and a box plot, eliminating out-of-limit values in the distribution characteristics of the original data according to the normal fluctuation range of the measurement point data, filling in the missing data values in the original data, using DFA to detrend the original data to eliminate the data trend item, using the K-means clustering algorithm to perform cluster analysis on the original data, and determining an abnormal judgment threshold, using the abnormal judgment threshold to compare with the data set density to determine whether the data set density is less than the abnormal judgment threshold, if so, the corresponding original data is abnormal data, otherwise it is normal data.
[0004] However, the above patent application documents do not take into account the characteristics of data changes (such as short-term fluctuations or cyclical changes) and the lag in threshold judgment, resulting in a decrease in the accuracy of data anomaly detection. Summary of the invention
[0005] In order to solve the above-mentioned technical problem of decreased accuracy in detecting anomalies in data, the present invention provides solutions in the following aspects.
[0006] In a first aspect, a data intelligent analysis method comprises:
[0007] Sort the collected data in chronological order and construct a curve showing the data changing over time;
[0008] Decomposing the curve into three components, namely, a trend term, a seasonal term, and a residual term, and dividing the curve into a plurality of segmented data segments based on the characteristics of the trend term and the seasonal term, and respectively calculating the abnormality degree of the trend term, the abnormality degree of the seasonal term, and the abnormality degree of the residual term corresponding to each segmented data segment;
[0009] The comprehensive abnormality degree of each divided data segment is calculated, and the abnormality degree of the residual term is used for correction to obtain the corrected comprehensive abnormality degree of each divided data segment. The corrected comprehensive abnormality degree satisfies the relationship: ; In the formula, For the The comprehensive abnormality degree after correction of the divided data segments, is the comprehensive abnormality degree, For the The abnormal degree of the seasonal item corresponding to the divided data segment, For the The abnormal degree of the trend item corresponding to the divided data segment, For the The abnormality of the residual term of the data segment, is the maximum value;
[0010] When the abnormality of the divided data segment is greater than or equal to the preset threshold, the alarm mechanism is triggered.
[0011] The present invention constructs a curve to intuitively display the changing trend of the collected data, further decomposes the curve into a trend term, a seasonal term and a residual term, and divides the curve into a plurality of divided data segments, and respectively calculates the abnormal degree of the trend term, the abnormal degree of the seasonal term and the abnormal degree of the residual term corresponding to the divided data segments;
[0012] Among them, although the seasonal term and trend term can reflect the cyclical fluctuations of the long-term changes in the data, they cannot completely capture all the anomalies in the data. The residual term, as the remaining part of the data, contains all the information except the trend and seasonal factors. According to the above correction method, it not only takes into account the dominant role of the main anomaly on the comprehensive anomaly degree, but also gives the residual term an important role in anomaly detection, thereby improving the accuracy and reliability of data evaluation.
[0013] Preferably, the decomposition adopts STL decomposition method.
[0014] Preferably, dividing the curve into a plurality of data segments by combining the features of the trend term and the season term comprises:
[0015] The turning points of the trend item and the seasonal item are identified respectively, the curve is divided into a plurality of trend segments according to the turning points of the trend item, each trend segment is divided into a plurality of sub-segments according to the turning points of the seasonal item, and thus a plurality of divided data segments are obtained.
[0016] Different data segments may represent different trend stages or seasonal patterns. By dividing the data into segments, you can more clearly see the differences in data at different stages and patterns, so as to conduct more targeted analysis and processing.
[0017] Preferably, the autocorrelation function is used to obtain all the cycles and cycle lengths of the curve, and all the cycles are clustered to obtain multiple cycle categories. The center point of the category with the largest number of elements in the category is used as the reference point. The abnormality degree of each cycle of the seasonal item corresponding to the data segment is:
[0018] ; In the formula, For the of the categories The degree of abnormality of the period represented by the element, For the The number of elements in the category, is the number of all cycles, For the of the categories The distance from the element to the reference point, is the maximum distance from all elements in all categories to the reference point, is the constructed window size, For the window Parameters of the cycle;
[0019] The abnormal degree of the cycle is mapped to the divided data segments, and then the abnormal degree of the seasonal item corresponding to each divided data segment is obtained.
[0020] In periodic data affected by noise, the use of autocorrelation function can show obvious periodicity. By calculating the degree of abnormality of each cycle, the degree of deviation of each cycle from the normal cycle can be quantitatively evaluated. Cycles with higher degrees of abnormality may represent abnormal events or noise interference in the data.
[0021] Preferably, if the window A cycle is a normal cycle, then The value is 1, otherwise the value is 0.
[0022] By setting If the value is 0 or 1, we can make a detailed evaluation of each cycle. This evaluation method not only considers the distance from the cycle to the reference point, but also considers the state of the cycle itself (normal or abnormal), making the evaluation results of the data more comprehensive and accurate.
[0023] Preferably, the trend item component is divided into multiple initial data segments, the average slope of each initial data segment is calculated, the average slope of each initial data segment is clustered, the value of the center point of the category containing the largest number of elements in the category is used as the reference slope value of the data, and the abnormality degree of the initial data segment is calculated. The abnormality degree of the initial data segment satisfies the relationship:
[0024] ; In the formula, For the trend item The abnormality of the initial data segment, For the The initial data segment The anomaly score of the data, For the The total number of data on the initial data segment, For the The average slope of the data in the initial data segment, is the slope reference value, It is the maximum value of the difference between the corresponding slope average value and the slope reference value in all initial data segments; and then the abnormal degree of the initial data segment is mapped to the abnormal degree of the trend item of the corresponding divided data segment.
[0025] By dividing the initial data into segments and calculating their average slopes, we can capture the local characteristics of the data in more detail rather than just the overall trend, which helps to more accurately identify rising, falling, or stable trends in the data, especially when there is fluctuation or noise in the data.
[0026] Preferably, the abnormality degree of the residual item corresponding to the divided data segment satisfies the relationship:
[0027] ; In the formula, For the The abnormality of the residual term of the data segment, For the The mean of all data in the residual term of the divided data segment, For the The median of the means of all data in the residual terms of the data segments, is the maximum value of the data in the entire residual term, is the minimum value of the data in the entire residual term, For the The variance of all data in the residual term of the divided data segment, is the variance of the entire residual data.
[0028] By calculating the mean, median, and variance of the residual term and combining the maximum and minimum values of the entire residual term, data segments that do not conform to the overall trend can be more accurately identified. This method helps reduce false positives and improve the accuracy of anomaly detection.
[0029] In a second aspect, a data intelligent analysis system includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned data intelligent analysis method is implemented.
[0030] The beneficial effects of the present invention are:
[0031] By decomposing the data curve into three components: trend term, seasonal term and residual term, we can comprehensively identify anomalies in the data from different dimensions. The trend term focuses on the overall change trend of the data, the seasonal term analyzes the cyclical fluctuations of the data, and the residual term captures the random fluctuations in the data. By comprehensively considering the three, we can more comprehensively discover anomalies in the data.
[0032] According to the characteristics of trend items and seasonal items, data segments are flexibly divided. First, the turning points of trend items are identified to divide trend segments, and then each trend segment is subdivided into sub-segments according to the turning points of seasonal items. This division method can adapt to the characteristics and needs of different data, making the analysis more refined. According to the above correction method, the dominant role of the main anomaly on the comprehensive anomaly degree is considered, and the residual term is given an important role in anomaly detection, thereby improving the accuracy and reliability of data evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0034] Figure 1 It is a method flow chart of steps S1 to S4 in a data intelligent analysis method in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0036] The embodiment of the present invention discloses a data intelligent analysis method, referring to Figure 1 , including steps S1 to S4, which are specifically as follows:
[0037] S1: Sort the collected data in chronological order and construct a curve showing the data changing over time.
[0038] First, in the industrial production process, sensors and detection equipment are used to collect environmental parameters such as temperature, humidity, and pressure in real time.
[0039] Sort the collected data by timestamp from small to large, and then select appropriate tools (such as Excel, Matplotlib (Python library), Tableau, etc.) to construct a time change curve, where the timestamp is used as the horizontal axis (x-axis) and the collected data value is used as the vertical axis (y-axis).
[0040] By sorting the collected data in chronological order and constructing a curve showing the data changing over time, the data change process can be displayed intuitively, providing strong support for data analysis and decision-making.
[0041] Exemplarily, data for one month may be collected.
[0042] S2: Decompose the curve into three components: trend term, seasonal term and residual term, and divide the curve into multiple data segments based on the characteristics of the trend term and the seasonal term, and calculate the abnormality degree of the trend term, the abnormality degree of the seasonal term and the abnormality degree of the residual term corresponding to each data segment.
[0043] For the curve constructed by S1 above, the STL decomposition method is used to decompose it into three components: trend term, seasonal term and residual term. The STL decomposition method is a prior art and will not be described in detail here.
[0044] In addition to the STL decomposition method, the above-constructed curve can also be decomposed using classical decomposition methods (additive decomposition, multiplicative decomposition), empirical mode decomposition (EMD), etc.
[0045] Through the above decomposition, complex time series data can be broken down into simpler components. The trend term reflects the long-term trend of the data over time, the seasonal term reveals the repetitive pattern of the data in a fixed period, and the residual term contains random fluctuations and noise in the data.
[0046] The residual term usually contains random noise in the data. In some cases, the data may need to be cleaned to remove the influence of noise. Through decomposition, outliers or noise in the residual term can be identified, so that the original data can be corrected or smoothed to improve the quality and accuracy of the data. In time series data, there may be missing values. The decomposed trend term and seasonal term can provide a reference for filling missing values. For example, based on the long-term change law of the trend term and the cyclical pattern of the seasonal term, the missing values can be reasonably estimated to make the data more complete.
[0047] In STL decomposition, the trend term and the seasonal term reflect the long-term trend and periodic changes of the time series data respectively. In order to divide the original data more finely, the characteristics of these two components are further combined to determine the boundaries of the data segments, and the original data is divided into multiple data segments. The details are as follows:
[0048] First, identify all the maximum points, minimum points, and inflection points in the trend item data. These points usually represent important changes in the data trend, and then identify the peaks and valleys of periodic changes in the seasonal item data. These points can help us determine the seasonal cycle of the data. Based on these significant change points or turning points, the data can be divided into several large trend segments according to the trend item, and then in each trend segment, it can be further subdivided into smaller sub-segments according to the periodic changes of the seasonal item, thereby dividing the above curve into multiple segmented data segments.
[0049] For example, the data for one month is divided. Since the data time span is short, the daily change trend can be directly observed. According to the fluctuation of the data, the data for this month is divided into several small data segments (for example, if the data rises at the beginning of the month, is stable in the middle of the month, and falls at the end of the month, it can be divided into rising segments, stable segments, and falling segments accordingly). For the data for one month, seasonal analysis may be more reflected in the periodic changes within the week or day. If the data has a weekly repetitive pattern (such as the difference between weekdays and weekends), the data for this month can be divided by week (for example, the data can be divided into the first week, the second week, etc.).
[0050] After dividing the data into segments, the data characteristics within each segment can be analyzed more carefully to avoid the anomalies of local data being masked by overall trends or seasonal fluctuations. For example, in the machine tool cutting process, the cutting parameters may change differently in different processing stages. Segmented analysis can more accurately identify anomalies in each stage.
[0051] Among them, before the curve is divided more finely, the abnormal degree of trend item and the abnormal degree of seasonal item are analyzed respectively.
[0052] First, the autocorrelation function of the seasonal term is calculated using the autocorrelation analysis method to find the periodic characteristics of the data and determine the length of each cycle. For each cycle, a feature vector is constructed based on the extreme point and the cycle length. The feature vector of each cycle for:
[0053]
[0054] in, is the maximum value in the current cycle, is the minimum value in the current cycle, is the time length of the current cycle, is the time corresponding to the maximum value of the current cycle, The time corresponding to the minimum value of the current cycle, The start time of the current cycle.
[0055] The above feature vector is mapped to a point in multidimensional space, and the feature points corresponding to all cycles are clustered using the mean shift clustering algorithm. The clustering results divide the cycles into different categories, each of which represents a group of similar cycle features. The number of categories and the number of elements in each category are counted, and the center point of the category with the largest number of elements in the category is used as the reference point. All cycles corresponding to the elements in this category are considered normal cycles, and the degree of abnormality is not calculated. For the cycles corresponding to the elements in other categories, the degree of abnormality is calculated.
[0056] Specifically, the distance from the element in any category to the reference point is calculated; a window is constructed, and the window size is set to a certain ratio or multiple related to the cycle length, such as 1 to 3 times the cycle length, so that the window can cover several complete cycles centered around the current cycle. Then the abnormality of the above cycles other than the normal cycle is calculated, that is, the relationship is satisfied:
[0057]
[0058] In the formula, For the of the categories The degree of abnormality of the period represented by the element, For the The number of elements in the category, is the number of all cycles, For the of the categories The distance from the element to the reference point, is the maximum distance from all elements in all categories to the reference point, is the window size, For the window Parameters of the cycle.
[0059] Among them, if the window A cycle is a normal cycle, then The value is 1, otherwise the value is 0.
[0060] Reflects the proportion of cycle categories. When it is small, it means that the periodic anomaly corresponding to this category is more likely to be affected by noise and less likely to be a production anomaly; otherwise, it is more likely to be a production anomaly.
[0061] Reflects the distance from the reference point. The larger it is, the farther the element is from the reference point and the greater the abnormality of the element.
[0062] It reflects the normality of adjacent cycles. If the frequency of multiple adjacent cycles of the current cycle being normal cycles is high, it means that the possibility that the current cycle is a production abnormality or equipment abnormality is low; if the frequency is low, it means that the possibility that the current cycle is a production abnormality or equipment abnormality is high.
[0063] According to the above operation, the abnormality degree of all cycles can be obtained in the same way. After further detailed division, the abnormality degree of the cycle can be directly used as the abnormality degree of the seasonal item corresponding to the divided data segment. This is because although a cycle segment is subdivided into multiple data segments, these data segments still retain the seasonal characteristics and abnormality degree of the seasonal cycle to which they belong. In other words, no matter how the data segment is subdivided, the abnormality degree of its seasonal item is relative to the seasonal cycle to which it belongs.
[0064] Secondly, all the maximum points, minimum points and inflection points in the trend item component data are counted, and these points are regarded as the endpoints of the initial data segments, thereby dividing the trend item data into multiple initial data segments, and the isolation forest algorithm is used to calculate the residual common score of all trend item data, and the total number of data on each initial data segment in the trend item is obtained. The average slope of each initial data segment is calculated, and the mean shift clustering algorithm is used to cluster the average slope of each initial data segment, and the value of the intra-class center point of the category with the largest number of elements in the category is used as the reference slope value of the data.
[0065] Then the abnormal degree of the trend item satisfies the relationship:
[0066]
[0067] In the formula, For the trend item The abnormality of the initial data segment, For the The initial data segment The anomaly score of the data, For the The total number of data on the initial data segment, For the The average slope of the data in the initial data segment, is the slope reference value, It is the maximum value of the difference between the corresponding average slope value and the slope reference value in all initial data segments.
[0068] in, The larger it is, the greater the anomaly score of the current initial data segment in the entire trend item. Indicates The absolute value of the difference between the average slope of the initial data segment and the reference slope is normalized to the range of the maximum slope difference. The larger this value is, the greater the abnormality of the change trend of the current initial data segment in the entire trend item data, that is, the current initial data segment has a large abnormal change trend.
[0069] Similarly, the abnormality degree of all initial data segments can be obtained. For each subdivided data segment, determine which initial data segment it belongs to. Since the subdivided data segment is a further division of the initial data segment, each subdivided data segment can be traced back to the initial data segment to which it belongs. Once the initial data segment to which the subdivided data segment belongs is determined, the abnormality degree of the trend item of the subdivided data segment can be directly inherited from the abnormality degree of the initial data segment to which it belongs. This is because the subdivided data segment is still part of its seasonal cycle and maintains the corresponding seasonal characteristics and abnormality degree.
[0070] Finally, since the residuals of the entire curve data may be affected by many factors (seasonal changes, trend changes, sudden events, etc.), if the residual abnormality of the entire curve data is directly calculated, these subtle differences may be ignored. Therefore, after a more detailed division, the abnormality of the residual items of the divided data segments is calculated.
[0071] Specifically, obtain the The mean, median and variance of all data in the residual term of the data segment, then The abnormal degree of the residual term of the divided data segment satisfies the relationship:
[0072]
[0073] In the formula, For the The abnormality of the residual term of the data segment, For the The mean of all data in the residual term of the divided data segment, For the The median of the means of all data in the residual terms of the data segments, is the maximum value of the data in the entire residual term, is the minimum value of the data in the entire residual term, For the The variance of all data in the residual term of the divided data segment, is the variance of the entire residual data.
[0074] in, Evaluated the The degree of deviation from the central tendency of the divided data segments, This is to eliminate the dimensional impact between different partitioned data segments. In order to evaluate the The difference between the fluctuation degree of a divided data segment and the overall fluctuation degree. If the variance of the divided data segment is large, it means that the fluctuation degree of the divided data segment is high and may contain more abnormal information.
[0075] S3: Calculate the comprehensive abnormality degree of each divided data segment, and use the abnormality degree of the residual item to correct it, so as to obtain the corrected comprehensive abnormality degree of each divided data segment.
[0076] For each data segment, the comprehensive abnormality degree of the segment is obtained by combining the abnormality degree of the trend item and the abnormality degree of the seasonal item, and the abnormality degree of the residual item is used to make corrections to obtain the corrected comprehensive abnormality degree of each data segment.
[0077] Exemplarily, the comprehensive abnormality degree after the data segment is divided and corrected satisfies the relationship:
[0078]
[0079] In the formula, For the The comprehensive abnormality degree after correction of the divided data segments, is the comprehensive abnormality degree, For the The abnormal degree of the seasonal item corresponding to the divided data segment, For the The abnormal degree of the trend item corresponding to the divided data segment, For the The abnormality of the residual term of the data segment, is the maximum value.
[0080] Among them, The abnormal degree of the seasonal item corresponding to the divided data segment The abnormal degree of all cycles is obtained by calculating the above (the specific reasons have been mentioned above and will not be repeated here); The abnormal degree of the trend item corresponding to the divided data segment It is obtained by calculating the abnormality degree of the initial data segment as mentioned above (the specific reasons have been mentioned above and will not be repeated here).
[0081] Will As the exponent to correct the maximum value, The bigger the The smaller the The value of is reduced, that is, The greater the increase, the greater the correction; conversely, the smaller the correction.
[0082] In general, the degree of abnormality of each data segment in the original data is evaluated by combining the degree of abnormality of the seasonal term and the trend term and correcting it according to the degree of residual abnormality. The higher the degree of residual abnormality, the stronger the correction of the degree of abnormality; the lower the degree of residual abnormality, the weaker the correction of the degree of abnormality.
[0083] S4: When the abnormality of the divided data segment is greater than or equal to the preset threshold, the alarm mechanism is triggered.
[0084] According to the above S3, calculate the The corrected comprehensive abnormality degree of each divided data segment can be calculated in the same way as the corrected comprehensive abnormality degree of each divided data segment. An abnormality threshold is set (it is set to 0.5 based on the experience value and can be adjusted by the implementers based on the actual situation). When the abnormality degree of a divided data segment is greater than or equal to the abnormality threshold, it is determined that there is an abnormality in the divided data segment, that is, the production process or equipment at this time is abnormal, which triggers the alarm mechanism to remind the staff to perform maintenance.
[0085] An embodiment of the present invention further discloses a data intelligent analysis system, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the data intelligent analysis method according to the present invention is implemented.
[0086] The system also includes other components familiar to those skilled in the art, such as a communication bus and a communication interface. The configuration and functions of these components are known in the art and will not be described in detail here.
[0087] In the present invention, the aforementioned memory may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus or device. For example, a computer-readable storage medium may be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory RRAM (Resistive Random Access Memory), a dynamic random access memory DRAM (Dynamic Random Access Memory), a static random access memory SRAM (Static Random-Access Memory), an enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), a high-bandwidth memory HBM (High-Bandwidth Memory), a hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module or both. Any such computer storage medium may be part of a device or accessible or connectable to a device.
[0088] In the description of this specification, "plurality" or "several" means at least two, such as two, three or more, etc., unless otherwise clearly and specifically defined.
[0089] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.
Claims
1. A data intelligent analysis method, characterized in that: include: Sort the collected data in chronological order and construct a curve showing the data changing over time; Decomposing the curve into three components, namely, a trend term, a seasonal term, and a residual term, and dividing the curve into a plurality of segmented data segments based on the characteristics of the trend term and the seasonal term, and respectively calculating the abnormality degree of the trend term, the abnormality degree of the seasonal term, and the abnormality degree of the residual term corresponding to each segmented data segment; The comprehensive abnormality degree of each divided data segment is calculated, and the abnormality degree of the residual term is used for correction to obtain the corrected comprehensive abnormality degree of each divided data segment. The corrected comprehensive abnormality degree satisfies the relationship: ; In the formula, For the The comprehensive abnormality degree after correction of the divided data segments, is the comprehensive abnormality degree, For the The abnormal degree of the seasonal item corresponding to the divided data segment, For the The abnormal degree of the trend item corresponding to the divided data segment, For the The abnormality of the residual term of the data segment, is the maximum value; When the comprehensive abnormality degree of the divided data segment is greater than or equal to the preset threshold, the alarm mechanism is triggered.
2. A data intelligent analysis method according to claim 1, characterized in that: The decomposition adopts the STL decomposition method.
3. A data intelligent analysis method according to claim 2, characterized in that: The step of dividing the curve into a plurality of data segments by combining the features of the trend term and the season term comprises: The turning points of the trend item and the seasonal item are identified respectively, the curve is divided into a plurality of trend segments according to the turning points of the trend item, each trend segment is divided into a plurality of sub-segments according to the turning points of the seasonal item, and thus a plurality of divided data segments are obtained.
4. A data intelligent analysis method according to claim 3, characterized in that: Use the autocorrelation function to obtain all the cycles and cycle lengths of the curve, cluster all the cycles, and obtain multiple cycle categories. Take the center point of the category with the largest number of elements as the reference point. The abnormality degree of each cycle of the seasonal item corresponding to the data segment is: ; In the formula, For the of the categories The degree of abnormality of the period represented by the element, For the The number of elements in the category, is the number of all cycles, For the of the categories The distance from the element to the reference point, is the maximum distance from all elements in all categories to the reference point, is the constructed window size, For the window Parameters of the cycle; The abnormal degree of the cycle is mapped to the divided data segments, and then the abnormal degree of the seasonal item corresponding to each divided data segment is obtained.
5. A data intelligent analysis method according to claim 4, characterized in that: If the window A cycle is a normal cycle, then The value is 1, otherwise the value is 0.
6. A data intelligent analysis method according to claim 5, characterized in that: The trend item component is divided into multiple initial data segments, the average slope of each initial data segment is calculated, the average slope of each initial data segment is clustered, the value of the center point of the category with the largest number of elements in the category is used as the reference slope value of the data, and the abnormality of the initial data segment is calculated. The abnormality of the initial data segment satisfies the relationship: ; In the formula, For the trend item The abnormality of the initial data segment, For the The initial data segment The anomaly score of the data, For the The total number of data on the initial data segment, For the The average slope of the data in the initial data segment, is the slope reference value, It is the maximum value of the difference between the corresponding slope average value and the slope reference value in all initial data segments; and then the abnormal degree of the initial data segment is mapped to the abnormal degree of the trend item of the corresponding divided data segment.
7. A data intelligent analysis method according to claim 6, characterized in that: The abnormal degree of the residual item corresponding to the divided data segment satisfies the relationship: ; In the formula, For the The abnormality of the residual term of the data segment, For the The mean of all data in the residual term of the divided data segment, For the The median of the means of all data in the residual terms of the data segments, is the maximum value of the data in the entire residual term, is the minimum value of the data in the entire residual term, For the The variance of all data in the residual term of the divided data segment, is the variance of the entire residual data.
8. A data intelligent analysis system, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the data intelligent analysis method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Data analysis method based on big data
CN117708636A
Adaptive stability detection method for satellite seasonal fluctuation remote measurement
CN111680397A
Automatic task receiving system and method based on big data
CN113435763A