Method, device, server and medium for identifying data anomalies

By extracting data feature indicators and inputting the autonomous learning model, the problems of high artificial monitoring costs and poor accuracy in identifying large-scale data outliers in the prior art are solved, and more efficient and accurate data anomaly recognition is achieved.

CN113780329BActive Publication Date: 2025-05-23BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110366765.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-06
Publication Date
2025-05-23
Estimated Expiration
2041-04-06

AI Technical Summary

Technical Problem

In the prior art, when identifying outliers of large-scale data, artificial monitoring costs are high and the accuracy is poor, and it is difficult to accurately determine whether the data is abnormal by setting a fluctuation threshold.

Method used

By obtaining the target data sequence within the preset time period, determining its predicted value, extracting data characteristic indicators, and inputting them into the pre-trained autonomous learning model, generating prompt information for characterizing whether there are data exceptions.

Benefits of technology

It significantly reduces the workload of personnel, improves the accuracy of data anomaly identification, and meets training requirements while reducing the number of sample annotations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113780329B_ABST
    Figure CN113780329B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure disclose a method, device, server and medium for identifying data anomalies. A specific implementation of the method includes: obtaining a target data sequence within a preset time period; determining a predicted value corresponding to the target data sequence; extracting data feature indicators based on the target data sequence and the predicted value; inputting the data feature indicators into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is a data anomaly. This implementation reduces the workload of personnel and improves the accuracy of data anomaly identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, device, server, and medium for identifying data anomalies. Background Art

[0002] With the development of Internet technology, the scale of data is increasing. For large-scale data, how to identify outliers in a timely and effective manner is of great significance to the normal operation of business systems.

[0003] In the prior art, data changes are often monitored manually or fluctuation thresholds are set to determine whether the data is abnormal. However, manual monitoring requires a lot of manpower, and it is difficult to ensure that personnel can detect large-scale data anomalies in a timely manner. In actual scenarios, data usually fluctuates, and the accuracy of determining whether data is abnormal by setting fluctuation thresholds is poor. Summary of the invention

[0004] Embodiments of the present disclosure provide methods, devices, servers, and media for identifying data anomalies.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for identifying data anomalies, the method comprising: obtaining a target data sequence within a preset time period; determining a predicted value corresponding to the target data sequence; extracting data feature indicators based on the target data sequence and the predicted value; inputting the data feature indicators into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is data anomaly.

[0006] In a second aspect, an embodiment of the present disclosure provides a method for training an anomaly classification model, the method comprising: obtaining a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on data statistical features and time period comparison features within a historical time period, and the annotation values ​​are generated based on the method of the first aspect; taking the sample data features of the training samples in the training sample set as input, and taking the annotation values ​​corresponding to the input sample data features as expected output, to train a quasi-anomaly classification model; clustering the output values ​​corresponding to the training samples in the training sample set, and generating representative values ​​corresponding to categories of a target number of clusters as a reference for grade determination; generating an anomaly classification model based on the representative values ​​and the quasi-anomaly classification model, wherein the anomaly classification model is used to characterize the correspondence between data features and data anomaly levels.

[0007] In the third aspect, an embodiment of the present disclosure provides a device for identifying data anomalies, the device comprising: an acquisition unit, configured to acquire a target data sequence within a preset time period; a determination unit, configured to determine a predicted value corresponding to the target data sequence; an extraction unit, configured to extract data feature indicators based on the target data sequence and the predicted value; and a generation unit, configured to input the data feature indicators into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is data anomaly.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a device for training an anomaly classification model, the device comprising: a training sample acquisition unit, configured to acquire a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on data statistical features and time period comparison features within a historical time period, and the annotation values ​​are generated based on the method of claim 7; a training unit, configured to take the sample data features of the training samples in the training sample set as input, and take the annotation values ​​corresponding to the input sample data features as expected output, to train a quasi-anomaly classification model; a clustering unit, configured to cluster the output values ​​corresponding to the training samples in the training sample set, and generate representative values ​​corresponding to the categories of a target number of clusters as a reference for grade determination; a model generation unit, configured to generate an anomaly classification model based on the representative values ​​and the quasi-anomaly classification model, wherein the anomaly classification model is used to characterize the correspondence between data features and data anomaly levels.

[0009] In a fifth aspect, an embodiment of the present disclosure provides a server, comprising: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation method in the first aspect.

[0010] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0011] The method, device, server and medium for identifying data anomalies provided by the embodiments of the present disclosure extract data feature indicators based on the target data sequence and the predicted value of the determined target sequence, and then use the extracted data feature indicators as the input of the autonomous learning model trained by the machine learning method, so as to obtain prompt information for characterizing whether there is data anomaly. This significantly reduces the workload of personnel on the one hand, and improves the accuracy of data anomaly identification on the other hand by selecting data feature indicators and combining them with machine learning models. Moreover, since the amount of data in the prior art is usually large, the annotation cost required for ordinary machine learning model training is too high. The use of an autonomous learning model can meet the training requirements while reducing the amount of sample annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0013] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0014] Figure 2 is a flow chart of an embodiment of a method for identifying data anomalies according to the present disclosure;

[0015] Figure 3 is a schematic diagram of an application scenario of a method for identifying data anomalies according to an embodiment of the present disclosure;

[0016] Figure 4 is a flow chart of one embodiment of a method for training an anomaly classification model according to the present disclosure;

[0017] Figure 5 is a structural schematic diagram of an embodiment of a device for identifying data anomalies according to the present disclosure;

[0018] Figure 6 is a structural schematic diagram of an embodiment of an apparatus for training an anomaly classification model according to the present disclosure;

[0019] Figure 7 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0021] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0022] Figure 1 An exemplary architecture 100 is shown to which the method for identifying data anomalies or the apparatus for identifying data anomalies of the present disclosure can be applied.

[0023] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0024] The terminal devices 101, 102, 103 interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, information applications, etc.

[0025] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens and supporting human-computer interaction, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0026] The server 105 may be a server that provides various services, such as a background server that provides support for communication client applications on the terminal devices 101, 102, and 103. The background server may obtain the summary data generated by each client application, and perform corresponding processing (such as generating prompt information for indicating whether there is data anomaly) based on the above-mentioned summary data, and may also perform corresponding processing strategies based on the generated prompt information, such as pushing corresponding information to the selected terminal device.

[0027] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0028] It should be noted that the method for identifying data anomalies provided in the embodiments of the present disclosure is generally executed by the server 105 , and accordingly, the device for identifying data anomalies is generally disposed in the server 105 .

[0029] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0030] Continue to refer Figure 2 , shows a process 200 of an embodiment of a method for identifying data anomalies according to the present disclosure. The method for identifying data anomalies comprises the following steps:

[0031] Step 201, obtaining a target data sequence within a preset time period.

[0032] In this embodiment, the execution subject of the method for identifying data anomalies (such as Figure 1 The server 105 shown in the figure can obtain the target data sequence within the preset time period through a wired connection or a wireless connection. Specifically, the above-mentioned execution subject can obtain the target data sequence within the preset time period pre-stored locally, or it can be based on an electronic device (such as an electronic device) connected to it for communication. Figure 1 The interactive information sent by the terminal devices 101, 102, 103) shown in the figure generates a target data sequence within a preset time period.

[0033] In this embodiment, the elements in the above-mentioned target data sequence can be the target data corresponding to each sub-time period in the above-mentioned preset time period. As an example, the above-mentioned preset time period can be the time period of 12:00-17:59 on a certain date. Thus, the elements in the above-mentioned target data sequence can be, for example, target data corresponding to 12:00-12:59, 13:00-13:59, 14:00-14:59, 15:00-15:59, 16:00-16:59, and 17:00-17:59. The above-mentioned target data can be various types of data determined according to the actual application scenario. As an example, the above-mentioned target data can be the number of orders counted according to different commodity types or different sales channels (direct sales or agents) or different stores monitored by the backend server of the e-commerce system. As another example, the above-mentioned target data can also be the transaction data such as the transaction volume paid by different banks or the number of users logged in by different user terminals (such as PC terminals and APP terminals) monitored by the backend server of the electronic payment system. As another example, the target data may also include various application scenario data similar to periodic signals, such as the number of game login users, sound signals, heart wave signals, etc.

[0034] It should be noted that the above target data is usually periodic and the total amount in each period is basically constant.

[0035] In some optional implementations of this embodiment, the above execution subject may also obtain the target data sequence within a preset time period according to the following steps:

[0036] The first step is to obtain the original data sequence corresponding to each sub-time period within the preset time period.

[0037] In these implementations, the execution subject may obtain the original data sequence corresponding to each sub-time period within the preset time period through a wired connection or a wireless connection. Specifically, the execution subject may obtain the original data sequence corresponding to each sub-time period within the preset time period pre-stored locally, or may obtain the original data sequence corresponding to each sub-time period within the preset time period based on an electronic device (e.g. Figure 1 The interactive information sent by the terminal devices 101, 102, and 103 shown in the figure generates a raw data sequence corresponding to each sub-time period within the preset time period. The preset time period is, for example, 6 hours, and the sub-time period is, for example, the preset time period is evenly divided into 36 parts, that is, each sub-time period corresponds to 10 minutes. Each raw data in the raw data sequence can be, for example, the number of orders that occurred in the corresponding sub-time period.

[0038] As an example, the original data sequence corresponding to each sub-time period within the above preset time period may be [2, 5, 3, 2, 3, 1, 0, 1, 3, 2, 6, 8, 7, 1, 0, 1, 7, 2, 5, …].

[0039] The second step is to preprocess the original data sequence to generate the target data sequence within a preset time period.

[0040] In these implementations, the execution subject may preprocess the original data sequence obtained in the first step in various ways to generate a target data sequence within a preset time period. As an example, the preprocessing of the original data sequence may include smoothing the data using various data smoothing methods (e.g., moving average method).

[0041] Based on the above optional implementation methods, this solution can avoid the poor effect caused by directly using the original data for model training by preprocessing the original data sequence accordingly, thereby improving the effect of the method for identifying data anomalies.

[0042] Based on the above optional implementation, optionally, the above execution subject may also pre-process the original data sequence according to the following steps to generate a target data sequence within a preset time period:

[0043] S1. Obtain the minimum statistically detectable threshold corresponding to the original data sequence.

[0044] In these implementations, the execution subject may obtain the minimum statistically possible threshold value corresponding to the original data sequence in various ways, wherein the minimum statistically possible threshold value may be preset, for example, 21.

[0045] S2. Determine the reference value of the original data sequence.

[0046] In these implementations, the execution entity may first determine a reference value of the original data sequence. The reference value may be used to reflect the overall level of the original data sequence. As an example, the reference value may be an average value or a median value. For example, the execution entity may determine the average value 4 of the original data sequence [2,5,3,2,3,1,0,1,3,2,6,8,7,1,0,1,7,2,5,…] as the reference value.

[0047] S3. In response to determining that the reference value is less than the minimum statistically possible threshold, perform the following time period aggregation steps:

[0048] S31 , summing up the original data corresponding to each target number of adjacent sub-time periods in the original data sequence to generate a new original data sequence.

[0049] In these implementations, each new raw data in the above-mentioned new raw data sequence can be the sum of the original raw data corresponding to the corresponding target number of adjacent sub-time periods. As an example, the above-mentioned target number can be a preset value, or it can be a value that changes according to the iteration here (for example, the product of the preset value and the number of iterations). For example, the preset value can be 6, then the first loop will sum every 6 raw data in the raw data sequence, and the second loop will sum every 12 raw data in the raw data sequence (that is, every 2 raw data in the new raw data sequence), and so on. The above-mentioned execution entity can sum the 1st to 6th raw data in the above-mentioned raw sequence to obtain 16, sum the 7th to 12th raw data to obtain 20, sum the 13th to 18th raw data to obtain 18, and so on. Thereby generating a new raw data sequence [16, 20, 18, ...].

[0050] S32. Determine a reference value of a new original data sequence.

[0051] In these implementations, the execution subject may determine the reference value of the original data sequence generated in step S31 in a manner consistent with the aforementioned determination of the reference value. As an example, the reference value of the new original data sequence may be 17, for example.

[0052] S4. In response to determining that the reference value of the new original data sequence is less than the minimum statistically possible threshold, continue to perform the time period aggregation step.

[0053] In these implementations, the execution subject may determine whether the reference value of the new original sequence determined in step S32 is less than the minimum statistically possible threshold. In response to determining that it is less than, the execution subject may continue to perform the time period aggregation step, that is, perform steps S31-S32.

[0054] As an example, in response to determining that the reference value 17 of the new original data sequence is less than the minimum statistically detectable threshold 21, the execution entity may continue to execute the above steps S31 (e.g., generating a new original data sequence [36, 37, …]) and step S32 (e.g., determining the reference value 37 of the new original data sequence).

[0055] S5. In response to determining that the reference value of the new original data sequence is not less than the minimum statistically possible threshold, determine the new original data sequence as a target data sequence within a preset time period.

[0056] In these implementations, in response to determining that the reference value of the new original data sequence is not less than the minimum statistically possible threshold, the execution subject may determine the new original data sequence generated in step S32 as the target data sequence within a preset time period. Thus, each original data in the new original data sequence may be the sum of several adjacent original data in the original original sequence.

[0057] Based on the optional implementation method described in the above steps S1-S5, this solution amplifies the statistical dimension of the data to avoid the situation where it is difficult to reflect the data fluctuation trend due to the fact that some business data is relatively small in unit time (such as minutes) and is often 0, thereby more scientifically, accurately and explicitly reflecting the fluctuation trend of data in multiple time units.

[0058] Based on the optional implementation methods described in the first and second steps above, optionally, the execution subject may also pre-process the original data sequence according to the following steps to generate a target data sequence within a preset time period:

[0059] S'1. Determine the reference value of the original data sequence.

[0060] In these implementations, the execution entity may first determine a reference value of the original data sequence. The reference value may be used to reflect the overall level of the original data sequence. As an example, the reference value may be an average value or a median.

[0061] S'2. Based on the comparison between each original data in the original data sequence and the reference value, the original data in the original data sequence is remapped to generate mapped data.

[0062] In these implementations, based on each original data in the original data sequence and the reference value determined in step S'1, the execution subject may remap the original data in the original data sequence in various ways to generate mapped data.

[0063] As an example, the above execution entity can be remapped according to the following steps:

[0064] S'21. Determine the difference value (eg, difference or ratio) between the original data in the original data sequence and the reference value.

[0065] S'22. Obtain a preset segment multiple list and the number of elements included in the segment multiple list.

[0066] In these implementations, the elements in the above segment multiple list are arranged in descending order of value. As an example, the above segment multiple list can be [x, x / 2, x / 3, ... x / (x-2), x / (x-1), 1, (x-1) / x, (x-2) / x.., 2 / x, 1 / x, 0], where x∈N + And x≥3. The value of x can be preset according to the actual application scenario. For example, when x=3, the segment multiple list can be [3, 3 / 2, 1, 2 / 3, 1 / 3, 0].

[0067] S'23. Select an element that matches the difference value corresponding to the original data in the original data sequence from the preset segmentation multiple list as the target segmentation multiple.

[0068] In these implementations, the above-mentioned matching element can be, for example, the maximum value in the above-mentioned preset segmentation multiple list that is not greater than the above-mentioned difference value (e.g., ratio). As an example, the original data sequence can be [36, 37, 38, ...], and the reference value can be 37. For the first original data, the above-mentioned execution entity can use the ratio 36 / 37 as the difference value. Thus, the above-mentioned target segmentation multiple can be 2 / 3. Other data in the above-mentioned original data sequence can be deduced in the same way.

[0069] S'24. Generate mapped data corresponding to each original data in the original data sequence according to the target segmentation multiple and the associated information, the original data and the corresponding difference value.

[0070] In these implementations, the execution subject may generate the mapped data corresponding to each original data according to the target segmentation multiple corresponding to each original data in the original data sequence and the associated information, the original data and the corresponding difference value. The associated information may include but is not limited to at least one of the following: the maximum value of the remapped data range, the index of the target segmentation multiple in the segmentation multiple list, and the next segmentation multiple of the target segmentation multiple.

[0071] As an example, the mapped data score can be calculated by the following formula:

[0072] score=M-idx+(value-dif*list[idx]) / [(list[idx+1]-list[idx])*dif]

[0073] Wherein, M can be the maximum value of the remapped data range, that is, the remapped data range is [0, M]. idx can be the index of the target segmentation multiple in the above segmentation multiple list list. value can be the original data in the original data sequence. dif can be the above difference value (such as ratio).

[0074] It should be noted that other data in the above original data sequence can be deduced in this way, so as to generate mapped data corresponding to each original data in the original data sequence.

[0075] S'3. Based on the mapped data, generate a target data sequence within a preset time period.

[0076] In these implementations, based on the mapped data generated in step S2, the execution subject may generate a target data sequence within a preset time period in various ways. As an example, the execution subject may arrange the mapped data generated in a time sequence to generate a target data sequence within the preset time period. As another example, the execution subject may aggregate the target data within the preset time period of the mapped data arranged in a time sequence according to the time period described in steps S1-S5, and generate the target data within the preset time period according to the data corresponding to the aggregated time period.

[0077] Based on the optional implementation described in the above steps S'1-S'3, this solution remaps the original data with a value range of [0, +∞) according to the difference value generated based on the reference value, and retains the difference between the original data on the basis of ensuring that the range of the mapped data is reduced. Optionally, the risk of sharp points caused by proportional enlargement or reduction between the mapped data is reduced according to the above specific formula.

[0078] Step 202, determining the predicted value corresponding to the target data sequence.

[0079] In this embodiment, according to the data attribute information of the data set to be replaced obtained in step 201, the execution subject can determine the target candidate data set from the preset candidate data set set in various ways. As an example, the execution subject can generate a fitting curve of the target data sequence using various data fitting methods, and generate the predicted value corresponding to the target data sequence using the fitting curve. Based on the example of step 201, the predicted value corresponding to the target data sequence can be used to characterize the predicted value of the target data corresponding to 18:00-18:59.

[0080] In some optional implementations of this embodiment, the above-mentioned execution subject can input the target data sequence into a pre-trained time series prediction model to obtain a corresponding prediction value. Among them, the above-mentioned time series prediction model is usually trained based on the target training sample. The above-mentioned target training sample usually includes sample data with a time span not less than the time span corresponding to the above-mentioned target data sequence. Optionally, the above-mentioned target training sample includes a time span that is usually not less than twice the period corresponding to the above-mentioned target data sequence.

[0081] In these implementations, the above-mentioned time series prediction model may include, for example, an autoregressive integrated moving average model (ARIMA) and a long short-term memory (LSTM) network.

[0082] Step 203: extract data feature indicators based on the target data sequence and the predicted value.

[0083] In this embodiment, based on the target data sequence acquired in step 201 and the predicted value determined in step 202, the execution entity may extract data feature indicators in various ways. Among them, the data feature indicators may be used to characterize the characteristics of the target data sequence and the predicted value. As an example, the data feature indicators may include, for example, the number of occurrences of an outlier (e.g., 0). As another example, the data feature indicators may include, for example, the maximum number of consecutive occurrences of an outlier (e.g., 0). For example, if an outlier appears 3 times, 2 times, and 5 times consecutively, the maximum number of consecutive occurrences of the outlier is 5.

[0084] In some optional implementations of this embodiment, the above-mentioned execution subject may extract data characteristic indicators according to the following steps:

[0085] The first step is to obtain at least one historical data sequence matching a preset time period.

[0086] In these implementations, the historical time period corresponding to the above historical data sequence can match the above preset time period. As an example, the above-mentioned preset time period can be, for example, 12:00-18:00 on May 18, 2020. The historical time period corresponding to the above-mentioned at least one historical data series can include, but is not limited to, at least one of the following: 12:00-17:59 on May 18, 2019, 12:00-17:59 on April 18, 2020, 12:00-17:59 on May 17, 2020, 6:00-11:59 on May 18, 2020, 0:00-5:59 on May 18, 2020, 6:00-11:59 on May 18, 2020, 18:00-23:59 on May 18, 2019, 18:00-23:59 on April 18, 2020, and 18:00-23:59 on May 17, 2020.

[0087] The second step is to extract the first data feature indicator based on the target data sequence and at least one historical data sequence.

[0088] In these implementations, the first data characteristic indicator may be used to indicate the difference between the target data in a preset time period and the data in a historical time period. The first data characteristic indicator may include, but is not limited to, at least one of the following: a month-on-month difference, a year-on-year difference, and a year-on-year difference of the sum of data for multiple time periods. The difference value may be, for example, a ratio or difference between real values.

[0089] As an example, the difference value of the month-on-month value can be used to represent the difference in the real value between different historical time periods within the same day.

[0090] As another example, the difference value of the year-on-year value may represent the difference in real values ​​between the same historical time periods corresponding to different days.

[0091] As another example, the difference value of the sum of multiple time period data compared with the same period can, for example, represent the difference value between the sum of multiple data in the current time period and the previous m statistical periods in the previous x days and the sum of multiple data in the current time period of the day and the previous m statistical periods.

[0092] The third step is to extract the second data feature index based on the data in the target data sequence and the corresponding predicted value.

[0093] In these implementations, the second data characteristic indicator may be used to indicate the difference between the data in the same time period and the corresponding predicted value. The difference may be, for example, a difference or a ratio.

[0094] The fourth step is to extract the third data characteristic index based on the target data sequence and at least one historical data sequence.

[0095] In these implementations, the third data characteristic indicator may be used to indicate a data change trend within a time period. As an example, the data change trend may be represented by a slope corresponding to the data corresponding to the selected time period.

[0096] Step 204 , input the data feature index into the pre-trained autonomous learning model to generate prompt information for indicating whether there is data anomaly.

[0097] In this embodiment, the above-mentioned execution entity can input the data feature indicators generated by step 203 into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is data anomaly. Among them, the above-mentioned autonomous learning model can be obtained through active learning training based on a pre-constructed classifier. The above-mentioned classifier may include but is not limited to at least one of the following: naive Bayes, decision tree, logistic regression, support vector machine and neural network. This article adopts the XGBoost (eXtreme Gradient Boosting) algorithm as a classifier. The above-mentioned active learning method may include but is not limited to at least one of the following: generative member query, streaming active learning method, active learning method based on unlabeled sample pool, batch active learning method, semi-supervised active learning method, and learning method combined with generative adversarial network.

[0098] In some optional implementations of this embodiment, based on preprocessing including remapping, the above-mentioned execution entity can inversely transform the output values ​​calculated by the autonomous learning model during the training process, and transform the data range to [0, +∞).

[0099] In some optional implementations of this embodiment, the above execution subject may further continue to perform the following steps:

[0100] Step 205: In response to determining that the prompt information is used to indicate the existence of data anomaly, an alarm message is sent to the target end.

[0101] In these implementations, in response to determining that the prompt information is used to indicate the presence of data anomalies, the execution subject may send an alarm message to the target end, which may be a terminal used by a technician or a monitoring terminal of a central control system, which is not limited here.

[0102] Step 206: Receive alarm processing feedback information fed back by the target end.

[0103] In these implementations, the alarm processing feedback information may include at least one of the following: abnormal data severity level, alarm response time. The abnormal data severity level may be, for example, "high severity level", "medium severity level" or "low severity level" marked by the technical personnel using the target end. The alarm response time may be, for example, the time interval from sending the alarm information to receiving the "go to process" information sent by the target end.

[0104] Step 207: Generate training samples for training an abnormality classification model based on the alarm processing feedback information.

[0105] In these implementations, the training samples may include a label value for characterizing the severity of the anomaly. The label value may correspond to the severity level of the anomaly data and the alarm response time received in step 206. For example, the label value of "low severity level" and the alarm response time greater than 30 minutes may be 0.1; the label value of "medium severity level" and the alarm response time greater than 10 minutes and less than 30 minutes may be 0.5; the label value of "high severity level" and the alarm response time less than 10 minutes may be 1.0.

[0106] Based on the above optional implementation methods, this solution can send abnormal information to the designated recipient. Moreover, it can also supplement sample annotation data by recording the abnormal information recipient's feedback on the alarm information (including whether it is abnormal data and the severity of the problem, etc.), thereby reducing the cost of sample annotation.

[0107] Continue to see Figure 3 , Figure 3 is a schematic diagram of an application scenario of a method for identifying data anomalies according to an embodiment of the present disclosure. Figure 3 In the application scenario, a user registers a new user using terminal devices 301, 302, 303, etc. The server 304 can obtain a sequence 305 formed by the number of new user registrations every 15 minutes in the time period of 6:00-6:59. The server 304 determines the predicted value 20 (such as Figure 3 306). Based on the above sequence 305 and the predicted value 306, the server 304 can extract the data feature index 307. The above data feature index 307 can include, for example, the ratio of the number of new user registrations every 15 minutes in the time period of 6:00-6:59 on the previous day to the above sequence 306 (0.67, 1, 0.94, 1.08), the ratio of the actual value to the predicted value of the time period 6:45-6:59 is 0.9, the number of values ​​less than the preset value 10 is 1, and the number of values ​​0 is 0. The server 304 can input the above data feature index 307 into the pre-trained autonomous learning model to generate a prompt information 308 for characterizing the existence of data anomalies.

[0108] At present, one of the existing technologies is usually to manually monitor data changes or set fluctuation thresholds to determine whether the data is abnormal, resulting in excessively high labor costs and poor accuracy. The method provided by the above-mentioned embodiment of the present disclosure extracts data feature indicators based on the target data sequence and the predicted value of the determined target sequence, and then uses the extracted data feature indicators as the input of the autonomous learning model trained by the machine learning method, thereby obtaining prompt information for characterizing whether there is data anomaly. On the one hand, this significantly reduces the workload of personnel, and on the other hand, it improves the accuracy of data anomaly identification by selecting data feature indicators and combining them with machine learning models. Moreover, since the amount of data in the prior art is usually large, the annotation cost required for ordinary machine learning model training is too high. The use of an autonomous learning model can meet the training requirements while reducing the amount of sample annotation.

[0109] Further references Figure 4 , which shows a process 400 of an embodiment of a method for training an anomaly classification model. The process 400 of the method for training an anomaly classification model includes the following steps:

[0110] Step 401: Obtain a training sample set.

[0111] In this embodiment, the execution body (eg, Figure 1 The server 105 shown in the figure can obtain the training sample set in various ways. Among them, the training samples in the above training sample set can include sample data features and corresponding annotation values. The above sample data features can be generated based on the data statistical features and time period comparison features in the historical time period. Among them, the above data statistical features can include but are not limited to at least one of the following: at least one of the mean, variance, maximum value, and minimum value of the values ​​of the recent several time periods, at least one of the mean, variance, maximum value, and minimum value of the values ​​of the time period in the recent several days, the maximum number of consecutive values ​​of 0 in the recent several time periods, the number of 0s, and the number of consecutive alarms. The above data comparison features can include but are not limited to at least one of the following: the difference or ratio between the value of the current time period and the value of the previous N time periods, the difference or ratio between the value of the current time period and the value of the same time period in the previous N days, the difference or ratio between the data of the current time period and the model prediction value, and the slope value of the data associated with the current time period.

[0112] In this embodiment, the above-mentioned marked value can be generated based on the method described in steps 205 to 207 of the above-mentioned embodiment.

[0113] Step 402: Taking the sample data features of the training samples in the training sample set as input, taking the label values ​​corresponding to the input sample data features as expected output, and training to obtain a quasi-abnormal classification model.

[0114] In this embodiment, the execution subject may perform training based on the initial model by machine learning method using the training sample set obtained in step 401 to obtain the quasi-abnormal classification model. The initial model may include a regression model, for example.

[0115] Step 403: cluster the output values ​​corresponding to the training samples in the training sample set, and generate representative values ​​corresponding to the categories of the target number of clusters as references for level determination.

[0116] In this embodiment, the execution subject may cluster the output values ​​corresponding to the training samples in the training sample set, and generate representative values ​​corresponding to the categories of the target number of clusters as a reference for grade determination. The representative value may be, for example, the mean of the output values ​​belonging to the same category or the value corresponding to the cluster center of the category. As an example, the output values ​​corresponding to the training samples in the training sample set generate three clusters, and the corresponding representative values ​​are 0.7, 0.4, and 0.2, respectively. Then the execution subject may use the representative values ​​as a reference for grade determination.

[0117] It should be noted that the above output value is usually the output of the trained quasi-abnormal classification model or the output of the model in the later stage of training, rather than the output in the initial stage of training, so as to improve the accuracy of subsequent classification.

[0118] Step 404: Generate an abnormality classification model based on the representative value and the quasi-abnormality classification model.

[0119] In this embodiment, the execution subject may generate an abnormality grading model according to the representative value generated in step 403 and the quasi-abnormality grading model obtained in step 402. The abnormality grading model may be used to characterize the correspondence between data features and data abnormality levels. As an example, the execution subject may connect the quasi-abnormality grading model with the correspondence between the representative value and the data abnormality level (e.g., high, medium, low) to form the abnormality grading model.

[0120] from Figure 4 It can be seen that the process 400 of the method for training anomaly classification model in this embodiment embodies the steps of using training samples including labeled values ​​to train the model, and clustering the output values ​​corresponding to the training samples as a reference for grade determination. Therefore, the scheme described in this embodiment can autonomously select the classification threshold of the anomaly classification model according to the training samples, which is faster and more adaptable than manual setting or manual parameter adjustment, thereby helping to improve the accuracy of abnormal data abnormal grade identification.

[0121] Further references Figure 5As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for identifying data anomalies. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0122] like Figure 5 As shown, the apparatus 500 for identifying data anomalies provided in this embodiment includes an acquisition unit 501, a determination unit 502, an extraction unit 503 and a generation unit 504. The acquisition unit 501 is configured to acquire a target data sequence within a preset time period; the determination unit 502 is configured to determine a predicted value corresponding to the target data sequence; the extraction unit 503 is configured to extract data feature indicators based on the target data sequence and the predicted value; the generation unit 504 is configured to input the data feature indicators into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is a data anomaly.

[0123] In the apparatus 500 for identifying data anomalies, the specific processing of the acquisition unit 501, the determination unit 502, the extraction unit 503 and the generation unit 504 and the technical effects thereof can be referred to in Figure 2 The relevant descriptions of step 201, step 202, step 203 and step 204 in the corresponding embodiment are not repeated here.

[0124] In some optional implementations of this embodiment, the acquisition unit 501 may include: an acquisition module (not shown in the figure) and a generation module (not shown in the figure). The acquisition module may be configured to acquire the original data sequence corresponding to each sub-time period within a preset time period. The generation module may be configured to pre-process the original data sequence to generate a target data sequence within a preset time period.

[0125] In some optional implementations of the present embodiment, the above-mentioned generation module can be further configured to: obtain a minimum statistically detectable threshold value corresponding to the original data sequence; determine a reference value of the original data sequence; in response to determining that the reference value is less than the minimum statistically detectable threshold value, perform the following time period aggregation step: sum the original data corresponding to each target number of adjacent sub-time periods in the original data sequence to generate a new original data sequence; determine a reference value for the new original data sequence; in response to determining that the reference value of the new original data sequence is less than the minimum statistically detectable threshold value, continue to perform the time period aggregation step; in response to determining that the reference value of the new original data sequence is not less than the minimum statistically detectable threshold value, determine the new original data sequence as the target data sequence within a preset time period.

[0126] In some optional implementations of the present embodiment, the above-mentioned generation module can be further configured to: determine a reference value of the original data sequence; based on the comparison between each original data in the original data sequence and the reference value, remap the original data in the original data sequence to generate mapped data; based on the mapped data, generate a target data sequence within a preset time period.

[0127] In some optional implementations of this embodiment, the determination unit 502 may be further configured to: input the target data sequence into a pre-trained time series prediction model to obtain a corresponding prediction value. The time series prediction model may be trained based on a target training sample. The target training sample may include sample data whose time span is not less than the time span corresponding to the target data sequence.

[0128] In some optional implementations of this embodiment, the above-mentioned extraction unit 503 can be further configured to: obtain at least one historical data sequence matching a preset time period, wherein the historical time period corresponding to the historical data sequence matches the preset time period; extract a first data feature indicator based on the target data sequence and at least one historical data sequence, wherein the first data feature indicator is used to indicate the difference between the target data within the preset time period and the data within the historical time period; extract a second data feature indicator based on the data in the target data sequence and the corresponding predicted value, wherein the second data feature indicator is used to indicate the difference between the data within the same time period and the corresponding predicted value; extract a third data feature indicator based on the target data sequence and at least one historical data sequence, wherein the third data feature indicator is used to indicate the data change trend within the time period.

[0129] In some optional implementations of this embodiment, the device for identifying data anomalies may also include: a sending unit (not shown in the figure), configured to send alarm information to the target end in response to determining that the prompt information is used to characterize the existence of data anomalies; a receiving unit (not shown in the figure), configured to receive alarm processing feedback information fed back by the target end, wherein the alarm processing feedback information may include at least one of the following: abnormal data severity level, alarm response time; a sample generation unit (not shown in the figure), configured to generate training samples for training anomaly classification model based on the alarm processing feedback information, wherein the training samples may include labeled values ​​for characterizing the severity of the anomaly.

[0130] The device provided by the above-mentioned embodiment of the present disclosure extracts data feature indicators based on the target data sequence acquired by the acquisition unit 501 and the predicted value of the target sequence determined by the determination unit 502 through the extraction unit 503, and then uses the extracted data feature indicators as the input of the autonomous learning model trained by the machine learning method through the generation unit 504, so as to obtain prompt information for characterizing whether there is data anomaly. Thus, on the one hand, the workload of personnel is significantly reduced, and on the other hand, the accuracy of data anomaly identification is improved by selecting data feature indicators and combining them with machine learning models. Moreover, since the amount of data in the prior art is usually large, the annotation cost required for ordinary machine learning model training is too high. The use of autonomous learning models can meet the training requirements under the premise of reducing the amount of sample annotation.

[0131] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for training an abnormal classification model. Figure 4 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0132] like Figure 6 As shown, the apparatus 600 for identifying data anomalies provided in this embodiment includes a training sample acquisition unit 601, a training unit 602, a clustering unit 603 and a model generation unit 604. The training sample acquisition unit 601 is configured to acquire a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on the data statistical features and time period comparison features in the historical time period, and the annotation values ​​are generated based on the method described in the above embodiment; the training unit 602 is configured to take the sample data features of the training samples in the training sample set as input, and take the annotation values ​​corresponding to the input sample data features as the expected output, and train to obtain a quasi-abnormal classification model; the clustering unit 603 is configured to cluster the output values ​​corresponding to the training samples in the training sample set, and generate representative values ​​corresponding to the categories of the target number of clusters as a reference for grade determination; the model generation unit 604 is configured to generate an abnormal classification model according to the representative values ​​and the quasi-abnormal classification model, wherein the abnormal classification model is used to characterize the corresponding relationship between data features and data abnormality levels.

[0133] In the present embodiment, in the apparatus 600 for training an abnormal classification model, the specific processing of the training sample acquisition unit 601, the training unit 602, the clustering unit 603 and the model generation unit 604 and the technical effects thereof can be referred to respectively. Figure 4 The relevant descriptions of step 401, step 402, step 403 and step 404 in the corresponding embodiment are not repeated here.

[0134] The apparatus provided by the above-mentioned embodiment of the present disclosure uses the training samples including the labeled values ​​acquired by the acquisition unit 601 to perform model training through the training unit 602, and the clustering unit 603 clusters the output values ​​corresponding to the training samples as a reference for grade determination, and generates an abnormality grading model through the model generation unit 604. Therefore, the classification threshold of the abnormality grading model can be selected autonomously according to the training samples, which is faster and more adaptable than manual setting or manual parameter adjustment, thereby helping to improve the accuracy of abnormality grade identification of abnormal data.

[0135] Reference below Figure 7 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of server 105)700. Figure 7 The server shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0136] like Figure 7 As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0137] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Figure 7 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 7 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0138] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present application are executed.

[0139] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0140] The computer-readable medium may be included in the electronic device; or may exist independently without being installed in the server. The computer-readable medium carries one or more programs. When the one or more programs are executed by the server, the server: obtains a target data sequence within a preset time period; determines a predicted value corresponding to the target data sequence; extracts data feature indicators based on the target data sequence and the predicted value; inputs the data feature indicators into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is a data anomaly; or the server: obtains a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on data statistical features and time period comparison features within a historical time period, and the annotation values ​​are generated based on the method of the first aspect; the sample data features of the training samples in the training sample set are used as input, and the annotation values ​​corresponding to the input sample data features are used as expected outputs to train a quasi-abnormal grading model; clusters the output values ​​corresponding to the training samples in the training sample set, and generates representative values ​​corresponding to the categories of the target number of clusters as a reference for grade determination; generates an abnormal grading model based on the representative values ​​and the quasi-abnormal grading model, wherein the abnormal grading model is used to characterize the correspondence between data features and data anomaly levels.

[0141] Computer program code for performing the operations of the embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C", Python, or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0143] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be arranged in a processor, for example, may be described as: a processor including an acquisition unit, a determination unit, an extraction unit, and a generation unit. Alternatively, it may be described as: a processor including a training sample acquisition unit, a training unit, a clustering unit, and a model generation unit. Among them, the names of these units do not constitute a limitation on the units themselves in certain cases. For example, the acquisition unit may also be described as a "unit for acquiring a target data sequence within a preset time period."

[0144] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for identifying data anomalies, include: Obtain target data sequence within a preset time period; Determining a predicted value corresponding to the target data sequence; Acquire at least one historical data sequence matching the preset time period, wherein the historical time period corresponding to the historical data sequence matches the preset time period; extract a first data feature indicator based on the target data sequence and the at least one historical data sequence, wherein the first data feature indicator is used to indicate the difference between the target data within the preset time period and the data within the historical time period; extract a second data feature indicator based on the data in the target data sequence and the corresponding predicted value, wherein the second data feature indicator is used to indicate the difference between the data within the same time period and the corresponding predicted value; extract a third data feature indicator based on the target data sequence and the at least one historical data sequence, wherein the third data feature indicator is used to indicate the data change trend within the time period; The data feature indicators are input into a pre-trained autonomous learning model to generate prompt information for characterizing whether there is data anomaly.

2. The method according to claim 1, in, The step of obtaining a target data sequence within a preset time period includes: Obtaining the original data sequence corresponding to each sub-time period within the preset time period; The original data sequence is preprocessed to generate a target data sequence within the preset time period.

3. The method according to claim 2, in, The preprocessing of the original data sequence to generate the target data sequence within the preset time period includes: Obtaining a minimum statistically feasible threshold value corresponding to the original data sequence; Determining a reference value of the original data sequence; In response to determining that the reference value is less than the minimum statistically possible threshold, performing the following time period aggregation step: summing the original data corresponding to each target number of adjacent sub-time periods in the original data sequence to generate a new original data sequence; determining a reference value for the new original data sequence; In response to determining that the reference value of the new original data sequence is less than the minimum statistically detectable threshold, continuing to perform the time period aggregation step; In response to determining that the reference value of the new original data sequence is not less than the minimum statistically possible threshold, the new original data sequence is determined as the target data sequence within the preset time period.

4. The method according to claim 2, in, The preprocessing of the original data sequence to generate the target data sequence within the preset time period includes: Determining a reference value of the original data sequence; Based on the comparison between each original data in the original data sequence and the reference value, remapping the original data in the original data sequence to generate mapped data; Based on the mapped data, a target data sequence within the preset time period is generated.

5. The method according to claim 1, in, The determining the predicted value corresponding to the target data sequence includes: The target data sequence is input into a pre-trained time series prediction model to obtain a corresponding prediction value, wherein the time series prediction model is trained based on a target training sample, and the target training sample includes sample data with a time span not less than a time span corresponding to the target data sequence.

6. The method according to any one of claims 1 to 5, in, The method further comprises: In response to determining that the prompt information is used to indicate the existence of data anomaly, sending an alarm message to the target end; Receiving alarm processing feedback information fed back by the target end, wherein the alarm processing feedback information includes at least one of the following: abnormal data severity level, alarm response time; Based on the alarm processing feedback information, a training sample for training an abnormality classification model is generated, wherein the training sample includes a label value for characterizing the severity of the abnormality.

7. A method for training an anomaly classification model, include: Obtain a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on data statistical features and time period comparison features within a historical time period, and the annotation values ​​are generated based on the method of claim 6; Taking the sample data features of the training samples in the training sample set as input, taking the label values ​​corresponding to the input sample data features as expected output, and training to obtain a quasi-abnormal classification model; Clustering the output values ​​corresponding to the training samples in the training sample set, generating representative values ​​corresponding to the categories of the target number of clusters as a reference for level determination; The abnormality grading model is generated according to the representative value and the quasi-abnormality grading model, wherein the abnormality grading model is used to characterize the corresponding relationship between data features and data abnormality levels.

8. A device for training an anomaly classification model, include: A training sample acquisition unit, configured to acquire a training sample set, wherein the training samples in the training sample set include sample data features and corresponding annotation values, the sample data features are generated based on data statistical features and time period comparison features within a historical time period, and the annotation values ​​are generated based on the method of claim 6; A training unit is configured to take the sample data features of the training samples in the training sample set as input, take the label values ​​corresponding to the input sample data features as expected output, and train to obtain a quasi-abnormal classification model; A clustering unit configured to cluster the output values ​​corresponding to the training samples in the training sample set, and generate representative values ​​corresponding to the categories of a target number of clusters as a reference for level determination; The model generation unit is configured to generate the abnormality classification model according to the representative value and the quasi-abnormality classification model, wherein the abnormality classification model is used to characterize the corresponding relationship between data features and data abnormality levels.

9. A device for identifying data anomalies, include: An acquisition unit, configured to acquire a target data sequence within a preset time period; a determination unit, configured to determine a prediction value corresponding to the target data sequence; An extraction unit is configured to obtain at least one historical data sequence matching the preset time period, wherein the historical time period corresponding to the historical data sequence matches the preset time period; extract a first data feature indicator based on the target data sequence and the at least one historical data sequence, wherein the first data feature indicator is used to indicate the difference between the target data within the preset time period and the data within the historical time period; extract a second data feature indicator based on the data in the target data sequence and the corresponding predicted value, wherein the second data feature indicator is used to indicate the difference between the data within the same time period and the corresponding predicted value; extract a third data feature indicator based on the target data sequence and the at least one historical data sequence, wherein the third data feature indicator is used to indicate the data change trend within the time period; The generating unit is configured to input the data feature index into a pre-trained autonomous learning model to generate prompt information for indicating whether there is data anomaly.

10. A server, include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

11. A computer readable medium having a computer program stored thereon, in, When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A method and apparatus for generating information

    CN109388548A

  • Monitoring index abnormity detection method, model training method, device and equipment

    CN110008079A