Data processing method and device

By automatically identifying and deleting periodic abnormality detection points, the problems of errors and missed deletion in the existing technology are solved, and the risk control effect of risk control services is improved.

CN120123636APending Publication Date: 2025-06-10BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311685100.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In risk control scenarios, it is difficult for the prior art to automatically identify and delete abnormal detection points that occur periodically, resulting in errors and missed deletion, affecting business risk control.

Method used

By acquiring the first business data sequence and multiple reference business data sequences, their similarity is determined, and periodic detection points are automatically deleted according to the abnormal time series to avoid manual intervention.

Benefits of technology

It improves the accuracy of the detection results, avoids the error and missed deletion of abnormal detection points, and effectively controls the risk of actual business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123636A_ABST
    Figure CN120123636A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device. The method comprises the following steps: acquiring a first service data sequence and a plurality of reference service data sequences; determining the similarity between the first service data sequence and each reference service data sequence, and determining at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence; generating an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamp of each second data point; and deleting the first data point in the first service data sequence under the condition of determining that the target data point exists in the second data points of the at least one second service data sequence according to the abnormal time sequence, so that manual detection of periodically occurring abnormal detection points and manual deletion of the abnormal detection points are not needed, and the detection efficiency is improved. Error deletion and missing deletion of abnormal detection points are avoided, and risk control is effectively carried out on actual services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and in particular, to a data processing method and apparatus. Background Art

[0002] In a risk control scenario, a monitoring system inspects business data (such as login volume, transaction volume, etc.), and issues an alarm for the detected abnormal detection points. However, many of the abnormal detection points are periodically detected, that is, they will appear once every fixed period of time. The abnormally detected points that appear periodically do not affect the actual business. Therefore, it is necessary to exclude these periodically appearing abnormal detection points to avoid ineffective alarms.

[0003] In the related art, the periodically appearing abnormal detection points are detected manually and deleted manually. The labor cost is high and the accuracy of the detection result is low, resulting in incorrect deletion and missed deletion, thus affecting the actual business and being unable to effectively control the risks of the actual business. Summary of the Invention

[0004] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.

[0005] To this end, the present disclosure provides a data processing method and apparatus, which can generate an abnormal time series according to the first acquisition time of the abnormal first data point in the first business data sequence and the second acquisition time of the abnormal second data point in at least one second business data sequence similar to the first business data sequence. When it is determined according to the abnormal time series that the first data point in the first business data sequence is obtained by periodic detection, the first data point in the first business data sequence is automatically deleted, without manually detecting the periodically appearing abnormal detection points and manually deleting these abnormal detection points, improving the accuracy of the detection result, while avoiding incorrect deletion and missed deletion of abnormal detection points, and effectively controlling the risks of the actual business.

[0006] An embodiment of the first aspect of the present disclosure provides a data processing method, including: obtaining a first service data sequence and a plurality of reference service data sequences, where the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; determining the similarity between the first service data sequence and each of the reference service data sequences, and determining at least one second service data sequence from each of the reference service data sequences according to the similarity of each of the reference service data sequences; generating an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of the second data points; deleting the first data point in the first service data sequence when it is determined that there is a target data point among the second data points of the at least one second service data sequence according to the abnormal time sequence; where the acquisition timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same.

[0007] The data processing method of the embodiment of the present disclosure, by obtaining a first service data sequence and a plurality of reference service data sequences, where the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; determining the similarity between the first service data sequence and each of the reference service data sequences, and determining at least one second service data sequence from each of the reference service data sequences according to the similarity of each of the reference service data sequences; generating an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of the second data points; deleting the first data point in the first service data sequence when it is determined that there is a target data point among the second data points of the at least one second service data sequence according to the abnormal time sequence; where the acquisition timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same. Thus, according to the first acquisition moment of the abnormal first data point in the first service data sequence and the second acquisition moments of the abnormal second data points in at least one second service data sequence similar to the first service data sequence, an abnormal time sequence is generated, and when it is determined according to the abnormal time sequence that the first data point in the first service data sequence is obtained by periodic detection, the first data point in the first service data sequence is automatically deleted, improving the accuracy of the detection result, and there is no need to manually detect the periodically occurring abnormal detection points and manually delete these abnormal detection points, avoiding the misdeletion and omission of abnormal detection points, and effectively controlling the risks of the actual business.

[0008] A second aspect embodiment of the present disclosure provides a data processing apparatus, including: an acquisition module, configured to acquire a first service data sequence and a plurality of reference service data sequences, where the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; a determination module, configured to determine the similarity between the first service data sequence and each of the reference service data sequences, and determine at least one second service data sequence from each of the reference service data sequences according to the similarity of each of the reference service data sequences; a generation module, configured to generate an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each of the second data points; a deletion module, configured to delete the first data points in the first service data sequence when it is determined that there are target data points among the second data points of at least one second service data sequence according to the abnormal time sequence; where the acquisition timestamp interval between any two adjacent elements in the data sequence formed by the target data points and the first data points is the same.

[0009] A third aspect embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the data processing method described in the first aspect embodiment of the present disclosure is implemented.

[0010] A fourth aspect embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data processing method described in the first aspect embodiment of the present disclosure is implemented.

[0011] A fifth aspect embodiment of the present disclosure provides a computer program product. When the instruction processor in the computer program product executes, the data processing method described in the first aspect embodiment of the present disclosure is implemented.

[0012] Additional aspects and advantages of the present disclosure will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present disclosure. Description of the Drawings

[0013] The above and / or additional aspects and advantages of the present disclosure will become apparent and be easily understood from the following description of the embodiments in conjunction with the drawings, where:

[0014] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure;

[0015] Figure 2 It is a schematic flowchart of another data processing method provided by an embodiment of the present disclosure;

[0016] Figure 3Schematic flowchart of another data processing method provided by an embodiment of the present disclosure;

[0017] Figure 4 Schematic diagram of a distance matrix provided by an embodiment of the present disclosure;

[0018] Figure 5 Schematic diagram of path search for a distance matrix provided by an embodiment of the present disclosure;

[0019] Figure 6 Schematic diagram of the structure of a data processing device provided by an embodiment of the present disclosure;

[0020] Figure 7 Block diagram of an electronic device for data processing shown according to an exemplary embodiment. Detailed implementation manners

[0021] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0022] It should be noted that in the technical solution of the present disclosure, in terms of the collection, gathering, updating, analysis, processing, use, transmission, storage, etc. of the user's personal information, it complies with the provisions of relevant laws and regulations, is used for legal purposes, and does not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the security of the user's personal information, network security, and national security.

[0023] Currently, to detect whether a sequence is a periodic sequence, a common method is to divide the sequence into segments according to the period. Then, the similarity of each segment is judged. When the similarity is higher than the specified threshold, it can be determined that the sequence is a periodic sequence. Therefore, the periodicity problem can be solved by converting it into a problem of judging sequence similarity.

[0024] Currently, the commonly used methods for measuring sequence similarity are the Dynamic Time Warping (DTW) algorithm and the Longest Common Subsequence (LCS). Among them, DTW is a dynamic programming algorithm for calculating the similarity of two time series data (especially two sequence data with different time lengths); the Longest Common Subsequence (LCS) is used to find the longest subsequence among all sequences in a sequence set (usually two sequences), and calculating the proportion of this subsequence in the two sequences can be used to measure the similarity between the two sequences.

[0025] However, the DTW algorithm can be used to measure which two of the three sequences are more similar and cannot measure the similarity between two time series. The LCS algorithm can be well applied to measure the similarity between two sequences in general scenarios. However, in the anomaly detection scenario, what usually needs to be measured are two anomaly sequences, sudden increases or sudden decreases. At this time, the effectiveness of the LCS algorithm will be severely reduced.

[0026] In view of the above problems, the present disclosure proposes a data processing method and apparatus.

[0027] The following describes the data processing method and apparatus according to the embodiments of the present disclosure with reference to the accompanying drawings.

[0028] Figure 1 It is a schematic flowchart of a data processing method provided by an embodiment of the present disclosure.

[0029] As Figure 1 shown, the data processing method may include the following steps:

[0030] Step 101, obtain a first service data sequence and a plurality of reference service data sequences.

[0031] Among them, the first service data sequence includes at least one first data point of an anomaly, and the reference service data sequence includes at least one second data point of an anomaly.

[0032] As an example, service data is collected within a first set time period, and the collected data is sorted in the format of a time series to obtain a first service data sequence; service data is collected within a second set time period, and the collected data is sorted according to the time sequence to obtain reference service data. The second set time period may be before the first set time period.

[0033] Among them, the first service data sequence may include at least one first data point of an anomaly. The first data point of an anomaly may be determined according to a set number of data points before the first data point and / or a set number of data points after the first data point. For example, if the first data point in the first service data sequence is 10,000, the values of the 100 data points before the first data point in the first service data sequence are 1,000, and the values of the 100 data points after the first data point in the first service data sequence are also 1,000, then the first data point is an anomaly.

[0034] As an example, the reference service data sequence may include at least one second data point of an anomaly.

[0035] Step 102, determine the similarity between the first service data sequence and each reference service data sequence, and determine at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence.

[0036] In the embodiments of the present disclosure, at least one second service data sequence similar to the first service data sequence can be determined from each reference service data sequence.

[0037] As an example, the similarity between the first service data sequence and each reference service data sequence can be determined, and at least one second service data sequence can be determined from each reference service data sequence according to the similarity of each reference service data sequence, where the similarity between the second service data sequence and the first service data sequence is greater than a set similarity threshold (e.g., 0.85).

[0038] Step 103: Generate an abnormal time series according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each second data point.

[0039] In the embodiments of the present disclosure, a third acquisition timestamp can be obtained from the second acquisition timestamps of each second data point, and an abnormal time series can be generated according to the first acquisition timestamp and the third acquisition timestamp, where the acquisition moment indicated by the third acquisition timestamp is close to the acquisition moment indicated by the first acquisition timestamp.

[0040] As an example, it is determined whether there is a third acquisition timestamp among the second acquisition timestamps of each second data point; where the difference between the acquisition moment indicated by the third acquisition timestamp and the acquisition moment indicated by the first acquisition timestamp is less than a set difference threshold; in the case where there is a third acquisition timestamp among the second acquisition timestamps of each second data point, an abnormal time series is generated according to the first acquisition timestamp and the third acquisition timestamp.

[0041] That is to say, the acquisition moment indicated by each second acquisition timestamp is compared with the acquisition moment indicated by the first acquisition timestamp respectively to obtain the difference between the acquisition moment indicated by each second acquisition timestamp and the acquisition moment indicated by the first acquisition timestamp, and the second acquisition timestamp corresponding to the difference less than the set difference threshold is used as the third acquisition timestamp. The first acquisition timestamp and the third acquisition timestamp can be sorted according to the time sequence of the first acquisition timestamp and the third acquisition timestamp to obtain an abnormal time series.

[0042] Step 104: When it is determined that there is a target data point among the second data points of at least one second service data sequence according to the abnormal time series, delete the first data point in the first service data sequence.

[0043] Wherein, the acquisition timestamp interval between any two adjacent elements in the data sequence composed of the target data point and the first data point is the same.

[0044] As an example, according to the abnormal time series, it can be determined whether there is a target data point among the second data points of at least one second service data sequence, where the target data point is used to determine that the first data point is an abnormally collected data point in a periodic manner, and the collection timestamp intervals between any two adjacent elements in the data sequence composed of the target data point and the first data point are the same, that is, the first data point is collected periodically.

[0045] In the embodiments of the present disclosure, when it is determined according to the abnormal time series that there is a target data point among the second data points of at least one second service data sequence, the first data point in the first service data sequence can be deleted.

[0046] In summary, by obtaining a first service data sequence and multiple reference service data sequences, where the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; determining the similarity between the first service data sequence and each reference service data sequence, and determining at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence; generating an abnormal time series according to the first collection timestamp of the first data point and the second collection timestamps of each second data point; when it is determined according to the abnormal time series that there is a target data point among the second data points of at least one second service data sequence, deleting the first data point in the first service data sequence; where the collection timestamp intervals between any two adjacent elements in the data sequence composed of the target data point and the first data point are the same. Thus, according to the first collection time of the abnormal first data point in the first service data sequence and the second collection times of the abnormal second data points in at least one second service data sequence similar to the first service data sequence, an abnormal time series is generated, and when it is determined according to the abnormal time series that the first data point in the first service data sequence is obtained by periodic detection, the first data point in the first service data sequence is automatically deleted, without the need for manual detection of periodically occurring abnormal detection points and manual deletion of these abnormal detection points, improving the accuracy of the detection results, while avoiding misdeletion and omission of abnormal detection points, and effectively controlling the risks of actual services.

[0047] To clearly illustrate how the similarity between the first service data sequence and each reference service data sequence is determined in the above embodiments, the present disclosure proposes another data processing method.

[0048] Figure 2 It is a schematic flowchart of another data processing method provided by the embodiments of the present disclosure.

[0049] As Figure 2 shown, the data processing method may include the following steps:

[0050] Step 201, obtain a first service data sequence and multiple reference service data sequences.

[0051] Among them, the first service data sequence includes at least one abnormal first data point, and the reference service data sequence includes at least one abnormal second data point.

[0052] Step 202, obtain a first weight sequence corresponding to the first service data sequence, and obtain a second weight sequence corresponding to any one of the reference service data sequences.

[0053] To improve the sensitivity of abnormal data points, as an example, the similarity between the first service data sequence and any one of the reference service data sequences can be determined based on the first weight sequence corresponding to the first service data sequence and the second weight sequence corresponding to any one of the reference service data sequences.

[0054] Therefore, in the embodiments of the present disclosure, a first weight sequence corresponding to the first service data sequence can be obtained, and a second weight sequence corresponding to any one of the reference service data sequences can be obtained.

[0055] As an example, for any one first element in the first service data sequence, obtain multiple target elements from the first service data sequence, where the position difference between the target element and any one first element is less than a set threshold; determine the mean value of any one first element and the multiple target elements; and determine the first weight in the first weight sequence corresponding to any one first element according to any one first element and the mean value.

[0056] For example, the first weight sequence corresponding to the first service data sequence can be obtained according to the following Algorithm 1, and the logic corresponding to Algorithm 1 is:

[0057] (1) For any one first element in the first service data sequence, obtain multiple target elements from the first service data sequence, and determine the mean value of any one first element and the multiple target elements; where the position difference between the target element and any one first element is less than a set threshold;

[0058] (2) Obtain the absolute value of the difference between the any one first element and the mean value, use the set weight base b as the base, and the quotient y of the absolute value of the difference and the mean value as the exponent, and the first weight in the first weight sequence corresponding to the any one first element is the yth power of b.

[0059] Similarly, the second weight sequences corresponding to each reference service data sequence can be obtained according to Algorithm 1.

[0060] As an example, the magnitude bases of the data in the first service data sequence and the reference service data sequence may be different. To improve the accuracy of similarity calculation, the first service data sequence and the reference service data sequence can be normalized. Based on the normalized first service data sequence, the corresponding first weight sequence can be calculated. Based on the normalized reference service data sequence, the corresponding second weight sequence can be calculated, and based on the normalized first service data sequence, the normalized reference service data sequence, the first weight sequence, and the corresponding second weight sequence, the similarity between the normalized first service data sequence and the normalized reference service data sequence can be determined.

[0061] Among them, as an example, Algorithm 2 can be used to normalize the first service data sequence and each reference service data sequence respectively. The logic corresponding to Algorithm 2 is as follows:

[0062] (1) Obtain the maximum value and the minimum value of the first service data sequence, and determine the first difference between the maximum value and the minimum value;

[0063] (2) For any first element in the first service data sequence, compare the second difference between the any first element and the minimum value with the first difference to obtain the normalized any first element; according to the normalized any first element, generate the normalized first service data sequence.

[0064] Similarly, the normalized reference service data sequences corresponding to each reference service data sequence can be obtained according to Algorithm 2.

[0065] Step 203: Update the target parameter at least once according to the first service data sequence, any reference service data sequence, the first weight sequence, and the second weight sequence.

[0066] As an example, according to the first service data sequence, any second service data sequence, the first weight sequence, and the second weight sequence, perform at least one first loop process to update the target parameter.

[0067] Among them, the i-th first loop process includes: determining whether i is equal to the first length of the first service data sequence; in the case where i is equal to the first length, ending the first loop process; in the case where i is less than the first length, obtaining the (i - 1)-th first element in the first data sequence, and according to the (i - 1)-th first element, performing at least one second loop process to update the target parameter.

[0068] Among them, the j-th second loop process in at least one second loop process includes: determining whether j is equal to the second length of any second service data sequence; ending the second loop process when j is equal to the second length; when j is less than the second length, obtaining the (j - 1)-th second element in any reference service data sequence, and obtaining the data difference between the (i - 1)-th first element and the (j - 1)-th second element; determining the target standard deviation from the first standard deviation of each first element in the first service data sequence and the second standard deviation of each second element in any reference service data sequence; when the data difference is less than the target standard deviation, updating the target parameter updated in the (j - 1)-th second loop process in the (i - 1)-th first loop process according to the (i - 1)-th third element in the first weight sequence and the (j - 1)-th fourth element in the second weight sequence, so as to obtain the target parameter updated in the j-th second loop process in the i-th first loop process; when the data difference is greater than the target standard deviation, determining the target parameter updated in the j-th second loop process in the i-th first loop process according to the target parameter updated in the j-th second loop process in the (i - 1)-th first loop process and the target parameter updated in the (j - 1)-th second loop process in the i-th first loop process.

[0069] Step 204, determine the similarity between the first service data sequence and any reference service data sequence according to the target parameter, the first weight sequence and the second weight sequence obtained by the last update.

[0070] As an example, determine the first sum value of each element in the first weight sequence and the second sum value of each element in the second weight sequence, add the first sum value and the second sum value to obtain the third sum value, and use the ratio of the target parameter obtained by the last update to the third sum value as the similarity between the first service data sequence and the reference service data sequence. For example, the target parameter obtained by the last update is dp[l1][l2], and the similarity result between the first service data sequence and the reference service data sequence is result = dp[l1][l2] / (sum(w1) + sum(w2)), where w1 is the first weight sequence and w2 is the second weight sequence.

[0071] Step 205, determine at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence.

[0072] Step 206, generate an abnormal time series according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each second data point.

[0073] Step 207, when it is determined that there is a target data point among the second data points of at least one second service data sequence according to the abnormal time series, delete the first data point in the first service data sequence.

[0074] Among them, the acquisition timestamp intervals between any two adjacent elements in the data sequence composed of the target data point and the first data point are the same.

[0075] In summary, by obtaining the first weight sequence corresponding to the first service data sequence and obtaining the second weight sequence corresponding to any reference service data sequence; according to the first service data sequence, any reference service data sequence, the first weight sequence, and the second weight sequence, updating the target parameter at least once; according to the target parameter, the first weight sequence, and the second weight sequence obtained from the last update, determining the similarity between the first service data sequence and any reference service data sequence. Thus, by using the first weight sequence corresponding to the first service data sequence and the second weight sequence corresponding to any reference service data sequence, the similarity between the first service data sequence and any reference service data sequence can be effectively determined, and the sensitivity to abnormal data points is improved.

[0076] As an example, in the case where there is a target data point among the second data points of at least one second service data sequence determined according to the abnormal time sequence, before deleting the first data point in the first service data sequence, it is possible to determine whether there is a target data point among the second data points of at least one second service data sequence according to the abnormal time sequence. The following will be described in detail in conjunction with Figure 3 the embodiments.

[0077] Figure 3 It is a schematic flowchart of another data processing method provided by an embodiment of the present disclosure.

[0078] As Figure 3 shown, the data processing method may include the following steps:

[0079] Step 301, obtain a first service data sequence and multiple reference service data sequences.

[0080] Among them, the first service data sequence includes at least one abnormal first data point, and the reference service data sequence includes at least one abnormal second data point.

[0081] Step 302, determine the similarity between the first service data sequence and each reference service data sequence, and determine at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence.

[0082] Step 303, generate an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each second data point.

[0083] Step 304, construct an n×n distance matrix according to the abnormal time sequence.

[0084] Among them, n is used to indicate the number in the abnormal time series, and the element in the i-th row and j-th column of the distance matrix is determined according to the time difference between the j-th element and the i-th element in the abnormal time series.

[0085] For example, taking the abnormal time series [1680044651, 1680044772, 1680131051, 1680217452, 1680227772, 1680303852, 1680390252, 1680476652] as an example, the element in the i-th row and j-th column of the distance matrix is determined according to the time difference between the j-th element and the i-th element in the abnormal time series, and the distance matrix can be as Figure 4 shown.

[0086] Step 305: According to the distance matrix, determine whether there is a target data point among the second data points of at least one second service data sequence.

[0087] As an example, perform path search on each element of the distance matrix to obtain at least one search path. Among them, the difference between the elements indicated by the nodes in the search path is less than the set time offset; determine whether there is a target path among the at least one search path, where the number of nodes in the target path is the largest and the target path includes a target node, and the element indicated by the target node is determined according to the time difference between the first collection timestamp and the second collection timestamp of the target data point; in the case where there is a target path among the at least one search path, determine the target data point according to the target node in the target path; in the case where there is no target path among the at least one search path, determine that there is no target data point among the second data points of at least one second service data sequence.

[0088] As an example, a path search algorithm can be used to perform path search on each element of the distance matrix to obtain at least one search path, as Figure 5 shown. Among them, the logic of the path search algorithm is: sequentially determine the difference between each element of the distance matrix and the set period. When the difference is less than the set time offset, use this node as the node in the search path to be generated, and generate a search path according to the nodes in the search path to be generated. The set period is determined according to the difference between the first collection timestamp of the first data point and each third collection timestamp before the first collection timestamp in the abnormal time series. For example, it can be one day, two days, three days, etc.

[0089] In addition, it should be noted that after determining that the first data point in the first service data is obtained by periodic collection, the proportion of the third service data sequence that forms periodic collection with the first service data sequence indicated by the first service data sequence and the abnormal time sequence in the total data sequence can also be determined, and based on this proportion (for example, the proportion is greater than a set threshold), it can be determined whether the total data sequence is a periodic sequence.

[0090] Step 306, when it is determined that there is a target data point among the second data points of at least one second service data sequence, delete the first data point in the first service data sequence.

[0091] Wherein, the collection timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same.

[0092] In summary, by constructing an n×n distance matrix according to the abnormal time sequence, and judging whether there is a target data point among the second data points of at least one second service data sequence, wherein the collection timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same, thus, it can be effectively determined whether the first data point is obtained by periodic collection.

[0093] To implement the above embodiments, the present disclosure proposes a data processing device.

[0094] Figure 6 It is a schematic structural diagram of a data processing device provided by an embodiment of the present disclosure.

[0095] Such as Figure 6 As shown, the data processing device 600 includes: an acquisition module 610, a determination module 620, a generation module 630, and a deletion module 640.

[0096] Among them, an acquisition module 610 is configured to acquire a first service data sequence and a plurality of reference service data sequences, where the first service data sequence includes at least one first data point of an anomaly, and the reference service data sequences include at least one second data point of an anomaly; a determination module 620 is configured to determine the similarity between the first service data sequence and each reference service data sequence, and determine at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence; a generation module 630 is configured to generate an anomaly time series according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each second data point; a deletion module 640 is configured to delete the first data point in the first service data sequence when it is determined that there is a target data point among the second data points of at least one second service data sequence according to the anomaly time series; where the acquisition timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same.

[0097] As a possible implementation manner of an embodiment of the present disclosure, the determination module 620 is configured to acquire a first weight sequence corresponding to the first service data sequence, and acquire a second weight sequence corresponding to any reference service data sequence; update the target parameter at least once according to the first service data sequence, any reference service data sequence, the first weight sequence, and the second weight sequence; determine the similarity between the first service data sequence and any reference service data sequence according to the target parameter, the first weight sequence, and the second weight sequence obtained by the last update.

[0098] As a possible implementation manner of an embodiment of the present disclosure, the determination module 620 is further configured to, for any first element in the first service data sequence, acquire a plurality of target elements from the first service data sequence, where the position difference between the target element and any first element is less than a set threshold; determine the mean value of any first element and the plurality of target elements; determine the first weight in the first weight sequence corresponding to any first element according to any first element and the mean value.

[0099] As a possible implementation manner of an embodiment of the present disclosure, the determination module 620 is further configured to perform at least one first loop process according to the first service data sequence, any reference service data sequence, the first weight sequence, and the second weight sequence to update the target parameter; where the i-th first loop process includes: determining whether i is equal to the first length of the first service data sequence; ending the first loop process when i is equal to the first length; and when i is less than the first length, acquiring the (i - 1)-th first element in the first data sequence, and performing at least one second loop process according to the (i - 1)-th first element to update the target parameter.

[0100] As a possible implementation manner of an embodiment of the present disclosure, the j-th second loop process in at least one second loop process includes: determining whether j is equal to the second length of any reference service data sequence; ending the second loop process when j is equal to the second length; when j is less than the second length, obtaining the (j - 1)-th second element in any reference service data sequence, and obtaining the absolute value of the data difference between the (i - 1)-th first element and the (j - 1)-th second element; determining a target standard deviation from the first standard deviation of each first element in the first service data sequence and the second standard deviation of each second element in any reference service data sequence; when the absolute value of the data difference is less than the target standard deviation, updating the target parameter updated in the (j - 1)-th second loop process in the (i - 1)-th first loop process according to the (i - 1)-th third element in the first weight sequence and the (j - 1)-th fourth element in the second weight sequence, so as to obtain the target parameter updated in the j-th second loop process in the i-th first loop process; when the absolute value of the data difference is greater than the target standard deviation, determining the target parameter updated in the j-th second loop process in the i-th first loop process according to the target parameter updated in the j-th second loop process in the (i - 1)-th first loop process and the target parameter updated in the (j - 1)-th second loop process in the i-th first loop process.

[0101] As a possible implementation manner of an embodiment of the present disclosure, the generating module 630 is configured to determine whether there is a third acquisition timestamp in the second acquisition timestamps of each second data point, where the difference between the acquisition moment indicated by the third acquisition timestamp and the acquisition moment indicated by the first acquisition timestamp is less than a set difference threshold; when there is a third acquisition timestamp in the second acquisition timestamps of each second data point, generating an abnormal time series according to the first acquisition timestamp and the third acquisition timestamp.

[0102] As a possible implementation manner of an embodiment of the present disclosure, the data processing device 600 further includes: a constructing module and a judging module.

[0103] Wherein, the constructing module is configured to construct an n×n distance matrix according to the abnormal time series, where n is used to indicate the number in the abnormal time series, and the element in the i-th row and j-th column of the distance matrix is determined according to the time difference between the j-th element and the i-th element in the abnormal time series; the judging module is configured to judge whether there is a target data point in the second data points of at least one second service data sequence according to the distance matrix.

[0104] As a possible implementation of the embodiment of the present disclosure, a determination module is configured to perform path search on each element of the distance matrix to obtain at least one search path, where the difference between the elements indicated by the nodes in the search path is less than a set time offset; determine whether there is a target path from the at least one search path, where the target path has the largest number of nodes and the target path includes a target node, and the element indicated by the target node is determined according to the time difference between the first acquisition timestamp and the second acquisition timestamp of the target data point; in the case that there is a target path in the at least one search path, determine the target data point according to the target node in the target path; in the case that there is no target path in the at least one search path, determine that there is no target data point among the second data points of at least one second service data sequence.

[0105] The data processing device according to the embodiment of the present disclosure obtains a first service data sequence and a plurality of reference service data sequences, where the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; determines the similarity between the first service data sequence and each reference service data sequence, and determines at least one second service data sequence from each reference service data sequence according to the similarity of each reference service data sequence; generates an abnormal time series according to the first acquisition timestamp of the first data point and the second acquisition timestamp of each second data point; deletes the first data point in the first service data sequence in the case that it is determined that there is a target data point among the second data points of at least one second service data sequence according to the abnormal time series; where the acquisition timestamp interval between any two adjacent elements in the data sequence formed by the target data point and the first data point is the same. Thus, according to the first acquisition moment of the abnormal first data point in the first service data sequence and the second acquisition moment of the abnormal second data point in at least one second service data sequence similar to the first service data sequence, an abnormal time series is generated, and in the case that it is determined that the first data point in the first service data sequence is obtained by periodic detection according to the abnormal time series, the first data point in the first service data sequence is automatically deleted, without manual detection of periodically occurring abnormal detection points and manual deletion of these abnormal detection points, improving the accuracy of the detection result, while avoiding misdeletion and omission of abnormal detection points, and effectively controlling the risks of the actual service.

[0106] It should be noted that the foregoing explanation of the data processing method embodiment also applies to the data processing device of this embodiment, and will not be repeated here.

[0107] To implement the above embodiment, the present application also proposes an electronic device, as Figure 7 shown Figure 7It is a block diagram of an electronic device for data processing shown according to an exemplary embodiment.

[0108] As Figure 7 shown, the above-mentioned electronic device 700 includes:

[0109] A memory 710 and a processor 720, a bus 730 connecting different components (including the memory 710 and the processor 720), the memory 710 stores a computer program, and when the processor 720 executes the program, it implements the data processing method described in the embodiments of the present disclosure.

[0110] The bus 730 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0111] The electronic device 700 typically includes a variety of electronic device-readable media. These media can be any available media accessible by the electronic device 700, including volatile and non-volatile media, removable and non-removable media.

[0112] The memory 710 may further include a computer system-readable medium in the form of volatile memory, such as random access memory (RAM) 740 and / or cache memory 750. The electronic device 700 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 760 can be used to read and write non-removable, non-volatile magnetic media ( Figure 7 not shown, commonly referred to as a "hard disk drive"). Although Figure 7 not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to the bus 730 through one or more data media interfaces. The memory 710 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0113] A program / utilities 780 having a set (at least one) of program modules 770 can be stored, for example, in the memory 710. Such program modules 770 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 770 generally execute the functions and / or methods in the embodiments described in this disclosure.

[0114] The electronic device 700 can also communicate with one or more external devices 790 (such as a keyboard, a pointing device, a display, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 792. Moreover, the electronic device 700 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 793. As Figure 7 shown, the network adapter 793 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that although Figure 7 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0115] The processor 720 executes various functional applications and data processing by running the programs stored in the memory 710.

[0116] It should be noted that for the implementation process and technical principle of the electronic device in this embodiment, refer to the foregoing explanation of the data processing method of the embodiments of this disclosure, and details are not described herein again.

[0117] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data processing method described in the above embodiments is implemented.

[0118] To implement the above embodiments, the present disclosure also provides a computer program product. When the instructions in the computer program product are executed by a processor, the data processing method described in the above embodiments is executed.

[0119] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0120] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0121] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A data processing method, characterized in that, comprising: obtaining a first service data sequence and a plurality of reference service data sequences, wherein the first service data sequence includes at least one abnormal first data point, and the reference service data sequences include at least one abnormal second data point; determining the similarity between the first service data sequence and each of the reference service data sequences, and determining at least one second service data sequence from each of the reference service data sequences according to the similarity of each of the reference service data sequences; generating an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each of the second data points; when it is determined that there is a target data point among the second data points of the at least one second service data sequence according to the abnormal time sequence, deleting the first data point in the first service data sequence; wherein the acquisition timestamp interval between any two adjacent elements in the data sequence composed of the target data point and the first data point is the same.

2. The method according to claim 1, characterized in that, the determining the similarity between the first service data sequence and each of the reference service data sequences includes: obtaining a first weight sequence corresponding to the first service data sequence, and obtaining a second weight sequence corresponding to any one of the reference service data sequences; updating a target parameter at least once according to the first service data sequence, the any one of the reference service data sequences, the first weight sequence and the second weight sequence; determining the similarity between the first service data sequence and the any one of the reference service data sequences according to the target parameter obtained by the last update, the first weight sequence and the second weight sequence.

3. The method according to claim 2, characterized in that, the obtaining the first weight sequence corresponding to the first service data sequence includes: for any one first element in the first service data sequence, obtaining a plurality of target elements from the first service data sequence, wherein the position difference between the target element and the any one first element is less than a set threshold; determining the mean value of the any one first element and the plurality of target elements; determining the first weight in the first weight sequence corresponding to the any one first element according to the any one first element and the mean value.

4. The method according to claim 2, characterized in that, the updating the target parameter at least once according to the first service data sequence, the any one of the reference service data sequences, the first weight sequence and the second weight sequence includes: performing at least one first loop process according to the first service data sequence, the any one of the reference service data sequences, the first weight sequence and the second weight sequence to update the target parameter; wherein the i-th first loop process includes: judging whether i is equal to the first length of the first service data sequence; when i is equal to the first length, ending the first loop process; When i is less than the first length, obtain the (i - 1)-th first element in the first data sequence, and according to the (i - 1)-th first element, perform at least one second loop process to update the target parameter.

5. The method according to claim 4, wherein, the j-th second loop process in the at least one second loop process includes: judging whether j is equal to the second length of any of the reference service data sequences; when j is equal to the second length, end the second loop process; when j is less than the second length, obtain the (j - 1)-th second element in any of the reference service data sequences, and obtain the absolute value of the data difference between the (i - 1)-th first element and the (j - 1)-th second element; determine the target standard deviation from the first standard deviation of each first element in the first service data sequence and the second standard deviation of each second element in any of the reference service data sequences; when the absolute value of the data difference is less than the target standard deviation, update the target parameter updated in the (j - 1)-th second loop process in the (i - 1)-th first loop process according to the (i - 1)-th third element in the first weight sequence and the (j - 1)-th fourth element in the second weight sequence, so as to obtain the target parameter updated in the j-th second loop process in the i-th first loop process; when the absolute value of the data difference is greater than the target standard deviation, determine the target parameter updated in the j-th second loop process in the i-th first loop process according to the target parameter updated in the j-th second loop process in the (i - 1)-th first loop process and the target parameter updated in the (j - 1)-th second loop process in the i-th first loop process.

6. The method according to claim 1, wherein, the generating the abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of the respective second data points includes: judging whether there is a third acquisition timestamp among the second acquisition timestamps of the respective second data points, wherein the difference between the acquisition moment indicated by the third acquisition timestamp and the acquisition moment indicated by the first acquisition timestamp is less than a set difference threshold; when there is a third acquisition timestamp among the second acquisition timestamps of the respective second data points, generate an abnormal time sequence according to the first acquisition timestamp and the third acquisition timestamp.

7. The method according to claim 1, wherein, before deleting the first data point in the first service data sequence when it is determined that there is a target data point among the second data points of the at least one second service data sequence according to the abnormal time sequence, the method further includes: construct an n×n distance matrix according to the abnormal time sequence, where n is used to indicate the number in the abnormal time sequence, and the element in the i-th row and j-th column of the distance matrix is determined according to the time difference between the j-th element and the i-th element in the abnormal time sequence; Based on the distance matrix, determine whether there is a target data point among the second data points of the at least one second service data sequence.

8. The method according to claim 7, wherein, the determining whether there is a target data point among the second data points of the at least one second service data sequence based on the distance matrix includes: Performing path search on each element of the distance matrix to obtain at least one search path, wherein the difference between the elements indicated by the nodes in the search path is less than a set time offset; Determining whether there is a target path from the at least one search path, wherein the number of nodes in the target path is the largest and the target path includes a target node, and the element indicated by the target node is determined according to the time difference between the first acquisition timestamp and the second acquisition timestamp of the target data point; In the case where there is a target path in the at least one search path, determining the target data point according to the target node in the target path; In the case where there is no target path in the at least one search path, determining that there is no target data point among the second data points of the at least one second service data sequence.

9. A data processing device, wherein, comprising: An acquisition module for acquiring a first service data sequence and a plurality of reference service data sequences, wherein the first service data sequence includes at least one first data point that is abnormal, and the reference service data sequences include at least one second data point that is abnormal; A determination module for determining the similarity between the first service data sequence and each of the reference service data sequences, and determining at least one second service data sequence from each of the reference service data sequences according to the similarity of each of the reference service data sequences; A generation module for generating an abnormal time sequence according to the first acquisition timestamp of the first data point and the second acquisition timestamps of each of the second data points; A deletion module for deleting the first data point in the first service data sequence when it is determined that there is a target data point among the second data points of the at least one second service data sequence according to the abnormal time sequence; wherein the acquisition timestamp interval between any two adjacent elements in the data sequence composed of the target data point and the first data point is the same.

10. An electronic device, wherein, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, implementing the data processing method according to any one of claims 1-8.

11. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, implementing the data processing method according to any one of claims 1-8.