Job Abnormality Detection Method and Device for Intelligent Computing Power Network

By splitting and abnormal verification of the original running data of jobs in the intelligent computing network, the problem of misjudgment of job abnormal detection in the prior art is solved, and the accuracy of detection is improved.

CN119416134BActive Publication Date: 2025-06-17SUGON INFORMATION IND
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510026682.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-06-17
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

When detecting job abnormalities in intelligent computing power networks, the prior art is prone to misjudgment due to single type, which reduces the accuracy of the detection.

Method used

By splitting the original running data of the target job in the intelligent computing network, the local running data of each sub-period period is obtained, and the abnormal sub-period period is determined based on the similarity of adjacent sub-period periods, and abnormal verification is performed based on the historical running data to determine the abnormal detection result of the job.

Benefits of technology

Improve the accuracy of job abnormal detection and avoid misjudging reasonable data fluctuations as job abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416134B_ABST
    Figure CN119416134B_ABST
Patent Text Reader

Abstract

The present application relates to a method and apparatus for detecting job anomalies in an intelligent computing power network. The method includes: splitting the original operation data of a target job in an intelligent computing power network during a to-be-detected period to obtain local operation data corresponding to each sub-period within the to-be-detected period, determining abnormal sub-periods from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods, obtaining historical operation data within a historical detection period corresponding to the abnormal sub-periods, and determining an anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-periods. Using this method can improve the accuracy of job anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for detecting job anomalies in an intelligent computing power network. Background Art

[0002] With the expansion of the computing power scale and the wide application of intelligent computing power networks, in the process of job processing in intelligent computing power networks, in order to avoid waste of resources caused by job failures, it is necessary to detect anomalies in the operation of jobs.

[0003] Existing methods for detecting job anomalies generally directly apply the anomaly detection methods for jobs in ordinary networks to jobs in intelligent computing power networks. For example, a method of using a preset threshold is used to monitor certain types of operation data.

[0004] However, since the types of jobs in ordinary networks are relatively single, directly applying the anomaly detection methods for jobs in ordinary networks to jobs in intelligent computing power networks is prone to misjudgment, thereby reducing the accuracy of job anomaly detection. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method and device for detecting job anomalies in an intelligent computing power network that can improve the accuracy of job anomaly detection.

[0006] In a first aspect, this application provides a method for detecting job anomalies in an intelligent computing power network, including:

[0007] Split the original operation data of a target job in an intelligent computing power network during a to-be-detected period to obtain local operation data corresponding to each sub-period within the to-be-detected period;

[0008] Determine an abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0009] Obtain the historical operation data within the historical detection period corresponding to the abnormal sub-period;

[0010] Determine the anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0011] In the embodiments of this application, on the one hand, by determining the abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods, the rationality of determining the abnormal sub-period can be ensured; on the other hand, by further performing anomaly verification on the local operation data corresponding to the abnormal sub-period based on the historical operation data, misjudging reasonable data fluctuations as job anomalies can be avoided, thereby improving the accuracy of the anomaly detection result.

[0012] In one embodiment, the original operation data of a target job in the intelligent computing power network during a period to be detected is split to obtain local operation data corresponding to each sub-period within the period to be detected, including:

[0013] Perform first-order difference processing on the original operation data of the target job in the intelligent computing power network during the period to be detected to obtain the global operation data of the target job during the period to be detected; split the global operation data to obtain local operation data corresponding to each sub-period within the period to be detected.

[0014] In the embodiment of the present application, by performing first-order difference processing on the original operation data of the target job, the global operation data after removing the change trend is obtained, and then the global operation data is split to obtain the local operation data of each, which can make the local operation data more stable, thereby improving the accuracy of anomaly detection.

[0015] In one embodiment, according to the similarity between the local operation data corresponding to two adjacent sub-periods, an abnormal sub-period is determined from each sub-period, including:

[0016] For each sub-period, determine the similarity score of the sub-period according to the similarity between the local operation data corresponding to the sub-period and the local operation data corresponding to the neighbor sub-period; wherein, the neighbor sub-period is the sub-period adjacent to the sub-period among each sub-period, and the neighbor sub-period is before or after the sub-period; determine the abnormal sub-period from each sub-period according to the similarity scores of each sub-period.

[0017] In the embodiment of the present application, by determining the similarity score according to the similarity between adjacent local operation data and determining the abnormal sub-period from each sub-period according to the similarity scores of each sub-period, the rationality of determining the abnormal sub-period can be ensured.

[0018] In one embodiment, determining the similarity score of the sub-period according to the similarity between the local operation data corresponding to the sub-period and the local operation data corresponding to the neighbor sub-period includes:

[0019] Determine the data change trend of the local operation data corresponding to the sub-period and the data change trend of the local operation data corresponding to the neighbor sub-period; determine the similarity score of the sub-period according to the trend similarity between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighbor sub-period.

[0020] In the embodiment of the present application, by determining the similarity score of the sub-period according to the trend similarity between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighbor sub-period, the accuracy of determining the similarity score can be ensured.

[0021] In one embodiment, determining an abnormal sub-period from each sub-period according to the similarity scores of each sub-period includes:

[0022] Selecting an outlier similarity score from the similarity scores of each sub-period; taking the sub-period corresponding to the outlier similarity score as the abnormal sub-period.

[0023] In the embodiment of the present application, by selecting an outlier similarity score that deviates from the whole from the similarity scores of each sub-period and directly taking the sub-period corresponding to the outlier similarity score as the abnormal sub-period, the rationality of determining the abnormal sub-period can be ensured.

[0024] In one embodiment, obtaining historical operation data within a historical detection period corresponding to the abnormal sub-period includes:

[0025] Determining a historical detection period corresponding to the abnormal sub-period according to the time characteristics corresponding to the abnormal sub-period; selecting historical operation data that conforms to the job characteristics of the target job and is in a normal job state from each candidate operation data within the historical detection period.

[0026] In the embodiment of the present application, by determining the historical detection period based on the time feature dimension and screening out the historical operation data in the normal job state based on the job characteristics of the target job, the rationality and accuracy of determining the historical operation data can be ensured.

[0027] In one embodiment, determining an abnormal detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-period includes:

[0028] Determining the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period; determining the abnormal detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0029] In the embodiment of the present application, by determining the abnormal detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period, the accuracy of determining the abnormal detection result can be ensured.

[0030] In a second aspect, the present application further provides a job anomaly detection device for an intelligent computing power network, including:

[0031] A data splitting module, configured to split the original operation data of a target job in a to-be-detected period in the intelligent computing power network to obtain local operation data corresponding to each sub-period in the to-be-detected period;

[0032] An anomaly determination module, configured to determine an anomalous sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0033] A data acquisition module, configured to acquire historical operation data within a historical detection period corresponding to the anomalous sub-period;

[0034] An anomaly detection module, configured to determine an anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the anomalous sub-period.

[0035] In a third aspect, the present application further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0036] Split the original operation data of the target job in the to-be-detected period in the intelligent computing power network to obtain local operation data corresponding to each sub-period in the to-be-detected period;

[0037] Determine an anomalous sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0038] Acquire historical operation data within a historical detection period corresponding to the anomalous sub-period;

[0039] Determine an anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the anomalous sub-period.

[0040] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0041] Split the original operation data of the target job in the to-be-detected period in the intelligent computing power network to obtain local operation data corresponding to each sub-period in the to-be-detected period;

[0042] Determine an anomalous sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0043] Acquire historical operation data within a historical detection period corresponding to the anomalous sub-period;

[0044] Determine an anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the anomalous sub-period.

[0045] In a fifth aspect, the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0046] Split the original operation data of the target job in the intelligent computing power network during the period to be detected to obtain the local operation data corresponding to each sub-period within the period to be detected;

[0047] Determine the abnormal sub-periods from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0048] Obtain the historical operation data within the historical detection period corresponding to the abnormal sub-periods;

[0049] Determine the abnormal detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-periods.

[0050] The above method and device for job abnormal detection in the intelligent computing power network split the original operation data of the target job in the intelligent computing power network during the period to be detected to obtain the local operation data corresponding to each sub-period within the period to be detected, and determine the abnormal sub-periods from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods; subsequently, determine the abnormal detection result of the target job according to the historical operation data within the historical detection period corresponding to the abnormal sub-periods and the local operation data corresponding to the abnormal sub-periods. Compared with the related technology that only monitors the job operation status using a preset data threshold, using the above method, on the one hand, by determining the abnormal sub-periods from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods, the rationality of determining the abnormal sub-periods can be ensured; on the other hand, by further performing abnormal verification on the local operation data corresponding to the abnormal sub-periods based on the historical operation data, the misjudgment of reasonable data fluctuations as job abnormalities can be avoided, thereby improving the accuracy of the abnormal detection result. Description of the Drawings

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic flowchart of a method for job abnormal detection in the intelligent computing power network in an embodiment;

[0053] Figure 2 It is a schematic diagram of job abnormal detection in an embodiment;

[0054] Figure 3 It is a schematic flowchart of determining abnormal sub-periods in an embodiment;

[0055] Figure 4 Schematic flowchart of determining similarity score in one embodiment;

[0056] Figure 5 Schematic flowchart of determining similarity score in another embodiment;

[0057] Figure 6 Schematic flowchart of obtaining historical operation data in one embodiment;

[0058] Figure 7 Schematic flowchart of determining anomaly detection result in one embodiment;

[0059] Figure 8 Schematic flowchart of job anomaly detection method for intelligent computing power network in another embodiment;

[0060] Figure 9 Structural block diagram of job anomaly detection device for intelligent computing power network in one embodiment;

[0061] Figure 10 Internal structure diagram of computer device in one embodiment. Detailed implementation manners

[0062] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] With the expansion of the computing power scale and the wide application of intelligent computing power networks, in the process of job processing in intelligent computing power networks, in order to avoid waste of resources caused by job failures, it is necessary to perform anomaly detection on the operation of jobs.

[0064] In the existing job anomaly detection methods, generally, the anomaly detection methods for jobs in ordinary networks are directly applied to jobs in intelligent computing power networks. For example, a preset threshold method is used to monitor a certain type of operation data.

[0065] However, since the types of jobs in ordinary networks are relatively single, directly applying the anomaly detection methods for jobs in ordinary networks to jobs in intelligent computing power networks is likely to cause misjudgment, thereby reducing the accuracy of job anomaly detection.

[0066] Based on this, in an exemplary embodiment, as Figure 1 shown, a job anomaly detection method for intelligent computing power network is provided. Taking the application of this method to an anomaly detection device deployed in an intelligent computing power network as an example, the method specifically includes the following steps:

[0067] S101, split the original operation data of the target job in the to-be-detected period in the intelligent computing power network to obtain the local operation data corresponding to each sub-period in the to-be-detected period.

[0068] Among them, the target job is a job with abnormal detection requirements in the intelligent computing power network; the to-be-detected period is the period for which abnormal detection is required; the original operation data is the data generated when the target job runs in the to-be-detected period; the so-called local operation data is the operation data within each sub-period; the duration of each sub-period is the same.

[0069] In the embodiments of the present application, the user can submit each job to the intelligent computing power network through the user interface displayed on the terminal to run each job in the intelligent computing power network.

[0070] At the same time, in order to be able to more accurately monitor the operation status of each job running in the intelligent computing power network, in an optional implementation manner, the user can configure an abnormal detection policy for the job through the administrator interface displayed on the terminal.

[0071] Exemplarily, referring to Figure 2 the job abnormal detection schematic diagram shown, the user can modify each modifiable option in the monitoring policy displayed on the administrator page to configure an abnormal detection policy for the job. Among them, the modifiable options include but are not limited to the time window size (i.e., the time window for abnormal detection), the number of acquisitions (i.e., the number of data acquisitions within the time window), the abnormal detection interval (i.e., the to-be-detected interval can be customized), the data downsampling method (since the time series data has a high dimension and a large amount of data, therefore, a specific data downsampling method can be set to reduce the amount of data), and the data source (i.e., the job type of the target job can be customized), etc.

[0072] After configuring the abnormal detection policy, when it is determined that the to-be-detected period is a historical period, the original operation data of the target job in the to-be-detected period can be directly obtained based on the data source and the abnormal detection interval in the abnormal detection policy; subsequently, based on the time window size and the number of acquisitions, the original operation data is split to obtain the local operation data corresponding to each sub-period in the to-be-detected period.

[0073] When it is determined that the to-be-detected period is the current period, the local operation data corresponding to each sub-period can be collected in real time according to the data source, the time window size, and the number of acquisitions.

[0074] In another optional implementation manner, the user can directly send a job detection request including the job identifier of the target job and the to-be-detected period to the abnormal detection device through the request initiation interface displayed on the terminal.

[0075] After detecting a job detection request, the original operation data of the target job during the to-be-detected period can be obtained based on the job identifier; subsequently, based on a preset splitting duration, the original operation data during the to-be-detected period can be split into local operation data corresponding to each sub-period. Among them, the preset splitting duration can be determined by technicians based on historical experience or through a large number of experimental results, and this application does not limit it.

[0076] S102. Determine an abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods.

[0077] Among them, the abnormal sub-period is used to represent the sub-period in which the operation of the target job may be abnormal.

[0078] It can be understood that since job anomalies usually show long-term trends, therefore, according to the similarity between the local operation data corresponding to each sub-period, sub-periods with long-term different data distributions can be selected from each sub-interval as abnormal sub-periods.

[0079] In an alternative embodiment, the local operation data corresponding to each sub-period can be sequentially input into a trained first anomaly determination model in chronological order. The first anomaly determination model sequentially determines an abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods in time sequence and outputs it.

[0080] In another alternative embodiment, since the target job in the normal operation state should have similar operation characteristics in each sub-period, therefore, for each sub-period, by analyzing whether there are similar data characteristics between the local operation data corresponding to this sub-period and the local operation data corresponding to the neighboring sub-period in time sequence, it can be determined whether this sub-period is an abnormal sub-period.

[0081] For example, if there are no similar data characteristics between the local operation data A' corresponding to sub-period A and the local operation data B' corresponding to sub-period B, it proves that there may be an abnormal sub-period in sub-period A and sub-period B.

[0082] S103. Obtain the historical operation data during the historical detection period corresponding to the abnormal sub-period.

[0083] Among them, the so-called historical detection period is used to represent the period in the historical period with the same period characteristics as the abnormal sub-period. For example, if the abnormal sub-period is from November 15, 2023 to November 30, 2023, the historical detection period can be from November 15, 2022 to November 30, 2022. The so-called historical operation data is the actual operation data of the job during the historical detection period.

[0084] It can be understood that, due to the diverse types of tasks executed on the intelligent computing power network, each following a unique operating mode, and considering that these tasks usually run for a long time, it is common for the task data to change sharply (suddenly increase or decrease) in the short term.

[0085] Based on this, in order to avoid misjudging the above situation as a task anomaly, in an optional embodiment, according to the time period characteristics of the abnormal sub-time period, historical operation data within the historical detection time period with the same time period characteristics as the abnormal sub-time period can be obtained from the database constructed based on the historical operation data of each task, so as to determine whether the abnormal fluctuation of the data within the abnormal sub-time period is a normal data fluctuation.

[0086] Exemplarily, if the target task has been in a running state within the historical detection time period, the real operation data of the target task within the historical detection time period is preferentially obtained as the historical operation data; if the target task is not in a running state within the historical detection time period, the real operation data of other tasks with the same task type as the target task within the historical detection time period can be obtained as the historical operation data.

[0087] It can be understood that, in order to ensure the reliability of data collection, only the data under the normal running state of the task can be collected as the positive reference; or, only the data under the abnormal running state of the task can be collected as the negative reference.

[0088] S104. Determine the anomaly detection result of the target task according to the historical operation data and the local operation data corresponding to the abnormal sub-time period.

[0089] Among them, the so-called anomaly detection result is the result of whether there is an anomaly in the target task within the time period to be detected. The anomaly detection result can be that the task runs abnormally or the task runs normally.

[0090] After obtaining the historical running tasks, the historical operation data can be used as a reference to determine whether there is an anomaly in the local operation data corresponding to the abnormal sub-time period, and then determine the anomaly detection result of the target task.

[0091] In an optional implementation manner, the historical operation data, the task running state corresponding to the historical operation data, and the local operation data corresponding to the abnormal sub-time period can be input into the trained anomaly judgment model at the same time. The anomaly judgment model outputs the anomaly detection result according to the historical operation data, the task running state, the local operation data, and the model parameters.

[0092] In another alternative implementation, the data fluctuation condition corresponding to the historical operation data can be compared with the data fluctuation condition of the local operation data corresponding to the abnormal sub-period. If the comparison result shows a high degree of consistency, it is determined that the operation state of the target job is consistent with the operation state of the job corresponding to the historical operation data; if the comparison result shows a low degree of consistency, it is determined that the operation state of the target job is opposite to the operation state of the job corresponding to the historical operation data.

[0093] For example, when the operation state of the historical operation data is that the job is running normally, if the data fluctuation condition of the historical operation data is consistent with the data fluctuation condition of the local operation data corresponding to the abnormal sub-period, it is determined that the abnormal detection result is that the job is running normally; if the data fluctuation condition of the historical operation data is inconsistent with the data fluctuation condition of the local operation data corresponding to the abnormal sub-period, it is determined that the abnormal detection result is that the job is running abnormally.

[0094] In the case where it is determined that the operation of the target job is abnormal, continue to refer to Figure 2 As shown in the job abnormal detection schematic diagram, job abnormal alarm information can be generated according to the period to be detected and the job identifier of the target job, and the abnormal alarm information can be fed back to the management end to prompt the management personnel to perform abnormal processing on the target job. At the same time, the local operation data corresponding to the abnormal sub-period can be input into the database as job abnormal data, which is convenient for subsequent reference data when the job is in an abnormal operation state.

[0095] In the above job abnormal detection method for the intelligent computing power network, by splitting the original operation data of the target job in the period to be detected in the intelligent computing power network, the local operation data corresponding to each sub-period in the period to be detected is obtained, and the abnormal sub-period is determined from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods; subsequently, according to the historical operation data in the historical detection period corresponding to the abnormal sub-period and the local operation data corresponding to the abnormal sub-period, the abnormal detection result of the target job is determined. Compared with the related technology where only a preset data threshold is used to monitor the operation state of the job, by using the above method, on the one hand, by determining the abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods, the rationality of determining the abnormal sub-period can be ensured; on the other hand, by further performing abnormal verification on the local operation data corresponding to the abnormal sub-period based on the historical operation data, the misjudgment of reasonable data fluctuations as job abnormalities can be avoided, thereby improving the accuracy of the abnormal detection result.

[0096] To ensure the accuracy of anomaly detection, based on the above embodiments, in an embodiment of the present application, an optional method for data processing is provided. Specifically, first-order difference processing is performed on the original operation data of the target job in the period to be detected in the intelligent computing power network to obtain the global operation data of the target job in the period to be detected; the global operation data is split to obtain the local operation data corresponding to each sub-period in the period to be detected.

[0097] Among them, the so-called global operation data is used to represent the overall operation data of the target job in the period to be detected.

[0098] To make the operation data more stable, after obtaining the original operation data of the target job in the period to be detected, first-order difference processing can be directly performed on the original operation data, and the processed original operation data is used as the global operation data to remove the long-term trend in the original operation data.

[0099] In an optional implementation manner, according to the preset time window size, the global operation data can be split in the time dimension in chronological order, and the split result is used as the local operation data corresponding to each sub-period in the period to be detected. At this time, there is no overlap between the sub-periods.

[0100] For example, if the period to be detected is 4:00 - 7:00 and the time window size is 1 hour, sub-periods 4:00 - 5:00, 5:00 - 6:00, and 6:00 - 7:00 can be obtained.

[0101] In another optional implementation manner, when the duration of the period to be detected is short, to increase the number of sub-periods, the global operation data in the period to be detected can be split in the time dimension based on the time window size and the window sliding distance to obtain the local operation data corresponding to each sub-period. At this time, there is overlap between the sub-periods.

[0102] For example, if the period to be detected is 4:00 - 7:00, the time window size is 1 hour, and the window sliding distance is 30 minutes, sub-periods 4:00 - 5:00, 4:30 - 5:30, 5:00 - 6:00, 5:30 - 6:30, and 6:00 - 7:00 can be obtained.

[0103] In the embodiment of the present application, by performing first-order difference processing on the original operation data of the target job, the global operation data after removing the change trend is obtained, and then the global operation data is split to obtain the local operation data of each part, which can make the local operation data more stable, thereby improving the accuracy of anomaly detection.

[0104] To ensure the accuracy of abnormal sub - time periods, based on the above - mentioned embodiments, in the embodiments of the present application, an optional method for determining abnormal sub - time periods is provided. As Figure 3 shown, it specifically includes the following steps:

[0105] S301. For each sub - time period, determine the similarity score of the sub - time period according to the similarity between the local operation data corresponding to the sub - time period and the local operation data corresponding to the neighboring sub - time period.

[0106] Among them, the neighboring sub - time period is the sub - time period adjacent to the sub - time period among all sub - time periods, and the neighboring sub - time period is before or after the sub - time period. It can be understood that in order to ensure the rationality of similarity evaluation, the neighboring sub - time periods of each sub - time period need to be defined as sub - time periods in the same direction.

[0107] For example, for sub - time period 1, sub - time period 2, and sub - time period 3 in chronological order, when defining that the neighboring sub - time period is before the sub - time period, the neighboring sub - time period of sub - time period 2 is sub - time period 1, and the neighboring sub - time period of sub - time period 3 is sub - time period 2; when defining that the neighboring sub - time period is after the sub - time period, the neighboring sub - time period of sub - time period 1 is sub - time period 2, and the neighboring sub - time period of sub - time period 2 is sub - time period 3.

[0108] In the case where there may be no neighboring sub - time periods for the first sub - time period or the last sub - time period, since the local operation data of each sub - time period has removed the data trend, the first sub - time period and the last sub - time period can be used as neighboring sub - time periods for each other.

[0109] For example, for sub - time period 1, sub - time period 2, and sub - time period 3 in chronological order, when defining that the neighboring sub - time period is before the sub - time period, the neighboring sub - time period of sub - time period 1 is sub - time period 3; when defining that the neighboring sub - time period is after the sub - time period, the neighboring sub - time period of sub - time period 3 is sub - time period 1.

[0110] In an optional implementation manner, for each sub - time period, the local operation data corresponding to the sub - time period and the local operation data corresponding to the neighboring sub - time period of the sub - time period can be directly input into the first similarity determination model constructed based on the similarity algorithm. The first similarity determination model outputs the similarity score of the sub - time period according to the local operation data and model parameters within the two sub - time periods. Among them, the similarity algorithm can be the Fast Dynamic Time Warping (FastDTW) algorithm.

[0111] In another alternative embodiment, for each sub-period, the first data feature of the local operation data corresponding to the sub-period and the second data feature of the local operation data corresponding to the neighboring sub-period of the sub-period can be determined respectively; subsequently, the similarity value between the first data feature and the second data feature is used as the similarity score of the sub-period.

[0112] S302. Determine abnormal sub-periods from each sub-period according to the similarity scores of each sub-period.

[0113] In an alternative embodiment, a standard similarity score can be generated based on the similarity scores of each sub-period. For example, the average value of the similarity scores of each sub-period can be used as the standard similarity score.

[0114] Subsequently, for each sub-period, the similarity score of the sub-period can be compared with the standard similarity score. If the difference between the similarity score of the sub-period and the standard similarity score is less than the preset score threshold, it is determined that the sub-period is a normal sub-period; if the difference between the similarity score of the sub-period and the standard similarity score is greater than or equal to the preset score threshold, it is determined that the sub-period is an abnormal sub-period. Among them, the score threshold can be determined based on the historical experience of relevant technical personnel or based on a large number of experimental results, and there is no limitation in this application.

[0115] In another alternative embodiment, for each sub-period, it can be determined whether the sub-period is an abnormal sub-period according to the deviation between the similarity score of the sub-period and the similarity scores of other sub-periods.

[0116] Exemplarily, if the deviation degree between the similarity score of a certain sub-period and the similarity scores of other sub-periods is large, it is determined that the sub-period is an abnormal sub-period; if the deviation degree between the similarity score of a certain sub-period and the similarity scores of other sub-periods is small, it is determined that the sub-period is a normal sub-period.

[0117] In the embodiments of the present application, by determining the similarity score according to the similarity between adjacent local operation data and determining abnormal sub-periods from each sub-period according to the similarity scores of each sub-period, the rationality of determining abnormal sub-periods can be ensured.

[0118] To ensure the accuracy of the similarity score, on the basis of the above embodiments, in the embodiments of the present application, an alternative way to determine the similarity score is provided, as Figure 4 shown, which specifically includes the following steps:

[0119] S401. Determine the data change trend of the local operation data corresponding to the sub-period and the data change trend of the local operation data corresponding to the neighboring sub-period.

[0120] Among them, the so-called data change trend is used to characterize the change of local operation data at each sampling moment.

[0121] In an alternative embodiment, the local operation data corresponding to the sub-period can be input into a trained trend determination model, and the trend determination model outputs the data change trend corresponding to the sub-period according to the local operation data corresponding to the sub-period and the model parameters.

[0122] Correspondingly, the local operation data corresponding to the neighbor sub-period can also be input into a trained trend determination model, and the trend determination model outputs the data change trend corresponding to the neighbor sub-period according to the local operation data corresponding to the neighbor sub-period and the model parameters.

[0123] In another alternative embodiment, the data change trend at each sampling moment within the sub-period can be determined according to the data fluctuation between the operation data at each sampling moment within the sub-period; correspondingly, the data change trend at each sampling moment within the neighbor sub-period can also be determined according to the data fluctuation between the operation data at each sampling moment within the neighbor sub-period. It can be understood that the number of sampling moments within the sub-period is the same as the number of sampling moments within the neighbor sub-period.

[0124] S402, determine the similarity score of the sub-period according to the trend similarity between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighbor sub-period.

[0125] In an alternative embodiment, the data change trend corresponding to the sub-period and the data change trend corresponding to the neighbor sub-period can be compared for consistency to obtain the trend similarity between the data change trends of the two sub-periods; subsequently, the similarity score of the sub-period is determined according to the trend similarity.

[0126] Exemplarily, for each sampling moment within the sub-period, the corresponding sampling moment can be determined from each sampling moment within the neighbor sub-period according to the arrangement result of the sampling moment in the time sequence; subsequently, the trend similarity at this sampling moment is determined according to the data change trend at this sampling moment and the data change trend at the corresponding sampling moment. Further, according to the trend similarities at each sampling moment within the sub-period, the trend similarity between the data change trends of the two sub-periods can be determined.

[0127] For example, the sub - time period T includes sampling times T1, T2, and T3, and the neighboring sub - time period S includes sampling times S1, S2, and S3. The trend similarity K1 can be determined according to the data change trend at T1 and S1, the trend similarity K2 can be determined according to the data change trend at T2 and S2, and the trend similarity K3 can be determined according to the data change trend at T3 and S3. Subsequently, the trend similarity between the sub - time period T and the neighboring sub - time period S is determined according to K1, K2, and K3.

[0128] After determining the trend similarity, the determined trend similarity can be used as an index to query in the pre - constructed correspondence between candidate trend similarities and candidate similarity scores, and then the similarity score of the sub - time period can be obtained.

[0129] In practical applications, it is possible that the overall data change trends of the local operation data in two sub - time periods are the same, but the data change times are different. If the above - mentioned method of comparing one by one based on sampling times is adopted, the accuracy of determining the similarity score may be reduced.

[0130] Based on this, in another alternative embodiment, the FastDTW algorithm can be used to determine the trend similarity between the sub - time period and the neighboring sub - time period according to the common trend shape between the data change trend corresponding to the sub - time period and the data change trend corresponding to the neighboring sub - time period. Subsequently, the similarity score of the sub - time period is determined according to the trend similarity.

[0131] In the embodiments of the present application, by determining the similarity score of the sub - time period according to the trend similarity between the data change trend corresponding to the sub - time period and the data change trend corresponding to the neighboring sub - time period, the accuracy of determining the similarity score can be ensured.

[0132] In order to ensure the rationality of determining the abnormal sub - time period, on the basis of the above - mentioned embodiments, in the embodiments of the present application, another alternative way to determine the similarity score is provided, as Figure 5 shown, and specifically includes the following steps:

[0133] S501, select the outlier similarity score from the similarity scores of each sub - time period.

[0134] Among them, the so - called outlier similarity score is used to represent the similarity score that deviates from other similarity scores among the similarity scores of each sub - time period.

[0135] In an alternative embodiment, the similarity scores of each sub - time period can be mapped to a two - dimensional coordinate system to obtain the similarity score points of each sub - time period. Subsequently, outlier mining processing is performed on each similarity score point to obtain the outlier score points that deviate from other similarity score points, and the similarity score corresponding to the outlier score point is used as the outlier similarity score.

[0136] For example, the Z-Score method can be adopted to determine the deviation degree between each similarity scoring point and other similarity scoring points, so as to determine the outlier scoring points.

[0137] S502. Use the sub-period corresponding to the outlier similarity score as the abnormal sub-period.

[0138] It can be understood that since each similarity score is associated with a sub-period, and the deviation degree of the similarity score can characterize the abnormal situation of the local operation data within the sub-period to a certain extent, therefore, in an optional implementation manner, after determining the outlier similarity score, the sub-period corresponding to the outlier similarity score can be directly used as the abnormal sub-period.

[0139] In the embodiments of the present application, by selecting the outlier similarity scores that deviate from the whole from the similarity scores of each sub-period, and directly using the sub-period corresponding to the outlier similarity score as the abnormal sub-period, the rationality of determining the abnormal sub-period can be ensured.

[0140] To ensure the rationality of the determined historical operation data, on the basis of the above embodiments, in the embodiments of the present application, an optional way to obtain historical operation data is provided, as Figure 6 shown, which specifically includes the following steps:

[0141] S601. Determine the historical detection period corresponding to the abnormal sub-period according to the time characteristics corresponding to the abnormal sub-period.

[0142] Among them, the so-called historical detection period is the period within the historical period that has the same time characteristics as the abnormal sub-period.

[0143] It can be understood that in order to improve the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period, in an optional implementation manner, the historical detection period corresponding to the abnormal sub-period can be determined from the historical period based on the time characteristics corresponding to the abnormal sub-period. Among them, in the case of multiple historical detection periods, the historical detection period that is closer to the abnormal sub-period in the time sequence is preferably selected.

[0144] For example, when the abnormal sub-period is October 2024, the historical detection periods corresponding to the abnormal sub-period can be October 2023 and October 2022. Among them, October 2023 is preferably selected as the historical detection period.

[0145] S602. Select the historical operation data that conforms to the job characteristics of the target job and is in the normal job state from each candidate operation data within the historical detection period.

[0146] Among them, the so-called candidate operation data is the operation data generated during the operation of each job within the historical detection period.

[0147] It can be understood that since the types of jobs running on the intelligent computing power network are rich and diverse, and each has different operation characteristics, when selecting historical operation data, it is necessary to select the operation data of other jobs similar to the target job within the historical detection period as the historical operation data.

[0148] In addition, since there are many reasons for job anomalies, in the embodiments of the present application, only the operation data generated when the job is in a normal operation state is extracted as a positive reference to determine whether there is abnormal operation data within the abnormal sub-period.

[0149] In an optional implementation manner, the job characteristics of the target job under each preset dimension can be extracted; subsequently, based on the job characteristics, the candidate operation data that conforms to the job characteristics and is in a normal operation state is screened out from the candidate operation data within the historical detection period as the historical operation data.

[0150] Among them, the preset dimensions may include but are not limited to user ID (user_id), group ID (group_id), queue (queue), number of nodes (node_count), number of CPUs (cpu_count), and application name (application), etc.

[0151] For example, a data screening template corresponding to the target job can be constructed based on the user ID, group ID, queue, number of nodes, number of CPUs, and application name of the target job; subsequently, the candidate operation data in a normal operation state is screened out from the candidate operation data within the historical detection period by using the data screening template corresponding to the target job as the historical operation data.

[0152] In the embodiments of the present application, by determining the historical detection period based on the time feature dimension and screening out the historical operation data in a normal operation state based on the job characteristics of the target job, the rationality and accuracy of the determination of the historical operation data can be ensured.

[0153] To ensure the accuracy of the determination of the anomaly detection result, on the basis of the above embodiments, in the embodiments of the present application, an optional way to determine the anomaly detection result is provided, as Figure 7 shown, which specifically includes the following steps:

[0154] S701, determine the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0155] In an alternative embodiment, the historical operation data and the local operation data corresponding to the abnormal sub-period can be input into the trained second similarity determination model at the same time. The second similarity determination model outputs the similarity between the historical operation data and the local operation data according to the historical operation data, the local operation data, and the model parameters.

[0156] In another alternative embodiment, in order to save the time for model parameter tuning, the historical operation data and the local operation data corresponding to the abnormal sub-period can be directly processed by sampling a statistical probability-based outlier detection algorithm (Copula-Based Outlier Detection, COPOD) to obtain the similarity between the historical operation data and the local operation data. Among them, COPOD can be used to effectively model the dependence relationship between multiple random variables.

[0157] S702. Determine the abnormal detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0158] In an alternative embodiment, since the historical operation data is the data generated when the job is in a normal operation state, the abnormal detection result of the target job can be determined according to the magnitude relationship between the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period and the similarity threshold. Among them, the similarity threshold can be determined based on the historical experience of relevant technicians or based on a large number of experimental results, and this is not limited in this application.

[0159] Exemplarily, when the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period is greater than the similarity threshold, it can be determined that the abnormal detection result of the target job is that the job runs normally; when the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period is less than or equal to the similarity threshold, it can be determined that the abnormal detection result of the target job is that the job runs abnormally.

[0160] In the embodiments of the present application, by determining the abnormal detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period, the accuracy of determining the abnormal detection result can be ensured.

[0161] Figure 8 FIG. is a schematic flowchart of a job abnormal detection method for an intelligent computing power network in another embodiment. On the basis of the above embodiments, this embodiment provides an alternative example of a job abnormal detection method for an intelligent computing power network. Combining Figure 8 , the specific implementation process is as follows:

[0162] S801, perform first-order difference processing on the original operation data of the target job in the to-be-detected period in the intelligent computing power network to obtain the global operation data of the target job in the to-be-detected period.

[0163] S802, split the global operation data to obtain the local operation data corresponding to each sub-period in the to-be-detected period.

[0164] S803, determine the data change trend of the local operation data in each sub-period, and the data change trend of the local operation data in the corresponding neighbor sub-periods of each sub-period.

[0165] S804, determine the similarity score of each sub-period according to the trend similarity between the data change trend of each sub-period and the data change trend of the corresponding neighbor sub-data of each sub-period.

[0166] Among them, the neighbor sub-period is the sub-period adjacent to the sub-period in each sub-period, and the neighbor sub-period is before or after the sub-period.

[0167] S805, select the outlier similarity score from the similarity scores of each sub-period, and use the sub-period corresponding to the outlier similarity score as the abnormal sub-period.

[0168] S806, determine the historical detection period corresponding to the abnormal sub-period according to the time characteristics corresponding to the abnormal sub-period.

[0169] S807, select the historical operation data that conforms to the job characteristics of the target job and is in the normal job state from each candidate operation data in the historical detection period.

[0170] S808, determine the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0171] S809, determine the abnormal detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0172] The specific processes of the above S801 - S809 can refer to the description of the above method embodiments, and their implementation principles and technical effects are similar, which will not be elaborated here.

[0173] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0174] Based on the same inventive concept, an embodiment of the present application further provides a job anomaly detection device for an intelligent computing power network for implementing the job anomaly detection method for an intelligent computing power network described above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the job anomaly detection device for an intelligent computing power network provided below can refer to the limitations on the job anomaly detection method for an intelligent computing power network in the above text, and will not be repeated here.

[0175] In an exemplary embodiment, as Figure 9 shown, a job anomaly detection device 1 for an intelligent computing power network is provided, including: a data splitting module 10, an anomaly determination module 20, a data acquisition module 30, and an anomaly detection module 40, where:

[0176] The data splitting module 10 is configured to split the original operation data of the target job in the intelligent computing power network during the to-be-detected period to obtain local operation data corresponding to each sub-period during the to-be-detected period;

[0177] The anomaly determination module 20 is configured to determine an abnormal sub-period from each sub-period according to the similarity between the local operation data corresponding to two adjacent sub-periods;

[0178] The data acquisition module 30 is configured to acquire historical operation data during the historical detection period corresponding to the abnormal sub-period;

[0179] The anomaly detection module 40 is configured to determine an anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-period.

[0180] In an exemplary embodiment, the data splitting module 10 is specifically configured to:

[0181] Perform first-order difference processing on the original operation data of the target job during the period to be detected in the intelligent computing power network to obtain the global operation data of the target job during the period to be detected; split the global operation data to obtain the local operation data corresponding to each sub-period during the period to be detected.

[0182] In an exemplary embodiment, the anomaly determination module 20 includes:

[0183] A score determination unit, configured to, for each sub-period, determine the similarity score of the sub-period according to the similarity between the local operation data corresponding to the sub-period and the local operation data corresponding to the neighboring sub-period; wherein, the neighboring sub-period is the sub-period adjacent to the sub-period among each sub-period, and the neighboring sub-period is before or after the sub-period.

[0184] An anomaly determination unit, configured to determine the anomaly sub-period from each sub-period according to the similarity scores of each sub-period.

[0185] In an exemplary embodiment, the score determination unit is specifically configured to:

[0186] Determine the data change trend of the local operation data corresponding to the sub-period, and the data change trend of the local operation data corresponding to the neighboring sub-period; determine the similarity score of the sub-period according to the trend similarity between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighboring sub-period.

[0187] In an exemplary embodiment, the anomaly determination unit is specifically configured to:

[0188] Select the outlier similarity score from the similarity scores of each sub-period; use the sub-period corresponding to the outlier similarity score as the anomaly sub-period.

[0189] In an exemplary embodiment, the data acquisition module 30 is specifically configured to:

[0190] Determine the historical detection period corresponding to the anomaly sub-period according to the time characteristics corresponding to the anomaly sub-period; select the historical operation data that conforms to the job characteristics of the target job and is in the normal job state from each candidate operation data within the historical detection period.

[0191] In an exemplary embodiment, the anomaly detection module 40 is specifically configured to:

[0192] Determine the similarity between the historical operation data and the local operation data corresponding to the anomaly sub-period; determine the anomaly detection result of the target job according to the similarity between the historical operation data and the local operation data corresponding to the anomaly sub-period.

[0193] Each module in the above-mentioned job anomaly detection device for the intelligent computing power network can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0194] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store historical operation data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for job anomaly detection in the intelligent computing power network.

[0195] Those skilled in the art can understand that Figure 10 the structure shown in

[0196] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0197] In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in each of the above method embodiments.

[0198] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in each of the above method embodiments.

[0199] It should be noted that the data involved in this application (including but not limited to historical operation data, etc.) are all data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0200] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0201] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0202] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for detecting abnormal operation of an intelligent computing network, characterized in that: The method comprises: Performing first-order difference processing on the original operation data of the target job in the intelligent computing power network during the period to be detected, to obtain the global operation data of the target job during the period to be detected; According to the time window size and the window sliding distance, the global operation data is split in the time dimension to obtain the local operation data corresponding to each sub-period in the time period to be detected; wherein the duration of each sub-period is the same; For each sub-period, a similarity score of the sub-period is determined according to the trend similarity between the data change trend of the local operation data corresponding to the sub-period and the data change trend of the local operation data corresponding to the neighbor sub-period; wherein the neighbor sub-period is a sub-period adjacent to the sub-period in each sub-period, and the neighbor sub-period is located before or after the sub-period; According to the similarity scores of each sub-period, an abnormal sub-period is determined from each sub-period; Obtaining historical operation data within the historical detection period corresponding to the abnormal sub-period; An abnormality detection result of the target job is determined according to the historical operation data and the local operation data corresponding to the abnormal sub-period.

2. The method according to claim 1, characterized in that The method further comprises: Determine the data change trend of the local operation data corresponding to the sub-period according to the data fluctuation between the operation data at each sampling time in the local operation data corresponding to the sub-period; and According to the data fluctuation between the operation data at each sampling time in the local operation data corresponding to the neighbor sub-period, the data change trend of the local operation data corresponding to the neighbor sub-period is determined.

3. The method according to claim 1, characterized in that Determining the similarity score of the sub-period according to the trend similarity between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighboring sub-period includes: Adopting the Fast Dynamic Time Warping FastDTW algorithm, according to the common trend shape between the data change trend corresponding to the sub-period and the data change trend corresponding to the neighboring sub-period, determining the trend similarity between the sub-period and the neighboring sub-period; A similarity score for the sub-period is determined based on the trend similarity.

4. The method according to claim 1, characterized in that: The step of determining an abnormal sub-period from each sub-period according to the similarity score of each sub-period includes: From the similarity scores of each sub-period, select the outlier similarity score; The sub-period corresponding to the outlier similarity score is taken as the abnormal sub-period.

5. The method according to claim 1, characterized in that The acquiring of historical operation data in the historical detection period corresponding to the abnormal sub-period includes: Determine, according to the time feature corresponding to the abnormal sub-period, a historical detection period corresponding to the abnormal sub-period; From each candidate operation data within the historical detection period, historical operation data that meets the operation characteristics of the target operation and is in a normal operation state is selected.

6. The method according to claim 1, characterized in that The determining, according to the historical operation data and the local operation data corresponding to the abnormal sub-period, the abnormality detection result of the target job includes: Determining the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period; The abnormality detection result of the target job is determined according to the similarity between the historical operation data and the local operation data corresponding to the abnormal sub-period.

7. A device for detecting abnormal operation of an intelligent computing network, characterized in that: The device comprises: A data splitting module is used to perform first-order difference processing on the original operation data of the target job in the intelligent computing power network within the period to be detected, so as to obtain the global operation data of the target job within the period to be detected; split the global operation data in the time dimension according to the preset time window size and window sliding distance, so as to obtain the local operation data corresponding to each sub-period within the period to be detected; wherein the duration of each sub-period is the same; an abnormality determination module, for each sub-period, determining a similarity score of the sub-period according to a trend similarity between a data change trend of local operation data corresponding to the sub-period and a data change trend of local operation data corresponding to a neighboring sub-period; wherein the neighboring sub-period is a sub-period adjacent to the sub-period in each sub-period, and the neighboring sub-period is located before or after the sub-period; determining an abnormal sub-period from each sub-period according to the similarity score of each sub-period; A data acquisition module, used to acquire historical operation data in a historical detection period corresponding to the abnormal sub-period; The anomaly detection module is used to determine the anomaly detection result of the target job according to the historical operation data and the local operation data corresponding to the abnormal sub-period.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Detection method and device for network anomaly traffic and computer readable storage medium

    CN109327345A

  • Base station fault detection method and device, equipment and storage medium

    CN114793345A

  • Intelligent analysis method for civil engineering detection data based on machine learning

    CN118035772A