Detection method for interface response time anomaly, and related device
Through the dynamic threshold method and the confidence interval characteristics of the positive distribution, the lag problem of interface response time abnormal detection is solved, and more accurate interface response time abnormal judgment is achieved, adapting to data distribution changes, and reducing false alarms.
Patent Information
- Application Number
- PCT/CN2024/135008
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, the interface system in civil aviation ticket trading business has a lag in the detection of interface response time abnormality detection, making it difficult to accurately judge the abnormality of interface response time, especially when the data distribution changes with time, the false alarm rate of the static threshold is high.
The dynamic threshold method is adopted, by obtaining the current interface response time data, cleaning and using the interface response time abnormal detection model matched by the interface code, combining the confidence interval characteristics of the positive distribution to perform abnormal detection, and dynamically update the model to adapt to the data distribution changes.
It realizes more accurate interface response time abnormal detection, reduces false alarms, improves the real-time and accuracy of detection, and adapts to the dynamic changes in data distribution.
Smart Images

Figure CN2024135008_03072025_PF_FP_ABST
Abstract
Description
A method and related equipment for detecting abnormal interface response time
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311861580.9 and application name “A method for detecting interface response time anomalies and related equipment”, the entire contents of which are incorporated by reference in this disclosure. Technical Field
[0002] The present disclosure relates to the field of aviation transportation technology, and more particularly to a method for detecting abnormal interface response time and related equipment. Background Art
[0003] Currently, interface systems in civil aviation ticket transactions can experience anomalies such as sudden unresponsiveness and numerous timeouts, impacting core systems and, in turn, the stability of civil aviation operations. Currently, static thresholds are commonly used to detect interface response time anomalies. However, because data distribution fluctuates over time, due to environmental changes, seasonal influences, and other factors, determining an appropriate static threshold is difficult, resulting in a high error rate for static thresholds. For example, when an interface response time anomaly occurs, the anomaly alarm generated by the static threshold often lags behind. Summary of the Invention
[0004] In view of this, the present disclosure provides a method and related equipment for detecting interface response time anomalies, so as to more accurately analyze the distribution of interface response time data through dynamic thresholds, thereby more accurately judging the situation of interface response time anomalies.
[0005] A method for detecting abnormal interface response time, comprising:
[0006] Get the current interface response time data;
[0007] Clean the current interface response time data to obtain the system response time, the system call background response time, and the interface code;
[0008] Obtaining an interface response time anomaly detection model that matches the interface code, wherein the interface response time anomaly detection model is trained using a preset anomaly detection algorithm on a historical data set of interface response times within a preset time period before the current moment, and the preset anomaly detection algorithm is related to a normal distribution;
[0009] The system response time and the system call background response time are input into the interface response time anomaly detection model for anomaly detection, and the interface response time anomaly detection result is obtained according to the confidence interval characteristics of the normal distribution.
[0010] A device for detecting abnormal interface response time, comprising:
[0011] A data acquisition unit, configured to acquire current interface response time data;
[0012] a cleaning unit configured to clean the current interface response time data to obtain the system response time, the system call background response time, and the interface code;
[0013] a model acquisition unit configured to acquire an interface response time anomaly detection model that matches the interface code, wherein the interface response time anomaly detection model is trained using a preset anomaly detection algorithm on a historical data set of interface response times within a preset time period before a current moment, and the preset anomaly detection algorithm is related to a normal distribution;
[0014] The anomaly detection unit is configured to input the system's own response time and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtain the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
[0015] An electronic device, comprising: a memory and a processor;
[0016] The memory is used to store at least one instruction;
[0017] The processor is configured to execute the at least one instruction to implement the above-mentioned method for detecting abnormal interface response time.
[0018] A computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the method for detecting abnormal interface response time is implemented.
[0019] From the above technical solution, it can be seen that the present disclosure provides a method and related equipment for detecting interface response time anomalies, which obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the interface response time anomaly. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0021] FIG1 is a flow chart of a method for detecting abnormal interface response time disclosed in an embodiment of the present disclosure;
[0022] FIG2 is a flow chart of a method for obtaining an interface response time anomaly detection model that matches an interface code, disclosed in an embodiment of the present disclosure;
[0023] FIG3 is a flow chart of a training method for an interface response time anomaly detection model disclosed in an embodiment of the present disclosure;
[0024] FIG4 (1) is a bell-shaped curve diagram of a SAT system call background response time disclosed in an embodiment of the present disclosure;
[0025] FIG4 (2) is a bell-shaped curve diagram of the response time of a SAT system itself disclosed in an embodiment of the present disclosure;
[0026] FIG5 is a distribution diagram of anomaly detection of a SAT system call background response time and a SAT system response time according to an embodiment of the present disclosure;
[0027] FIG6 is a diagram of an interface response time anomaly detection alarm data chart disclosed in an embodiment of the present disclosure;
[0028] FIG7 is a distribution diagram of anomaly detection of a low-frequency call disclosed in an embodiment of the present disclosure;
[0029] FIG8 is a schematic diagram of a device for detecting abnormal interface response time according to an embodiment of the present disclosure;
[0030] FIG9 is a schematic structural diagram of an electronic device disclosed in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0032] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] The embodiment of the present disclosure discloses a method for detecting anomalies in interface response time and related equipment, which obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the abnormality of the interface response time.
[0035] 1 is a flow chart of a method for detecting abnormal interface response time disclosed in an embodiment of the present disclosure. The method includes:
[0036] Step S101: Acquire current interface response time data.
[0037] In actual applications, the current interface response time data may be interface response time data within a certain time period (eg, 1 minute).
[0038] The current interface response time data may include: request date and time (date), interface name (ServiceName), system call background response time (BackSystemCostTime), system response time (ProcessTime), overall response time (TotalCostTime), response ID (Requestid), transaction ID (Txnid), SAT processing request IP address (HostIP), request upstream system name (AppKey), request status (CallResult), error code (ErrorCode), initial error code (NewErrorCode), interface code (ServiceCode).
[0039] Step S102: Clean the current interface response time data to obtain the system response time, the system call background response time, and the interface code.
[0040] This embodiment cleans the current interface response time data, retaining only the system's own response time, the system call background response time, and the interface code.
[0041] The system in the system response time and the system call backend response time may be a SAT (Service API of TravelSky) system. Accordingly, the system response time may be the SAT system response time, and the system call backend response time may be the SAT system call backend response time.
[0042] Step S103: Acquire an interface response time anomaly detection model that matches the interface code.
[0043] It should be noted that, since the interface name is a character string and cannot be calculated, this embodiment uses the interface code instead of the interface name, and the interface code and the interface name are in a one-to-one correspondence.
[0044] Assuming that the interface code is 1002, the matched interface response time anomaly detection model is the interface response time anomaly detection model with contamination=0.005.
[0045] Contamination is a parameter setting of the EllipticEnvelope algorithm. When the value of "contamination" is close to 0, it means that there are few outliers in the data and normal values are the majority. When the value of "contamination" is close to 1, it means that there are relatively many outliers in the data and normal values are the minority.
[0046] The interface response time anomaly detection model is trained on a historical data set of interface response time within a preset time period before the current moment using a preset anomaly detection algorithm, where the preset anomaly detection algorithm is related to the normal distribution.
[0047] The preset time period is determined based on actual needs, for example, six days. This means that when training the interface response time anomaly detection model, the historical interface response time data from six days ago can be randomly sampled and updated 12 times daily in the morning, afternoon, and evening time periods to generate a historical interface response time dataset. In actual applications, this collected historical interface response time dataset can be combined with a low-frequency supplementary dataset to create a new training dataset.
[0048] Regular machine learning is performed using the Elliptic Envelope anomaly detection algorithm to update the model, thereby achieving a fully automatic intelligent AI learning model.
[0049] The Elliptic Envelope algorithm is a statistical method for detecting outliers in multidimensional datasets. Its purpose is to identify data points that deviate from a normal distribution model. The Elliptic Envelope algorithm is often used for anomaly detection and outlier removal to help identify unusual patterns or unusual behavior in data.
[0050] Step S104: input the system response time and the system call background response time into the interface response time anomaly detection model to perform anomaly detection, and obtain the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
[0051] The present disclosure adopts a mechanism in which the system automatically samples data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, thereby realizing a dynamic threshold.
[0052] Assuming that most system requests currently respond within 3 to 5 seconds, the normal distribution after data modeling will also fall within 3 to 5 seconds. During interface response time anomaly detection, interface response time data less than 3 seconds and greater than 5 seconds will be identified as abnormal, thus forming a dynamic threshold of 3 to 5 seconds.
[0053] The system changes based on the reasons for the changes in interface response time. The response time of most requests becomes 5 to 10 seconds. At this time, the system automatically updates the training data by sampling data in the preset time period. After re-modeling, the normal distribution interval will also change to 5 to 10 seconds. At this time, the dynamic threshold will also dynamically change to 5 to 10 seconds.
[0054] Based on this, the present disclosure obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
[0055] In summary, the present disclosure provides a method for detecting interface response time anomalies, which obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the interface response time anomaly.
[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0057] The interface response time anomaly detection result obtained in the above embodiment can be further analyzed using a decision tree method for secondary confirmation. When the interface response time anomaly is confirmed again, an abnormal data graph can be drawn.
[0058] Therefore, to further optimize the above embodiment, the detection method may further include the following steps regarding the process of secondary determination of the interface response time anomaly detection result using a decision tree:
[0059] Identify all abnormal points from the interface response time anomaly detection results;
[0060] Determine the abnormal ratio of all abnormal points to the abnormality detection result of the interface response time, and the abnormal time interval of all abnormal points;
[0061] When the abnormality ratio exceeds a preset abnormality ratio, and / or the abnormal time interval satisfies a preset time interval restriction condition, it is determined that the interface response time corresponding to the current interface response time data is abnormal.
[0062] This paper verifies the interface response time anomaly detection results through a decision tree to further improve the accuracy of anomaly detection. The decision tree in this paper mainly determines the anomaly ratio (the ratio of abnormal requests to total requests) and the anomaly time interval (the time interval between two anomalies).
[0063] In practical applications, after the interface response time anomaly detection results obtained by the interface response time anomaly detection model are verified using a decision tree, an anomaly data graph can be drawn, and the interface response time anomaly data and the anomaly data graph can be output as an alarm in the form of an email.
[0064] To further optimize the above embodiment, referring to FIG2 , a flow chart of a method for obtaining an interface response time anomaly detection model that matches an interface code disclosed in an embodiment of the present disclosure is shown. The method includes:
[0065] Step S201: Determine whether an interface response time anomaly detection model matching the interface code exists. If yes, execute step S202; if not, execute step S203.
[0066] For example, if the interface codes are 1001, 1002, and 1003, then determine whether there are three models: model_1001, model_1002, and model_1003.
[0067] Step S202: directly obtain an interface response time anomaly detection model.
[0068] In actual applications, for interface codes 1001, 1002, and 1003, the content of the interface response time anomaly detection model obtained is:
[0069] {'1001':'EllipticEnvelope(contamination=0.001)','1002':'EllipticEnvelope(contamination=0.005)'},'1003':'EllipticEnvelope(contamination=0.008)'}
[0070] Step S203: Train an interface response time anomaly detection model based on a historical data set of interface response times within a preset time period before the current moment.
[0071] To further optimize the above embodiment, see FIG3 , which is a flow chart of a method for training an interface response time anomaly detection model disclosed in an embodiment of the present disclosure. The method includes:
[0072] Step S301: Acquire multiple random sampling times based on the morning, noon and evening time ranges of a preset time period.
[0073] In practical applications, the value of the preset time period and the value of the random sampling time are determined according to actual needs, and this disclosure does not limit them here.
[0074] For example, the preset time period is from 2023-12-11 10:00 to 2023-12-11 11:00. Within this time period, several random times are generated: "2023-12-11 10:13", "2023-12-11 10:22", and "2023-12-11 10:47". Based on the 3 minutes before and after these random times (preset), the database is used to obtain the historical data of the interface response time to form a historical interface response time data set.
[0075] Step S302: Based on the interface information to be sampled and the random sampling time, a historical data set of the interface response time within the preset time period before the current moment is obtained from a database to obtain an initial training data set.
[0076] In this embodiment, the various historical data in the interface response time historical data set may include: request date and time (date), interface name (ServiceName), system call background response time (BackSystemCostTime), system itself response time (ProcessTime), overall response time (TotalCostTime), response ID (Requestid), transaction ID (Txnid), SAT processing request IP address (HostIP), request upstream system name (AppKey), request status (CallResult), error code (ErrorCode), initial error code (NewErrorCode), interface code (ServiceCode).
[0077] Step S303: Clean and filter the initial training data set to obtain a training data set.
[0078] The training data set includes: a system response time set, a system call background response time set, and an interface code set.
[0079] Among them, each system response time in the system response time set may be a SAT system call background response time, and each system call background response time in the system call background response time set may be a SAT system response time.
[0080] Based on the interface code dimension, the corresponding mean μ and standard deviation σ of the SAT system call background response time and the SAT system's own response time are calculated respectively.
[0081] The probability density function is:
[0082] f(x)=1 / (σ√(2π))*exp(-((x-μ) 2 ) / (2σ 2 ))
[0083] Where μ is the mean of the distribution, σ 2 is the variance of the distribution.
[0084] We can obtain the bell-shaped curve of the SAT system call background response time as shown in Figure 4(1), and the bell-shaped curve of the SAT system's own response time as shown in Figure 4(2).
[0085] Finally, the multivariate Gaussian distribution model training using the EllipticEnvelope anomaly detection algorithm will produce 1 to N models (according to the number of backend interface types called by SAT).
[0086] These models are written into dictionaries with the interface code as the key and the model as the value (for example: {'1005':'EllipticEnvelope(contamination=0.003)','1006':'EllipticEnvelope(contamination=0.003)'},'1007':'EllipticEnvelope(contamination=0.003)'}) and provided to the detection program for use.
[0087] Step S304: Based on the classification of each historical interface code in the interface code set, a preset anomaly detection algorithm is used to train a multivariate Gaussian distribution model for the historical response time of the system itself and the historical response time of the system call background corresponding to each historical interface code to obtain an interface response time anomaly detection model corresponding to each historical interface code.
[0088] Therefore, each historical interface code in the interface code set in the present disclosure corresponds to an interface response time anomaly detection module. Therefore, according to the interface code in the current interface response time data, the corresponding interface response time anomaly detection module can be uniquely determined.
[0089] It's important to note that the model training process is performed daily. Request data from civil aviation-related transaction systems is cyclical. Interface response times vary based on changes in request volume from Monday to Sunday (for example, travel increases from Friday to Sunday, resulting in a corresponding increase in request volume). Training data is updated by moving forward six days from the current date (if today is Wednesday, data from the morning, noon, and evening of last Thursday is collected). The model is rebuilt using the above steps. The previously established interface response time anomaly detection model at 23:59:59 can be replaced, ensuring that the most up-to-date model for interface response time anomaly detection is used. Simultaneously, based on daily training data updates, the model (built using the Elliptic Envelope anomaly detection algorithm, which is based on a normal distribution and updates the corresponding normal distribution as training data is updated) is iterated. The abnormal range for interface response time is also adjusted to implement dynamic thresholds.
[0090] For example, the normal response time range is between 10-500ms. Any time outside this range is considered an interface response time anomaly. If server hardware issues cause this range to shift to between 150-700ms and cannot be quickly restored, the anomaly detection program will automatically collect the latest interface response time data as a training set and update the interface response time anomaly detection model.
[0091] (1) Example 1
[0092] Two normal or high-frequency call interfaces: DirectVerify and Availability are tested. The interface codes (ServiceCode) of the interfaces are: 1006 and 1007 respectively, the abnormality rate is: 0.05, and the alarm recipient is.
[0093] (1) The real-time detection module uses the DirectVerify and Availability interfaces as query conditions to call the data collection module to query the SAT Elsticsearch database for data within one minute. The obtained JSON information is cleaned and filtered. A DataFrame containing date, ServiceName, BackSystemCostTime, SATProcessTime, TotalCostTime, Requestid, Txnid, HostIP, AppKey, CallResult, ErrorCode, NewErrorCode, and ServiceCode is generated.
[0094] A DataFrame is a data structure commonly used in data analysis and data science. It is a two-dimensional table, similar to a spreadsheet or database table, that can be used to store and manipulate structured data. DataFrames are most commonly used in data analysis libraries in programming languages.
[0095] (2) Then call the database to query DirectVerify. The interface codes of the Availability interface are: 1006 and 1007. At this time, the real-time detection module will call the model dictionary generated by the AI machine learning training module and obtain the training model corresponding to the interface through the Key: value method. Substitute the BackSystemCostTime and SATProcessTime data of the two interfaces into the models 1006 and 1007 for prediction. As shown in Figure 5, the anomaly detection distribution diagram of the SAT system call background response time and the SAT system's own response time detects anomalies formed by abnormal values that deviate from the normal distribution.
[0096] (3) Trigger the result analysis and mapping module. At this time, the data of these abnormal points only contains three dimensions: BackSystemCostTime, SATProcessTime, and ServiceCode (interface code). It is necessary to restore the dimensions with the original DataFrame through index query to form an abnormal DataFrame. Determine the abnormal rate t:
[0097] t = a / d;
[0098] Where a is the number of data with the abnormal DataFrame interface code of 1006, and d is the number of data with the original DataFrame interface code of 1006.
[0099] Determine whether t is greater than the abnormal rate (0.05) of the interface with interface code 1006 in the database. If it is greater, it is determined to be abnormal, triggering the result email push module to alarm, and sending the alarm email to "abctest@hotmail.com". The email content is the interface response time abnormality detection alarm data chart shown in Figure 6, including abnormal interface information and data icon information, and the attachment contains detailed information on the interface abnormality, etc.
[0100] (II) Example 2
[0101] A low-frequency call interface, GetAvailability, is detected. The interface code of the interface is 1005, the abnormality rate is 0.03, and the alarm receiver is abctest@hotmail.com.
[0102] (1) The real-time detection module uses the GetAvailability interface as the query condition to call the data collection module to query the SAT Elsticsearch database for data within one minute. If there is a call request for this interface within one minute, the obtained JSON information is cleaned and filtered. A DataFrame containing date, ServiceName, BackSystemCostTime, SATProcessTime, TotalCostTime, Requestid, Txnid, HostIP, AppKey, CallResult, ErrorCode, NewErrorCode, and ServiceCode is generated. If there is no such request, the detection process is skipped.
[0103] (2) Generally, for low-frequency call interfaces, the basic data lacks the time-consuming data of the interface, resulting in no corresponding model generation, making it impossible to make predictions. In this case, the prediction program will add the prepared prediction data of the interface to the low-frequency supplementary data set. The supplementary data set will be continuously supplemented by the real-time detection module until it reaches the preset upper limit, thereby achieving the purpose of self-learning to improve the training data of the AI machine learning training module.
[0104] (3) The database is called again to query the interface ID of the GetAvailability interface: 1005. At this time, the real-time detection module will call the model dictionary generated by the AI machine learning training module and obtain the training model corresponding to the interface through the Key: value method. The BackSystemCostTime and SATProcessTime data of the two interfaces are respectively substituted into the model 1005 for prediction. As shown in Figure 7, the anomaly detection distribution diagram of low-frequency calls shows the anomaly points formed by the detected outliers that deviate from the normal distribution.
[0105] (4) Trigger the result analysis and mapping module. At this time, the data of these abnormal points only contains three dimensions: BackSystemCostTime, SATProcessTime, and Interface ID. It is necessary to restore the dimensions with the original DataFrame through index query to form an abnormal DataFrame. Determine the abnormality rate t:
[0106] t = a / d;
[0107] Where a is the number of data with the abnormal DataFrame interface ID of 1005, and d is the number of data with the original DataFrame interface ID of 1005.
[0108] Determine whether t is greater than the exception rate (0.03) for interface ID 1005 in the database. If so, it is considered a quasi-anomaly. At this time, because the overall number of requests for low-frequency interface calls is low, the exception rate may be extremely high. A secondary confirmation is performed to weight the number of abnormal requests and the overall response time to reduce the high exception rate caused by the request volume and avoid misjudgment. After the anomaly is confirmed, the result email push module is triggered to issue an alarm. The email content contains the abnormal interface information and data icon information, and the attachment contains all detailed information about the interface anomaly.
[0109] (III) Example 3
[0110] The AOS-IET front end calls the management module HTTP interface to perform insert, update, and delete operations on the backend database.
[0111] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0112] Corresponding to the above method embodiment, the present disclosure also discloses a device for detecting abnormal interface response time.
[0113] 8 , which is a schematic diagram of a device for detecting abnormal interface response time disclosed in an embodiment of the present disclosure, the device may include:
[0114] The data acquisition unit 401 is configured to acquire current interface response time data.
[0115] In actual applications, the current interface response time data may be interface response time data within a certain time period (eg, 1 minute).
[0116] The current interface response time data may include: request date and time (date), interface name (ServiceName), system call background response time (BackSystemCostTime), system response time (ProcessTime), overall response time (TotalCostTime), response ID (Requestid), transaction ID (Txnid), SAT processing request IP address (HostIP), request upstream system name (AppKey), request status (CallResult), error code (ErrorCode), initial error code (NewErrorCode), interface code (ServiceCode).
[0117] The cleaning unit 402 is configured to clean the current interface response time data to obtain the system response time, the system call background response time and the interface code.
[0118] This embodiment cleans the current interface response time data, retaining only the system's own response time, the system call background response time, and the interface code.
[0119] The system in the system response time and the system call backend response time may be a SAT (Service API of TravelSky) system. Accordingly, the system response time may be the SAT system response time, and the system call backend response time may be the SAT system call backend response time.
[0120] The model acquisition unit 403 is configured to acquire an interface response time anomaly detection model that matches the interface code.
[0121] The interface response time anomaly detection model is trained on a historical data set of interface response time within a preset time period before the current moment using a preset anomaly detection algorithm, and the preset anomaly detection algorithm is related to the normal distribution.
[0122] The anomaly detection unit 404 is configured to input the system's own response time and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtain the interface response time anomaly detection result based on the confidence interval characteristics of the normal distribution.
[0123] The present disclosure adopts a mechanism in which the system automatically samples data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, thereby realizing a dynamic threshold.
[0124] Assuming that most system requests currently respond within 3 to 5 seconds, the normal distribution after data modeling will also fall within 3 to 5 seconds. During interface response time anomaly detection, interface response time data less than 3 seconds and greater than 5 seconds will be identified as abnormal, thus forming a dynamic threshold of 3 to 5 seconds.
[0125] The system changes based on the reasons for the changes in interface response time. The response time of most requests becomes 5 to 10 seconds. At this time, the system automatically updates the training data by sampling data in the preset time period. After re-modeling, the normal distribution interval will also change to 5 to 10 seconds. At this time, the dynamic threshold will also dynamically change to 5 to 10 seconds.
[0126] Based on this, the present disclosure obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
[0127] In summary, the present disclosure discloses a device for detecting anomalies in interface response time, which obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the abnormality of the interface response time.
[0128] The interface response time anomaly detection result obtained in the above embodiment can be further analyzed using a decision tree method for secondary confirmation. When the interface response time anomaly is confirmed again, an abnormal data graph can be drawn.
[0129] Therefore, the detection device may further include:
[0130] an abnormal point determination unit, configured to determine all abnormal points from the interface response time abnormality detection result;
[0131] an abnormal parameter determination unit, configured to determine an abnormal proportion of all abnormal points in the interface response time abnormality detection result, and an abnormal time interval of all abnormal points;
[0132] The abnormal result determination unit is configured to determine that the interface response time corresponding to the current interface response time data is abnormal when the abnormal proportion exceeds a preset abnormal proportion and / or the abnormal time interval meets a preset time interval restriction condition.
[0133] This paper verifies the interface response time anomaly detection results through a decision tree to further improve the accuracy of anomaly detection. The decision tree in this paper mainly determines the anomaly ratio (the ratio of abnormal requests to total requests) and the anomaly time interval (the time interval between two anomalies).
[0134] In practical applications, after the interface response time anomaly detection results obtained by the interface response time anomaly detection model are verified using a decision tree, an anomaly data graph can be drawn, and the interface response time anomaly data and the anomaly data graph can be output as an alarm in the form of an email.
[0135] To further optimize the above embodiment, the model acquisition unit 403 may be specifically used to:
[0136] Determining whether the interface response time anomaly detection model matching the interface code exists;
[0137] If yes, directly obtain the interface response time anomaly detection model;
[0138] If not, the interface response time anomaly detection model is obtained by training based on the interface response time history data set within the preset time period before the current moment.
[0139] To further optimize the above embodiment, the detection device may further include:
[0140] A model training unit, configured to train the interface response time anomaly detection model;
[0141] The model training unit is specifically used for:
[0142] Based on the morning, noon and evening time ranges of the preset time period, a plurality of random sampling times are obtained;
[0143] Based on the interface information to be sampled and the random sampling time, obtaining the interface response time history data set within the preset time period before the current moment from a database to obtain an initial training data set;
[0144] Cleaning and filtering the initial training data set to obtain a training data set, wherein the training data set includes: a system response time set, a system call background response time set, and an interface code set;
[0145] Taking each historical interface code in the interface code set as the classification basis, the preset anomaly detection algorithm is used to train the multivariate Gaussian distribution model for the system's own historical response time and the system call background historical response time corresponding to each historical interface code to obtain the interface response time anomaly detection model corresponding to each historical interface code.
[0146] Therefore, each historical interface code in the interface code set in the present disclosure corresponds to an interface response time anomaly detection module. Therefore, according to the interface code in the current interface response time data, the corresponding interface response time anomaly detection module can be uniquely determined.
[0147] It's important to note that the model training process is performed daily. Request data from civil aviation-related transaction systems is cyclical. Interface response times vary based on changes in request volume from Monday to Sunday (for example, travel increases from Friday to Sunday, resulting in a corresponding increase in request volume). Training data is updated by moving forward six days from the current date (if today is Wednesday, data from the morning, noon, and evening of last Thursday is collected). The model is rebuilt using the above steps. The previously established interface response time anomaly detection model at 23:59:59 can be replaced, ensuring that the most up-to-date model for interface response time anomaly detection is used. Simultaneously, based on daily training data updates, the model (built using the Elliptic Envelope anomaly detection algorithm, which is based on a normal distribution and updates the corresponding normal distribution as training data is updated) is iterated. The abnormal range for interface response time is also adjusted to implement dynamic thresholds.
[0148] For example, the normal response time range is between 10-500ms. Any time outside this range is considered an interface response time anomaly. If server hardware issues cause this range to shift to between 150-700ms and cannot be quickly restored, the anomaly detection program will automatically collect the latest interface response time data as a training set and update the interface response time anomaly detection model.
[0149] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0150] Corresponding to the above embodiment, as shown in FIG9 , the present disclosure further provides an electronic device, which may include: a processor 1 and a memory 2;
[0151] The processor 1 and the memory 2 communicate with each other via a communication bus 3.
[0152] Processor 1, configured to execute at least one instruction;
[0153] Memory 2, configured to store at least one instruction;
[0154] The processor 1 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present disclosure.
[0155] The memory 2 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0156] The processor executes at least one instruction to implement the steps shown in the embodiment of the method for detecting abnormal interface response time.
[0157] In summary, the present disclosure provides an electronic device that obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the abnormality of the interface response time.
[0158] Corresponding to the above embodiment, the present disclosure further provides a computer-readable storage medium, which stores at least one instruction. When the at least one instruction is executed by a processor, the steps shown in the embodiment of the method for detecting an abnormal interface response time are implemented.
[0159] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0160] In summary, the present disclosure provides a computer-readable storage medium, which obtains the current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the abnormality of the interface response time.
[0161] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
[0162] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination.
[0163] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure. Industrial Applicability
[0164] The present disclosure provides a method and related equipment for detecting anomalies in interface response time, which obtains current interface response time data, cleans it to obtain the system response time itself, the system call background response time and the interface code, uses the interface code to obtain the corresponding interface response time anomaly detection model, and then inputs the system response time itself and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtains the interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution. The present disclosure adopts a mechanism for automatically sampling data in a preset time period to update the interface response time anomaly detection model after the interface response time changes, and the corresponding normal distribution interval also changes accordingly, that is, the dynamic threshold changes. Since the dynamic threshold takes into account the historical data information in the preset time period before the current moment, compared with the static threshold in the traditional solution, it can more accurately analyze the distribution of the interface response time data, thereby more accurately judging the abnormality of the interface response time.
Claims
1. A method for detecting abnormal interface response time, comprising: Obtaining current interface response time data; Cleaning the current interface response time data to obtain the system's own response time, the system's call-backend response time, and the interface code; Obtaining an interface response time anomaly detection model that matches the interface code, wherein the interface response time anomaly detection model is trained by a preset anomaly detection algorithm for an interface response time historical data set within a preset time period before the current moment, and the preset anomaly detection algorithm is related to the normal distribution; Inputting the system's own response time and the system's call-backend response time into the interface response time anomaly detection model for anomaly detection, and obtaining an interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
2. The detection method according to claim 1, wherein, It further includes: Determining all abnormal points from the interface response time anomaly detection result; Determining the abnormal occupancy ratio of all the abnormal points in the interface response time anomaly detection result, and the abnormal time interval of all the abnormal points; When the abnormal occupancy ratio exceeds a preset abnormal occupancy ratio, and / or the abnormal time interval meets a preset time interval limit condition, determining that the interface response time corresponding to the current interface response time data is abnormal.
3. The detection method according to claim 1, wherein The obtaining of the interface response time anomaly detection model that matches the interface code includes: Judging whether there is an interface response time anomaly detection model that matches the interface code; If so, directly obtaining the interface response time anomaly detection model; If not, training the interface response time anomaly detection model based on the interface response time historical data set within the preset time period before the current moment.
4. The detection method according to claim 1, wherein, The training process of the interface response time anomaly detection model includes: Based on the morning, noon, and evening time ranges of the preset time period, obtaining multiple random sampling times; Based on the interface information to be sampled and the random sampling times, obtaining the interface response time historical data set within the preset time period before the current moment from the database to obtain an initial training data set; Cleaning and filtering the initial training data set to obtain a training data set, wherein the training data set includes: a set of the system's own response times, a set of the system's call-backend response times, and a set of interface codes; Taking each historical interface code in the set of interface codes as a classification basis, training a multivariate Gaussian distribution model for the system's own historical response time and the system's call-backend historical response time corresponding to each historical interface code by using the preset anomaly detection algorithm, to obtain an interface response time anomaly detection model corresponding to each historical interface code.
5. A device for detecting abnormal interface response time, comprising: A data acquisition unit configured to obtain current interface response time data; A cleaning unit configured to clean the current interface response time data to obtain the system's own response time, the system's call-backend response time, and the interface code; A model acquisition unit, configured to acquire an interface response time anomaly detection model that matches the interface code, where the interface response time anomaly detection model is trained from an interface response time historical data set within a preset time period before the current moment by using a preset anomaly detection algorithm, and the preset anomaly detection algorithm is related to the normal distribution; An anomaly detection unit, configured to input the system's own response time and the system call background response time into the interface response time anomaly detection model for anomaly detection, and obtain an interface response time anomaly detection result according to the confidence interval characteristics of the normal distribution.
6. The detection device according to claim 5, wherein, It further includes: An anomaly point determination unit, configured to determine all anomaly points from the interface response time anomaly detection result; An anomaly parameter determination unit, configured to determine the anomaly ratio of all the anomaly points in the interface response time anomaly detection result, and the anomaly time interval of all the anomaly points; An anomaly result determination unit, configured to determine that the interface response time corresponding to the current interface response time data is abnormal when the anomaly ratio exceeds a preset anomaly ratio and / or the anomaly time interval meets a preset time interval limit condition.
7. The detection device according to claim 5, wherein, The model acquisition unit is specifically used for: Judging whether there is an interface response time anomaly detection model that matches the interface code; If so, directly acquire the interface response time anomaly detection model; If not, train the interface response time anomaly detection model based on the interface response time historical data set within the preset time period before the current moment.
8. The detection device according to claim 5, wherein, It further includes: A model training unit, configured to train the interface response time anomaly detection model; The model training unit is specifically used for: Based on the morning, noon, and evening time ranges of the preset time period, obtain multiple random sampling times; Based on the interface information to be sampled and the random sampling times, obtain the interface response time historical data set within the preset time period before the current moment from the database to obtain an initial training data set; Clean and filter the initial training data set to obtain a training data set, where the training data set includes: a system's own response time set, a system call background response time set, and an interface code set; Taking each historical interface code in the interface code set as a classification basis, perform multivariate Gaussian distribution model training on the system's own historical response time and the system call background historical response time corresponding to each historical interface code by using the preset anomaly detection algorithm, and obtain an interface response time anomaly detection model corresponding to each historical interface code.
9. An electronic device, wherein, The electronic device includes: a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement the method for detecting interface response time anomalies according to any one of claims 1 to 4.
10. A computer-readable storage medium, wherein, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the method for detecting interface response time anomalies according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method and device for detecting abnormal response time
CN111666187A
Interface anomaly detection method and device
CN115426180A
Abnormality detection method, device and system, electronic equipment and storage medium
CN116028255A
Application anomaly detection method and device, equipment and medium
CN117131405A
Interface response time abnormity detection method and related equipment
CN117827562A