Component change abnormality detection method, electronic device, and medium

By counting the spatial distribution values and historical performance of abnormal events in the cloud computing platform, and judging the anomalies of component changes, the detection failure problem caused by relying on manual rules in the existing technology is solved, and higher robustness and accuracy are achieved.

WO2025163414A1PCT designated stage Publication Date: 2025-08-07CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050494
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-29
Filing Date
2025-01-17
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

In the prior art, when computing components change, cloud computing platforms rely on abnormal detection rules set by manual experience, resulting in the rules being invalid when new types of components are changed, and they cannot effectively detect component changes abnormalities, affecting platform performance.

Method used

By counting the abnormal events of the changed server in the cloud computing platform server cluster, calculate the spatial distribution value of the abnormal events, and based on the reference value range of historical abnormal events, determine whether the component change is an abnormal change, and determine the server that exceeds the reference abnormal events.

Benefits of technology

Improves the robustness and accuracy of component change abnormality detection, reduces dependence on manual preset rules, and can promptly detect abnormal servers with damaged performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050494_07082025_PF_FP_ABST
    Figure IB2025050494_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a component change abnormality detection method, an electronic device, and a medium. The component change abnormality detection method comprises: counting abnormal events that occur in changed servers, which have undergone a component change, in a server cluster of a cloud computing platform; calculating spatial distribution values of the abnormal events, and on the basis of the spatial distribution values, determining whether the component change is an abnormal change; and if the component change is an abnormal change, on the basis of the condition about whether an over-reference abnormal event occurs in the changed servers, determining an abnormal server from among the changed servers, wherein the over-reference abnormal event is an abnormal event with an event attribute value beyond a corresponding reference value range, and the reference value range is obtained on the basis of attribute values of abnormal events that occur in a historical stage. The embodiments of the present disclosure can effectively improve the robustness and accuracy of component change abnormality detection.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]This disclosure claims priority to Chinese patent application number 202410123859.5, filed with the Patent Office of the People's Republic of China on January 29, 2024, entitled "Method, Apparatus, Electronic Device, and Medium for Detecting Component Change Anomalies," the entire contents of which are incorporated herein by reference. Technical Field: Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, electronic device, and medium for detecting component change anomalies. Background: Cloud computing platforms, as the infrastructure of the Internet, have a variety of computing components at their bottom layer, including storage, networking, central processing units (CPUs), and management and control. To meet the performance and stability requirements of cloud computing platforms, these computing components will also undergo changes due to upgrades and iterations. Changes to computing components may result in anomalies. For example, component changes may cause abnormal events to occur on a large number of servers in the cloud computing platform. Component change anomalies can have serious consequences. Therefore, component change anomaly detection is necessary during the computing component change process to promptly identify abnormal changes and the abnormal servers whose performance has been compromised during these abnormal changes, allowing for subsequent effective interception measures. Related technologies primarily rely on manual experience to set anomaly detection rules. For example, for a certain component change, if ten servers in a computing platform simultaneously experience startup failures, a component change anomaly is determined. This approach relies heavily on manually pre-set rules. When new types of component changes occur, these pre-set rules may not be effective. Therefore, a more effective component change anomaly detection solution is urgently needed. SUMMARY OF THE INVENTION In view of this, embodiments of the present disclosure provide a component change anomaly detection method, electronic device, and medium to at least partially address the aforementioned issues. According to a first aspect of an embodiment of the present disclosure, a component change anomaly detection method is provided, comprising: counting abnormal events occurring on changed servers that have undergone component changes in a server cluster of a cloud computing platform; calculating spatial distribution values ​​of the abnormal events, and determining whether the component change is an abnormal change based on the spatial distribution values; if it is an abnormal change, determining an abnormal server from the changed servers based on whether an over-baseline abnormal event occurs on the changed servers; the over-baseline abnormal event is an abnormal event in which an event attribute value exceeds a corresponding benchmark value range; the benchmark value range is obtained based on attribute values ​​of abnormal events occurring in a historical stage.According to a second aspect of an embodiment of the present disclosure, another component change anomaly detection method is provided, comprising: receiving a component change notification and obtaining an abnormal event that occurs on a changed server in a cloud computing platform after the component change; calculating a spatial distribution value of the abnormal event, and determining whether the component change is an abnormal change based on the spatial distribution value; if it is an abnormal change, determining an abnormal server from the changed servers based on whether an over-baseline abnormal event occurs on the changed server; the over-baseline abnormal event is an abnormal event in which an event attribute value exceeds a corresponding baseline value range; the baseline value range is obtained based on attribute values ​​of abnormal events that occurred in a historical stage; and returning an abnormal change notification, wherein the abnormal change notification includes identification information of the abnormal server. According to a third aspect of an embodiment of the present disclosure, another component change anomaly detection method is provided, comprising: receiving component change log information and abnormal event log information for a server cluster of a cloud computing platform; determining, for each component change recorded in the component change log information, a corresponding abnormal event information segment from the abnormal event log information; calculating, based on the abnormal event information segment, a spatial distribution value of abnormal events corresponding to each component change, and determining an abnormal change from each component change based on the spatial distribution value; for the abnormal change, determining an abnormal server from the changed servers based on whether an over-baseline abnormal event occurs in the changed server; the over-baseline abnormal event is an abnormal event in which an event attribute value exceeds a corresponding benchmark value range; the benchmark value range is obtained based on attribute values ​​of abnormal events that occurred in a historical stage; and returning an abnormal change notification, wherein the abnormal change notification includes identification information of the abnormal change and identification information of the abnormal server corresponding to the abnormal change. According to a fourth aspect of an embodiment of the present disclosure, a component change anomaly detection device is provided, comprising: a statistical module for counting abnormal events occurring in a changed server that has undergone component changes in a server cluster of a cloud computing platform; an abnormal change determination module for calculating a spatial distribution value of the abnormal event and determining whether the component change is an abnormal change based on the spatial distribution value; an abnormal server determination module for determining an abnormal server from the changed servers based on whether an over-baseline abnormal event occurs in the changed server if it is an abnormal change; the over-baseline abnormal event is an abnormal event in which an event attribute value exceeds a corresponding benchmark value range; the benchmark value range is obtained based on the attribute values ​​of abnormal events occurring in a historical stage.According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; the memory is configured to store at least one executable instruction, wherein the executable instruction causes the processor to perform operations corresponding to the method described in any one of the first to third aspects. According to a sixth aspect of the embodiments of the present disclosure, a computer storage medium is provided, storing a computer program, which, when executed by the processor, implements the method described in any one of the first to third aspects. The component change anomaly detection method, electronic device, and medium provided in the embodiments of the present disclosure, after obtaining an abnormal event occurring in a changed service in a server cluster, calculates a spatial distribution value of the abnormal event at the server cluster spatial domain level, and then performs component change anomaly detection based on the spatial distribution value. Furthermore, when a component change is detected as an abnormal change, a benchmark value range of the abnormal event attribute value is obtained at the time domain level based on the historical performance of the abnormal event, and then, based on the benchmark value range, determines an abnormal event exceeding the benchmark and an abnormal server. The disclosed embodiments reduce reliance on manually preset rules during component change anomaly detection. Instead, they use the spatial and historical performance of abnormal events as criteria for determining anomalies at the current stage. In other words, they primarily predict the current trend based on the historical performance of abnormal events. Therefore, the disclosed embodiments can effectively improve the robustness and accuracy of component change anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS To more clearly illustrate the disclosed embodiments or technical solutions in the prior art, the following briefly describes the drawings required for use in the embodiments or prior art descriptions. Obviously, the drawings described below represent only some of the embodiments described in the disclosed embodiments. Persons skilled in the art will be able to derive other drawings based on these drawings. Figure 1 is a step flow chart of a component change anomaly detection method according to Example 1 of the present disclosure; Figure 2 is a detection process schematic diagram corresponding to the embodiment shown in Figure 1; Figure 3 is a step flow chart of a component change anomaly detection method according to Example 2 of the present disclosure; Figure 4 is a step flow chart of a component change anomaly detection method according to Example 3 of the present disclosure; Figure 5 is a structural block diagram of a component change anomaly detection device according to Example 4 of the present disclosure; Figure 6 is a structural block diagram of a component change anomaly detection device according to Example 5 of the present disclosure; Figure 7 is a structural block diagram of a component change anomaly detection device according to Example 6 of the present disclosure; Figure 8 is a structural schematic diagram of an electronic device according to Example 7 of the present disclosure.DETAILED DESCRIPTION To help those skilled in the art better understand the technical solutions in the embodiments of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments represent only a portion of the embodiments of the present disclosure, and are not exhaustive. All other embodiments derived by persons skilled in the art based on the embodiments of the present disclosure should fall within the scope of protection of the embodiments of the present disclosure. Example 1 refers to FIG1 , which is a flowchart of a component change anomaly detection method according to Example 1 of the present disclosure. Specifically, the component change anomaly detection method provided in this embodiment includes the following steps: Step 102: Count abnormal events occurring on servers in a server cluster of a cloud computing platform that have undergone component changes. Specifically, when performing component changes, a phased release is typically used to release the new component to the server cluster to test the performance of the new component. For example, a cloud computing platform's server cluster contains 100 servers (i.e., physical machines). When a component is changed, the changed component can be selectively used on some servers (e.g., 40 of them), while the remaining servers continue to use the unchanged component. The changed servers in this step may refer to servers in the server cluster that have already installed the new component, such as the 40 servers in the above example. This step counts abnormal events that occur within a preset period of time after the component change on these servers that have already installed the new component. Specifically, abnormal events can be counted for each changed server in the server cluster. In this embodiment, the type of component that is changed and the specific content of the change are not limited. For example, it can be a storage component, a network component, a load balancing component, etc. The specific content of the abnormal event is also not limited in this embodiment. The server conditions that determine when an abnormal event has occurred can be pre-defined based on actual circumstances. For example, at least one of the following conditions can be defined as an abnormal event: service startup / stop, memory overflow, unexpected process exit, network packet loss, memory fragmentation, bandwidth surge, network disconnection, system downtime, restart failure, etc. Step 104 calculates the spatial distribution value of the abnormal event and, based on the spatial distribution value, determines whether the component change is an abnormal change. The spatial distribution value represents the ratio of the number of servers experiencing abnormal events among the changed servers to the number of servers experiencing abnormal events in the server cluster.Specifically, for each abnormal event obtained in step 102, this step can be used to calculate a spatial distribution value, thereby obtaining the spatial distribution value corresponding to each abnormal event. For example, for abnormal event A occurring in a changed server, the number of servers in the changed server that experienced abnormal event A can be obtained, assuming it is a. The number of servers in the server cluster of the cloud computing platform that experienced abnormal event A can also be obtained, assuming it is b. Based on the value of a / b, the spatial distribution value corresponding to abnormal event A can be obtained. For example, a / b can be determined as the spatial distribution value corresponding to abnormal event A, or a certain function can be performed on a / b, and the resulting function value can be determined as the spatial distribution value corresponding to abnormal event A. The spatial distribution value of the abnormal event obtained in this step can be positively correlated with the ratio of the number of servers in the changed server that experienced the abnormal event to the number of servers in the server cluster that experienced the abnormal event, or it can be negatively correlated with this ratio. When the two are positively correlated, the calculated spatial distribution value can be compared with a first preset spatial distribution threshold. When the spatial distribution value exceeds the first preset spatial distribution threshold, the component change is determined to be an abnormal change. The first preset spatial distribution threshold can be customized based on experience, for example, 0.9. When the two are negatively correlated, the calculated spatial distribution value can be compared with a second preset spatial distribution threshold. When the spatial distribution value is less than the second preset spatial distribution threshold, the component change is determined to be an abnormal change. The second preset spatial distribution threshold can also be customized based on experience, for example, 0.1. Since the abnormal event occurring in the changed server can be one or more types, the number of calculated spatial distribution values ​​can also be one or more. When there is only one spatial distribution value, whether it is an abnormal change can be determined based on the magnitude relationship between the spatial distribution value and a preset spatial distribution threshold. When there are multiple spatial distribution values, whether it is an abnormal change can be determined based on the magnitude relationship between each of the multiple spatial distribution values ​​and the preset spatial distribution threshold. For example, it can be set that when all spatial distribution values ​​are greater than the preset spatial distribution threshold, it is determined to be a normal change. Conversely, when there are a preset number of spatial distribution values ​​less than the preset spatial distribution threshold, it is determined to be an abnormal change. This preset number can be customized, for example, to any natural number. Theoretically, some abnormal event will "randomly" occur on a server at any moment.However, it is understandable that when an abnormal event associated with (caused by) a component change occurs, it tends to occur primarily on the servers that have undergone the change. Therefore, the present disclosure calculates a spatial distribution value representing the proportion of changed servers experiencing abnormal events in the server cluster, and uses this value to determine whether a component change is an abnormal change. Compared to determining whether an abnormal change is an abnormal change based on the number of changed servers experiencing abnormal events, this approach can effectively improve the robustness and accuracy of detecting abnormal component changes. In step 106, if the change is an abnormal change, abnormal servers are identified from the changed servers based on whether any abnormal events exceeding a baseline occur. An abnormal event exceeding a baseline is an abnormal event whose event attribute value exceeds a corresponding baseline value range. The baseline value range is derived based on the attribute values ​​of abnormal events that occurred historically. The baseline value range can be used to represent the typical value range of event attribute values ​​for abnormal events. Specifically, an abnormal server can refer to a server with impaired performance. Abnormal events that occur on the changed server during the current component change process are typically likely to have occurred in the past. However, because their event attribute values ​​may have been within a relatively normal range in those past periods, their occurrence in those periods did not significantly damage the abnormal server's performance. Therefore, in this step, based on the attribute values ​​of the abnormal event that occurred in the past, statistical analysis (such as using box plots or normal distribution analysis) can be performed to determine the relatively normal range, which is also the benchmark value range in this step. Furthermore, by comparing the time attribute values ​​of the abnormal event during the current component change process to see if they are within the benchmark value range, it is determined whether the current occurrence of the abnormal event has a tendency to cause server performance impairment, that is, whether the abnormal event is an above-baseline abnormal event. Specifically, if the event attribute values ​​of the abnormal event exceed the corresponding benchmark value range, it indicates that the abnormal event is an above-baseline abnormal event. The occurrence of the abnormal time during the current component change has a tendency to cause server performance impairment, and the changed server that experienced the above-baseline abnormal event is determined to be an abnormal server with impaired performance.In this step, by comparing the event attribute values ​​of the abnormal events caused by the component change with the baseline value range for abnormal events obtained based on historical data, we further identify abnormal events that may cause server performance impairment, and then determine the changed server experiencing such abnormal events as abnormal servers. This process uses the historical performance of abnormal events as the basis for determining current trends, thereby improving the accuracy of abnormal server identification and, in other words, further enhancing the robustness and accuracy of anomaly detection. In the disclosed embodiments, the specific content of the event attribute values ​​of abnormal events is not limited and can be customized based on the specific content of the abnormal event. For example, for abnormal events such as restart failure, network loss, and service start and stop, the event attribute values ​​may include: occurrence frequency; for abnormal events such as bandwidth sudden change, the event attribute values ​​may include: occurrence frequency, specific value for each occurrence (specific bandwidth sudden change in each occurrence), etc. The component change anomaly detection solution provided by the disclosed embodiments, after obtaining abnormal events occurring in changed services within a server cluster, calculates the spatial distribution of abnormal events at the server cluster spatial level. Component change anomaly detection is then performed based on this spatial distribution. Furthermore, when a component change is detected as abnormal, a baseline value range for the abnormal event attribute is determined based on the historical performance of the abnormal event at the temporal level. Based on this baseline value range, abnormal events exceeding the baseline and abnormal servers are identified. This disclosed embodiment reduces reliance on manually pre-set rules during component change anomaly detection, instead using the spatial and historical performance of abnormal events as the criteria for determining anomalies at the current stage. In other words, the historical performance of abnormal events is primarily used to predict the current trend. Therefore, this disclosed embodiment can effectively improve the robustness and accuracy of component change anomaly detection. The component change anomaly detection method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to servers and personal computers (PCs). Optionally, in some embodiments, determining an abnormal server from the changed servers based on whether an abnormal event exceeding the baseline occurs in the changed server includes: obtaining a baseline value range corresponding to the abnormal event occurring in the changed server; determining whether an event attribute value of the abnormal event occurring in the changed server exceeds the corresponding baseline value range; if so, determining that the abnormal event is an abnormal event exceeding the baseline, and determining the abnormal server from the changed servers where the abnormal event exceeding the baseline occurs.Specifically, the benchmark value range corresponding to an abnormal event can be a range within which the event attribute values ​​of the abnormal event typically fall, obtained through statistical analysis of the event attribute values ​​of the abnormal event occurring in one or more historical periods. The following describes the process of obtaining the benchmark value range: Typically, there is an association between component changes and abnormal events. That is, when a component change occurs, there is a high probability that one or more abnormal events will occur in the server. Therefore, in the disclosed embodiments, historical data can be used to pre-establish an association between component changes and abnormal events. Furthermore, for abnormal events associated with component changes, a corresponding benchmark value range can be pre-statistically obtained. In this way, when determining whether an abnormal event exceeds the benchmark, the benchmark value range obtained in advance can be used to determine whether the abnormal event corresponds to a component change. Therefore, optionally, in some embodiments, obtaining a baseline value range corresponding to an abnormal event occurring in a changed server may include: determining, based on a pre-acquired association relationship between component changes and abnormal events, whether the abnormal event occurring in the changed server is an associated abnormal event associated with the component change; if so, obtaining a pre-calculated baseline value range corresponding to the associated abnormal event. Furthermore, in some embodiments, establishing an association relationship between component changes and abnormal events may include: obtaining first and second historical attribute values ​​of each abnormal event occurring during a historical component change process; wherein the first historical attribute value is an attribute value of an abnormal event occurring in the server before the component change; and the second historical attribute value is an attribute value of an abnormal event occurring in the server after the component change; performing variance analysis on the first and second historical attribute values ​​of each abnormal event to obtain an associated abnormal event associated with the component change, thereby establishing an association relationship between the component change and the abnormal event. Optionally, in some embodiments, calculating a baseline value range corresponding to an associated abnormal event may include: generating a baseline value range for the associated abnormal event based on the first and second historical attribute values ​​of the associated abnormal event. Specifically, when the event attribute values ​​of an abnormal event differ significantly before and after a component change, this indicates an association between the component change and the abnormal event. Therefore, in the above process, after obtaining the characteristic attribute values ​​(first historical attribute values) of the abnormal event before the component change and the characteristic attribute values ​​(second historical attribute values) of the abnormal event after the component change, variance analysis can be used to determine whether the difference in the event attribute values ​​of the abnormal event before and after the component change is significant (whether it is significant), thereby establishing an association between the component change and the abnormal event.Furthermore, a component change operation is more likely to result in an associated abnormal event in the server. Therefore, after the component change is performed on the server, the associated abnormal event is also more likely to occur in the server. In the disclosed embodiments, exceeding-baseline abnormal event determination is performed based on the baseline value range corresponding to the abnormal event. Therefore, after the association relationship is established, a baseline value range for generating the associated abnormal event can be pre-established for the associated abnormal event for use in subsequent exceeding-baseline abnormal event determination. Specifically, a statistical analysis can be performed on the first and second historical attribute values ​​to obtain the baseline value range for the associated abnormal event. In the disclosed embodiments, the specific method for statistical analysis is not limited and can be customized based on actual circumstances. For example, a box plot can be used to perform statistical analysis on the first and second historical attribute values ​​to obtain the baseline value range. Alternatively, a normal distribution can be used to perform statistical analysis on the first and second historical attribute values ​​to obtain the baseline value range. Optionally, in some embodiments, after determining whether the abnormal event occurring in the changed server is an associated abnormal event associated with a component change, the method further includes: if the abnormal event occurring in the changed server is an abnormal event not associated with the component change, obtaining historical attribute values ​​of the non-associated abnormal event within a first preset time period before the component change as third historical attribute values; and generating a benchmark value range corresponding to the non-associated abnormal event based on the third historical attribute values. Specifically, as described above, for an associated abnormal event associated with a component change, a benchmark value range corresponding to the associated abnormal event can be pre-generated based on the first historical attribute values ​​and the second historical attribute values ​​during the process of establishing the association relationship. For non-associated abnormal events, in embodiments of the present disclosure, statistical analysis can be performed based on the historical attribute values ​​(third historical attribute values) of the non-associated abnormal event within the first preset time period before the component change to obtain a benchmark value range for the non-associated abnormal event. In embodiments of the present disclosure, the specific method for performing the statistical analysis is not limited and can be customized based on actual circumstances. For example, a box plot method may be used to perform statistical analysis on the third historical attribute values ​​to obtain the aforementioned benchmark value range. Alternatively, a normal distribution method may be used to perform statistical analysis on the third historical attribute values ​​to obtain the corresponding benchmark value range.Furthermore, in the embodiments of the present disclosure, the process of identifying abnormal servers from modified servers experiencing abnormal events exceeding the baseline can include: determining the modified server experiencing the abnormal event exceeding the baseline as an abnormal server; or further screening the modified servers experiencing the abnormal event exceeding the baseline to identify abnormal servers. In the embodiments of the present disclosure, the specific screening rules are not limited. For example, abnormal servers can be screened based on the number of types of abnormal events exceeding the baseline, such as determining modified servers experiencing N or more abnormal events as abnormal servers. Abnormal servers can also be screened based on the degree to which event attribute values ​​of abnormal events exceeding the baseline exceedance range. Furthermore, abnormality assessments of abnormal events exceeding the baseline can be performed using other dimensions, and abnormal servers can be screened based on the resulting assessment values. Optionally, in some embodiments, the process of identifying an abnormal server from among modified servers experiencing an abnormal event exceeding baseline may include: obtaining multiple evaluation dimension information corresponding to the abnormal event exceeding baseline, as well as weights for each evaluation dimension; performing a weighted summation based on each evaluation dimension information and its weights to obtain an abnormality assessment value for the abnormal event exceeding baseline; determining a target abnormal event from among the abnormal events exceeding baseline based on the abnormality assessment value; and determining the modified server experiencing the target abnormal event as an abnormal server. Specifically, in embodiments of the present disclosure, the specific content of the evaluation dimension information corresponding to the abnormal event exceeding baseline may be customized based on actual circumstances. For example, the evaluation dimension information may include at least one of the following: the time difference between component modification and the first occurrence of the abnormal event exceeding baseline, the number of occurrences of the abnormal event exceeding baseline within a preset statistical window, the duration of the abnormal event exceeding baseline within the preset statistical window, the impact caused by the abnormal event exceeding baseline, the proportion of Elastic Compute Service (ECS) instances among the servers experiencing the abnormal event exceeding baseline, and the severity level of the abnormal event exceeding baseline. Furthermore, in the embodiment of the present disclosure, the analytic hierarchy process (AHP) may be used to calculate the weight values ​​of the above-mentioned evaluation dimension information, and a weighted sum is performed based on the obtained weight values ​​to obtain the abnormality evaluation value of the abnormal event exceeding the benchmark.In the above-described embodiments of the present disclosure, the abnormality degree of a benchmark abnormal event is evaluated from different evaluation dimensions using multiple evaluation dimension information corresponding to the benchmark abnormal event. This results in an abnormality assessment value that more accurately reflects the abnormality degree of the benchmark abnormal event. Furthermore, based on the abnormality assessment values ​​of each benchmark abnormal event, a target abnormal event is identified from the benchmark abnormal events, with the target abnormal event being considered as an abnormal event with a higher abnormality degree. Subsequently, a modified server experiencing the abnormal event with a higher abnormality degree is identified as an abnormal server with impaired performance. This improves the accuracy of abnormal server detection. Furthermore, when evaluating the abnormality degree of the benchmark abnormal event, the AHP analysis method is used to assign different weights to different dimensional information. This allows for a focus on high-priority dimensional information while also considering less likely dimensional information, further enhancing the comprehensiveness and accuracy of the abnormality assessment. Optionally, in some embodiments, the method further includes: clustering each abnormal event exceeding the baseline that occurs on each abnormal server to obtain a clustering result; the clustering result includes cluster category information and information about the number of abnormal servers included in each cluster category; and triggering a component change interception operation when the number of abnormal servers included in a cluster category exceeds a preset threshold. Specifically, an abnormal component change that has been determined to be an abnormal change may also cause a large number of servers in the cloud computing platform to simultaneously experience abnormal events of the same type. In this case, the abnormal component change can be said to have generated a clustered change risk for the cloud computing platform. The occurrence of a clustered change risk can have serious consequences. Therefore, when determining that an abnormal change has occurred, it is further determined whether the abnormal change will generate a clustered change risk, thereby promptly discovering and intercepting the risk. In the disclosed embodiments, by clustering each abnormal event exceeding the baseline that occurs on each abnormal server in the server cluster that has been determined to have performance impairment, multiple categories are obtained, as well as the number of abnormal servers included in each category. When a large number of abnormal servers in a certain category are identified, abnormal component changes are considered to have created a clustered change risk for the cloud computing platform. This triggers a component change interception operation. In the disclosed embodiments, the specific content of the interception operation is not limited and can be customized based on actual circumstances. For example, a clustered change risk notification can be sent to inform publishers of which component release has been identified as having a clustered change risk, as well as how many servers that have been modified have experienced a certain type of severe abnormal event (component release) as a result of the component release. This facilitates publishers to implement corresponding component change interception measures.Optionally, in some embodiments, each abnormal event is provided with tag information; clustering the abnormal events exceeding the baseline that occur on each abnormal server to obtain a clustering result includes: performing similarity calculations on the abnormal events exceeding the baseline that occur on each abnormal server based on the tag information to obtain a similarity calculation result; and clustering the abnormal events exceeding the baseline that occur on each abnormal server based on the similarity calculation result to obtain a clustering result. Specifically, in the embodiments of the present disclosure, the specific content of the tag information is not limited and can be customized according to actual circumstances. For example, the tag information may include at least one of the following: associated product form information, event category information, and event text information. The associated product form information may include information about the product associated with the abnormal event, such as whether the product is a graphics processing unit (GPU) or a CPU. Exemplarily, the event category information may include network abnormal events, motherboard abnormal events, disk abnormal events, and so on. In embodiments of the present disclosure, label information is pre-set for abnormal events. Furthermore, during clustering, clustering can be performed based on the label information in addition to clustering based on the text data describing the abnormal events. This compensates for the limited flexibility of text clustering and effectively improves the robustness of clustering of abnormal events exceeding the baseline. Optionally, in some embodiments, before clustering the abnormal events occurring on each abnormal server, the following steps may be performed: obtaining attribute values ​​of each abnormal event occurring on the abnormal server within a second preset time period before and after a component change; calculating attribute value fluctuation parameters for each abnormal event occurring on the abnormal server within the second preset time period; the attribute value fluctuation parameters characterizing the degree of fluctuation of the attribute values ​​of the abnormal event; determining key abnormal events from each abnormal event occurring on the abnormal server based on the attribute value fluctuation parameters; and clustering the abnormal events occurring on each abnormal server, including clustering the key abnormal events occurring on each abnormal server. Specifically, embodiments of the present disclosure do not limit the specific calculation method employed for calculating the attribute value fluctuation parameters, which may be customized based on actual circumstances. For example, an algorithm such as Squeeze (a type of root cause location algorithm) may be used to calculate the above attribute value fluctuation parameters.In the above-described embodiment of the present disclosure, after identifying an abnormal event exceeding the baseline, the attribute value fluctuation parameters of each abnormal event exceeding the baseline within a second preset time period before and after the component change are calculated. Furthermore, key abnormal events are identified based on these attribute value fluctuation parameters. Subsequently, clustering is performed based on these key abnormal events. Because the attribute value fluctuation parameters characterize the degree of fluctuation in the attribute values ​​of an abnormal event exceeding the baseline, the abnormal event that reflects the criticality of the current component change, i.e., the key abnormal event, can be identified from among the abnormal events exceeding the baseline based on these attribute value fluctuation parameters. In other words, the present embodiment analyzes the attribute value fluctuation parameters of each abnormal event exceeding the baseline that occurs on a normal server to identify the root cause of the server being identified as an abnormal server, i.e., the primary abnormal event that led to the server being identified as an abnormal server. See Figure 2, which is a schematic diagram of the detection process corresponding to the embodiment shown in Figure 1. The following describes the component change anomaly detection method according to the above-described embodiment of the present disclosure in conjunction with Figure 2. The detection process can be divided into two parts: the first part, the offline calculation of historical data; and the second part, the real-time component change anomaly detection. The first part, offline calculation of historical data, specifically involves obtaining historical component change information and the corresponding abnormal events that occurred on the changed servers. Based on this information, offline calculations are performed to determine the correlation between component changes and abnormal events, as well as the baseline value range corresponding to the associated abnormal events.The second part, real-time component change anomaly detection, specifically includes: obtaining real-time component change information and abnormal events occurring on the corresponding changed servers; calculating the spatial distribution value of the abnormal event based on the obtained information as a spatial feature, and determining whether the real-time component change is an abnormal change; if it is determined to be an abnormal change, then determining whether the associated abnormal event is an abnormal event exceeding the baseline based on the association relationship calculated in the first part, the benchmark value range corresponding to the associated abnormal event, and the event attribute value of the abnormal event; or, based on a real-time orthogonal method, calculating the benchmark value range corresponding to the non-associated abnormal event, and then determining whether the non-associated abnormal event is an abnormal event exceeding the baseline based on the benchmark value range corresponding to the non-associated abnormal event and the event attribute value of the abnormal event; In addition, for component changes detected as abnormal changes (suspected anomalies in Figure 2), root cause analysis can be performed based on the attribute value fluctuation parameters of the abnormal event exceeding the baseline, that is, determining the key abnormal events reflecting the current component change problem based on the attribute value fluctuation parameters of the abnormal event exceeding the baseline; and clustering the key abnormal events in each abnormal server. This generates a clustering result containing category information and the number of abnormal servers in each category. Based on the clustering result, a determination is then made as to whether a cluster anomaly exists (whether it causes a clustered change risk). If the current component change is determined to be a cluster anomaly, a notification can be sent to the component publisher to initiate a corresponding circuit-breaking operation, or the component change can be tagged to form component change history data for subsequent use. Example 2: Referring to FIG3 , FIG3 is a flowchart of the steps of a component change anomaly detection method according to Example 2 of the present disclosure. The component change anomaly detection method shown in FIG3 can be applied in the following scenario: When a specific component change is made to a server in a cloud computing platform, anomaly detection can be performed using the method shown in FIG3 to detect whether the specific component change is an abnormal change. If the specific component change is detected as an abnormal change, the abnormal server whose performance was compromised by the abnormal change is further identified. An abnormal change notification containing abnormal server identification information is then returned to inform the server that the specific component change is an abnormal change and which specific servers have had their performance compromised by the abnormal change. Specifically, the component change anomaly detection method provided in this embodiment includes the following steps: Step 302: Receive a component change notification and obtain abnormal events that occur on the changed server in the cloud computing platform after the component change. Step 304: Calculate the spatial distribution value of the abnormal events and, based on the spatial distribution value, determine whether the component change is an abnormal change.In step 306, if the change is an abnormal change, the abnormal server is identified from the changed servers based on whether the changed server has experienced an abnormal event exceeding the baseline. In step 308, an abnormal change notification is returned, which includes the identification information of the abnormal server. The component change anomaly detection solution provided in the disclosed embodiments, upon receiving a change notification regarding a specific component change, obtains abnormal events that occurred on the changed server in the cloud computing platform after the component change. The spatial distribution value of the abnormal events is then calculated at the server cluster spatial domain level. Component change anomaly detection is then performed based on this spatial distribution value. Furthermore, if the component change is detected as an abnormal change, a benchmark value range for the abnormal event attribute value is determined based on the historical performance of abnormal events at the temporal domain level. The abnormal event exceeding the baseline and the abnormal server are then determined based on this benchmark value range. The disclosed embodiments reduce reliance on manually pre-set rules during component change anomaly detection. Instead, they use the spatial and historical performance of abnormal events as criteria for determining anomalies at the current stage. In other words, they primarily use the historical performance of abnormal events to predict the current trend. Therefore, the disclosed embodiments can effectively improve the robustness and accuracy of component change anomaly detection. The component change anomaly detection method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to servers and PCs. For the third embodiment, refer to FIG4 , which is a flowchart of the steps of the component change anomaly detection method according to the third embodiment of the disclosed embodiment. Because there are many types of computing components, and a specific type of computing component may also undergo a high number of version changes, a server cluster in a cloud computing platform may experience a series of component change processes. Therefore, the application scenario for component change anomaly detection shown in FIG4 can be: during the component change process, component change log information recording various component changes is continuously acquired. Accordingly, abnormal event log information for the server cluster in the cloud computing platform is also synchronously acquired. Furthermore, using the methods provided in the embodiments of the present disclosure, anomaly detection is performed on each component change recorded in the component change log information. When an abnormal change is detected, identification information of the abnormal change and identification information of the abnormal server corresponding to the abnormal change are returned. Specifically, the component change anomaly detection method provided in this embodiment includes the following steps: Step 402: Receive component change log information and abnormal event log information for the server cluster in the cloud computing platform. The component change log information is used to record various component changes made to the servers in the cloud computing platform.Step 404: For each component change recorded in the component change log, the corresponding abnormal event information segment is determined from the abnormal event log. The abnormal event information segment corresponding to each component change describes an abnormal event that occurred on the changed server during the corresponding component change. Step 406: Based on the abnormal event information segment, the spatial distribution value of the abnormal events corresponding to each component change is calculated. Based on the spatial distribution value, abnormal changes are determined from each component change. Step 408: For abnormal changes, the abnormal server is determined from the changed servers based on whether the changed server has experienced an abnormal event exceeding the baseline. Step 410: An abnormal change notification is returned, which includes the identification information of the abnormal change and the identification information of the abnormal server corresponding to the abnormal change. The component change anomaly detection solution provided by the disclosed embodiments, during the process of performing a series of component changes on a server cluster within a cloud computing platform, obtains component change log information recording various component changes and abnormal event log information for the server cluster within the cloud computing platform. Based on this information, the solution performs anomaly detection on each of the component changes, and provides identification information for each detected abnormal change and the identification information of the abnormal server corresponding to the abnormal change. In the above embodiment, when performing anomaly detection on each component change, the spatial distribution value of abnormal events is calculated at the spatial domain level of the server cluster, and component change anomaly detection is performed based on this spatial distribution value. Furthermore, when an abnormal change is detected, a baseline value range for the abnormal event attribute value is determined based on the historical performance of the abnormal event at the temporal level. Based on this baseline value range, abnormal events exceeding the baseline and abnormal servers are identified. The disclosed embodiments reduce reliance on manually pre-set rules during component change anomaly detection. Instead, they use the spatial and historical performance of abnormal events as criteria for determining anomalies at the current stage. In other words, they primarily use the historical performance of abnormal events to predict the current trend. Therefore, the disclosed embodiments can effectively improve the robustness and accuracy of component change anomaly detection. The component change anomaly detection method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to servers and PCs. Example 4: Figure 5 is a block diagram of a component change anomaly detection device according to Example 4 of the disclosed embodiments.The component change anomaly detection device provided by the embodiment of the present disclosure includes: a statistical module 502, which is used to count abnormal events that have occurred in changed servers that have undergone component changes in a server cluster of a cloud computing platform; an abnormal change determination module 504, which is used to calculate the spatial distribution value of the abnormal events and determine whether the component change is an abnormal change based on the spatial distribution value; an abnormal server determination module 506, which is used to determine an abnormal server from the changed servers based on whether an abnormal event exceeding the baseline has occurred in the changed servers if the change is an abnormal change; an abnormal event exceeding the baseline is an abnormal event in which the event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values ​​of abnormal events that have occurred in a historical stage. Optionally, in some embodiments, when performing the step of determining an abnormal server from the changed servers based on whether an abnormal event exceeding a baseline has occurred on the changed server, the abnormal server determination module 506 is specifically configured to: obtain a baseline value range corresponding to the abnormal event occurring on the changed server; determine whether the event attribute value of the abnormal event occurring on the changed server exceeds the corresponding baseline value range; if so, determine that the abnormal event is an abnormal event exceeding a baseline, and determine the abnormal server from the changed servers that have experienced the abnormal event exceeding a baseline. Optionally, in some embodiments, when performing the step of obtaining a baseline value range corresponding to the abnormal event occurring on the changed server, the abnormal server determination module 506 is specifically configured to: determine whether the abnormal event occurring on the changed server is an associated abnormal event associated with the component change based on a pre-acquired association between the component change and the abnormal event; and if so, obtain a pre-calculated baseline value range corresponding to the associated abnormal event. Optionally, in some embodiments, the component change anomaly detection device further includes: an association relationship establishment module, configured to obtain a first historical attribute value and a second historical attribute value of each abnormal event occurring during a component change process in a historical stage; wherein the first historical attribute value is an attribute value of an abnormal event occurring in the server before the component change; and the second historical attribute value is an attribute value of an abnormal event occurring in the server after the component change; performing variance analysis on the first historical attribute value and the second historical attribute value of each abnormal event to obtain an associated abnormal event associated with the component change, so as to establish an association relationship between the component change and the abnormal event; and a baseline value range calculation module, configured to generate a baseline value range for the associated abnormal event based on the first historical attribute value and the second historical attribute value of the associated abnormal event.Optionally, in some embodiments, the abnormal server determination module 506 is further configured to: after determining whether the abnormal event occurring in the changed server is an associated abnormal event associated with a component change, if the abnormal event occurring in the changed server is an abnormal event not associated with the component change, obtain historical attribute values ​​of the non-associated abnormal event within a first preset time period before the component change as a third historical attribute value; and generate a baseline value range corresponding to the non-associated abnormal event based on the third historical attribute value. Optionally, in some embodiments, when determining an abnormal server from the changed servers in which an abnormal event exceeding a baseline occurred, the abnormal server determination module 506 is further configured to: obtain multiple evaluation dimension information corresponding to the abnormal event exceeding a baseline, and a weight value for each evaluation dimension; perform a weighted sum based on each evaluation dimension information and the weight value for each evaluation dimension information to obtain an abnormality assessment value for the abnormal event exceeding a baseline; determine a target abnormal event from the abnormal event exceeding a baseline based on the abnormality assessment value; and determine the changed server in which the target abnormal event occurred as an abnormal server. Optionally, in some embodiments, the component change anomaly detection apparatus further includes: a clustering module configured to cluster each abnormal event exceeding a baseline occurring on each abnormal server to obtain a clustering result; the clustering result includes cluster category information and information about the number of abnormal servers included in each cluster category; and triggering a component change interception operation when the number of abnormal servers included in a cluster category exceeds a preset threshold. Optionally, in some embodiments, each abnormal event is provided with label information; the clustering module is specifically configured to: perform similarity calculation on each abnormal event exceeding a baseline occurring on each abnormal server based on the label information to obtain a similarity calculation result; and perform clustering on each abnormal event exceeding a baseline occurring on each abnormal server based on the similarity calculation result to obtain a clustering result. Optionally, in some embodiments, the component change anomaly detection device further includes: a key abnormal event determination module, which is used to obtain the attribute values ​​of each super-baseline abnormal event occurring on each abnormal server within a second preset time period before and after the component change before clustering the super-baseline abnormal events occurring on each abnormal server; calculate the attribute value fluctuation parameter of each super-baseline abnormal event within the second preset time period; the attribute value fluctuation parameter represents the degree of fluctuation of the attribute value of the super-baseline abnormal event; based on the attribute value fluctuation parameter, determine the key abnormal event from each super-baseline abnormal event occurring on the abnormal server; and a clustering module, which is specifically used to: perform a clustering operation on each key abnormal event occurring on each abnormal server to obtain a clustering result.The component change anomaly detection device of this embodiment is used to implement the component change anomaly detection method described in the first embodiment above and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here. Furthermore, the functional implementation of each module in the component change anomaly detection device of this embodiment can be referenced to the corresponding descriptions in the method embodiment above and will not be described in detail here. Example 5 FIG6 is a block diagram of a component change anomaly detection device according to the fifth embodiment of the present disclosure. The component change anomaly detection device provided in this embodiment includes: a notification receiving module 602 for receiving component change notifications and obtaining abnormal events that occur on the changed servers in the cloud computing platform after the component change; a first detection module 604 for calculating the spatial distribution value of the abnormal events and, based on the spatial distribution value, determining whether the component change is an abnormal change; if it is an abnormal change, identifying the abnormal server from the changed servers based on whether the changed servers have abnormal events exceeding the baseline; and a first notification returning module 606 for returning an abnormal change notification, which includes identification information of the abnormal server. The component change anomaly detection device of this embodiment is used to implement the component change anomaly detection method described in the second embodiment above and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here. Furthermore, the functional implementation of each module in the component change anomaly detection device of this embodiment can be referenced to the corresponding descriptions in the aforementioned method embodiments, and will not be repeated here. Example Six: Figure 7 is a block diagram of a component change anomaly detection device according to Example Six of the present disclosure. The component change anomaly detection device provided in this embodiment includes: a log information receiving module 702, configured to receive component change log information and abnormal event log information for a server cluster of a cloud computing platform; a second detection module 704, configured to determine, from the abnormal event log information, corresponding abnormal event information segments for each component change recorded in the component change log information; calculate, based on the abnormal event information segments, the spatial distribution value of abnormal events corresponding to each component change, and determine abnormal changes from each component change based on the spatial distribution value; for abnormal changes, determine abnormal servers from among the changed servers based on whether the changed servers experience abnormal events exceeding the baseline; and a second notification returning module 706, configured to return an abnormal change notification, the abnormal change notification including identification information of the abnormal change and identification information of the abnormal server corresponding to the abnormal change. The component change anomaly detection device of this embodiment is used to implement the corresponding component change anomaly detection method in the aforementioned embodiment 3, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.Furthermore, the functional implementation of each module in the component change anomaly detection device of this embodiment can refer to the corresponding descriptions in the aforementioned method embodiments and will not be repeated here. With reference to FIG8 , Embodiment 7 shows a schematic structural diagram of an electronic device according to Embodiment 7 of the present disclosure. The present disclosure does not limit the specific implementation of the electronic device. As shown in FIG8 , the electronic device may include a processor 802, a communications interface 804, a memory 806, and a communication bus 808. The processor 802, communications interface 804, and memory 806 communicate with each other via a communications bus 808. The communications interface 804 is used to communicate with other electronic devices. The processor 802 is configured to execute a program 810, which may specifically perform the relevant steps in the aforementioned component change anomaly detection method embodiments. Specifically, the program 810 may include program code, which includes computer operating instructions. The processor 802 may be a CPU, an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present disclosure. The smart device includes one or more processors, which may be of the same type, such as one or more CPUs, or different types, such as one or more CPUs and one or more ASICs. Memory 806 is used to store program 810. Memory 806 may include high-speed RAM or non-volatile memory, such as at least one disk drive. Program 810 may include multiple computer instructions. Specifically, program 810 may cause processor 802 to perform operations corresponding to the component change anomaly detection method described in any of the aforementioned method embodiments. The specific implementation of each step in program 810 can be found in the descriptions of the corresponding steps and units in the aforementioned method embodiments, and corresponding beneficial effects are achieved, so detailed description is omitted here. Those skilled in the art will clearly understand that, for ease and brevity of description, the specific operating processes of the devices and modules described above can be found in the descriptions of the corresponding processes in the aforementioned method embodiments, and detailed description is omitted here. The present disclosure also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the method described in any of the aforementioned method embodiments.The computer storage medium includes, but is not limited to, a compact disc read-only memory (CD-ROM), random access memory (RAM), a floppy disk, a hard disk, or a magneto-optical disk. The disclosed embodiments also provide a computer program product comprising computer instructions that instruct a computing device to perform operations corresponding to any of the component change anomaly detection methods described in the aforementioned method embodiments. Furthermore, it should be noted that all user-related information (including, but not limited to, user device information, user personal information, etc.) and data (including, but not limited to, sample data used for model training, data used for analysis, stored data, and displayed data, etc.) involved in the disclosed embodiments are authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for the user to choose to authorize or reject. It should be noted that, depending on implementation needs, the various components / steps described in the embodiments of this disclosure may be split into more components / steps, or two or more components / steps or partial operations of components / steps may be combined into new components / steps to achieve the objectives of the embodiments of this disclosure. The methods according to the embodiments of this disclosure described above may be implemented in hardware or firmware, or as software or computer code that can be stored on a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or non-transitory machine-readable medium downloaded via a network and then stored on a local recording medium. Thus, the methods described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)). It will be understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., random access memory (RAM), read-only memory (ROM), flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods described herein are implemented.Furthermore, when a general-purpose computer accesses the code for implementing the methods described herein, the execution of the code transforms the general-purpose computer into a special-purpose computer for executing the methods described herein. Those skilled in the art will appreciate that the units and method steps described in the various examples in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals skilled in the art may use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of the present disclosure. The above embodiments are intended only to illustrate the embodiments of the present disclosure and are not intended to limit them. Persons skilled in the relevant technical fields may make various changes and modifications without departing from the spirit and scope of the embodiments of the present disclosure. Therefore, all equivalent technical solutions are also within the scope of the embodiments of the present disclosure, and the scope of patent protection for the embodiments of the present disclosure shall be defined by the claims.

Claims

Claims 1. A component change anomaly detection method, comprising: Collect statistics on abnormal events that occur on servers that have undergone component changes in the server cluster of the cloud computing platform; calculating a spatial distribution value of the abnormal event, and determining whether the component change is an abnormal change based on the spatial distribution value; If it is the abnormal change, determining the abnormal server from the changed servers based on whether an abnormal event exceeding a baseline occurs on the changed servers; The super-baseline abnormal event is an abnormal event in which the event attribute value exceeds the corresponding benchmark value range; The benchmark value range is obtained based on the attribute values of abnormal events that occurred in historical stages.

2. The method according to claim 1, wherein: The method of determining an abnormal server from the changed servers based on whether an abnormal event exceeding a baseline occurs on the changed server includes: obtaining a baseline value range corresponding to the abnormal event occurring on the changed server; determining whether an event attribute value of the abnormal event occurring on the changed server exceeds the corresponding baseline value range; if so, determining that the abnormal event is an abnormal event exceeding a baseline, and determining the abnormal server from the changed servers where the abnormal event exceeding the baseline occurs.

3. The method according to claim 2, wherein: The obtaining of a benchmark value range corresponding to the abnormal event occurring in the changed server includes: determining, based on a pre-obtained association relationship between component changes and abnormal events, whether the abnormal event occurring in the changed server is an associated abnormal event associated with the component change; and if so, obtaining a pre-calculated benchmark value range corresponding to the associated abnormal event.

4. The method according to claim 3, wherein: The process of establishing the association relationship between the component change and the abnormal event includes: obtaining the first historical attribute value and the second historical attribute value of each abnormal event that occurred during the component change process in the historical stage; wherein the first historical attribute value is the attribute value of the abnormal event that occurred in the server before the component change; the second historical attribute value is the attribute value of the abnormal event that occurred in the server after the component change; performing variance analysis on the first historical attribute value and the second historical attribute value of each abnormal event to obtain associated abnormal events associated with the component change, so as to establish the association relationship between the component change and the abnormal event; the process of calculating the benchmark value range corresponding to the associated abnormal event includes: generating the benchmark value range of the associated abnormal event based on the first historical attribute value and the second historical attribute value of the associated abnormal event.

5. The method according to any one of claims 3 to 4, wherein: After determining whether the abnormal event occurring in the changed server is an associated abnormal event associated with the component change, the method further includes: if the abnormal event occurring in the changed server is an abnormal event unrelated to the component change, obtaining a historical attribute value of the unrelated abnormal event within a first preset time period before the component change is performed as a third historical attribute value; Based on the third historical attribute value, a reference value range corresponding to the unrelated abnormal event is generated.

6. The method according to any one of claims 2 to 5, wherein: The method of determining an abnormal server from the changed servers where the super-baseline abnormal event occurs includes: obtaining multiple evaluation dimension information corresponding to the super-baseline abnormal event, and a weight value of each evaluation dimension information; performing a weighted sum based on each evaluation dimension information and the weight value of each evaluation dimension information to obtain an abnormal evaluation value of the super-baseline abnormal event; determining a target abnormal event from the super-baseline abnormal event based on the abnormal evaluation value; and determining the changed server where the target abnormal event occurs as an abnormal server.

7. The method according to any one of claims 1 to 6, wherein: The method further includes: clustering each abnormal event exceeding the baseline occurring on each abnormal server to obtain a clustering result; the clustering result includes cluster category information and information about the number of abnormal servers included in each cluster category; and triggering a component change interception operation when the number of abnormal servers included in a cluster category exceeds a preset number threshold.

8. The method according to claim 7, wherein: Each abnormal event is provided with label information; the clustering operation on each abnormal event exceeding the baseline that occurs on each abnormal server to obtain a clustering result includes: performing a similarity calculation on each abnormal event exceeding the baseline that occurs on each abnormal server based on the label information to obtain a similarity calculation result; and performing a clustering operation on each abnormal event exceeding the baseline that occurs on each abnormal server based on the similarity calculation result to obtain a clustering result.

9. The method according to any one of claims 7 to 8, wherein: Before clustering the super-benchmark abnormal events occurring in the abnormal servers, the method further includes: obtaining the attribute values of the super-benchmark abnormal events occurring in the abnormal servers within a second preset time period before and after the component change; calculating the attribute value fluctuation parameters of the super-benchmark abnormal events within the second preset time period; the attribute value fluctuation parameters characterize the degree of fluctuation of the attribute values of the super-benchmark abnormal events; based on the attribute value fluctuation parameters, determining key abnormal events from the super-benchmark abnormal events occurring in the abnormal servers; the clustering operation on the super-benchmark abnormal events occurring in the abnormal servers includes: clustering the key abnormal events occurring in the abnormal servers.

10. A component change anomaly detection method, comprising: Receive component change notifications and obtain abnormal events that occur on the changed servers in the cloud computing platform after the component changes; Calculating the spatial distribution value of the abnormal event, and determining whether the component change is an abnormal change based on the spatial distribution value; If the change is abnormal, the abnormal server is determined from the changed servers based on whether an abnormal event exceeding the baseline occurs on the changed server. The abnormal event exceeding the baseline is an abnormal event in which the event attribute value exceeds the corresponding baseline value range. The baseline value range is obtained based on the attribute values of abnormal events that occurred in the historical period. Returns an abnormal change notification, wherein the abnormal change notification includes identification information of the abnormal server.

11. A component change anomaly detection method, comprising: Receive component change log information and abnormal event log information for the server cluster of the cloud computing platform; For each component change recorded in the component change log information, determining a corresponding abnormal event information segment from the abnormal event log information; Calculating a spatial distribution value of abnormal events corresponding to each component change based on the abnormal event information fragments, and determining an abnormal change from each component change based on the spatial distribution value; For the abnormal change, based on whether an abnormal event exceeding a baseline occurs on the changed server, an abnormal server is determined from the changed servers; the abnormal event exceeding a baseline is an abnormal event in which an event attribute value exceeds a corresponding baseline value range; the baseline value range is obtained based on the attribute values of abnormal events that occurred in a historical period; Returns an abnormal change notification, where the abnormal change notification includes identification information of the abnormal change and identification information of the abnormal server corresponding to the abnormal change.

12. An electronic device comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store at least one execution instruction, where the executable instruction enables the processor to perform an operation corresponding to the method according to any one of claims 1 to 11.

13. A computer storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.

14. A computer program product comprising computer instructions, wherein the computer instructions instruct a computing device to perform operations corresponding to the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Intelligent detection method and detection system for server exception of hybrid strategy

    CN111061620A

  • Running method and device of server cluster, equipment and storage medium

    CN115001956A

  • Abnormality detection method and device, equipment, storage medium and program

    CN115033453A

  • Anomaly detection method and cloud network platform

    CN115514620A