Component change anomaly detection method and device, electronic equipment and medium

By calculating the spatial distribution value and historical performance of abnormal events in the cloud computing platform and determining the abnormal events beyond the benchmark, the accuracy and robustness of component change abnormality detection in the prior art are solved, and timely discovery of abnormal servers and prediction of performance impaired trends are achieved.

CN120389961APending Publication Date: 2025-07-29HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410123859.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, component change abnormal detection heavily relies on manual experience setting rules, resulting in the rules being invalid when new types of components are changed, and it is impossible to effectively detect abnormal changes in cloud computing platforms.

Method used

By counting exception events that have been changed in the server cluster of cloud computing platform servers, the spatial distribution value of exception events is calculated, and the benchmark exception events are determined based on the reference value range of historical exception events, thereby determining the abnormal server and reducing dependence on manual preset rules.

Benefits of technology

Improves the robustness and accuracy of component change abnormality detection, and can promptly detect abnormal servers with damaged performance, reducing the potential risks of cluster change.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120389961A_ABST
    Figure CN120389961A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a component change anomaly detection method and device, electronic equipment and a storage medium. The component change anomaly detection method comprises the following steps: counting abnormal events occurring in a changed server subjected to component change in a server cluster of a cloud computing platform; calculating a spatial distribution value of the abnormal event, and determining whether the component change is an abnormal change based on the spatial distribution value; if the change is the abnormal change, determining an abnormal server from the changed servers based on whether the changed servers have a standard exceeding abnormal event or not; the exceeded benchmark abnormal event is an abnormal event of which the event attribute value exceeds the corresponding benchmark value range; the reference value range is obtained based on the attribute value of the abnormal event occurring in the historical stage. According to the embodiment of the invention, the robustness and accuracy of component change anomaly detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to a method, apparatus, electronic device, and storage medium for detecting abnormal component changes. Background Art

[0002] As the infrastructure of the Internet, the underlying layer of a cloud computing platform involves various computing components such as storage, network, CPU, and management and control. To meet the performance iteration and stability requirements of the cloud computing platform, the above-mentioned computing components will also be changed due to upgrades and iterations.

[0003] Abnormalities may occur during the change of computing components. For example, component changes may cause abnormal events to occur in a relatively large number of servers in the cloud computing platform. Abnormal component changes may cause serious consequences. Therefore, during the process of computing component changes, it is necessary to detect abnormal component changes, so as to timely discover abnormal changes and the abnormal servers with impaired performance in this abnormal change, so as to take effective measures to intercept them subsequently.

[0004] In the related art, it mainly relies on manual experience to set abnormal discovery rules. For example, for a certain component change, if 10 servers in the computing platform simultaneously experience the abnormal event of failed startup, it is determined that the component change is abnormal.

[0005] The above method seriously relies on manually preset rules. When new types of component changes occur, the originally set rules may not work. Therefore, there is an urgent need for a more effective component change anomaly detection solution. Summary of the Invention

[0006] In view of this, embodiments of the present application provide a component change anomaly detection solution to at least partially solve the above problems.

[0007] According to the first aspect of the embodiments of the present application, a method for detecting abnormal component changes is provided, including:

[0008] Counting abnormal events that occur in the changed servers in the server cluster of the cloud computing platform where component changes have been made;

[0009] Calculating the spatial distribution value of the abnormal events, and based on the spatial distribution value, determining whether the component change is an abnormal change;

[0010] If it is the abnormal change, based on whether the changed server has an over-benchmark abnormal event, determining the abnormal server from the changed servers; the over-benchmark abnormal event is an abnormal event whose event attribute value exceeds the corresponding benchmark value range; the benchmark value range is obtained based on the attribute values of abnormal events that occurred in the historical stage.

[0011] According to a second aspect of the embodiments of the present application, another method for detecting abnormal component changes is provided, including:

[0012] Receiving a component change notice, and obtaining abnormal events that occur in the changed servers in the cloud computing platform after the components are changed;

[0013] Calculating a spatial distribution value of the abnormal events, and determining whether the component change is an abnormal change based on the spatial distribution value;

[0014] If it is an abnormal change, determining abnormal servers from the changed servers based on whether the changed servers have ultra-baseline abnormal events; the ultra-baseline abnormal event is an abnormal event whose event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage;

[0015] Returning an abnormal change notice, where the abnormal change notice contains the identification information of the abnormal servers.

[0016] According to a third aspect of the embodiments of the present application, yet another method for detecting abnormal component changes is provided, including:

[0017] Receiving component change log information and abnormal event log information for the server cluster of the cloud computing platform;

[0018] For each component change recorded in the component change log information, determining a corresponding abnormal event information segment from the abnormal event log information;

[0019] Based on the abnormal event information segment, calculating a spatial distribution value of the abnormal events corresponding to each component change, and determining an abnormal change from each component change based on the spatial distribution value;

[0020] For the abnormal change, determining abnormal servers from the changed servers based on whether the changed servers have ultra-baseline abnormal events; the ultra-baseline abnormal event is an abnormal event whose event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage;

[0021] Returning an abnormal change notice, where the abnormal change notice contains the identification information of the abnormal change and the identification information of the abnormal servers corresponding to the abnormal change.

[0022] According to a fourth aspect of the embodiments of the present application, a device for detecting abnormal component changes is provided, including:

[0023] A statistics module, configured to count abnormal events that occur in the changed servers in the server cluster of the cloud computing platform where component changes have been made;

[0024] An abnormal change determination module, which calculates the spatial distribution value of the abnormal event and determines whether the component change is an abnormal change based on the spatial distribution value;

[0025] An abnormal server determination module, which, if it is an abnormal change, determines an abnormal server from the changed servers based on whether a super-baseline abnormal event occurs in the changed servers; the super-baseline abnormal event is an abnormal event whose event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage.

[0026] According to the fifth aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method described in any one of the first aspect to the third aspect.

[0027] According to the sixth aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the first aspect to the third aspect.

[0028] In the component change abnormal detection solution provided by the embodiments of the present application, after obtaining the abnormal events that occur in the changed services in the server cluster, the spatial distribution value of the abnormal events is calculated from the spatial domain level of the server cluster, and then the abnormal detection of the component change is performed based on the above spatial distribution value; in addition, when it is detected that the component change is an abnormal change, then from the time domain level, the baseline value range of the abnormal event attribute value is obtained according to the historical performance of the abnormal event, and then the super-baseline abnormal event and the abnormal server are determined based on the above baseline value range. In the embodiments of the present application, the dependence on artificially preset rules in the process of component change abnormal detection is weakened, but the spatial performance and historical performance of the abnormal event are used as the abnormal determination criteria in the current stage, or rather, the trend performance in the current stage is mainly predicted through the historical performance of the abnormal event. Therefore, the embodiments of the present application can effectively improve the robustness and accuracy of component change abnormal detection. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0030] Figure 1It is a flowchart of steps of a component change anomaly detection method according to Embodiment 1 of the present application;

[0031] Figure 2 It is Figure 1 a schematic diagram of the detection process corresponding to the illustrated embodiment;

[0032] Figure 3 It is a flowchart of steps of a component change anomaly detection method according to Embodiment 2 of the present application;

[0033] Figure 4 It is a flowchart of steps of a component change anomaly detection method according to Embodiment 3 of the present application;

[0034] Figure 5 It is a structural block diagram of a component change anomaly detection device according to Embodiment 4 of the present application;

[0035] Figure 6 It is a structural block diagram of a component change anomaly detection device according to Embodiment 5 of the present application;

[0036] Figure 7 It is a structural block diagram of a component change anomaly detection device according to Embodiment 6 of the present application;

[0037] Figure 8 It is a schematic diagram of the structure of an electronic device according to Embodiment 7 of the present application. Detailed implementation manners

[0038] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0039] Embodiment 1

[0040] Referring to Figure 1 , Figure 1 it is a flowchart of steps of a component change anomaly detection method according to Embodiment 1 of the present application.

[0041] Specifically, the component change anomaly detection method provided in this embodiment includes the following steps:

[0042] Step 102, count the abnormal events that occur in the changed servers that have undergone component changes in the server cluster of the cloud computing platform.

[0043] Specifically, when making component changes, a gray release method is usually adopted to release new components in the server cluster to test the performance of the new components. For example: The server cluster of the cloud computing platform contains 100 servers (i.e., physical machines). When making a change to a certain component, the changed component can be selectively used in some servers (such as 40 of them), while the unchanged component continues to be used in the remaining servers. The changed servers in this step can refer to the servers in the server cluster that have installed the new component, such as the 40 servers in the above example.

[0044] What is counted in this step is the abnormal events that occur in the servers that have installed the new component within a preset duration after the component change. Specifically, for each changed server in the server cluster, the abnormal events that occur in it can be counted separately.

[0045] In the embodiments of the present application, neither the type of the component that has changed nor the specific change content is limited. For example: It can be a storage component, a network component, a load balancing component, etc. in the server.

[0046] In the embodiments of the present application, the specific content of the abnormal event is not limited either. It can be preset according to the actual situation when the server shows what kind of situation to determine that the server has an abnormal event. For example: The following at least one situation can be set as an abnormal event: service start and stop, memory overflow, process unexpected exit, network packet loss, memory fragmentation, bandwidth mutation, network disconnection, downtime, restart failure, etc.

[0047] Step 104, calculate the spatial distribution value of the abnormal event, and based on the spatial distribution value, determine whether the component change is an abnormal change.

[0048] Among them, the spatial distribution value represents the ratio of the number of servers with abnormal events in the changed servers to the number of servers with abnormal events in the server cluster.

[0049] Specifically, for each abnormal event obtained through step 102, the spatial distribution value can be calculated respectively by this step, and the spatial distribution value corresponding to each abnormal event can be obtained.

[0050] For example, for the abnormal event A that occurs in the changed server, the number of servers in the changed server where the abnormal event A occurs can be obtained, assumed to be a. At the same time, the number of servers in the server cluster in the cloud computing platform where the abnormal event A occurs can be obtained, assumed to be b. Then, based on the value of a / b, the spatial distribution value corresponding to the abnormal event A can be obtained. For example, a / b can be determined as the spatial distribution value corresponding to the abnormal event A, or a certain function operation can be performed on a / b, and the obtained function value can be determined as the spatial distribution value corresponding to the abnormal event A.

[0051] The spatial distribution value of the abnormal event obtained through this step can be positively correlated with the ratio of the number of servers with abnormal events in the changed server to the number of servers with abnormal events in the server cluster, or negatively correlated with the above ratio.

[0052] When the two are positively correlated, the calculated spatial distribution value can be compared with the first preset spatial distribution threshold. When the spatial distribution value is greater than the first preset spatial distribution threshold, it is determined that the component change is an abnormal change. Among them, the first preset spatial distribution threshold can be custom-set according to experience, for example, taking the value of 0.9; when the two are negatively correlated, the calculated spatial distribution value can be compared with the second preset spatial distribution threshold. When the spatial distribution value is less than the second preset spatial distribution threshold, it is determined that the component change is an abnormal change. Among them, the second preset spatial distribution threshold can also be custom-set according to experience, for example, taking the value of 0.1.

[0053] Since the types of abnormal events that occur in the changed server can be one or multiple, correspondingly, the number of calculated spatial distribution values may also be one or multiple. When the number of spatial distribution values is one, the size relationship between the spatial distribution value and the preset spatial distribution threshold can be used to determine whether it is an abnormal change; when the number of spatial distribution values is multiple, the size relationship between each of the above multiple spatial distribution values and the preset spatial distribution threshold can be used to determine whether it is an abnormal change. For example, it can be set that when each spatial distribution value is greater than the preset spatial distribution threshold, it is determined as a normal change. On the contrary, when there are a preset number of spatial distribution values less than the preset spatial distribution threshold, it is determined as an abnormal change. The above preset number can be custom-set, for example, set as any natural number.

[0054] In theory, some abnormal event will "randomly" occur on a certain server at any moment. However, it is understandable that when an abnormal event associated with (caused by) component changes occurs, it should concentrate on those servers that have been changed. Therefore, in this application, by calculating the spatial distribution value representing the proportion of the changed servers where abnormal events occur in the server cluster, and then determining whether the component change is an abnormal change through this value. Compared with the method of determining whether it is an abnormal change simply by the number of changed servers where abnormal events occur, the above method can effectively improve the robustness and accuracy of component abnormal change detection.

[0055] Step 106, if it is an abnormal change, then determine the abnormal servers from the changed servers based on whether a super-benchmark abnormal event occurs on the changed servers.

[0056] Among them, the super-benchmark abnormal event is an abnormal event whose event attribute value exceeds the corresponding benchmark value range; the benchmark value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage, and the benchmark value range can be used to represent the normal value range of the event attribute values of the abnormal events.

[0057] Specifically, the abnormal servers can refer to the servers with impaired performance.

[0058] For the abnormal events that occur on the changed servers during the current component change process, they usually occurred with a high probability in the historical stage. However, since their event attribute values may be within a relatively normal range in the historical stage, the occurrence of these abnormal events in the historical stage did not cause significant damage to the performance of the abnormal servers. Therefore, in this step, for the abnormal events, based on the attribute values of these abnormal events that occurred in the historical stage, the above relatively normal range, that is, the benchmark value range in this step, can be obtained through statistical analysis (such as using box plots, normal distribution analysis, etc.); and then by comparing whether the time attribute value of the abnormal event during the current component change process is within the above benchmark value range, it is judged whether the current occurrence of this abnormal event has a tendency to cause server performance damage, that is, to determine whether this abnormal event is a super-benchmark abnormal event.

[0059] Specifically, when the event attribute value of the abnormal event exceeds the corresponding benchmark value range, it indicates that this abnormal event is a super-benchmark abnormal event, and the occurrence of this abnormal event during the current component change has a tendency to cause server performance damage. The changed servers where the above super-benchmark abnormal events occur are determined as abnormal servers with impaired performance.

[0060] In this step, by comparing the event attribute values of the abnormal events caused by this component change with the baseline value range of abnormal events obtained based on historical data, it is further determined among each abnormal event that there will be an over-baseline abnormal event that will cause the server performance to be damaged. Then, the changed server where the above-baseline abnormal event occurs is determined as an abnormal server. The above process uses the historical performance of abnormal events as the basis for determining the current trend, thereby improving the accuracy of abnormal server determination, that is, further improving the robustness and accuracy of anomaly detection.

[0061] In the embodiments of the present application, the specific content of the event attribute value of an abnormal event is not limited and can be customized according to the specific content of the abnormal event. For example, for abnormal events such as restart failure, network disconnection, service start and stop, the event attribute value may include: frequency of occurrence; for abnormal events such as bandwidth mutation, the event attribute value may include: frequency of occurrence, specific value of a single occurrence (specific bandwidth mutation amount in a single occurrence), etc.

[0062] The component change anomaly detection solution provided by the embodiment of the present application, after obtaining the abnormal event that has occurred in the changed service in the server cluster, calculates the spatial distribution value of the abnormal event from the spatial domain level of the server cluster, and then performs component change anomaly detection based on the above spatial distribution value; in addition, when the component change is detected to be an abnormal change, the benchmark value range of the abnormal event attribute value is obtained from the time domain level according to the historical performance of the abnormal event, and then the super-baseline abnormal event and the abnormal server are determined based on the above benchmark value range. The embodiment of the present application weakens the dependence on manually preset rules in the component change anomaly detection process, and instead uses the spatial performance and historical performance of the abnormal event as the abnormality judgment standard for the current stage, or in other words, mainly predicts the trend performance of the current stage through the historical performance of the abnormal event. Therefore, the embodiment of the present application can effectively improve the robustness and accuracy of component change anomaly detection.

[0063] The component change anomaly detection method of this embodiment can be executed by any appropriate electronic device with data processing capabilities, including but not limited to: a server and a PC.

[0064] Optionally, in some embodiments, determining an abnormal server from the changed servers based on whether an abnormal event exceeding a baseline occurs on the changed servers includes:

[0065] Get the baseline value range corresponding to the abnormal events that occurred on the changed server;

[0066] Determine whether the event attribute value of the abnormal event occurring in the changed server exceeds the corresponding benchmark value range;

[0067] If so, determine the abnormal event as an abnormal event exceeding the benchmark, and determine the abnormal server from the changed servers where the abnormal event exceeding the benchmark occurs.

[0068] Specifically, the benchmark value range corresponding to the abnormal event can be obtained through statistical analysis of the event attribute values of the abnormal event that occurred in a certain or certain historical stages, and the event attribute values of the abnormal event usually fall within this range.

[0069] The following describes the process of obtaining the benchmark value range:

[0070] Generally, there is an association relationship between component changes and abnormal events. That is to say, when a certain component change occurs, it is very likely that a certain or certain abnormal events will occur in the server. Therefore, in the embodiments of the present application, historical data can be used to establish an association relationship between component changes and abnormal events in advance. Furthermore, for the associated abnormal events associated with component changes, the corresponding benchmark value range can be statistically obtained in advance. In this way, when judging abnormal events exceeding the benchmark, for the associated abnormal events corresponding to component changes, the above-mentioned benchmark value range obtained in advance can be used to determine abnormal events exceeding the benchmark.

[0071] Therefore, optionally, in some embodiments, obtaining the benchmark value range corresponding to the abnormal event that occurs in the changed server may include:

[0072] According to the association relationship between component changes and abnormal events obtained in advance, determine whether the abnormal event that occurs in the changed server is an associated abnormal event associated with component changes;

[0073] If so, obtain the benchmark value range corresponding to the associated abnormal event calculated in advance.

[0074] Furthermore, in some embodiments, the process of establishing the association relationship between component changes and abnormal events may include:

[0075] Obtain the first historical attribute value and the second historical attribute value of each abnormal event that occurs during the component change in the historical stage; where the first historical attribute value is the attribute value of the abnormal event that occurs in the server before the component change; the second historical attribute value is the attribute value of the abnormal event that occurs in the server after the component change;

[0076] Perform variance analysis on the first historical attribute value and the second historical attribute value of each abnormal event to obtain the associated abnormal events associated with component changes, so as to establish the association relationship between component changes and abnormal events;

[0077] Optionally, in some embodiments, the process of calculating the benchmark value range corresponding to the associated abnormal event may include:

[0078] Generate a reference value range for the associated abnormal event based on the first historical attribute value and the second historical attribute value of the associated abnormal event.

[0079] Specifically, when the difference in the event attribute values of the abnormal event is large before and after the component change, it indicates that there is an association between the component change and the abnormal event. Therefore, in the above process, after obtaining the characteristic attribute value (the first historical attribute value) of the abnormal event before the component change and the characteristic attribute value (the second historical attribute value) of the abnormal event after the component change, an analysis of variance can be used to determine whether the difference in the event attribute values of the abnormal event before and after the component change is large (whether it is significant), and then an association between the component change and the abnormal event can be established.

[0080] In addition, for a certain component change operation, the possibility of an associated abnormal event occurring in the server is relatively high. Therefore, when the above component change is performed in the server, the possibility of the above associated abnormal event occurring in the server is also relatively high. In the embodiments of the present application, the determination of the abnormal event exceeding the reference is based on the reference value range corresponding to the abnormal event. Therefore, after the association is established, for the associated abnormal event among them, a reference value range for generating the associated abnormal event can be established in advance for use in subsequent operations of determining the abnormal event exceeding the reference.

[0081] Specifically, statistical analysis can be performed on the above first historical attribute value and second historical attribute value to obtain the reference value range of the associated abnormal event. In the embodiments of the present application, the specific method of statistical analysis is not limited and can be custom-set according to the actual situation. For example: the box plot method can be used to perform statistical analysis on the first historical attribute value and the second historical attribute value to obtain the above reference value range; or, the normal distribution method can also be used to perform statistical analysis on the first historical attribute value and the second historical attribute value to obtain the above reference value range.

[0082] Optionally, in some embodiments, after determining whether the abnormal event occurring in the changed server is an associated abnormal event related to the component change, the method further includes:

[0083] If the abnormal event occurring in the changed server is a non-associated abnormal event of the component change, obtain the historical attribute value of the non-associated abnormal event within the first preset time period before the component change as the third historical attribute value;

[0084] Generate a reference value range corresponding to the non-associated abnormal event based on the third historical attribute value.

[0085] Specifically, as described above, for an associated abnormal event associated with component changes, during the process of establishing the association relationship, a reference value range corresponding to the associated abnormal event can be pre-generated based on the first historical attribute value and the second historical attribute value. For a non-associated abnormal event, in the embodiments of the present application, statistical analysis can be performed based on the historical attribute values (the third historical attribute values) of the non-associated abnormal event within the first preset time period before the component change, so as to obtain the reference value range of the non-associated abnormal event.

[0086] In the embodiments of the present application, there is no limitation on the specific manner of the above statistical analysis, and it can be custom-set according to the actual situation. For example: The box plot method can be used to perform statistical analysis on the third historical attribute value, so as to obtain the above reference value range; or, the normal distribution method can also be used to perform statistical analysis on the third historical attribute value, so as to obtain the corresponding reference value range.

[0087] In addition, in the embodiments of the present application, the process of determining the abnormal server from the changed servers where the ultra-reference abnormal event occurs can be: determining the changed server where the ultra-reference abnormal event occurs as the abnormal server; or, re-screening from the changed servers where the ultra-reference abnormal event occurs to obtain the abnormal server. In the embodiments of the present application, there is no limitation on the specific screening rule. For example: The abnormal server can be screened from the changed servers based on the number of types of ultra-reference abnormal events that occur. For example, the changed server where more than N types of ultra-reference abnormal events occur is determined as the abnormal server; or, the degree to which the event attribute value of the ultra-reference abnormal event exceeds the corresponding reference value range can be used to screen the abnormal server from it; or, the ultra-reference abnormal event can be evaluated for abnormalities from other dimensions, and then based on the obtained evaluation value, the abnormal server can be screened from it.

[0088] Optionally, in some embodiments, the process of determining the abnormal server from the changed servers where the ultra-reference abnormal event occurs may include:

[0089] Obtain multiple evaluation dimension information corresponding to the ultra-reference abnormal event, and the weight values of each evaluation dimension information;

[0090] Based on each evaluation dimension information and the weight values of each evaluation dimension information, perform weighted summation to obtain the abnormal evaluation value of the ultra-reference abnormal event;

[0091] According to the abnormal evaluation value, determine the target abnormal event from the ultra-reference abnormal event;

[0092] Determine the changed server where the target abnormal event occurs as the abnormal server.

[0093] Specifically, in the embodiments of the present application, the specific content of the evaluation dimension information corresponding to the ultra-baseline abnormal event can be customized according to the actual situation. For example, the evaluation dimension information may include at least one of the following: the time difference information from component change to the first occurrence of the ultra-baseline abnormal event, the number of occurrences of the ultra-baseline abnormal event within the preset statistical window time, the duration of the ultra-baseline abnormal event within the preset statistical window time, the impact information caused by the ultra-baseline abnormal event, the proportion information of ECS (Elastic Compute Service) instances in the server where the ultra-baseline abnormal event occurs, and the severity level information of the ultra-baseline abnormal event.

[0094] Further, in the embodiments of the present application, the AHP analysis method (The analytic hierarchy process) can be used to calculate the weight values of the above-mentioned evaluation dimension information, and weighted summation is performed based on the obtained weight values to obtain the abnormal evaluation value of the ultra-baseline abnormal event.

[0095] In the above embodiments of the present application, through the multiple evaluation dimension information corresponding to the baseline abnormal event, the abnormal degree of the baseline abnormal event is evaluated from different evaluation dimensions, so as to obtain an abnormal evaluation value that can more accurately reflect the abnormal degree of the baseline abnormal event. Furthermore, based on the abnormal evaluation values of each baseline abnormal event, the target abnormal event is determined from the baseline abnormal events, and the target abnormal event is used as the abnormal event with a higher abnormal degree. Then, the changed server where the above abnormal event with a higher abnormal degree occurs is determined as the abnormal server with impaired performance. Therefore, the accuracy of abnormal server detection can be improved.

[0096] Further, when evaluating the abnormal degree of the baseline abnormal event, the AHP analysis method gives different weights to different dimension information. In this way, while focusing on the high-priority dimension information, the low-probability dimension information can be taken into account, further improving the comprehensiveness and accuracy of the abnormal evaluation.

[0097] Optionally, in some embodiments, the method further includes:

[0098] Performing a clustering operation on each ultra-baseline abnormal event occurring in each abnormal server to obtain a clustering result; the clustering result includes clustering category information and the number information of abnormal servers included in each clustering category;

[0099] When the number of abnormal servers included in the clustering category exceeds the preset number threshold, a component change interception operation is triggered.

[0100] Specifically, for an abnormal component change that has been determined to be an abnormal change, it may also cause the same type of abnormal events to occur simultaneously in a relatively large number of servers in the cloud computing platform. When this situation occurs, it can be said that the above-mentioned abnormal component change poses a potential risk of aggregated change to the cloud computing platform. The occurrence of the potential risk of aggregated change may cause serious consequences. Therefore, when an abnormal change is determined, it is possible to further determine whether the abnormal change will cause a potential risk of aggregated change, so as to detect and intercept this potential risk in a timely manner.

[0101] In the embodiments of the present application, clustering operations are performed on each super-baseline abnormal event that occurs in each abnormal server determined to have impaired performance in the server cluster, to obtain multiple categories, and the number of abnormal servers included in each category. When the number of abnormal servers included in a certain category is relatively large, it is considered that the abnormal component change poses a potential risk of aggregated change to the cloud computing platform. After that, an interception operation for the component change can be triggered. In the embodiments of the present application, the specific content of the interception operation is not limited and can be custom-set according to the actual situation. For example: an aggregated change potential risk notification can be sent to inform the release personnel which specific component release is determined to have a potential risk of aggregated change, and how many changed servers have occurred a relatively serious abnormal event of a certain type (component release) caused by this component release, so as to facilitate the release personnel to perform corresponding component change interception measures.

[0102] Optionally, in some embodiments, each abnormal event is set with tag information; clustering operations are performed on each super-baseline abnormal event that occurs in each abnormal server to obtain a clustering result, including:

[0103] Based on the tag information, similarity calculations are performed on each super-baseline abnormal event that occurs in each abnormal server to obtain a similarity calculation result;

[0104] Based on the similarity calculation result, clustering operations are performed on each super-baseline abnormal event that occurs in each abnormal server to obtain a clustering result.

[0105] Specifically, in the embodiments of the present application, the specific content of the tag information is not limited and can be custom-set according to the actual situation. For example, the tag information may include at least one of the following: associated product form information, event category information, event text information. Among them, the associated product form information may be: information about the product associated with the abnormal event, such as: the product is a GPU or a CPU, etc. Exemplarily, the event category information may be: network-class abnormal events, motherboard-class abnormal events, disk-class abnormal times, etc.

[0106] In the embodiments of the present application, tag information is preset for abnormal events. Furthermore, when performing a clustering operation, a clustering operation can be performed based on the above tag information on the basis of clustering based on the text data used to describe abnormal events, so as to make up for the problem of poor flexibility existing in the text clustering method, thereby effectively improving the robustness of the clustering of outlier abnormal events.

[0107] Optionally, in some embodiments, before performing a clustering operation on each outlier abnormal event that occurs in each abnormal server, it may further include:

[0108] Obtain the attribute values of each outlier abnormal event that occurs in the abnormal server within a second preset time period before and after the component change;

[0109] Calculate the attribute value fluctuation parameter of each outlier abnormal event within the second preset time period; the attribute value fluctuation parameter characterizes the degree of fluctuation of the attribute value of the outlier abnormal event;

[0110] Based on the attribute value fluctuation parameter, determine the key abnormal events from each outlier abnormal event that occurs in the abnormal server;

[0111] Performing a clustering operation on each outlier abnormal event that occurs in each abnormal server includes:

[0112] Performing a clustering operation on each key abnormal event that occurs in each abnormal server.

[0113] Specifically, the embodiments of the present application do not limit the specific calculation method adopted when calculating the attribute value fluctuation parameter, and can be custom-set according to the actual situation. For example: algorithms such as Squeeze (a squeezing algorithm, one of the root cause location algorithms) can be used to calculate the above attribute value fluctuation parameter.

[0114] In the above embodiments of the present application, after determining the outlier abnormal event, the attribute value fluctuation parameter of each outlier abnormal event within the second preset time period before and after the component change is calculated. Furthermore, based on the above attribute value fluctuation parameter, the key abnormal events are determined. After that, when performing clustering, it is based on the above key abnormal events.

[0115] Since the attribute value fluctuation parameter characterizes the degree of fluctuation of the attribute value of the outlier abnormal event, therefore, the key outlier abnormal events that can reflect the current component change can be found from each outlier abnormal event based on the attribute value fluctuation parameter, that is, the above key abnormal events. That is to say, in the embodiments of the present application, by analyzing the attribute value fluctuation parameters of each outlier abnormal event that occurs in the abnormal server, the root cause for the server to be determined as an abnormal server is found, that is: which abnormal event mainly causes the server to be determined as an abnormal server.

[0116] SeeFigure 2 , Figure 2 is Figure 1 the schematic diagram of the detection process corresponding to the illustrated embodiment. The following will describe the process of the component change anomaly detection method of the above embodiments of the present application in conjunction with Figure 2 :

[0117] The detection process can be divided into the following two parts: The first part is the offline calculation part of historical data; the second part is the real-time component change anomaly detection part.

[0118] Among them, the first part, the offline calculation part of historical data, specifically includes: obtaining the component change information in the historical stage, and the abnormal events that occurred in the corresponding changed servers; based on the above information obtained, performing offline calculation to obtain the correlation relationship between component changes and abnormal events, and the benchmark value range corresponding to the associated abnormal events.

[0119] The second part, the real-time component change anomaly detection part, specifically includes: obtaining the real-time component change information, and the abnormal events that occurred in the corresponding changed servers; calculating the spatial distribution value of the abnormal events based on the above information obtained as the spatial feature to determine whether the real-time component change is an abnormal change; if it is determined to be an abnormal change, then based on the correlation relationship calculated in the first part above, the benchmark value range corresponding to the associated abnormal events, and the event attribute value of the abnormal event, determining whether the associated abnormal event is a super-benchmark abnormal event, or based on the real-time orthogonal method, statistically obtaining the benchmark value range corresponding to the non-associated abnormal events, and then based on the benchmark value range corresponding to the non-associated abnormal events and the event attribute value of the abnormal event, determining whether the non-associated abnormal events are super-benchmark abnormal events; in addition, for the detected abnormal component change (the suspected anomaly in Figure 2 ), root cause analysis can be performed based on the attribute value fluctuation parameter of the super-benchmark abnormal event, that is: determining the key abnormal event reflecting the current component change problem based on the attribute value fluctuation parameter of the super-benchmark abnormal event; clustering based on the key abnormal events in each abnormal server to obtain a clustering result including category information and the number of abnormal servers included in each category; and then determining whether it is a clustering anomaly (whether it causes a potential risk of aggregated change) based on the clustering result; when it is determined that the current component change is a clustering anomaly, a notification can be sent to the component release personnel to perform the corresponding fuse operation, or the current component change can be marked to form component change historical data for subsequent use.

[0120] Embodiment 2

[0121] Refer to Figure 3 , Figure 3 is the step flowchart of a component change anomaly detection method according to Embodiment 2 of the present application.Figure 3 The application scenario of component change anomaly detection shown can be: when a specific component change is made to a server in a cloud computing platform, the method shown can be used for anomaly detection to detect whether the above specific component change is an abnormal change. And when the above specific component change is detected as an abnormal change, further determine the abnormal servers with impaired performance in this abnormal change, and then return an abnormal change notice containing the identification information of the abnormal servers to inform that the above specific component change is an abnormal change and which specific server performances are impaired due to this abnormal change. Figure 3 Specifically, the component change anomaly detection method provided in this embodiment includes the following steps:

[0122] Specifically, the component change anomaly detection method provided in this embodiment includes the following steps:

[0123] Step 302, receive a component change notice and obtain the abnormal events that occur in the changed servers in the cloud computing platform after the component change.

[0124] Step 304, calculate the spatial distribution value of the abnormal events, and based on the spatial distribution value, determine whether the component change is an abnormal change. [[ID=**14]]

[0125] Step 306, if it is an abnormal change, then determine the abnormal servers from the changed servers based on whether a hyperbenchmark abnormal event occurs in the changed servers.

[0126] Step 308, return an abnormal change notice, and the abnormal change notice contains the identification information of the abnormal servers.

[0127] In the component change anomaly detection solution provided in the embodiments of the present application, when a change notice of a specific component change is received, the abnormal events that occur in the changed servers in the cloud computing platform after the component change are obtained. Then, from the spatial domain level of the server cluster, the spatial distribution value of the abnormal events is calculated, and then the anomaly detection of the component change is performed based on the above spatial distribution value. In addition, when it is detected that this component change is an abnormal change, from the time domain level, the benchmark value range of the abnormal event attribute value is obtained according to the historical performance of the abnormal event, and then the hyperbenchmark abnormal event and the abnormal servers are determined based on the above benchmark value range. In the embodiments of the present application, the dependence on artificially preset rules in the component change anomaly detection process is weakened, and the spatial performance and historical performance of the abnormal events are used as the current stage anomaly determination criteria. Or rather, the trend performance of the current stage is mainly predicted through the historical performance of the abnormal events. Therefore, the embodiments of the present application can effectively improve the robustness and accuracy of component change anomaly detection.

[0128] The component change anomaly detection method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, PCs, etc.

[0129] Embodiment III

[0130] Refer to Figure 4 , Figure 4 As shown in Figure 4 is a flowchart of the steps of a component change anomaly detection method according to Embodiment III of the present application. Since there are many types of computing components, and for a specific type of computing component, the number of version changes may also be large. Based on the above situation, for the server cluster in the cloud computing platform, there may be a series of component change processes. Therefore, Figure 4 The application scenario of the component change anomaly detection shown in Figure 4 can be: during the process of component change, continuously obtain the component change log information recording various component change situations. Correspondingly, also synchronously obtain the abnormal event log information for the server cluster of the cloud computing platform. Then, through the method provided in the embodiments of the present application, perform anomaly detection on each component change recorded in the component change log information respectively. And when an abnormal change is detected, return the identification information of the abnormal change, as well as the identification information of the abnormal server corresponding to the abnormal change.

[0131] Specifically, the component change anomaly detection method provided in this embodiment includes the following steps:

[0132] Step 402, receive the component change log information and the abnormal event log information for the server cluster of the cloud computing platform.

[0133] Among them, the component change log information is used to record various component changes made to the servers in the cloud computing platform.

[0134] Step 404, for each component change recorded in the component change log information, determine the corresponding abnormal event information segment from the abnormal event log information.

[0135] Among them, the abnormal event information segment corresponding to the component change is used to describe the abnormal events that occurred on the changed servers during the corresponding component change process.

[0136] Step 406, based on the abnormal event information segment, calculate the spatial distribution value of the abnormal event corresponding to each component change, and based on the spatial distribution value, determine the abnormal change from each component change.

[0137] Step 408, for the abnormal change, determine the abnormal server from the changed servers based on whether a super-reference abnormal event occurred on the changed server.

[0138] Step 410, return an abnormal change notice, and the abnormal change notice includes the identification information of the abnormal change, as well as the identification information of the abnormal server corresponding to the abnormal change.

[0139] In the component change anomaly detection solution provided by the embodiments of the present application, during a series of component changes to the server cluster in the cloud computing platform, component change log information recording various component change situations and anomaly event log information for the server cluster of the cloud computing platform are obtained. Then, based on the above information, anomaly detection is respectively performed on the above series of component changes, and the identification information of the detected abnormal changes and the identification information of the abnormal servers corresponding to the abnormal changes are obtained.

[0140] In the above embodiment, when performing anomaly detection on each component change, from the spatial domain level of the server cluster, the spatial distribution value of the anomaly event is calculated, and then anomaly detection of the component change is performed based on the above spatial distribution value; in addition, when an abnormal change is detected, from the time domain level, the reference value range of the anomaly event attribute value is obtained according to the historical performance of the anomaly event, and then the abnormal events beyond the reference and the abnormal servers are determined based on the above reference value range. In the embodiments of the present application, the dependence on artificially preset rules in the component change anomaly detection process is weakened, and the spatial performance and historical performance of the anomaly event are used as the current stage anomaly determination criteria. Or rather, the trend performance in the current stage is mainly predicted through the historical performance of the anomaly event. Therefore, the embodiments of the present application can effectively improve the robustness and accuracy of component change anomaly detection.

[0141] The component change anomaly detection method of this embodiment can be executed by any suitable electronic device with data processing capabilities, including but not limited to: servers, PCs, etc.

[0142] Embodiment 4

[0143] Figure 5 It is a structural block diagram of a component change anomaly detection device according to Embodiment 4 of the present application. The component change anomaly detection device provided by the embodiments of the present application includes:

[0144] A statistics module 502, configured to count the anomaly events that occur in the changed servers with component changes in the server cluster of the cloud computing platform;

[0145] An abnormal change determination module 504, configured to calculate the spatial distribution value of the anomaly event and determine whether the component change is an abnormal change based on the spatial distribution value;

[0146] An abnormal server determination module 506, configured to, if it is an abnormal change, determine the abnormal server from the changed servers based on whether the changed servers have abnormal events beyond the reference; an abnormal event beyond the reference is an abnormal event whose event attribute value exceeds the corresponding reference value range; the reference value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage.

[0147] Optionally, in some embodiments, the abnormal server determination module 506, when performing the step of determining an abnormal server from the changed servers based on whether a hyperbenchmark abnormal event occurs in the changed servers, specifically is used for:

[0148] Obtain the benchmark value range corresponding to the abnormal event that occurs in the changed server;

[0149] Determine whether the event attribute value of the abnormal event that occurs in the changed server exceeds the corresponding benchmark value range;

[0150] If so, determine the abnormal event as a hyperbenchmark abnormal event, and determine an abnormal server from the changed servers where the hyperbenchmark abnormal event occurs.

[0151] Optionally, in some embodiments, the abnormal server determination module 506, when performing the step of obtaining the benchmark value range corresponding to the abnormal event that occurs in the changed server, specifically is used for:

[0152] According to the pre-obtained association relationship between component changes and abnormal events, determine whether the abnormal event that occurs in the changed server is an associated abnormal event associated with the component change;

[0153] If so, obtain the benchmark value range corresponding to the associated abnormal event calculated in advance.

[0154] Optionally, in some embodiments, the component change abnormal detection device further includes:

[0155] An association relationship establishment module, configured to obtain the first historical attribute value and the second historical attribute value of each abnormal event that occurs during the component change in the historical stage; wherein, the first historical attribute value is the attribute value of the abnormal event that occurs in the server before the component change; the second historical attribute value is the attribute value of the abnormal event that occurs in the server after the component change; perform variance analysis on the first historical attribute value and the second historical attribute value of each abnormal event to obtain an associated abnormal event associated with the component change, so as to establish an association relationship between the component change and the abnormal event;

[0156] A benchmark value range calculation module, configured to generate a benchmark value range of the associated abnormal event based on the first historical attribute value and the second historical attribute value of the associated abnormal event.

[0157] Optionally, in some embodiments, the abnormal server determination module 506 is further used for: after performing the step of determining whether the abnormal event that occurs in the changed server is an associated abnormal event associated with the component change, if the abnormal event that occurs in the changed server is a non-associated abnormal event of the component change, obtain the historical attribute value of the non-associated abnormal event within the first preset time period before the component change, as the third historical attribute value;

[0158] Generate a baseline value range corresponding to the non-associated exception event based on the third historical attribute value.

[0159] Optionally, in some embodiments, the exception server determination module 506, when performing the step of determining the exception server from the changed servers with supra-baseline exception events, is specifically configured to:

[0160] Obtain multiple evaluation dimension information corresponding to the supra-baseline exception event, and the weight value of each evaluation dimension information;

[0161] Based on each evaluation dimension information and the weight value of each evaluation dimension information, perform weighted summation to obtain an exception evaluation value of the supra-baseline exception event;

[0162] Determine a target exception event from the supra-baseline exception events according to the exception evaluation value;

[0163] Determine the changed server where the target exception event occurs as the exception server.

[0164] Optionally, in some embodiments, the component change exception detection device further includes:

[0165] A clustering module, configured to perform a clustering operation on each supra-baseline exception event occurring in each exception server to obtain a clustering result; the clustering result includes clustering category information and the number information of the exception servers included in each clustering category; when the number of exception servers included in a clustering category exceeds a preset number threshold, trigger a component change interception operation.

[0166] Optionally, in some embodiments, each exception event is provided with label information; the clustering module is specifically configured to:

[0167] Based on the label information, calculate the similarity of each supra-baseline exception event occurring in each exception server to obtain a similarity calculation result;

[0168] Based on the similarity calculation result, perform a clustering operation on each supra-baseline exception event occurring in each exception server to obtain a clustering result.

[0169] Optionally, in some embodiments, the component change exception detection device further includes:

[0170] A key abnormal event determination module is configured to, before clustering the ultra-baseline abnormal events that occur in each abnormal server, obtain the attribute values of the ultra-baseline abnormal events that occur in the abnormal server within a second preset time period before and after the component change; calculate the attribute value fluctuation parameter of each ultra-baseline abnormal event within the second preset time period; the attribute value fluctuation parameter characterizes the fluctuation degree of the attribute value of the ultra-baseline abnormal event; based on the attribute value fluctuation parameter, determine the key abnormal events from the ultra-baseline abnormal events that occur in the abnormal server.

[0171] A clustering module is specifically configured to: perform a clustering operation on the key abnormal events that occur in each abnormal server to obtain a clustering result.

[0172] The component change abnormal detection device in this embodiment is used to implement the corresponding component change abnormal detection method in the foregoing Embodiment 1, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the component change abnormal detection device in this embodiment can be referred to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0173] Embodiment 5

[0174] Figure 6 FIG. is a structural block diagram of a component change abnormal detection device according to Embodiment 5 of the present application. The component change abnormal detection device provided in the embodiment of the present application includes:

[0175] A notification receiving module 602 is configured to receive a component change notification and obtain the abnormal events that occur in the changed servers in the cloud computing platform after the component change.

[0176] A first detection module 604 is configured to calculate the spatial distribution value of the abnormal events, and based on the spatial distribution value, determine whether the component change is an abnormal change; if it is an abnormal change, determine the abnormal server from the changed servers based on whether the ultra-baseline abnormal events occur in the changed servers.

[0177] A first notification return module 606 is configured to return an abnormal change notification, and the abnormal change notification includes the identification information of the abnormal server.

[0178] The component change abnormal detection device in this embodiment is used to implement the corresponding component change abnormal detection method in the foregoing Embodiment 2, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the component change abnormal detection device in this embodiment can be referred to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0179] Embodiment 6

[0180] Figure 7It is a structural block diagram of a component change anomaly detection device according to Embodiment 6 of the present application. The component change anomaly detection device provided by the embodiment of the present application includes:

[0181] A log information receiving module 702, configured to receive component change log information and, for the server cluster of the cloud computing platform, anomaly event log information;

[0182] A second detection module 704, configured to, for each component change recorded in the component change log information, determine a corresponding anomaly event information segment from the anomaly event log information; calculate a spatial distribution value of the anomaly event corresponding to each component change based on the anomaly event information segment, and determine an abnormal change from each component change based on the spatial distribution value; for the abnormal change, determine an abnormal server from the changed servers based on whether a super-reference anomaly event occurs in the changed servers;

[0183] A second notification return module 706, configured to return an abnormal change notification, where the abnormal change notification includes identification information of the abnormal change and identification information of the abnormal server corresponding to the abnormal change.

[0184] The component change anomaly detection device of this embodiment is used to implement the corresponding component change anomaly detection method in the foregoing Embodiment 3, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the component change anomaly detection device of this embodiment can refer to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0185] Embodiment 7

[0186] Refer to Figure 8 , which shows a structural schematic diagram of an electronic device according to Embodiment 7 of the present application. The specific implementation of the electronic device is not limited in the specific embodiment of the present application.

[0187] As Figure 8 shown, the electronic device may include: a processor 802, a communication interface 804, a memory 806, and a communication bus 808.

[0188] Wherein:

[0189] The processor 802, the communication interface 804, and the memory 806 communicate with each other through the communication bus 808.

[0190] The communication interface 804 is configured to communicate with other electronic devices.

[0191] A processor 802 is configured to execute a program 810, and specifically, can execute the relevant steps in the embodiment of the component change anomaly detection method described above.

[0192] Specifically, the program 810 may include program code, and the program code includes computer operation instructions.

[0193] The processor 802 may be a CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0194] A memory 806 is configured to store the program 810. The memory 806 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0195] The program 810 may include multiple computer instructions. Specifically, the program 810 may cause the processor 802 to execute the operations corresponding to the component change anomaly detection method described in any one of the foregoing multiple method embodiments through the multiple computer instructions.

[0196] For the specific implementation of each step in the program 810, reference may be made to the corresponding steps and descriptions in the corresponding units in the foregoing method embodiments, and they have corresponding beneficial effects, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, and will not be elaborated here.

[0197] The embodiments of the present application further provide a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in any one of the foregoing multiple method embodiments. The computer storage medium includes but is not limited to: Compact Disc Read-Only Memory (CD-ROM), Random Access Memory (RAM), floppy disk, hard disk, or magneto-optical disk, etc.

[0198] The embodiments of the present application further provide a computer program product, including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to any one of the component change anomaly detection methods in the foregoing multiple method embodiments.

[0199] In addition, it should be noted that the information related to users (including but not limited to user device information, user personal information, etc.) and data (including but not limited to sample data for training the model, data for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data that have been authorized by the users or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

[0200] It should be pointed out that according to the needs of implementation, each component / step described in the embodiments of this application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0201] The methods according to the embodiments of this application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through the network and will be stored in a local recording medium. Thus, the methods described herein can be stored on such a software process on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA)). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a Random Access Memory (RAM), a Read-Only Memory (ROM), a flash memory, etc.) that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods described herein are implemented. In addition, when a general-purpose computer accesses the code for implementing the methods shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0202] Those of ordinary skill in the art will appreciate that the units and method steps of the examples described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled artisans may use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0203] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A method for detecting abnormal component changes, comprising: Counting abnormal events that occur in the changed servers with component changes in the server cluster of the cloud computing platform; Calculating the spatial distribution value of the abnormal events, and based on the spatial distribution value, determining whether the component change is an abnormal change; If it is the abnormal change, determining abnormal servers from the changed servers based on whether the changed servers have super-baseline abnormal events; The super-baseline abnormal event is an abnormal event whose event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage.

2. The method according to claim 1, wherein, The determining abnormal servers from the changed servers based on whether the changed servers have super-baseline abnormal events includes: Obtaining the baseline value range corresponding to the abnormal events that occur in the changed servers; Judging whether the event attribute value of the abnormal events that occur in the changed servers exceeds the corresponding baseline value range; If so, determining the abnormal event as a super-baseline abnormal event, and determining abnormal servers from the changed servers where the super-baseline abnormal event occurs.

3. The method according to claim 2, wherein The obtaining the baseline value range corresponding to the abnormal events that occur in the changed servers includes: According to the pre-obtained correlation relationship between component changes and abnormal events, judging whether the abnormal events that occur in the changed servers are associated abnormal events related to component changes; If so, obtaining the baseline value range corresponding to the associated abnormal event calculated in advance.

4. The method according to claim 3, wherein The establishment process of the correlation relationship between component changes and abnormal events includes: Obtaining the first historical attribute value and the second historical attribute value of each abnormal event that occurred during the component change in the historical stage; wherein, the first historical attribute value is the attribute value of the abnormal event that occurred in the server before the component change; the second historical attribute value is the attribute value of the abnormal event that occurred in the server after the component change; Performing variance analysis on the first historical attribute value and the second historical attribute value of each abnormal event to obtain associated abnormal events related to component changes, so as to establish the correlation relationship between component changes and abnormal events; The calculation process of the baseline value range corresponding to the associated abnormal event includes: Generating the baseline value range of the associated abnormal event based on the first historical attribute value and the second historical attribute value of the associated abnormal event.

5. The method according to claim 3, wherein After judging whether the abnormal events that occur in the changed servers are associated abnormal events related to component changes, the method further includes: If the abnormal events that occur in the changed servers are non-associated abnormal events of the component change, obtaining the historical attribute value of the non-associated abnormal event within the first preset time period before the component change as the third historical attribute value; Generating the baseline value range corresponding to the non-associated abnormal event based on the third historical attribute value.

6. The method according to claim 2, wherein The determining abnormal servers from the changed servers where the super-baseline abnormal event occurs includes: Obtaining multiple evaluation dimension information corresponding to the super-baseline abnormal event, and the weight value of each evaluation dimension information; Based on the information of each evaluation dimension and the weight value of each evaluation dimension, perform weighted summation to obtain the anomaly evaluation value of the super-baseline anomaly event; Determine the target anomaly event from the super-baseline anomaly events according to the anomaly evaluation value; Determine the changed servers where the target anomaly event occurs as anomaly servers.

7. The method according to claim 1, wherein The method further includes: Perform a clustering operation on each super-baseline anomaly event that occurs in each anomaly server to obtain a clustering result; the clustering result includes clustering category information and the number information of the anomaly servers included in each clustering category; When the number of anomaly servers included in a clustering category exceeds a preset number threshold, trigger a component change interception operation.

8. The method according to claim 7, wherein Each anomaly event is set with label information; The performing a clustering operation on each super-baseline anomaly event that occurs in each anomaly server to obtain a clustering result includes: Based on the label information, calculate the similarity of each super-baseline anomaly event that occurs in each anomaly server to obtain a similarity calculation result; Based on the similarity calculation result, perform a clustering operation on each super-baseline anomaly event that occurs in each anomaly server to obtain a clustering result.

9. The method according to claim 7, wherein Before performing the clustering operation on each super-baseline anomaly event that occurs in each anomaly server, the method further includes: Obtain the attribute values of each super-baseline anomaly event that occurs in the anomaly server within a second preset time period before and after the component change; Calculate the attribute value fluctuation parameter of each super-baseline anomaly event within the second preset time period; the attribute value fluctuation parameter characterizes the fluctuation degree of the attribute value of the super-baseline anomaly event; Based on the attribute value fluctuation parameter, determine the key anomaly events from each super-baseline anomaly event that occurs in the anomaly server; The performing a clustering operation on each super-baseline anomaly event that occurs in each anomaly server includes: Perform a clustering operation on each key anomaly event that occurs in each anomaly server.

10. A component change anomaly detection method, including: Receive a component change notification, and obtain the anomaly events that occur in the changed servers in the cloud computing platform after the component change; Calculate the spatial distribution value of the anomaly events, and based on the spatial distribution value, determine whether the component change is an abnormal change; If it is an abnormal change, then based on whether the changed server has a super-baseline anomaly event, determine the anomaly server from the changed servers; The super-baseline anomaly event is an anomaly event whose event attribute value exceeds the corresponding baseline value range; the baseline value range is obtained based on the attribute values of the anomaly events that occurred in the historical stage; Return an abnormal change notification, and the abnormal change notification includes the identification information of the anomaly server.

11. A component change anomaly detection method, including: Receive component change log information and anomaly event log information for the server cluster of the cloud computing platform; For each component change recorded in the component change log information, determine the corresponding anomaly event information segment from the anomaly event log information; Based on the anomaly event information segment, calculate the spatial distribution value of the anomaly events corresponding to each component change, and based on the spatial distribution value, determine the abnormal change from each component change; For the abnormal change, based on whether a hyperbenchmark abnormal event occurs in the changed server, determine the abnormal server from the changed servers; The hyperbenchmark abnormal event is an abnormal event whose event attribute value exceeds the corresponding benchmark value range; the benchmark value range is obtained based on the attribute values of the abnormal events that occurred in the historical stage; Return an abnormal change notice, where the abnormal change notice includes the identification information of the abnormal change and the identification information of the abnormal server corresponding to the abnormal change.

12. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the method according to any one of claims 1-11.

13. A computer storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method according to any one of claims 1-11.

14. A computer program product, including computer instructions, where the computer instructions instruct a computing device to perform the operations corresponding to the method according to any one of claims 1-11.