A Fast Fault Identification Method Based on TF-IDF Weighted Cosine Similarity Algorithm

By using the TF-IDF weighted cosine similarity algorithm in the fault definition system, the fault data sample size in the index library is expanded, and the similarity push processing solution is used to solve the problem of definition failure caused by the limited number of fault data in the prior art, and the success rate and efficiency of fault definition are improved.

CN117707819BActive Publication Date: 2025-05-09SUZHOU GAIYA INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311595077.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-09
Estimated Expiration
2043-11-27

AI Technical Summary

Technical Problem

In the prior art, the number of fault data that can be compared in the database is limited, resulting in some faults not being successfully defined.

Method used

The fault fast definition method based on TF-IDF weighted cosine similarity algorithm is adopted. By injecting faults and collecting benchmark data, the sample size in the index library is expanded, and the processing solution is pushed using the concept of similarity.

Benefits of technology

Improves the success rate of fault definition, reduces the possibility of misdiagnosis, and reduces maintenance and management burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117707819B_ABST
    Figure CN117707819B_ABST
Patent Text Reader

Abstract

The present application relates to a method for rapid fault definition based on the TF‑IDF weighted cosine similarity algorithm, including: injecting faults based on a pre-built fault simulation model, collecting benchmark data corresponding to preset fault indicators during the fault injection process, and storing the benchmark data in a preset indicator library; receiving a fault definition instruction, and comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on the pre-built fault definition model, and determining the similarity; based on the similarity, pushing a processing plan; wherein, based on the similarity, pushing a processing plan at least includes: if there is target benchmark data similar to the fault data, then pushing the preset processing plan corresponding to the target benchmark data. The present application can continuously and autonomously introduce faults, collect fault indicator-related data, automatically expand the number of fault samples available for comparison, and can also use the TF‑IDF weighted cosine similarity algorithm to achieve rapid definition and location of faults.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault diagnosis, and in particular to a method for rapid fault definition based on a TF-IDF weighted cosine similarity algorithm. Background Art

[0002] As current technological trends evolve towards large-scale clusters, ultra-complex distributed systems, and microservice architectures, the scale of computing has expanded, the number of nodes has increased dramatically, and the dependencies between nodes have become increasingly complex, resulting in a sharp increase in the probability of failure. At the same time, the types of failures are also becoming more diverse. Therefore, how to quickly diagnose and define failures when they are discovered has become a hot topic of discussion today.

[0003] The current fault definition method generally stores relevant fault data and corresponding processing solutions for multiple different faults in a database in advance. For example, the fault data can exist in text form, such as including the cause of the fault, the type of fault, the location of the fault, etc.; or, the fault data can also be expressed as the performance data and status information corresponding to the equipment when the fault occurs; when an actual fault occurs, the fault data corresponding to the actual fault can be compared one by one with the fault data in the database to find the consistent fault data and feedback the corresponding processing solution to achieve a successful match, thereby facilitating maintenance personnel to maintain the actual fault that has occurred.

[0004] However, in combination with the above-mentioned technical solutions for fault diagnosis, it can be seen that the search for the above-mentioned faults is limited by the number of fault data in the database. When the number of fault data available for comparison in the database is limited, it is easy to cause the inability to find fault data from the database that matches the actual fault data, thereby resulting in definition failure. Therefore, the above-mentioned fault definition solution is not perfect and needs to be improved urgently. Summary of the Invention

[0005] In order to improve the technical problem that some faults cannot be successfully defined due to the limited amount of fault data in the database available for comparison with actual fault data, the present application provides a fast fault definition method based on a TF-IDF weighted cosine similarity algorithm.

[0006] In the first aspect, the present application provides a method for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm, which adopts the following technical solutions:

[0007] Injecting a fault based on a pre-built fault simulation model, during which time benchmark data corresponding to a preset fault indicator is collected and stored in a preset indicator library;

[0008] receiving a fault delineation instruction, and comparing the fault data in the fault delineation instruction with the benchmark data in the indicator library one by one based on a pre-built fault delineation model, and determining similarity;

[0009] Based on the similarity, push a processing solution;

[0010] The push processing solution based on the similarity at least includes:

[0011] If there is target reference data similar to the fault data, a preset processing solution corresponding to the target reference data is pushed.

[0012] By adopting the above-mentioned technical solution, faults are injected, and the data corresponding to the injected faults are stored as benchmark data in the indicator library, so as to expand the sample size that can be used for comparison with the fault data of the actual fault (i.e., the fault corresponding to the fault definition instruction), and the success rate of definition is improved by expanding the sample size; in addition, compared with the existing comparison scheme that needs to find consistent fault data, the present application proposes the concept of similarity, that is, as long as there is benchmark data similar to the current actual fault (it does not need to be completely consistent, just similar), the benchmark data will be used as the target benchmark data, and its corresponding processing solution will be pushed, thereby further improving the success rate of definition and improving the problem that some faults cannot be defined due to the limited sample size available for comparison.

[0013] Optionally, the step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity may include:

[0014] Converting the reference data and the fault data in the fault definition instruction into a vector form;

[0015] The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes:

[0016] Based on the preset cosine similarity algorithm, the vectorized benchmark data and fault data are compared and the cosine similarity is calculated.

[0017] By adopting the above technical solution, cosine similarity is a measurement method for measuring the similarity between two vectors, which is calculated based on the cosine value of the angle between the vectors. The value range of cosine similarity is between -1 and 1, where 1 indicates complete similarity, -1 indicates complete dissimilarity, and 0 indicates no linear correlation. Generally, the closer the cosine similarity is to 1, the more similar the two vectors are.

[0018] Optionally, converting the reference data and the fault data in the fault delimitation instruction into a vector form includes:

[0019] Based on the preset TF-IDF algorithm, the benchmark data and fault data are converted into vector form.

[0020] By adopting the above technical solution, the TF-IDF weighted cosine similarity algorithm can reduce the weight of common words and increase the weight of key words. During the algorithm process, TF-IDF is used to assign different weights to different feature words based on the importance and frequency of occurrence of different feature words corresponding to the benchmark data and fault data, so that the benchmark data and fault data can be represented in vector form based on the weight values ​​of different feature words.

[0021] Optionally, the indicator library includes several indicator vector libraries, and different indicator vector libraries are used to store benchmark data in vector form corresponding to different types of faults; each indicator vector library corresponds to a label attribute for distinguishing other indicator vector libraries, and the processing scheme includes at least the label attribute;

[0022] The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes:

[0023] The fault data is compared with the benchmark data in each indicator vector library in a preset order, and the similarity is determined.

[0024] By adopting the above technical solution, all benchmark data are classified using the indicator vector library. The basis for classification can be determined according to actual needs. For example, each layer corresponds to an indicator vector library according to the layer of the component to which the fault belongs (such as the IaaS layer, PaaS layer, and SaaS layer). After finding benchmark data similar to the fault data in the fault definition instruction, the label attributes will be fed back in the processing plan, and the label attributes will be used to realize fault location, so that maintenance personnel can more clearly know the source of the fault (that is, the layer to which it belongs).

[0025] Optionally, the method further includes:

[0026] Calculate the similarity between any two reference data and add the similarity to a preset comparison table;

[0027] The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes:

[0028] Filtering reference data from the indicator library, wherein the reference data refers to benchmark data that meets preset conditions;

[0029] Comparing the fault data with the reference data and calculating a reference similarity;

[0030] If the reference similarity is not less than a preset similarity threshold, it means that the reference data is similar to the fault data, and the reference data is used as the target benchmark data;

[0031] If the reference similarity is lower than a preset similarity threshold, then, in combination with a preset comparison table, the similarities corresponding to other benchmark data other than the reference data are subtracted from the reference similarity to obtain a similarity difference;

[0032] Based on the similarity difference and multiple preset comparison echelons, other benchmark data except the reference data are added to the corresponding comparison echelon. Each comparison echelon has a pre-defined difference range and a comparison priority. All benchmark data in the comparison echelon with a higher priority will be compared with the fault data first.

[0033] Based on the priority, the benchmark data in the comparison echelon is compared with the fault data in sequence to determine the similarity.

[0034] By adopting the above technical solution, the comparison order is limited and adjusted. Specifically, one or more reference data can be selected from the benchmark data, and the fault data can be compared with the reference data first, and the similarity between the two (i.e., the reference similarity) is calculated. If A exists in other benchmark data, and the similarity between A and the reference data is close to the reference similarity (i.e., the similarity difference is small), then A and the fault data are very likely to be similar. Therefore, this application will classify all benchmark data other than the reference data based on the size of the similarity difference, and classify them into different comparison echelons. Then, according to the priority size of the comparison echelon, the benchmark data in the echelon will be compared with the fault data in sequence to obtain the similarity.

[0035] Optionally, if there is target reference data similar to the fault data, pushing a preset processing solution corresponding to the target reference data, which includes:

[0036] Traversing all similarities in order, and comparing the similarities one by one with a preset similarity threshold;

[0037] If there is a first similarity that is not less than a preset similarity threshold, the reference data corresponding to the largest first similarity is used as the target reference data.

[0038] By adopting the above-mentioned technical solution, the present application proposes to traverse all benchmark data when searching for benchmark data similar to fault data, and preferably selects benchmark data with a similarity greater than a preset similarity threshold and the highest similarity as the target benchmark data, and pushes the processing solution corresponding to the target benchmark data to the user, so that the target benchmark data finally determined is the benchmark data closest to the fault data, and the processing solution finally pushed can help solve the fault corresponding to the fault data with the greatest probability.

[0039] Optionally, several concurrent fault sets are preset, each of which contains at least several benchmark data, and all benchmark data in the same concurrent fault set meet the following conditions: when a fault corresponding to one of the benchmark data occurs, faults corresponding to other benchmark data in the corresponding concurrent fault set will be triggered;

[0040] The method further comprises:

[0041] After the target benchmark data is determined, if the target benchmark data exists in any concurrent fault set, a new priority comparison set is created. The priority comparison set includes all benchmark data in the concurrent fault set where the target benchmark data is located, and a validity period is set. After the validity period of the newly created priority comparison set has expired, the priority comparison set is deleted.

[0042] The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes:

[0043] If there is a priority comparison set, the fault data is first compared with all the benchmark data in the priority comparison set and the similarity is determined; if no target benchmark data similar to the fault data is found, the fault data is then compared with all other benchmark data except the priority comparison set and the similarity is determined for each of them;

[0044] If there is no priority comparison set, the fault data is compared with the benchmark data in the indicator library one by one, and the similarity is determined.

[0045] By adopting the above technical solution, the occurrence of some faults will lead to other faults. Therefore, the benchmark data corresponding to such faults can be collected in the same concurrent fault set. When the target benchmark data matched by a certain fault exists in any concurrent fault set A, then within the effective period thereafter, if other faults B are newly discovered, the fault data of fault B will be prioritized for similarity comparison with other fault data in concurrent fault set A (i.e., the priority comparison set). After the comparison is completed, it will be compared with other benchmark data in the indicator library. That is, the comparison order of fault data and benchmark data is adjusted based on the correlation between faults, so as to more quickly define the target benchmark data for the fault.

[0046] In a second aspect, the present application provides a rapid fault definition system based on the TF-IDF weighted cosine similarity algorithm, comprising:

[0047] A fault injection module injects faults based on a pre-built fault simulation model. During the fault injection process, the module collects benchmark data corresponding to preset fault indicators and stores the benchmark data in a preset indicator library.

[0048] a fault definition module, configured to receive a fault definition instruction, compare the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on a pre-built fault definition model, and determine the similarity;

[0049] The solution pushing module is used to push a processing solution based on the similarity; specifically, if there is target reference data similar to the fault data, then push the preset processing solution corresponding to the target reference data.

[0050] In a third aspect, the present application provides a device for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm, comprising a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and execute any method in the first aspect.

[0051] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and execute any one of the methods in the first aspect.

[0052] In summary, this application includes at least one of the following beneficial technical effects:

[0053] 1. In this application, the use of the TF-IDF weighted cosine similarity algorithm makes fault location more accurate. Compared with traditional methods, this algorithm can more accurately identify the source of the fault and reduce the possibility of misdiagnosis.

[0054] 2. Furthermore, this application uses fault injection technology, which can automatically expand the indicator vector library without manual management, intervention, or input of large amounts of fault information. This reduces the required human resources, eases the maintenance and management burden, and speeds up fault diagnosis.

[0055] 3. Furthermore, traditional troubleshooting may take a lot of time to check and eliminate each layer, while the application adopts a layered deduction and measurement method to locate faults in sequence, reducing the time required for troubleshooting and helping to quickly determine the root cause of the problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0057] Figure 1 This is a flow chart of a method for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm disclosed in an embodiment of the present application.

[0058] Figure 2 This is a flow chart of a rapid fault definition system based on the TF-IDF weighted cosine similarity algorithm disclosed in an embodiment of the present application.

[0059] Explanation of the accompanying symbols: 1. Fault injection module; 2. Fault definition module; 3. Solution push module. DETAILED DESCRIPTION

[0060] The following is combined with Figure 1-2 This application is described in further detail.

[0061] An embodiment of the present application discloses a method for rapid fault definition based on a TF-IDF weighted cosine similarity algorithm (hereinafter referred to as the rapid fault definition method). This method continuously injects faults using a preset fault injection technology (such as the Chaos Mesh fault injection technology), collects benchmark data corresponding to preset fault indicators from the injected faults, vectorizes the benchmark data using TF-IDF, and then adds the benchmark data to an indicator library, so that the benchmark data in the indicator library is continuously expanded as the number of injected faults increases. When a fault actually occurs, the rapid fault definition method disclosed in this application uses the TF-IDF weighted cosine similarity algorithm to screen out benchmark data similar to the current fault data from the indicator library, thereby defining the fault. This achieves the following effect: by expanding the amount of benchmark data in the indicator library available for comparison with the fault data, it helps to understand different types of faults, provides more support for fault diagnosis, and helps improve the success rate of fault definition.

[0062] Specifically, the execution body of the above-mentioned fault rapid definition method is a fault rapid definition system based on the TF-IDF weighted cosine similarity algorithm (hereinafter referred to as the fault definition system). Figure 1 Specifically describe the specific process steps of the fault definition system to implement the rapid fault definition method:

[0063] S101 , injecting a fault based on a pre-built fault simulation model. During the fault injection process, collecting benchmark data corresponding to preset fault indicators, and storing the benchmark data in a preset indicator library.

[0064] In practice, the embodiments of the present application use the Chaos Mesh platform to inject faults into a preset distributed system. Chaos Mesh can implement comprehensive fault injection, covering various complex system components on Kubernetes. Accordingly, the injected fault types may include Pod failures, network failures, file I / O failures, kernel failures, HTTP failures, etc. Different fault types correspond to different fault behaviors (for example, the fault behaviors corresponding to network failures may include: loss, delay, etc., and the fault behaviors corresponding to HTTP failures may include: response abort, request delay, etc.). In addition, the rapid fault definition system can also customize fault types to meet the needs of various fault scenarios.

[0065] Fault injection can be performed using Chaos Mesh's default input methods, including the following: 1. Using command-line tools such as kubectl; 2. Using the client; and 3. Using the Chaos Dashboard WebUI to operate and observe chaos experiments. Furthermore, fault injection can be scheduled (e.g., real-time) or manually triggered.

[0066] During the fault injection process, the fault demarcation system will call multiple monitoring sources (such as promtheus, a distributed data endpoint monitoring system; ELK log analysis system, etc.) to monitor and capture indicator data (such as performance data and status data) of each component in the preset distributed system.

[0067] Optionally, the preset fault indicators can be application-level indicators and container-level indicators. Application-level indicators can be collected from the reference logic of the service process and mainly reflect the service quality. The corresponding indicator data may include: average latency over the past 10 seconds, number of requests processed over the past second, etc. Container-level indicators are collected from the running environment of the service container and mainly reflect the virtualized resource usage of the service. The corresponding indicator data may include: average CPU time percentage over the past minute, current memory usage of the working machine, etc.

[0068] The fault definition system processes the aforementioned indicator data to form benchmark data. It then vectorizes the benchmark data based on a preset TF-IDF algorithm and stores it in a preset indicator library. Specifically, data processing can include performing word segmentation on the text-based indicator data and pre-processing to remove illegal characters, thereby representing the benchmark data as a word set containing keywords (keywords can be words or characters) corresponding to the preset fault indicators. In this embodiment, the indicator library is pre-divided into multiple indicator vector libraries based on hierarchy, and different indicator vector libraries are used to store vector-based benchmark data corresponding to different types of faults. Each indicator vector library has a corresponding label attribute that distinguishes it from other indicator vector libraries, namely, the Iaas layer indicator vector library, the Paas layer indicator vector library, the Saas layer indicator vector library, and the Unknown indicator vector library. This enables categorized storage of the benchmark data and facilitates differentiation of injected faults when multiple faults are injected simultaneously.

[0069] S102 , receiving a fault definition instruction, and comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on a pre-built fault definition model, and determining the similarity.

[0070] The following steps are included before “comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity” in S102:

[0071] Based on the preset TF-IDF algorithm, the benchmark data and fault data are converted into vector form.

[0072] The step of “comparing the fault data in the fault definition instruction with the benchmark data in the index library one by one and determining the similarity” in S102 specifically includes:

[0073] Based on the preset cosine similarity algorithm, the vectorized benchmark data and fault data are compared and the cosine similarity is calculated.

[0074] In practice, when a fault actually occurs, a fault definition instruction is triggered. This instruction contains at least the fault data corresponding to the preset fault indicators. When the fault definition instruction is generated, the fault definition system first converts both the baseline data and the fault data into vector form using the TF-IDF algorithm. It then uses a preset cosine similarity algorithm to calculate the cosine of the angle between the two vectors. This cosine of the angle is then used to determine the similarity between the baseline data and the fault data. (The TF-IDF algorithm, the cosine similarity algorithm, and their corresponding formulas are all prior art and will not be elaborated on here.)

[0075] Specifically, the value range of the cosine value of the angle is [-1, 1], where 1 indicates complete similarity, -1 indicates complete dissimilarity, and 0 indicates no linear correlation. Therefore, the closer the cosine similarity is to 1, the more similar the two vectors are. Therefore, in this application, the fault definition system will pre-store a similarity threshold, and the similarity threshold is a value in [-1, 1]. This application defaults to: when the cosine similarity obtained corresponding to the baseline data and the fault data is greater than the similarity threshold, the baseline data is considered to be similar to the fault data, and thus the fault corresponding to the baseline data and the fault data can be considered to be the same fault.

[0076] Optionally, in other embodiments, when comparing the vectorized baseline data with the fault data, the method for rapid fault definition further includes:

[0077] In a preset order, the fault data is compared with the benchmark data in each indicator vector library and the similarity is determined.

[0078] In practice, the indicator library has been divided into multiple indicator vector libraries (IaaS layer, PaaS layer, SaaS layer) based on the different layers of the fault component. Accordingly, when comparing benchmark data with fault data, the comparison order can also be limited. As mentioned above, the benchmark data in the indicator vector library is compared one by one, using the indicator vector library as a unit. For example, the benchmark data can be compared layer by layer in the order of IaaS layer > PaaS layer > SaaS layer > Unknown layer. That is, the benchmark data in the IaaS layer is compared first, and after the comparison is completed, the PaaS layer is compared...

[0079] Optionally, in addition to the above limitations on the comparison order, this application also proposes the following limitations:

[0080] Accordingly, the rapid fault definition method also includes:

[0081] Calculate the similarity between any two benchmark data and add the similarity to the preset comparison table.

[0082] The step of “comparing the fault data in the fault definition instruction with the benchmark data in the index library one by one and determining the similarity” in S102 specifically includes:

[0083] Filter out reference data from the indicator library. Reference data refers to benchmark data that meets preset conditions.

[0084] Compare the fault data with the reference data and calculate the reference similarity;

[0085] If the reference similarity is not less than the preset similarity threshold, it means that the reference data is similar to the fault data, and the reference data is used as the target benchmark data;

[0086] If the reference similarity is lower than the preset similarity threshold, then the similarity corresponding to the other benchmark data except the reference data is subtracted from the reference similarity in combination with the preset comparison table to obtain the similarity difference;

[0087] Based on the similarity difference and multiple preset comparison echelons, other benchmark data except the reference data are added to the corresponding comparison echelon. Each comparison echelon has a pre-defined difference range and a comparison priority. All benchmark data in the comparison echelon with a higher priority will be compared with the fault data first.

[0088] Based on the priority, the benchmark data in the comparison echelon is compared with the fault data in sequence to determine the similarity.

[0089] In practice, a preset comparison table is used to store the similarity between any two reference data sets. The "any two reference data sets" mentioned here can refer to any two reference data sets belonging to the same indicator vector library. When fault data needs to be compared with the reference data set one by one, the fault confinement system selects reference data from the indicator library. The reference data can be any one or more reference data sets randomly selected from the indicator library by the fault confinement system. The fault definition system calculates the cosine similarity between the reference data and the fault data (i.e., the reference similarity mentioned above). When the reference similarity is not less than the similarity threshold, it means that the reference data is similar to the fault data; when the reference similarity is less than the similarity threshold, it means that the reference data is not similar to the fault data. At this time, the fault definition system will combine the comparison table to calculate the difference between the reference similarity and the similarity in the comparison table (i.e., the similarity difference), and use the similarity difference to realize the classification of all benchmark data. Specifically, all benchmark data with similarity differences in the same preset difference range are added to the same comparison echelon. Each comparison echelon corresponds to a priority. The higher the priority of the comparison echelon, the benchmark data contained in it will be compared with the fault data first and the cosine similarity will be calculated. It can be seen that the solution mentioned in this application is mainly based on the following idea: if the similarity between A (fault data) and B (reference data) is X, the similarity between C (benchmark data) and B (reference data) is Y, and X and Y are similar (that is, the similarity difference is small), then it is very likely that C and A are similar. Therefore, C and A can be compared and the cosine similarity can be calculated first to verify whether C and A are truly similar, so as to improve the efficiency of finding benchmark data similar to the fault data.

[0090] S103, based on the similarity, push a processing solution;

[0091] Among them, S103 specifically includes:

[0092] If there is target benchmark data similar to the fault data, the preset processing solution corresponding to the target benchmark data is pushed;

[0093] If it does not exist, the fault data will be added to the Unkown indicator vector library and a reminder message will be pushed.

[0094] In implementation, the target benchmark data refers to benchmark data whose cosine similarity value with the fault data is not less than a preset similarity threshold, that is, benchmark data similar to the fault data. If the target benchmark data exists, the fault definition system will push or display the corresponding processing solution of the target benchmark data to the user terminal based on the pre-stored correspondence table of each benchmark data and its corresponding processing solution. The processing solution specifically refers to the repair solution for the fault. If the target benchmark data does not exist, the fault definition system will add the fault data to the Unknown indicator vector library and push or display a reminder message similar to "No similar fault found" or "Unknown fault" to the user terminal. In other embodiments, when the user terminal solves and repairs the corresponding fault in the Unknown indicator vector library, the fault definition system can provide a way for manual entry of the processing solution and selection of the benchmark data corresponding to the processing solution. The user terminal can store and update the corresponding relationship between the corresponding processing solution and the benchmark data based on the aforementioned manual entry.

[0095] Optionally, before “if there is target reference data similar to the fault data, pushing a preset processing solution corresponding to the target reference data” in S103, the following steps may also be included:

[0096] Traverse all similarities in order and compare the similarities one by one with the preset similarity threshold;

[0097] If there is a first similarity that is not less than a preset similarity threshold, the reference data corresponding to the largest first similarity is used as the target reference data.

[0098] In practice, the first similarity refers to a similarity that is no less than a preset similarity threshold. Similarity refers to the cosine similarity between the fault data and any reference data. If multiple first similarities exist, the fault confinement system selects the reference data corresponding to the highest first similarity as the target reference data, prioritizing the reference data most similar to the fault data so that the final feedback solution successfully resolves the fault corresponding to the fault data.

[0099] Optionally, the fault confinement system further presets a number of concurrent fault sets, each of which contains at least a number of benchmark data, and all benchmark data in the same concurrent fault set satisfy the following conditions: when a fault corresponding to one of the benchmark data occurs, it will trigger faults corresponding to other benchmark data in the corresponding concurrent fault set;

[0100] The method for rapid fault definition disclosed in this application also includes:

[0101] After the target benchmark data is determined, if the target benchmark data exists in any concurrent fault set, a new priority comparison set is created. The priority comparison set contains all the benchmark data in the concurrent fault set where the target benchmark data is located, and a validity period is set. After the validity period of the new set, the priority comparison set is deleted.

[0102] The step of “comparing the fault data in the fault definition instruction with the benchmark data in the index library one by one and determining the similarity” in S102 specifically includes:

[0103] If there is a priority comparison set, the fault data is first compared with all the benchmark data in the priority comparison set and the similarity is determined; if no target benchmark data similar to the fault data is found, the fault data is then compared with all other benchmark data except the priority comparison set and the similarity is determined for each of them;

[0104] If there is no priority comparison set, the fault data will be compared with the benchmark data in the indicator library one by one, and the similarity will be determined.

[0105] In implementation, a concurrent fault set contains several benchmark data, and all benchmark data stored in the same concurrent fault set meet the following requirements: when a fault corresponding to one of the benchmark data occurs, within a preset effective time, it is highly likely to trigger the faults corresponding to other benchmark data in the concurrent fault set.

[0106] Therefore, whenever the fault confinement system finds target benchmark data corresponding to fault data, it confirms whether the target benchmark data exists in any concurrent fault set. If so, it starts counting and takes the union of the benchmark data in all concurrent fault sets containing the target benchmark data to form a priority comparison set. That is, the benchmark data in the priority comparison set is the benchmark data contained in all concurrent fault sets containing the target benchmark data. The priority comparison set is a newly generated set that is automatically deleted after its validity period. The creation or deletion of a priority comparison set does not affect all concurrent fault sets; all concurrent fault sets and the benchmark data they contain always exist.

[0107] The fault definition system uses a priority comparison set to divide all benchmark data into two categories. If a fault definition instruction is received during the timing process, the fault data in the fault definition instruction will be compared with the benchmark data in the priority comparison set first. After the comparison is completed, if the target benchmark data does not exist, the fault data will be compared with other benchmark data outside the priority comparison set to determine whether the target benchmark data exists. During this process, if the timing reaches the preset effective time, the fault definition system will delete the priority comparison set and compare the benchmark data with the fault data one by one according to the comparison rules described above (for example, the benchmark data in each indicator vector library will be compared with the fault data one by one in sequence). This further improves the search efficiency of the target benchmark data.

[0108] The present application also discloses a fault rapid definition system based on the TF-IDF weighted cosine similarity algorithm. Figure 2 The fault rapid definition system based on the TF-IDF weighted cosine similarity algorithm includes:

[0109] Fault injection module 1 injects faults based on a pre-built fault simulation model. During the fault injection process, it collects benchmark data corresponding to preset fault indicators and stores the benchmark data in a preset indicator library.

[0110] Fault definition module 2 is used to receive a fault definition instruction, compare the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on the pre-built fault definition model, and determine the similarity;

[0111] The solution pushing module 3 is used to push a processing solution based on similarity; specifically, if there is target reference data similar to the fault data, then push the preset processing solution corresponding to the target reference data.

[0112] Optionally, the fault delimitation module 2 is further configured to convert the reference data and the fault data in the fault delimitation instruction into vector form, and to compare the vectorized reference data and the fault data based on a preset cosine similarity algorithm, and to calculate the cosine similarity.

[0113] Optionally, the fault definition module 2 is further configured to convert the benchmark data and the fault data into a vector form based on a preset TF-IDF algorithm.

[0114] Optionally, the indicator library includes several indicator vector libraries, and different indicator vector libraries are used to store benchmark data in vector form corresponding to different types of faults; each indicator vector library corresponds to a label attribute used to distinguish other indicator vector libraries, and the processing plan includes at least the label attribute; the fault definition module is also used to compare the fault data with the benchmark data in each indicator vector library in a preset order and determine the similarity.

[0115] Optionally, a similarity comparison module is used to calculate the similarity between any two benchmark data and add the similarity to a preset comparison table.

[0116] The fault definition module 2 is also used to filter out reference data from the indicator library, where reference data refers to benchmark data that meets preset conditions; it is also used to compare the fault data with the reference data and calculate the reference similarity; if the reference similarity is not less than the preset similarity threshold, it means that the reference data is similar to the fault data, and the reference data is used as the target benchmark data; if the reference similarity is lower than the preset similarity threshold, then in combination with the preset comparison table, the similarity corresponding to other benchmark data other than the reference data is subtracted from the reference similarity to obtain the similarity difference; based on the similarity difference and the preset multiple comparison echelons, the other benchmark data other than the reference data is added to the corresponding comparison echelon; wherein each comparison echelon is pre-defined with a difference range and a comparison priority, and all benchmark data in the comparison echelon with a higher priority are compared with the fault data first; it is also used to compare the benchmark data in the comparison echelon with the fault data in sequence based on the priority and determine the similarity.

[0117] Optionally, the solution pushing module 3 is further configured to sequentially traverse all similarities and compare the similarities with a preset similarity threshold one by one; and to use the benchmark data corresponding to the largest first similarity as the target benchmark data if there is a first similarity not less than the preset similarity threshold.

[0118] Optionally, several concurrent fault sets are preset, each concurrent fault set contains at least several benchmark data, and all benchmark data belonging to the same concurrent fault set meet the following requirements: when a fault corresponding to one of the benchmark data occurs, it will trigger the faults corresponding to other benchmark data in the corresponding concurrent fault set.

[0119] It also includes a concurrent fault determination module, which is used to, after determining the target benchmark data, if the target benchmark data exists in any concurrent fault set, take the union of all benchmark data in the concurrent fault set where the target benchmark data is located to generate a priority comparison set, and set the effective time; and after the effective time, delete the priority comparison set.

[0120] The fault definition module 2 is also used to, if there is a priority comparison set, first compare the fault data with all the benchmark data in the priority comparison set and determine the similarity; if no target benchmark data similar to the fault data is found, then compare the fault data with all other benchmark data except the priority comparison set and determine the similarity respectively; and if there is no priority comparison set, then compare the fault data with the benchmark data in the indicator library one by one and determine the similarity.

[0121] An embodiment of the present application also discloses a device for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm. The device for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm includes a memory and a processor. The memory stores a computer program that can be loaded by the processor and execute the rapid fault definition method based on the TF-IDF weighted cosine similarity algorithm as described above.

[0122] An embodiment of the present application further discloses a computer-readable storage medium storing a computer program that can be loaded by a processor and execute the above-mentioned method for rapid fault definition based on the TF-IDF weighted cosine similarity algorithm. The computer-readable storage medium includes, for example, various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0123] It should be noted that, in this document, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0124] The above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of protection of this application. Obviously, the embodiments described are only some of the embodiments of this application, not all of them. Based on these embodiments, all other embodiments obtained by persons of ordinary skill in the art without inventive effort are also within the scope of protection to be protected by this application.

Claims

1. A method for rapid fault definition based on TF-IDF weighted cosine similarity algorithm, characterized in that: include: Inject faults based on a pre-built fault simulation model. During the fault injection process, collect benchmark data corresponding to preset fault indicators and store the benchmark data in a preset indicator library; the preset fault indicators are specifically application-level indicators and container-level indicators; the application-level indicators are collected from the reference logic of the service process to reflect the service quality; the container-level indicators are collected from the operating environment of the service container to reflect the virtualized resource occupancy of the service; receiving a fault definition instruction, and comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on a pre-built fault definition model, and determining the similarity; Based on the similarity, push a processing solution; Wherein, the push processing solution based on the similarity at least includes: If there is target reference data similar to the fault data, the preset processing scheme corresponding to the target reference data is pushed; there are several concurrent fault sets preset, each of which contains at least several reference data, and all reference data belonging to the same concurrent fault set meet the following conditions: when a fault corresponding to one of the reference data occurs, it will trigger the faults corresponding to other reference data in the corresponding concurrent fault set; The method further comprises: After determining the target benchmark data, if the target benchmark data exists in any concurrent fault set, a new priority comparison set is created, the priority comparison set includes all benchmark data in the concurrent fault set where the target benchmark data is located, and the validity period is set; and after the new validity period, the priority comparison set is deleted; The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes: If there is a priority comparison set, the fault data is first compared with all the benchmark data in the priority comparison set, and the similarity is determined; if no target benchmark data similar to the fault data is found, the fault data is then compared with all other benchmark data except the priority comparison set, and the similarity is determined respectively; If there is no priority comparison set, the fault data is compared with the benchmark data in the indicator library one by one, and the similarity is determined; The method further comprises: Calculate the similarity between any two benchmark data and add the similarity to a preset comparison table; The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes: Filtering reference data from the indicator library, wherein the reference data refers to benchmark data that meets preset conditions; Comparing the fault data with the reference data, and calculating a reference similarity; If the reference similarity is not less than a preset similarity threshold, it means that the reference data is similar to the fault data, and the reference data is used as target benchmark data; If the reference similarity is lower than a preset similarity threshold, then in combination with a preset comparison table, the similarity corresponding to other benchmark data other than the reference data is subtracted from the reference similarity to obtain a similarity difference; Based on the similarity difference and the preset multiple comparison echelons, other benchmark data except the reference data are added to the corresponding comparison echelon; wherein each comparison echelon is pre-defined with a difference range and a comparison priority, and all benchmark data in the comparison echelon with a higher priority have a higher priority to be compared with the fault data; Based on the priority, the benchmark data in the comparison echelon is compared with the fault data in sequence to determine the similarity.

2. The method for rapid fault definition based on TF-IDF weighted cosine similarity algorithm according to claim 1, characterized in that: The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity may include: Converting the reference data and the fault data in the fault definition instruction into a vector form; The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes: Based on the preset cosine similarity algorithm, the vectorized benchmark data and fault data are compared, and the cosine similarity is calculated.

3. The method for rapid fault definition based on TF-IDF weighted cosine similarity algorithm according to claim 2, characterized in that: The reference data and the fault data in the fault definition instruction are converted into a vector form, including: Based on the preset TF-IDF algorithm, the benchmark data and fault data are converted into vector form.

4. The method for rapid fault definition based on TF-IDF weighted cosine similarity algorithm according to claim 1, characterized in that: The indicator library includes several indicator vector libraries, and different indicator vector libraries are used to store benchmark data in vector form corresponding to different types of faults; each indicator vector library corresponds to a label attribute for distinguishing other indicator vector libraries, and the processing scheme includes at least the label attribute; The step of comparing the fault data in the fault definition instruction with the benchmark data in the indicator library one by one and determining the similarity includes: The fault data is compared with the benchmark data in each indicator vector library in a preset order, and the similarity is determined.

5. The method for rapid fault definition based on TF-IDF weighted cosine similarity algorithm according to claim 1, characterized in that: If there is target reference data similar to the fault data, a preset processing solution corresponding to the target reference data is pushed, which includes: Traversing all similarities in order, and comparing the similarities with the preset similarity thresholds one by one; If there is a first similarity that is not less than a preset similarity threshold, the reference data corresponding to the largest first similarity is used as the target reference data.

6. A fault rapid definition system based on TF-IDF weighted cosine similarity algorithm, characterized in that: include, A fault injection module (1) injects a fault based on a pre-built fault simulation model. During the fault injection process, benchmark data corresponding to a preset fault indicator is collected and stored in a preset indicator library. The preset fault indicator is specifically an application-level indicator and a container-level indicator. The application-level indicator is collected from the reference logic of the service process and is used to reflect the service quality. The container-level indicator is collected from the operating environment of the service container and is used to reflect the virtualized resource occupancy of the service. A fault definition module (2) is used to receive a fault definition instruction, compare the fault data in the fault definition instruction with the benchmark data in the indicator library one by one based on a pre-built fault definition model, and determine the similarity; A solution pushing module (3) is used to push a processing solution based on the similarity; specifically, if there is target reference data similar to the fault data, push a preset processing solution corresponding to the target reference data; There are several concurrent fault sets preset, each of which contains at least several benchmark data, and all benchmark data belonging to the same concurrent fault set meet the following conditions: when a fault corresponding to one of the benchmark data occurs, the faults corresponding to other benchmark data in the corresponding concurrent fault set will be triggered; The system also includes a concurrent fault determination module, which is used for, after determining and obtaining the target benchmark data, if the target benchmark data exists in any concurrent fault set, taking the union of all benchmark data in the concurrent fault set where the target benchmark data is located to generate a priority comparison set, and setting a valid time; and after the valid time, deleting the priority comparison set; The fault definition module (2) is further configured to, if a priority comparison set exists, first compare the fault data with all the benchmark data in the priority comparison set and determine the similarity; if no target benchmark data similar to the fault data is found, then compare the fault data with all other benchmark data except the priority comparison set and determine the similarity respectively; and if no priority comparison set exists, then compare the fault data with the benchmark data in the indicator library one by one and determine the similarity; the system further comprises a similarity comparison module, configured to calculate the similarity of any two benchmark data and add the similarity to a preset comparison table; The fault definition module (2) is also used to select reference data from the index library, the reference data refers to benchmark data that meets preset conditions; compare the fault data with the reference data, and calculate the reference similarity; If the reference similarity is not less than a preset similarity threshold, it means that the reference data is similar to the fault data, and the reference data is used as target benchmark data; If the reference similarity is lower than a preset similarity threshold, then in combination with a preset comparison table, the similarity corresponding to other benchmark data other than the reference data is subtracted from the reference similarity to obtain a similarity difference; Based on the similarity difference and multiple preset comparison echelons, other benchmark data except the reference data are added to the corresponding comparison echelon; wherein each comparison echelon is pre-defined with a difference range and a comparison priority, and all benchmark data in the comparison echelon with a higher priority have higher priority in comparison with the fault data; based on the priority, the benchmark data in the comparison echelon are compared with the fault data in sequence to determine the similarity.

7. A fault rapid definition device based on TF-IDF weighted cosine similarity algorithm, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executes any one of the methods of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Text vectorization-based handling reference method in fault power failure first-aid repair event

    CN112711947A

  • Converter station fault strategy model training method and device, and converter station fault strategy model pushing method and device

    CN117009516A