Data labeling method and device for equipment log file fault traceability

By automating the processing of device log files through algorithms and utilizing the workspace window to statistically analyze the frequency and distance correlation of log data, the problem of low efficiency in manual annotation and inflexibility in artificial intelligence annotation is solved, enabling efficient and interpretable log file fault tracing.

CN120929833APending Publication Date: 2025-11-11ZHENGZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511043582.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In existing technologies, manual annotation of equipment log file faults is inefficient and prone to errors, while artificial intelligence annotation methods are inflexible and lack interpretability.

Method used

The device log files are processed automatically through algorithms. The frequency and distance correlation of log data are statistically analyzed using the workspace window. The positive, negative and comprehensive correlations between logs are calculated and automatically labeled.

Benefits of technology

It improves data processing efficiency, reduces manual annotation costs, can quickly adapt to new data, provides interpretable results, facilitates manual verification, and increases the credibility and transparency of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929833A_ABST
    Figure CN120929833A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of log file data analysis, and particularly relates to a data labeling method and device for equipment log file fault tracing. The method comprises the following steps: S1, traversing a to-be-labeled log data set sorted according to a reverse time order by utilizing a set working area window until the last piece of data in the working area window is the last piece of data of the to-be-labeled log data set; counting the occurrence times of all the log data of the counted categories in the working area window, and accumulating the times that the target categories are the same and the counted categories are the same as corresponding frequency related parameters; and S2, calculating frequency correlation among the log data according to the frequency correlation parameters, and labeling the to-be-labeled log data set according to the frequency correlation. The technical problems that in the prior art, a manual labeling method is low in efficiency and prone to making mistakes, and an artificial intelligence labeling method is not flexible and lacks interpretability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of log file data analysis technology, specifically relating to a data annotation method and apparatus for tracing the source of equipment log file faults. Background Technology

[0002] Medical devices, power systems, and similar equipment generate a large amount of log data during operation, recording crucial information such as operating status, parameter changes, and abnormal situations. This log data not only reflects the equipment's operating patterns but also contains potential failure modes. Therefore, log data is frequently used for equipment fault diagnosis.

[0003] Log data is a primary source of information for fault diagnosis. By employing system logs, RAS logs, and other methods, fault characteristics are extracted from the logs for fault diagnosis. Data annotation is a crucial step in log-based fault diagnosis, processing unprocessed log information and transforming it into machine-readable data. Fault classification is typically done manually, labeling logs according to the actual fault type. However, manual annotation methods are inefficient, susceptible to subjective bias, require strict data privacy protection, demand highly skilled annotators, and are prone to errors.

[0004] Besides manual annotation methods, artificial intelligence algorithms are also used for annotation. For example, the Chinese invention patent application CN107301118A, published on October 27, 2017, discloses a log-based automatic fault indicator annotation method and system. This method trains a fault indicator model that can be automatically annotated by using the correspondence between equipment indicator data and log data, and completes automatic annotation using the trained fault indicator model. However, artificial intelligence annotation methods rely on powerful computing resources, have inflexible model updates, and lack interpretability. Summary of the Invention

[0005] The purpose of this invention is to provide a data annotation method and apparatus for tracing faults in device log files, so as to solve the technical problems of low efficiency and error-proneness of manual annotation methods in the prior art, and the inflexibility and lack of interpretability of artificial intelligence annotation methods.

[0006] To solve the above-mentioned technical problems, the present invention provides a data annotation method for fault tracing in equipment log files, the technical solution of which is: a data annotation method for fault tracing in equipment log files, the method comprising:

[0007] S1. Use the set workspace window to traverse the log dataset to be labeled, which is sorted in reverse chronological order, until the last data in the workspace window is the last data in the log dataset to be labeled.

[0008] Count the number of times log data of all the counted categories appear in the work area window, and sum up the number of times that the target category is the same and the counted category is also the same as the corresponding frequency-related parameter;

[0009] The category being counted is a category that is different from the target category within the workspace window; the target category is the category of the first log data entry within the workspace window.

[0010] S2. Calculate the frequency correlation between log data based on the frequency correlation parameters, and label the log dataset to be labeled based on the frequency correlation.

[0011] The beneficial effects of the above technical solution are as follows: The technical solution of the data annotation method for fault tracing of equipment log files of the present invention belongs to an improved invention. Compared with manual annotation: The present invention automates the processing of large amounts of data through algorithms, significantly reducing the time and cost required for manual annotation and improving the efficiency of data processing. Compared with artificial intelligence annotation: The annotation algorithm of the present invention does not need to retrain the model when processing new datasets, and can adapt to new data more quickly, saving time and resources; moreover, the correlation between logs obtained by the present invention is interpretable, which facilitates manual inspection and verification, increasing the credibility and transparency of the results. The present invention solves the technical problems of low efficiency and error-proneness of existing manual annotation methods, and the inflexibility and lack of interpretability of artificial intelligence annotation methods.

[0012] Furthermore, the correlation includes positive correlation and / or negative correlation;

[0013] The positive correlation between LogB category and LogA category indicates the degree of correlation between log data of LogB category and LogA category that occur sequentially.

[0014] The inverse correlation between LogB category and LogA category indicates the degree of correlation between log data of LogB category and LogA category that do not occur sequentially.

[0015] Furthermore, the positive correlation between LogB category and LogA category log data is the ratio of the frequency correlation parameter of the target category being LogA and the statistical category being LogB to the work area window size.

[0016] Furthermore, the inverse correlation of log data of category LogB relative to category LogA is obtained by subtracting the frequency correlation parameter of category LogB (target category LogA) from the number of times the log data of category LogB appears in the log dataset to be labeled, and then dividing by the frequency correlation parameter of category LogB (target category LogA) and category LogB.

[0017] Furthermore, the frequency correlation also includes a comprehensive correlation; the comprehensive correlation is calculated according to the following formula:

[0018] OverallRelevance LogA,LogB =p×Relevance LogA,LogB -q×ReverseRelevance LogA,LogB

[0019] Among them, OverallRelevance LogA,LogB The overall relevance of log data of category LogB relative to category LogA; ReverseRelevance LogA,LogB The inverse correlation of log data of category LogB relative to category LogA; q is the inverse correlation coefficient; Relevance LogA,LogB represents the positive correlation between LogB category and LogA category log data; p is the correlation coefficient.

[0020] Furthermore, S1 also includes: calculating the distances of all log data of the same category within the workspace window relative to the target log data, and accumulating the distances of the target log data that are of the same category and are also of the same category as the distance-related parameter; the target log data is the first log data in the workspace window;

[0021] S2 further includes: calculating the distance correlation between log data based on the distance correlation parameter, wherein the distance correlation is used to label the log dataset to be labeled.

[0022] Further, the distance correlation includes the average distance between log data of category LogB and category LogA; the average distance is calculated according to the following formula:

[0023]

[0024] Among them, Avg_Distance LogA,LogB The average distance of log data in category LogB relative to category LogA; Count LogA,LogB For parameters where the target category is LogA and the category being counted is LogB; Distance LogA,LogB This is a distance-related parameter for log data whose statistical category is LogB and whose target log data category is LogA.

[0025] Furthermore, S1 also includes: calculating the distances of all log data of the same category within the workspace window relative to the target log data, and accumulating the distances of the target log data that are of the same category and are also of the same category as the distance-related parameter; the target log data is the first log data in the workspace window;

[0026] S2 further includes: calculating the distance correlation between log data based on the distance correlation parameter, wherein the distance correlation is used to label the log dataset to be labeled;

[0027] The distance correlation includes the average distance between LogB category log data and LogA category log data; the average distance is calculated according to the following formula:

[0028]

[0029] Among them, Avg_Distance LogA,LogB The average distance of log data in category LogB relative to category LogA; Count LogA,LogB For parameters where the target category is LogA and the category being counted is LogB; Distance LogA,LogB This is a distance-related parameter for log data whose statistical category is LogB and whose target log data category is LogA.

[0030] Furthermore, the process of labeling the log dataset to be labeled based on the aforementioned correlation includes:

[0031] S21. Obtain the correlation between all log categories using the following method:

[0032] For log data of category LogA, calculate the comprehensive correlation and average distance of log data of other categories relative to log data of category LogA;

[0033] The set of corresponding categories with comprehensive correlation not lower than the correlation threshold is sorted according to the average distance to obtain the correlation relationship of LogA categories;

[0034] S22. Label the log dataset to be labeled according to the correlation.

[0035] The present invention also provides a technical solution for a data annotation device for tracing device log file faults: a data annotation device for tracing device log file faults includes a processor, the processor being used to execute a computer program to implement the steps of the data annotation method for tracing device log file faults as described above. Attached Figure Description

[0036] Figure 1This is an algorithm flowchart illustrating an implementation method for data annotation in fault tracing of equipment log files according to the present invention.

[0037] Figure 2 This is a schematic diagram illustrating the implementation of the data annotation method for fault tracing in device log files according to the present invention. Detailed Implementation

[0038] Compared to manual annotation: This invention automates the processing of large amounts of data through algorithms, significantly reducing the time and cost required for manual annotation and improving data processing efficiency. Compared to AI annotation: The annotation algorithm of this invention does not require retraining the model when processing new datasets, enabling it to adapt to new data more quickly and saving time and resources; moreover, the correlations between the logs obtained by this invention are interpretable, facilitating manual inspection and verification, increasing the credibility and transparency of the results. This invention solves the technical problems of low efficiency and error-proneness in existing manual annotation methods, and the inflexibility and lack of interpretability in AI annotation methods.

[0039] Implementation method of data annotation for fault tracing in equipment log files:

[0040] The data source for this implementation is the medical device log file dataset (LogDataset, i.e., the log dataset to be labeled), which contains LogDatasetSize log data entries, and the data is sorted in reverse order according to the log occurrence time.

[0041] Data preprocessing: Cleaning and standardizing log data to remove invalid or duplicate data and ensure data integrity and consistency.

[0042] Data storage: The processed log data is stored in a database for subsequent processing and analysis.

[0043] Parameter settings: Set the size of the algorithm workspace (i.e., the workspace window): WorkSize; Set the log correlation threshold: RelevanceThreshold; Set the dictionary of correlation parameters between logs: RelevanceDict.

[0044] like Figure 1 As shown, the data annotation method for fault tracing in the device log file of this embodiment includes the following steps:

[0045] 1. Start executing the following algorithm from the first record in the LogDataset:

[0046] 1) Define the index of the first record in the current workspace in the LogDataset as StartIndex = 0;

[0047] 2) The data in the current workspace is LogDataset[StartIndex:StartIndex+WorkSize-1], which is a subset of LogDataset consisting of data from StartIndex to StartIndex+WorkSize.

[0048] 3) Count the occurrences of log data categories (i.e., the categories being counted) other than the category of the first data entry (i.e., the target category) within the statistical work area. TatgetUID,SourceUID Distance of log data in all categories except the first log entry from the workspace. TatgetUID,SourceUID The data is accumulated and recorded in the RelevanceDict dictionary, where TatgetUID represents the category identifier of the first log data in the workspace, and SourceUID is the category identifier of the currently analyzed log.

[0049] Count TatgetUID,SourceUID This indicates the number of times the log data of the SourceUID category appears in the workspace when the category identifier of the first log data in the workspace is TatgetUID, that is, the number of times the SourceUID category appears relative to the TatgetUID category;

[0050] Distance TatgetUID,SourceUID This indicates the distance between the first log data in the workspace with the category identifier TatgetUID and the first log data in the workspace with the corresponding category SourceUID, i.e., the distance between the SourceUID category and the TatgetUID category.

[0051] Because the log data of this invention is arranged in chronological order, therefore, Distance TatgetUID,SourceUID The actual physical meaning of the representation is: the time between the event corresponding to the TatgetUID category log and the event corresponding to the SourceUID category log, and the event corresponding to the TatgetUID category log is earlier than the event corresponding to the SourceUID category log.

[0052] 4) Set StartIndex = StartIndex + 1, and repeat steps 1) to 4) until the following expression is satisfied.

[0053] StartIndex+WorkSize-1>LogDatasetSize-1

[0054] The above formula means that the index StartIndex+WorkSize-1 of the last log data in the work area exceeds the index LogDatasetSize-1 of the last log data in the LogDataset dataset; indicating that all LogDatasetSize data in the LogDataset dataset have been statistically analyzed through the work area.

[0055] 2. Calculate the positive correlation and average distance between logs:

[0056] Positive correlation is used to represent the degree of correlation between log data of categories LogB and LogA that occur sequentially. It is the frequency correlation between category LogB and category LogA, and is expressed as the positive correlation between log data of category LogB and category LogA. LogA,LogB The average distance represents the distance between the two types of logs in the dataset, which can reflect the order in which logs occur to some extent, but cannot accurately represent the interval between log occurrences.

[0057] For a given log entry, let its UID (i.e., category) be LogA. The methods for calculating the positive correlation and average distance between the other log entries and the log entry itself are as follows:

[0058] Search the RelevanceDict dictionary for all Count and Distance values ​​with the index TatgetUID "LogA". The set of all possible values ​​for SourceUID in the found Count and Distance values ​​is the set of UIDs for all other logs related to LogA.

[0059] For an element LogB in the above set, the correlation between LogB and LogA, that is, the positive correlation between log data of category LogB and category LogA, is as follows:

[0060]

[0061] Where, represents the positive correlation between LogB category and LogA category log data; WorkSize is the size of the work area; Count LogA,LogB This represents the number of times the LogB category appears relative to the LogA category in the correlation parameter dictionary.

[0062] The average distance from LogB to LogA, that is, the average distance of log data of category LogB relative to category LogA, is:

[0063]

[0064] Among them, Avg_Distance LogA,LogB The average distance of log data in category LogB relative to category LogA; Count LogA,LogB The number of times the LogB category appears relative to the LogA category in the correlation parameter dictionary; Distance LogA,LogB This represents the distance of log data of category LogB relative to category LogA in the correlation parameter dictionary.

[0065] 3. Calculate the inverse correlation between logs:

[0066] Inverse correlation is used to represent the degree of correlation between log data of categories LogB and LogA that do not occur sequentially. Positive correlation only represents the association between other categories relative to category LogA. If log data of category LogB is frequently and evenly distributed in the LogDataset dataset, when calculating positive correlation, the following will occur: the positive correlation between category LogB and category LogA is very high, and the positive correlation between category LogB and other categories is also very high. In this case, the category that is truly associated with LogA will be covered by LogB, making it difficult to distinguish the category that is truly associated with LogA.

[0067] For a given log entry, let its UID (i.e., category) be LogA. The methods for calculating the positive correlation and average distance between the other log entries and the log entry itself are as follows:

[0068] Search the RelevanceDict dictionary for all Count and Distance values ​​with the index TatgetUID "LogA". The set of all possible values ​​for SourceUID in the found Count and Distance values ​​is the set of UIDs for all other logs related to LogA.

[0069] For an element LogB in the above set, calculate the number of times log data of category LogB appears in the entire LogDataset, denoted as AllCount. LogB The inverse correlation between LogB and LogA, that is, the inverse correlation of log data of category LogB relative to category LogA, is as follows:

[0070]

[0071] Among them, ReverseRelevance LogA,LogB This represents the inverse correlation of log data of category LogB relative to category LogA; Count LogA,LogBAllCount represents the number of times the LogB category appears relative to the LogA category in the correlation parameter dictionary. LogB The number of times log data of the LogB category appears in the log dataset to be labeled.

[0072] 4. Calculate the overall correlation between logs:

[0073] Given two log entries with UIDs LogA and LogB respectively, what is the overall correlation between log data of category LogB and log data of category LogA?

[0074] OverallRelevance LogA,LogB =p×Relevance LogA,LogB -q×ReverseRelevance LogA,LogB

[0075] Among them, OverallRelevance LogA,LogB The overall relevance of log data of category LogB relative to category LogA; ReverseRelevance LogA,LogB The inverse correlation of log data of category LogB relative to category LogA; q is the inverse correlation coefficient; Relevance LogA,LogB represents the positive correlation between LogB category and LogA category log data; p is the correlation coefficient.

[0076] 5. Relationships between output logs:

[0077] For a given log entry, let its UID be LogA. Find all log entries with the index SourceUID of LogA and their overall relevance (OverallRelevance). Remove entries with a relevance value less than the RelevanceThreshold (the relevance threshold). The set of all remaining values ​​for the index SourceUID represents the set of other UIDs with high relevance to log entries of the LogA category. Sort these UIDs by their average distance (Avg_Distance) to obtain the correlation between the log entries.

[0078] 6. Label the LogDataset according to the relationships between the logs mentioned above.

[0079] 7. Visualization of data results:

[0080] Use data visualization tools (such as Matplotlib, Tableau, etc.) to visualize the correlations and annotation results between logs on medical devices.

[0081] Visualized content:

[0082] Log Relationship Graph: Displays the correlation between logs, representing logs and their relationships in the form of nodes and edges.

[0083] Annotation results chart: Displays the distribution of annotation results, showing the number of logs in different categories in the form of bar chart or pie chart.

[0084] Trend Analysis Chart: Displays the trend of log data over time, showing the dynamic changes of log data in the form of a line chart.

[0085] The data annotation method for fault tracing in device log files described in this embodiment is based on the following modules, which can be either software or hardware modules, such as... Figure 2 As shown.

[0086] Data acquisition module: Located at the front end of the device, it is responsible for collecting log file data from the medical equipment. It acquires log data via a network interface or by connecting directly to the medical equipment.

[0087] Data storage module: Located at the center of the device, it is connected to the data acquisition module via an internal data bus. It receives log data from the data acquisition module and stores it in the database.

[0088] Data analysis module: Located at the back end of the device, it connects to the data storage module via an internal data bus. It reads log data from the data storage module and performs related analysis and calculations.

[0089] Data visualization module: Located at the front of the device, it connects to the data analysis module via an internal data bus. It receives the analysis results from the data analysis module and transforms them into visual charts and graphs for display to the user.

[0090] After data annotation is completed, it can not only be visualized as described above to assist engineers and other professionals in fault diagnosis, but it can also be used to build fault diagnosis models. That is, based on the annotated log data, automated fault diagnosis can be achieved through existing machine learning algorithms suitable for fault diagnosis. The specific implementation process can be referred to existing technologies, and will not be described in detail in this embodiment.

[0091] Implementation method of data annotation device for fault tracing in equipment log files:

[0092] A data annotation apparatus for tracing device log file faults includes a processor, which executes a computer program to implement the steps of the data annotation method for tracing device log file faults as described above. The specific data annotation method for tracing device log file faults has been described in sufficient detail in the above-described embodiments and will not be repeated here.

[0093] Specifically, a processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. A processor can also be a processor that supports the Advanced Reduced Instruction Set Machine (ARM) architecture.

[0094] The present invention has the following advantages:

[0095] 1. Improved efficiency: This invention automates the processing of large amounts of data through algorithms, significantly reducing the time and cost required for manual annotation and improving the efficiency of data processing.

[0096] 2. Reduced hardware costs: This invention has low computational resource requirements and can run in resource-constrained environments, thus reducing hardware costs. The algorithm has a fast processing speed, enabling rapid annotation of large amounts of data to meet rapidly growing data demands.

[0097] 3. Improve the credibility and transparency of the results: The correlation between logs obtained by this invention is interpretable, which facilitates manual inspection and verification, thereby increasing the credibility and transparency of the results.

[0098] 4. Rapid Adaptation to New Data: This invention eliminates the need to retrain the model when processing new datasets, enabling faster adaptation and saving time and resources. The algorithm requires minimal computing resources and can operate in resource-constrained environments. Automated processing eliminates the need for annotators with advanced expertise, reducing the demands on them. This allows for wider adoption and scaling of annotation work.

[0099] 5. Improved Annotation Quality: This invention avoids the subjective influence of manual annotation through objective data statistics and calculations, thus improving the consistency and accuracy of the annotation results. The algorithm is based on objective data statistics and calculations, avoiding the subjective influence of manual annotation. The consistency and accuracy of the annotation results are guaranteed.

[0100] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still make modifications to the technical solutions described in the foregoing embodiments without creative effort, or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data annotation method for fault tracing in equipment log files, characterized in that, The method includes: S1. Use the set workspace window to traverse the log dataset to be labeled, which is sorted in reverse chronological order, until the last data in the workspace window is the last data in the log dataset to be labeled. Count the number of times log data of all the counted categories appear in the work area window, and sum up the number of times that the target category is the same and the counted category is also the same as the corresponding frequency-related parameter; The category being counted is a category that is different from the target category within the workspace window; the target category is the category of the first log data entry within the workspace window. S2. Calculate the frequency correlation between log data based on the frequency correlation parameters, and label the log dataset to be labeled based on the frequency correlation.

2. The data annotation method for fault tracing in equipment log files according to claim 1, characterized in that, The frequency correlation includes positive correlation and / or negative correlation; The positive correlation between LogB category and LogA category indicates the degree of correlation between log data of LogB category and LogA category that occur sequentially. The inverse correlation between LogB category and LogA category indicates the degree of correlation between log data of LogB category and LogA category that do not occur sequentially.

3. The data annotation method for fault tracing in equipment log files according to claim 2, characterized in that, The positive correlation between LogB category and LogA category log data is the ratio of the frequency correlation parameter of the target category LogA and the statistical category LogB to the work area window size.

4. The data annotation method for fault tracing in equipment log files according to claim 2, characterized in that, The inverse correlation of LogB category log data relative to LogA category is obtained by subtracting the frequency correlation parameter of LogB category log data from the number of times it appears in the log dataset to be labeled from the frequency correlation parameter of LogA category and LogB category, and then dividing by the frequency correlation parameter of LogA category and LogB category.

5. The data annotation method for fault tracing in equipment log files according to claim 2, 3, or 4, characterized in that, The frequency correlation also includes a comprehensive correlation; the comprehensive correlation is calculated according to the following formula: OverallRelevance LogA,LogB =p×Relevance LogA,LogB -q×ReverseRelevance LogA,LogB Among them, OverallRelevance LogA,LogB The overall relevance of log data of category LogB relative to category LogA; ReverseRelevance LogA,LogB The inverse correlation of log data of category LogB relative to category LogA; q is the inverse correlation coefficient; Relevance LogA,LogB represents the positive correlation between LogB category and LogA category log data; p is the correlation coefficient.

6. The data annotation method for fault tracing in equipment log files according to claim 1, characterized in that, S1 further includes: calculating the distances of all log data of the same category within the workspace window relative to the target log data, and accumulating the distances of the target log data that are of the same category and are also of the same category as the distance-related parameter; the target log data is the first log data in the workspace window; S2 further includes: calculating the distance correlation between log data based on the distance correlation parameter, wherein the distance correlation is used to label the log dataset to be labeled.

7. The data annotation method for fault tracing in equipment log files according to claim 6, characterized in that, The distance correlation includes the average distance between LogB category log data and LogA category log data; the average distance is calculated according to the following formula: Among them, Avg_Distance LogA,LogB The average distance of log data in category LogB relative to category LogA; Count LogA,LogB For parameters where the target category is LogA and the category being counted is LogB; Distance LogA,LogB This is a distance-related parameter for log data whose statistical category is LogB and whose target log data category is LogA.

8. The data annotation method for fault tracing in equipment log files according to claim 5, characterized in that, S1 further includes: calculating the distances of all log data of the same category within the workspace window relative to the target log data, and accumulating the distances of the target log data that are of the same category and are also of the same category as the distance-related parameter; the target log data is the first log data in the workspace window; S2 further includes: calculating the distance correlation between log data based on the distance correlation parameter, wherein the distance correlation is used to label the log dataset to be labeled; The distance correlation includes the average distance between LogB category log data and LogA category log data; the average distance is calculated according to the following formula: Among them, Avg_Distance LogA,LogB The average distance of log data in category LogB relative to category LogA; Count LogA,LogB For parameters where the target category is LogA and the category being counted is LogB; Distance LogA,LogB This is a distance-related parameter for log data whose statistical category is LogB and whose target log data category is LogA.

9. The data annotation method for fault tracing in equipment log files according to claim 7, characterized in that, The process of labeling the log dataset to be labeled based on the aforementioned correlation includes: S21. Obtain the correlation between all log categories using the following method: For log data of category LogA, calculate the comprehensive correlation and average distance of log data of other categories relative to log data of category LogA; The set of corresponding categories with comprehensive correlation not lower than the correlation threshold is sorted according to the average distance to obtain the correlation relationship of LogA categories; S22. Label the log dataset to be labeled according to the correlation.

10. A data annotation device for tracing faults in device log files, comprising a processor, characterized in that, The processor is used to execute a computer program to implement the steps of the data annotation method for fault tracing of device log files as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Log-based fault index automatic tagging method and system

    CN107301118A