Intelligent operation and maintenance method based on large model

By filtering and segmenting data sources in big data clusters and utilizing a variety of indicators such as mean, correlation, and risk, the accuracy problem of multi-source data mining in big data cluster operation and maintenance, which has not been effectively addressed by existing technologies, has been solved, thereby improving detection effectiveness and data security.

CN121029555BActive Publication Date: 2026-02-10北京科杰科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511157722.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2026-02-10
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

In existing technologies, big data cluster operation and maintenance rely on a single model and real-time data, without screening the training data, making it difficult to fully explore the potential of multi-source data, resulting in poor accuracy of the generated detection reports.

Method used

By acquiring labeled log data from several data sources, we select and classify feature data sources using indicators such as data richness mean, log association mean, text interaction coefficient, and sensitive keyword ratio. We also encrypt the data based on its risk level, dynamically adjust the log collection interval, and generate a detection report.

Benefits of technology

It improved model training efficiency and data quality, enhanced operational accuracy, adapted to the characteristics of different data sources, improved the timeliness and effectiveness of anomaly detection, and ensured data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029555B_ABST
    Figure CN121029555B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and more particularly to an intelligent operation and maintenance method based on a large model, comprising: obtaining labeled log data corresponding to a plurality of data sources; directly selecting or dividing and selecting the data sources according to the data richness mean value and the log correlation mean value to select a plurality of labeled log data; determining the data risk degree corresponding to each selected labeled log data according to the sensitive keyword proportion, and locally or globally encrypting according to the data risk degree; taking the encrypted labeled log data as training data and transmitting it to the target platform to train the pre-training model; determining the log collection interval according to the processing reference value of the target big data cluster, and adjusting the log collection interval based on the interval abnormal comparison value and the abnormal number reference value to obtain log text data. The present application can improve the accuracy of the detection report generation, thereby reducing the operation and maintenance cost and risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an intelligent operation and maintenance method based on a large model. Background Technology

[0002] As the core infrastructure of distributed computing, the stability of big data clusters is crucial for ensuring the continuous and stable operation of enterprise businesses. Currently, the operation and maintenance of big data clusters, as well as task analysis, largely rely on manual processing. This not only leads to high operation and maintenance costs but also presents significant challenges. More seriously, any anomalies in the big data cluster service can rapidly disrupt and severely impact enterprise business processes and production. Therefore, improving the accuracy and efficiency of big data cluster operation and maintenance to reduce costs and risks is a pressing issue for those skilled in the art.

[0003] Chinese Patent Publication No. CN117852636A discloses a method for updating equipment operation and maintenance knowledge based on a large model, including: S1, data collection; S2, data preprocessing; S3, selecting the LLaMa2 open-source pre-trained model as the base large model; S4, incrementally training the pre-trained base large model; S5, designing fine-tuning objectives and fine-tuning training data, and fine-tuning the incrementally trained model; S6, evaluating the fine-tuned model; S7, adjusting and improving the model based on the evaluation results; and S8, model deployment and application. However, the above solution has the following problems: it relies only on a single model and real-time data, does not screen the training data, and is difficult to fully explore the potential of multi-source data, resulting in poor accuracy of the generated detection report. Summary of the Invention

[0004] To address this, the present invention provides an intelligent operation and maintenance method based on a large model, which overcomes the problem in the prior art that relies solely on a single model and real-time data, fails to screen training data, makes it difficult to fully exploit the potential of multi-source data, and results in poor accuracy of the generated detection report.

[0005] To achieve the above objectives, this invention provides an intelligent operation and maintenance method based on a large model, comprising:

[0006] Obtain labeled log data corresponding to several data sources;

[0007] Based on the data richness mean and the log correlation mean, feature data sources are directly selected or data sources are divided and selected to select a number of labeled log data;

[0008] In the selection of data source partitioning, the association partitioning is determined based on the text interaction coefficient, and several log combinations are obtained by association partitioning according to the time series correlation degree or text correlation degree. The selection analysis is then carried out based on the data source feature coefficient and the proportion of association combinations.

[0009] The data risk level of each labeled log data is determined based on the proportion of sensitive keywords, and partial or overall encryption is performed according to the data risk level.

[0010] Encrypted labeled log data is used as training data and transmitted to the target platform for training the pre-trained model;

[0011] The log collection interval is determined based on the processing reference value of the target big data cluster, and the log collection interval is adjusted based on the interval anomaly comparison value and the anomaly number reference value to obtain log text data.

[0012] The acquired log data is input into the pre-trained model to generate a detection report.

[0013] Furthermore, if the average data richness is greater than or equal to the preset average data richness or the average log association is greater than or equal to the preset average log association, then the feature data source is directly selected.

[0014] In the feature data source selection, the labeled log data from data sources with a comprehensive characterization value greater than the preset comprehensive characterization value are used as training data.

[0015] Furthermore, if the average data richness is less than the preset average data richness and the average log association is less than the preset average log association, then data source segmentation and selection will be performed.

[0016] Furthermore, based on the text interaction coefficient, the association segmentation is determined according to either temporal or textual correlation, including:

[0017] When performing association partitioning for each data source, or when performing association partitioning for a single data source,

[0018] If the text interaction coefficient is less than the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the text correlation degree.

[0019] If the text interaction coefficient is greater than or equal to the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the time-series correlation degree.

[0020] Furthermore, the temporal correlation degree is determined based on the point influence correlation degree and the root cause difference coefficient;

[0021] The temporal correlation degree is positively correlated with the point influence correlation degree, and the temporal correlation degree is negatively correlated with the root cause difference coefficient.

[0022] Furthermore, selection analysis is conducted based on the characteristic coefficients of the data source and the proportion of associated combinations, including:

[0023] The number of logs selected for each data source is determined based on the characteristic coefficients of the data source, and the number of logs selected for each combination is determined based on the proportion of associated combinations of each log combination. The number of logs selected for each unstable data source is adjusted to be increased based on the feature comparison deviation value of each unstable data source.

[0024] The increase in the number of logs selected for a single unstable data source is positively correlated with the feature comparison deviation value corresponding to that unstable data source.

[0025] The unstable data source is a data source whose feature comparison deviation value is greater than a preset feature comparison deviation value.

[0026] Furthermore, the data risk level corresponding to each selected labeled log data is determined based on the proportion of sensitive keywords, including:

[0027] For a single labeled log data,

[0028] If the percentage of sensitive keywords is greater than or equal to the preset percentage of sensitive keywords, the data risk level is determined based on the percentage of sensitive keywords.

[0029] If the proportion of sensitive keywords is less than the preset proportion of sensitive keywords, the data risk level is determined based on the sensitivity relevance and the proportion of sensitive keywords.

[0030] Furthermore, based on the level of data risk, partial or overall encryption can be implemented, including:

[0031] If the data risk level is greater than or equal to the preset data risk level, then overall encryption will be performed;

[0032] If the data risk level is less than the preset data risk level, then local encryption will be performed.

[0033] Furthermore, the log collection interval is determined based on the processing reference values ​​of the target big data cluster;

[0034] The processing reference value is determined based on the task density index and the task processing time consumption rate.

[0035] Furthermore, under the preset adjustment conditions, the log collection interval is reduced based on the interval anomaly comparison value;

[0036] The preset adjustment condition is that the interval anomaly comparison value is greater than the preset interval anomaly comparison value, and the decrease in the log collection interval is positively correlated with the interval anomaly comparison value.

[0037] Compared with the prior art, the beneficial effects of the present invention are as follows: In the technical solution of the present invention, the correlation degree of the labeled log data of each data source is effectively reflected by the data enrichment mean and the log correlation mean. Then, feature data sources are directly selected or data sources are selected by division based on the data enrichment mean and the log correlation mean. Direct selection of feature data sources can efficiently utilize high-quality data for training, reduce redundant operations, and improve model training efficiency. Data source division selection can focus on more relevant parts, improve data quality and correlation, help to discover potential patterns, improve model performance, and thus improve the accuracy of operation and maintenance strategy determination.

[0038] Furthermore, this invention effectively reflects the textual correlation of labeled log data in the data source through the textual interaction coefficient. Then, the correlation is determined based on the textual interaction coefficient according to the temporal correlation or textual correlation. This can better adapt to the characteristics of different data sources, form a more representative and regular training dataset, and help improve the quality and effect of model training.

[0039] Furthermore, this invention determines the number of logs selected for each data source by using the feature coefficients of the data source, enabling targeted data selection for different data sources and improving the effectiveness of data selection. The number of logs selected for each combination is determined based on the proportion of related combinations, fully considering the correlation between log combinations and making the selected training data more representative and relevant. For unstable data sources with feature alignment deviation values ​​greater than a preset value, the number of logs selected can be increased based on the magnitude of the feature alignment deviation. This dynamic adjustment mechanism effectively addresses the instability of data sources, improves the ability to process abnormal or unstable data, and thus improves the overall quality of the training data.

[0040] Furthermore, this invention combines the proportion of sensitive keywords and the degree of sensitivity association to determine the data risk level, making the assessment of data risk level more comprehensive and accurate. The data risk level effectively reflects the sensitivity of information in the labeled log data, and then different encryption strategies are adopted according to the data risk level, realizing flexible adaptation of encryption methods. This can ensure data security while avoiding the waste of resources caused by over-encrypting all data.

[0041] Furthermore, this invention comprehensively considers the task density index and task processing time of the big data cluster, enabling the log collection interval to be dynamically adjusted according to the actual load of the cluster. This allows it to adapt to the operational needs of the big data cluster under different load scenarios. Under preset adjustment conditions, the log collection interval is reduced based on the interval anomaly comparison value, ensuring that the log collection frequency can be quickly increased when anomalies occur, timely capturing of abnormal information, and improving the timeliness and effectiveness of monitoring. Attached Figure Description

[0042] Figure 1 This is a schematic diagram of the intelligent operation and maintenance method based on a large model according to the present invention;

[0043] Figure 2 This is a flowchart illustrating the process of directly selecting feature data sources or selecting data source partitions based on data richness and log correlation mean in this invention.

[0044] Figure 3 This is a flowchart illustrating the process of determining association division based on temporal or textual correlation coefficients according to the present invention.

[0045] Figure 4 This is a flowchart illustrating the present invention's method of partial or overall encryption based on data risk level. Detailed Implementation

[0046] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0047] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0048] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0049] Please see Figures 1 to 4 As shown, this invention provides an intelligent operation and maintenance method based on a large model, comprising:

[0050] Obtain labeled log data corresponding to several data sources;

[0051] Based on the data richness mean and the log correlation mean, feature data sources are directly selected or data sources are divided and selected to select a number of labeled log data;

[0052] In the selection of data source partitioning, the association partitioning is determined based on the text interaction coefficient, and several log combinations are obtained by association partitioning according to the time series correlation degree or text correlation degree. The selection analysis is then carried out based on the data source feature coefficient and the proportion of association combinations.

[0053] The data risk level of each labeled log data is determined based on the proportion of sensitive keywords, and partial or overall encryption is performed according to the data risk level.

[0054] Encrypted labeled log data is used as training data and transmitted to the target platform for training the pre-trained model;

[0055] The log collection interval is determined based on the processing reference value of the target big data cluster, and the log collection interval is adjusted based on the interval anomaly comparison value and the anomaly number reference value to obtain log text data.

[0056] The acquired log data is input into the pre-trained model to generate a detection report.

[0057] The application scenario of this invention is the anomaly detection of log text data in a target big data cluster;

[0058] The target big data cluster is a computing environment containing several processing nodes. Each processing node is a computer capable of handling large-scale processing tasks. Each processing node corresponds to several processing tasks, including but not limited to data backup and recovery, data indexing and retrieval, and data mining. This is content that is easy for those skilled in the art to understand, and will not be elaborated on in detail.

[0059] This invention includes several data sources, each of which includes several labeled log data. Each labeled log data includes log text data, entities, event types, time-series correlation diagrams, and detection reports. The detection reports include fault types and their impact ranges. Fault types include, but are not limited to, hard disk damage, memory failure, and network interface card failure. The impact range is the processing nodes where the anomalies occur. This is content that is easily understood by those skilled in the art and will not be elaborated further.

[0060] Log text data refers to the raw text records generated during the operation of the target big data cluster. Entities are entity information extracted from the log text data, including but not limited to node IP, error code, task ID, and service name. Event types are key events reflecting key behaviors or state changes during the operation of the target big data cluster, including but not limited to "disk I / O timeout", "HDFS replica loss", and "task retry surge". In the time series correlation graph, each event type in the log text data is a point. If two event types occur consecutively within 5 minutes, there is an edge between them. Each edge has a corresponding weight coefficient. The formula for calculating the influence coefficient of a single edge is W = 1 / (1+Δt), where Δt is the time difference between the two event types corresponding to the edge, in minutes. The root cause influence reference value of a single point is the average value of the influence coefficients of each edge connected to that point in the time series correlation graph.

[0061] Encrypted labeled log data is used as training data and transmitted to the target platform for training the pre-trained model;

[0062] The encrypted log data is sent to the target platform via an HTTPS POST request. The target platform is a private cloud platform built by the user based on open source technologies such as OpenStack, which is used to receive the encrypted labeled log data and train the model.

[0063] The process of training a pre-trained model includes, but is not limited to, data preprocessing, dataset partitioning, and setting hyperparameters. These are common techniques used by those skilled in the art and will not be elaborated upon here.

[0064] The pre-trained model takes log text data as input and detection reports as output.

[0065] This invention includes several historical records, each recording at least one instance of anomaly detection in log text data within a big data cluster, including the average data richness, average log association, comprehensive representation value, and text interaction coefficient. Each historical record also has a corresponding pass / fail marker, indicating whether the anomaly detection process in the log text data of the big data cluster meets user requirements. These markers can be manually recorded. It is understood that users can determine whether the anomaly detection process in the log text data of the big data cluster meets their requirements based on self-defined indicators. These self-defined indicators can be, but are not limited to, the number of errors, which will not be elaborated upon here. The number of errors refers to the number of times the fault type is incorrectly identified in the detection report generated by the pre-trained model.

[0066] Specifically, if the average data richness is greater than or equal to the preset average data richness or the average log association is greater than or equal to the preset average log association, then the feature data source is directly selected.

[0067] In the feature data source selection, the labeled log data from data sources with a comprehensive characterization value greater than the preset comprehensive characterization value are used as training data.

[0068] Among them, the average data richness is the average of the data richness corresponding to each data source. The data richness corresponding to a single data source = the richness coefficient corresponding to the data source / the average of the richness coefficients corresponding to each data source + the reference value of the data quantity corresponding to the data source / the average of the reference values ​​of the data quantity corresponding to each data source.

[0069] For a single data source, the data source is designated as the target data source, and other data sources other than the target data source are designated as reference data sources. Different keywords appearing in the log text data corresponding to each labeled log data of the target data source are designated as target keywords, and different keywords appearing in the log text data corresponding to each labeled log data of the reference data source are designated as reference keywords.

[0070] The richness coefficient corresponding to the target data source = the number of keywords that are the same in the target keywords and the reference keywords / the total number of reference keywords; the keywords are identified through NLP technology, which is a common technique used by those skilled in the art and will not be elaborated on in detail.

[0071] The reference value for the number of data points corresponding to a single data source is the number of labeled log data points in that data source.

[0072] The values ​​of the preset data richness mean and the preset log association mean can be determined by the user according to the actual application scenario. The smaller the values ​​of the preset data richness mean and the preset log association mean, the greater the user's need for data source segmentation and selection. A method for determining the values ​​of the preset data richness mean and the preset log association mean is provided. The historical records of the user's data source segmentation and selection are detected, and the average value of the data richness mean and the average value of the log association mean corresponding to the historical records that meet the user's needs are respectively recorded as the preset data richness mean and the preset log association mean.

[0073] The log correlation mean is the average log correlation degree of each data source. The log correlation degree of the target data source = the number of different event types appearing in the target data source / the total number of different event types appearing in all data sources.

[0074] The comprehensive representation value corresponding to a single data source = data richness of the data source / mean data richness + log correlation of the data source / mean log correlation;

[0075] The preset comprehensive characterization value can be determined by the user based on the actual application scenario. The greater the user's demand for improving the accuracy of data selection quality, the larger the preset comprehensive characterization value should be. A method for determining the preset comprehensive characterization value is provided, which detects the historical records of direct selection of feature data sources, and records the average value of the comprehensive characterization values ​​corresponding to each data source selected in the historical records that meet the user's needs as the preset comprehensive characterization value.

[0076] Specifically, if the average data richness is less than the preset average data richness and the average log association is less than the preset average log association, then data source segmentation and selection will be performed.

[0077] Understandably, the data richness mean measures the richness of the data source, and the log correlation mean reflects the correlation between event types in the data source. If the data richness mean is greater than or equal to the preset data richness mean or the log correlation mean is greater than or equal to the preset log correlation mean, it indicates that the overall quality of the data source is good, and feature data sources can be directly selected from them. This simplifies the data source selection process, improves data selection efficiency, and ensures the quality of the selected data sources.

[0078] If the average data richness is less than the preset average data richness and the average log association is less than the preset average log association, it indicates that the overall information content and richness of the data source are insufficient. Direct use may not meet the requirements of analysis or training. Therefore, it is necessary to divide and select the data source in order to mine the data and improve the relevance and accuracy of the analysis.

[0079] Specifically, the association classification based on text interaction coefficients, whether according to temporal or textual correlation, includes:

[0080] When performing association partitioning for each data source, or when performing association partitioning for a single data source,

[0081] If the text interaction coefficient is less than the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the text correlation degree.

[0082] If the text interaction coefficient is greater than or equal to the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the time-series correlation.

[0083] Among them, the text interaction coefficient corresponding to a single data source is the average of the text association mean of each labeled log data in that data source;

[0084] The average text association of a single labeled log data is the average of the text association between that labeled log data and the other labeled log data in the data source. For any two labeled log data, the text association between the two labeled log data is the larger of the number of event types that both of the labeled log data exist in and the number of event types that both of the labeled log data exist in.

[0085] The value of the preset text interaction coefficient can be determined by the user according to the actual application scenario. The larger the value of the preset text interaction coefficient, the greater the user's need to classify and associate the labeled log data corresponding to the data source based on the text correlation. A method for determining the value of the preset text interaction coefficient is provided, which detects the historical records of classifying and associating the labeled log data corresponding to the data source, and records the average value of the text interaction coefficients corresponding to the historical records that can meet the user's needs as the preset text interaction coefficient.

[0086] The association analysis is performed on the labeled log data corresponding to a single data source based on textual or temporal correlation. This includes: performing association analysis on each labeled log data in the data source; when performing association analysis on a single labeled log data, the labeled log data is recorded as the target labeled log data, and other labeled log data other than the target labeled log data are recorded as reference labeled log data; the reference labeled log data and the target labeled log data that meet the preset conditions are recorded into a log group, and the association analysis continues on the labeled log data that are not recorded into the log group until all labeled log data are recorded into the corresponding log group, at which point the association analysis stops.

[0087] It should be noted that if the labeled log data corresponding to a single data source is associated and divided according to the text correlation degree, the reference labeled log data that meets the preset conditions is the reference labeled log data whose text correlation degree with the target labeled log data is greater than the preset text correlation degree.

[0088] If the labeled log data corresponding to a single data source is associated and divided according to the temporal correlation degree, then the reference labeled log data that meets the preset conditions is the reference labeled log data whose temporal correlation degree with the target labeled log data is greater than the preset temporal correlation degree;

[0089] The values ​​of preset text correlation and preset time-series correlation can be determined by the user based on the actual application scenario. The greater the user's need to improve the accuracy of log combination segmentation, the larger the values ​​of preset text correlation and preset time-series correlation will be. A method for determining the values ​​of preset text correlation and preset time-series correlation is provided, which is the average of the reference text correlation for each log combination in the historical records that can meet the user's needs and are associated with the labeled log data corresponding to the data source based on text correlation. The average of the reference time-series correlation for each log combination in the historical records that can meet the user's needs and are associated with the labeled log data corresponding to the data source based on time-series correlation is also recorded as the preset time-series correlation. The reference text correlation for a single log combination is the text correlation for any two labeled log data in that log combination, and the reference time-series correlation for a single log combination is the time-series correlation for any two labeled log data in that log combination.

[0090] Understandably, the text interaction coefficient reflects the strength of the correlation between event types among labeled log data in a single data source. If the text interaction coefficient is less than the preset text interaction coefficient, it indicates that the overall correlation between event types among labeled log data is weak. In this case, the text interaction coefficient is used to classify the labeled log data corresponding to the data source based on the text interaction coefficient. This can accurately focus on log combinations that have significant overlap in event types, ensuring that effective correlation at the content level is captured.

[0091] If the text interaction coefficient is greater than or equal to the preset text interaction coefficient, it indicates that there is a strong correlation between the event types among the labeled log data. At this time, the correlation based on the time sequence correlation can further highlight the correlation between the log data in terms of the impact of event types and potential root causes.

[0092] Specifically, the temporal correlation degree is determined based on the point influence correlation degree and the root cause difference coefficient;

[0093] The temporal correlation degree is positively correlated with the point influence correlation degree, and the temporal correlation degree is negatively correlated with the root cause difference coefficient.

[0094] Specifically, for any two labeled log data, the same event type in the time sequence association graph corresponding to the two labeled log data is recorded as a reference type, and each reference type in any labeled log data is sorted in order of the occurrence time of the event type from earliest to latest and recorded as a reference sequence;

[0095] The correlation between the points corresponding to any two labeled log data is the average of the similarity of the influence of the points corresponding to each reference type of the two labeled log data.

[0096] The impact similarity of a single reference type = the total number of identical edges corresponding to the reference type / the larger value among the number of connection edges corresponding to the two labeled log data. Identical edges are those that connect points corresponding to the reference type in the time series association graphs corresponding to the two labeled log data, and the other points they connect also correspond to the same event type.

[0097] For a single reference type, the number of edges corresponding to the points of that reference type in the time-series correlation graph corresponding to a single labeled log data is recorded as the number of connection edges corresponding to that reference type in the labeled log data;

[0098] The formula for calculating the root cause difference coefficient ε between any two labeled log data is: Where n is the number of the same event type in the two labeled log data, i = 1, 2, ..., n, a i Let b be the root cause reference value for the point corresponding to the i-th event type in a reference sequence of labeled log data. i The root cause influence reference value for the point corresponding to the i-th event type in the reference sequence of another labeled log data;

[0099] The temporal correlation between any two labeled log data is equal to the point influence correlation / preset point influence correlation + (1 - root cause difference coefficient / preset root cause difference coefficient).

[0100] The values ​​of the preset point influence correlation degree and the preset root cause difference coefficient can be determined by the user based on the actual application scenario. The greater the user's need for improving the accuracy of time series correlation determination, the smaller the values ​​of the preset point influence correlation degree and the preset root cause difference coefficient should be. A method for determining the values ​​of the preset point influence correlation degree and the preset root cause difference coefficient is provided, which are respectively denoted as the preset point influence correlation degree and the preset root cause difference coefficient, corresponding to each historical record that can meet the user's needs. The reference point influence correlation degree and the reference root cause difference coefficient corresponding to a single historical record are respectively the point influence correlation degree and the root cause difference coefficient corresponding to any two labeled log data in that historical record.

[0101] Understandably, point impact correlation reflects the similarity of labeled log data in terms of event type impact, while root cause difference coefficient reflects the degree of difference in potential root causes. Integrating point impact correlation and root cause difference coefficient together can comprehensively reflect the degree of correlation between labeled log data.

[0102] Specifically, the selection and analysis are based on the characteristic coefficients of the data source and the proportion of associated combinations, including:

[0103] The number of logs selected for each data source is determined based on the characteristic coefficients of the data source, and the number of logs selected for each combination is determined based on the proportion of associated combinations of each log combination. The number of logs selected for each unstable data source is adjusted to be increased based on the feature comparison deviation value of each unstable data source.

[0104] The increase in the number of logs selected for a single unstable data source is positively correlated with the feature comparison deviation value corresponding to that unstable data source.

[0105] The unstable data source is a data source whose feature comparison deviation value is greater than a preset feature comparison deviation value.

[0106] Wherein, the data source feature coefficient corresponding to a single data source = number of log combinations / preset number of combinations + (1 - combination similarity / preset combination similarity);

[0107] The number of log combinations corresponding to a single data source is the total number of log combinations contained in that data source, and the preset number of combinations is the average number of log combinations corresponding to each data source.

[0108] The combined similarity corresponding to a single data source is the average of the similarity mean values ​​corresponding to each log combination in that data source. For a single data source, log combinations from other data sources outside of that data source are denoted as analysis log combinations. The average similarity of a single log combination in that data source is the average of the similarity reference values ​​corresponding to that log combination and each analysis log combination. For any two log combinations, they are denoted as the first combination and the second combination, respectively. The similarity reference value corresponding to the two log combinations is the maximum value among the similarity thresholds corresponding to each labeled log data in the first combination. The similarity threshold corresponding to a single labeled log data in the first combination is the maximum value among the sub-similarity thresholds corresponding to that labeled log data and each labeled log data in the second combination.

[0109] The sub-similarity threshold for any two labeled log data is calculated as: text correlation / preset text correlation + temporal correlation / preset temporal correlation.

[0110] The number of logs selected for a single data source is the number of labeled log data selected from that data source as training data;

[0111] The number of logs selected for a single data source is the smallest integer greater than or equal to n0, where n0 = (the feature coefficient of the data source corresponding to this data source / the sum of the feature coefficients of the data sources corresponding to each data source) × the average number of training data selected for each historical record that can meet the user's needs.

[0112] The percentage of associated combinations corresponding to a single log combination in a single data source = the number of analytical log combinations whose similarity reference value is greater than the preset similarity reference value / the total number of analytical log combinations;

[0113] The user can determine the value of the preset similarity reference value according to the actual application scenario. The greater the user's need to improve the accuracy of determining the proportion of associated combinations, the larger the value of the preset similarity reference value will be. One preset similarity reference value is provided, and the preset similarity reference value is 1.7.

[0114] The number of combinations selected for a single log combination is the number of labeled log data selected from that log combination as training data;

[0115] The number of combinations selected for a single log combination in a single data source is n, where n is the smallest integer greater than or equal to n1. n1 = (the proportion of associated combinations corresponding to this log combination / the sum of the proportions of associated combinations corresponding to each log combination in this data source) × the number of logs selected for this data source.

[0116] It should be noted that if the number of labeled log data in a single log combination is less than n, then all labeled log data in that log combination will be used as training data. After selecting training data for each log combination in a single data source, labeled log data will be randomly selected until the total amount of selected training data reaches the number of logs selected for that data source.

[0117] The feature alignment deviation value for a single data source = the average value of the feature reference values ​​corresponding to the historical records that can meet the user's needs - the feature reference value corresponding to that data source;

[0118] The feature reference value corresponding to a single data source = (the number of logs selected for this data source / the average number of logs selected for each data source) / (the feature coefficient of this data source / the average feature coefficient of each data source + the average proportion of the associated combination corresponding to each log combination in this data source / the average proportion of the associated combination corresponding to each log combination in all data sources).

[0119] The increment value for the number of logs selected for a single data source is the smallest integer greater than or equal to z0, where z0 = (feature comparison deviation value corresponding to the data source / average value of feature reference values ​​corresponding to historical records that can meet user needs) × the initial number of logs selected for the data source.

[0120] Annotated log data selected from various data sources will be used as training data;

[0121] The user can determine the preset feature comparison deviation value according to the actual application scenario. The greater the user's need to improve the selection accuracy, the smaller the preset feature comparison deviation value will be. One preset feature comparison deviation value is provided, with a preset feature comparison deviation value of 0.

[0122] Specifically, the data risk level of each selected labeled log data is determined based on the proportion of sensitive keywords, including:

[0123] For a single labeled log data,

[0124] If the percentage of sensitive keywords is greater than or equal to the preset percentage of sensitive keywords, the data risk level is determined based on the percentage of sensitive keywords.

[0125] If the proportion of sensitive keywords is less than the preset proportion of sensitive keywords, the data risk level is determined based on the sensitivity relevance and the proportion of sensitive keywords.

[0126] The present invention includes a sensitive keyword database, which stores a number of sensitive keywords, including but not limited to passwords, keys, IP addresses and MAC addresses, which will not be elaborated further.

[0127] The percentage of sensitive keywords corresponding to a single labeled log data = the number of sensitive keywords appearing in the log text data corresponding to that labeled log data / the total number of keywords contained in the log text data corresponding to that labeled log data;

[0128] The user can determine the preset sensitive keyword percentage based on the actual application scenario. The greater the user's need to reduce the risk of information leakage, the smaller the preset sensitive keyword percentage should be. One preset sensitive keyword percentage is provided, which is 40%.

[0129] When determining the data risk level based on the proportion of sensitive keywords, the data risk level = the proportion of sensitive keywords / the average proportion of sensitive keywords corresponding to each labeled log data;

[0130] When determining the data risk level based on the sensitivity correlation and the proportion of sensitive keywords, the data risk level = 0.5 × sensitivity correlation / average sensitivity correlation of each labeled log data + 0.5 × proportion of sensitive keywords / average proportion of sensitive keywords of each labeled log data;

[0131] The sensitivity correlation of a single labeled log data is the average number of sensitive collocations corresponding to each common keyword appearing in the log text data corresponding to that labeled log data;

[0132] Ordinary keywords are keywords other than sensitive keywords;

[0133] The log text data corresponding to each selected training data point is recorded as the reference text.

[0134] The number of sensitive collocations corresponding to a single common keyword = the number of sensitive texts corresponding to that common keyword / the number of reference texts containing that common keyword;

[0135] If a single reference text containing the common keyword is detected, and the reference text contains a sensitive keyword whose interval keyword count is less than the preset interval keyword count, then the reference text is recorded as the sensitive text corresponding to the common keyword.

[0136] The number of keywords between any two keywords is the number of keywords that are located between the two keywords. The preset number of keywords can be determined by the user according to the actual application scenario. The greater the user's need to improve the accuracy of sensitive text judgment, the smaller the preset number of keywords should be. One preset number of keywords is provided, which is 10.

[0137] It is understandable that the proportion of sensitive keywords can effectively reflect the sensitivity of a single labeled log data. When the proportion of sensitive keywords is greater than or equal to the preset proportion of sensitive keywords, it means that the log text data contains a lot of sensitive information. Directly determining the data risk level based on the proportion of sensitive keywords can accurately reflect its high-risk characteristics.

[0138] If the proportion of sensitive keywords is less than the preset proportion of sensitive keywords, it means that there is less sensitive information contained, but there may be a potential connection between ordinary keywords and sensitive keywords. Combining the degree of sensitivity connection with the proportion of sensitive keywords to comprehensively determine the data risk can more comprehensively assess the risk of log data and avoid missing potential risks by simply relying on the proportion of sensitive keywords.

[0139] Specifically, encryption can be performed locally or globally based on the level of data risk, including:

[0140] If the data risk level is greater than or equal to the preset data risk level, then overall encryption will be performed;

[0141] If the data risk level is less than the preset data risk level, then local encryption will be performed.

[0142] The user can determine the value of the preset data risk level according to the actual application scenario. The smaller the value of the preset data risk level, the greater the user's need for overall encryption. A method for determining the value of the preset data risk level is provided, which detects the user's history of overall encryption and records the average value of the data risk level corresponding to the history that meets the user's needs as the preset data risk level.

[0143] When performing overall encryption on a single labeled log data, all keywords in the log text data of that labeled log data are encrypted. When performing partial encryption, all sensitive keywords in the labeled log data are encrypted, and a preset number of ordinary keywords in the log text data of that labeled log data are also encrypted. The preset number for a single labeled log data is calculated as follows: (Data risk level of the labeled log data / Average data risk level of each labeled log data in the historical records that can be partially encrypted and meet user needs) × Average number of ordinary keywords encrypted for each labeled log data in the historical records that can be partially encrypted and meet user needs. When encrypting the preset number of ordinary keywords, ordinary keywords are selected in descending order of the number of sensitive collocations corresponding to the ordinary keywords until the preset number is reached.

[0144] When encrypting keywords, the AES algorithm is used, which is a common technique in this field and will not be elaborated on in detail.

[0145] Understandably, if the data risk level is greater than or equal to the preset data risk level, it indicates that the log text data contains a large amount of sensitive information or is highly related to sensitive information. Overall encryption can protect data security to the greatest extent and prevent the leakage of sensitive information.

[0146] If the data risk level is less than the preset data risk level, it means that the log text data is relatively low in sensitivity. Local encryption is adopted to protect sensitive information while reducing the resource utilization of encryption processing, thereby improving processing efficiency.

[0147] Specifically, the log collection interval is determined based on the processing reference values ​​of the target big data cluster;

[0148] The processing reference value is determined based on the task density index and the task processing time consumption rate.

[0149] Wherein, the task density index = the total number of tasks processed in the target big data cluster / the number of processing nodes in the target big data cluster;

[0150] The task processing time rate is the average of the subtask processing time rates corresponding to each processing node within the preset time period. It can be understood that the greater the user's demand for improving the accuracy of task processing time rate confirmation, the longer the preset time period will be. A method for setting the length of the preset time period is provided. The preset time period is 5 minutes. The time 5 minutes earlier than the current time is recorded as the reference time, and the time period between the reference time and the current time is recorded as the preset time period.

[0151] The subtask processing time of a single processing node = 5 min / the number of processing tasks processed by the processing node within a preset time period;

[0152] The log collection interval is the time length of log data corresponding to the target big data cluster collected between two consecutive collections.

[0153] The log collection interval is T, where T = processing reference value / average processing reference value corresponding to the historical records that can meet user needs × average log collection interval corresponding to the historical records that can meet user needs;

[0154] Processing reference value = Task density index / Average of task density indices corresponding to historical records that can meet user needs + Task processing time rate / Average of task processing time rates corresponding to historical records that can meet user needs.

[0155] Understandably, the task density index reflects the average task load of each processing node, while the task processing time rate reflects the efficiency and time consumption of the cluster in processing tasks. Determining the processing reference value of the target big data cluster based on the task density index and task processing time rate can comprehensively reflect the actual operating status of the target big data cluster, thereby dynamically adjusting the log collection interval. This can ensure the effectiveness of log text data collection while making reasonable use of system resources.

[0156] Specifically, under the preset adjustment conditions, the log collection interval is reduced based on the interval anomaly comparison value;

[0157] The preset adjustment condition is that the interval anomaly comparison value is greater than the preset interval anomaly comparison value, and the decrease in the log collection interval is positively correlated with the interval anomaly comparison value.

[0158] Wherein, the interval anomaly comparison value = the average value of the log collection intervals corresponding to the historical records that can meet the user's needs - the log collection interval corresponding to the target big data cluster;

[0159] The user can determine the preset interval anomaly comparison value according to the actual application scenario. The greater the user's need to improve monitoring efficiency, the smaller the preset interval anomaly comparison value should be. One preset interval anomaly comparison value is provided, which is 2 seconds.

[0160] The reduction in log collection interval = (interval anomaly comparison value / preset interval anomaly comparison value - 1) × T;

[0161] Understandably, the interval anomaly comparison value reflects the difference between the user's historical log collection interval and the current target cluster log collection interval. The larger the difference, the greater the deviation between the current collection interval and the user's needs, and the more adjustment is required.

[0162] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A smart operation and maintenance method based on a large model, characterized in that, include: Obtain labeled log data corresponding to several data sources; Based on the data richness mean and the log correlation mean, feature data sources are directly selected or data sources are divided and selected to select a number of labeled log data; In the selection of data source partitioning, the association partitioning is determined based on the text interaction coefficient, and several log combinations are obtained by association partitioning according to the time series correlation degree or text correlation degree. The selection analysis is then carried out based on the data source feature coefficient and the proportion of association combinations. The data risk level of each labeled log data is determined based on the proportion of sensitive keywords, and partial or overall encryption is performed according to the data risk level. Encrypted labeled log data is used as training data and transmitted to the target platform for training the pre-trained model; The log collection interval is determined based on the processing reference value of the target big data cluster, and the log collection interval is adjusted based on the interval anomaly comparison value and the anomaly number reference value to obtain log text data. The acquired log data is input into the trained pre-trained model to generate a detection report; Based on text interaction coefficients, association classification is determined according to temporal or textual correlation, including: When performing association partitioning for each data source, or when performing association partitioning for a single data source, If the text interaction coefficient is less than the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the text correlation degree. If the text interaction coefficient is greater than or equal to the preset text interaction coefficient, then the labeled log data corresponding to the data source will be associated and divided according to the time-series correlation degree. The temporal correlation degree is determined based on the point influence correlation degree and the root cause difference coefficient; Among them, the temporal correlation degree and the point influence correlation degree are positively correlated, and the temporal correlation degree and the root cause difference coefficient are negatively correlated; The selection and analysis are based on the characteristic coefficients of the data source and the proportion of associated combinations, including: The number of logs selected for each data source is determined based on the characteristic coefficients of the data source, and the number of logs selected for each combination is determined based on the proportion of associated combinations of each log combination. The number of logs selected for each unstable data source is adjusted to be increased based on the feature comparison deviation value of each unstable data source. The increase in the number of logs selected for a single unstable data source is positively correlated with the feature comparison deviation value corresponding to that unstable data source. The unstable data source is a data source whose feature comparison deviation value is greater than a preset feature comparison deviation value.

2. The intelligent operation and maintenance method based on a large model according to claim 1, characterized in that, If the average data richness is greater than or equal to the preset average data richness or the average log association is greater than or equal to the preset average log association, then the feature data source is directly selected. In the feature data source selection, the labeled log data from data sources with a comprehensive characterization value greater than the preset comprehensive characterization value are used as training data.

3. The intelligent operation and maintenance method based on a large model according to claim 1, characterized in that, If the average data richness is less than the preset average data richness and the average log association is less than the preset average log association, then data source segmentation and selection will be performed.

4. The intelligent operation and maintenance method based on a large model according to claim 1, characterized in that, The data risk level of each selected labeled log data is determined based on the proportion of sensitive keywords, including: For a single labeled log data, If the percentage of sensitive keywords is greater than or equal to the preset percentage of sensitive keywords, the data risk level is determined based on the percentage of sensitive keywords. If the proportion of sensitive keywords is less than the preset proportion of sensitive keywords, the data risk level is determined based on the sensitivity relevance and the proportion of sensitive keywords.

5. The intelligent operation and maintenance method based on a large model according to claim 4, characterized in that, Encryption can be performed locally or globally based on the level of data risk, including: If the data risk level is greater than or equal to the preset data risk level, then overall encryption will be performed; If the data risk level is less than the preset data risk level, then local encryption will be performed.

6. The intelligent operation and maintenance method based on a large model according to claim 1, characterized in that, The log collection interval is determined based on the processing reference values ​​of the target big data cluster; The processing reference value is determined based on the task density index and the task processing time consumption rate.

7. The intelligent operation and maintenance method based on a large model according to claim 6, characterized in that, Under preset adjustment conditions, the log collection interval is reduced based on the interval anomaly comparison value. The preset adjustment condition is that the interval anomaly comparison value is greater than the preset interval anomaly comparison value, and the decrease in the log collection interval is positively correlated with the interval anomaly comparison value.

Citation Information

Patent Citations

  • Equipment operation and maintenance knowledge updating method based on large model

    CN117852636A

  • Log quality assessment method, device, server and storage medium

    CN108121645A

  • System anomaly detection method based on program log data

    CN115604003A