Abnormal desensitization log recognition model training method and device, equipment and medium
By training anomaly detection model, learning the characteristic distribution of normal desensitization logs, identifying different abnormal desensitization logs, solving the problem of difficult to identify unconventional format log data in the prior art, and achieving effective abnormal desensitization log recognition.
Patent Information
- Application Number
- CN202510083099.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for the prior art to effectively identify whether unanticipated log data (such as log data in unconventional formats, etc.) is an abnormal desensitized log, especially in the absence of real abnormal desensitized log samples.
By obtaining a collection of sample desensitization logs including normal desensitization logs, training anomaly detection model, and learning the characteristic distribution of normal desensitization logs, thereby identifying abnormal desensitization logs that are different from the characteristic distribution of normal desensitization logs.
It realizes the correct identification of whether the expected log data is an abnormal desensitized log without the need for abnormal samples, and improves the effective identification ability of abnormal desensitized logs.
Smart Images

Figure CN120011812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a training method for an abnormal desensitized log recognition model, an abnormal desensitized log recognition method, a device, equipment, a medium and a program product. Background Art
[0002] For log data involving user privacy, it is often necessary to desensitize the logs. For example, use encryption algorithms to convert IP addresses, account names, passwords, phone numbers, etc. in log data to protect sensitive information during data analysis and storage. In order to prevent information leakage due to log desensitization failures, it is necessary to further identify abnormal desensitized logs after desensitizing the logs.
[0003] In related technologies, it is usually necessary to predefine identification rules based on expert experience, such as predefining the data types, data formats, key fields, etc. that need to be desensitized, and then use the predefined identification rules to identify abnormal desensitized logs. However, due to the limitations of expert experience, predefined identification rules often cannot correctly identify unexpected log data (such as log data with unusual formats, etc.) as abnormal desensitized logs. It is necessary to provide a training program for the abnormal desensitized log identification model to effectively identify abnormal desensitized logs. Summary of the invention
[0004] The embodiments of this specification provide a training method, device, equipment, medium and program product for an abnormal desensitized log recognition model, which can effectively identify abnormal desensitized logs.
[0005] In a first aspect, an embodiment of this specification provides a method for training an abnormal desensitization log recognition model, including:
[0006] Obtain a first sample feature vector set corresponding to the first sample desensitized log set; the first sample desensitized log set includes normal desensitized logs;
[0007] Based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector, the anomaly detection model is trained to obtain a first anomaly desensitized log recognition model.
[0008] In a possible implementation, based on each first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector, an anomaly detection model is trained to obtain an abnormal desensitized log recognition model, including:
[0009] Inputting at least one first sample feature vector into an anomaly detection model to obtain a first prediction label corresponding to each first sample feature vector;
[0010] Determining a first classification loss of an anomaly detection model according to each first prediction label and each first sample label;
[0011] The parameters of the anomaly detection model are adjusted based on the first classification loss, and the step of obtaining the first sample feature vector set corresponding to the first sample desensitized log set is performed again until the first training stop condition is reached, so as to obtain the first abnormal desensitized log recognition model.
[0012] In a possible implementation, before obtaining the first sample feature vector set corresponding to the first sample anonymized log set, the method further includes:
[0013] Input a batch of first sample desensitized logs into a feature extraction model to obtain a first sample feature vector corresponding to each first sample desensitized log in the batch of first sample desensitized logs;
[0014] Building a first log feature library based on each first sample feature vector and a first sample label carried by each first sample feature vector;
[0015] Obtain the first sample feature vector set corresponding to the first sample anonymized log set, including:
[0016] Obtain a first sample feature vector set corresponding to the first sample desensitized log set from the first log feature library.
[0017] In a possible implementation, the method further includes:
[0018] Obtain a second sample feature vector set corresponding to the second sample desensitized log set; the data format and / or data semantics of each second sample desensitized log in the second sample desensitized log set are different from those of each first sample desensitized log in the first sample desensitized log set; the second sample desensitized log set includes normal desensitized logs;
[0019] Based on at least one second sample feature vector in the second sample feature vector set and the second sample label carried by each second sample feature vector, the first abnormal desensitized log recognition model is fine-tuned to obtain a second abnormal desensitized log recognition model.
[0020] In a possible implementation, based on at least one second sample feature vector in the second sample feature vector set and a second sample label carried by each second sample feature vector, the first abnormal desensitization log recognition model is fine-tuned to obtain a second abnormal desensitization log recognition model, including:
[0021] Inputting at least one second sample feature vector into the first abnormal desensitization log recognition model to obtain a second prediction label corresponding to each second sample feature vector;
[0022] Determine a second classification loss of the first abnormal desensitization log recognition model according to each second prediction label and each second sample label;
[0023] The parameters of the first abnormal desensitized log recognition model are adjusted based on the second classification loss, and the step of obtaining the second sample feature vector set corresponding to the second sample desensitized log set is performed again until the second training stop condition is reached to obtain the second abnormal desensitized log recognition model.
[0024] In a second aspect, the embodiment of this specification provides a method for identifying abnormal desensitized logs, including:
[0025] Obtain a first target feature vector set corresponding to a first target desensitized log set to be identified;
[0026] Inputting at least one first target feature vector in the first target feature vector set into the first abnormal desensitization log recognition model to obtain a first identification label corresponding to each first target desensitization log;
[0027] Determine, according to each first identification tag, an abnormal desensitized log in the first target desensitized log set;
[0028] Among them, the first abnormal desensitized log recognition model is trained using the method provided in the first aspect of the embodiment of this specification.
[0029] In a possible implementation, after determining the abnormal desensitized log in the first target desensitized log set according to each first identification tag, the method further includes:
[0030] Generates abnormal log warning prompt information for optimizing abnormal desensitized logs.
[0031] In a third aspect, the embodiment of this specification provides a training device for an abnormal desensitization log recognition model, including:
[0032] A first acquisition module is used to acquire a first sample feature vector set corresponding to a first sample desensitized log set; the first sample desensitized log set includes a normal desensitized log;
[0033] The training module is used to train the anomaly detection model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector to obtain a first anomaly desensitized log recognition model.
[0034] In a fourth aspect, an embodiment of the present specification provides an abnormal desensitization log identification device, including:
[0035] A second acquisition module is used to acquire a first target feature vector set corresponding to a first target desensitized log set to be identified;
[0036] An identification module, used for inputting at least one first target feature vector in the first target feature vector set into a first abnormal desensitization log identification model to obtain a first identification label corresponding to each first target desensitization log;
[0037] A determination module, used to determine the abnormal desensitized log in the first target desensitized log set according to each first identification tag;
[0038] Among them, the first abnormal desensitized log recognition model is trained using the method provided in the first aspect of the embodiment of this specification.
[0039] In a fifth aspect, an embodiment of the present specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps provided in the first aspect or the second aspect of the embodiment of the present specification.
[0040] In a sixth aspect, an embodiment of this specification provides a computer program product, including a computer program; when the above-mentioned computer program is executed by a processor, the method steps provided in the first aspect or the second aspect of the embodiment of this specification are implemented.
[0041] The training method, device, electronic device, computer storage medium and computer program product of the above abnormal desensitized log recognition model obtains the first sample feature vector set corresponding to the first sample desensitized log set, and trains the first abnormal desensitized log recognition model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector. When it is impossible to collect real abnormal desensitized logs, the model can learn the feature distribution of normal desensitized logs through the first sample desensitized log set including normal desensitized logs, thereby identifying abnormal desensitized logs with different feature distributions from normal desensitized logs. The entire training process of the abnormal desensitized log recognition model can correctly identify whether unexpected log data (such as log data with an unusual format, etc.) is an abnormal desensitized log without the need for abnormal samples, and can enable the trained first abnormal desensitized log recognition model to effectively identify abnormal desensitized logs.
[0042] The above-mentioned abnormal desensitized log identification method, device, electronic device, computer storage medium and computer program product obtain the first target feature vector set corresponding to the first target desensitized log set to be identified, input at least one first target feature vector in the first target feature vector set into the first abnormal desensitized log identification model, obtain the first identification label corresponding to each first target desensitized log, and determine the abnormal desensitized log in the first target desensitized log set according to each first identification label. The first abnormal desensitized log identification model obtained by the above-mentioned training can be used to correctly identify whether unexpected log data (such as log data in an unusual format, etc.) is an abnormal desensitized log, thereby realizing effective identification of abnormal desensitized logs. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 A schematic diagram of the architecture of a method for training an abnormal desensitization log recognition model provided by an exemplary embodiment of this specification;
[0045] Figure 2 A flowchart of a method for training an abnormal desensitization log recognition model provided by an exemplary embodiment of this specification;
[0046] Figure 3 A flowchart of another method for training an abnormal desensitization log recognition model provided as an exemplary embodiment of this specification;
[0047] Figure 4 A flowchart of another method for training an abnormal desensitization log recognition model provided as an exemplary embodiment of this specification;
[0048] Figure 5 A flowchart of another method for training an abnormal desensitization log recognition model provided as an exemplary embodiment of this specification;
[0049] Figure 6 A flowchart of a method for identifying abnormal desensitized logs provided as an exemplary embodiment of this specification;
[0050] Figure 7 A schematic diagram of the architecture of a training model for identifying abnormal desensitized logs and a method for identifying abnormal desensitized logs provided as an exemplary embodiment of this specification;
[0051] Figure 8A schematic diagram of the architecture of another abnormal desensitized log identification model training and abnormal desensitized log identification method provided for an exemplary embodiment of this specification;
[0052] Fig. 9 A schematic diagram of the structure of a training device for an abnormal desensitization log recognition model provided by an exemplary embodiment of this specification;
[0053] Fig.10 A schematic diagram of the structure of an abnormal desensitization log identification device provided by an exemplary embodiment of this specification;
[0054] Fig.11 The present invention is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this specification more clear, the following is a further detailed description of this specification in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this specification and are not used to limit this specification.
[0056] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previously associated objects are in an "or" relationship.
[0057] See also Figure 1 , is a schematic diagram of the architecture of a training method for an abnormal desensitization log recognition model provided by an exemplary embodiment of this specification. Among them, the terminal 10 communicates with the server 20 through a network. The data storage system can store data that the server 20 needs to process. The data storage system can be integrated on the server 20, or it can be placed on the cloud or other network servers.
[0058] In some possible embodiments, the training method of the abnormal desensitized log recognition model provided in this specification can be jointly executed by the terminal 10 and the server 20. Specifically, the server 20 responds to the training sample acquisition command sent by the terminal 10 to obtain the first sample feature vector set corresponding to the first sample desensitized log set; the first sample desensitized log set includes normal desensitized logs; the server 20 responds to the model training command sent by the terminal 10, and trains the anomaly detection model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector, to obtain the first abnormal desensitized log recognition model. Accordingly, the training device of the abnormal desensitized log recognition model can also be respectively set in the terminal 10 and the server 20.
[0059] In some possible embodiments, the training method of the abnormal desensitization log recognition model provided in this specification can be executed by the terminal 10. Accordingly, the training device of the abnormal desensitization log recognition model can also be set in the terminal 10.
[0060] In some possible embodiments, the training method of the abnormal desensitization log recognition model provided in this specification may be executed by the server 20. Accordingly, the training device of the abnormal desensitization log recognition model may also be set in the server 20.
[0061] It is worth noting that the terminal 10 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices, etc. The server 20 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0062] In one embodiment, Figure 2 As shown in the figure, a training method for an abnormal desensitized log recognition model is provided, and the method is applied to Figure 1 The server 20 in the example is used as an example to illustrate, and the following steps are included:
[0063] S202: Obtain a first sample feature vector set corresponding to a first sample desensitized log set; the first sample desensitized log set includes normal desensitized logs.
[0064] The first sample desensitized log set includes at least one pre-collected first sample desensitized log, and each first sample desensitized log may be, but is not limited to, formatted and output by the log tool class provided by the log desensitization component. The first sample feature vector set includes at least one first sample feature vector for training the first abnormal desensitized log recognition model, and each first sample feature vector is extracted from the features of each first sample desensitized log.
[0065] Understandably, since abnormal desensitized logs (i.e., un-desensitized logs) account for a very small proportion of online real data, it is often difficult to collect sufficiently abundant abnormal desensitized logs as training samples. Therefore, this embodiment obtains a first abnormal desensitized log recognition model based on anomaly detection model training, and trains the model by using a first sample desensitized log set including normal desensitized logs as a training data set, so that the model learns the characteristic distribution of normal desensitized logs, thereby identifying abnormal desensitized logs with different characteristic distribution from normal desensitized logs.
[0066] Optionally, the server 20 obtains a first sample feature vector set corresponding to the first sample desensitized log set from a pre-built first log feature library. The first log feature library stores a plurality of pre-collected first sample desensitized logs, and the first sample desensitized log set includes at least one first sample desensitized log stored in the first log feature library.
[0067] S204: Based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector, an anomaly detection model is trained to obtain a first anomaly desensitized log recognition model.
[0068] Among them, the first sample label carried by each first sample feature vector is used to characterize the first sample desensitized log type (including normal desensitized log and abnormal desensitized log) corresponding to each first sample feature vector. The anomaly detection model can be, but is not limited to, an autoencoder model, a generative adversarial network model (GAN), a convolutional neural network model (CNN), a deep belief network model (DBN), etc.
[0069] Optionally, the server 20 inputs at least one first sample feature vector into an anomaly detection model (such as an autoencoder model, etc.) after random initialization parameters, obtains a first prediction label corresponding to each first sample feature vector, determines a first classification loss of the anomaly detection model based on each first prediction label and each first sample label, adjusts the parameters of the anomaly detection model based on the first classification loss, and executes the step of obtaining a first sample feature vector set corresponding to a first sample desensitized log set again, until the first training stop condition is reached, and obtains a first abnormal desensitized log recognition model. The first training stop condition may be, but is not limited to, the first abnormal desensitized log recognition model converging or reaching a first preset number of iterations.
[0070] The above abnormal desensitized log identification method obtains the first sample feature vector set corresponding to the first sample desensitized log set, and trains the first abnormal desensitized log identification model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector. When it is impossible to collect real abnormal desensitized logs, the model can learn the feature distribution of normal desensitized logs through the first sample desensitized log set including normal desensitized logs, thereby identifying abnormal desensitized logs with different feature distributions from normal desensitized logs. The entire training process of the abnormal desensitized log identification model can correctly identify whether unexpected log data (such as log data in an unusual format, etc.) is an abnormal desensitized log without the need for abnormal samples, and can enable the trained first abnormal desensitized log identification model to effectively identify abnormal desensitized logs.
[0071] In one embodiment, Figure 3 As shown in the figure, another training method for abnormal desensitization log recognition model is provided, which is applied to Figure 1 The server 20 in the example is used as an example to illustrate, and the following steps are included:
[0072] S302: Input a batch of first sample desensitized logs into a feature extraction model to obtain a first sample feature vector corresponding to each first sample desensitized log in the batch of first sample desensitized logs.
[0073] Among them, the feature extraction model can be but is not limited to a term frequency-inverse document frequency (Term Frequency-Inverse Document Frequency, TF-IDF) model, a word vectorization (Word to Vector, Word2Vec) model, etc.
[0074] Optionally, the server 20 pre-collects a batch of first sample desensitized logs from at least one online log system, and performs data preprocessing on the collected batch of first sample desensitized logs, removes redundant text content in the batch of first sample desensitized logs, such as removing system status information including machine IP, link information, and log printing time, and only retains text content with semantic value in the batch of first sample desensitized logs. Then, the first sample desensitized logs after data preprocessing are input into a feature extraction model (such as a TF-IDF model) to extract the first sample feature vectors corresponding to each first sample desensitized log in the batch of first sample desensitized logs.
[0075] S304: Construct a first log feature library based on each first sample feature vector and the first sample label carried by each first sample feature vector.
[0076] Among them, the first sample label carried by each first sample feature vector is used to characterize the first sample desensitized log type (including normal desensitized log and abnormal desensitized log) corresponding to each first sample feature vector. The first log feature library stores a training data set for training the first abnormal desensitized log recognition model. During the model training process, the server 20 can directly retrieve the first sample feature vector set corresponding to the first sample desensitized log set from the first log feature library without re-extracting the first sample feature vector corresponding to the original first sample desensitized log.
[0077] Optionally, the server 20 constructs a first log feature library based on each first sample feature vector output by the feature extraction model, and a first sample label used to characterize the first sample desensitized log type corresponding to each first sample feature vector, so that during the model training process, at least one first sample feature vector (i.e., a set of first sample feature vectors) carrying the first sample label can be directly extracted from the constructed first log feature library. Specifically, the server 20 can store each first sample feature vector in association with the first sample label corresponding to each first sample feature vector, construct a first log feature library, and divide the training data stored in the first log feature library into multiple training batches before model training, i.e., into multiple sets of first sample feature vectors, so that during the model training process, the first sample feature vector set can be directly extracted from the constructed first log feature library to be input into the anomaly detection model for training.
[0078] S306: Obtain a first sample feature vector set corresponding to the first sample desensitized log set from the first log feature library.
[0079] It is understandable that due to the limitation of computing resources, the full amount of training data stored in the first log feature library cannot be input into the anomaly detection model for training during the model training process. Therefore, the server 20 divides the multiple first sample feature vectors carrying the first sample label stored in the first log feature library into multiple training batches in advance, that is, into multiple first sample feature vector sets, each of which includes at least one first sample feature vector carrying the first sample label. During the model training process, a method of inputting batches into the anomaly detection model for training is adopted, and one of the first sample feature vector sets in the multiple first sample feature vector sets is selected in turn from the first log feature library, and inputted into the anomaly detection model for training.
[0080] Specifically, after the server 20 divides the multiple first sample feature vectors carrying the first sample labels stored in the first log feature library into multiple first sample feature vector sets, it also sets training rounds. Completing each training round represents that the anomaly detection model has traversed all the training data stored in the first log feature library, that is, it has traversed the multiple divided first sample feature vector sets. Exemplarily, assuming that the multiple first sample feature vectors carrying the first sample labels stored in the first log feature library are divided into N batches, and the training rounds are set to M, then when training the model, it is necessary to traverse all the first sample feature vectors stored in the first log feature library at least N times, and each traversal requires the divided N first sample feature vector sets to be respectively input into the anomaly detection model for training. Normally, the training round M is greater than 1, that is, each first sample feature vector set needs to be repeatedly input into the anomaly detection model for training.
[0081] In this embodiment, based on the feature extraction model, multiple first sample feature vectors corresponding to a batch of first sample desensitized logs are extracted, and a first log feature library is constructed according to multiple first sample feature vectors and the first sample labels they carry, so that during the model training process, the first sample feature vectors required for each training can be directly selected from the first log feature library without using the feature extraction model to re-extract the first sample feature vectors corresponding to each first sample desensitized log each time. The above process can simplify the data processing flow, realize data reuse, and thus improve the efficiency of model training.
[0082] S308: Input at least one first sample feature vector into an anomaly detection model to obtain a first prediction label corresponding to each first sample feature vector.
[0083] Optionally, the server 20 inputs the multiple first sample feature vector sets stored in the first log feature library into the anomaly detection model one by one for training. In the process of inputting each first sample feature vector set into the anomaly detection model each time, specifically, at least one first sample feature vector contained in each first sample feature vector set is input into the anomaly detection model with randomly initialized parameters for forward propagation to obtain the first prediction label corresponding to each first sample feature vector output by the anomaly detection model. It can be understood that the first prediction label corresponding to each first sample feature vector is used to characterize the prediction type (including normal desensitized log and abnormal desensitized log) of the first sample desensitized log corresponding to each first sample feature vector.
[0084] S310: Determine a first classification loss of an anomaly detection model according to each first prediction label and each first sample label.
[0085] Optionally, the server 20 calculates the total classification loss based on the first prediction labels corresponding to all the first sample feature vectors input into the anomaly detection model each time, and the first sample labels corresponding to all the first sample feature vectors, and uses it as the first classification loss of the anomaly detection model. Among them, the first classification loss can be but is not limited to being calculated by the cross-entropy loss function (Cross-Entropy Loss). It can be understood that the smaller the first classification loss, the more accurate the prediction result of the anomaly detection model for the first sample desensitized log type.
[0086] S312: Adjust parameters of the anomaly detection model based on the first classification loss.
[0087] Optionally, the server 20 performs back propagation based on the calculated first classification loss, and adjusts the parameters of the anomaly detection model through optimization algorithms such as gradient descent, so as to minimize the first classification loss, so that the anomaly detection model can predict the type of the first sample desensitized log more accurately.
[0088] S314: Determine whether the first training stop condition is met. If yes, execute S316; if no, execute S306.
[0089] The first training stop condition may be, but is not limited to, the current anomaly detection model converging, reaching a first preset number of iterations, etc.
[0090] Optionally, the server 20 determines whether the anomaly detection model has converged by observing the first classification loss output by each iterative anomaly detection model. If the current anomaly detection model has converged, the current anomaly detection model is used as the first anomaly desensitized log recognition model; if the current anomaly detection model has not converged, the first sample feature vector set corresponding to the first sample desensitized log set is obtained from the first log feature library again for the next iteration until the anomaly detection model converges.
[0091] S316: Use the current anomaly detection model as the first anomaly desensitized log recognition model.
[0092] In this embodiment, by constructing a first log feature library, and dividing the multiple first sample feature vectors carrying the first sample labels stored in the first log feature library into multiple training batches and inputting them into the anomaly detection model, the parameters of the anomaly detection model are continuously adjusted according to the first classification loss obtained from each training of the anomaly detection model until the model converges, thereby obtaining the first abnormal desensitized log recognition model, which can not only effectively improve the accuracy of the prediction results of the first abnormal desensitized log recognition model, but also improve the training efficiency of the first abnormal desensitized log recognition model.
[0093] The training method of the above-mentioned abnormal desensitized log recognition model collects real log data from at least one online log system for training, uses a feature extraction model to pre-extract multiple first sample feature vectors corresponding to a batch of first sample desensitized logs, and constructs a first log feature library according to the multiple first sample feature vectors and the first sample labels they carry. In the model training process, the multiple first sample feature vectors carrying the first sample labels stored in the first log feature library are divided into multiple training batches and input into the anomaly detection model, so as to continuously adjust the parameters of the anomaly detection model according to the first classification loss obtained from each training of the anomaly detection model until the model converges, thereby obtaining the first abnormal desensitized log recognition model, which can not only effectively improve the accuracy of the prediction results of the first abnormal desensitized log recognition model, but also improve the efficiency of the training of the first abnormal desensitized log recognition model.
[0094] In one embodiment, Figure 4 As shown in Figure 2, another training method for abnormal desensitization log recognition model is provided, which is applied to Figure 1 The server 20 in the example is used as an example to illustrate, and the following steps are included:
[0095] S402: Obtain a second sample feature vector set corresponding to the second sample desensitized log set.
[0096] Among them, the data format and / or data semantics of each second sample desensitized log in the second sample desensitized log set are different from those of each first sample desensitized log in the first sample desensitized log set; the second sample desensitized log set includes normal desensitized logs.
[0097] It is understandable that due to system updates and other reasons, the data format and / or data semantics of the desensitized logs in the online log system may change. In this case, it is necessary to obtain at least one second sample desensitized log with a different data format and / or data semantics from the first sample desensitized log from the online log system, and fine-tune the first abnormal desensitized log recognition model based on the log features of the second sample desensitized log. Figure 2 or Figure 3 The model obtained by training the abnormal desensitization log recognition model shown in the figure.
[0098] Optionally, the server 20 obtains a second sample feature vector set corresponding to the second sample desensitized log set from a pre-built second log feature library. The second log feature library stores a plurality of pre-collected second sample desensitized logs, and the second sample desensitized log set includes at least one second sample desensitized log stored in the second log feature library.
[0099] S404: Based on at least one second sample feature vector in the second sample feature vector set and a second sample label carried by each second sample feature vector, fine-tune the first abnormal desensitized log recognition model to obtain a second abnormal desensitized log recognition model.
[0100] Optionally, the server 20 inputs at least one second sample feature vector into the first abnormal desensitization log recognition model, obtains the second prediction label corresponding to each second sample feature vector, determines the second classification loss of the first abnormal desensitization log recognition model according to each second prediction label and each second sample label, adjusts the parameters of the first abnormal desensitization log recognition model based on the second classification loss, and executes the step of obtaining the second sample feature vector set corresponding to the second sample desensitization log set again, until the second training stop condition is reached, and obtains the second abnormal desensitization log recognition model. The second training stop condition may be, but is not limited to, the convergence of the first abnormal desensitization log recognition model or reaching the second preset number of iterations.
[0101] The above-mentioned abnormal desensitized log identification method obtains a second sample desensitized log set whose data format and / or data semantics are different from the first sample desensitized log, and trains the model through at least one second sample feature vector in the second sample feature vector set and the second sample label carried by each second sample feature vector, and performs model fine-tuning on the basis of the pre-trained first abnormal desensitized log identification model, which can improve the training efficiency of the abnormal desensitized log identification model, quickly adapt to the updated system, and thus effectively identify the updated abnormal desensitized logs.
[0102] In one embodiment, Figure 5 As shown in Figure 2, another training method for abnormal desensitization log recognition model is provided, which is applied to Figure 1 The server 20 in the example is used as an example to illustrate, and the following steps are included:
[0103] S502: Input a batch of second sample desensitized logs into a feature extraction model to obtain a second sample feature vector corresponding to each second sample desensitized log in the batch of second sample desensitized logs.
[0104] Optionally, the server 20 pre-collects a batch of second sample desensitized logs whose data format and / or data semantics are different from the first sample desensitized logs, and performs data preprocessing on the collected batch of second sample desensitized logs, removing redundant text content in the batch of second sample desensitized logs, such as removing system status information including machine IP, link information, and log printing time, and only retaining text content with semantic value in the batch of second sample desensitized logs. Then, the second sample desensitized logs after data preprocessing are input into a feature extraction model (such as a TF-IDF model) to extract the second sample feature vectors corresponding to each second sample desensitized log in the batch of second sample desensitized logs.
[0105] S504: Construct a second log feature library based on each second sample feature vector and the second sample label carried by each second sample feature vector.
[0106] Among them, the second sample label carried by each second sample feature vector is used to characterize the second sample desensitized log type (including normal desensitized log and abnormal desensitized log) corresponding to each second sample feature vector. The second log feature library stores a training data set for fine-tuning the second abnormal desensitized log recognition model. During the model fine-tuning process, the server 20 can directly retrieve the second sample feature vector set corresponding to the second sample desensitized log set from the second log feature library without re-extracting the second sample feature vector corresponding to the original second sample desensitized log.
[0107] Optionally, the server 20 constructs a second log feature library based on each second sample feature vector output by the feature extraction model, and a second sample label used to characterize the second sample desensitized log type corresponding to each second sample feature vector, so that during the model training process, at least one second sample feature vector (i.e., a set of second sample feature vectors) carrying a second sample label can be directly extracted from the constructed second log feature library. Specifically, the server 20 can store each second sample feature vector in association with the second sample label corresponding to each second sample feature vector, construct a second log feature library, and divide the training data stored in the second log feature library into multiple training batches, i.e., into multiple sets of second sample feature vectors before fine-tuning the model, so that during the model fine-tuning process, the second sample feature vector set can be directly extracted from the constructed second log feature library to input into the first abnormal desensitized log recognition model for fine-tuning.
[0108] S506: Obtain a second sample feature vector set corresponding to the second sample desensitized log set from the second log feature library.
[0109] Optionally, the server 20 divides the multiple second sample feature vectors carrying the second sample label stored in the second log feature library into multiple training batches, that is, into multiple second sample feature vector sets, each of which includes at least one second sample feature vector carrying the second sample label. In the process of fine-tuning the model, the method of inputting batches into the first abnormal desensitized log recognition model for fine-tuning is adopted, and one of the multiple second sample feature vector sets is selected from the second log feature library in turn, and inputted into the first abnormal desensitized log recognition model for fine-tuning.
[0110] Specifically, after the server 20 divides the multiple second sample feature vectors carrying the second sample labels stored in the second log feature library into multiple second sample feature vector sets, it also sets training rounds. Completing each training round represents that the first abnormal desensitization log recognition model has traversed all the training data stored in the second log feature library, that is, it has traversed the multiple divided second sample feature vector sets.
[0111] In this embodiment, based on the feature extraction model, multiple second sample feature vectors corresponding to a batch of second sample desensitized logs are extracted, and a second log feature library is constructed according to multiple second sample feature vectors and the second sample labels they carry, so that in the process of model fine-tuning, the second sample feature vectors required for each fine-tuning can be directly selected from the second log feature library without using the feature extraction model to re-extract the second sample feature vectors corresponding to each second sample desensitized log each time. The above process can simplify the data processing flow, realize data reuse, and thus improve the efficiency of model training.
[0112] S508: Input at least one second sample feature vector into the first abnormal desensitization log recognition model to obtain a second prediction label corresponding to each second feature vector.
[0113] Optionally, the server 20 inputs at least one second sample feature vector contained in each second sample feature vector set into a first abnormal desensitization log recognition model with predetermined parameters for forward propagation, so that the first abnormal desensitization log recognition model outputs a second prediction label corresponding to each second sample feature vector. It can be understood that the second prediction label corresponding to each second sample feature vector is used to characterize the prediction type (including normal desensitization log and abnormal desensitization log) of the second sample desensitization log corresponding to each second sample feature vector.
[0114] S510: Determine a second classification loss of the first abnormal desensitized log recognition model according to each second prediction label and each second sample label.
[0115] Optionally, the server 20 calculates the total classification loss based on the second prediction labels corresponding to all the second sample feature vectors input into the first abnormal desensitization log recognition model, and the second sample labels corresponding to all the second sample feature vectors, and uses it as the second classification loss of the first abnormal desensitization log recognition model. Among them, the second classification loss can be but is not limited to being calculated by the cross entropy loss function. It can be understood that the smaller the second classification loss, the more accurate the prediction result of the first abnormal desensitization log recognition model for the second sample desensitization log type.
[0116] S512: Adjust parameters of the first abnormal desensitization log recognition model based on the second classification loss.
[0117] Optionally, the server 20 performs back propagation based on the calculated second classification loss, and adjusts the parameters of the first abnormal desensitized log recognition model through optimization algorithms such as gradient descent, thereby minimizing the second classification loss, making the first abnormal desensitized log recognition model more accurate in predicting the type of the second sample desensitized log.
[0118] S514: Determine whether the second training stop condition is met. If yes, execute S516; if no, execute S506.
[0119] Among them, the second training stop condition can be but is not limited to the convergence of the current first abnormal desensitization log recognition model, reaching the second preset number of iterations, etc.
[0120] Optionally, the server 20 determines whether the first abnormal desensitization log recognition model has converged by observing the second classification loss output by the first abnormal desensitization log recognition model in each iteration. If the current first abnormal desensitization log recognition model has converged, the current first abnormal desensitization log recognition model is used as the second abnormal desensitization log recognition model; if the current first abnormal desensitization log recognition model has not converged, the second sample feature vector set corresponding to the second sample desensitization log set is obtained from the second log feature library again to perform the next iteration until the first abnormal desensitization log recognition model converges.
[0121] S516: Use the current first abnormal desensitization log recognition model as the second abnormal desensitization log recognition model.
[0122] In this embodiment, a second log feature library is constructed, and multiple second sample feature vectors carrying second sample labels stored in the second log feature library are divided into multiple training batches and input into a first abnormal desensitized log recognition model with predetermined parameters, so as to continuously adjust the parameters of the first abnormal desensitized log recognition model according to the second classification loss obtained from each training of the first abnormal desensitized log recognition model until the first abnormal desensitized log recognition model converges, thereby obtaining a second abnormal desensitized log recognition model, which can not only quickly adapt to the updated system logs through model fine-tuning, but also ensure the accuracy of the second abnormal desensitized log recognition model in identifying abnormal desensitized logs.
[0123] In one embodiment, Figure 6 As shown, a method for identifying abnormal desensitized logs is provided, and the method is applied to Figure 1 The server 20 in the example is used as an example to illustrate, and the following steps are included:
[0124] S602: Obtain a first target feature vector set corresponding to a first target desensitized log set to be identified.
[0125] Among them, the first target desensitized log set includes at least one asynchronously acquired first target desensitized log, and each first target desensitized log can be, but is not limited to, formatted and output by the log tool class provided by the log desensitization component. The first target feature vector set includes at least one first sample feature vector for input into the first abnormal desensitized log recognition model to identify the corresponding first target desensitized log type, and each first target feature vector is extracted from the features of each first target desensitized log.
[0126] Optionally, the server 20 obtains the first target desensitized log set to be identified in at least one online log system in an asynchronous manner, and obtains the first target feature vector set corresponding to the first target desensitized log set to be identified through a feature extraction model (such as a TF-IDF model). Specifically, the server 20 asynchronously collects at least one first target desensitized log in the online log system through a log collection service or a timed export of log files, and performs data preprocessing on the collected at least one first target desensitized log, removes redundant text content in at least one first target desensitized log, such as removing system status information including machine IP, link information, and log printing time, and only retains text content with semantic value in at least one first target desensitized log. Then, the preprocessed at least one first target desensitized log is input into the feature extraction model to obtain at least one first target feature vector to be input into the trained first abnormal desensitized log recognition model.
[0127] In this embodiment, the server collects a first target desensitized log set to be identified in at least one online log system in an asynchronous manner, and can collect and identify the logs to be identified in the online log system without affecting the normal operation of the online log system, thereby improving the reliability and stability of abnormal desensitized log identification.
[0128] S604: Input at least one first target feature vector in the first target feature vector set into the first abnormal desensitization log recognition model to obtain a first identification label corresponding to each first target desensitization log.
[0129] Among them, the first abnormal desensitized log recognition model adopts Figure 2 or Figure 3 The model obtained by training the abnormal desensitized log identification model training method shown. The first identification label corresponding to each first target desensitized log is used to characterize the first target desensitized log type (including normal desensitized log and abnormal desensitized log) predicted by the first abnormal desensitized log identification model.
[0130] Optionally, the server 20 inputs at least one first target feature vector in the first target feature vector set into a trained first abnormal desensitization log recognition model to obtain a first identification label corresponding to each first target desensitization log output by the first abnormal desensitization log recognition model.
[0131] S606: Determine the abnormal desensitized log in the first target desensitized log set according to each first identification tag.
[0132] It can be understood that the first identification label can be, but is not limited to, a binary label, etc. In the case where the first identification label is a binary label, the first identification label is 0, which means that the first target desensitized log type predicted by the first abnormal desensitized log recognition model is a normal desensitized log, and the first identification label is 1, which means that the first target desensitized log type predicted by the first abnormal desensitized log recognition model is an abnormal desensitized log.
[0133] In this embodiment, by obtaining the first target feature vector set corresponding to the first target desensitized log set to be identified, at least one first target feature vector in the first target feature vector set is input into the first abnormal desensitized log recognition model to obtain the first identification label corresponding to each first target desensitized log, and according to each first identification label, the abnormal desensitized log in the first target desensitized log set is determined. The first abnormal desensitized log recognition model obtained by the above training can be used to correctly identify whether the unexpected log data is an abnormal desensitized log, thereby realizing effective identification of abnormal desensitized logs.
[0134] S608: Generate abnormal log warning prompt information for prompting optimization of abnormal desensitized logs.
[0135] Optionally, when the server 20 identifies the presence of abnormal desensitized logs in the first target desensitized log set, it generates abnormal log alarm prompt information for prompting the optimization of the abnormal desensitized logs, such as "Log A is an abnormal desensitized log, please re-desensitize Log A", and sends the abnormal log alarm prompt information to the terminal for display.
[0136] In this embodiment, after identifying abnormal desensitized logs, abnormal log alarm prompt information is generated to prompt users to optimize abnormal desensitized logs, which can effectively prevent key data leakage and improve the reliability of abnormal desensitized log identification and the security of log data.
[0137] The above abnormal desensitized log identification method improves the reliability and stability of abnormal desensitized log identification by asynchronously collecting the first target desensitized log set to be identified in the online log system; determines the abnormal desensitized logs in the first target desensitized log set through the pre-trained first abnormal desensitized log identification model, thereby improving the accuracy of abnormal desensitized log identification; and generates abnormal log warning prompt information after identifying the abnormal log, thereby improving the security of log data. The above method can realize the effective identification of abnormal desensitized logs.
[0138] In order to explain in detail the training method of the abnormal desensitization log recognition model and the technical solution of the abnormal desensitization log recognition method in one or more embodiments of this specification, the following will use specific application examples and combine Figure 7 and Figure 8 The whole processing process is described, which specifically includes the following steps:
[0139] 1. The server builds the first log feature library based on a batch of first sample anonymized logs. The specific processing is as follows:
[0140] a) Collect a batch of first sample desensitized logs from at least one online log system (such as system A, system B, etc.); the batch of first sample desensitized logs includes normal desensitized logs.
[0141] b) Perform data preprocessing on the collected first batch of desensitized logs, and remove system status information such as machine IP, link information, log printing time, etc. in the first batch of desensitized logs.
[0142] c) Inputting the preprocessed batch of first sample desensitized logs into the feature extraction model to obtain the first sample feature vector corresponding to each first sample desensitized log in the batch of first sample desensitized logs.
[0143] d) collecting first sample feature vectors output by the feature extraction model, and constructing a first log feature library based on each first sample feature vector and a first sample label carried by each first sample feature vector.
[0144] 2. The server obtains a first sample feature vector set from the first log feature library, and trains the anomaly detection model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector to obtain a first anomaly desensitized log recognition model. The specific processing is as follows:
[0145] a) Inputting at least one first sample feature vector into an anomaly detection model (such as an autoencoder model) to obtain a first prediction label corresponding to each first sample feature vector.
[0146] b) Determine a first classification loss (such as a cross entropy loss) of the anomaly detection model based on each first prediction label and each first sample label.
[0147] c) Adjust the parameters of the anomaly detection model based on the first classification loss, and execute a) in step 2 again until the first training stop condition is reached (such as the anomaly detection model converges), and obtain the first anomaly desensitized log recognition model.
[0148] 3. The server identifies the abnormal desensitized logs in the online log system based on the trained first abnormal desensitized log identification model. The specific processing is as follows:
[0149] a) A first target desensitized log set to be identified, which is generated by at least one online log system (such as system A, etc.) is obtained in an asynchronous manner; the first target desensitized log set includes at least one first target desensitized log.
[0150] b) Perform data preprocessing on the acquired first target desensitized log set to remove system status information such as machine IP, link information, log printing time, etc. in the first target desensitized log set.
[0151] c) Input the preprocessed first target desensitized log set into the feature extraction model to obtain a first target feature vector set corresponding to the first target desensitized log set.
[0152] d) Inputting at least one first target feature vector in the first target feature vector set into the first abnormal desensitized log recognition model to obtain a first identification label corresponding to each first target desensitized log.
[0153] e) Determine the abnormal desensitized logs in the first target desensitized log set according to each first identification tag.
[0154] 4. When the server identifies that there are abnormal desensitized logs in the first target desensitized log set, it generates abnormal log warning prompt information for prompting optimization of the abnormal desensitized logs, and sends it to the terminal for display.
[0155] 5. When the data format and / or data semantics of the desensitized logs in the online log system change, the server builds a second log feature library based on a batch of second sample desensitized logs. The specific processing is as follows:
[0156] a) Collect a batch of second sample desensitized logs from at least one updated online log system (such as system A', system B', etc.); the batch of second sample desensitized logs includes normal desensitized logs, and the data format and / or data semantics of each second sample desensitized log in the batch of second sample desensitized logs are different from those of each first sample desensitized log.
[0157] b) Perform data preprocessing on the collected batch of second sample anonymized logs to remove system status information such as machine IP, link information, log printing time, etc. in the batch of second sample anonymized logs.
[0158] c) Inputting the preprocessed batch of second sample desensitized logs into the feature extraction model to obtain the second sample feature vector corresponding to each second sample desensitized log in the batch of second sample desensitized logs.
[0159] 6. The server obtains a second sample feature vector set from the second log feature library, and fine-tunes the first abnormal desensitized log recognition model based on at least one second sample feature vector in the second sample feature vector set and the second sample label carried by each second sample feature vector, to obtain a second abnormal desensitized log recognition model. The specific processing is as follows:
[0160] a) Inputting at least one second sample feature vector into the first abnormal desensitization log recognition model to obtain a second prediction label corresponding to each second sample feature vector.
[0161] b) Determine a second classification loss (such as a cross entropy loss) of the first abnormal desensitization log recognition model according to each second prediction label and each second sample label.
[0162] c) Adjust the parameters of the first abnormal desensitization log recognition model based on the second classification loss, and execute a) in step 5 again until the second training stop condition is reached (such as the first abnormal desensitization log recognition model converges), and obtain the second abnormal desensitization log recognition model.
[0163] 7. The server identifies the abnormal desensitized logs in the updated online log system based on the trained second abnormal desensitized log identification model. The specific processing is as follows:
[0164] a) A second target desensitized log set to be identified, which is produced by at least one updated online log system (such as system A', etc.) is obtained in an asynchronous manner; the second target desensitized log set includes at least one second target desensitized log, and the data format and / or data semantics of each second target desensitized log in the second target desensitized log set are different from those of each first target desensitized log in the first target desensitized log set.
[0165] b) Perform data preprocessing on the acquired second target anonymized log set to remove system status information such as machine IP, link information, log printing time, etc. in the second target anonymized log set.
[0166] c) Input the preprocessed second target desensitized log set into the feature extraction model to obtain a second target feature vector set corresponding to the second target desensitized log set.
[0167] d) Inputting at least one second target feature vector in the second target feature vector set into the second abnormal desensitization log recognition model to obtain a second identification label corresponding to each second target desensitization log.
[0168] 8. When the server identifies that there are abnormal desensitized logs in the second target desensitized log set, it generates abnormal log warning prompt information for prompting optimization of the abnormal desensitized logs, and sends it to the terminal for display.
[0169] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0170] The invention concept of the training method based on the above abnormal desensitization log recognition model is as follows: Fig. 9 As shown, the embodiment of this specification also provides a training device 900 for an abnormal desensitization log recognition model for implementing the training method of the abnormal desensitization log recognition model involved above. The training device 900 for the abnormal desensitization log recognition model includes:
[0171] The first acquisition module 901 is used to acquire a first sample feature vector set corresponding to a first sample desensitized log set; the first sample desensitized log set includes a normal desensitized log;
[0172] The training module 902 is used to train the anomaly detection model based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector to obtain a first anomaly desensitized log recognition model.
[0173] In one possible implementation, the training module 902 is specifically used to: input at least one first sample feature vector into the anomaly detection model to obtain a first prediction label corresponding to each first sample feature vector; determine a first classification loss of the anomaly detection model based on each first prediction label and each first sample label; adjust the parameters of the anomaly detection model based on the first classification loss, and execute again the step of obtaining a first sample feature vector set corresponding to the first sample desensitized log set until the first training stop condition is reached, thereby obtaining a first abnormal desensitized log recognition model.
[0174] In a possible implementation, the training device 900 for the abnormal desensitized log recognition model also includes a log feature library construction module, which is used to input a batch of first sample desensitized logs into a feature extraction model to obtain a first sample feature vector corresponding to each first sample desensitized log in a batch of first sample desensitized logs; based on each first sample feature vector and the first sample label carried by each first sample feature vector, a first log feature library is constructed; the first acquisition module 901 is specifically used to: obtain a first sample feature vector set corresponding to the first sample desensitized log set from the first log feature library.
[0175] In one possible implementation, the training device 900 for the abnormal desensitized log recognition model also includes a model fine-tuning module, which is used to obtain a second sample feature vector set corresponding to the second sample desensitized log set; each second sample desensitized log in the second sample desensitized log set has a different data format and / or data semantics from each first sample desensitized log in the first sample desensitized log set; the second sample desensitized log set includes normal desensitized logs; based on at least one second sample feature vector in the second sample feature vector set and a second sample label carried by each second sample feature vector, the first abnormal desensitized log recognition model is fine-tuned to obtain a second abnormal desensitized log recognition model.
[0176] In one possible implementation, the model fine-tuning module is specifically used to: input at least one second sample feature vector into the first abnormal desensitized log recognition model to obtain a second prediction label corresponding to each second sample feature vector; determine the second classification loss of the first abnormal desensitized log recognition model according to each second prediction label and each second sample label; adjust the parameters of the first abnormal desensitized log recognition model based on the second classification loss, and execute again the step of obtaining a second sample feature vector set corresponding to the second sample desensitized log set, until the second training stop condition is reached, and obtain the second abnormal desensitized log recognition model.
[0177] Based on the inventive concept of the above abnormal desensitized log identification method, Fig.10As shown, the embodiment of this specification also provides an abnormal desensitization log identification device 100 for implementing the abnormal desensitization log identification method involved above. The abnormal desensitization log identification device 100 includes:
[0178] The second acquisition module 101 is used to acquire a first target feature vector set corresponding to a first target desensitized log set to be identified;
[0179] The identification module 102 is used to input at least one first target feature vector in the first target feature vector set into the first abnormal desensitization log identification model to obtain a first identification label corresponding to each first target desensitization log; wherein the first abnormal desensitization log identification model is trained using the training method of the abnormal desensitization log identification model;
[0180] The determination module 103 is used to determine the abnormal desensitized log in the first target desensitized log set according to each first identification tag.
[0181] In a possible implementation, the abnormal desensitized log identification device 100 further includes a generation module for generating abnormal log warning prompt information for prompting optimization of abnormal desensitized logs.
[0182] The various modules in the training device 900 of the abnormal desensitization log recognition model and the abnormal desensitization log recognition device 100 can be implemented in whole or in part by software, hardware, and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0183] The embodiment of the present specification also provides an electronic device, which may be a server, and its internal structure diagram may be as follows: Fig.11As shown. The electronic device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The application database of the electronic device is used to store abnormal desensitization log data. The input / output interface of the electronic device is used to exchange information between the processor and an external device. The communication interface of the electronic device is used to communicate with an external terminal through a network connection. The processor of the electronic device executes a computer program to implement a training method for an abnormal desensitization log recognition model or an abnormal desensitization log recognition method.
[0184] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of this specification, and does not constitute a limitation on the electronic device to which the scheme of this specification is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0185] In a possible implementation, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0186] In a possible implementation manner, a computer storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0187] In a possible implementation manner, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0188] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of this specification is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (Digital Subscriber Line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that contains one or more available media integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0189] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes. In the absence of conflict, the technical features in this embodiment and the implementation scheme can be combined arbitrarily.
[0190] The embodiments described above are merely preferred embodiments of this specification and are not intended to limit the scope of this specification. Without departing from the design spirit of this specification, various modifications and improvements made to the technical solutions of this specification by ordinary technicians in this field should fall within the scope of protection determined by the claims.
[0191] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions. For example, the "first sample desensitized log set", "first target desensitized log set", "second sample desensitized log set", "second target desensitized log set" and so on involved in this specification are all obtained with full authorization.
[0192] The above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A training method for an abnormal desensitization log recognition model, the method comprising: Obtain a first sample feature vector set corresponding to a first sample desensitized log set; the first sample desensitized log set includes a normal desensitized log; Based on at least one first sample feature vector in the first sample feature vector set and the first sample label carried by each of the first sample feature vectors, the anomaly detection model is trained to obtain a first anomaly desensitized log recognition model.
2. The method according to claim 1, wherein the anomaly detection model is trained based on each first sample feature vector in the first sample feature vector set and the first sample label carried by each first sample feature vector to obtain an abnormal desensitized log recognition model, comprising: Inputting the at least one first sample feature vector into the anomaly detection model to obtain a first prediction label corresponding to each first sample feature vector; Determining a first classification loss of the anomaly detection model according to each of the first prediction labels and each of the first sample labels; The parameters of the anomaly detection model are adjusted based on the first classification loss, and the step of obtaining the first sample feature vector set corresponding to the first sample desensitized log set is performed again until the first training stop condition is reached, so as to obtain the first anomaly desensitized log recognition model.
3. The method according to claim 1, before obtaining the first sample feature vector set corresponding to the first sample anonymized log set, the method further comprises: Input a batch of first sample desensitized logs into a feature extraction model to obtain a first sample feature vector corresponding to each first sample desensitized log in the batch of first sample desensitized logs; Building a first log feature library based on each of the first sample feature vectors and a first sample label carried by each of the first sample feature vectors; The obtaining a first sample feature vector set corresponding to the first sample desensitized log set includes: Obtain a first sample feature vector set corresponding to the first sample desensitized log set from the first log feature library.
4. The method of claim 1, further comprising: Obtain a second sample feature vector set corresponding to the second sample desensitized log set; The data format and / or data semantics of each second sample desensitized log in the second sample desensitized log set are different from those of each first sample desensitized log in the first sample desensitized log set; The second sample desensitized log set includes normal desensitized logs; Based on at least one second sample feature vector in the second sample feature vector set and a second sample label carried by each second sample feature vector, the first abnormal desensitized log recognition model is fine-tuned to obtain a second abnormal desensitized log recognition model.
5. The method according to claim 4, wherein the first abnormal desensitization log recognition model is fine-tuned based on at least one second sample feature vector in the second sample feature vector set and a second sample label carried by each second sample feature vector to obtain a second abnormal desensitization log recognition model, comprising: Inputting the at least one second sample feature vector into the first abnormal desensitization log recognition model to obtain a second prediction label corresponding to each second sample feature vector; Determine a second classification loss of the first abnormal desensitized log recognition model according to each of the second prediction labels and each of the second sample labels; Adjust the parameters of the first abnormal desensitized log recognition model based on the second classification loss, and execute the step of obtaining the second sample feature vector set corresponding to the second sample desensitized log set again until the second training stop condition is reached to obtain the second abnormal desensitized log recognition model.
6. A method for identifying abnormal desensitized logs, the method comprising: Obtain a first target feature vector set corresponding to a first target desensitized log set to be identified; Inputting at least one first target feature vector in the first target feature vector set into the first abnormal desensitization log recognition model to obtain a first identification label corresponding to each first target desensitization log; Determine, according to each of the first identification tags, abnormal desensitized logs in the first target desensitized log set; The first abnormal desensitized log recognition model is trained by using the training method of the abnormal desensitized log recognition model described in any one of claims 1 to 5.
7. The method according to claim 6, after determining the abnormal desensitized logs in the first target desensitized log set according to each of the first identification tags, the method further comprises: Generate abnormal log warning prompt information for prompting optimization of the abnormal desensitized log.
8. A training device for an abnormal desensitization log recognition model, the device comprising: A first acquisition module, used to acquire a first sample feature vector set corresponding to a first sample desensitized log set; The first sample desensitized log set includes normal desensitized logs; A training module is used to train an anomaly detection model based on at least one first sample feature vector in the first sample feature vector set and a first sample label carried by each first sample feature vector to obtain a first anomaly desensitized log recognition model.
9. An abnormal desensitized log identification device, the device comprising: A second acquisition module is used to acquire a first target feature vector set corresponding to a first target desensitized log set to be identified; an identification module, used for inputting at least one first target feature vector in the first target feature vector set into the first abnormal desensitization log identification model to obtain a first identification label corresponding to each first target desensitization log; A determination module, used to determine the abnormal desensitized log in the first target desensitized log set according to each of the first identification tags; The first abnormal desensitized log recognition model is trained by using the training method of the abnormal desensitized log recognition model described in any one of claims 1 to 5.
10. An electronic device comprising: Processor and memory; The memory stores a computer program, and when the processor executes the computer program, the method steps of any one of claims 1-5 or 6-7 are implemented.
11. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps as claimed in any one of claims 1 to 5 or 6 to 7.
12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 or 6 to 7 are implemented.