A computer security risk management method
By building an annotation model and encryption protection mechanism for sensitive words of log files, sensitive data protection problems in computer system logs are solved, automatic identification and encryption protection is realized, and efficiency and security are improved.
Patent Information
- Application Number
- CN202210173901.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively protect sensitive fields and privacy data in computer system logs, resulting in illegal users who may use this information to cause losses.
By constructing a labeling model for sensitive words in the log file, the sensitive words in the log file are automatically marked, and the target value score of sensitive words is calculated based on the hierarchical division and frequency of occurrence. When the score is greater than the threshold, the log file paragraph is encrypted and sent to the administrator client through the three-level encryption and key chain mechanism, and the sensitive words are replaced after the administrator confirms.
It realizes automatic identification and encryption protection of sensitive data in computer system logs, reduces manual participation, improves efficiency, and enhances data security through three-level encryption and key chain mechanisms, reducing the risk of sensitive data leakage.
Smart Images

Figure CN114547654B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent decision-making, and particularly relates to a computer security risk management method. Background Art
[0002] Computer system logs. In modern society, in order to maintain the operation status of its own system resources, a computer system generally has a corresponding log recording system for date and timestamp information of daily events or misoperation alarms. A so-called log refers to an ordered set in time of certain operations of a specified object of the system and their operation results. Each log file consists of log records, and each log record describes a single system event. Usually, a system log is a text file that can be directly read by users, which contains a timestamp and an information or other information specific to the subsystem.
[0003] Log files record necessary and valuable information for activities related to IT resources such as servers, workstations, firewalls, and application software, which is very important for system monitoring, querying, reporting, and security auditing. The records in log files can be used for the following purposes: monitoring system resources; auditing user behavior; alerting for suspicious behavior; determining the scope of intrusion behavior; providing help for system recovery.
[0004] However, sensitive fields or privacy data will inevitably appear in log files, such as names, ID numbers, addresses, phone numbers, bank accounts, email addresses, passwords, medical information, educational backgrounds, etc. If these sensitive fields cannot be effectively protected and are used by illegal users, it will cause great losses to users and pose a relatively high risk. Summary of the Invention
[0005] The purpose of the present invention is to provide a computer security risk management method for solving or improving the above problems in view of the above deficiencies in the prior art.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A computer security risk management method, which includes the following steps:
[0008] S1. Obtain multiple log files of a computer, preprocess the log files, and build a labeling model for sensitive words in the log files based on the preprocessed log files;
[0009] S2. Input the log files to be labeled into the labeling model, output multiple labeled log file paragraphs, and obtain a target set of sensitive words;
[0010] S3. Identify the sensitive words in the target set of sensitive words and classify the sensitive words by level;
[0011] S4. Assign weights to different sensitive words according to the divided levels and the frequencies of the sensitive words, and calculate the target value scores of the sensitive words;
[0012] S5. If the calculated target value score of the sensitive word is greater than the threshold, encrypt the corresponding log file paragraph to form an encrypted package, and at the same time form a decryption key package for decrypting the encrypted package, and send the encrypted package and the decryption key package to the administrator client;
[0013] S6. The administrator client decrypts the encrypted package using the decryption key package through the encrypted communication protocol to obtain the log file paragraph corresponding to the sensitive word;
[0014] S7. After the administrator confirms the sensitive words in the log file paragraph, replace the original sensitive words with non-sensitive words.
[0015] A further technical solution is that in step S1, multiple log files of the computer are obtained, and the log files are preprocessed to obtain target log files, including:
[0016] S1.1. Define multiple sensitive words according to the user's personalized needs;
[0017] S1.2. According to the start position and end position of each paragraph in the log file, split the log file into multiple log paragraphs, traverse the log paragraphs, filter the log fields with sensitive words, and use the log fields with sensitive words as the target logs; at the same time, use the supervised learning method for text annotation;
[0018] S1.3. Convert each log file in the target log into a preset file format, and according to the time series, use the log paragraphs with sensitive words as the input data x i , use the log paragraphs after text annotation as the output data y i , where i is the corresponding time series moment;
[0019] S1.4. Merge the formatted log files to obtain the labeled model training dataset Q = {(x 1 , y 1 ), (x 2 , y 2 ), (x 3 , y 3 )... (x i , y i )};
[0020] S1.5. Based on the supervised learning of neural networks, construct a labeling model T;
[0021] T = P(Y 1 , Y 2, Y 3 …Y i |X 1 , X 2 , X 3 …X i )
[0022] Among them, X i is all possible log paragraphs, that is, the new input paragraph sequence; Y i is the possible tags of all log paragraphs, that is, the output tag sequence corresponding to the new input sequence; P is the probability of log paragraph tags.
[0023] A further technical solution is that in step S2, multiple annotated log file paragraphs are integrated according to the time series to obtain the target set of sensitive words.
[0024] A further technical solution is that in step S3, the sensitive words in the target set of sensitive words are identified, and the sensitive words are classified by level, including:
[0025] According to the type or attribute of the sensitive words, the sensitive words are divided into four levels: the first level, the second level, the third level, and the fourth level.
[0026] A further technical solution is that in step S4, according to the divided levels and the frequency of occurrence of sensitive words, different sensitive words are weighted, and the target value score of sensitive words is calculated, including:
[0027] F = N 1 *M 1 +N 2 *M 2 +N 3 *M 3 +N 4 *M 4 +ΔD
[0028] F1 = 2TP / (2TP + FP)
[0029] Among them, F is the target value score of sensitive words, N 1 , N 2 , N 3 and N 4 are the occurrence times of sensitive words in the first level, the second level, the third level, and the fourth level respectively, M 1 , M 2 , M 3 and M 4 are the weights of sensitive words in the first level, the second level, the third level, and the fourth level respectively; F1 is the prediction accuracy of the target value score, TP is the number of correct predictions, and FP is the number of wrong predictions;
[0030] ΔD is the correction value of the target value score, that is, the difference between the calculated value and the actual value of the target value score of sensitive words. When the difference is not within the preset difference range, the counter increments the FP value by 1; when the difference is within the preset difference range, the counter increments the TP value by 1.
[0031] A further technical solution is that if the calculated target value score of sensitive words in step S5 is greater than the threshold, the corresponding log file paragraph is encrypted to form an encrypted package, and at the same time, a decryption key package for decrypting the encrypted package is formed, and the encrypted package and the decryption key package are sent to the administrator client, including:
[0032] S5.1. Randomly disassemble the corresponding log file into n sub-transmission data packets, and randomly divide the n sub-transmission data packets into three subset data packets;
[0033] S5.2. Classify the three subset data packets and name them the first-level data packet, the second-level data packet, and the third-level data packet;
[0034] S5.3. Send a request to the server to allocate three encrypted sub-servers corresponding to the three subset data packets, and each sub-server corresponds to a subset data packet, including the first-level server, the second-level server, and the third-level server;
[0035] S5.4. The first-level server generates a first key for decrypting the first subset data packet;
[0036] The second-level server generates a second key for decrypting the second subset data packet;
[0037] The third-level server generates a third key for decrypting the third subset data packet;
[0038] S5.5. Store the three keys in different IP addresses, and at the same time perform three-key level chaining, and transmit the three keys to the administrator client respectively using three different custom communication protocols.
[0039] A further technical solution is that each subset data packet does not contain a complete log paragraph. Each complete log paragraph is split, and the sensitive words are distributed in at least two subset data packets.
[0040] A further technical solution is that step S5 performs three-key level chaining, including:
[0041] When the first subset data packet is not decrypted, the second subset data packet is in a locked state, and the second key cannot decrypt the second subset data packet; when and only when the first subset data packet is decrypted by the first key, the second subset data packet is unlocked, and the second key is used to decrypt the second subset data;
[0042] When the second subset of data packets is not decrypted, the third subset of data packets is in a locked state, and the third key cannot decrypt the third subset of data packets; only when the second subset of data packets is decrypted by the second key, the third subset of data packets is unlocked, and the third key is used to decrypt the third subset of data.
[0043] The computer security risk management method provided by the present invention has the following beneficial effects:
[0044] The present invention realizes the automatic annotation of log files by constructing an annotation model for sensitive words in log files, reduces manual participation, and improves the efficiency of annotating sensitive words in log files.
[0045] The present invention classifies sensitive words in log files, calculates the target value score of sensitive words according to the classification, and when the score value is greater than a preset value, the log file paragraph where the corresponding sensitive word is located is sent to the administrator client.
[0046] In order to prevent sensitive words from being stolen by others during transmission, the present invention encrypts the corresponding log file paragraph to form an encrypted package, and at the same time forms a decryption key package for decrypting the encrypted package, and sends it to the administrator client to reduce the leakage of encrypted sensitive words.
[0047] The present invention adopts three-level encryption and interlocks the three-level encryption to increase the encryption level.
[0048] After the administrator confirms the sensitive words, the present invention replaces the sensitive words in the log files in the computer with non-sensitive words to realize the management of computer security risks. Brief Description of the Drawings
[0049] Figure 1 It is a flowchart of the computer security risk management method. Detailed Embodiments
[0050] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0051] According to an embodiment of the present application, referring to Figure 1 , the computer security risk management method of this solution includes:
[0052] Step S1: Obtain multiple log files of the computer, preprocess the log files, and build an annotation model for sensitive words in the log files, which specifically includes:
[0053] Step S1.1: Define multiple sensitive words according to the user's personalized needs, such as name, ID number, address, phone number, bank account number, email, password, medical information, educational background, or other user-defined words;
[0054] Step S1.2: Split the log files into multiple log paragraphs according to the start position and end position of each paragraph in the log files. Traverse the log paragraphs, filter out the log fields with sensitive words, and use the log fields with sensitive words as the target logs; at the same time, use the supervised learning method for text annotation.
[0055] In this embodiment, the log files are split into paragraphs, and each paragraph is a complete objective narrative log, which is convenient for later encryption and splitting of the logs.
[0056] This embodiment uses model construction to automatically annotate newly input log files or log paragraphs through the model, so as to improve the recognition and annotation efficiency of sensitive words;
[0057] Step S1.3: Convert each log file in the target logs into a preset file format, and use the log paragraphs with sensitive words as the input data x i , use the log paragraphs after text annotation as the output data y i , where i is the corresponding time series moment;
[0058] Step S1.4: Merge the formatted log files to obtain the annotation model training dataset Q = {(x 1 , y 1 ), (x 2 , y 2 ), (x 3 , y 3 )…(x i , y i )};
[0059] Step S1.5: Build an annotation model T based on the supervised learning of neural networks;
[0060] T = P(Y 1 , Y 2 , Y 3 …Y i |X 1 , X 2 , X 3 …X i )
[0061] Among them, X i represents all possible log paragraphs, that is, the new input paragraph sequence; Y i represents all possible tags of the log paragraphs, that is, the output tag sequence corresponding to the new input sequence; P is the probability of the log paragraph tags.
[0062] Step S2: Input the log file to be annotated into the annotation model, output multiple annotated log file paragraphs, and obtain the target sensitive word set; in this embodiment, based on this annotation model T, a new log paragraph can be used as the input and brought into the model to output the log paragraph with sensitive words annotated.
[0063] Among them, multiple annotated log file paragraphs are integrated according to the time series to obtain the target sensitive word set.
[0064] Step S3: Identify the sensitive words in the target sensitive word set, and classify the sensitive words. This classification can be based on the user's personalized definition.
[0065] In this embodiment, according to the types or attributes of the sensitive words, the sensitive words are divided into four levels: the first level, the second level, the third level, and the fourth level. The higher the level, the higher the degree of sensitivity it represents.
[0066] Step S4: According to the classified levels and the frequencies of the sensitive words appearing, assign weights to different sensitive words and calculate the target value scores of the sensitive words, which specifically includes:
[0067] F = N 1 *M 1 +N 2 *M 2 +N 3 *M 3 +N 4 *M 4 +ΔD
[0068] F1 = 2TP / (2TP + FP)
[0069] Among them, F is the target value score of the sensitive word, N 1 , N 2 , N 3 and N 4 are the occurrence times of the sensitive words in the first level, the second level, the third level, and the fourth level respectively, M 1 , M 2 , M 3 and M 4 are the weights of the sensitive words in the first level, the second level, the third level, and the fourth level respectively; F1 is the prediction accuracy of the target value score, TP is the number of correct predictions, and FP is the number of wrong predictions;
[0070] ΔD is the correction value of the target value score, that is, the difference between the calculated value and the actual value of the target value score of sensitive words. When the difference is not within the preset difference range, the counter increments the FP value by 1; when the difference is within the preset difference range, the counter increments the TP value by 1.
[0071] In this embodiment, F1 is the prediction accuracy rate of the target value score. When the accuracy rate is greater than the preset value, it proves the correctness of the calculation of the target value score F of sensitive words.
[0072] The preset value for calculating the prediction accuracy rate of the F1 target value score in this embodiment is 96.56%. When the value of F1 is greater than 96.56%, the calculation of the target value score F of sensitive words is representative.
[0073] Step S5: If the calculated target value score of sensitive words is greater than the threshold, encrypt the corresponding log file paragraph to form an encrypted package, and at the same time form a decryption key package for decrypting the encrypted package, and send the encrypted package and the decryption key package to the administrator client. Specifically, it includes:
[0074] Step S5.1: Randomly disassemble the corresponding log file into n sub-transmission data packets, and randomly divide the n sub-transmission data packets into three subset data packets;
[0075] Step S5.2: Perform level division on the three subset data packets and name them the first-level data packet, the second-level data packet, and the third-level data packet;
[0076] Step S5.3: Send a request to the server to allocate three encrypted sub-servers corresponding to the three subset data packets, and each sub-server corresponds to a subset data packet, including the first-level server, the second-level server, and the third-level server;
[0077] Step S5.4: The first-level server generates a first key for decrypting the first subset data packet;
[0078] The second-level server generates a second key for decrypting the second subset data packet;
[0079] The third-level server generates a third key for decrypting the third subset data packet;
[0080] Step S5.5: Store the three keys in different IP addresses, and at the same time perform three-key level chaining, and transmit the three keys to the administrator client respectively using three different custom communication protocols.
[0081] A further technical solution of this embodiment is that each subset data packet does not contain a complete log paragraph. The logs of each complete paragraph are split, and sensitive words are distributed in at least two subset data packets. Only when all subset data packets are decrypted can a complete log paragraph with sensitive words be formed, so as to protect the implicit information of the computer to the greatest extent and reduce the risk of implicit information leakage of the computer.
[0082] Step S5 performs three-level key chaining, including:
[0083] When the first subset data packet is not decrypted, the second subset data packet is in a locked state, and the second key cannot decrypt the second subset data packet; when and only when the first subset data packet is decrypted by the first key, the second subset data packet is unlocked, and the second key is used to decrypt the second subset data.
[0084] When the second subset data packet is not decrypted, the third subset data packet is in a locked state, and the third key cannot decrypt the third subset data packet; when and only when the second subset data packet is decrypted by the second key, the third subset data packet is unlocked, and the third key is used to decrypt the third subset data.
[0085] That is, this embodiment chains the three subset data. Even if the key of one subset data packet is stolen, the remaining two subset data packets cannot be unlocked. As can be seen from the above, one subset data packet only contains partial log paragraphs and cannot form complete sensitive words. Therefore, in this case, the implicit information of the user cannot be stolen either.
[0086] Step S6: The administrator client decrypts the encrypted packet using the decryption key packet through the encrypted communication protocol to obtain the log file paragraph corresponding to the sensitive word;
[0087] Specifically:
[0088] The administrator or user obtains three keys at different addresses on three different servers according to three different communication protocols;
[0089] Use the first key to decrypt the first subset data packet. At this time, the second subset data packet is unlocked;
[0090] Use the second key to decrypt the second subset data packet. At this time, the third subset data packet is unlocked;
[0091] Use the third key to decrypt the third subset data packet. At this time, all three subset data packets are unlocked;
[0092] After all three subset data packets are unlocked, they are combined according to the time sequence to form a complete log file with sensitive words.
[0093] Step S7: After the administrator confirms the sensitive words in the log file paragraph, replace the original sensitive words with non-sensitive words, that is, accurately locate the log file on the computer and modify it.
[0094] Although the specific embodiments of the invention have been described in detail with reference to the accompanying drawings, it should not be construed as a limitation on the protection scope of this patent. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of this patent.
Claims
1. A computer security risk management method, characterized in that, it includes the following steps: S1. Obtain multiple log files of the computer, preprocess the log files, and build an annotation model for sensitive words in the log files based on the preprocessed log files; S2. Input the log files to be annotated into the annotation model, output multiple annotated log file paragraphs, and obtain a set of target sensitive words; S3. Identify the sensitive words in the set of target sensitive words and classify the sensitive words by level; S4. Assign weights to different sensitive words according to the classified levels and the frequencies of the sensitive words, and calculate the target value scores of the sensitive words; S5. If the calculated target value score of the sensitive word is greater than the threshold, encrypt the corresponding log file paragraph to form an encrypted package, and at the same time form a decryption key package for decrypting the encrypted package, and send the encrypted package and the decryption key package to the administrator client; S6. The administrator client decrypts the encrypted package using the decryption key package through an encrypted communication protocol to obtain the log file paragraph corresponding to the sensitive word; S7. After the administrator confirms the sensitive words in the log file paragraph, replace the original sensitive words with non-sensitive words; In S5, if the calculated target value score of the sensitive word is greater than the threshold, encrypt the corresponding log file paragraph to form an encrypted package, and at the same time form a decryption key package for decrypting the encrypted package, and send the encrypted package and the decryption key package to the administrator client, including: S5.
1. Randomly disassemble the corresponding log file paragraph into n sub-transmission data packets, and randomly divide the n sub-transmission data packets into three subset data packets; S5.
2. Classify the three subset data packets and name them the first-level data packet, the second-level data packet, and the third-level data packet; S5.
3. Send a request to the server to allocate three encrypted sub-servers corresponding to the three subset data packets, and each sub-server corresponds to a subset data packet, including the first-level server, the second-level server, and the third-level server; S5.
4. The first-level server generates a first key for decrypting the first subset data packet; The second-level server generates a second key for decrypting the second subset data packet; The third-level server generates a third key for decrypting the third subset data packet; S5.
5. Store the three keys in different IP addresses, and at the same time perform three-level key chaining, and transmit the three keys to the administrator client respectively using three different custom communication protocols.
2. The computer security risk management method according to claim 1, characterized in that: In S2, the multiple annotated log file paragraphs are integrated according to the time series to obtain a set of target sensitive words.
3. The computer security risk management method according to claim 1, characterized in that, In S3, identifying the sensitive words in the set of target sensitive words and classifying the sensitive words by level includes: Classify the sensitive words into four levels, the first level, the second level, the third level, and the fourth level according to the types or attributes of the sensitive words.
4. The computer security risk management method according to claim 1, It is characterized in that in S4, according to the divided levels and the occurrence frequencies of sensitive words, weights are assigned to different sensitive words, and the target value scores of sensitive words are calculated, including: F = N 1 *M 1 +N 2 *M 2 +N 3 *M 3 +N 4 *M 4 +ΔD F1 = 2TP / (2TP + FP) Among them, F is the target value score of sensitive words, N 1 , N 2 , N 3 and N 4 are the occurrence times of sensitive words in the first level, second level, third level, and fourth level respectively, M 1 , M 2 , M 3 and M 4 are the weights of sensitive words in the first level, second level, third level, and fourth level respectively; F1 is the prediction accuracy of the target value score, TP is the number of correct predictions, and FP is the number of incorrect predictions; ΔD is the target value score correction value, that is, the difference between the calculated value and the actual value of the target value score of the sensitive word. When the difference is not within the preset difference range, the counter increments the FP value by 1; when the difference is within the preset difference range, the counter increments the TP value by 1.
5. The computer security risk management method according to claim 1 It is characterized in that each subset data packet does not contain a complete log paragraph. The logs of each complete paragraph are split, and the sensitive words are distributed in at least two subset data packets.
6. The computer security risk management method according to claim 1 It is characterized in that in S5, three key level linkages are performed, including: when the first subset data packet is not decrypted, the second subset data packet is in a locked state, and the second key cannot decrypt the second subset data packet; when and only when the first subset data packet is decrypted by the first key, the second subset data packet is unlocked, and the second key is used to decrypt the second subset data; when the second subset data packet is not decrypted, the third subset data packet is in a locked state, and the third key cannot decrypt the third subset data packet; when and only when the second subset data packet is decrypted by the second key, the third subset data packet is unlocked, and the third key is used to decrypt the third subset data.
Citation Information
Patent Citations
Method and system for encrypting cloud server files and cloud server
CN103442061A
Log file sensitive field monitoring method and device and computer equipment
CN110377479A
Log file encryption method and device, storage medium and electronic equipment
CN112788012A