Log desensitization method and device, equipment and medium
By setting the preset intercepted character set and the key name set to be desensitized, the log data is divided and desensitized based on the preset desensitization rules, the problem that regular expressions in the prior art cannot adapt to the dynamic log format, and efficient and accurate log desensitization is achieved.
Patent Information
- Application Number
- CN202411880684.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-23
AI Technical Summary
When using regular expressions to desensitize logs in the prior art, it is impossible to flexibly adapt to the dynamically changing log format, resulting in the inability to accurately match and desensitize related information.
By setting the preset intercepted character set and the preset key name set to be desensitized, the log data is divided to generate the first string data containing sensitive characters and the second string data containing no sensitive characters. Then, based on the preset desensitization character positioning rules and desensitization rules, desensitization characters are desensitized, and a target log is finally generated.
It improves the efficiency and accuracy of log desensitization, enhances the flexibility of desensitization methods, enables it to adapt to dynamic and variable log formats, avoids the omission of sensitive information, and solves the problem of low regular expression matching efficiency and desensitization accuracy.
Smart Images

Figure CN120030583A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular, to a log desensitization method, device, equipment and medium. Background Art
[0002] During the operation of a business system, there may be business processing processes involving personal sensitive information such as names and mobile phone numbers. In the logs of a transaction system, there are exact personal sensitives such as names, mobile phone numbers, and bank card numbers. Once these information are leaked, it may bring inestimable losses to users. In the era when personal information is gradually emphasized, log data desensitization is the focus of enterprise attention.
[0003] In related technologies, when performing log desensitization, regular expressions are usually used for log desensitization. However, regular expressions are static and can only process logs with fixed formats and structures. When facing logs with unfixed and dynamically changing information, regular expressions cannot flexibly adapt to these dynamically changing formats, resulting in the inability to accurately match and desensitize relevant information. Summary of the Invention
[0004] The present invention provides a log desensitization method, device, electronic equipment and medium to solve the technical problem that in the process of using regular expressions for log desensitization, it is impossible to flexibly adapt to dynamically changing formats, resulting in the inability to accurately match and desensitize relevant information.
[0005] In a first aspect, a log desensitization method is provided, including:
[0006] Obtain log data to be desensitized;
[0007] Based on a preset character set for truncation and a preset set of keys to be desensitized, divide the log data to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, where the string data includes multiple characters between two truncation characters;
[0008] Based on a preset desensitized character positioning rule, determine at least one character to be desensitized in each first string data, and desensitize each character to be desensitized based on a preset desensitization rule;
[0009] Generate target logs based on the at least one desensitized first string data, the at least one second string data, and the truncation characters between adjacent two string data.
[0010] In a second aspect, a log desensitization device is provided, including:
[0011] An obtaining module, configured to obtain log data to be desensitized;
[0012] A first generating module, configured to divide the log data based on a preset intercepted character set and a preset key name set to be desensitized, and generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, wherein the string data includes a plurality of characters between two intercepted characters;
[0013] A determination module, configured to determine at least one to-be-desensitized character in each first character string data based on a preset desensitization character location rule, and desensitize each to-be-desensitized character based on the preset desensitization rule;
[0014] The second generating module is used to generate a target log based on the desensitized at least one first string data, at least one second string data and intercepted characters between two adjacent string data.
[0015] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned log desensitization method when executing the computer program.
[0016] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned log desensitization method are implemented.
[0017] In the scheme implemented by the above-mentioned log desensitization method, device, electronic device and storage medium, the log content is quickly segmented by the set intercepted character set, and the pre-set key name set to be desensitized is used to quickly identify whether each string needs to be desensitized, so as to divide the log data into the first string data that needs to be desensitized and the second string data that does not need to be desensitized. Afterwards, according to the desensitization position matching rule, the characters to be desensitized in different string lengths are located, and finally the preset desensitization rules are adopted to perform targeted desensitization processing on the characters corresponding to each position to be desensitized, and the desensitized target log is obtained. The present application matches sensitive information, calculates the position of sensitive information and desensitizes sensitive information by presetting log segmentation rules, sensitive data location rules and desensitization rules, which improves the desensitization efficiency and the flexibility of log desensitization, so that the desensitization method can adapt to the dynamic and changeable log format and avoid the omission of sensitive information. It solves the problem of repeated regular operations, low matching efficiency and desensitization accuracy in the process of using regular expressions to match and identify desensitized content in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0019] Figure 1 It is a flow chart of a log desensitization method in one embodiment of the present invention;
[0020] Figure 2 It is a schematic flow chart of a log desensitization method according to an embodiment of the present invention;
[0021] Figure 3 It is a structural schematic diagram of a log desensitization device in one embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be understood that the drawings in the present invention are only for the purpose of illustration and description and are not used to limit the scope of protection of the present invention.
[0023] In addition, it should be understood that the schematic drawings are not drawn to scale. The flow chart used in the present invention shows the operations implemented according to some embodiments of the present invention. It should be understood that the operations of the flow chart can be implemented out of order, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flow chart under the guidance of the content of the present invention, and can also remove one or more operations from the flow chart.
[0024] In addition, the embodiments described in the present invention are only some embodiments of the present invention, rather than all embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.
[0025] It should be noted that the term "comprising" will be used in the embodiments of the present invention to indicate the existence of the features declared thereafter, but does not exclude the addition of other features. It should also be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, it should also be noted that the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0026] The case is described in detail below with reference to the accompanying drawings in the specification.
[0027] In the embodiments of this specification, log desensitization refers to the deformation of certain sensitive information in the log through desensitization rules to achieve reliable protection of sensitive privacy data. In the case of customer security data or some commercial sensitive data, the real data is transformed and provided for use without violating the system rules. For example, personal information such as ID card number, mobile phone number, card number, customer number, etc. need to be desensitized.
[0028] A large amount of sensitive data is retained in the log systems corresponding to different application systems and devices. For example, the ID number, user name, user ID, etc. carried in the user's session, such as: bank card number, client IP address, server IP, etc. in the request structured query language (SQL) statement. When the log needs to be queried and analyzed, these sensitive data will be exposed to unauthorized users, resulting in information security risks. In related technologies, sensitive information in the log is usually matched by regular expressions, and then uniform desensitization processing is performed based on the matched strings.
[0029] However, different sensitive information is generally matched using different regular expressions. The logs in the application system may contain multiple sensitive information. In order to ensure that sensitive information in the logs can be desensitized, it is necessary to traverse the rules of the regular expression and match the string multiple times. This method is time-consuming and labor-intensive, has low performance efficiency, and there is a risk that some sensitive information may not be desensitized and encrypted due to omissions. In addition, regular expressions are static and can only process logs with fixed formats and structures. When faced with non-fixed and dynamically changing log information, regular expressions cannot flexibly adapt to these dynamically changing formats, resulting in the inability to accurately match and desensitize related information.
[0030] Based on the above problems, the embodiments of this specification provide a log desensitization method, device, equipment and medium to improve the efficiency of log desensitization while ensuring the accuracy and flexibility of log desensitization.
[0031] See also Figure 1 This embodiment of the present invention provides a log desensitization method, which specifically includes the following steps:
[0032] S10: Obtain the log data to be desensitized.
[0033] It is understandable that the execution subject of the present invention may be a log desensitization device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0034] Among them, log data involving sensitive information is obtained, some of the log data are sensitive numbers that need to be desensitized, and some of the log data are non-sensitive data that do not need to be desensitized.
[0035] In actual application scenarios, taking the financial field as an example, when the business system executes the reprint business, the generated log is likely to record some real user information, such as name, bank card number, etc. In order to ensure that user information is not leaked, the log needs to be desensitized. The server automatically collects log data before log output by deploying a log agent, where the collected log data can be a line of log or a log file.
[0036] S20: dividing the log data based on a preset intercepted character set and a preset key name set to be desensitized, and generating at least one first string data containing sensitive information and at least one second string data not containing sensitive information;
[0037] The character string data includes a plurality of character data between two intercepted characters.
[0038] In this step, the preset interception character set is a plurality of pre-set format special characters corresponding to the log data. The log data is traversed, and the preset interception character set is used to sequentially obtain a plurality of character data between two intercepted characters in the log data as the string data to be analyzed. Thereafter, each string data to be analyzed is parsed using the preset set of key names to be desensitized to obtain a string containing sensitive information, which is marked as the first string data; and a string not containing sensitive information, which is marked as the second string data. According to the above method, the obtained plurality of string data to be analyzed are parsed in turn to divide the log data into a plurality of first string data containing sensitive information and a plurality of second string data not containing sensitive information.
[0039] It is understandable that the number of at least one first string data obtained by segmenting the log data can be one or more. For example, if the log data contains only one field that needs to be desensitized, the number of the first string data is one; if the log data contains multiple fields that need to be desensitized, the number of the first string data is multiple. Similarly, the number of at least one second string data can be one or more.
[0040] Optionally, the preset interception character set includes, but is not limited to: paired tag start characters (such as '<[{(
"), paired tag end characters (such as '>]})
[0041] Furthermore, if the log data is in unstructured text format, such as "User John Doe loggedin with email johndoe@example.com and phone number 123-456-7890", the log data exists in free text format and does not contain special characters. In order to ensure the accuracy of sensitive data identification, natural language processing (NLP) technology can be used to perform word segmentation and entity recognition on the text to generate key=value pairs to convert unstructured log data into structured log data.
[0042] Through the above method, the log data is accurately divided into multiple strings using the preset interception characters, which improves the accuracy of subsequent field parsing. Then, the preset key name set to be desensitized is used to accurately identify the string containing sensitive information, which facilitates the subsequent special desensitization processing and improves the desensitization efficiency.
[0043] In an embodiment of the present application, a specific log data segmentation scheme is provided. In S20, that is, based on a preset character set for truncation and a preset set of keys to be desensitized, the log data is divided to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters. The specific steps include the following steps S21 - S23:
[0044] S21: Based on the preset character set for truncation, determine multiple starting characters and the corresponding ending characters for each starting character in the log data;
[0045] Among them, the starting character is the next character adjacent to the truncation character, and the ending character is the previous character adjacent to the truncation character;
[0046] S22: Generate multiple string data to be analyzed based on the multiple starting characters and multiple ending characters;
[0047] For steps S21 - S22, obtain the first character of the log data. Based on the preset character set for truncation, determine whether the first character is a truncation character. If the first character is a truncation character, use the next character adjacent to it as the starting character; if the first character is not a truncation character, use the first character as the starting character. Then, obtain the next character adjacent to the first character and determine whether it is a truncation character. In the above manner, until a truncation character is obtained, use the previous character adjacent to the truncation character as the ending character corresponding to the first starting character, and use the multiple characters between the first starting character and the ending character as a string data to be analyzed. Then, obtain the starting character of the next non - truncation character and its corresponding ending character, and generate the string data to be analyzed. In the above manner, traverse the log data to obtain multiple string data to be analyzed.
[0048] Exemplarily, the JSON object of the user's purchase of commodity information is: {"userName":"Mr. Wang","userPhone":15311111111,"goodsName":"Package 1"}. Based on the preset character set for truncation, determine the truncation character "{". Then the first starting character is "u", and the ending character corresponding to the first starting character is "e", and the string data to be analyzed is "userName". Then, obtain the starting character "Wang" of the next non - truncation character, and the ending character corresponding to this starting character is "Sheng". In the above manner, divide the log data into multiple string data to be analyzed: "userName", "Mr. Wang", "userPhone", "15311111111", "goodsName", "Package 1".
[0049] Through the above method, the starting position and the ending position are determined by using the preset interception character set, the character string to be analyzed is determined from the log data, and the valid data is extracted to improve the efficiency of log desensitization.
[0050] S23: Based on a preset set of key names to be desensitized, the plurality of character string data to be analyzed are divided into at least one first character string data and at least one second character string data.
[0051] In this step, in order to improve the efficiency of sensitive data determination, a set of key names to be desensitized corresponding to the sensitive data types is pre-set. After obtaining the key names to be desensitized contained in multiple string data to be analyzed, it can be determined that the next string data to be analyzed adjacent to the key names to be desensitized contains sensitive information. At this time, the next string data to be analyzed adjacent to the key names to be desensitized is marked as the first string data. Other string data do not contain sensitive information and are marked as the second string data.
[0052] In one embodiment of the present application, a specific sensitive data screening solution is provided. In S23, based on a preset set of key names to be desensitized, multiple character string data to be analyzed are divided into at least one first character string data and at least one second character string data, specifically including the following steps S231-S233:
[0053] S231: Obtain at least one target key name contained in the plurality of character string data to be analyzed;
[0054] S232: For any target key name, match the target key name with a preset set of key names to be desensitized. If the preset set of key names to be desensitized includes the target key name, obtain the next character string data to be analyzed adjacent to the target key name as the first character string data;
[0055] S233: Obtain other string data except the at least one first string data from the plurality of string data to be analyzed as second string data.
[0056] For steps S231-S233, traverse multiple string data to be analyzed, obtain at least one target key name contained therein, and compare each target key name with the preset key name set to be desensitized in turn. If the preset key name set to be desensitized contains the target key name, it means that the value corresponding to the target key name is sensitive information, and it can be determined that the next string data to be analyzed adjacent to the target key name needs to be desensitized. Therefore, the next string data to be analyzed adjacent to the target key name is obtained, and the string data is the value corresponding to the target key name, which is used as the first string data to be desensitized. Furthermore, if the target key name is not included in the key name set to be desensitized, it means that the value corresponding to the target key name is non-sensitive information, and there is no need to desensitize the string data adjacent to the target key name. According to the above method, after determining at least one first string data, it can be determined that the other string data in the multiple string data to be analyzed except at least one first string data is the second string data.
[0057] Exemplarily, the multiple string data to be analyzed are: "userName", "Mr. Wang", "userPhone", "15311111111", "goodsName", "Package 1". Based on the preset key name set to be desensitized, it is determined that the multiple string data to be analyzed include the key name set to be desensitized "userName" and "userPhone". At this time, the values corresponding to each key name set to be desensitized, "Mr. Wang" and "15311111111", are obtained, which are the first string data to be desensitized. The other string data "userName", "userPhone", "goodsName", and "Package 1" are marked as the second string data.
[0058] Through the above method, by identifying the key-value pairs (key=value) in the log content and combining the type of data to be desensitized, sensitive information can be quickly located and the efficiency of sensitive information matching can be improved.
[0059] S30: Determine at least one to-be-desensitized character in each first character string data based on a preset desensitization character location rule, and desensitize each to-be-desensitized character based on the preset desensitization rule;
[0060] In this step, after obtaining at least one first string data containing sensitive information, it needs to be desensitized. In the related art, a regular matching method is usually adopted to match from the head of the log. When a qualified string is matched, it will continue to match backward until the end of the log. However, this method will not only increase invalid matches, but may also match errors. For example, the string "11111111123" is a continuous string. It is not a mobile phone number, but it may be matched as a mobile phone number for desensitization. Based on the above problems, the present application proposes to pre-set desensitized data matching rules, wherein the preset desensitized character positioning rules include a mapping relationship between different string lengths and desensitized character coordinates, and obtain at least one desensitized character in each first string to be desensitized through coordinate matching operations. Thereafter, the pre-set desensitization rules are used to desensitize the selected characters to be desensitized in the first string data.
[0061] Through the above method, the coordinates of the desensitized characters can be flexibly adjusted according to different string lengths, which significantly improves the accuracy and consistency of desensitization, reduces misjudgments and errors, and improves the efficiency of desensitization processing.
[0062] In one embodiment of the present application, a specific desensitized character positioning scheme is provided. In S30, that is, based on a preset desensitized character positioning rule, at least one to-be-desensitized character in each first character string data is determined, and based on the preset desensitization rule, each to-be-desensitized character is desensitized, specifically including the following steps S31-S33:
[0063] S31: Obtain the string length of each first string data;
[0064] S32: Determine at least one character to be desensitized in each first character string data based on a preset desensitized character location rule and a character string length;
[0065] The preset desensitized character positioning rule includes the corresponding relationship between the character string length and the coordinates of the character to be desensitized;
[0066] S33: Desensitize each character to be desensitized based on a preset desensitization rule.
[0067] For steps S31-S33, for any first character string data, obtain its character string length, and based on the character string length, determine the coordinates of the character to be desensitized corresponding to the character string length in the preset desensitized character positioning rule, and then locate at least one character to be desensitized in the first character string data. Finally, desensitize each character to be desensitized based on the preset desensitization rule.
[0068] In one embodiment of the present application, a specific desensitized character positioning scheme is provided. In S32, that is, based on a preset desensitized character positioning rule and a character string length, at least one character to be desensitized in each first character string data is determined, specifically including the following steps S321-S323:
[0069] S321: For any first character string data, obtaining at least one target coordinate position corresponding to the character string length in a preset desensitized character location rule;
[0070] S322: Acquire multiple coordinate positions of multiple characters in the first character string data;
[0071] S323: Determine at least one character to be desensitized in the first character string data based on at least one target coordinate value and a plurality of coordinate positions.
[0072] For steps S321-S323, for any first character string data, at least one target coordinate value of the character to be desensitized corresponding to the character string length of the character string data is determined in the preset desensitized character location rule. Thereafter, the coordinate value corresponding to each character in the first character string data is obtained, and the at least one target coordinate value is matched with the multiple coordinate values to determine each character to be desensitized in the first character string data.
[0073] Optionally, preset desensitized character positioning rules are set in advance according to the business needs of the application system to obtain the coordinate position of the character to be desensitized in any string data through coordinate operation. The rule may contain multiple operation rules, separated by semicolons. Each rule contains 3 elements, starting coordinates, ending coordinates, and matching string length; the elements are separated by commas. When the coordinate value > 0, it means calculating from the starting position; when the coordinate value is less than 0, it means calculating from the ending position, and when the coordinate value = 0, it means calculating from the corresponding position. Exemplarily, the preset desensitized character positioning rules are {4, -1, 16; 1, 2, 4; 1, 0}. The above rules represent that the string length range is set to: L≥16; 16<L<4; L≤4. When the string length is greater than or equal to 16, the characters from the fifth to the second to last digit of the string data are desensitized. For example, if the string data is "A123456789123456", it will be "A123************6" after desensitization. When the string length is between 4 and 16, the characters from the first to the third digits are desensitized. For example, if the string data is "A1234", it will be "***34" after desensitization. When the string length is less than or equal to 4, the content after the first digit is desensitized (for example, if the string is A12, it will be A** after desensitization).
[0074] Through the above method, the string length of the string data is realized, the position of the desensitized characters is dynamically adjusted, and the flexibility of character desensitization is improved.
[0075] In one embodiment of the present application, a specific desensitized character positioning scheme is provided. In S33, that is, based on a preset desensitization rule, each character to be desensitized is desensitized, specifically including the following steps S331 - S333:
[0076] S331: For any first string data, obtain the data type of the character to be desensitized;
[0077] S332: In the preset desensitization rule, determine the preset desensitization method corresponding to the data type;
[0078] S333: Use the preset desensitization method to desensitize each character to be desensitized.
[0079] For steps S331 - S333, desensitization methods for different data types are preset. For any first string, based on the data type of the character to be desensitized in the string data, determine the preset desensitization method in the preset desensitization rule, and then use the preset desensitization method to desensitize each character to be desensitized. For example, if the data to be desensitized is a phone number, the replace direct replacement method is used.
[0080] Through the above method, the most suitable desensitization method is selected for different sensitive data types to ensure that the desensitization effect meets the expectations.
[0081] S40: Generate a target log based on the at least one first string data after desensitization, the at least one second string data, and the intercept characters between adjacent two string data.
[0082] In this step, each first string data after desensitization and each second string data that does not need to be desensitized are recombined in the format of the original log, and the original intercept characters are added between the combined strings to form the target log after desensitization.
[0083] Through the above method, it is ensured that the generated target log retains the structure and integrity of the original log, while effectively protecting sensitive information, so that the log still has the value of being reused after desensitization.
[0084] In one embodiment of the present application, a specific desensitized character positioning scheme is provided. In S30, that is, based on a preset desensitized character positioning rule, determine at least one character to be desensitized in each first string data, and based on a preset desensitization rule, desensitize each character to be desensitized, specifically including the following steps S41 - S43:
[0085] S41: For any first string data, write the first string data into the first preset buffer area and desensitize the first string data in the first preset buffer area;
[0086] S42: writing the desensitized first character string data into a second preset buffer area, and writing the second character string data adjacent to the first character string data and the intercepted characters therebetween into the second preset buffer area in the order of the log data;
[0087] S43: Summarize the desensitized at least one first string data, at least one second string data, and intercepted characters between two adjacent string data contained in the second preset buffer area in the order of the log data to obtain a target log.
[0088] For steps S41-S44, a string desensitization buffer matchBuf (i.e., the first preset buffer area) is maintained, and the string between the intercepted characters is segmented and written to the buffer area for subsequent desensitization processing. At the same time, a string result buffer retBuf (i.e., the second preset buffer area) is maintained, and the content that has been desensitized or does not need to be desensitized is written to the buffer area. Finally, it is only necessary to output the content of the buffer area, so that the present application can use the cache method to perform log desensitization management. Specifically, in the log desensitization process, traverse from the first character, when the first string data is generated, the first string data is written to the first buffer area for desensitization processing, the first string data after desensitization is written to the second buffer area, the next intercepted character adjacent to the first string data is written to the second buffer area, and then the next second string data adjacent to the next intercepted character is written to the second buffer area, so that the string contained in the second buffer area is the original log data order. Traverse the log in the above manner until the end of the log, summarize and organize all string data in the second buffer area, and output the target log after desensitization.
[0089] Through the above method, by setting the desensitizing cache area and the result cache area, the log cannot be desensitized, and the relevance of the log data before and after desensitization is guaranteed, so that the log still has the value of reuse after desensitization.
[0090] In actual application scenarios, such as Figure 2As shown in the figure, it is a schematic block diagram of the log desensitization process. Specifically, during the log desensitization process, after obtaining an un-desensitized log content line, starting from the first character (cItem), traverse to determine whether the current character is a special character (i.e., a preset truncation character). If not, write the character cltem into the buffer matchBuf (i.e., the first preset buffer), and set the previously matched character as preChar. Then, determine whether there is a next character in the log data. If not, write the remaining content in the buffer matchBuf into the buffer retBuf (i.e., the second preset buffer). Further, if the current character is a special character, determine whether the current character is an end tag or a separator. If so, it means that the content before this special character is the string data to be analyzed. At this time, determine whether there is a corresponding value of matchKey in the buffer matchBuf. If there is, output the content of the buffer matchBuf as matchValue, desensitize matchValue according to matchKey, and write the desensitized value into the buffer retBuf; if there is no corresponding value of matchKey in the buffer matchBuf, output the content of the buffer to matchKey, check matchKey to determine whether it needs to be desensitized, and at the same time write the content in matchBuf into retBuf, clear the buffer matchBuf, wait for the next value to be desensitized, and write the special character into retBuf. Determine whether there is a next character. If so, return to the step of determining whether it is a special character; if there is no next character, the remaining content in the buffer matchBuf can be written into the buffer retBuf, and the entire traversal process ends. At this time, the content in the buffer retBuf is the desensitized log content in the same order as the original log data. By outputting the content of retBuf, the desensitized target log can be obtained. Optionally, after obtaining the log data, the string to be desensitized can be converted into an array of characters to be processed, and each character (cItem) is obtained for judgment each time. It should be noted that for logs in XML format, the log format is usually <cifname> Customer Name< / cifname> , in order to improve the matching efficiency of sensitive fields, the acquisition of matchValue needs to be judged at the node "whether it is possible to skip the matchBuf judgment". If it is < / tag, it can also represent the end character, so that the value corresponding to cifName can be obtained normally. At the same time, the > after < / is considered a tag that can be skipped, and no desensitization judgment is performed on it, so that the content in xml tag format can be parsed normally.
[0091] It can be seen that in the above scheme, the log content is quickly segmented by the set intercepted character set, and the pre-set key name set to be desensitized is used to quickly identify whether each string needs to be desensitized, so as to realize the division of log data into the first string data that needs to be desensitized and the second string data that does not need to be desensitized. Afterwards, according to the desensitization position matching rule, the characters to be desensitized in different string lengths are located, and finally the preset desensitization rules are adopted to perform targeted desensitization processing on the characters corresponding to each position to be desensitized, and the desensitized target log is obtained. The present application matches sensitive information, calculates the position of sensitive information, and desensitizes sensitive information by presetting log segmentation rules, sensitive data location rules, and desensitization rules. While improving the desensitization efficiency, it also improves the flexibility of log desensitization, so that the desensitization method can adapt to the dynamic and changeable log format, avoid the omission of sensitive information, and solve the problem of repeated regular operations, low matching efficiency and desensitization accuracy in the process of using regular expressions to match and identify desensitized content in the related technology.
[0092] In one embodiment, a log desensitization device is provided, and the log desensitization device corresponds one-to-one to the log desensitization method in the above embodiment. Figure 3 As shown, the log desensitization device includes: an acquisition module 101, a first generation module 102, a determination module 103, and a second generation module 104. The functional modules are described in detail as follows:
[0093] The acquisition module 101 is used to acquire the log data to be desensitized;
[0094] A first generating module 102, configured to divide the log data based on a preset intercepted character set and a preset to-be-desensitized key name set, and generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, wherein the string data includes a plurality of characters between two intercepted characters;
[0095] The determination module 103 is used to determine at least one to-be-desensitized character in each first character string data based on a preset desensitization character location rule, and desensitize each to-be-desensitized character based on the preset desensitization rule;
[0096] The second generating module 104 is used to generate a target log based on the desensitized at least one first string data, at least one second string data and intercepted characters between two adjacent string data.
[0097] In one embodiment, the first generating module 102 specifically includes:
[0098] A first determining unit is used to determine, based on a preset intercepted character set, a plurality of starting characters and an ending character corresponding to each starting character in the log data, wherein the starting character is the next character adjacent to the intercepted character, and the ending character is the previous character adjacent to the intercepted character;
[0099] A first generating unit, used for generating a plurality of character string data to be analyzed based on a plurality of starting characters and a plurality of ending characters;
[0100] The second generating unit is used to divide the multiple character string data to be analyzed into at least one first character string data and at least one second character string data based on a preset set of key names to be desensitized.
[0101] In one embodiment, the second generating unit is specifically configured to:
[0102] Obtain at least one target key name contained in multiple string data to be analyzed;
[0103] For any target key name, match the target key name with a preset set of key names to be desensitized. If the preset set of key names to be desensitized contains the target key name, obtain the next character string data to be analyzed adjacent to the target key name as the first character string data;
[0104] The other character string data except the at least one first character string data among the plurality of character string data to be analyzed are obtained as the second character string data.
[0105] In one embodiment, the determination module 103 specifically includes:
[0106] An acquiring unit, used for acquiring the character string length of each first character string data;
[0107] A second determination unit is used to determine at least one to-be-desensitized character in each first character string data based on a preset desensitized character location rule and a character string length, wherein the preset desensitized character location rule includes a correspondence between a character string length and a coordinate of a to-be-desensitized character;
[0108] The desensitization unit is used to desensitize each character to be desensitized based on a preset desensitization rule.
[0109] In one embodiment, the second determining unit is specifically configured to:
[0110] For any first character string data, obtaining at least one target coordinate position corresponding to the character string length in a preset desensitized character positioning rule;
[0111] Obtaining multiple coordinate positions of multiple characters in the first character string data;
[0112] At least one to-be-desensitized character is determined in the first character string data based on at least one target coordinate position and a plurality of coordinate positions.
[0113] In one embodiment, the desensitization unit is specifically used for:
[0114] For any first character string data, obtain the data type of the character to be desensitized;
[0115] In the preset desensitization rules, determine the preset desensitization method corresponding to the data type;
[0116] Use the preset desensitization method to desensitize each character to be desensitized.
[0117] In one embodiment, the second generating module 104 specifically includes:
[0118] A first writing unit, configured to write any first character string data into a first preset buffer area, and desensitize the first character string data in the first preset buffer area;
[0119] A second writing unit is used to write the desensitized first character string data into a second preset buffer area, and write the second character string data adjacent to the first character string data and the intercepted characters therebetween into the second preset buffer area in accordance with the order of the log data;
[0120] The third generating unit is used to summarize the desensitized at least one first string data, at least one second string data and the intercepted characters between two adjacent string data included in the second preset buffer area in the order of log data to obtain a target log.
[0121] The present invention provides a log desensitization device, which quickly segments the log content through a set interception character set, and uses a pre-set key name set to be desensitized to quickly identify whether each character string needs to be desensitized, so as to realize the division of log data into a first character string data that needs to be desensitized and a second character string data that does not need to be desensitized. Afterwards, according to the desensitization position matching rule, the characters to be desensitized in different character string lengths are located, and finally the preset desensitization rule is adopted to perform desensitization processing on the characters corresponding to each position to be desensitized in a targeted manner, and obtain the desensitized target log. The present application matches sensitive information, calculates the position of sensitive information, and desensitizes sensitive information by presetting log segmentation rules, sensitive data positioning rules, and desensitization rules, which improves the desensitization efficiency and the flexibility of log desensitization, so that the desensitization method can adapt to the dynamic and changeable log format, avoid the omission of sensitive information, and solve the problem of repeated regular operations, low matching efficiency and desensitization accuracy in the process of using regular expressions to match and identify desensitized content in the related art.
[0122] For the specific limitations of the log desensitization device, reference may be made to the limitations of the log desensitization method in the foregoing text, which will not be elaborated herein. Each module in the foregoing log desensitization device may be implemented in whole or in part by software, hardware, and their combination. Each of the foregoing modules may be embedded in or independent of a processor in an electronic device in the form of hardware, or may be stored in a memory in the electronic device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to each of the foregoing modules.
[0123] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0124] Obtain log data to be desensitized;
[0125] Based on a preset character set for interception and a preset set of keys to be desensitized, divide the log data to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, where the string data includes multiple characters between two intercepted characters;
[0126] Based on a preset desensitized character positioning rule, determine at least one character to be desensitized in each first string data, and desensitize each character to be desensitized based on a preset desensitization rule;
[0127] Generate a target log based on the at least one desensitized first string data, the at least one second string data, and the intercepted characters between two adjacent string data.
[0128] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0129] Obtain log data to be desensitized;
[0130] Based on a preset character set for interception and a preset set of keys to be desensitized, divide the log data to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, where the string data includes multiple characters between two intercepted characters;
[0131] Based on a preset desensitized character positioning rule, determine at least one character to be desensitized in each first string data, and desensitize each character to be desensitized based on a preset desensitization rule;
[0132] Generate a target log based on the at least one desensitized first string data, the at least one second string data, and the intercepted characters between two adjacent string data.
[0133] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or electronic device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0134] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0135] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0136] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A log desensitization method, characterized in that: include: Get the log data to be desensitized; Based on a preset intercepted character set and a preset to-be-desensitized key name set, the log data is divided to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, wherein the string data includes a plurality of characters between two intercepted characters; Based on a preset desensitization character location rule, determine at least one to-be-desensitized character in each first character string data, and desensitize each to-be-desensitized character based on the preset desensitization rule; A target log is generated based on the desensitized at least one first string data, the at least one second string data, and intercepted characters between two adjacent string data.
2. The method according to claim 1, characterized in that The step of dividing the log data based on the preset intercepted character set and the preset key name set to be desensitized to generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters specifically includes: Based on the preset interception character set, in the log data, a plurality of starting characters and an ending character corresponding to each starting character are determined, wherein the starting character is the next character adjacent to the interception character, and the ending character is the previous character adjacent to the interception character; Based on the multiple starting characters and the multiple ending characters, generate multiple character string data to be analyzed; Based on the preset key name set to be desensitized, the plurality of character string data to be analyzed are divided into the at least one first character string data and the at least one second character string data.
3. The method according to claim 2, characterized in that The step of dividing the plurality of character string data to be analyzed into the at least one first character string data and the at least one second character string data based on the preset to-be-desensitized key name set specifically includes: Obtain at least one target key name contained in the plurality of character string data to be analyzed; For any target key name, match the target key name with the preset key name set to be desensitized. If the preset key name set to be desensitized contains the target key name, obtain the next character string data to be analyzed adjacent to the target key name as the first character string data; The other character string data except the at least one first character string data among the plurality of character string data to be analyzed are obtained as the second character string data.
4. The method according to claim 1, characterized in that: The step of determining at least one to-be-desensitized character in each first character string data based on a preset desensitization character location rule, and desensitizing each to-be-desensitized character based on the preset desensitization rule specifically includes: Get the string length of each first string data; Based on the preset desensitized character positioning rule and the character string length, determining at least one character to be desensitized in each of the first character string data, wherein the preset desensitized character positioning rule includes a corresponding relationship between the character string length and the coordinates of the character to be desensitized; Based on the preset desensitization rule, each character to be desensitized is desensitized.
5. The method according to claim 4, characterized in that The step of determining at least one to-be-desensitized character in each first character string data based on the preset desensitized character location rule and the character string length specifically includes: For any first character string data, obtaining at least one target coordinate position corresponding to the character string length in the preset desensitized character positioning rule; Obtaining multiple coordinate positions of multiple characters in the first character string data; Based on the at least one target coordinate position and the multiple coordinate positions, the at least one character to be desensitized is determined in the first character string data.
6. The method according to claim 4, characterized in that The step of desensitizing each character to be desensitized based on the preset desensitization rule specifically includes: For any first character string data, obtain the data type of the character to be desensitized; In the preset desensitization rule, determine the preset desensitization method corresponding to the data type; The preset desensitization method is used to desensitize each of the characters to be desensitized.
7. The method according to any one of claims 1 to 6, characterized in that The step of generating a target log based on the desensitized at least one first string data, at least one second string data, and intercepted characters between two adjacent string data specifically includes: For any first character string data, write the first character string data into a first preset buffer area, and desensitize the first character string data in the first preset buffer area; Writing the desensitized first string data into a second preset buffer area, and writing the second string data adjacent to the first string data and the intercepted characters between the two into the second preset buffer area in the order of log data; According to the order of the log data, at least one first string data, at least one second string data and intercepted characters between two adjacent string data included in the second preset buffer area are summarized to obtain the target log.
8. A log desensitization device, characterized in that: include: The acquisition module is used to obtain the log data to be desensitized; A first generating module, configured to divide the log data based on a preset intercepted character set and a preset key name set to be desensitized, and generate at least one first string data containing sensitive characters and at least one second string data not containing sensitive characters, wherein the string data includes a plurality of characters between two intercepted characters; A determination module, configured to determine at least one to-be-desensitized character in each first character string data based on a preset desensitization character location rule, and desensitize each to-be-desensitized character based on the preset desensitization rule; The second generating module is used to generate a target log based on the desensitized at least one first string data, at least one second string data and intercepted characters between two adjacent string data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the log desensitization method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the log desensitization method as described in any one of claims 1 to 7 are implemented.