A log processing method, an electronic device, a storage medium and a program product

By deploying a cluster of proxy nodes in the data center for local log collection and sharding processing, combined with load balancing and parallel desensitization, the high latency and low efficiency problems in multi-data center log processing are solved, achieving efficient log processing and resource optimization.

CN120512358BActive Publication Date: 2025-10-24LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511006627.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-24
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

In existing technologies, log processing in multi-data center scenarios suffers from high latency and low processing efficiency, especially the cross-data center transmission performance bottleneck caused by reliance on centralized nodes.

Method used

A proxy node cluster is deployed in the data center. Logs are collected and processed locally through the proxy nodes. A load balancing mechanism is used to migrate high-load tasks to low-load nodes for processing. At the same time, a layered processing engine and rule hot loading module are used for parallel desensitization processing.

Benefits of technology

It improves the efficiency of log processing and resource utilization, avoids the waste of bandwidth transmitted across data centers, increases throughput, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512358B_ABST
    Figure CN120512358B_ABST
Patent Text Reader

Abstract

The application discloses a log processing method, an electronic device, a storage medium and a program product, and relates to the technical field of computers, and comprises the following steps: acquiring logs collected by each agent node in an agent node cluster, performing sharding on the logs to obtain first log segments, and distributing the first log segments to each agent node for desensitization processing; performing sharding on the to-be-processed logs of a first target agent node to obtain second log segments, and distributing the second log segments to a second target agent node, so that the second target agent node performs desensitization processing on the second log segments. According to the method, the agent node cluster is deployed in a data center, local log collection and desensitization processing are performed by each agent node in the agent node cluster, the logs are processed in the data center, the processing efficiency and the resource utilization rate can be improved, and bandwidth waste caused by cross-data-center log transmission is avoided. In addition, the method coordinates the tasks of the agent nodes, and can optimize resource utilization.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and in particular to a log processing method, an electronic device, a storage medium and a program product. BACKGROUND

[0002] Log desensitization is a technology for processing sensitive information in logs, aiming to preserve log usability while preventing sensitive data leakage and protecting user privacy and information security. With the acceleration of enterprise digital transformation, the demand for log desensitization in multi-data center scenarios is growing. Traditional solutions rely on centralized nodes for log collection and desensitization, requiring the transmission of massive logs across data centers to the central node for processing, which can easily form a performance bottleneck, resulting in high latency and low log processing efficiency. Therefore, how to solve the above technical defects has become a technical problem to be solved by those skilled in the art. SUMMARY

[0003] The present application provides a log processing method, an electronic device, a storage medium and a program product to at least solve the problem of high latency and low log processing efficiency in related technologies.

[0004] The present application provides a log processing method, comprising:

[0005] Obtaining logs collected by each agent node in an agent node cluster; the agent node cluster is deployed in a data center;

[0006] Sharding the logs to obtain first log segments, and distributing the first log segments to each agent node for desensitization processing;

[0007] Determining a first target agent node from the agent node cluster; the first target agent node is an agent node whose load meets a first preset condition;

[0008] Selecting a second target agent node from the agent node cluster; the second target agent node is an agent node whose load meets a second preset condition;

[0009] Sharding the logs to be processed of the first target agent node to obtain second log segments, and distributing the second log segments to the second target agent node, so that the second target agent node performs desensitization processing on the second log segments.

[0010] The present application also provides a log processing device, comprising:

[0011] An obtaining unit configured to obtain logs collected by each agent node in an agent node cluster; the agent node cluster is deployed in a data center;

[0012] The first fragmentation unit is configured to fragment the log to obtain a first log segment, and distribute the first log segment to each agent node for desensitization processing.

[0013] The determining unit is configured to determine a first target agent node from the agent node cluster; the first target agent node is an agent node whose load satisfies a first preset condition.

[0014] The selecting unit is configured to select a second target agent node from the agent node cluster; the second target agent node is an agent node whose load satisfies a second preset condition.

[0015] The second fragmentation unit is configured to fragment the log to be processed of the first target agent node to obtain a second log segment, and distribute the second log segment to the second target agent node, so that the second target agent node performs desensitization processing on the second log segment.

[0016] The application further provides an electronic device, including a memory configured to store a computer program, and a processor configured to execute the computer program to implement the steps of any of the log processing methods.

[0017] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of any of the log processing methods.

[0018] The application further provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the steps of any of the log processing methods.

[0019] The application deploys an agent node cluster in a data center, and each agent node in the agent node cluster collects and desensitizes local logs, so that the logs are processed in the data center, the processing efficiency and resource utilization are improved, and bandwidth waste caused by cross-data-center transmission of logs is avoided. In addition, the application distributes the fragmented logs to the agent nodes, and the agent nodes process in parallel, so that the throughput is improved. Furthermore, the application coordinates the tasks of the agent nodes, and migrates the tasks of the agent nodes with high load to the agent nodes with low load for processing, so that the resource utilization is optimized. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] In order to more clearly illustrate the embodiments of the application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.Figure 1 A flowchart of a log processing method provided by an embodiment of the present application is shown in FIG. 1.

[0022] Figure 2 A schematic diagram of a log processing system provided by an embodiment of the present application is shown in FIG. 2.

[0023] Figure 3 A schematic diagram of a log processing apparatus provided by an embodiment of the present application is shown in FIG. 3.

[0024] Figure 4 A schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 4. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0026] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0027] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0028] Embodiments of the present application provide a log processing method, which will be described in detail in combination with the execution flow of the method.

[0029] Reference Figure 1 As shown in FIG. 1, the log processing method provided by the embodiments of the present application includes:

[0030] S101: acquiring logs collected by each agent node in an agent node cluster; the agent node cluster is deployed in a data center.

[0031] The agent node cluster is deployed in the data center, and the agent node cluster includes a plurality of agent nodes. The number of agent nodes in the agent node cluster can be set differently, which is not limited by the present application.

[0032] The agent node is responsible for local collection, cleaning and desensitization processing of logs.

[0033] In some embodiments, obtaining the logs collected by each agent node in the cluster of agent nodes comprises:

[0034] reading the logs collected by the agent node from a message queue.

[0035] The agent node obtains local files through a tool, monitors changes in the local files through file listening, and sends newly added logs to a message queue kafka. The logs collected by the agent node can be read from the message queue.

[0036] S102: Sharding the logs to obtain first log segments, and distributing the first log segments to each agent node for desensitization processing.

[0037] Reading logs from a message queue, and cutting the read logs into independent log segments. The log segments obtained in step S101 are referred to as first log segments, so as to distinguish the log segments obtained in step S105 (referred to as second log segments). It should be noted that the terms "first" and "second" are not used to limit size, order, priority, etc.

[0038] Distribute the first log segments obtained by sharding to each agent node. Each agent node independently processes the assigned log segments to achieve parallel processing. After processing, write the processing result to the central storage. For example, write to Elasticsearch.

[0039] The agent node is built-in with a rule engine, which performs desensitization processing on the received first log segments through a function call mechanism. Exemplarily, the desensitization rule can be as follows:

[0040] rules:

[0041] #Original IP rule…

[0042] -id: "rule_employee_id"

[0043] log_type: "hr" #Only for HR system logs

[0044] field_type: "employee_id"

[0045] pattern: "E\d{5}" #Employee ID format: E + 5 digits

[0046] mode: "debug" #Debug mode: partial desensitization

[0047] action: "partial_mask".

[0048] When the agent node starts, load the de-identification rules in the de-identification strategy library as function pointers or regular expression templates.

[0049] In some embodiments, the de-identification strategy library stores the de-identification rules in the form of key-value pairs; the key is a log type, and the value is a de-identification function or a regular expression.

[0050] In this embodiment, the de-identification strategy library stores the de-identification rules in the form of key-value pairs, the key is a log type: field type (for example, audit: ip), and the value is a de-identification function or a regular expression. The agent node directly calls the de-identification rules through the function name or function pointer.

[0051] In some embodiments, the log is divided into a first log segment, and the first log segment is distributed to each agent node for de-identification processing, including:

[0052] According to a first preset time window, the log is divided into the first log segment, and the first log segment is distributed to each agent node in the agent node cluster; the first log segment of the same time window is processed by the same agent node.

[0053] In this embodiment, the division mode for the real-time log stream is as follows: the log is divided into independent log segments according to a pre-set fixed time window, and each log segment includes all log entries in the corresponding time window. The log segment obtained by division is mapped to the agent node, so as to ensure that the log segment of the same time window is processed by the fixed agent node.

[0054] The division mode according to the time window can be as follows:

[0055] .

[0056] In the formula, Hash represents a time conversion formula, which can convert time into a number; ShardTimestamp is time information of a message; Mod is a remainder function; and AgentNums is the number of agent nodes.

[0057] For example, Hash is a function of converting time information into seconds, so as to give the log segment of the same second to the same agent node for processing. AgentNums is set to 3, and the value of Hash (ShardTimestamp) is taken modulo 3. The remainder is 1, the log segment corresponding to the time information is distributed to the first agent node, the remainder is 2, the log segment corresponding to the time information is distributed to the second agent node, and so on.

[0058] It can be understood that the Hash function can be set in different modes to control the time interval.

[0059] S103: determining a first target agent node from the agent node cluster; the first target agent node is an agent node whose load satisfies a first preset condition.

[0060] S104: selecting a second target agent node from the agent node cluster; the second target agent node is an agent node whose load satisfies a second preset condition.

[0061] S105: performing sharding on the to-be-processed log of the first target agent node to obtain a second log segment, and distributing the second log segment to the second target agent node, so that the second target agent node performs desensitization processing on the second log segment.

[0062] Steps S103 to S105 aim to coordinate the tasks of agent nodes, and achieve resource balancing through task migration. Specifically, a first target agent node is determined from the agent node cluster. The first target agent node is a high-load agent node. The second target agent node is a low-load node.

[0063] In some embodiments, determining the first target agent node from the agent node cluster comprises:

[0064] calculating a load score of the agent node according to a preset load indicator of the agent node;

[0065] if the load score of the agent node exceeds a preset threshold, determining the agent node as the first target agent node.

[0066] The preset load indicator can include CPU usage, memory occupancy, to-be-processed log queue length, etc. Each agent node in the agent node cluster periodically reports its CPU usage, memory occupancy, to-be-processed log queue length, etc. According to each preset load indicator of each agent node, a load score of each agent node is calculated. The agent node whose load score exceeds a preset threshold (for example, 80%) is determined as the first target agent node. It can be understood that the number of the first target agent nodes can be one or more.

[0067] In some embodiments, calculating a load score of the agent node according to a preset load indicator of the agent node comprises:

[0068] performing weighted summation on the preset load indicator of the agent node to obtain the load score of the agent node.

[0069] Taking the preset load indicators including CPU usage, memory occupancy, and to-be-processed log queue length as an example, the way of calculating the load score of the agent node is as follows:

[0070] .

[0071] LoadScore represents a load score, CPU represents a CPU usage rate, Memory represents a memory occupancy rate, and QueueLength represents a length of a log queue to be processed. The weight corresponding to each preset load index can be preset and dynamically adjusted.

[0072] In some embodiments, selecting the second target proxy node from the proxy node cluster comprises:

[0073] Selecting a preset number of proxy nodes with the lowest load scores from the proxy node cluster as the second target proxy nodes.

[0074] For example, two proxy nodes with the lowest load scores are selected from the proxy node cluster as the second target proxy nodes.

[0075] It can be understood that the load scores of the selected preset number of proxy nodes can be the same. For example, the preset number is 3, the load score of the first proxy node, the load score of the second proxy node, and the load score of the third proxy node are the same and the lowest, and the first proxy node, the second proxy node, and the third proxy node are selected as the second target proxy nodes. The load scores of the selected preset number of proxy nodes can also be different. For example, the preset number is 3, the load score of the first proxy node, the load score of the second proxy node, and the load score of the third proxy node are different, and the load score of the first proxy node is the lowest, the load score of the second proxy node is the second lowest, and the load score of the third proxy node is only higher than those of the first proxy node and the second proxy node, and the first proxy node, the second proxy node, and the third proxy node are selected as the second target proxy nodes.

[0076] The log shards to be processed by the first target proxy node are divided to obtain second log segments, the second log segments are distributed to the second target proxy nodes for desensitization processing through RPC (Remote Procedure Call) or a message queue, and metadata (for example, log source and processing progress) is synchronized.

[0077] In some embodiments, dividing the logs to be processed by the first target proxy node to obtain second log segments comprises:

[0078] The logs to be processed by the first target proxy node are divided according to a second preset time window to obtain the second log segments.

[0079] The part of the to-be-processed log of the first target proxy node can be fragmented, and then the second log segments obtained by fragmentation are distributed to the second target proxy node for processing. At this time, the first target proxy node retains part of the to-be-processed log, and the part of the to-be-processed log is processed by the first target proxy node. The entire to-be-processed log of the first target proxy node can also be fragmented, and then the second log segments obtained by fragmentation are distributed to the second target proxy node for processing. At this time, the first target proxy node does not retain any to-be-processed log.

[0080] In some embodiments, distributing the second log segments to the second target proxy node comprises:

[0081] The second log segments are evenly distributed to the second target proxy node, or the second log segments are distributed to the second target proxy node according to the load scores of the second target proxy nodes, based on the principle that the lower the load score, the more second log segments are allocated.

[0082] For example, the second target proxy node is 3, and the to-be-processed log of the first target proxy node is divided into 3 parts after fragmentation. Each second target proxy node is responsible for one part.

[0083] For example, the second target proxy node includes a No. 1 proxy node and a No. 2 proxy node. The load score of the No. 1 proxy node is higher than that of the No. 2 proxy node. After the to-be-processed log of the first target proxy node is fragmented, the number of second log segments distributed to the No. 1 proxy node is less than that of the No. 2 proxy node.

[0084] In some embodiments, the method further comprises:

[0085] The target log is cut into a plurality of log blocks;

[0086] The log blocks are allocated to idle proxy nodes for desensitization processing;

[0087] The processing results of the idle proxy nodes are aggregated and a report is generated.

[0088] This embodiment is directed to a batch log processing scenario, and a target log is processed. The target log refers to a log that is not collected by a proxy node and uploaded to a message queue. For example, the target log is a log directly uploaded by a user. The target log can be cut according to size or number of rows to obtain a plurality of log blocks. The log blocks obtained by cutting are allocated to idle proxy nodes for desensitization processing. The proxy nodes return processing results after processing. The processing results returned by the proxy nodes are aggregated and a report is generated.

[0089] In some embodiments, the metadata of the log comprises a log type and a source, and the log type or the source has a corresponding relationship with a de-identification mode, so that the agent node determines the de-identification mode according to the metadata of the log, and performs de-identification processing by using the determined de-identification mode; the de-identification mode comprises a first de-identification mode and a second de-identification mode; the first de-identification mode is a mode in which all sensitive fields of the log are hidden; and the second de-identification mode is a mode in which part of the sensitive fields of the log are hidden.

[0090] The agent node sends the collected log to the message queue, and labels the log when sending, and the label specifies the de-identification mode used by the log. The de-identification mode comprises a first de-identification mode and a second de-identification mode; the first de-identification mode is a mode in which all sensitive fields of the log are hidden; and the second de-identification mode is a mode in which part of the sensitive fields of the log are hidden. The de-identification mode can have a corresponding relationship with the log type or the source. For example, the source comprises a production environment and a debugging environment. The de-identification mode corresponding to the production environment is the first de-identification mode. The de-identification mode corresponding to the debugging environment is the second de-identification mode.

[0091] When the agent node performs log de-identification processing, the de-identification mode is selected according to the log metadata, and then the corresponding de-identification mode is used for de-identification processing.

[0092] For example, in the strict mode, that is, the first de-identification mode:

[0093] The sensitive fields (such as IP, ID number) are replaced with irreversible hash values (SHA-256).

[0094] For example: 192.168.1.1 is de-identified to a3f4b2...c89d.

[0095] In the debugging mode, that is, the second de-identification mode:

[0096] Part of the de-identification, part of the information is retained for debugging.

[0097] For example: the IP address is de-identified to 192.168.xx.xx, and the mobile phone number is de-identified to 138****5678.

[0098] According to the log type or the source to determine the de-identification mode, the de-identification strength can be dynamically adjusted to meet different de-identification needs.

[0099] In some embodiments, it further comprises:

[0100] Listen to the change event of the rule configuration file;

[0101] When the change event is listened to, the new rule is parsed and recorded to the rule table;

[0102] Replace the old rule pointer by atomic operation;

[0103] Syntax check of new rules;

[0104] If the check fails, roll back to the previous version.

[0105] This embodiment aims to dynamically update the de-sensitization rules without restarting the service.

[0106] The rule update process includes:

[0107] Listen to the change event of the rule configuration file, and use inotify or Watchdog library to realize real-time listening.

[0108] After detecting file changes, parse new rules and load them into the rule table in memory. Replace the old rule pointer through atomic operation to avoid process interruption.

[0109] Syntax check of new rules, check the rules of the configuration file, and roll back to the previous version when failed.

[0110] This embodiment automatically maintains de-sensitization rules, which can avoid the problems of insufficient de-sensitization coverage, service restart or manual rule loading caused by manual rule updating.

[0111] In some embodiments, it further includes:

[0112] Call a preset model to identify unidentified log segments;

[0113] Generate de-sensitization rules according to the identification results output by the large model;

[0114] Send the de-sensitization rules to the console for review, and when the de-sensitization rules are approved, inject the de-sensitization rules into the de-sensitization policy library, and broadcast to each proxy node.

[0115] This embodiment aims to identify unknown sensitive data and generate de-sensitization rules. The processing flow includes:

[0116] When the proxy node finds that the field does not match any de-sensitization rule, trigger an analysis request.

[0117] Input the unmatched log segment into a large model, for example, LLM (Large Language Model), and prompt examples, such as:

[0118] .

[0119] The model returns a structured result, for example:

[0120] {“TempToken”: “hexadecimal”, “IP”: “IPv4”}.

[0121] According to the large model output, a regular expression is automatically generated (for example, \b0x[0-9A-F]{4,}\b matches a hexadecimal token).

[0122] The new rule is submitted to the administrator console for review, and after passing, it is injected into the desensitization strategy library and broadcast to all proxy nodes.

[0123] Through tools such as ZooKeeper / Consul, configuration files can be dynamically pushed to all proxy nodes to dynamically inject new sensitive entity types through configuration files, and take effect immediately.

[0124] Through large model technology, logs that are not recognized by desensitization rules are identified, and new rules are generated, which can continuously improve desensitization rules, form a rule library closed loop iteration, break through the limitations of regular matching, better cope with new sensitive data formats (such as mixed language logs), and improve desensitization accuracy.

[0125] Reference Figure 2 As shown, the log processing method provided by the embodiments of the application can be implemented by a log processing system, which includes a proxy node cluster, a hierarchical processing engine, a rule hot loading module, a load balancing module, a desensitization strategy library, and an intelligent analysis module.

[0126] The proxy node cluster includes a plurality of proxy nodes, which are responsible for local collection, cleaning, and desensitization of logs. The proxy nodes obtain local files through tools. Through file listening, changes in local files are monitored, and after changes are obtained, the newly added logs are sent to the message queue kafka. The proxy node has a built-in rule engine, which performs desensitization operations through a function call mechanism. The proxy node loads the desensitization rules in the desensitization strategy library when it starts, and stores them as function pointers or regular expression templates.

[0127] The proxy node can select a desensitization mode (strict mode / debug mode) according to log metadata (such as log_type=audit, env=production). If the field matches an existing rule (such as IP address regular \d{1,3}\.\d{1,3}\.\d{1,3}\.\d{1,3}), the predefined desensitization function is called. If the field does not match the rule, a rule missing event is triggered, and a large model analysis process is started.

[0128] The hierarchical processing engine supports real-time log stream sharding processing and batch log concurrent processing to improve throughput.

[0129] Real-time log stream sharding processing:

[0130] The hierarchical processing engine divides the log into independent shards, i.e. the first log segment in the foregoing, according to a fixed time window, and each shard contains all log entries within the time window. The hierarchical processing engine assigns the shards to corresponding proxy nodes, ensuring that the logs of the same time window are processed by the same proxy node. Each proxy node independently processes the assigned shard, and writes the result to the central storage after processing is completed.

[0131] Batch log concurrent processing:

[0132] The hierarchical processing engine divides a large log file into multiple blocks according to size or number of lines. The blocks are assigned to idle proxy nodes for processing, and the task status is recorded. The proxy nodes return the results after processing is completed, and the hierarchical processing engine aggregates the data and generates a final report.

[0133] The load balancing module is responsible for dynamically coordinating the task load of the proxy nodes, and achieving resource balance through task migration. Each proxy node periodically reports CPU usage, memory occupancy, and length of the queue of logs to be processed. The load balancing module maintains a global load table, and calculates the load score of the proxy nodes according to the weight.

[0134] When the load score of a certain proxy node exceeds a threshold value, the following operations are performed:

[0135] The logs to be processed of the proxy node are divided into multiple shards according to time windows. The proxy node with the lowest load is selected according to the load table. The shards are sent to the selected proxy node through RPC or a message queue, and the metadata of the logs (such as log source and processing progress) is synchronized.

[0136] The desensitization strategy library stores multi-level desensitization rules, and supports dynamic calling by the proxy nodes according to log types or sources. The desensitization strategy library stores the rules in the form of key-value pairs, with the key being the log type:field type, and the value being a desensitization function or a regular expression. The proxy nodes directly call the rules through function names or function pointers.

[0137] The rule hot loading module is responsible for dynamically updating the desensitization rules without restarting the service.

[0138] The rule updating process includes:

[0139] Configuration file monitoring: The proxy node monitors the change event of the rule configuration file, and uses inotify or Watchdog library to realize real-time monitoring.

[0140] Memory hot loading:

[0141] After detecting the file change, the new rules are parsed and loaded into the rule table in the memory. The old rule pointer is replaced through atomic operation to avoid interruption during processing.

[0142] Validation of effectiveness:

[0143] The agent node performs syntax checking on the new rule, checks the rule of the configuration file, and rolls back to the previous version when the checking fails.

[0144] The intelligent analysis module automatically identifies unknown sensitive data in combination with a large model, and generates a desensitization rule.

[0145] When the agent node finds that the field does not match any rule, an analysis request is triggered.

[0146] The unmatched log segment is input to the large model, and the model returns a structured result. According to the output of the large model, a regular expression is automatically generated. The new rule is submitted to the administrator console for review, and after passing the review, the desensitization strategy library is injected and broadcast to all agent nodes. The rule hot loading module can dynamically push the configuration file to all agent nodes through tools such as ZooKeeper / Consul.

[0147] As a specific implementation, the layered processing engine can be deployed on each agent node, the rule hot loading module can be deployed on each agent node, the intelligent analysis module can be deployed on each agent node, the load balancing module can be deployed on the center node in the agent node cluster, and the desensitization strategy library can be deployed on the center node in the agent node cluster.

[0148] In summary, the present application deploys an agent node cluster in a data center, and each agent node in the agent node cluster collects and desensitizes local logs. The logs are processed in the data center, which can improve processing efficiency and resource utilization, and avoid bandwidth waste caused by cross-data-center transmission of logs. In addition, the present application distributes the log fragments to the agent nodes for parallel processing, which can improve throughput. Furthermore, the present application coordinates the tasks of the agent nodes, and migrates the tasks of high-load agent nodes to low-load agent nodes for processing, which can optimize resource utilization.

[0149] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and a general hardware platform as necessary, and of course, it can also be realized by hardware, but in many cases, the former is a better implementation.

[0150] The embodiments of the present application also provide a log processing device. Referring to FIG. 1, the device includes: Figure 3

[0151] An acquisition unit 10 is configured to acquire logs collected by each agent node in an agent node cluster; and the agent node cluster is deployed in a data center.

[0152] A first fragmentation unit 20 is configured to fragment the logs to obtain first log fragments, and distribute the first log fragments to each agent node for desensitization processing.​

[0153] The determining unit 30 is configured to determine a first target proxy node from the proxy node cluster; the first target proxy node is a proxy node whose load satisfies a first preset condition.

[0154] The selecting unit 40 is configured to select a second target proxy node from the proxy node cluster; the second target proxy node is a proxy node whose load satisfies a second preset condition.

[0155] The second fragmenting unit 50 is configured to fragment the to-be-processed log of the first target proxy node to obtain a second log fragment, and distribute the second log fragment to the second target proxy node, so that the second target proxy node performs the desensitization processing on the second log fragment.

[0156] In the above embodiment, as a specific implementation, the metadata of the log includes a log type and a source, and the log type or the source has a corresponding relationship with a desensitization mode, so that the proxy node determines the desensitization mode according to the metadata of the log, and performs the desensitization processing by using the determined desensitization mode; the desensitization mode includes a first desensitization mode and a second desensitization mode; the first desensitization mode is a mode in which all sensitive fields of the log are hidden; and the second desensitization mode is a mode in which part of the sensitive fields of the log are hidden.

[0157] In the above embodiment, as a specific implementation, the method further includes:

[0158] The monitoring unit is configured to monitor a change event of the rule configuration file.

[0159] The parsing unit is configured to, when the change event is monitored, parse a new rule and record the new rule to a rule table.

[0160] The replacing unit is configured to replace an old rule pointer by using an atomic operation.

[0161] The checking unit is configured to perform a syntax check on the new rule.

[0162] The rollback unit is configured to, if the check fails, roll back to a previous version.

[0163] In the above embodiment, as a specific implementation, the method further includes:

[0164] The calling unit is configured to call a preset model to identify the unidentified log fragment.

[0165] The generating unit is configured to generate a desensitization rule according to an identification result output by the preset model.

[0166] An injection unit is configured to send the desensitization rule to a console for review, and when the desensitization rule passes the review, inject the desensitization rule into a desensitization policy library and broadcast to each of the proxy nodes.

[0167] In the above embodiment, as a specific implementation, the desensitization policy library stores the desensitization rule in the form of a key-value pair; the key is a log type, and the value is a desensitization function or a regular expression.

[0168] In the above embodiment, as a specific implementation, the first slicing unit 20 is configured to:

[0169] slice the log according to a first preset time window to obtain the first log segment, and distribute the first log segment to each of the proxy nodes in the proxy node cluster; the first log segments of the same time window are processed by the same proxy node.

[0170] In the above embodiment, as a specific implementation, the determination unit 30 includes:

[0171] A calculation subunit is configured to calculate a load score of the proxy node according to a preset load index of the proxy node;

[0172] A determination subunit is configured to determine the proxy node as the first target proxy node if the load score of the proxy node exceeds a preset threshold.

[0173] In the above embodiment, as a specific implementation, the calculation subunit is configured to:

[0174] weight and sum the preset load index of the proxy node to obtain the load score of the proxy node.

[0175] In the above embodiment, as a specific implementation, the selection unit 40 is configured to:

[0176] select a preset number of proxy nodes with the lowest load scores from the proxy node cluster as the second target proxy nodes.

[0177] In the above embodiment, as a specific implementation, the second slicing unit 50 is configured to:

[0178] slice the to-be-processed log of the first target proxy node according to a second preset time window to obtain the second log segment.

[0179] In the above embodiment, as a specific implementation, the system further includes:

[0180] A third slicing unit is configured to divide the target log into a plurality of log blocks.

[0181] an allocation unit configured to allocate the log blocks to idle proxy nodes for desensitization processing;

[0182] an aggregation unit configured to aggregate processing results of the idle proxy nodes and generate a report.

[0183] In the above embodiments, as a specific implementation, the obtaining unit 10 is configured to:

[0184] read the logs collected by the proxy nodes from the message queue.

[0185] The log processing apparatus provided in the present application deploys a proxy node cluster in a data center, and each proxy node in the proxy node cluster collects and desensitizes local logs. The logs are processed in the data center, which can improve processing efficiency and resource utilization and avoid bandwidth waste caused by cross-data-center transmission of logs. In addition, the log processing apparatus provided in the present application distributes log shards to proxy nodes for parallel processing by the proxy nodes, which can improve throughput. Furthermore, the log processing apparatus provided in the present application coordinates tasks of the proxy nodes, migrates tasks of high-load proxy nodes to low-load proxy nodes for processing, and can optimize resource utilization.

[0186] The features of the embodiments of the log processing apparatus can be understood with reference to the related descriptions of the embodiments of the log processing method, which will not be described herein.

[0187] The embodiments of the present application also provide an electronic device. As shown in the accompanying drawings, the electronic device includes a memory 1 and a processor 2. The memory 1 stores a computer program. The processor 2 is configured to run the computer program to perform the steps in any of the above log processing method embodiments. Figure 4

[0188] The electronic device provided in the present application deploys a proxy node cluster in a data center, and each proxy node in the proxy node cluster collects and desensitizes local logs. The logs are processed in the data center, which can improve processing efficiency and resource utilization and avoid bandwidth waste caused by cross-data-center transmission of logs. In addition, the electronic device provided in the present application distributes log shards to proxy nodes for parallel processing by the proxy nodes, which can improve throughput. Furthermore, the electronic device provided in the present application coordinates tasks of the proxy nodes, migrates tasks of high-load proxy nodes to low-load proxy nodes for processing, and can optimize resource utilization.

[0189] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. The computer program is configured to perform the steps in any of the above log processing method embodiments when running.​

[0190] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0191] Embodiments of the present application also provide a computer program product, which includes a computer program, and the computer program, when executed by a processor, implements the steps in any of the log processing method embodiments described above.

[0192] Embodiments of the present application also provide another computer program product, which includes a non-volatile computer readable storage medium, and the non-volatile computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps in any of the log processing method embodiments described above.

[0193] The skilled in the art can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0194] The above describes in detail a log processing method, an electronic device, a storage medium and a program product provided by the present application. The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above description of the examples is only used to help understand the method and its core idea of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A log processing method characterized by, The method comprises the following steps: obtaining logs collected by each agent node in an agent node cluster; the agent node cluster is deployed in a data center; the agent node is responsible for local collection and desensitization processing of the logs; sharding the logs to obtain first log segments, and distributing the first log segments to each agent node for desensitization processing; each agent node independently processes the assigned log segments to realize parallel processing, and writes the processing result to the central storage after processing is completed; determining a first target agent node from the agent node cluster; the first target agent node is an agent node whose load meets a first preset condition; selecting a second target agent node from the agent node cluster; the second target agent node is an agent node whose load meets a second preset condition; sharding the logs to be processed of the first target agent node to obtain second log segments, and distributing the second log segments to the second target agent node, so that the second target agent node performs desensitization processing on the second log segments; calling a preset model to identify unidentified log segments; generating a desensitization rule according to the identification result output by the preset model; sending the desensitization rule to a control console for review, and when the desensitization rule passes the review, injecting the desensitization rule into a desensitization strategy library and broadcasting it to each agent node.

2. The log processing method according to claim 1, characterized by, The metadata of the logs includes log types and sources, and the log types or sources have a corresponding relationship with desensitization modes, so that the agent node determines the desensitization mode according to the metadata of the logs, and performs desensitization processing by using the determined desensitization mode; the desensitization mode includes a first desensitization mode and a second desensitization mode; the first desensitization mode is a mode in which all sensitive fields of the logs are hidden; the second desensitization mode is a mode in which part of the sensitive fields of the logs are hidden.

3. The log processing method of claim 1, wherein, Further comprising: listening to change events of a rule configuration file; when a change event is listened to, parsing a new rule and recording it to a rule table; replacing an old rule pointer through atomic operation; performing syntax checking on the new rule; if the checking fails, rolling back to the previous version.

4. The log processing method of claim 1, wherein, The desensitization strategy library stores desensitization rules in the form of key-value pairs; the key is the log type, and the value is the desensitization function or regular expression.

5. The log processing method of claim 1, wherein, The method of sharding the logs to obtain first log segments, and distributing the first log segments to each agent node for desensitization processing comprises: sharding the logs according to a first preset time window to obtain the first log segments, and distributing the first log segments to each agent node in the agent node cluster; the first log segments of the same time window are processed by the same agent node.

6. The log processing method of claim 1, wherein, Determining a first target agent node from the agent node cluster comprises: calculating the load score of the agent node according to the preset load index of the agent node; if the load score of the agent node exceeds a preset threshold, the agent node is determined as the first target agent node.

7. The log processing method of claim 6, wherein, Calculating the load score of the agent node according to the preset load index of the agent node comprises: weighting and summing the preset load index of the agent node to obtain the load score of the agent node.

8. The log processing method of claim 1, wherein, The selecting a second target agent node from the agent node cluster comprises: The second target agent node is selected from a preset number of agent nodes with the lowest load score in the agent node cluster.

9. The log processing method of claim 1, wherein, The second log segment is obtained by sharding the to-be-processed log of the first target agent node, comprising: The second log segment is obtained by sharding the to-be-processed log of the first target agent node according to a second preset time window.

10. The log processing method of claim 1, wherein, Further comprising: The target log is divided into a plurality of log blocks; The log blocks are assigned to idle agent nodes for desensitization processing; The processing results of the idle agent nodes are aggregated and a report is generated.

11. The log processing method of claim 1, wherein, The log collected by each agent node in the agent node cluster comprises: The log collected by the agent node is read from the message queue.

12. An electronic device, comprising: Comprising: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the log processing method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium, and when the computer program is executed by the processor, the steps of the log processing method according to any one of claims 1 to 11 are implemented.

14. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the log processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and system for desensitizing sensitive data of contract file

    CN112800460A

  • Log writing method and device, storage medium and electronic equipment

    CN117009192A

  • Log data processing method and device, equipment, storage medium and program product

    CN118394713A