Intercept log deduplication method, device and equipment of terminal protection system
By using the method of cyclic queue and sliding time window in the terminal protection system, the duplicate logs are judged using feature metadata groups and hash algorithms, which solves the problem of overlapping fixed window logs and realizes efficient log deduplication of the terminal protection system.
Patent Information
- Application Number
- CN202510865709.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
When existing terminal protection systems use fixed time windows to deduplicate logs, duplicate logs in adjacent time windows may not be recognized, causing log overlap and affecting user experience.
The circular queue and sliding time window method are used to obtain the characteristic metadata group of the intercept log set, and the hash algorithm is used to generate the feature identifier, and the historical feature metadata group is traversed in the time slice order in the loop queue, and the duplicate feature identifier is judged. According to the timing identification difference value and the time window period, whether to discard or anchor the feature metadata group and log are determined to achieve dynamic deduplication.
It effectively solves the merge problem of the same audit or intercept logs under the sliding time window, avoids overlapping of fixed window logs, and improves the accuracy and efficiency of log deduplication.
Smart Images

Figure CN120371797A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of log deduplication, and particularly relates to a method, device, and equipment for deduplicating intercepted logs of a terminal protection system. Background Art
[0002] In a terminal protection system under Windows, when user-mode software operates on a PE file (Portable Executable), multiple kernel-mode driver events will be triggered, resulting in multiple audit or intercepted logs. Deduplication within a fixed time window is a simple and effective merging method, but duplicate logs may still be generated in adjacent time windows, causing certain troubles to users. For example, assume that the fixed time window period is 1 s, and t1 and t2 are two consecutive fixed time windows; if there is log A1 in the last 0.1 s of the t1 fixed time window and log A2 in the first 0.1 s of the t2 time window, when log A1 and log A2 are duplicate logs, even if the time interval between log A1 and log A2 is 0.2 s, due to the use of a fixed time window, duplicate log A1 may not be detected when log A2 is obtained, and deduplication of log A2 cannot be achieved. Therefore, the fixed time window only deduplicates within that window, and logs in two adjacent windows cannot be deduplicated even if the time interval may be less than 1 s. Summary of the Invention
[0003] To solve the above problems, the present invention provides a method, device, and equipment for deduplicating intercepted logs of a terminal protection system, which solves the merging of the same audit or intercepted logs under a sliding time window of the terminal protection system and solves the troubles caused by possible overlap of fixed-window logs.
[0004] In a first aspect, a method for deduplicating intercepted logs of a terminal protection system provided by the present invention includes: S1, obtaining an intercepted log set, determining a characteristic metadata group for each intercepted log in the intercepted log set, and obtaining a characteristic metadata group set corresponding to the intercepted log set; wherein, the characteristic metadata group includes a characteristic identifier and a timing identifier of the intercepted log; S2. For each feature metadata group in the set of feature metadata groups, traverse the historical feature metadata group set in each node of the circular queue in the order of time slices starting from the tail node, and determine whether there is a duplicate feature identifier. If there is a duplicate feature identifier, determine whether to discard or anchor the target feature metadata group and the target interception log corresponding to the duplicate feature identifier according to whether the difference between the position of the traversed node and the time sequence identifier exceeds the time window period. If no duplicate feature identifier is found, anchor the target feature metadata group and the target interception log. Wherein, the time sequence identifier difference is the difference between the time sequence identifier in the target feature metadata group and the time sequence identifier corresponding to the duplicate feature identifier in the historical feature metadata group set, and the historical feature metadata group sets with a preset time slice span are stored in each node of the circular queue according to a preset first data structure. S3. Insert the anchored feature metadata group into the tail node of the circular queue according to the preset first data structure. S4. Repeat steps S1 - S3 to continuously process the new interception log set.
[0005] In an alternative embodiment, the capacity of the circular queue C is: (1) In formula (1), C represents the number of nodes in the circular queue; represents the time slice span value, represents the time window period value. Wherein, is an integer multiple of.
[0006] In an alternative embodiment, step S2 includes: S21. Compare the feature identifier of the target interception log with the historical feature identifiers in the historical feature metadata group set in the current node to determine whether there is a duplicate feature identifier. S22. Determine whether the traversal of the circular queue is over. If the traversal is not over, determine whether the historical feature metadata group set of the C-th node is traversed. S23. If the historical feature metadata group set of the C-th node is traversed, determine whether there is a duplicate feature identifier. If there is, determine whether the time sequence identifier difference exceeds the preset time window period. If it does not exceed the preset time window period, discard the target feature metadata group and the target log. If it exceeds the preset time window period, anchor the target feature metadata group and the target log corresponding to the duplicate feature identifier. S24. If the historical feature metadata group set of non - the Cth node is traversed, determine whether there is a duplicate feature identifier; if there is a duplicate feature identifier, discard the target feature metadata group and the target log; if there is no duplicate feature identifier, update the existing node and return to step S21 to continue traversing.
[0007] In an alternative embodiment, the feature identifier is generated by using a hashing algorithm based on the path, creation time, last modification time, and file size of the intercepted log.
[0008] In an alternative embodiment, the hashing algorithm is the MD5 algorithm.
[0009] In an alternative embodiment, the time - series identifier is the timestamp when the intercepted log is intercepted.
[0010] In an alternative embodiment, the first data structure is one of a red - black tree or a hash table.
[0011] In an alternative embodiment, the target feature metadata group further includes the number of duplicate logs.
[0012] In a second aspect, a deduplication device for intercepted logs of a terminal protection system provided by the present invention includes: A data acquisition module, configured to acquire an intercepted log set, determine the feature metadata group of each intercepted log in the intercepted log set, and obtain the feature metadata group set corresponding to the intercepted log set; wherein, the feature metadata group includes a feature identifier and a time - series identifier of the intercepted log. A traversal module, configured to, for each feature metadata group in the feature metadata group set, traverse the historical feature metadata group set in each node of the circular queue starting from the tail node in the order of time slices, and determine whether there is a duplicate feature identifier; if there is a duplicate feature identifier, determine whether to discard or anchor the target feature metadata group and the target intercepted log corresponding to the duplicate feature identifier according to whether the difference between the node position where the traversal reaches and the time - series identifier exceeds the time - window period; if no duplicate feature identifier is found, anchor the target feature metadata group and the target intercepted log; wherein, the time - series identifier difference is the difference between the time - series identifier in the target feature metadata group and the time - series identifier corresponding to the duplicate feature identifier in the historical feature metadata group set, and each node in the circular queue stores the historical feature metadata group set with a preset time - slice span according to a preset first data structure. A queue update module, configured to insert the anchored target feature metadata group into the tail node of the circular queue according to the preset first data structure. A repetition module, configured to repeat the data acquisition module, the traversal module, and the queue update module to continuously process a new intercepted log set.
[0013] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of the foregoing embodiments are implemented.
[0014] In a fourth aspect, the present invention provides a computer-readable medium having non-volatile program code executable by a processor, and the program code causes the processor to execute the method according to any one of the foregoing embodiments.
[0015] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention are as follows: For the method, device, and equipment for deduplicating interception logs of the terminal protection system of the present invention, first, by obtaining an interception log set, the characteristic metadata group of each log in the interception log set is determined to obtain a characteristic metadata group set corresponding to the interception log set; then, for each characteristic metadata group in the characteristic metadata group set, starting from the tail node in the order of time slices, each node in the circular queue is traversed to check the historical characteristic metadata group set therein to determine whether there is a duplicate characteristic identifier; if there is a duplicate characteristic identifier, it is determined whether to discard or anchor the target characteristic metadata group and the target interception log corresponding to the duplicate characteristic identifier according to whether the difference between the traversed node position and the time sequence identifier exceeds the time window period. If no duplicate characteristic identifier is found, the target characteristic metadata group and the target interception log are anchored; finally, the anchored target characteristic metadata group is inserted into the tail node of the circular queue according to a preset first data structure; this process is repeated to continuously process new interception log sets; the present invention utilizes the elimination strategy of the circular queue and the dynamic sliding control strategy of the time window to achieve log deduplication, solve the merging of the same audit or interception logs in the terminal protection system under the sliding time window, and solve the problem of possible overlap of fixed window logs. Description of the Drawings
[0016] Figure 1 It is a schematic flowchart of the method for deduplicating interception logs of the terminal protection system provided by the embodiments of the present invention; Figure 2 It is a schematic diagram of the principle of storing MD5 logs in a circular queue provided by the embodiments of the present invention; Figure 3 It is a schematic diagram of the principle of traversing a circular queue provided by the embodiments of the present invention; Figure 4 It is another schematic diagram of the principle of traversing a circular queue provided by the embodiments of the present invention; Figure 5 It is a schematic flowchart of step S2 provided by the embodiments of the present invention; Figure 6 It is a schematic flowchart of the working process of the timer provided by the embodiments of the present invention; Figure 7Schematic diagram of the system principle of the interception log deduplication device for the terminal protection system provided by the embodiment of the present invention; Figure 8 Schematic diagram of the system principle of the electronic device provided by the embodiment of the present invention.
[0017] In the figure: 100 - data acquisition module; 200 - traversal module; 300 - queue update module; 400 - duplicate module; 1000 - electronic device; 1001 - communication interface; 1002 - processor; 1003 - memory; 1004 - bus. Specific embodiments
[0018] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0019] Refer to Figure 1 , a method for deduplicating interception logs of a terminal protection system, including the following steps S1 - step S4.
[0020] S1, obtain an interception log set, determine the characteristic metadata group of each interception log in the interception log set, and obtain a characteristic metadata group set corresponding to the interception log set; wherein, the characteristic metadata group includes the characteristic identifier and the time sequence identifier of the interception log.
[0021] Specifically, the execution subject of this embodiment is the terminal protection system. When the terminal protection system intercepts an interception log or a batch of interception logs, the batch of interception logs obtained in this embodiment is called an interception log set.
[0022] In this embodiment, a circular queue is used as the basic data structure, and each node of the circular queue stores data using a first data structure. Among them, the first data structure can be a search tree. Preferably, the search tree is implemented using a red - black tree. The red - black tree has balanced insertion and search performance, which is convenient for quick insertion and search, and the memory usage is flexible. In this embodiment, for the initially obtained interception logs, their characteristic data is extracted and stored as characteristic metadata groups to obtain a historical characteristic metadata group set. Here, the characteristic metadata group includes the characteristic identifier and the time sequence identifier of the interception log. The characteristic identifier is generated using a hash algorithm based on four elements: the path, creation time, last modification time, and file size of the interception log. These four elements can uniquely represent the interception log. Among them, the hash algorithm is the MD5 algorithm (Message Digest Algorithm 5), and the time sequence identifier is the timestamp when the interception log is intercepted.
[0023] For example, when the interception log is a PE file, the MD5 value of the four - element group (path, creation time, last modification time, and file size) of the PE files generated within the time window of the search tree elements and the log generation time (the timestamp when the PE file is intercepted).
[0024] In addition, the circular queue does not store MD5 logs (i.e., feature metadata groups) without limit, but has a limit. The C of the circular queue is set as: (1) In formula (1), C represents the number of nodes in the circular queue; represents the time slice span value, represents the time window period value; among them, is an integer multiple of.
[0025] In this embodiment, in addition to dividing the time window, the time slice is also used to measure the nodes of the circular queue (1 node stores the MD5 logs of 1 time slice).
[0026] For example, if the time window period value is 2s and the time slice span value is 1s, then the capacity C of the circular queue is 3. In this way, the circular queue only needs to store the MD5 logs of 3 nodes, and one node corresponds to the logs of 1 time slice. In this way, in the subsequent process of log deduplication, when the intercepted log obtained at 0.9s is considered, since the logs of 3s (1 node corresponds to the logs of 1s) are retained, the time window can be used to deduplicate the logs between 0.9s and 2.9s.
[0027] S2. For each feature metadata group in the set of feature metadata groups, traverse the historical feature metadata group sets in each node of the circular queue starting from the tail node in the order of time slices, and determine whether there are duplicate feature identifiers; if there are duplicate feature identifiers, then determine whether to discard or anchor the target feature metadata group and the target intercepted log corresponding to the duplicate feature identifier according to whether the difference between the node position where the traversal arrives and the time sequence identifier exceeds the time window period. If no duplicate feature identifier is found, then anchor the target feature metadata group and the target intercepted log; among them, the time sequence identifier difference is the difference between the time sequence identifier in the target feature metadata group and the time sequence identifier corresponding to the duplicate feature identifier in the historical feature metadata group set, and the historical feature metadata group sets with a preset time slice span are stored in each node of the circular queue according to a preset first data structure.
[0028] Preferably, referring to Figure 5 , step S2 includes the following steps S21 to step S24.
[0029] S21. Compare the feature identifier of the target intercepted log with the historical feature identifiers in the historical feature metadata group set in the current node to determine whether there are duplicate feature identifiers.
[0030] S22. Determine whether the traversal of the circular queue is over. If the traversal is not over, then determine whether the historical feature metadata group set of the Cth node is traversed.
[0031] S23, if the historical feature metadata group set of the C-th node is traversed, go to step S231; otherwise, go to step S24.
[0032] S231, determine whether there is a duplicate feature identifier. If there is, perform the following steps S232 to S233; if not, directly perform step S234.
[0033] Step S232, determine whether the difference in the timing identifier exceeds the preset time window period.
[0034] Step S233, if it does not exceed the time window period, discard the target feature metadata group and the target log.
[0035] Step S234, if it exceeds the time window period, anchor the target feature metadata group and the target log corresponding to the duplicate feature identifier to the current time slice and record the log (insert it into the tail node of the queue).
[0036] S24, if the historical feature metadata group set of a node other than the C-th node is traversed, determine whether there is a duplicate feature identifier; if there is a duplicate feature identifier, discard the target feature metadata group and the target log; update the node and return to step S21 to continue traversing; if there is no duplicate feature identifier, update the existing node and return to step S21 to continue traversing.
[0037] Specifically, the nodes of the circular queue represent time windows, and the red-black tree stores the MD5 logs generated within the time window under this window.
[0038] During traversal, use a timer to time the time slice, rotate cyclically in the order of time slices, and always keep the data in the circular queue new in and old out. In a possible embodiment, the process of the timer is as Figure 6 shown, including the following steps S61 to S64.
[0039] Step S61, the timer starts.
[0040] Step S62, determine whether it is the time slice; if so, perform step S63, otherwise sleep.
[0041] Step S63, insert data (i.e., the target feature metadata group) into the tail of the circular queue.
[0042] Step S64, determine whether the capacity C of the circular queue is reached; if so, go to step S65; if not, go to step S62.
[0043] Step S65, pop the head node of the queue. Go to step S62.
[0044] When new logs arrive, trigger the action of querying. For example, for target intercepted logs, query each node of the circular queue. If there is no duplicate feature identifier among all nodes, that is, there is no identical intercepted log, report the target intercepted log to the latest node. If there is a duplicate intercepted log, discard the target intercepted log.
[0045] The time window is a sliding window at the business level. If the time window period is 2s, that is, the business time window is 2s, and the update period of the circular queue node is 1s (i.e., the time slice span is 1s). Every 1s (one time slice), the node at the farthest 2s should be discarded. Since the circular queue has nodes from 1 - 2s, 0 - 1s, and a node from 2 - 3s is also needed. In this way, if the current time is 0.6s, the MD5 logs from 2 - 2.6s can be traversed. Therefore, the capacity of the circular queue is set to 2 + 1 = 3 nodes.
[0046] For the red - black tree, this embodiment further elaborates with the update period of the circular queue node being 1s. When the circular queue of this embodiment stores the feature metadata group, the node is updated once every 1s. That is, one node of the circular queue manages all intercepted logs within 1s. If there are multiple intercepted logs within this 1s, then these multiple logs should all belong to this node.
[0047] The storage structure of the red - black tree is convenient for sorting and searching. When traversing for searching, the traversing direction is from the tail node of the queue to the head node of the queue. As Figure 4 shown, search from right to left. The right side is the node of the most recently occurred time period. The node to the left is the node of the previous time slice, and the left - most is the node of the time slice farthest away. Here, the search traversal means searching in a certain node to see if there is an MD5 value in the red - black tree that is the same as the MD5 value of the target intercepted log. If there is the same MD5 value, discard the target intercepted log, that is, the feature metadata group corresponding to the target intercepted log is not inserted into the node of the circular queue. If there is no same MD5 value, insert the feature metadata group corresponding to the target intercepted log into the circular queue node. In this way, the information about whether the same target intercepted log exists within the adjacent window period (0 - 3s) is always maintained.
[0048] In addition, during traversal in this embodiment, at different nodes, the processing methods for the found identical MD5 logs are different. That is, it is necessary to determine whether to discard or retain the target feature metadata group and the target interception log according to whether the difference between the node position traversed and the time sequence identifier exceeds the time window period. If the same MD5 log is found at the first to the (C - 1)th nodes, the target interception log is directly discarded. However, if the same MD5 log is found at the Cth node, it is necessary to judge the difference in time sequence identifiers. If the time sequence difference is greater than the time window period value, the target interception log is directly retained. If the time sequence difference is less than or equal to the time window period value, the target interception log is discarded.
[0049] For example, the target interception log was obtained at 9:15:26 seconds (s). When traversing the circular queue from right to left, if the same MD5 log is not found within 0 - 1 s and is found within 1 - 2 s, then the target interception log is directly discarded. If the same MD5 log is found within 2 - 3 s and the timestamp of the same MD5 log is 9:15:06 seconds, then the time sequence identifier difference is 2 s, and the target interception log is discarded.
[0050] The operating principle of the circular queue in this embodiment is to insert from one end (new MD5 logs) and pop from the other end (discard old MD5 logs). On the one hand, the new intercepted logs are judged according to the aforementioned steps. When the log needs to be inserted into the circular queue, the log is inserted into the red - black tree of the tail node and stored in chronological order. On the other hand, the system starts an independent timer to trigger queue rotation at a fixed period, inserts a new node at the tail, and removes the earliest node that exceeds the time window period. In this embodiment, it is inserted from the head node and popped from the tail node. That is, the retained feature metadata group is inserted from the tail node, and the old set of feature metadata groups is popped from the head node. In this way, the capacity of the circular queue is always kept within the capacity value C.
[0051] For example, during traversal, since a node is inserted at the tail of the circular queue, then a node is popped from the head of the circular queue, always keeping only 3 nodes at a certain moment, and all logs within 1 s are stored under each node. In this way, the queue capacity is set within 3, and for the intercepted logs obtained at a certain moment, it is only necessary to compare whether there is the same MD5 value among the MD5 logs within 0 - 1 s, 1 - 2 s, and 2 - 3 s.
[0052] The term "anchoring" is used to describe a special processing method for duplicate logs, that is, fixing the duplicate feature metadata group and its corresponding interception log in the current circular queue node, so that it remains in the storage state during the subsequent sliding of the time window until the preset discard condition is met.
[0053] S3, insert the anchored feature metadata group into the tail node of the circular queue according to the preset first data structure.
[0054] Specifically, when new logs arrive, a query operation is triggered to query the circular queue node. If there is no duplicate feature identifier in this node (red - black tree), the intercepted log is reported and inserted into the set; if there is a duplicate feature identifier, the target intercepted log is discarded.
[0055] S4. Repeat steps S1 - S3 to continuously process the new intercepted log set.
[0056] When a new intercepted log set is obtained, repeat steps S1 - S3 to traverse the circular queue. First, check whether the first C - 1 nodes have the same MD5 value. If so, discard the target intercepted log. If not, then check the C - th node. When the same MD5 value is found, it is necessary to calculate the difference in time - sequence identifiers by calculating the difference in timestamps, and compare the difference in time - sequence identifiers with the time - window period value. If the difference in time - sequence identifiers is greater than the time - window period value, the target intercepted log and its MD5 log are retained; if the difference in time - sequence identifiers is less than or equal to the time - window period value, the target intercepted log and its MD5 log are discarded.
[0057] The following further illustrates the establishment process and traversal process of the circular queue through the accompanying drawings and embodiments. Taking the time - window period value as 2s, the time - slice span value as 1s, and the capacity C of the circular queue as 3 as an example, continue the description. In the time window of 0 - 1s, at time t1, a feature metadata group (MD5 - 1, (t1, cnt1)) is generated, at time t2, a feature metadata group (MD5 - 2, (t2, cnt2)) is generated, and at time t3, a feature metadata group (MD5 - 3, (t3, cnt3)) is generated. Then the generated red - black tree is as Figure 2 shown.
[0058] Within the time window of 0 - 2s, query the elements of the circular queue in the order from near to far (the queue is from right to left). If there is a duplicate feature identifier in the red - black tree, discard it; if there is no duplicate feature identifier in the red - black tree, report it and insert it into the nearest set tree at the end of the queue. As Figure 3 shown, in the time period of 1 - 2s, a feature metadata group (MD5 - 1, (t1, cnt1)), a feature metadata group (MD5 - 2, (t2, cnt2)), and a feature metadata group (MD5 - 3, (t3, cnt3)) are generated. In the time period of 0 - 1s, a feature metadata group (MD5 - 4, (t4, cnt4)) and a feature metadata group (MD5 - 5, (t5, cnt5)) are generated. At the current time, the feature metadata group of the target intercepted log is (MD5 - 1, (t1’, cnt1’)). Then when traversing the circular queue, if the feature metadata group (MD5 - 1, (t1’, cnt1’)) is found within the 1 - 2s time window, the target intercepted log and the feature metadata group (MD5 - 1, (t1’, cnt1’)) are discarded.
[0059] When the interception log at the 3rd second arrives, traverse the circular queue and query sequentially from the nearest to the farthest (from right to left in the queue). Among the C - 1 elements, if it exists, discard it; if it does not exist, insert it into the rightmost element set. After finding it in the Cth element, compare whether the difference in time sequence identifiers (the difference between the time identifier in the target feature metadata group and the time identifier of the duplicate log) exceeds the time window of 2 seconds. If it exceeds, it can be inserted into the 0 - 1 second elements; otherwise, discard it. For example, as Figure 4 shown, at the 2nd - 3rd second, the feature metadata groups (MD5 - 1, (t1, cnt1)), (MD5 - 2, (t2, cnt2)), and (MD5 - 3, (t3, cnt3)) are generated, and at the 1st - 2nd second, the feature metadata groups (MD5 - 4, (t4, cnt4)) and (MD5 - 5, (t5, cnt5)) are generated. Adding the current time, there are the feature metadata groups MD5 - 1 and MD5 - 8. There are duplicate logs between them, and the difference between the time stamps t” of the two and the time stamp t’ of the duplicate log is greater than 2 seconds. Then, the feature metadata groups MD5 - 1 and MD5 - 8 will be inserted at the 0 - 1 second.
[0060] Among them, the queue from right to left means querying from the right (the direction of entering the queue) to the left (the direction of leaving the queue) of the circular queue. The rightmost is always the MD5 log generated within the nearest 1 second, and going left is the MD5 log of the second - nearest 1 second, and the leftmost is the MD5 log of the farthest 1 second. If a duplicate feature identifier is found within 1 second, the current interception log will not be recorded. If no duplicate feature identifier is found, continue to traverse the next 1 second (the next node). Note that under each node is a red - black tree, and the logs generated within the current time window (1 second).
[0061] In an alternative embodiment, the target feature metadata group further includes the number of duplicate logs.
[0062] For example, in (MD5 - 1, (t1, cnt1)), the cnt1 attribute is used to represent the number of duplicate logs, that is, one log corresponds to one cnt, and the cnt of MD5 - 1 is cnt1. The cnt in the subsequent feature metadata groups all represents the number of logs repeatedly hit in this time window. After adding this attribute, the ability to accurately record persistent attack behaviors can be achieved. For all the above situations where logs need to be recorded for hits, perform one more action: increment the count by 1, and for newly inserted data, the initial value of cnt is 1.
[0063] For example, Figure 5In step S234, since there are duplicate feature identifiers, a step of "hit count is 1" is added in step S234 (for the newly inserted data, the initial value of cnt is 1). After step S24 (before step S241), a new step is added: update the log hit count by adding 1, that is, increment the count of cnt1 by 1.
[0064] In this embodiment, a circular queue is used to deduplicate log data, which solves the merging of the same audit or interception logs in the terminal protection system under the sliding time window and the trouble caused by the possible overlap of fixed window logs.
[0065] See Figure 7 , an interception log deduplication device for a terminal protection system provided by an embodiment of the present invention includes a data acquisition module 100, a traversal module 200, a queue update module 300, and a duplication module 400. The data acquisition module 100 is used to acquire an interception log set, determine the feature metadata group of each interception log in the interception log set, and obtain a feature metadata group set corresponding to the interception log set; wherein, the feature metadata group includes a feature identifier and a time sequence identifier of the interception log. The traversal module 200 is used to traverse each feature metadata group in the feature metadata group set, starting from the tail node of the circular queue in the order of time slices, and judge whether there is a duplicate feature identifier in the historical feature metadata group set of each node in the circular queue; if there is a duplicate feature identifier, determine whether to discard or anchor the target feature metadata group and the target interception log corresponding to the duplicate feature identifier according to whether the difference between the position of the traversed node and the time sequence identifier exceeds the time window period. If no duplicate feature identifier is found, anchor the target feature metadata group and the target interception log; wherein, the time sequence identifier difference is the difference between the time sequence identifier in the target feature metadata group and the time sequence identifier corresponding to the duplicate feature identifier in the historical feature metadata group set, and each node in the circular queue stores the historical feature metadata group set with a preset time slice span according to a preset first data structure. The queue update module 300 is used to insert the anchored target feature metadata group into the tail node of the circular queue according to the preset first data structure. The duplication module 400 is used to repeat the data acquisition module, the traversal module, and the queue update module to continuously process the new interception log set.
[0066] In an optional embodiment, the capacity of the circular queue C is: (1) In formula (1), C represents the number of nodes in the circular queue; represents the time slice span value, represents the time window period value; wherein, is an integer multiple of.
[0067] In an alternative embodiment, the traversal module 200 includes a feature identifier comparison module, a first judgment module, a second judgment module, and a third judgment module. The feature identifier comparison module is used to compare the feature identifier of the target interception log with the historical feature identifiers in the historical feature metadata set of the current node to determine whether there are duplicate feature identifiers. The first judgment module is used to judge whether the traversal loop queue ends. If the traversal has not ended, it judges whether it has traversed to the historical feature metadata set of the C-th node. The second judgment module is used to, if it has traversed to the historical feature metadata set of the C-th node, judge whether there are duplicate feature identifiers. If there are, it judges whether the time sequence identifier difference exceeds a preset time window period; if it does not exceed the preset time window period, it discards the target feature metadata group and the target log; if it exceeds the preset time window period, it anchors the target feature metadata group and the target log corresponding to the duplicate feature identifier. The third judgment module is used to, if it has traversed to the historical feature metadata set of a node other than the C-th node, judge whether there are duplicate feature identifiers; if there are duplicate feature identifiers, it discards the target feature metadata group and the target log; if there are no duplicate feature identifiers, it updates the existing node and returns to the feature identifier comparison module to continue the traversal.
[0068] In an alternative embodiment, the feature identifier is generated by using a hash algorithm and based on the path, creation time, last modification time, and file size of the interception log.
[0069] In an alternative embodiment, the hash algorithm is the MD5 algorithm.
[0070] In an alternative embodiment, the time sequence identifier is the timestamp when the interception log is intercepted.
[0071] In an alternative embodiment, the first data structure is one of a red-black tree or a hash table.
[0072] In an alternative embodiment, the target feature metadata group further includes the number of duplicate logs.
[0073] By using the device provided in the embodiment of the present application, since the device has the same inventive concept as the above method provided in the embodiment of the present application, on the premise that the method can solve the technical problem, the device can also solve the technical problem, which will not be elaborated here.
[0074] Refer to Figure 8, an embodiment of the present invention further provides an electronic device 1000, including a communication interface 1001, a processor 1002, a memory 1003, and a bus 1004. The processor 1002, the communication interface 1001, and the memory 1003 are connected through the bus 1004. The memory 1003 is configured to store a computer program for supporting the processor 1002 to execute the method for deduplicating interception logs of the above terminal protection system, and the processor 1002 is configured to execute the program stored in the memory 1003.
[0075] Optionally, an embodiment of the present invention further provides a computer-readable medium having non-volatile program code executable by a processor 1002, and the program code causes the processor 1002 to execute the method for deduplicating interception logs of the terminal protection system in the above embodiment.
[0076] As is known by common technical knowledge, the present invention can be implemented by other embodiments without departing from its spiritual essence or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.
Claims
1. A method for deduplicating interception logs of a terminal protection system, characterized in that, Including: S1. Obtain an interception log set, determine the characteristic metadata group of each interception log in the interception log set, and obtain a characteristic metadata group set corresponding to the interception log set; wherein, the characteristic metadata group includes a characteristic identifier and a time sequence identifier of the interception log; S2. For each characteristic metadata group in the characteristic metadata group set, traverse the historical characteristic metadata group set in each node of the circular queue in the order of time slices starting from the tail node, and determine whether there is a duplicate characteristic identifier; if there is a duplicate characteristic identifier, determine whether to discard or anchor the target characteristic metadata group and the target interception log corresponding to the duplicate characteristic identifier according to whether the difference between the traversed node position and the time sequence identifier exceeds the time window period. If no duplicate characteristic identifier is found, anchor the target characteristic metadata group and the target interception log; wherein, the time sequence identifier difference is the difference between the time sequence identifier in the target characteristic metadata group and the time sequence identifier corresponding to the duplicate characteristic identifier in the historical characteristic metadata group set, and each node in the circular queue stores the historical characteristic metadata group set with a preset time slice span according to a preset first data structure; S3. Insert the anchored target characteristic metadata group into the tail node of the circular queue according to the preset first data structure; S4. Repeat steps S1 - S3 to continuously process the new interception log set.
2. The method for deduplicating interception logs of the terminal protection system according to claim 1, wherein The capacity of the circular queue C is as follows: (1) In formula (1), C represents the number of nodes in the circular queue; represents the time slice span value, represents the time window period value; where is an integer multiple of.
3. The deduplication method for the interception log of the terminal protection system according to claim 2, characterized in that, The S2 includes: S21. Compare the characteristic identifier of the target interception log with the historical characteristic identifiers in the historical characteristic metadata group set in the current node to determine whether there is a duplicate characteristic identifier; S22. Determine whether the traversal of the circular queue ends. If the traversal has not ended, determine whether it has traversed to the historical characteristic metadata group set of the Cth node; S23. If it has traversed to the historical characteristic metadata group set of the Cth node, determine whether there is a duplicate characteristic identifier. If there is, determine whether the time sequence identifier difference exceeds the preset time window period; if it does not exceed the preset time window period, discard the target characteristic metadata group and the target log; if it exceeds the preset time window period, anchor the target characteristic metadata group and the target log corresponding to the duplicate characteristic identifier; S24. If it has traversed to the historical characteristic metadata group set of a non - Cth node, determine whether there is a duplicate characteristic identifier; if there is a duplicate characteristic identifier, discard the target characteristic metadata group and the target log; if there is no duplicate characteristic identifier, update the existing node and return to step S21 to continue traversing.
4. The deduplication method for the interception logs of the terminal protection system according to claim 1, characterized in that, The characteristic identifier is generated by using a hash algorithm based on the path, creation time, last modification time, and file size of the interception log.
5. The deduplication method for the interception logs of the terminal protection system according to claim 4, characterized in that, The hash algorithm is the MD5 algorithm.
6. The deduplication method for the interception log of the terminal protection system according to claim 1, characterized in that, The time sequence identifier is the timestamp when the interception log is intercepted.
7. The deduplication method for the interception logs of the terminal protection system according to claim 1, wherein The first data structure is one of a red - black tree or a hash table.
8. The duplicate removal method for the interception log of the terminal protection system according to claim 1, characterized in that The target characteristic metadata group further includes the number of duplicate logs.
9. A de-duplication device for interception logs of a terminal protection system, characterized in that, Including: A data acquisition module, configured to obtain an interception log set, determine the characteristic metadata group of each interception log in the interception log set, and obtain a characteristic metadata group set corresponding to the interception log set; wherein, the characteristic metadata group includes a characteristic identifier and a time sequence identifier of the interception log; A traversal module, configured to traverse, for each feature metadata group in the feature metadata group set, the historical feature metadata group set in each node of the circular queue in the order of time slices starting from the tail node, and determine whether there is a duplicate feature identifier; if there is a duplicate feature identifier, determine whether to discard or anchor the target feature metadata group and the target interception log corresponding to the duplicate feature identifier according to whether the difference between the traversed node position and the timing identifier exceeds the time window period. If no duplicate feature identifier is found, anchor the target feature metadata group and the target interception log; wherein, the timing identifier difference is the difference between the timing identifier in the target feature metadata group and the timing identifier corresponding to the duplicate feature identifier in the historical feature metadata group set, and each node in the circular queue stores the historical feature metadata group set with a preset time slice span according to a preset first data structure; A queue update module, configured to insert the anchored target feature metadata group into the tail node of the circular queue according to a preset first data structure; A repetition module, configured to repeatedly execute the data acquisition module, the traversal module, and the queue update module to continuously process a new interception log set.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Method, system and device for preventing log storm on DCS controller and storage medium
CN117056188A
Real time searching and reporting
US20120197928A1