Method, device, equipment and medium for constructing time series graph

By constructing and analyzing the timing map of traffic data information, the problem of the inability to detect complex network attacks in the prior art is solved, and higher security map accuracy and detection of latent attack behavior are achieved.

CN114547491BActive Publication Date: 2025-05-16BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210204857.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-03
Publication Date
2025-05-16
Estimated Expiration
2042-03-03

AI Technical Summary

Technical Problem

In the prior art, probe devices cannot detect high-complexity network attacks, resulting in insufficient accuracy of the security map.

Method used

By obtaining the traffic data information within the preset time period, the original sequence is constructed, and sampling it is obtained multiple subsequences. Based on the distance distribution between each subsequence and the original sequence, the information gain of each subsequence is determined, and the target sequence is obtained from it after comparison. Use the target sequence as the node of the timing graph and use its timing relationship in the original sequence as the edge to build the timing graph.

Benefits of technology

Improves the accuracy of the timing map, can detect network attacks that cannot be detected by device probes, and solves long-term latent attack behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547491B_ABST
    Figure CN114547491B_ABST
Patent Text Reader

Abstract

The disclosed embodiments relate to a method, device, equipment and medium for constructing a time series graph, wherein the method comprises: obtaining traffic data information within a preset time period; constructing an original sequence based on the traffic data information; sampling and processing the original sequence to obtain multiple subsequences, and determining the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence; comparing the information gains of each subsequence, and obtaining a target sequence from multiple subsequences based on the comparison results; constructing a time series graph by using the target sequence as a node of the time series graph, and using the time series relationship of the target sequence in the same original sequence as an edge of the time series graph. In the disclosed embodiments, the network security attack or access path of the time series is converted into the expression form of the time series graph, and the network attack that cannot be detected by the device probe can be fed back on the time series graph, thereby improving the accuracy of the time series graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, equipment and medium for constructing a time series graph. Background Art

[0002] With the development of computer technology, network security has become increasingly important. Security graphs can intuitively express data and behaviors related to network security.

[0003] In the related technology, the network can be security checked through probe devices. If an abnormal situation is detected, the probe device will issue an alarm message, parse the alarm message, extract the source Internet Protocol IP address, destination Internet Protocol IP address and event, and build a security map based on the above triples.

[0004] However, for the above technical solution, some highly complex network attacks can bypass the detection of the probe device, resulting in the probe device being unable to issue an alarm, thus causing the security map to be insufficiently accurate. Summary of the invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, device, equipment and medium for constructing a time series graph.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for constructing a time series graph, the method comprising:

[0007] Obtain traffic data information within a preset time period;

[0008] Based on the traffic data information, construct an original sequence;

[0009] Sampling the original sequence to obtain a plurality of subsequences, and determining the information gain of each subsequence based on a distance distribution between each subsequence and the original sequence;

[0010] Comparing the information gains of the subsequences, and obtaining a target sequence from the multiple subsequences based on the comparison results;

[0011] The target sequence is used as a node of a time series graph, and the time series relationship of the target sequence in the same original sequence is used as an edge of the time series graph to construct the time series graph.

[0012] In an optional implementation manner, constructing the original sequence based on the traffic data information includes:

[0013] Parsing the traffic data information to obtain the number of accesses of the access terminal in each sub-time period of the preset time period;

[0014] Based on the number of accesses and the timing relationship between the sub-time periods, an original sequence corresponding to each of the access ends is constructed.

[0015] In an optional implementation manner, the sampling process of the original sequence to obtain a plurality of subsequences includes:

[0016] The original sequence is sampled according to a preset sliding window and a preset sliding distance to obtain the multiple subsequences.

[0017] In an optional implementation, determining the information gain of each of the subsequences based on the distance distribution between each of the subsequences and the original sequence includes:

[0018] constructing a first sequence set based on the subsequences, and calculating a first distance distribution from the first sequence set to each of the original sequences;

[0019] Eliminate the subsequence currently being processed in the first sequence set to obtain a second sequence set, and calculate a second distance distribution from the second sequence set to each of the original sequences;

[0020] The information gain of the subsequence currently being processed is determined according to the first distance distribution and the second distance distribution.

[0021] In an optional implementation manner, the calculating a first distance distribution from the first sequence set to each of the original sequences includes:

[0022] Acquire a first sampling subsequence determined by sampling the original sequence currently being processed;

[0023] Calculating a first distance between each of the first sampling subsequences and a subsequence in the first sequence set, and taking a minimum value of the first distances as a first target distance between a subsequence in the first sequence set and the original sequence currently being processed;

[0024] calculating the first target distance between each of the original sequences and each of the subsequences in the first sequence set, and determining the first distance distribution according to the first target distance;

[0025] The calculating a second distance distribution from the second sequence set to each of the original sequences comprises:

[0026] Acquire a second sampling subsequence determined by sampling the original sequence currently being processed;

[0027] Calculating a second distance between each of the second sampling subsequences and a subsequence in the second sequence set, and taking a minimum value of the second distances as a second target distance between the subsequence in the second sequence set and the original sequence currently being processed;

[0028] The second target distance between each of the original sequences and each of the subsequences in the second sequence set is calculated, and the second distance distribution is determined according to the second target distance.

[0029] In an optional implementation, the target sequence is used as a node of a time series graph, and the time series relationship of the target sequence in the same original sequence is used as an edge of the time series graph to construct the time series graph, including:

[0030] Obtaining a homologous sequence in the target sequence that belongs to the same original sequence;

[0031] Searching the homologous sequences in the original sequence to obtain a time sequence identifier corresponding to each homologous sequence;

[0032] The homologous sequences are used as nodes of the time sequence map, and the connection relationship between the homologous sequences is determined according to the time sequence identifier corresponding to each of the homologous sequences to construct the time sequence map.

[0033] In a second aspect, an embodiment of the present disclosure further provides a device for constructing a time series graph, the device comprising:

[0034] An acquisition module is used to obtain flow data information within a preset time period;

[0035] A first construction module, used to construct an original sequence based on the traffic data information;

[0036] A determination module, configured to perform sampling processing on the original sequence to obtain a plurality of subsequences, and determine the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence;

[0037] A comparison module, used for comparing the information gains of the subsequences, and obtaining a target sequence from the multiple subsequences based on the comparison result;

[0038] The second construction module is used to construct the time series graph by taking the target sequence as a node of the time series graph and taking the time series relationship of the target sequence in the same original sequence as an edge of the time series graph.

[0039] In a third aspect, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device implements the above method.

[0040] In a fourth aspect, the present disclosure provides a device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method when executing the computer program.

[0041] In a fifth aspect, the present disclosure provides a computer program product, wherein the computer program product comprises a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.

[0042] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:

[0043] The method for constructing a time series graph of the embodiment of the present disclosure obtains traffic data information within a preset time period; constructs an original sequence based on the traffic data information; samples the original sequence to obtain multiple subsequences, and determines the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence; compares the information gains of each subsequence, and obtains a target sequence from multiple subsequences based on the comparison results; uses the target sequence as a node of the time series graph, and uses the time series relationship of the target sequence in the same original sequence as an edge of the time series graph to construct a time series graph. It can be seen that the embodiment of the present disclosure can extract the pattern of traffic data information within a preset time period, and obtain a time series graph in combination with the time series characteristics, convert the network security attack or access path of the time series into the expression form of the time series graph, construct a time series behavior on the time series graph, and can feedback the network attack that the device probe cannot detect on the time series graph, thereby improving the accuracy of the time series graph, and better utilizing the information contained in the traffic data information, and can better solve the latent attack behavior of a long period of time. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0045] Figure 1 A schematic diagram of a flow chart of a method for constructing a time series graph provided in an embodiment of the present disclosure;

[0046] Figure 2 A schematic diagram of a flow chart of another method for constructing a timing graph provided in an embodiment of the present disclosure;

[0047] Figure 3 A schematic diagram of a method for determining a timing identifier provided by an embodiment of the present disclosure;

[0048] Figure 4 A schematic diagram of the structure of a timing graph construction device provided in an embodiment of the present disclosure;

[0049] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0050] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0051] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0052] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0053] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0054] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0055] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0056] In order to solve the above problems, the embodiments of the present disclosure provide a method for constructing a timing graph, which is introduced below in conjunction with specific embodiments.

[0057] Figure 1The present invention provides a flow chart of a method for constructing a time sequence map, which can be executed by a device for constructing a time sequence map, wherein the device can be implemented by software and / or hardware, and can generally be integrated in an electronic device. Figure 1 As shown, the method includes:

[0058] Step 101, obtaining flow data information within a preset time period.

[0059] In the field of network security, especially in the intranet access environment, most access modes follow certain timing rules. Such timing rules can be displayed in the form of timing graphs.

[0060] Specifically, first, the flow data information within a preset time period is obtained. In this embodiment, the preset time period can be set according to the application scenario, and this embodiment is not limited. The preset time period can be a longer time period, such as half a year or a year. The flow data information records the number of access initiations of one or more access terminals within the time period. The flow data information can be a network communication log such as a netflow protocol log or an http log.

[0061] Step 102: construct an original sequence based on the traffic data information.

[0062] After obtaining the flow data information, the flow data information can be parsed to construct the original sequence. There are many methods for constructing the original sequence, which can be selected according to the application scenario. This embodiment does not limit it. The example is described as follows:

[0063] In an optional implementation manner, dividing the preset time period into a plurality of sub-time periods, and constructing the original sequence based on the sub-time periods specifically includes:

[0064] First, the traffic data information is parsed to obtain the number of accesses of the access terminal in each sub-time period of the preset time period.

[0065] The flow data information is parsed to obtain information such as access time and source Internet Protocol (IP) address in the flow information, and then the number of accesses of the access end in each sub-time period is obtained based on the information. The access end can be determined according to the application scenario, and this embodiment does not limit it. For example, the access end can be the source IP address.

[0066] It should be noted that the length of the above-mentioned sub-time period can be determined based on the length of the preset time period. In an optional implementation, the length of the sub-time period can be set to be relatively short. Specifically, a proportional threshold can be preset, and the ratio of the length of the preset time period to the length of the sub-time period must be greater than the proportional threshold. For example, the proportional threshold can be 150. If the preset time period is half a year, the sub-time period can be 1 day.

[0067] Furthermore, based on the number of accesses and the temporal relationship between each sub-time period, an original sequence corresponding to each access end is constructed.

[0068] After determining the number of accesses in each sub-time period, the number of accesses of the same access terminal in multiple sub-time periods may be extracted, and the multiple number of accesses may be sorted according to the time sequence relationship between each sub-time period to obtain the corresponding original sequence.

[0069] For example, if we parse the traffic data information and count the number of visits initiated by IP1 every day within 10 days, and the statistical results are in chronological order: 11, 2, 13, 5, 8, 6, 3, 10, 9, 4. Then the original sequence constructed is {11, 2, 13, 5, 8, 6, 3, 10, 9, 4}

[0070] In another optional implementation, the source IP address, destination Internet Protocol (IP) address and time in the traffic data information can be parsed, and the original sequence can be constructed based on the number of accesses initiated from the target source IP address to the destination IP address in each sub-time period of a preset time period.

[0071] Step 103: Sampling the original sequence to obtain multiple subsequences, and determining the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence.

[0072] After obtaining the original sequence, each original sequence is sampled to obtain multiple subsequences. In an optional implementation, the original sequence can be sampled by a preset sliding window and sliding distance to obtain multiple subsequences. It should be noted that the multiple subsequences can be sampled from different original sequences. The length of the sliding window and the sliding distance can be set according to the application scenario. For example, IP1{11,2,13,5,8,6,3,10,9,4} can be sampled with a sliding window of 7 and a sliding distance of 1, and the four subsequences obtained are: {11,2,13,5,8,6,3}, {2,13,5,8,6,3,10}, {13,5,8,6,3,10,9}, {5,8,6,3,10,9,4}. Optionally, the multiple subsequences obtained by sampling can also be deduplicated.

[0073] In this embodiment, the distance between the subsequence and the original sequence is calculated. The specific calculation method is not limited in this embodiment, and those skilled in the art can make corresponding settings according to scenario requirements. The distance between each subsequence and each original sequence can be determined by calculation, and a distance distribution can be constructed based on the distance, so as to determine the information gain of each subsequence according to the distance distribution.

[0074] In an optional implementation, the distance distribution including the currently processed subsequence and the distance distribution excluding the currently processed subsequence can be calculated separately, so as to determine the information gain of the currently processed subsequence based on the information difference between the two distance distributions, and each subsequence is processed to obtain the information gain corresponding to each subsequence.

[0075] Step 104: compare the information gains of the subsequences, and obtain a target sequence from the multiple subsequences based on the comparison results.

[0076] Furthermore, the information gain of each subsequence can be numerically compared. If the numerical value is larger, it means that the amount of information carried by the subsequence is larger and the representativeness of the subsequence is stronger. Therefore, the target sequence can be obtained from multiple subsequences based on the comparison result.

[0077] In an optional implementation, an integer N may be preset, and the subsequences may be sorted from large to small according to information gain, and the first N subsequences may be obtained as the target sequence.

[0078] Step 105, using the target sequence as a node of the time series graph, and using the time series relationship of the target sequence in the same original sequence as an edge of the time series graph, to construct a time series graph.

[0079] It can be understood that the graph is composed of nodes and edges. In this embodiment, a time series graph can be constructed based on the target sequence obtained by screening. Specifically, the target sequence can be used as a node of the time series graph, and the target sequence can be matched on the same original sequence. The successfully matched target sequences are sorted according to the time sequence, and edges are established between two adjacent target sequences after sorting, thereby determining the edges of the time series graph and then constructing the time series graph. In an optional implementation, the edge between adjacent target sequences can be a directed line pointing from a target sequence with an earlier time sequence to a target sequence with a later time sequence.

[0080] Taking the original sequence corresponding to IP1 as {11, 2, 13, 5, 8, 6, 3, 10, 9, 4} as an example, assuming that the first target sequence is {11, 2, 13, 5, 8, 6, 3} and the second target sequence is {13, 5, 8, 6, 3, 10, 9}, it can be determined that on IP1, the timing of the first target sequence is earlier than the timing of the second target sequence. The first target sequence and the second target sequence can be used as two nodes of the timing graph, and a line from the first target sequence to the second target sequence can be established as an edge of the timing graph.

[0081] In summary, the method for constructing a time series graph in the embodiment of the present disclosure obtains traffic data information within a preset time period; constructs an original sequence based on the traffic data information; samples the original sequence to obtain multiple subsequences, and determines the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence; compares the information gains of each subsequence, and obtains a target sequence from multiple subsequences based on the comparison results; uses the target sequence as a node of the time series graph, and uses the time series relationship of the target sequence in the same original sequence as the edge of the time series graph to construct a time series graph. It can be seen that the embodiment of the present disclosure can extract the pattern of traffic data information within a preset time period, and obtain a time series graph by combining the time series characteristics, convert the network security attack or access path of the time series into the expression form of the time series graph, construct a time series behavior on the time series graph, and can feedback the network attack that the device probe cannot detect on the time series graph, thereby improving the accuracy of the time series graph, and better utilizing the information contained in the traffic data information, and can better solve the latent attack behavior of a long period of time. In addition, it can be applied to scenarios such as big data security analysis or threat hunting.

[0082] Based on the above embodiments, Figure 2 A flow chart of another method for constructing a timing graph provided in an embodiment of the present disclosure is shown in FIG. Figure 2 As shown, the information gain of each subsequence is determined based on the distance distribution between each subsequence and the original sequence, including the following steps:

[0083] Step 201, obtaining flow data information within a preset time period.

[0084] Step 202: construct an original sequence based on the traffic data information.

[0085] In this embodiment, the information that can be obtained by parsing the traffic data information includes but is not limited to: one or more of source IP address information, destination IP address information, port information, protocol information, traffic size information, and time information, and then the original sequence is constructed based on the information obtained from the above analysis.

[0086] Step 203: Sampling the original sequence to obtain multiple subsequences, constructing a first sequence set based on the subsequences, and calculating a first distance distribution from the first sequence set to each original sequence.

[0087] In this embodiment, the first distance distribution is a distance distribution including the currently processed subsequence, and the second distance distribution is a distance distribution excluding the currently processed subsequence. Therefore, the information gain of the currently processed subsequence can be determined according to the change in the amount of information between the first distance distribution and the second distance distribution, specifically including:

[0088] In this embodiment, the first sequence set includes subsequences obtained by sampling each original sequence, and a first distance distribution from the first sequence set to each original sequence is determined by calculating the distance from each subsequence in the first sequence set to each original sequence. The first distance distribution can reflect the distribution of distances between the subsequences in the first sequence set and the original sequences.

[0089] In an optional implementation manner, the calculating of the first distance distribution from the first sequence set to each original sequence specifically includes:

[0090] Step a1: obtaining a first sampling subsequence determined by sampling the original sequence currently being processed.

[0091] In this embodiment, the first sampling subsequence is a subsequence included in the original sequence currently being processed. In an optional implementation, the correspondence between the original sequence and the subsequence can be recorded in a preset list, and the preset list can be searched according to the original sequence currently being processed to determine the first sampling subsequence corresponding to the original sequence currently being processed. In another optional implementation, the original sequence currently being processed can be sampled according to a preset sliding window and a preset sliding distance to obtain the corresponding first sampling sequence.

[0092] Step a2: calculating a first distance between each first sampling subsequence and a subsequence in the first sequence set, and taking a minimum value of the first distances as a first target distance between a subsequence in the first sequence set and a currently processed original sequence.

[0093] In this embodiment, the first distance is: the Euclidean distance between the first sampling subsequence and the subsequence in the first sequence set. Multiple first sampling subsequences included in the currently processed original sequence are calculated to obtain the first distance corresponding to each first sampling subsequence, and the minimum value of the multiple first distances is used as the first target distance between the currently processed original sequence and the subsequence in the first sequence set.

[0094] For example, if there are N source IP addresses, then there are N elements in the original sequence set, and the original sequence set OrgS = {OrgS i ,i∈[1,2,…N]}, if the original sequence currently being processed is OrgS i , the first sampling subsequence is SubS k , then there is the first sampling subsequence SubS k ∈OrgS i , the subsequence in the first sequence set is SubS j , SubS k With SubS j The first distance is d(SubS k ,SubS j ),but:

[0095]

[0096] Among them, SubS k,r SubS k The rth value of SubS j,r SubS j The rth value of OrgS i With SubS j The first target distance is d(OrgS i ,SubS j ),but:

[0097]

[0098] Step a3: calculating a first target distance between each original sequence and each subsequence in the first sequence set, and determining a first distance distribution according to the first target distance.

[0099] In an optional implementation, a first target distance between the original sequence currently being processed and each subsequence in the first sequence set is calculated, and the calculation is performed for each original sequence to obtain the first target distance between each original sequence and each subsequence in the first sequence set, thereby forming a first distance distribution based on the multiple first target distances.

[0100] Step 204: remove the currently processed subsequence in the first sequence set to obtain a second sequence set, and calculate a second distance distribution from the second sequence set to each original sequence.

[0101] In this embodiment, the subsequence currently being processed is the subsequence currently being calculated for information gain, and the subsequence currently being processed may be removed from the first sequence set to obtain a second sequence set, and the second distance distribution from the second sequence set to each original sequence is determined by calculating the distance from each subsequence in the second sequence set to each original sequence. The second distance distribution can reflect the distribution of the distances between the subsequences in the second sequence set and the original sequences.

[0102] In an optional implementation manner, the calculating of the second distance distribution from the second sequence set to each original sequence specifically includes:

[0103] Step b1, obtaining a second sampling subsequence determined by sampling the original sequence currently being processed.

[0104] In this embodiment, the second sampling subsequence is a subsequence included in the original sequence currently being processed. In an optional implementation, the correspondence between the original sequence and the subsequence can be recorded in a preset list, and the preset list can be searched according to the original sequence currently being processed to determine the second sampling subsequence corresponding to the original sequence currently being processed. In another optional implementation, the original sequence currently being processed can be sampled according to a preset sliding window and a preset sliding distance to obtain the corresponding second sampling sequence.

[0105] Step b2, calculating the second distance between each second sampling subsequence and the subsequence in the second sequence set, and taking the minimum value of the second distances as the second target distance between the subsequence in the second sequence set and the original sequence currently being processed.

[0106] In this embodiment, the second distance is: the Euclidean distance between the second sampling subsequence and the subsequence in the second sequence set. The multiple second sampling subsequences included in the currently processed original sequence are calculated to obtain the second distance corresponding to each second sampling subsequence, and the minimum value of the multiple second distances is used as the second target distance between the currently processed original sequence and the subsequence in the second sequence set.

[0107] For example, if the original sequence currently being processed is OrgS i ', the second sampling subsequence is SubS k ′, then there is a second sampling subsequence SubS k ′∈OrgS i ', the subsequence in the second sequence set is SubS j ′,SubS k ′ and SubS j The second distance of ' is d(SubS k ′,SubS j '),but:

[0108]

[0109] Among them, SubS k,r ′ indicates SubS k ′, the rth value, SubS j,r ′ indicates SubS j ′ is the rth value. i ′ and SubS j The second target distance of ′ is b(OrgS i ′,SubS j '),but:

[0110]

[0111] Step b3: calculating a second target distance between each original sequence and each subsequence in the second sequence set, and determining a second distance distribution according to the second target distance.

[0112] In an optional implementation, the second target distance between the original sequence currently being processed and each subsequence in the second sequence set is calculated, and the calculation is performed for each original sequence to obtain the second target distance between each original sequence and each subsequence in the second sequence set, thereby forming a second distance distribution based on the multiple second target distances.

[0113] Step 205: Determine the information gain of the subsequence currently being processed according to the first distance distribution and the second distance distribution.

[0114] Information gain can be used to measure the difference between different probability distributions, so as to determine the amount of information contained in the corresponding data. In this embodiment, one of the reasons why the first distance distribution is different from the second distance distribution is that the first sequence set corresponding to the first distance distribution includes the subsequence currently being processed, while the second sequence set corresponding to the second distance distribution does not include the subsequence currently being processed. Therefore, the change in the amount of information between the first distance distribution and the second distance distribution can be calculated, and the information gain of the subsequence currently being processed can be determined based on the calculation result.

[0115] Step 206: compare the information gains of the subsequences, and obtain a target sequence from the multiple subsequences based on the comparison results.

[0116] In an optional implementation, a quantity threshold N may be preset, where N is a positive integer, and each subsequence is sorted from large to small according to the value of the information gain, and the first N subsequences are taken as the target sequence.

[0117] In another optional implementation, a quantity threshold M may be preset, where M is a positive integer. By comparing the information gains corresponding to the subsequences, the one with the largest information gain is determined to be the target sequence, and the target sequence is removed from the first sequence set, the first sequence set is updated, the information gains corresponding to the subsequences in the first sequence set are recalculated, the one with the largest information gain is determined to be the target sequence, and the target sequence set is then updated according to the latest target sequence obtained until the number of target sequences obtained is greater than or equal to M.

[0118] Optionally, the obtained target sequence may also be identified by a serial number.

[0119] Step 207, obtaining homologous sequences in the target sequence that belong to the same original sequence.

[0120] In this embodiment, the target sequence can be used to match the original sequence, so as to screen out homologous sequences belonging to the same original sequence from the target sequence.

[0121] In an optional implementation, if the correspondence between the original sequence and the subsequence is pre-recorded, the subsequences in the correspondence can be screened according to the target sequence to determine the homologous sequences belonging to the same original sequence.

[0122] Step 208: search for homologous sequences in the original sequence to obtain a time sequence identifier corresponding to each homologous sequence.

[0123] Furthermore, the homologous sequence can be searched in the original sequence, and the timing mark corresponding to each homologous sequence can be determined according to the position of the homologous sequence in the original sequence. It should be noted that a homologous sequence may be searched multiple times in the same original sequence, so a homologous sequence can correspond to multiple timing marks.

[0124] Step 208: Using homologous sequences as nodes of a time series graph, determining the connection relationship between homologous sequences according to the time series identifier corresponding to each homologous sequence, and constructing a time series graph.

[0125] Each homologous sequence has a corresponding timing identifier, so the connection relationship between homologous sequences can be determined according to the timing identifier. There are many methods for confirming the connection relationship, which are not limited in this embodiment. For example, homologous sequences can be connected in pairs according to the order of the timing identifiers from first to last.

[0126] Figure 3 A schematic diagram of a method for determining a timing mark provided by an embodiment of the present disclosure, such as Figure 3As shown, the original sequence is {11, 2, 13, 5, 8, 6, 3, 10, 9, 4}, the first homologous sequence is {11, 2, 13, 5, 8, 6, 3}, and the second homologous sequence is {13, 5, 8, 6, 3, 10, 9}. In the original sequence, the value at the front corresponds to the front time sequence, and according to the time sequence relationship, the first homologous sequence is ahead of the second homologous sequence, so the time sequence identifier of the first homologous sequence can be recorded as 1, and the time sequence identifier of the second homologous sequence can be recorded as 2, wherein the larger the time sequence identifier, the later the time sequence. The first homologous sequence and the second homologous sequence can be used as nodes of the time sequence map, and a directed connection from the first homologous sequence to the second homologous sequence is established, thereby constructing a time sequence map.

[0127] In summary, the method for constructing a time series graph in the embodiment of the present disclosure can calculate the information gain of each subsequence based on the distance, so as to extract representative target sequences, thereby improving the accuracy of the data source for constructing the graph, and constructing a time series graph according to the time series relationship between the target sequences, so that the time series graph can reflect the time series relationship between the target sequences, thereby enhancing the richness of the information contained in the time series graph.

[0128] Figure 4 This is a schematic diagram of the structure of a timing graph construction device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can generally be integrated in an electronic device. Figure 4 As shown, the device comprises:

[0129] The acquisition module 401 is used to acquire the flow data information within a preset time period;

[0130] A first construction module 402 is used to construct an original sequence based on the traffic data information;

[0131] A determination module 403 is configured to perform sampling processing on the original sequence to obtain multiple subsequences, and determine the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence;

[0132] A comparison module 404 is used to compare the information gains of the subsequences and obtain a target sequence from the multiple subsequences based on the comparison results;

[0133] The second construction module 405 is used to construct the time series graph by taking the target sequence as a node of the time series graph and taking the time series relationship of the target sequence in the same original sequence as an edge of the time series graph.

[0134] Optionally, the first building module 402 is used to:

[0135] Parsing the traffic data information to obtain the number of accesses of the access terminal in each sub-time period of the preset time period;

[0136] Based on the number of accesses and the timing relationship between the sub-time periods, an original sequence corresponding to each of the access ends is constructed.

[0137] Optionally, the determining module 403 includes:

[0138] The sampling unit is used to perform sampling processing on the original sequence according to a preset sliding window and a preset sliding distance to obtain the multiple subsequences.

[0139] Optionally, the determining module 403 includes:

[0140] A first calculation unit, configured to construct a first sequence set based on the subsequences, and calculate a first distance distribution from the first sequence set to each of the original sequences;

[0141] a second calculation unit, configured to remove the subsequence currently being processed in the first sequence set, obtain a second sequence set, and calculate a second distance distribution from the second sequence set to each of the original sequences;

[0142] A determining unit is used to determine the information gain of the subsequence currently being processed according to the first distance distribution and the second distance distribution.

[0143] Optionally, the first computing unit is configured to:

[0144] Acquire a first sampling subsequence determined by sampling the original sequence currently being processed;

[0145] Calculating a first distance between each of the first sampling subsequences and a subsequence in the first sequence set, and taking a minimum value of the first distances as a first target distance between a subsequence in the first sequence set and the original sequence currently being processed;

[0146] calculating the first target distance between each of the original sequences and each of the subsequences in the first sequence set, and determining the first distance distribution according to the first target distance;

[0147] The second computing unit is used for:

[0148] Acquire a second sampling subsequence determined by sampling the original sequence currently being processed;

[0149] Calculating a second distance between each of the second sampling subsequences and a subsequence in the second sequence set, and taking a minimum value of the second distances as a second target distance between the subsequence in the second sequence set and the original sequence currently being processed;

[0150] The second target distance between each of the original sequences and each of the subsequences in the second sequence set is calculated, and the second distance distribution is determined according to the second target distance.

[0151] Optionally, the second building module 405 is used to:

[0152] Obtaining a homologous sequence in the target sequence that belongs to the same original sequence;

[0153] Searching the homologous sequences in the original sequence to obtain a time sequence identifier corresponding to each homologous sequence;

[0154] The homologous sequences are used as nodes of the time sequence map, and the connection relationship between the homologous sequences is determined according to the time sequence identifier corresponding to each of the homologous sequences to construct the time sequence map.

[0155] The timing graph construction device provided in the embodiments of the present disclosure can execute the timing graph construction method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0156] In order to implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program / instruction, which implements the timing graph construction method in the above embodiments when the computer program / instruction is executed by a processor.

[0157] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.

[0158] The following specific reference Figure 5 , which shows a schematic diagram of the structure of an electronic device 500 suitable for implementing the embodiment of the present disclosure. The electronic device 500 in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0159] like Figure 5As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0160] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0161] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the timing diagram construction method of the embodiment of the present disclosure are executed.

[0162] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0163] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0164] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0165] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains traffic data information within a preset time period; constructs an original sequence based on the traffic data information; samples the original sequence to obtain multiple subsequences, and determines the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence; compares the information gains of each subsequence, and obtains a target sequence from the multiple subsequences based on the comparison results; uses the target sequence as a node of a time series graph, and uses the time series relationship of the target sequence in the same original sequence as an edge of the time series graph to construct a time series graph. It can be seen that the embodiment of the present disclosure can extract the traffic data information within a preset time period, and obtain a time series graph by combining the time series characteristics, convert the network security attack or access path of the time series into the expression form of the time series graph, construct a time series behavior on the time series graph, and can feedback the network attack that the device probe cannot detect on the time series graph, thereby improving the accuracy of the time series graph, and better utilizing the information contained in the traffic data information, and can better solve the latent attack behavior of a long period of time.

[0166] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0167] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0168] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.

[0169] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0170] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0171] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0172] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0173] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. A method for constructing a time series graph, characterized in that: include: Acquire flow data information within a preset time period; wherein the flow data information includes the number of accesses of one or more access terminals within the preset time period; Based on the traffic data information, construct an original sequence; Sampling the original sequence to obtain a plurality of subsequences, and determining the information gain of each subsequence based on a distance distribution between each subsequence and the original sequence; Comparing the information gains of the subsequences, and obtaining a target sequence from the multiple subsequences based on the comparison results; The target sequence is used as a node of a time series graph, and the time series relationship of the target sequence in the same original sequence is used as an edge of the time series graph to construct the time series graph; The determining the information gain of each of the subsequences based on the distance distribution between each of the subsequences and the original sequence includes: A first sequence set is constructed based on the subsequences, and a first distance distribution from the first sequence set to each of the original sequences is calculated; wherein the calculating of the first distance distribution from the first sequence set to each of the original sequences comprises: obtaining a first sampling subsequence determined by sampling the original sequence currently being processed; calculating a first distance between each of the first sampling subsequences and a subsequence in the first sequence set, and taking a minimum value of the first distances as a first target distance between a subsequence in the first sequence set and the original sequence currently being processed; calculating the first target distance between each of the original sequences and each of the subsequences in the first sequence set, and determining the first distance distribution according to the first target distance; Eliminate the subsequence currently being processed in the first sequence set to obtain a second sequence set, and calculate a second distance distribution from the second sequence set to each of the original sequences; wherein calculating the second distance distribution from the second sequence set to each of the original sequences includes: obtaining a second sampling subsequence determined by sampling the currently processed original sequence; calculating a second distance between each of the second sampling subsequences and a subsequence in the second sequence set, and taking a minimum value of the second distances as a second target distance between a subsequence in the second sequence set and the currently processed original sequence; calculating the second target distance between each of the original sequences and each of the subsequences in the second sequence set, and determining the second distance distribution according to the second target distance; The information gain of the subsequence currently being processed is determined according to the first distance distribution and the second distance distribution.

2. The method according to claim 1, characterized in that The constructing the original sequence based on the traffic data information includes: Parsing the traffic data information to obtain the number of accesses of the access terminal in each sub-time period of the preset time period; Based on the number of accesses and the timing relationship between the sub-time periods, an original sequence corresponding to each of the access ends is constructed.

3. The method according to claim 1, characterized in that The sampling process of the original sequence to obtain a plurality of subsequences includes: The original sequence is sampled according to a preset sliding window and a preset sliding distance to obtain the multiple subsequences.

4. The method according to claim 1, characterized in that The step of using the target sequence as a node of a time series graph and using the time series relationship of the target sequence in the same original sequence as an edge of the time series graph to construct the time series graph includes: Obtaining a homologous sequence in the target sequence that belongs to the same original sequence; Searching the homologous sequences in the original sequence to obtain a time sequence identifier corresponding to each homologous sequence; The homologous sequences are used as nodes of the time sequence map, and the connection relationship between the homologous sequences is determined according to the time sequence identifier corresponding to each of the homologous sequences to construct the time sequence map.

5. A time series graph construction device, characterized in that: include: An acquisition module, used to acquire flow data information within a preset time period; wherein the flow data information includes the number of accesses of one or more access terminals within the preset time period; A first construction module, used to construct an original sequence based on the traffic data information; A determination module, configured to perform sampling processing on the original sequence to obtain a plurality of subsequences, and determine the information gain of each subsequence based on the distance distribution between each subsequence and the original sequence; A comparison module, used for comparing the information gains of the subsequences, and obtaining a target sequence from the multiple subsequences based on the comparison result; A second construction module is used to construct the time series graph by taking the target sequence as a node of the time series graph and taking the time series relationship of the target sequence in the same original sequence as an edge of the time series graph; The determining the information gain of each of the subsequences based on the distance distribution between each of the subsequences and the original sequence includes: A first sequence set is constructed based on the subsequences, and a first distance distribution from the first sequence set to each of the original sequences is calculated; wherein the calculating of the first distance distribution from the first sequence set to each of the original sequences comprises: obtaining a first sampling subsequence determined by sampling the original sequence currently being processed; calculating a first distance between each of the first sampling subsequences and a subsequence in the first sequence set, and taking a minimum value of the first distances as a first target distance between a subsequence in the first sequence set and the original sequence currently being processed; calculating the first target distance between each of the original sequences and each of the subsequences in the first sequence set, and determining the first distance distribution according to the first target distance; Eliminate the subsequence currently being processed in the first sequence set to obtain a second sequence set, and calculate a second distance distribution from the second sequence set to each of the original sequences; wherein calculating the second distance distribution from the second sequence set to each of the original sequences includes: obtaining a second sampling subsequence determined by sampling the currently processed original sequence; calculating a second distance between each of the second sampling subsequences and a subsequence in the second sequence set, and taking a minimum value of the second distances as a second target distance between a subsequence in the second sequence set and the currently processed original sequence; calculating the second target distance between each of the original sequences and each of the subsequences in the second sequence set, and determining the second distance distribution according to the second target distance; The information gain of the subsequence currently being processed is determined according to the first distance distribution and the second distance distribution.

6. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the timing diagram construction method described in any one of claims 1-4 above.

7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the timing graph construction method described in any one of claims 1 to 4 above.

8. A computer program product, characterized in that The computer program product includes a computer program / instruction, and when the computer program / instruction is executed by a processor, the method for constructing a timing graph as described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Time series early classification method and device based on shapes

    CN110389975A

  • Systems and methods for determining sequence

    US20200082913A1