Flow analysis-fused malicious software detection method, system and device and storage medium
By constructing a multidimensional traffic feature analysis method, the problem of insufficient accuracy and robustness of malware detection in existing technologies is solved, and efficient identification and intelligent detection of diverse malware are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies rely on single-dimensional traffic metrics in malware detection, lacking multi-dimensional feature analysis. This results in insufficient accuracy and robustness of detection results, poor adaptability in dynamic network environments, high false positive and false negative rates, and low levels of intelligence and automation.
By constructing a node distribution structure mapping under time and space coordinates, and integrating the combination rules of time density and spatial clustering difference, abnormal behavior is identified, effective communication path structure is extracted, and a label arrangement path order index is generated to achieve multi-dimensional feature analysis and intelligent detection.
It improves the accuracy and stability of malware detection, enhances the coverage of diverse malware, reduces false positive and false negative rates, and improves detection efficiency and scalability.
Smart Images

Figure CN121750320A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network security, and in particular to a malware detection method and system fusing traffic analysis, a device and a storage medium. BACKGROUND
[0002] The technical field of malware detection includes technologies related to the identification of potential threats, traffic behavior analysis, and security protection in network environments. With the diversification of network attack methods, malware detection methods based on traffic analysis have attracted attention due to their ability to monitor network behavior in real time and detect abnormal activities. The core technology in this field involves the collection and processing of network traffic data, through techniques such as feature extraction, behavior modeling, and pattern recognition, to achieve accurate detection of malware. Detection techniques based on traffic analysis are widely used in network security protection to improve system security and reduce potential risks.
[0003] Among them, the malware detection method and system based on traffic analysis refers to the extraction of key features in traffic through deep analysis of network communication traffic, and the identification of potential malicious behavior by combining machine learning or rule matching techniques. It targets abnormal patterns and malicious software communication features in network traffic, involving steps such as traffic preprocessing, feature selection, and classification. Specifically, through the analysis of traffic data, the communication behavior, data transmission patterns, and abnormal traffic features of malware can be identified to determine whether there is a security threat. This technology combines multidimensional data analysis and intelligent algorithms to cope with complex network attack scenarios to some extent.
[0004] Existing technologies rely on single-dimensional traffic indicators such as entropy rate or specific behavior patterns for feature extraction, lacking comprehensive analysis capabilities for multi-dimensional traffic features, which may affect the accuracy and robustness of detection results. In addition, in a dynamically changing network environment, the application of static benchmarks or fixed rules may not fully adapt to the needs of actual scenarios, resulting in high false positive and false negative rates. At the same time, some technical solutions mainly focus on specific types of malware, with limited scope of application, and may have insufficient coverage when facing diverse malware types. In addition, the preprocessing and feature extraction process of traffic data relies more on manual rules, with low intelligence and automation, which may limit detection efficiency and scalability. SUMMARY
[0005] To solve the technical problems existing in the prior art, the embodiments of the present application provide a malware detection method and system fusing traffic analysis, a device and a storage medium. The technical solution is as follows: A malware detection method fusing traffic analysis, comprising the following steps: S1: Obtain key features of communication behavior in network traffic, extract their distribution density and changing trend in time series, determine whether the time distribution is periodic, calculate the degree of spatial clustering, and generate a comparison map of traffic feature distribution. S2: Call the density and clustering values in the traffic characteristic distribution comparison chart, establish a sequence according to the time window, analyze the continuous fluctuation of characteristic points, filter abnormal density segments, mark the time and space combination range, and generate a set of coordinates of abnormal behavior segments. S3: Obtain node connection information of communication path in network traffic data, identify changes in transmission direction and data volume of adjacent nodes, determine the degree of concentration of main path, filter coherent segments, and generate a continuous structure diagram of communication path transmission. S4: Call the coordinate information of the abnormal behavior segment coordinate set and the continuous structure diagram of the communication path transmission, determine whether the path coincides with the abnormal area, identify invalid paths, extract valid structure boundaries, and generate a filtered communication path structure coordinate range diagram. S5: Call the outline range in the coordinate range map of the filtered communication path structure, extract the center of the internal node, calculate the angle and distance of any node block, classify them according to the consistency of arrangement and determine the connection order, and generate the label arrangement path order index.
[0006] As a further embodiment of the present invention, the traffic feature distribution comparison map includes a time distribution density sequence, spatial clustering judgment parameters, node connection structure graph, and distribution mapping coordinate set; the abnormal behavior segment coordinate set includes a time window index number, abnormal segment boundary, cumulative fluctuation features, and merged segment identifier; the communication path transmission continuous structure graph includes transmission direction distribution, data volume change intensity, path trajectory, and continuous structure node group; the filtered communication path structure coordinate range graph includes abnormal segment removal markers, boundary connected regions, circumscribed rectangle range, and communication path coordinate index; and the label arrangement path order index includes a center point location set, arrangement direction angle, path connection order, and label classification number.
[0007] As a further aspect of the present invention, the step of obtaining S1 is as follows: S101: Obtain key feature points of communication behavior in network traffic data, call the timestamp and spatial coordinate set in traffic data, extract nodes where the data volume changes abruptly in the region of each feature point, filter the node set at the boundary of the feature point by the cluster density of the location coordinates, group the nodes into permutation sequences according to the time series and spatial arrangement order, mark the index position according to the order of the nodes in the time and spatial coordinate system in each group sequence, and generate node permutation coordinate sequence values. S102: Based on the node arrangement coordinate sequence value, calculate the set of time interval values of adjacent nodes in the time series, take the time difference values of all node pairs to form an interval sequence, perform a one-time difference operation on the difference values of adjacent intervals in the interval sequence, determine whether all differences are constant values, establish the interval difference state matrix of nodes in the time series, and generate the time equal arithmetic distribution stability coefficient. S103: Based on the node arrangement coordinate sequence value and the time equal distribution stability coefficient, extract the spatial coordinate sequence of the first node in the spatial data, measure the set of spatial distance values of each adjacent node, calculate the difference sequence between all distance values, construct a continuous aggregation change curve, connect the original positions of each node in sequence to form a line segment structure, superimpose it on the time and spatial coordinate space and mark the aggregation trend range, and generate a traffic characteristic distribution comparison map.
[0008] As a further aspect of the present invention, the step of obtaining S2 is as follows: S201: Call the time distribution density and spatial clustering values recorded in the traffic feature distribution comparison chart, assign a sequence number to each time window block in the traffic data, and form an ordered number group according to the block arrangement order in time and spatial coordinates. Based on the block corresponding to each number in the group, extract the time distribution density of its location as a numerical sequence and establish the block distribution sequence group value. S202: Based on the block distribution sequence group value, perform item-by-item difference accumulation operation on the time distribution density value corresponding to adjacent numbered blocks, compare the cumulative distribution density value with the total difference of all time distribution density values in the current numbered segment, determine whether there is a continuous fluctuation trend in the current numbered segment, and call the spatial clustering degree value recorded in the traffic feature distribution comparison chart to obtain the difference between the maximum and minimum values of the spatial clustering degree of the corresponding numbered segment, calculate the ratio of this difference to the average value of all spatial clustering degree differences in the segment, and generate the fluctuation change amplitude coefficient; S203: Based on the block distribution sequence group value and fluctuation change amplitude coefficient, identify numbered segments with continuous fluctuation trend and spatial clustering fluctuation amplitude exceeding a set threshold, extract all block number sets that meet the conditions, aggregate the coordinate boundary points corresponding to each number in the set in terms of time and space range, map the aggregated boundary range to the time and space coordinate system, and establish a set of coordinates for abnormal behavior segments.
[0009] As a further aspect of the present invention, the step of obtaining S3 is as follows: S301: Obtain node connection information of communication path in network traffic data, collect transmission value matrix of communication path under data volume channel, traverse the transmission value of each node and its neighboring nodes, calculate the transmission difference between each group of nodes, record the direction identifier of transmission value change, construct transmission jump amplitude sequence under each direction, determine whether each jump amplitude value in the sequence exceeds the average transmission change value corresponding to that direction, and generate direction transmission jump amplitude value. S302: Based on the direction transmission jump amplitude value, extract the node groups in each direction whose jump amplitude continuously exceeds the average value, classify and aggregate these node groups according to the direction identifier, perform line segment connection on the nodes in each group that have the same jump direction and are continuously arranged, construct a set of node trajectories for continuously changing paths, draw all trajectory line segments in time and space coordinate space and generate a path merging layer, and establish a continuous structure diagram of communication path transmission.
[0010] As a further aspect of the present invention, the step of obtaining S4 is as follows: S401: Call the coordinate set of the abnormal behavior segment and the coordinate range in the continuous structure diagram of the communication path transmission, perform overlap detection on each boundary coordinate point in the two sets of regions, extract the horizontal and vertical ranges of all boundary points in the time and space coordinate system, compare whether there is an intersection area in the same coordinate axis direction, determine whether the time and space overlap is not less than the set intersection node threshold, and generate a region intersection matching mark value. S402: Based on the region intersection matching mark value, filter out the region pairs that simultaneously meet the intersection condition, extract the coordinate segments in each pair that belong to the abnormal behavior feature region and overlap with the communication path transmission range as abnormal blocks, remove all coordinate sets with intersection marks, and retain the communication path connected coordinate point set without intersection marks to generate a non-overlapping communication path connected path coordinate group. S403: Based on the coordinate group of the non-overlapping communication path connected paths, extract the minimum and maximum coordinate values of the boundaries in each path, construct the bounding rectangle containing the path region according to the time and space coordinate system, calculate the minimum bounding rectangle of each group of boundary points, and map the set of rectangle ranges to the time and space coordinate region layer to establish a filtered communication path structure coordinate range map.
[0011] As a further aspect of the present invention, the step of obtaining S5 is as follows: S501: Call the contour boundary information recorded in the filtered communication path structure coordinate range map, limit the node area covered by each boundary rectangle in time and space coordinates as the search range, extract the contour coordinate set of the communication path in each range, calculate the geometric center point of all node coordinates in each contour boundary, organize the node structure center points in the communication path area according to the coordinate order, and generate the node area center point coordinate group. S502: Based on the coordinate group of the center point of the node region, calculate the directional angle and spatial distance between any two center points, use the directional angle and distance of all center point pairs as a set of combined parameters to identify the arrangement trend, call the directional angle difference between each center point and its adjacent points to determine whether its arrangement direction has continuous characteristics, and generate a set of directional offset filtering paths. S503: Based on the direction offset, filter the path set, extract the set of center point connection paths that meet the continuous direction feature, number the points in the path according to the relative order of the points in time and space coordinates, map all the numbered paths to a unified layer and merge them into a single sorted node structure, and establish a label arrangement path order index.
[0012] A malware detection system integrating traffic analysis, the system comprising: The feature distribution mapping module obtains the coordinates of key nodes in the communication behavior of network traffic data, extracts the time distribution density and determines whether it forms an arithmetic sequence, calculates the coordinate difference of the spatial starting point in the whole time and space coordinates, connects all nodes to construct a structural graphic mapping, and generates a traffic feature distribution comparison map by combining the displacement and distribution sequence in the time and space directions. The abnormal behavior segment calculation module calls the distribution density and aggregation value in the traffic feature distribution comparison map, establishes a time window number sequence, judges the cumulative trend of distribution density in each window, filters segments whose spatial aggregation fluctuation exceeds the set threshold, marks the combined areas of abnormal behavior at the same time, and generates a set of abnormal behavior segment coordinates. The communication path transmission identification module acquires node connection information and transmission values of the communication path, analyzes the direction of change of transmission difference between adjacent nodes, determines whether continuous transmission is concentrated in the main path, filters out node segments with coherent directions and connects them into a path structure, and generates a continuous structure diagram of communication path transmission. The abnormal segment removal and judgment module calls the coordinate information of the abnormal behavior segment coordinate set and the continuous structure diagram of the communication path transmission, performs overlap judgment on the corresponding areas, marks the overlapping segments as abnormal paths and removes them, extracts the boundary contour of the remaining path, and generates a filtered communication path structure coordinate range diagram. The tag path sequence construction module calls the boundary data of the filtered communication path structure coordinate range map, extracts the center coordinates of the node blocks, calculates the angle and distance between any two points, classifies them into permutation chains according to the degree of offset, constructs paths according to the connection order, and generates a tag permutation path sequence index.
[0013] A malware detection device integrating traffic analysis includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the malware detection system integrating traffic analysis.
[0014] A malware detection storage medium integrating traffic analysis is disclosed, wherein a computer program is stored thereon, and when the computer program is executed by a processor, it implements the steps of the malware detection method integrating traffic analysis. The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this invention, by constructing a node distribution structure mapping under time and spatial coordinates and integrating the combination rules of time density and spatial clustering difference, a basis for judging abnormal behavior can be formed. Abnormal segments are identified by grouping based on block cumulative values and permutation sequences. Communication trajectories are constructed by combining transmission change amplitude and jump direction, and composite abnormal regions are filtered out by the intersection of spatiotemporal segments. Node centers are extracted within defined boundaries and classified according to permutation offset, and the actual path arrangement is analyzed, enhancing the coherence of structure recognition and the stability of behavior judgment. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the workflow of the present invention; Figure 2 This is a detailed flowchart of S1 of the present invention; Figure 3 This is a detailed flowchart of the S2 process of the present invention; Figure 4 This is a detailed flowchart of the S3 process of the present invention; Figure 5 This is a detailed flowchart of the S4 process of the present invention; Figure 6 This is a detailed flowchart of S5 of the present invention; Figure 7 This is a system flowchart of the present invention. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0017] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Example
[0018] Please see Figure 1 This invention provides a technical solution: a malware detection method integrating traffic analysis, comprising the following steps: S1: Obtain key feature points of communication behavior in network traffic data, extract the distribution density and change trend of each feature point in the time series, determine whether the time distribution of feature points shows a periodic pattern, calculate the degree of clustering of feature points in the spatial dimension, construct a feature distribution mapping layer, and generate a traffic feature distribution comparison map. S2: Call the distribution density and clustering values in the traffic feature distribution comparison chart, establish a sequence according to the time window, perform continuous fluctuation analysis on the changing trend of feature points in each window, filter the abnormal distribution density segments, mark the combination range of time and space abnormal segments, and generate a set of coordinates of abnormal behavior segments. S3: Obtain node connection information of communication path in network traffic data, identify the transmission direction and data volume change trend between adjacent nodes, determine whether continuous transmission is concentrated in the main path, filter coherent segments of transmission direction and form structural path, and generate a continuous structure diagram of communication path transmission. S4: Call the coordinate set of abnormal behavior segment and the coordinate information of the continuous structure diagram of the communication path transmission, perform position overlap judgment on the abnormal area and the transmission path, identify the intersection segment as an invalid path, retain the remaining transmission paths and extract the structural boundary, and generate a filtered communication path structure coordinate range diagram. S5: Call the outline range of the filtered communication path structure coordinate range map, extract the center position of the internal nodes, perform angle and distance calculations on any two node blocks, classify the node blocks according to the consistency of arrangement and form a connection order, and generate a label arrangement path order index.
[0019] The traffic characteristic distribution comparison map includes time distribution density sequence, spatial clustering judgment parameters, node connection structure diagram and distribution mapping coordinate set. The abnormal behavior segment coordinate set includes time window index number, abnormal segment boundary, fluctuation cumulative characteristics and merged segment identifier. The communication path transmission continuous structure diagram includes transmission direction distribution, data volume change intensity, path trajectory and continuous structure node group. The filtered communication path structure coordinate range diagram includes abnormal segment removal mark, boundary connected region, circumscribed rectangle range and communication path coordinate index. The label arrangement path order index includes center point location set, arrangement direction angle, path connection order and label classification number.
[0020] Please see Figure 2 The steps to obtain S1 are as follows: S101: Obtain key feature points of communication behavior in network traffic data, call the timestamp and spatial coordinate set in traffic data, extract nodes where the data volume changes abruptly in the region of each feature point, filter the node set at the boundary of the feature point by the cluster density of the location coordinates, group the nodes into permutation sequences according to the time series and spatial arrangement order, mark the index position according to the order of the nodes in the time and spatial coordinate system in each group sequence, and generate node permutation coordinate sequence values. Based on the collected traffic data, the sending time and node location coordinates recorded in each data packet are extracted. These time and coordinates are combined as two dimensions to form a time-space mapped node set. For each node, a fixed radius area is defined centered on its location. The number of data packets appearing within this area per unit time is counted. The number in the current time period is compared with the data volume in the previous time period. If the difference is greater than a fixed value, such as 20 data packets, it is determined as a data mutation, and the node is recorded. Subsequently, among all data mutation nodes, the number of data mutation nodes per unit area is counted. If the density of a region exceeds 5 nodes per 100 square meters, the region is considered a mutation node. The data mutation density is relatively high, belonging to the feature point region. The boundary of this region is redefined, and the range is expanded by a certain distance within the boundary. The set of neighboring nodes is re-extracted to complete the boundary node screening. Then, these boundary nodes are sorted from earliest to latest according to their sending time, and their horizontal and vertical coordinates in space are simultaneously sorted in ascending order. This results in a node arrangement structure that follows the sorting rules on both time and space axes. Furthermore, each node is assigned a sequence number, starting from 1 and increasing sequentially according to the sorting results. For example, the first node with the leftmost position in space and the earliest time is numbered 1, and so on until the last node. Finally, an arrangement coordinate list is formed in which each node has position and sequence number information.
[0021] S102: Based on the node arrangement coordinate sequence value, calculate the set of time interval values of adjacent nodes in the time series, take the time difference of all node pairs to form an interval sequence, perform a one-time difference operation on the difference of adjacent intervals in the interval sequence, determine whether all differences are constant values, establish the interval difference state matrix of nodes in the time series, and generate the time equal arithmetic distribution stability coefficient. Based on the generated node coordinate sequence values, the timestamps of the nodes are extracted sequentially. The time differences between two adjacent timestamps are compared in sequence, and all time differences are summarized to form a continuous time interval set. Then, the differences between two adjacent time intervals in this set are subtracted again to obtain a set of difference sequences. For this set of difference sequences, each item is judged to see if it is all a certain fixed value or if its variation is within a certain range. For example, if each difference is less than 1 second, it is considered to be within an acceptable fluctuation range and can be regarded as constant. All differences are counted to see if they meet the condition. If all of them meet the condition, it means that the time interval distribution of the node has an arithmetic progression. A judgment matrix of the time difference distribution state is generated to describe the stability of the change of each time interval. For example, if there are 10 sets of differences, and 9 of them meet the condition that the variation is within 1 second, it means that the sequence has high stability. The number of sets that meet the condition is divided by the total number to obtain the stability ratio, which is used as the stability coefficient of the time distribution. The closer the coefficient is to 1, the more stable the interval is. If there are multiple differences that do not meet the condition, the coefficient is reduced accordingly. Finally, this coefficient is used to represent the arithmetic progression stability state of the current node's time sequence.
[0022] S103: Based on the node arrangement coordinate sequence value and the stability coefficient of the time arithmetic distribution, extract the spatial coordinate sequence of the first node in the spatial data, measure the set of spatial distance values of each adjacent node, calculate the difference sequence between all distance values, and construct a continuous aggregation change curve. Connect the nodes sequentially according to their original positions to form a line segment structure, superimpose it on the time and spatial coordinate space and mark the aggregation trend range to generate a traffic characteristic distribution comparison map. Based on the aforementioned node arrangement coordinate sequence values and time distribution stability coefficient, the first group of nodes in each group is selected, and their spatial coordinates are extracted sequentially. For each pair of adjacent nodes, the actual distance between them is calculated. The distance is derived using the horizontal and vertical coordinates of the location points to obtain the distance values between all adjacent nodes. These distances are arranged according to the order in which the nodes appear, and the difference between two adjacent distance values is processed to obtain a distance change sequence. Then, a visual clustering change curve is constructed based on the changing trend of these differences. If the difference gradually increases, it indicates that the distance between nodes is getting farther and the clustering degree is decreasing; conversely, it indicates that the clustering is increasing. Subsequently, the nodes are connected sequentially according to their original positions to form a line segment structure, forming a path direction diagram. This diagram is superimposed on a two-dimensional coordinate system of time and space, with the timestamp of each node as the vertical axis and the location coordinate as the horizontal axis, forming a graphical coordinate system. Each line segment in the diagram is numbered and marked with its direction, ultimately forming a complete traffic path trajectory diagram. The strength range of the clustering degree can be marked by the density changes between different path segments, forming the final traffic characteristic distribution comparison diagram.
[0023] Please see Figure 3The steps to obtain S2 are as follows: S201: Call the time distribution density and spatial clustering values recorded in the traffic feature distribution comparison chart, assign a sequence number to each time window block in the traffic data, and form an ordered number group according to the block arrangement order in time and spatial coordinates. Based on the block corresponding to each number in the group, extract the time distribution density of its location as a numerical sequence and establish the block distribution sequence group value. The entire traffic data is divided into continuous and non-overlapping time windows on the time axis. Each window can be set to a length of 5 seconds and numbered sequentially as T1, T2, T3, etc. Then, each time window is spatially divided according to its location within its spatial region. Each spatial block is set as a 100m × 100m grid area based on the horizontal and vertical coordinate intervals. The combination of time windows and spatial grids forms a two-dimensional block. For example, in the time window numbered T1, the spatial blocks are numbered sequentially from left to right as S1, S2, S3, etc. Each T+S combination forms a uniquely numbered block, with numbering methods such as T1S1, T1S2...T2S1, T2S2, and so on. Then, all blocks are numbered in an ordered sequence according to the time window from front to back, each spatial coordinate from left to right, and top to bottom. For example, a numbering sequence might be: T1S1, T1S2, T1S3. For each numbered block, such as T2S1, T2S2, T2S3, etc., the time distribution density value of its corresponding position is extracted from the traffic characteristic distribution comparison map. This value is the number of data packets received in the spatial block per unit time. The unit can be set to "data packets per second". For example, the time distribution density corresponding to a block number T3S2 is 15 packets per second. This value is recorded under the corresponding numbered block. Finally, all numbered blocks and their corresponding time distribution density values are combined into a continuous numerical sequence according to the numbering order. This sequence is the block distribution sequence group value.
[0024] S202: Based on the block distribution sequence group value, perform item-by-item difference accumulation operation on the time distribution density value corresponding to adjacent numbered blocks, compare the cumulative distribution density value with the total difference of all time distribution density values in the current numbered segment, determine whether there is a continuous fluctuation trend in the current numbered segment, and call the spatial clustering degree value recorded in the traffic feature distribution comparison chart to obtain the difference between the maximum and minimum values of the spatial clustering degree of the corresponding numbered segment, calculate the ratio of this difference to the average value of all spatial clustering degree differences in the segment, and generate the fluctuation change amplitude coefficient; For every two adjacent numbered blocks, the time distribution density values are subtracted sequentially according to their numbering order to obtain a sequence of density change values. These change values are then accumulated to obtain the cumulative change value of the segment containing the current number. For example, if the block numbering group is T2S1, T2S2, and T2S3, with time distribution density values of 12, 18, and 15 respectively, adjacent differences are 6 and -3, and the cumulative difference is 3, this cumulative difference is subtracted from the total difference between all time distribution density values in the current segment and the average value in the segment. If the difference between the cumulative difference and the total difference is large, i.e., greater than a set difference threshold of 5, it is determined that the segment may have a continuous fluctuation trend. The system retrieves the spatial clustering value of each block recorded in the traffic feature distribution comparison chart, obtains the maximum and minimum clustering values in the current segment (e.g., the maximum is 0.85 and the minimum is 0.25), calculates the difference between them as 0.6, and then calculates the difference between all adjacent blocks from the clustering values of all blocks in the segment. The average of all differences is then calculated (e.g., the average difference is 0.15). The maximum and minimum difference of 0.6 are divided by the average difference of 0.15 to obtain a ratio of 4.0, which is recorded as the fluctuation range coefficient. If this coefficient exceeds a certain threshold, such as 3.0, it is considered that the clustering fluctuation range of the segment is large, thus completing the generation of the fluctuation range coefficient for the segment.
[0025] S203: Based on the block distribution sequence group value and fluctuation change amplitude coefficient, identify numbered segments with continuous fluctuation trends and spatial clustering fluctuation amplitude exceeding a set threshold, extract all block number sets that meet the conditions, aggregate the coordinate boundary points corresponding to each number in the set in terms of time and space range, map the aggregated boundary range to the time and space coordinate system, and establish a set of coordinates for abnormal behavior segments. The process involves selecting paragraphs from all paragraphs that meet two criteria: first, their temporal distribution density exhibits a continuous fluctuation trend, meaning the cumulative difference within the paragraph exceeds a preset difference threshold (e.g., a cumulative difference of 9 and a threshold of 5); second, the spatial clustering fluctuation coefficient of the paragraph exceeds a set baseline value (e.g., a threshold of 3.0 and a current coefficient of 4.5). For all eligible paragraphs, their numbered block sets are extracted (e.g., T3S2, T3S3, T4S1, etc.). Then, the four boundary angle coordinates are extracted from each corresponding spatial block. The system records the start and end times of the corresponding time windows, merges all boundary corner coordinates and their corresponding time windows, obtains the actual range of the entire segment in time and space, and forms an aggregated boundary data set. This set is then mapped onto a two-dimensional coordinate graph of time and space. For example, the rectangular area formed by T3S2 to T4S1 is drawn in the area of time segment 3 to 4, space segment X-axis 100 meters to 300 meters, and space segment Y-axis 200 meters to 400 meters, and is uniformly marked as an abnormal behavior area, thus forming a set of coordinates for the abnormal behavior segment.
[0026] Please see Figure 4 The steps to obtain S3 are as follows: S301: Obtain node connection information of communication path in network traffic data, collect transmission value matrix of communication path under data volume channel, traverse the transmission value of each node and its neighboring nodes, calculate the transmission difference between each group of nodes, record the direction identifier of transmission value change, construct transmission jump amplitude sequence under each direction, determine whether each jump amplitude value in the sequence exceeds the average transmission change value corresponding to that direction, and generate direction transmission jump amplitude value. Extract the source and destination node addresses from all communication pairs in the raw network traffic data to form a set of node connection pairs, such as node A and node B, node B and node C, etc. Each connection pair corresponds to a communication path. Record the number of data packets transmitted between each connection pair per unit time into a transmission value matrix. Each row in this matrix represents a starting node, and each column represents a destination node. The elements in the matrix are the transmission amounts from the starting node to the destination node. For example, the transmission amount from node A to node B is 120 data packets, and from node B to node C is 95 data packets. Then, traverse each node and its neighboring nodes... Using the current node as the central node, sequentially search for all its directly communicating target nodes, read the transmission volume values between the central node and its adjacent nodes one by one, and calculate the transmission difference between the current node and each neighboring node. For example, if the transmission volume of node A is 120 and that of node B is 95, the difference is 25. If the difference is positive, it is marked as rising towards the target direction; if it is negative, it is marked as falling. Simultaneously, record the direction corresponding to each transmission difference. For example, from A to B is eastward rising, and from B to C is northward falling. Collect and organize the transmission differences in each direction to construct a jump amplitude sequence for that direction. For example, the eastward jump amplitude sequence is [25, [30, 18], and [12, -15, 20] for the north direction. Then, in each sequence, it is determined whether each jump amplitude exceeds the average value of all jump amplitude values in that direction. If the current jump amplitude is greater than the average value of all jump amplitude values in that direction, for example, when the average value is 20, a certain jump amplitude is 30, then the value is marked as a valid jump. Finally, all amplitudes greater than the average jump value in that direction are extracted from the jump amplitude sequences of all directions and marked as directional transmission jump amplitude values.
[0027] S302: Based on the directional transmission jump amplitude value, extract the node groups in each direction where the jump amplitude continuously exceeds the average value, classify and aggregate these node groups according to the directional identifier, perform line segment connection on the nodes in each group that have the same jump direction and are arranged continuously, construct a set of node trajectories for continuously changing paths, draw all trajectory line segments in time and space coordinate space and generate a path merging layer, and establish a continuous structure diagram of communication path transmission. From the jump sequence in each direction, select node groups where the jump amplitude consecutively exceeds the corresponding average value. For example, in the eastward jump amplitude sequence, if node group ABC has jump values of 25, 28, and 30, all greater than the average value of 20 for that direction, then this group of nodes is extracted. All node groups that meet the criteria are categorized and aggregated according to their direction, such as eastward group, northward group, westward group, etc. Each direction group contains multiple consecutive jump node links. For each group of nodes, determine whether their jump directions are consistent. If the direction identifiers are consistent and the nodes are adjacent in order, for example, nodes X, Y, and Z are all eastward... The process involves connecting these nodes sequentially with straight lines to create a continuously changing path, such as connecting from X to Y and then to Z, forming the path XYZ. All the resulting path segments are then organized and plotted on a spatial coordinate graph according to the spatial coordinates of each node and the corresponding time value. The horizontal axis represents spatial location, and the vertical axis represents time progress, forming a layer that displays the transmission trajectory of the communication nodes. Path segments in different directions are then superimposed on the graph, with different colors used to represent different directions, forming a complete path merging layer. Finally, a continuous structure diagram of the communication path transmission is obtained.
[0028] Please see Figure 5 The steps to obtain S4 are as follows: S401: Call the coordinate set of the abnormal behavior segment and the coordinate range in the continuous structure diagram of the communication path transmission, perform overlap detection on each boundary coordinate point in the two sets of regions, extract the horizontal and vertical ranges of all boundary points in the time and space coordinate system, compare whether there is an intersection area in the same coordinate axis direction, determine whether the time and space overlap is not less than the set intersection node threshold, and generate the region intersection matching mark value. Read the boundary coordinates of each segment in both sets of data. The coordinates of the abnormal behavior segments are represented by rectangular boundaries. Each segment contains four boundary values: top, bottom, left, and right. The horizontal and vertical coordinate ranges of each path segment in the communication path structure are also recorded. For each pair of regions, compare whether their horizontal coordinate ranges overlap. The criterion is whether there is an intersection between the minimum and maximum values of the two. For example, if the horizontal range of an abnormal region is 100 to 200, and the horizontal range of the communication path is 180 to 250, they have an intersection of 20 to 200. The same method is used to compare... If there is an intersection in the vertical range, it means that the two regions actually overlap in spatial coordinates. Then, compare the start and end times of the two regions on the time axis to see if they intersect. For example, if one period is from the 5th to the 10th period and the other is from the 9th to the 14th period, it is determined that there is a temporal intersection. Record the number of points that simultaneously meet the intersection conditions in both space and time. If the number of intersection nodes is not less than the set threshold (which can be set to 5 points), mark the regions that meet the conditions as matching regions based on the intersection judgment results, and finally generate the region intersection matching mark value.
[0029] S402: Based on the region intersection matching mark value, filter out the region pairs that simultaneously meet the intersection condition, extract the coordinate segments in each pair that belong to the abnormal behavior feature region and the communication path transmission range as abnormal blocks, remove all coordinate sets with intersection marks, and retain the communication path connected coordinate point set without intersection marks to generate non-overlapping communication path connected path coordinate group. Based on the region intersection matching marker value, all abnormal regions and path region pairs that meet the intersection conditions are filtered out. For each pair of intersection regions, the spatial segments that actually overlap are extracted, that is, path segments whose horizontal and vertical coordinates are both within the intersection interval. For example, if the X-axis 120 to 180 and the Y-axis 150 to 200 in the communication path are located in the abnormal block, then this segment is recorded as an abnormal block. At the same time, the path segments that have been marked as intersecting are removed from the communication path coordinate set, and only the path segments that have not participated in any intersection are retained. The coordinates of all nodes in these remaining path segments are summarized and recombined into a continuous connected path point set to generate a non-overlapping communication path connected path coordinate group.
[0030] S403: Based on the coordinate group of the connected paths of non-overlapping communication paths, extract the minimum and maximum coordinate values of the boundaries in each path, construct the bounding rectangle containing the path region according to the time and space coordinate system, calculate the minimum bounding rectangle of each group of boundary points, and map the set of rectangle ranges to the time and space coordinate region layer to establish a filtered communication path structure coordinate range map. Based on the coordinate group of the non-overlapping communication path, the minimum and maximum values of the boundary are extracted for each path. For example, if the minimum X coordinate of a path is 300 and the maximum is 420, and the minimum Y coordinate is 100 and the maximum is 250, then the bounding rectangle range of the path is determined to be 300 to 420 horizontally and 100 to 250 vertically. If the path spans multiple time windows, the earliest and latest time numbers are recorded to form the start and end boundaries of the time axis. All path boundaries are constructed into rectangles. Each rectangle is composed of the start and end time of the corresponding path and the spatial range. Then, each rectangular coordinate segment is mapped onto a two-dimensional layer combining space and time to finally form a filtered communication path structure coordinate range map.
[0031] Please see Figure 6 The steps to obtain S5 are as follows: S501: Call the contour boundary information recorded in the filtered communication path structure coordinate range map, limit the node area covered by each boundary rectangle in time and space coordinates as the search range, extract the contour coordinate set of the communication path in each range, calculate the geometric center point of all node coordinates in each contour boundary, organize the node structure center points in the communication path area according to the coordinate order, and generate the node area center point coordinate group. The coordinates of the top-left and bottom-right corners of each bounding rectangle are read as the boundary ranges of the time and space axes. These rectangle boundaries are used as constraints to sequentially scan the coordinates of all nodes covered within the rectangles, forming a corresponding node list within each rectangle. For example, if 8 communication nodes are extracted in a certain area, their spatial positions are distributed within the range of 200 to 260 on the x-axis and 300 to 360 on the y-axis. Then, the average of the x and y coordinates of each node is calculated to obtain the coordinates of the geometric center point in that area, which is the structural center position of all nodes in space. The geometric center points generated in each rectangle are sorted from left to right and from front to back according to their coordinates. The sorting method is first arranged from early to late according to the time dimension, and then arranged in ascending order of the horizontal axis coordinate values within the same time period. Finally, a complete set of node region center point coordinates is constructed, which is used for subsequent arrangement feature extraction tasks.
[0032] S502: Based on the coordinate group of the center point of the node region, calculate the directional angle and spatial distance between any two center points, use the directional angle and distance of all center point pairs as a set of combined parameters to identify the arrangement trend, call the directional angle difference between each center point and its adjacent points to determine whether its arrangement direction has continuous characteristics, and generate a set of directional offset filtering paths. Based on the coordinate set of the center point of the node region, any two center points are selected sequentially, and their horizontal and vertical coordinate differences are extracted. The directional change angle between the two points is calculated based on these two differences, and the straight-line distance between the two points is recorded as another parameter value. These angles and distances are combined into a set of parameters. Then, the directional angle difference between each center point and its adjacent points is compared. For example, if the current center point angle is 45 degrees and the next point angle is 47 degrees, the directional angle difference between the two points is 2 degrees. If the difference is less than a set threshold, such as 5 degrees, it is determined that the two points are basically aligned in the same direction; otherwise, it is determined to be a sudden change in direction. This process is repeated for the entire center point sequence to identify point sequence segments with directional continuity characteristics. For example, if the directional difference of 5 consecutive points is less than 5 degrees, it is classified as a continuous path. Finally, the generation of the directional offset filtering path set is completed.
[0033] S503: Filter the path set according to the direction offset, extract the set of center point connection paths that meet the continuous direction feature, number the points in the path according to the relative order of time and space coordinates, map all numbered paths to a unified layer and merge them into a single sorted node structure, and establish a label-arranged path order index. The path set is filtered based on directional offset, and all center point path segments that meet the continuity condition are extracted. The points contained in each path are renumbered according to time and spatial order. For example, the five points in the first path are numbered from morning to night as P01 to P05, and the points in the second path are numbered as P06 to P10, ensuring that the numbering order is not repeated in the overall structure. Then, all numbered points are connected according to coordinate relationship and drawn into a unified two-dimensional layer. This layer uses time as the vertical axis and space as the horizontal axis, and merges all paths into a sorted structure according to the numbering order, finally forming a label-arranged path order index.
[0034] Please see Figure 7 A malware detection system that integrates traffic analysis, the system comprising: The feature distribution mapping module obtains the coordinates of key nodes in the communication behavior of network traffic data, extracts the time distribution density and determines whether it forms an arithmetic sequence, calculates the coordinate difference of the spatial starting point in the whole time and space coordinates, connects all nodes to construct a structural graphic mapping, and generates a traffic feature distribution comparison map by combining the displacement and distribution sequence in the time and space directions. The abnormal behavior segment calculation module calls the distribution density and clustering values in the traffic feature distribution comparison chart, establishes a time window number sequence, judges the cumulative trend of distribution density within each window, filters segments whose spatial clustering fluctuations exceed the set threshold, marks the combined areas of abnormal behavior, and generates a set of abnormal behavior segment coordinates. The communication path transmission identification module acquires node connection information and transmission values of the communication path, analyzes the direction of change of transmission difference between adjacent nodes, determines whether continuous transmission is concentrated in the main path, filters out node segments with coherent directions and connects them into a path structure, and generates a continuous structure diagram of communication path transmission. The abnormal segment removal and judgment module calls the coordinate information of the abnormal behavior segment coordinate set and the continuous structure diagram of the communication path transmission, performs overlap judgment on the corresponding areas, marks the overlapping segments as abnormal paths and removes them, extracts the boundary contour of the remaining path, and generates a filtered communication path structure coordinate range diagram. The tag path sequence construction module calls the boundary data of the filtered communication path structure coordinate range map, extracts the center coordinates of the node blocks, calculates the angle and distance between any two points, classifies them into permutation chains according to the degree of offset, constructs paths according to the connection order, and generates a tag permutation path order index.
[0035] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for detecting malware that integrates traffic analysis, characterized in that, Includes the following steps: S1: Obtain key features of communication behavior in network traffic, extract their distribution density and changing trend in time series, determine whether the time distribution is periodic, calculate the degree of spatial clustering, and generate a comparison map of traffic feature distribution. S2: Call the density and clustering values in the traffic characteristic distribution comparison chart, establish a sequence according to the time window, analyze the continuous fluctuation of characteristic points, filter abnormal density segments, mark the time and space combination range, and generate a set of coordinates of abnormal behavior segments. S3: Obtain node connection information of communication path in network traffic data, identify changes in transmission direction and data volume of adjacent nodes, determine the degree of concentration of main path, filter coherent segments, and generate a continuous structure diagram of communication path transmission. S4: Call the coordinate information of the abnormal behavior segment coordinate set and the continuous structure diagram of the communication path transmission, determine whether the path coincides with the abnormal area, identify invalid paths, extract valid structure boundaries, and generate a filtered communication path structure coordinate range diagram. S5: Call the outline range in the coordinate range map of the filtered communication path structure, extract the center of the internal node, calculate the angle and distance of any node block, classify them according to the consistency of arrangement and determine the connection order, and generate the label arrangement path order index.
2. The malware detection method based on fused traffic analysis according to claim 1, characterized in that: The traffic feature distribution comparison map includes a time distribution density sequence, spatial clustering judgment parameters, node connection structure graph, and distribution mapping coordinate set. The abnormal behavior segment coordinate set includes time window index number, abnormal segment boundary, fluctuation cumulative feature, and merged segment identifier. The communication path transmission continuous structure graph includes transmission direction distribution, data volume change intensity, path trajectory, and continuous structure node group. The filtered communication path structure coordinate range graph includes abnormal segment removal mark, boundary connected region, circumscribed rectangle range, and communication path coordinate index. The label arrangement path order index includes center point location set, arrangement direction angle, path connection order, and label classification number.
3. The malware detection method based on fused traffic analysis according to claim 1, characterized in that, The steps for obtaining S1 are as follows: S101: Obtain key feature points of communication behavior in network traffic data, call the timestamp and spatial coordinate set in traffic data, extract nodes where the data volume changes abruptly in the region of each feature point, filter the node set at the boundary of the feature point by the cluster density of the location coordinates, group the nodes into permutation sequences according to the time series and spatial arrangement order, mark the index position according to the order of the nodes in the time and spatial coordinate system in each group sequence, and generate node permutation coordinate sequence values. S102: Based on the node arrangement coordinate sequence value, calculate the set of time interval values of adjacent nodes in the time series, take the time difference values of all node pairs to form an interval sequence, perform a one-time difference operation on the difference values of adjacent intervals in the interval sequence, determine whether all differences are constant values, establish the interval difference state matrix of nodes in the time series, and generate the time equal arithmetic distribution stability coefficient. S103: Based on the node arrangement coordinate sequence value and the time equal distribution stability coefficient, extract the spatial coordinate sequence of the first node in the spatial data, measure the set of spatial distance values of each adjacent node, calculate the difference sequence between all distance values, construct a continuous aggregation change curve, connect the original positions of each node in sequence to form a line segment structure, superimpose it on the time and spatial coordinate space and mark the aggregation trend range, and generate a traffic characteristic distribution comparison map.
4. The malware detection method based on fused traffic analysis according to claim 1, characterized in that, The steps for obtaining S2 are as follows: S201: Call the time distribution density and spatial clustering values recorded in the traffic feature distribution comparison chart, assign a sequence number to each time window block in the traffic data, and form an ordered number group according to the block arrangement order in time and spatial coordinates. Based on the block corresponding to each number in the group, extract the time distribution density of its location as a numerical sequence and establish the block distribution sequence group value. S202: Based on the block distribution sequence group value, perform item-by-item difference accumulation operation on the time distribution density value corresponding to adjacent numbered blocks, compare the cumulative distribution density value with the total difference of all time distribution density values in the current numbered segment, determine whether there is a continuous fluctuation trend in the current numbered segment, and call the spatial clustering degree value recorded in the traffic feature distribution comparison chart to obtain the difference between the maximum and minimum values of the spatial clustering degree of the corresponding numbered segment, calculate the ratio of this difference to the average value of all spatial clustering degree differences in the segment, and generate the fluctuation change amplitude coefficient; S203: Based on the block distribution sequence group value and fluctuation change amplitude coefficient, identify numbered segments with continuous fluctuation trend and spatial clustering fluctuation amplitude exceeding a set threshold, extract all block number sets that meet the conditions, aggregate the coordinate boundary points corresponding to each number in the set in terms of time and space range, map the aggregated boundary range to the time and space coordinate system, and establish a set of coordinates for abnormal behavior segments.
5. The malware detection method based on fused traffic analysis according to claim 1, characterized in that, The steps for obtaining S3 are as follows: S301: Obtain node connection information of communication path in network traffic data, collect transmission value matrix of communication path under data volume channel, traverse the transmission value of each node and its neighboring nodes, calculate the transmission difference between each group of nodes, record the direction identifier of transmission value change, construct transmission jump amplitude sequence under each direction, determine whether each jump amplitude value in the sequence exceeds the average transmission change value corresponding to that direction, and generate direction transmission jump amplitude value. S302: Based on the direction transmission jump amplitude value, extract the node groups in each direction whose jump amplitude continuously exceeds the average value, classify and aggregate these node groups according to the direction identifier, perform line segment connection on the nodes in each group that have the same jump direction and are continuously arranged, construct a set of node trajectories for continuously changing paths, draw all trajectory line segments in time and space coordinate space and generate a path merging layer, and establish a continuous structure diagram of communication path transmission.
6. The malware detection method based on fused traffic analysis according to claim 1, characterized in that, The steps for obtaining S4 are as follows: S401: Call the coordinate set of the abnormal behavior segment and the coordinate range in the continuous structure diagram of the communication path transmission, perform overlap detection on each boundary coordinate point in the two sets of regions, extract the horizontal and vertical ranges of all boundary points in the time and space coordinate system, compare whether there is an intersection area in the same coordinate axis direction, determine whether the time and space overlap is not less than the set intersection node threshold, and generate a region intersection matching mark value. S402: Based on the region intersection matching mark value, filter out the region pairs that simultaneously meet the intersection condition, extract the coordinate segments in each pair that belong to the abnormal behavior feature region and overlap with the communication path transmission range as abnormal blocks, remove all coordinate sets with intersection marks, and retain the communication path connected coordinate point set without intersection marks to generate a non-overlapping communication path connected path coordinate group. S403: Based on the coordinate group of the non-overlapping communication path connected paths, extract the minimum and maximum coordinate values of the boundaries in each path, construct the bounding rectangle containing the path region according to the time and space coordinate system, calculate the minimum bounding rectangle of each group of boundary points, and map the set of rectangle ranges to the time and space coordinate region layer to establish a filtered communication path structure coordinate range map.
7. The malware detection method based on fused traffic analysis according to claim 1, characterized in that, The steps for obtaining S5 are as follows: S501: Call the contour boundary information recorded in the filtered communication path structure coordinate range map, limit the node area covered by each boundary rectangle in time and space coordinates as the search range, extract the contour coordinate set of the communication path in each range, calculate the geometric center point of all node coordinates in each contour boundary, organize the node structure center points in the communication path area according to the coordinate order, and generate the node area center point coordinate group. S502: Based on the coordinate group of the center point of the node region, calculate the directional angle and spatial distance between any two center points, use the directional angle and distance of all center point pairs as a set of combined parameters to identify the arrangement trend, call the directional angle difference between each center point and its adjacent points to determine whether its arrangement direction has continuous characteristics, and generate a set of directional offset filtering paths. S503: Based on the direction offset, filter the path set, extract the set of center point connection paths that meet the continuous direction feature, number the points in the path according to the relative order of the points in time and space coordinates, map all the numbered paths to a unified layer and merge them into a single sorted node structure, and establish a label arrangement path order index.
8. A malware detection system integrating traffic analysis, characterized in that, The system is used in the malware detection method based on the fused traffic analysis according to any one of claims 1-7, the system comprising: The feature distribution mapping module obtains the coordinates of key nodes in the communication behavior of network traffic data, extracts the time distribution density and determines whether it forms an arithmetic sequence, calculates the coordinate difference of the spatial starting point in the whole time and space coordinates, connects all nodes to construct a structural graphic mapping, and generates a traffic feature distribution comparison map by combining the displacement and distribution sequence in the time and space directions. The abnormal behavior segment calculation module calls the distribution density and aggregation value in the traffic feature distribution comparison map, establishes a time window number sequence, judges the cumulative trend of distribution density in each window, filters segments whose spatial aggregation fluctuation exceeds the set threshold, marks the combined areas of abnormal behavior at the same time, and generates a set of abnormal behavior segment coordinates. The communication path transmission identification module acquires node connection information and transmission values of the communication path, analyzes the direction of change of transmission difference between adjacent nodes, determines whether continuous transmission is concentrated in the main path, filters out node segments with coherent directions and connects them into a path structure, and generates a continuous structure diagram of communication path transmission. The abnormal segment removal and judgment module calls the coordinate information of the abnormal behavior segment coordinate set and the continuous structure diagram of the communication path transmission, performs overlap judgment on the corresponding areas, marks the overlapping segments as abnormal paths and removes them, extracts the boundary contour of the remaining path, and generates a filtered communication path structure coordinate range diagram. The tag path sequence construction module calls the boundary data of the filtered communication path structure coordinate range map, extracts the center coordinates of the node blocks, calculates the angle and distance between any two points, classifies them into permutation chains according to the degree of offset, constructs paths according to the connection order, and generates a tag permutation path sequence index.
9. A malware detection device integrating traffic analysis, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, it implements the malware detection system with fused traffic analysis as described in any one of claims 8.
10. A malware detection storage medium integrating traffic analysis, wherein a computer program is stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the malware detection method based on fused traffic analysis as described in any one of claims 1 to 7.