LSTM-based fraud-related network flow identification method and server
By applying LSTM-based timing modeling and adaptive weight evaluation methods in network traffic analysis, the problem of difficult to identify variable fraud methods in the prior art is solved, and a higher recognition accuracy and lower false alarm rate are achieved.
Patent Information
- Application Number
- CN202510133836.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing network traffic analysis methods are difficult to adapt to the ever-changing communication methods and operating methods of criminal gangs, resulting in a large number of suspicious communication behaviors not being identified, and simple rule matching will produce a large number of false alarms.
The fraud-related network traffic recognition method is adopted based on LSTM, and the network traffic data is time-series modeled, dynamic features of communication behavior are extracted, and feature evaluation is performed in combination with adaptive weight coefficients to improve the recognition accuracy.
Effectively improve the accuracy of identification of changing fraud-related communication behaviors, reduce the false positive rate, and automatically adjust the distribution of feature importance to adapt to new threats.
Smart Images

Figure CN119945789A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of network traffic analysis, and in particular to a method and server for identifying fraudulent network traffic based on LSTM. Background Art
[0002] With the development of Internet technology, online fraud activities are becoming increasingly rampant, causing huge economic losses to society and individuals. In order to effectively combat online fraud, it is necessary to analyze massive amounts of network traffic data and promptly discover and identify criminal gangs.
[0003] At present, the mainstream network traffic analysis method adopts rule-based detection technology. This technology scans and matches network traffic data in real time by presetting a series of judgment criteria, including communication address library, port sequence, protocol type, etc. When traffic that meets the criteria is detected, the system will issue an early warning message.
[0004] However, criminal gangs continue to change their communication methods and operating methods, using techniques such as address rotation and multi-layer proxy to evade detection. Detection methods based on fixed rules are difficult to adapt to such changes, resulting in a large number of suspicious communication behaviors not being discovered. At the same time, simple rule matching will also generate a large number of false positives, making it difficult for analysts to deal with them. Summary of the invention
[0005] The present application provides a LSTM-based fraud-related network traffic identification method and server, which are used to improve the accuracy of identifying constantly changing suspicious network communication behaviors.
[0006] In the first aspect, the present application provides a method for identifying fraudulent network traffic based on LSTM, including: collecting network traffic data including communication address, communication time, protocol type and data packet size in real time through a network traffic collection device to obtain original traffic data to be analyzed; performing statistical analysis on the data of each communication address in the original traffic data within a preset time window to obtain communication feature data including the number of communications per unit time, average data packet size, number of target addresses and proportion of protocol types; using an LSTM network to perform time series modeling on a sequence of communication events arranged in chronological order for each communication address in the original traffic data to obtain behavioral feature data that characterizes the temporal variation law of communication behavior; obtaining marked fraudulent groups from a historical feature library; The communication feature data and the behavior feature data are used to calculate the similarity between the communication feature data and the behavior feature data and the feature template to obtain the feature matching degree; the deviation value between the communication feature data and the behavior feature data and the corresponding features in the normal behavior feature library is calculated to obtain the feature abnormality degree; based on the feature matching degree and the feature abnormality degree, an adaptive weight coefficient is used for weighted combination to obtain a suspicion score, wherein the adaptive weight coefficient is dynamically adjusted according to the historical warning accuracy rate; for the communication feature data and the behavior feature data corresponding to the suspicious communication address whose suspicion score exceeds the preset suspicion threshold, the gang feature description in the fraud-related gang feature template with the highest matching degree is selected, and the first warning result including the suspicious communication address, the gang suspicion level and the gang feature description is output.
[0007] By adopting the above technical solution, the embodiment of the present application constructs a multi-level fraud-related network traffic identification architecture. On this basis, static communication features and dynamic behavior features are extracted for each communication address, where the communication features reflect the overall statistical laws of communication behavior, and the behavior features capture the temporal change pattern of the communication event sequence through the LSTM network. These two types of features are compared bidirectionally with the fraud-related gang feature template and the normal behavior feature library to obtain the feature matching degree and feature abnormality, respectively, and dynamically weighted and combined through adaptive weight coefficients to form a closed-loop feature evaluation mechanism. Since the weight coefficient will be automatically adjusted according to the historical warning accuracy, the system can continuously optimize the importance distribution of the features, thereby effectively improving the recognition accuracy of the ever-changing fraud-related communication behaviors.
[0008] In combination with some embodiments of the first aspect, in some embodiments, the LSTM network is used to perform time series modeling on the communication event sequence arranged in chronological order for each communication address in the original traffic data to obtain behavioral characteristic data that characterizes the temporal change law of communication behavior, specifically including: encoding the communication events of each communication address within a preset time window in chronological order into an event vector sequence, wherein each event vector contains the target address type, communication time interval and data packet size at that moment, and the event vector sequence reflects the behavior change process of the communication address; using the LSTM network to analyze the change relationship between adjacent event vectors, extracting the communication target switching law, time interval change trend and data packet size change pattern, and obtaining a hidden layer state sequence that characterizes the dynamic change of communication behavior; performing weighted averaging on the hidden layer state sequence in the time dimension, with the weight decreasing with the time interval, to obtain the behavioral characteristic data that characterizes the evolution law of communication behavior.
[0009] By adopting the above technical solution, the embodiment of the present application establishes a refined time series feature extraction mechanism. By encoding communication events into event vectors containing multi-dimensional information, the communication behavior characteristics at each time point are fully preserved. By analyzing the changing relationship between adjacent event vectors, the LSTM network can not only capture the switching mode of the communication object, but also identify the changing trend of the time interval and the size of the data packet, thereby effectively extracting the dynamic characteristics of the communication behavior. The time-weighted average method is used for the hidden state sequence, so that recent behavior changes have a higher weight. This decreasing weight design makes the behavior feature data more accurately reflect the latest evolution trend of the communication behavior, thereby improving the real-time identification capability of fraudulent communication behavior.
[0010] In combination with some embodiments of the first aspect, in some embodiments, the method obtains a marked feature template of a fraud gang from a historical feature library, calculates the similarity between the communication feature data and the behavior feature data and the feature template, and obtains a feature matching degree, specifically including: obtaining a marked feature template of a fraud gang from the historical feature library, wherein each feature template of a fraud gang includes a communication feature template and a behavior feature template, the communication feature template includes statistical distribution parameters of the number of communications per unit time, the average data packet size, the number of target addresses, and the proportion of protocol types, and the behavior feature template includes a feature vector of a target address switching rule, a communication time interval change trend, and a data packet size change pattern; respectively calculating the Mahalanobis distance between the communication feature data and the communication feature template, and the dynamic time warping distance between the behavior feature data and the behavior feature template, wherein the dynamic time warping distance is obtained by constructing a cumulative cost matrix, recording the Euclidean distance between the feature vectors of corresponding positions of two sequences at each matrix position, and using a dynamic programming method to find the minimum cumulative cost path; calculating a comprehensive similarity based on the Mahalanobis distance and the dynamic time warping distance to obtain the feature matching degree.
[0011] By adopting the above technical solution, the embodiment of the present application constructs a two-layer feature matching framework. The communication feature template characterizes the overall communication mode of the fraud gang through statistical distribution parameters, while the behavior feature template retains the typical time series change characteristics. When matching features, the similarity of communication features is measured by the Mahalanobis distance, which takes into account the correlation between features and can more accurately measure the differences in the multidimensional feature space. For the matching of behavioral features, the dynamic time warping distance is used to allow sequences to be flexibly aligned in the time dimension. The optimal alignment method is found through the cumulative cost matrix and dynamic programming algorithm, which effectively solves the problem of asynchronous communication behavior on the time scale and improves the recognition accuracy of fraud gangs with similar behavior patterns but offset timing.
[0012] In combination with some embodiments of the first aspect, in some embodiments, based on the feature matching degree and the feature abnormality, an adaptive weight coefficient is used for weighted combination to obtain a suspicion score, specifically including: calculating the distinguishing contribution of the feature matching degree and the feature abnormality to the warning result for each record in the historical warning data, wherein the distinguishing contribution is obtained by calculating the distribution difference of the feature matching degree and the feature abnormality in the accurate warning sample and the false alarm sample; constructing a feature importance scoring model based on the distinguishing contribution, calculating the initial weight coefficients of the feature matching degree and the feature abnormality, and the sum of the weight coefficients is 1; after each round of warning result verification, adding the newly added accurate warning samples to the feature importance scoring model, updating the distinguishing contribution in real time, and dynamically adjusting the weight coefficient; using the updated weight coefficient to weighted sum the feature matching degree and the feature abnormality to obtain the suspicion score.
[0013] By adopting the above technical solution, the embodiment of the present application realizes an adaptive feature weight adjustment mechanism. By analyzing the distribution differences of feature matching degree and feature abnormality degree in accurate warning samples and false alarm samples, the distinguishing contribution of different features to the warning results is quantitatively evaluated. The feature importance scoring model constructed based on this distinguishing contribution makes the allocation of initial weights more targeted. As new warning results are continuously verified, the system can automatically absorb newly added accurate warning samples, update the distinguishing contribution of features in real time and dynamically adjust the weight coefficient, forming a continuously optimized closed-loop feedback mechanism, which effectively reduces the false alarm rate and improves the accuracy of the warning.
[0014] In combination with some embodiments of the first aspect, in some embodiments, the method also includes: for any two communication addresses in the suspicious communication address set whose suspicion score exceeds a preset suspicion threshold, calculating the first Euclidean distance between their respective corresponding communication feature data and the second Euclidean distance between their behavior feature data, when the first Euclidean distance and the second Euclidean distance are both less than the preset distance threshold, classifying the two communication addresses as the same suspicious gang; calculating the cosine similarity between the average feature vector of each suspicious gang and the feature template of the fraud-related gang, selecting the gang feature description in the feature template with the largest cosine similarity, and outputting a second warning result including a list of suspicious gang member addresses, the gang suspicion level and the gang feature description.
[0015] By adopting the above technical solution, the embodiment of the present application establishes a multi-dimensional gang identification mechanism. By simultaneously calculating the Euclidean distance of communication features and behavioral features, the similarity between communication addresses is evaluated from both static and dynamic dimensions. Only when the distances in both dimensions meet the threshold requirements are they divided into the same gang. This double constraint ensures the accuracy of gang division. For each identified gang, by calculating the cosine similarity between its average feature vector and the template, not only the numerical difference of the feature vector is considered, but also the consistency of the vector direction is paid attention to, so that the gang's crime characteristics can be matched more accurately, improving the accuracy of identifying fraud-related gangs.
[0016] In combination with some embodiments of the first aspect, in some embodiments, the method further includes: extracting all suspicious communication records within a preset time range from the original traffic data corresponding to the communication address of the first suspicious group, and sorting the suspicious communication records in chronological order; the first suspicious group is any suspicious group; performing semantic analysis on each suspicious communication record to identify key event types therein, including financial transactions, information transmission, identity authentication, and instruction issuance; classifying the suspicious communication records according to the key event types to obtain an event association graph for each category, in which nodes represent communication records and edges represent the chronological order of events. sequence; extract suspicious communication records that match the corresponding stages in the preset fraud-related gang crime pattern template from the event association diagram of each category, and store them as a first crime record sequence associated with the first suspicious gang, including identity authentication records in the information collection stage, instruction issuance records in the implementation stage, and fund transaction records in the completion stage; when the first suspicious gang is marked as a confirmed fraud-related gang, generate a structured evidence document based on the first crime record sequence, the structured evidence document contains the timestamp, communication content and related party address of each suspicious communication record in the first crime record sequence, and the list of addresses of suspicious gang members in the second early warning result.
[0017] By adopting the above technical solution, the embodiment of the present application constructs a complete evidence chain extraction framework. By semantically analyzing suspicious communication records and classifying them by event type, multiple event association diagrams are formed, which clearly show the temporal relationship of different types of events. By matching these events with the preset crime mode template, the key records of each crime stage such as information collection, implementation and completion can be accurately identified. This stage-based evidence extraction method not only ensures the integrity of the evidence, but also reflects the continuity of the crime process through temporal association. The structured evidence document finally generated contains both detailed communication records and the association relationship between gang members, providing reliable evidence support for subsequent case handling.
[0018] In combination with some embodiments of the first aspect, in some embodiments, the suspicious communication records that match the corresponding stages in the preset fraud gang crime pattern template are extracted from the event association graphs of each category and stored as the first crime record sequence associated with the first suspicious gang, specifically including: identifying records with a successful communication status in the fund transaction category in the event association graph as key evidence nodes, each key evidence node corresponding to a fund transaction; based on the address of the associated party of each key evidence node, extracting the identity authentication record and instruction issuance record corresponding to the address in the event association graphs of other categories, and combining the extracted records to form a complete chain of evidence; merging the records in all evidence chains and arranging them in ascending order by timestamp to obtain the first crime record sequence.
[0019] By adopting the above technical solution, the embodiment of the present application establishes an evidence chain construction mechanism with fund transactions as the core. By taking the successful fund transaction records as the key evidence nodes, a clear starting point for evidence extraction is established. Based on these key nodes, the related identity authentication and instruction issuance records are traced back in other categories of events through the related party addresses, which not only ensures that each evidence chain has actual fund transaction support, but also completely restores the preparation process before the transaction. Multiple evidence chains are merged in time to form a complete crime record sequence, which clearly shows the evolution of the entire crime process and improves the integrity and credibility of the evidence chain.
[0020] In a second aspect, the present application provides a server, comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the server to execute the method described in the first aspect and any possible implementation of the first aspect.
[0021] In a third aspect, the present application provides a computer program product comprising instructions, which, when executed on a server, enables the server to execute the method described in the first aspect and any possible implementation of the first aspect.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium comprising instructions, which, when executed on a server, causes the server to execute the method described in the first aspect and any possible implementation of the first aspect.
[0023] It is understandable that the server provided in the second aspect, the computer program product provided in the third aspect, and the computer-readable storage medium provided in the fourth aspect are all used to execute the method provided in the embodiment of the present application. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method, which will not be repeated here.
[0024] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Since the LSTM network is used to model the communication event sequence in time series, and the adaptive weight coefficient is used to dynamically combine the feature matching degree and the feature anomaly degree, the technical problem that the rule matching method is difficult to adapt to the ever-changing fraud methods is effectively solved, thereby achieving the technical effect of improving the accuracy of identifying fraudulent communication behaviors.
[0025] 2. By calculating the characteristic distance between suspicious communication addresses, the group is clustered, and the group feature description is output based on the similarity between the average feature vector of the group and the feature template of the fraud group. This method can effectively identify the fraud group with similar communication behavior and crime characteristics. This method realizes the automatic clustering of the fraud group through the characteristic distance measurement, and improves the efficiency of identifying the gang crime.
[0026] 3. By adopting the technical means based on fund transaction records as key evidence nodes and tracing back relevant identity authentication and instruction issuance records through the addresses of related parties, the technical problems of incomplete and lack of relevance of the evidence chain are effectively solved, thereby achieving the technical effect of improving the integrity and credibility of the evidence chain. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a structural diagram of an applicable system architecture of the fraudulent network traffic identification method based on LSTM in the embodiment of the present application; Figure 2 It is a flowchart of a method for identifying fraudulent network traffic based on LSTM in an embodiment of the present application; Figure 3A and Figure 3BIt is another flowchart of the method for identifying fraudulent network traffic based on LSTM in an embodiment of the present application; Figure 4 It is a schematic diagram of an exemplary hardware structure of a server in an embodiment of the present application. DETAILED DESCRIPTION
[0028] Figure 1 It is a structural diagram of an applicable system architecture of the LSTM-based fraud-related network traffic identification method in the embodiment of the present application.
[0029] See also Figure 1 The system includes a network traffic collection device, a server and an early warning display terminal. The network traffic collection device is deployed at the Internet exit and connected to the server through a network interface to collect network communication data; the server is connected to a storage device through a data bus, and the storage device contains a feature library. The server is also connected to the early warning display terminal through a network interface; the early warning display terminal is used to receive and display analysis results.
[0030] The terminal may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server may be implemented as an independent server or a server cluster consisting of multiple servers.
[0031] In related technologies, fraudulent network traffic can be identified by using a rule-based matching solution. For example, fixed IP address blacklists, suspicious port sequences, and protocol type combinations are set to perform real-time scanning of network traffic. However, this solution has obvious disadvantages: when criminal gangs use dynamic IPs, changing communication ports, and other methods to evade detection, fixed rules cannot adapt in time, resulting in a large number of suspicious communication behaviors not being discovered.
[0032] The advantages of the LSTM-based fraud-related network traffic identification method in the embodiment of this application can be clearly demonstrated through the following specific examples: Assume that in the network traffic monitoring system of a provincial anti-fraud center, the communication behavior of an IP address (denoted as IP_A) is monitored within 24 hours. First, the system collects the original traffic data of the IP address in real time, including: communication records with 50 different target addresses, timestamps of each communication, protocol type used (HTTP / HTTPS), and data packet size and other information.
[0033] In the preset 1-hour time window, the system statistically analyzed the communication data of IP_A and obtained the communication characteristic data: on average, it established connections with 3 new target addresses every 10 minutes, the average packet size of a single communication was 2KB, and the HTTPS protocol accounted for 80%. These statistical characteristics initially showed a certain degree of abnormality, but these static characteristics alone could not determine the possibility of fraud.
[0034] Next, the system encodes IP_A's communication events in each time window into an event vector sequence and inputs it into the LSTM network for time series modeling. By analyzing the communication behavior of 24 consecutive time windows, the LSTM network captures a significant behavior pattern: the communication process between IP_A and each target address follows a similar time series pattern - first, multiple small data packets of HTTPS communication in a short period of time (suspected information detection), then large data packets with longer intervals (suspected information acquisition), and finally frequent small data packet interactions (suspected instruction delivery). This dynamic behavior feature data shows the evolution of communication behavior.
[0035] The system obtains the marked fraud gang feature templates from the historical feature library for matching analysis. One of the fraud gang feature templates shows a highly similar behavior pattern: phased information collection and instruction delivery through the HTTPS protocol. The calculated feature matching degree is 0.85, indicating that IP_A's communication behavior is highly similar to the fraud gang template.
[0036] At the same time, the system compares IP_A's features with the normal behavior feature library. Normal network access usually shows: the communication frequency with the target address is relatively stable, the packet size is evenly distributed, and the protocol type is relatively fixed. However, IP_A's behavior obviously deviates from these features, and the calculated feature abnormality is 0.78.
[0037] The system uses an adaptive weight coefficient (dynamically adjusted based on the accuracy of historical warnings, the current feature matching weight is 0.6, and the feature anomaly weight is 0.4) to weight the two indicators and obtain a final suspicious score of 0.82, which exceeds the preset suspicious threshold of 0.75. The system then outputs the first warning result, marking IP_A as a highly suspicious fraud-related communication address, and attaches a characteristic description of the fraud gang: "a telecommunications network fraud gang characterized by information collection and remote control."
[0038] This example clearly demonstrates the creative value of this solution: through the time series modeling of the LSTM network, the system can identify dynamic behavior patterns that are difficult to detect with traditional rule matching methods; through the dual measurement of feature matching and abnormality, a more comprehensive assessment of suspicion is provided; through the dynamic adjustment of adaptive weights, rapid adaptation to changing fraud methods is achieved. In the end, this method successfully identified communication behaviors with typical fraud characteristics, which are likely to be ignored in traditional rule matching methods.
[0039] The following describes the fraudulent network traffic identification method based on LSTM in the embodiment of the present application in combination with the above system architecture and examples: See also Figure 2 , is a flow chart of a method for identifying fraudulent network traffic based on LSTM in an embodiment of the present application. The method comprises the following steps: S201, collecting network traffic data including communication address, communication time, protocol type and data packet size in real time through a network traffic collection device to obtain original traffic data to be analyzed; Among them, network traffic collection equipment refers to the data collection device deployed at the network exit, which is used to capture and record network communication data; the communication address refers to the source address and destination address involved in the communication; the communication time refers to the specific time when each communication occurs; the protocol type is used to indicate the protocol used for network communication, such as HTTP, HTTPS, etc.; the data packet size refers to the amount of data transmitted in a single communication; the raw traffic data refers to the unprocessed raw network communication records.
[0040] After the network traffic monitoring system is started, the system needs to continuously obtain network communication data as the basis for subsequent analysis. Specifically, the network traffic collection device intercepts all data packets passing through the network interface in real time, parses the header information of the data packet to extract the communication address, timestamp, protocol type and other fields, and calculates the byte size of the data packet. This information is organized into a standard format of traffic records and stored in the system's data cache area to form the original traffic data set to be analyzed.
[0041] S202, performing statistical analysis on the data of each communication address in the original traffic data within a preset time window to obtain communication characteristic data including the number of communications per unit time, the average data packet size, the number of target addresses, and the proportion of protocol types; Among them, the preset time window represents a fixed time interval for statistical analysis, such as 1 hour or 24 hours; the number of communications per unit time refers to the communication frequency within the standard time unit; the average packet size represents the arithmetic mean of all packet sizes within the time window; the number of target addresses is used to represent the number of different target addresses connected to the communication address; the proportion of protocol types refers to the percentage of various network protocols in the total number of communications; and the communication feature data represents the set of feature indicators obtained after statistical analysis of the original traffic data.
[0042] After completing the collection of raw traffic data, it is necessary to conduct preliminary statistical analysis on the data to extract communication behavior characteristics. Specifically, the system first groups the raw traffic data according to the communication address, then traverses all communication records of each communication address within the preset time window, counts the number of communications and calculates the frequency per unit time, accumulates the size of the data packet and calculates the average, counts the number of different target addresses, calculates the number of times each type of protocol is used and converts it into a percentage, and finally organizes these statistical indicators into structured communication characteristic data.
[0043] S203, using an LSTM network to perform time series modeling on a communication event sequence arranged in chronological order for each communication address in the original traffic data, to obtain behavior feature data that characterizes a time series change rule of the communication behavior; Among them, LSTM network represents long short-term memory neural network, which is a deep learning model specially used to process sequence data; communication event sequence refers to a series of communication records arranged in chronological order; time series modeling is used to represent the mathematical description of the time dependency in sequence data; behavioral feature data refers to data extracted by deep learning model that can reflect the dynamic change characteristics of communication behavior.
[0044] After obtaining the statistical characteristics of the original traffic data, it is necessary to further analyze the dynamic characteristics of communication behavior over time. Specifically, the system first sorts the communication records of each communication address according to the timestamp to form a time series data sequence, and then inputs these sequences into the pre-trained LSTM network. The time series dependency in the sequence is extracted through the multi-layer nonlinear transformation of the network, and finally a feature vector that can characterize the dynamic change law of communication behavior is obtained.
[0045] In some embodiments, LSTM-based temporal feature extraction can be implemented in a variety of ways: optionally, first feature encode the communication events, convert each event into a vector containing information such as the target address type, time interval, and packet size, and then build a multi-layer LSTM network to selectively memorize and forget historical information through a gating mechanism, and finally extract a vector representation representing the temporal features from the hidden state of the network; optionally, a bidirectional LSTM network structure is used, while considering the forward and backward dependencies of the sequence, and weighted combination of features at different times through an attention mechanism, ultimately obtaining a more comprehensive temporal feature representation. It is understandable that other methods can also be used to implement LSTM-based temporal feature extraction, which is not limited here.
[0046] S204, obtaining a marked fraud gang feature template from a historical feature database, calculating the similarity between the communication feature data and the behavior feature data and the feature template, and obtaining a feature matching degree; Among them, the historical feature library refers to the database that stores the confirmed communication features of fraud gangs; the feature template of the fraud gang refers to the verified set of communication behavior features of typical fraud gangs; the similarity is used to indicate the degree of closeness between two feature vectors; the feature matching degree indicates the degree of matching between the communication behavior to be analyzed and the known fraud pattern.
[0047] After extracting the communication features and behavior features, it is necessary to compare and analyze these features with the known fraud gang features. Specifically, the system first reads all the marked fraud gang feature templates from the historical feature library, and then calculates the similarity between the communication feature data and behavior feature data of the communication address to be analyzed and these templates, and uses a weighted method to combine the similarities of the two types of features into the final feature matching degree, which is used to measure the similarity between the behavior pattern of the communication address and the known fraud gang.
[0048] In some embodiments, the calculation of feature matching can be achieved in a variety of ways: Optionally, for communication feature data, the Mahalanobis distance is used to calculate its similarity with the template feature, which takes into account the correlation between features. For behavioral feature data, the dynamic time warping distance is used to calculate the sequence similarity, and finally the comprehensive matching degree is obtained by weighted average; Optionally, a multi-layer perceptron network is constructed, and the features to be analyzed and the template features are respectively input into the network, and the nonlinear relationship between the features is learned through the network. The output layer uses the sigmoid function to map the similarity to between 0 and 1, indicating the degree of matching. It is understandable that other methods can also be used to calculate feature matching, which are not limited here.
[0049] S205, calculating the deviation values of the communication feature data and the behavior feature data from the corresponding features in the normal behavior feature library to obtain the feature abnormality degree; Among them, the normal behavior feature library refers to the database that stores normal network communication behavior features; the deviation value refers to the degree of difference between the feature to be analyzed and the normal behavior feature; the feature abnormality indicates the degree to which the communication behavior deviates from the normal pattern; and the corresponding features are used to represent feature data of the same type.
[0050] After completing the match with the fraud-related template, it is also necessary to evaluate the degree of difference between the communication behavior and the normal behavior. Specifically, the system obtains the corresponding type of feature data from the normal behavior feature library, and calculates the deviation between the communication feature data and the behavior feature data of the communication address to be analyzed and the normal features, and combines these deviation values to obtain the feature abnormality, which is used to quantify the degree of deviation between the behavior pattern of the communication address and the normal communication behavior.
[0051] S206. Based on the feature matching degree and the feature abnormality degree, an adaptive weight coefficient is used for weighted combination to obtain a suspicion score, wherein the adaptive weight coefficient is dynamically adjusted according to the historical warning accuracy rate; Among them, the adaptive weight coefficient represents the dynamically adjusted feature combination weight; the historical warning accuracy rate refers to the proportion of correct warnings confirmed in the system's historical warning results; the suspicion score represents the comprehensive suspicion level of the communication behavior; and dynamic adjustment is used to indicate that the weight coefficient will be continuously updated as the system runs.
[0052] After obtaining the feature matching degree and feature anomaly degree, these two indicators need to be reasonably combined to obtain the final evaluation result. Specifically, the system analyzes the contribution of feature matching degree and feature anomaly degree to the warning results based on the accurate warning samples and false alarm samples in the historical warning data, and determines the initial weight coefficient accordingly. Then, as new warning results are continuously verified, the system automatically adjusts the weight coefficient to improve the accuracy of the warning, and finally obtains a score reflecting the suspicious degree of the communication behavior through weighted combination.
[0053] S207. For the communication feature data and behavior feature data corresponding to the suspicious communication address whose suspicion score exceeds the preset suspicion threshold, select the gang feature description in the fraud gang feature template with the highest matching degree, and output the first warning result including the suspicious communication address, the gang suspicion level and the gang feature description.
[0054] Among them, the preset suspicious threshold represents the scoring limit for the system to judge the communication behavior as suspicious; the gang characteristic description refers to the textual description of the communication behavior characteristics of the fraud gang; the gang suspicion level is used to represent the grading result of the degree of suspicion; the first warning result represents the system's warning output for a single suspicious communication address.
[0055] After calculating the suspicion score, it is necessary to generate warning information for communication behaviors that exceed the threshold. Specifically, the system first compares the suspicion score with the preset suspicion threshold. For communication addresses that exceed the threshold, it re-searches the matching degree with each fraud gang feature template, selects the gang feature description in the template with the highest matching degree, and determines the gang's suspicion level based on the suspicion score. Finally, this information is organized into a standard format for warning result output.
[0056] In some embodiments, the generation of early warning results can be achieved in a variety of ways: optionally, a multi-level early warning level system is constructed, and the degree of suspicion is divided into three levels: high, medium, and low according to the interval range of the suspicion score. For each suspicious communication address, the most matching gang feature description is extracted, and structured early warning information is generated by combining the suspicion level and feature description; optionally, an early warning template system is used to pre-define early warning description templates for different types of fraud-related behaviors, and the key information fields in the template are automatically filled according to the matched gang characteristics and suspicion level to generate standardized early warning results. It is understandable that other methods can also be used to generate early warning results, which are not limited here.
[0057] In the embodiment of the present application, by constructing a feature template library for fraud gangs and combining a dual evaluation mechanism of feature matching degree and feature abnormality degree, accurate identification of suspicious communication behaviors is achieved, effectively solving the technical problem of single feature matching being prone to misjudgment in the prior art, and significantly improving the accuracy of identifying fraud gangs.
[0058] In order to better illustrate the technical solution of the embodiment of the present application, another embodiment of the present application is described in detail below in combination with more specific implementation and detailed processing flow.
[0059] See also Figure 3A and Figure 3B , which is another flow chart of the fraud-related network traffic identification method based on LSTM in an embodiment of the present application.
[0060] S301, collecting network traffic data including communication address, communication time, protocol type and data packet size in real time through a network traffic collection device to obtain original traffic data to be analyzed; For example, in the network monitoring system of a provincial anti-fraud center, a network traffic collection device model NTA-6000 is deployed, which collects communication data of the network exit in real time through the mirror port. From 10:00:00 to 11:00:00 on January 15, 2024, the device captured the original traffic data from the IP address "192.168.1.100", including source IP, destination IP, communication timestamp, protocol type (HTTPS 80%, HTTP 20%) and packet size (200 bytes to 5KB) and other information.
[0061] S302, performing statistical analysis on the data of each communication address in the original traffic data within a preset time window to obtain communication characteristic data including the number of communications per unit time, the average data packet size, the number of target addresses, and the proportion of protocol types; Continuing with the above example, the system uses a 10-minute sliding time window to perform statistical analysis on the raw traffic data of the IP address "192.168.1.100", and obtains the following communication characteristic data: the IP establishes 5 connections with 3 different target addresses on average per minute; the average data packet size is 2KB, of which the average packet size for communication with "203.0.113.50" is 4KB, and the average size for communication with other addresses is 1.5KB; connections are established with an average of 7 different IP addresses in each 10-minute window; in terms of the proportion of protocol types, the HTTPS protocol accounts for 80% (mainly used for large data packet communication with "203.0.113.50"), and the HTTP protocol accounts for 20% (mainly used for small data packet communication with other addresses).
[0062] In some embodiments, the communication characteristic data obtained may also contain many other special data: for example, the proportion of encrypted traffic, scanning behavior indicators reflecting the frequency of attempting to connect to multiple target addresses in a short period of time, connection retry rate, data packet size distribution, address geographical distribution and other characteristics.
[0063] Preferably, in some embodiments, in order to further improve the accuracy of identifying fraud-related gangs, the communication feature data also includes the zero-load packet ratio used to count the proportion of control packets of remote control behaviors without data load, and the number of concurrent connections used to reflect the number of active connections maintained by the fraud-related gangs in communicating with multiple targets at the same time.
[0064] Specifically, statistics on the proportion of zero-load packets may include: extracting the load information of each data packet from the original traffic data, dividing the data packets into two categories: zero-load packets (such as TCP ACK packets, SYN packets and other control packets) and non-zero-load packets, and recording the timestamp of each zero-load packet; performing weighted statistics on the proportion of zero-load packets of each communication address within the preset time window according to the divided time period, and different time periods have different weights. For example, 24 hours are divided into 12 time periods, and each time period is assigned a different benchmark weight, where the weight of zero-load packets in the late night period (such as 2-4 am) is 1.5, because there are fewer normal businesses in this period; the weight of zero-load packets in the working period (such as 9 am to 6 pm) is 0.8, because there are more heartbeat packets generated by normal business in this period.
[0065] For the number of concurrent connections, the time overlap of concurrent connections can be counted, specifically, it can include: extracting TCP / UDP session information from the original traffic data, including the start time, end time, source address, destination address and other fields of the session, and building a unique session identifier based on these fields; arranging all sessions of each communication address in the preset time window according to the time axis, calculating the length of the time overlap interval of any two sessions, and counting the number of active sessions that exist at each moment; calculating the maximum number of concurrent connections, the average number of concurrent connections, and the duration distribution of concurrent connections in the time window based on the statistical results, where the duration is classified and counted according to short connections (less than 1 minute), medium connections (1-10 minutes) and long connections (greater than 10 minutes), and finally obtaining concurrent connection characteristic indicators that reflect the collaborative communication behavior of the group.
[0066] S303, encoding the communication events of each communication address within the preset time window into an event vector sequence in chronological order, wherein each event vector includes the target address type, communication time interval and data packet size at that moment, and the event vector sequence reflects the behavior change process of the communication address; Among them, the event vector represents a vector representation that encodes the communication event features into numerical form; the target address type refers to the address category of the communication object; the communication time interval represents the time difference between two adjacent communications; and the event vector sequence is used to represent multiple event vectors arranged in chronological order.
[0067] In some embodiments, vectorized encoding of communication events may be implemented in the following manner: Optionally, a feature engineering method is used to type encode the target address, map the IP address segment to a predefined address type label, calculate the time difference between adjacent communication events, and normalize the data packet size; Alternatively, an autoencoder network is used to learn compact representations of features through a multi-layer neural network.
[0068] Continuing with the above example, the system encodes the communication events of the IP address "192.168.1.100" in the time window from 10:00:00 to 10:10:00. The specific process is as follows: First, "203.0.113.50" is mapped to type 1 (high-frequency communication object), and other addresses are mapped to type 2 (medium frequency) or type 3 (low frequency) according to the communication frequency; then the time difference from the last communication is recorded, such as [2 seconds, 5 seconds, 1 second, 3 seconds, etc.]; at the same time, the data packet size of each communication is recorded, such as [4KB, 1.5KB, 1.5KB, 4KB, etc.]; finally, the event vector sequence in the following form is obtained: [1, 2 seconds, 4KB]→[2, 5 seconds, 1.5KB]→[3, 1 second, 1.5KB]→[1, 3 seconds, 4KB], etc.
[0069] S304, using the LSTM network to analyze the change relationship between adjacent event vectors, extract the communication target switching law, time interval change trend and data packet size change pattern, and obtain a hidden state sequence that characterizes the dynamic change of communication behavior; Among them, the changing relationship between adjacent event vectors represents the changing characteristics of continuous communication events; the communication target switching law refers to the timing pattern of communication with different target addresses; the time interval change trend represents the dynamic change of communication frequency; the data packet size change pattern is used to represent the changing characteristics of data transmission volume; the hidden layer state sequence refers to the state vector inside the LSTM network that represents the sequence characteristics.
[0070] In some embodiments, dynamic feature extraction based on LSTM can be implemented in the following ways: Optionally, build a multi-layer LSTM network, where each layer focuses on temporal features of different scales, and combine the features of each layer through residual connections; Optionally, an attention mechanism is introduced into the LSTM network to assign different importance weights to event vectors at different times.
[0071] Continuing with the above example, the system uses a three-layer LSTM network (128 hidden units in each layer) to process the event vector sequence of "192.168.1.100" and extracts the following dynamic features: the IP follows a cyclic communication target switching pattern of "1→2→3→1"; in terms of time intervals, the communication interval with high-frequency targets is stable at 2-3 seconds, and the communication interval with other targets fluctuates between 1-5 seconds; in terms of data packet size, large packets of 4KB are maintained for communication with high-frequency targets, while small packets of about 1.5KB are maintained for communication with other targets.
[0072] S305, performing weighted averaging on the hidden layer state sequence in the time dimension, with the weight decreasing with the time interval, to obtain the behavior feature data that describes the evolution law of the communication behavior; Among them, weighted average means assigning different weights to each state vector in the sequence for averaging; the time dimension refers to the time axis direction of the sequence data; the weight decreases with the time interval, indicating that the state at a closer moment has a higher weight; the behavioral feature data is used to represent the timing characteristics of the communication behavior.
[0073] After obtaining the hidden state sequence of the LSTM network, these time series features need to be integrated into a feature vector of fixed dimension. Specifically, the system calculates the weight coefficient for each state vector in the hidden state sequence according to the time interval between its corresponding moment and the current moment. The larger the time interval, the smaller the weight. Then, all state vectors are weighted averaged according to the weights to obtain a vector representation that can comprehensively characterize the time series features of communication behavior.
[0074] Continuing with the above example, the system uses the exponential decay function w(t) = exp(-t / τ) to calculate the weight, where τ is the characteristic decay time constant of 300 seconds. The state vectors in the last 5 minutes are weighted by time, and the state weights in the last minute range from 1.0 to 0.98, decreasing in sequence; through weighted averaging, a 128-dimensional behavior feature vector is finally obtained, in which the first 32 dimensions encode the switching mode of the communication target, the middle 48 dimensions represent the change law of the time interval, and the last 48 dimensions describe the distribution characteristics of the packet size.
[0075] S306. Obtain a marked fraud gang feature template from the historical feature library, wherein each fraud gang feature template includes a communication feature template and a behavior feature template, wherein the communication feature template includes statistical distribution parameters of the number of communications per unit time, the average data packet size, the number of target addresses, and the proportion of protocol types, and the behavior feature template includes a target address switching rule, a communication time interval change trend, and a feature vector of a data packet size change pattern; Among them, the characteristic template of the fraud gang represents the verified characteristic description of a typical fraud gang; the statistical distribution parameter refers to the numerical indicator that describes the statistical law of the characteristic; the characteristic vector is used to represent the numerical representation of the behavioral characteristics; the communication characteristic template and the behavioral characteristic template correspond to the static statistical characteristics and the dynamic time series characteristics respectively.
[0076] For example, the characteristic template of the fraud gang numbered "FT-2023001" contains the following information: in terms of communication characteristic template, the average number of communications per unit time is 5.2 times (standard deviation 0.8), the average data packet size is 2.1KB (standard deviation 0.3KB), the average number of target addresses is 8 (standard deviation 2), HTTPS accounts for 75-85%, and HTTP accounts for 15-25%; in terms of behavioral characteristic template, a 128-dimensional vector is used to describe the cyclic switching pattern of "master control node→data node→detection node→master control node".
[0077] S307, respectively calculating the Mahalanobis distance between the communication feature data and the communication feature template, and the dynamic time warping distance between the behavior feature data and the behavior feature template, wherein the dynamic time warping distance is obtained by constructing a cumulative cost matrix, recording the Euclidean distance between the feature vectors of the corresponding positions of the two sequences at each matrix position, and using a dynamic programming method to find the minimum cumulative cost path; Among them, Mahalanobis distance represents a distance measurement method that takes into account feature correlation; Euclidean distance refers to the straight-line distance between two points in the vector space; the cumulative cost matrix is used to record the cumulative distance value in the sequence alignment process; and the minimum cumulative cost path represents the path corresponding to the optimal alignment of the two sequences.
[0078] After obtaining the feature template, it is necessary to calculate the similarity between the feature to be analyzed and the template feature. Specifically, the system uses the Mahalanobis distance to measure static communication features, which takes into account the correlation and scale differences between features; for dynamic behavioral features, the dynamic time warping distance is used for measurement. By constructing a cumulative cost matrix and using a dynamic programming algorithm to find the optimal alignment path, the system effectively handles the problems of inconsistent sequence lengths and asynchronous time scales.
[0079] In some embodiments, the calculation of feature distance can be achieved in multiple ways: Optionally, for the Mahalanobis distance calculation of the communication features, the covariance matrix of the features is first estimated, and then the standardized distance between the feature vectors is calculated based on the covariance matrix. This method can handle the correlation and dimensionality differences between the features; Optionally, for dynamic time warping distance calculation of behavioral features, first define a local distance measurement method to calculate the feature similarity of corresponding positions in the sequence, then use a dynamic programming algorithm to find the optimal path in the cumulative cost matrix, and finally obtain the sequence distance that takes time elasticity into account.
[0080] It is understandable that other methods may be used to calculate the feature distance, which is not limited here.
[0081] Continuing with the above example, the system calculates the distance between the feature of the IP address "192.168.1.100" and the "FT-2023001" template: For the communication feature, the feature vector X=[5.0, 2.0, 7.0, 0.8] is first constructed (representing the number of communications, packet size, number of addresses, and HTTPS proportion, respectively), and the template mean vector μ=[5.2, 2.1, 8.0, 0.8] is then calculated based on the covariance matrix Σ (a 4×4 matrix containing the correlation between the features) to obtain the Mahalanobis distance d=1.2; for the behavior feature, the system constructs a 10×10 cumulative cost matrix, and uses the dynamic programming method to calculate the optimal alignment path of the two 128-dimensional feature vector sequences, and finally obtains a DTW distance of 1.5. The above distance calculation results show that the behavioral characteristics of the IP address have a high similarity with the template of the fraud gang. This is because the Mahalanobis distance is 1.2 and the dynamic time warping distance is 1.5, both of which are lower than the typical thresholds of normal communication behavior (Mahalanobis distance threshold 3.0, dynamic time warping distance threshold 4.0).
[0082] S308, calculating the comprehensive similarity based on the Mahalanobis distance and the dynamic time warping distance to obtain the feature matching degree; Among them, the comprehensive similarity refers to the overall similarity obtained by combining multiple distance metrics; the feature matching degree refers to the overall matching degree between the feature to be analyzed and the template feature; the calculation of the comprehensive similarity is used to unify the distance measurement results of different types to the same metric standard.
[0083] After obtaining the distance measurement results of communication features and behavioral features respectively, it is necessary to reasonably combine these results to obtain the overall matching degree. Specifically, the system first normalizes the Mahalanobis distance and dynamic time warping distance to unify their value ranges to the interval [0,1], then sets the weight coefficient according to the importance of the two features, calculates the comprehensive similarity through weighted combination, and finally obtains the matching index that reflects the overall matching degree of the features.
[0084] In some embodiments, the calculation of comprehensive similarity can be achieved in a variety of ways: Optionally, a distance-based similarity conversion method is used to convert the normalized Mahalanobis distance and dynamic time warping distance into similarity values through an exponential function, and then the two similarities are combined using a weighted average method, and the weight can be determined to an optimal value through cross-validation; Optionally, a multi-layer perceptron network is constructed, and the two distances are used as input features. The network learns the optimal feature combination method, and the output layer uses the sigmoid function to map the result to the [0,1] interval to represent the comprehensive matching degree.
[0085] It is understandable that other methods may be used to calculate the comprehensive similarity, which is not limited here.
[0086] Continuing with the above example, the system converts the Mahalanobis distance and dynamic time warping distance into comprehensive similarity: first, the distance is normalized, the Mahalanobis distance d1=1.2, the reference threshold T1=3.0, and after normalization s1=exp(-d1 / T1)=0.67; the DTW distance d2=1.5, the reference threshold T2=4.0, and after normalization s2=exp(-d2 / T2)=0.69; then a weighted combination is performed, the communication feature weight w1=0.4 (static feature), the behavior feature weight w2=0.6 (dynamic feature is more important), and finally the feature matching degree is obtained = w1×s1+w2×s2=0.4×0.67+0.6×0.69=0.682, indicating that the communication behavior of the IP address has a similarity of 68.2% with the fraud gang template "FT-2023001".
[0087] S309, calculating the deviation values of the communication feature data and the behavior feature data from the corresponding features in the normal behavior feature library to obtain the feature abnormality degree; Among them, the normal behavior feature library represents a data set that stores normal network communication behavior features; the deviation value refers to the degree of difference between the feature to be analyzed and the normal behavior feature; the feature abnormality is used to indicate the degree to which the communication behavior deviates from the normal pattern; and the corresponding feature represents feature data of the same type.
[0088] After completing the matching analysis with the fraud template, it is also necessary to evaluate the degree of difference between the communication behavior and the normal behavior. Specifically, the system obtains the corresponding type of feature data from the normal behavior feature library, and calculates the deviation between the communication feature data and the behavior feature data of the communication address to be analyzed and the normal features, and combines these deviation values to obtain the feature abnormality, which is used to quantify the degree of deviation between the behavior pattern of the communication address and the normal communication behavior.
[0089] In some embodiments, the calculation of feature abnormality can be achieved in a variety of ways: Optionally, based on the statistical distribution of normal behavior characteristics, the Z score of the feature to be analyzed is calculated, which represents the multiple of the standard deviation of the feature from the mean. After the Z scores of the communication characteristics and the behavior characteristics are calculated respectively, the comprehensive abnormality is obtained through weighted combination; Optionally, a local anomaly factor algorithm is used to calculate the local density ratio of the feature to be analyzed relative to its k nearest neighbors. The larger the density ratio, the more abnormal the feature. Finally, the anomaly scores of different features are normalized and combined to obtain the feature anomaly degree.
[0090] It is understandable that other methods may be used to calculate the characteristic abnormality, which is not limited here.
[0091] Continuing with the above example, the system extracts the baseline features from the normal behavior feature library and calculates the feature deviation of the IP address "192.168.1.100": For the communication feature deviation, based on the Z score, the number of communications per unit time Z1=(5.0-3.2) / 0.5=3.6, the average packet size Z2=(2.0-1.5) / 0.2=2.5, the number of target addresses Z3=(7.0-4.0) / 1.0=3.0, and the proportion of protocol types Z4=(0.8-0.6) / 0.1=2.0, The communication characteristic abnormality is 2.775. For the behavioral characteristic deviation, based on the local abnormality factor calculation, k=5 nearest neighbors are selected, and the local reachable density of the current IP is calculated to be lrd=0.3, the average local reachable density of the neighbors is lrd_ref=0.8, and the behavioral characteristic abnormality is 2.667. The final comprehensive abnormality calculation result is: 0.4×2.775+0.6×2.667=2.71, indicating that the behavior of the IP address deviates significantly from the normal mode (it is generally believed that an abnormality >2.0 indicates an obvious abnormality).
[0092] S310, calculating the distinguishing contribution of the feature matching degree and the feature abnormality degree to the warning result for each record in the historical warning data, wherein the distinguishing contribution is obtained by calculating the distribution difference of the feature matching degree and the feature abnormality degree in the accurate warning sample and the false alarm sample; Among them, historical warning data refers to the set of warning records generated by the system in history; accurate warning samples refer to warning records that are confirmed to be correct; false alarm samples refer to warning records that are confirmed to be wrong; and discrimination contribution is used to represent the ability of features to distinguish between correct warnings and false alarms.
[0093] After obtaining the feature matching degree and feature anomaly degree, it is necessary to evaluate the contribution of these two types of features to the accuracy of warnings. Specifically, the system first extracts accurate warning samples and false alarm samples from historical warning data, and then analyzes the distribution characteristics of these two types of samples in terms of feature matching degree and feature anomaly degree. By calculating the degree of difference between the distributions, the contribution of each type of feature to distinguishing correct warnings from false alarms is quantified.
[0094] In some embodiments, the calculation of the differentiated contribution can be achieved in a variety of ways: Optionally, the information gain method is used to divide the feature value into multiple intervals according to the threshold value, and the information gain brought by each feature in distinguishing accurate warning samples from false alarm samples is calculated. The larger the gain value, the greater the distinguishing contribution of the feature. Optionally, Fisher discriminant analysis is used to calculate the ratio of the between-class variance to the within-class variance of the feature in the two categories of accurate warning samples and false alarm samples. The larger the ratio, the stronger the distinguishing ability of the feature.
[0095] It is understandable that other methods may be used to calculate the differentiated contributions, which are not limited here.
[0096] Continuing with the above example, the system analyzes the most recent 1,000 historical warning records (700 of which are accurate warnings and 300 are false alarms), and analyzes the distribution of feature matching. The mean of accurate warning samples is μ1=0.65, the standard deviation is σ1=0.12, the mean of false alarm samples is μ2=0.45, the standard deviation is σ2=0.15, and the distribution difference D1=(μ1-μ2) / √((σ1²+σ2²) / 2)=1.47 is calculated; the distribution of feature anomaly is analyzed, and the mean of accurate warning samples is μ3=2.50 , standard deviation σ3=0.40, false alarm sample mean μ4=1.80, standard deviation σ4=0.45, calculated distribution difference D2=(μ3-μ4) / √((σ3²+σ4²) / 2)=1.65; finally calculated distinction contribution: feature matching distinction contribution C1=D1 / (D1+D2)=0.471, feature abnormality distinction contribution C2=D2 / (D1+D2)=0.529, analysis shows that the contribution of feature abnormality in distinguishing accurate warnings from false alarms is slightly higher than that of feature matching.
[0097] S311, constructing a feature importance scoring model based on the distinction contribution, calculating the initial weight coefficients of the feature matching degree and the feature abnormality degree, and the sum of the weight coefficients is 1; Among them, the feature importance scoring model represents a mathematical model used to evaluate the importance of features; the initial weight coefficient refers to the initial weight coefficient when the features are combined; the sum of the weight coefficients is 1, which represents the normalization constraint of all weight coefficients; the feature importance is used to indicate the degree of influence of different features on the warning results.
[0098] After obtaining the distinguishing contribution of the feature, it is necessary to establish a reasonable feature importance evaluation mechanism. Specifically, the system builds a scoring model based on the distinguishing contribution of the feature. The model takes into account the performance of the feature in historical data, quantifies the importance of the feature through mathematical modeling, and finally calculates the initial weight coefficients of feature matching and feature anomaly, which meet the normalization constraints.
[0099] In some embodiments, feature importance evaluation can be achieved in a variety of ways: Optionally, an entropy-based feature importance assessment method is used to convert the distinguishing contribution into information entropy, and the initial weight is obtained through normalization. The feature with a larger entropy value obtains a higher weight, ensuring that important features play a greater role in the combination; Optionally, a logistic regression model is used, with the distinguishing contribution as a feature and the warning result as a label for training, and the feature coefficients are extracted from the trained model and normalized by the softmax function to obtain the initial weights.
[0100] It is understandable that other methods may be used to implement feature importance evaluation, which is not limited here.
[0101] Continuing with the above example, the system builds a feature importance scoring model based on distinguishing contributions: first calculate the base weights, feature matching weight w1_base=C1=0.471, feature anomaly weight w2_base=C2=0.529; then adjust based on historical accuracy, feature matching historical accuracy p1=0.85, feature anomaly historical accuracy p2=0.88, calculate adjustment coefficients k1=p1 / (p1+p2)=0.491, k2=p2 / (p1+p2)=0.509; the final initial weights are: feature matching weight w1=(w1_base×k1) / (w1_base×k1+w2_base×k2)=0.465, feature abnormality weight w2=(w2_base×k2) / (w1_base×k1+w2_base×k2)=0.535, verification w1+w2=1.000, these weights reflect the relative importance of the two types of features in early warning judgment.
[0102] S312. After each round of early warning result verification, the newly added accurate early warning samples are added to the feature importance scoring model, the distinction contribution is updated in real time, and the weight coefficient is dynamically adjusted; Among them, warning result verification means manual confirmation of the warning generated by the system; newly added accurate warning samples refer to the newly verified correct warning records; real-time update means timely incorporating the information of new samples into the calculation; dynamic adjustment is used to indicate that the weight coefficient will be updated with the addition of new samples.
[0103] During the operation of the system, it is necessary to continuously absorb new warning verification results to optimize feature weights. Specifically, when new warning results are verified as accurate warnings, the system adds these new samples to the feature importance scoring model, recalculates the distinguishing contribution of the features, and adjusts the weight coefficients of feature matching and feature anomaly according to the updated contribution values to achieve dynamic optimization of weights.
[0104] In some embodiments, dynamic adjustment of weights can be achieved in a variety of ways: Optionally, an incremental learning method is used. When a new accurate warning sample is added, only the statistics related to the new sample are updated, and then the distinction contribution is recalculated based on the updated statistical features. The weight coefficient is adjusted by the gradient descent method to maintain the real-time performance of the model; Optionally, an online learning algorithm is used to model the weight adjustment as a sequential decision problem. Whenever there is a new warning verification result, the model parameters are updated according to whether the prediction is accurate, and the weight coefficient is optimized by minimizing the cumulative loss.
[0105] It is understandable that other methods may be used to achieve dynamic adjustment of weights, which are not limited here.
[0106] Continuing with the above example, in the latest round of early warning verification, the system added verification information of 50 early warning results (including 35 accurate warnings and 15 false alarms). After analyzing the newly added samples, the feature matching degree distribution is accurate samples (0.67±0.11) and false alarm samples (0.43±0.14), and the feature anomaly degree distribution is accurate samples (2.55±0.38) and false alarm samples (1.75±0.42). When updating the distinguishing contribution, the new difference degree of feature matching degree D1_new = 1.52 (an increase from the original value of 1.47), and the new difference degree of feature anomaly degree D2_new = 1.63 (an increase from the original value of 1.6 5 slightly decreased), the updated distinction contribution is C1_new=0.482, C2_new=0.518; in terms of dynamic adjustment of weights, the feature matching accuracy in the newly added samples is p1_new=35 / 50=0.70, the feature anomaly accuracy is p2_new=38 / 50=0.76, and the historical weight attenuation factor α=0.9 is used. The final updated weights are: feature matching weight w1_new=α×0.465+(1-α)×0.482=0.467, feature anomaly weight w2_new=α×0.535+(1-α)×0.518=0.533.
[0107] S313, using the updated weight coefficient to perform weighted summation on the feature matching degree and the feature abnormality degree to obtain the suspicion score; Among them, weighted sum means linearly combining multiple features according to weights; suspicion score refers to a numerical indicator that comprehensively reflects the degree of suspicion of communication behavior; the updated weight coefficient represents the latest weight value after dynamic adjustment.
[0108] After completing the update of the weight coefficient, it is necessary to reasonably combine the scores of different features to obtain the final result. Specifically, the system uses the updated weight coefficient to perform a weighted summation operation on the feature matching degree and feature anomaly degree to obtain a comprehensive suspicion score, which takes into account both the similarity between the communication behavior and the known fraud-related patterns and reflects the degree of deviation from normal behavior.
[0109] In some embodiments, weighted combination of features can be achieved in a variety of ways: Optionally, a linear weighting method is used to multiply the feature matching degree and the feature abnormality degree by the corresponding weight coefficients and then add them together to obtain a normalized suspiciousness score, where a higher score indicates a more suspicious communication behavior; Optionally, a nonlinear combination method is used to capture the synergistic effect between features by designing feature interaction terms, and the importance of the interaction terms is determined according to the updated weights, and finally a suspicion score that takes feature correlation into account is obtained.
[0110] It is understandable that other methods may be used to implement weighted combination of features, which are not limited here.
[0111] Continuing with the above example, the system uses the updated weights to calculate the suspicion score of the IP address "192.168.1.100": first prepare the feature values, feature matching degree M=0.682 (the degree of matching with the template of the fraud gang), feature anomaly degree A=2.71 (the degree of deviation from normal behavior); then perform feature normalization, the feature matching degree remains the original value (already in the range [0,1]), and the feature anomaly degree is normalized to A'=min(1,A / 3)=min(1,2.71 / 3)=0.903; finally perform weighted combination, feature matching degree weight w1=0.467, feature anomaly degree weight w2=0.533, and suspicion score S=w1×M+w2×A'=0.467×0.682+0.533×0.903=0.800. This score exceeds the system preset suspicion threshold of 0.75, indicating that the IP address has a high possibility of being involved in fraud.
[0112] S314, for the communication feature data and behavior feature data corresponding to the suspicious communication address whose suspicion score exceeds the preset suspicion threshold, select the gang feature description in the fraud gang feature template with the highest matching degree, and output the first warning result including the suspicious communication address, the gang suspicion level and the gang feature description; Among them, the preset suspicious threshold represents the scoring limit for the system to judge the communication behavior as suspicious; the gang characteristic description refers to the textual description of the communication behavior characteristics of the fraud gang; the gang suspicion level is used to represent the grading result of the degree of suspicion; the first warning result represents the system's warning output for a single suspicious communication address.
[0113] After calculating the suspicion score, it is necessary to generate warning information for communication behaviors that exceed the threshold. Specifically, the system first compares the suspicion score with the preset suspicion threshold. For communication addresses that exceed the threshold, it re-searches the matching degree with each fraud gang feature template, selects the gang feature description in the template with the highest matching degree, and determines the gang's suspicion level based on the suspicion score. Finally, this information is organized into a standard format for warning result output.
[0114] In some embodiments, the generation of early warning results can be achieved in a variety of ways: Optionally, a multi-level warning level system is constructed to divide the degree of suspicion into three levels: high, medium, and low according to the interval range of the suspicion score. For each suspicious communication address, the most matching gang feature description is extracted, and structured warning information is generated by combining the suspicion level and feature description; Optionally, an early warning template system is used to predefine early warning description templates for different types of fraud-related behaviors, and automatically fill in key information fields in the template based on the matched gang characteristics and degree of suspicion to generate standardized early warning results.
[0115] It is understandable that other methods may be used to generate warning results, which are not limited here.
[0116] Continuing with the above example, since the suspicion score of the IP address "192.168.1.100" is 0.800, which exceeds the preset threshold of 0.75, the system generates the first warning result: the suspicious communication address information includes the IP address 192.168.1.100, the first discovery time 2024-01-15 10:00:00, and the continuous monitoring time of 60 minutes; the gang suspicion level determination is divided according to the suspicion score range (0.75-0.85 is moderate suspicion Level 2, 0.85-0.95 is highly suspicious Level 3, and >0.95 is extremely suspicious Level 4). The current score of 0.800 corresponds to moderate suspicion (Level 2). 2); Description of gang characteristics (from template FT-2023001), including crime characteristics (adopting a layered architecture of "master control node-data node-detection node"), communication mode (stable communication with the master control node for large data packets, and detection communication with other nodes for small data packets), and behavioral characteristics (having an obvious periodic communication pattern, suspected to be used for remote control and data transmission).
[0117] S315, for any two communication addresses in the suspicious communication address set whose suspicion scores exceed a preset suspicion threshold, calculate a first Euclidean distance between their corresponding communication feature data and a second Euclidean distance between their corresponding behavior feature data, and when both the first Euclidean distance and the second Euclidean distance are less than a preset distance threshold, classify the two communication addresses as the same suspicious group; Among them, the suspicious communication address set represents all communication addresses judged to be suspicious; the first Euclidean distance refers to the distance measurement between communication feature vectors; the second Euclidean distance represents the distance measurement between behavior feature vectors; the preset distance threshold is used to represent the distance limit for judging whether two communication addresses belong to the same gang.
[0118] Continuing with the above example, the system analyzes three suspicious IP addresses: - IP1 (192.168.1.100) and IP2 (192.168.1.101): d1 = 0.42, d2 = 0.38 - IP1 and IP3 (192.168.1.102): d1=0.36, d2=0.41 - IP2 and IP3: d1=0.45, d2=0.44 All distances are less than the threshold 0.5 and are classified into the same group.
[0119] S316, calculating the cosine similarity between the average feature vector of each suspicious gang and the feature template of the fraud gang, selecting the gang feature description in the feature template with the largest cosine similarity, and outputting a second warning result including a list of addresses of suspicious gang members, the gang suspicion level, and the gang feature description; Among them, the average feature vector represents the mean of the characteristics of all gang members; cosine similarity refers to an indicator that measures the degree of similarity between the directions of two vectors; the gang member address list is used to represent all communication addresses belonging to the same gang; the second warning result represents the system's overall warning output for suspicious gangs.
[0120] After completing the division of suspicious gangs, the overall characteristics of each gang need to be analyzed. Specifically, the system first calculates the feature mean of all members in each suspicious gang to obtain the average feature vector of the gang, then calculates the cosine similarity between the vector and each template in the feature template library of fraud gangs, selects the gang feature description corresponding to the template with the highest similarity, and finally generates an early warning result containing gang member information, suspicious level and feature description.
[0121] In some embodiments, the matching and early warning of gang characteristics can be achieved in a variety of ways: Optionally, a weighted average method is used to calculate the gang characteristics, and a weight is set according to the suspiciousness score of each member to highlight the characteristic contribution of highly suspicious members. Then, the most matching gang characteristic template is found through cosine similarity to generate gang-level warning information; Optionally, an integrated matching method is used to match the characteristics of each member of the gang with the template library respectively, the number of matches of different templates is counted, and the template with the highest matching frequency and the largest average similarity is selected as the source of the gang feature description.
[0122] It is understandable that other methods may be used to achieve gang feature matching and early warning, which are not limited here.
[0123] Continuing with the above example, the system performs feature matching analysis on the identified suspicious groups: first, the average feature vector of the group is calculated. The average communication feature value is [4.9, 2.0, 7.0, 0.80], and the average behavior feature value is the arithmetic average of the behavior feature vectors of the three IPs. Then, the cosine similarity is calculated with the templates in the historical feature template library. The communication feature similarity of template FT-2023001 is 0.985, the behavior feature similarity is 0.942, and the comprehensive similarity is 0.960. The communication feature similarity of template FT-2023002 is 0.876, and the behavior feature similarity is 0.942. The feature similarity is 0.834, the comprehensive similarity is 0.852, the communication feature similarity of template FT-2023003 is 0.823, the behavior feature similarity is 0.795, and the comprehensive similarity is 0.807; finally, the second warning result is generated: the list of suspicious gang members includes 192.168.1.100 (master control node, suspicious degree 0.800), 192.168.1.101 (data node, suspicious degree 0.785), 192.168.1.102 (detection node, suspicious degree 0.790), and the gang's suspicious level is Level 2 (moderately suspicious). The gang’s characteristics (from the most matching template FT-2023001) include organizational characteristics (using a three-tier architecture, including master control, data processing and network detection functions), communication characteristics (frequent communication between nodes, and obvious stratification of data packet size distribution), and behavioral characteristics (having a stable periodic communication pattern, suspected of being used for collaborative crime). The warning time is 2024-01-15 11:00:00. The associated warning shows that it has highly similar characteristics to a telecommunications network fraud gang uncovered in December 2023 (case number: 2023120001).
[0124] S317, extracting all suspicious communication records within a preset time range from the original traffic data corresponding to the communication address of the first suspicious group, and sorting the suspicious communication records in chronological order; the first suspicious group is any suspicious group; Among them, the first suspicious group represents the target group to be analyzed; the preset time range refers to the time window for analyzing communication records; the suspicious communication records represent network communication data with suspicious characteristics; and the time sequence sorting is used to indicate the arrangement in the order of the time when the communication occurred.
[0125] After identifying a suspicious group, it is necessary to conduct an in-depth analysis of the group's specific communication behavior. Specifically, the system first determines the time range for analysis, then extracts the communication records of all members of the group from the original traffic data, performs a preliminary screening of these records to retain communication data with suspicious characteristics, and finally sorts all suspicious communication records by timestamp to form an ordered sequence that reflects the temporal characteristics of the group's activities.
[0126] In some embodiments, the extraction and sorting of suspicious communication records can be achieved in a variety of ways: Optionally, a multi-condition filtering method is used to set filtering conditions such as communication frequency, data volume, and protocol type, and multiple rounds of filtering are performed on the original communication records to retain records that meet suspicious characteristics, and then use a quick sort algorithm to sort them by timestamp; Optionally, a time window analysis method is used to divide the preset time range into multiple continuous time windows, identify abnormal communication patterns in each window, extract corresponding communication records, and finally merge the results of each window while maintaining the time sequence relationship.
[0127] It is understandable that other methods may be used to extract and sort suspicious communication records, which are not limited here.
[0128] Continuing with the above example, the system extracts the communication records of the identified suspicious gangs from 10:00:00 to 11:00:00: First, 300 records of IP1 (192.168.1.100), 280 records of IP2 (192.168.1.101), and 260 records of IP3 (192.168.1.102) are extracted from the original data; then suspicious records are screened based on feature rules, including Rule 1 (HTTPS communication with data packet size > 3KB), Rule 2 (frequent small packet communication within a short period of time < 5 seconds), and Rule 3 (direct communication with known fraud-related IPs); after screening, 150 records of IP1 are retained (mainly large data packet HTTP S communication), IP2 retains 120 records (mainly data processing related communications), IP3 retains 100 records (mainly exploratory communications); finally, they are sorted in chronological order, such as 10:00:15.023-IP1→203.0.113.50: HTTPS, 4KB; 10:00:16.157-IP2→203.0.113.51: HTTPS, 1.5KB; 10:00:17.089-IP3→203.0.113.52: HTTP, 1.5KB; 10:00:18.234-IP1→203.0.113.50: HTTPS, 4KB, etc., a total of 370 suspicious communication records are obtained.
[0129] S318. Perform semantic analysis on each suspicious communication record to identify key event types, including fund transactions, information transmission, identity authentication, and instruction issuance; Among them, semantic analysis refers to parsing the meaning of the communication content; key event type refers to the business type corresponding to the communication behavior; capital transaction refers to the communication involving the flow of funds; information transmission refers to the communication of data exchange; identity authentication refers to the communication of identity verification; instruction issuance is used to represent the communication of control commands.
[0130] After obtaining the orderly suspicious communication records, it is necessary to understand the specific business meaning reflected in each record. Specifically, the system conducts in-depth analysis of the content of each communication record, and classifies the communication behaviors into different types such as financial transactions, information transmission, identity authentication, and instruction issuance through methods such as feature matching and pattern recognition, so as to facilitate the subsequent restoration of the gang's crime process.
[0131] In some embodiments, semantic analysis of communication records can be achieved in multiple ways: Optionally, a rule matching method is used to predefine a feature rule base for different types of communication events, including protocol features, data packet features, and behavior pattern features, and the event type of the communication record is determined through multi-layer rule matching; Optionally, a deep learning model is used to take various features of communication records as input, and the type of communication event is identified through a pre-trained multi-classification model. The model can capture complex feature combinations and temporal dependencies.
[0132] It is understandable that other methods may be used to implement semantic analysis of communication records, which are not limited here.
[0133] Continuing with the above example, the system performs semantic analysis on 370 suspicious communication records: first, feature rules are defined, including fund transaction features (HTTPS protocol, data packet size > 3KB, fixed target address such as payment gateway, typical request-response mode), information transmission features (large number of continuous data packets, two-way data flow, stable data packet size), identity authentication features (short request packet < 1KB, standard authentication process sequence, fixed time interval) and command issuance features (small data packet 1-2KB, mainly one-way communication, bursty communication mode). The analysis results show that among the 150 records of IP1 (192.168.1.100), there are 30 fund transactions (such as 10:15:23.456 initiating a transaction request to the payment gateway, 10:15:24.789 receiving a transaction response, etc.) and 120 command issuances (such as 10:00:15.023 sending a control command to IP2, 10:00:18.234 sending a command to IP3, etc.). Send task assignment to IP3, etc.); The 120 records of IP2 (192.168.1.101) include 90 information transmissions (such as 10:20:34.567 batch data upload, 10:20:35.890 data synchronization confirmation, etc.) and 30 identity authentications (such as 10:10:45.678 session establishment request, 10:10:46.901 authentication response, etc.); 100 records of IP3 (192.168.1.102) It includes 70 information transmissions (such as 10:30:12.345 detection data reporting, 10:30:13.678 status synchronization, etc.) and 30 identity authentications (such as 10:05:56.789 authentication request, 10:05:57.012 session confirmation, etc.); the overall distribution is: 30 fund transactions (8.1%), 160 information transmissions (43.2%), 60 identity authentications (16.2%), and 120 instructions (32.5%).
[0134] S319, classifying the suspicious communication records according to the key event types, and obtaining an event association graph for each category, wherein nodes represent communication records and edges represent the time sequence of events; Among them, the event association graph represents a directed graph that describes the association relationship between similar events; a node refers to the basic unit in the graph that represents a single communication record; an edge is used to represent the temporal relationship between nodes; and a time sequence represents the temporal sequence relationship of the occurrence of events.
[0135] After completing the semantic analysis of the communication records, it is necessary to build the association structure of different types of events. Specifically, the system first groups the communication records according to the identified event types, and then uses the communication records as nodes within each category to establish the connection relationship between nodes based on the timestamps, forming a directed graph structure that reflects the temporal evolution of events, which is convenient for analyzing the development laws of different types of events.
[0136] In some embodiments, the event correlation graph can be constructed in a variety of ways: Optionally, a time-sequence graph construction method is used to connect the communication records of the same type in chronological order, and at the same time, the relationship between the two communicating parties is considered, and directed edges with time attributes are added between the nodes to form a complete event evolution link; Optionally, a hierarchical graph structure is used to add hierarchical connections based on the basic temporal relationship and the organizational relationship of the communication address, so as to construct a multi-layer event association graph that reflects both the temporal sequence and the organizational structure.
[0137] It is understandable that other methods may be used to construct the event association graph, which is not limited here.
[0138] Continuing with the above example, the system constructs a correlation graph for the 370 suspicious communication records by event type: In the fund transaction correlation graph (30 nodes), the node attributes include timestamp, source IP (192.168.1.100), target IP (payment gateway), data packet size (4KB), and typical paths such as 10:15:23.456 [transaction request] → 10:15:24.789 [transaction response], 10:25:34.567 [transaction request] → 10:25:35.890 [transaction response], etc. Response], etc., a total of 15 complete transactions; in the information transmission association diagram (160 nodes), the node attributes include timestamp, source IP (192.168.1.101 / 102), target IP (data server), data packet size (1.5KB), typical paths such as 10:20:34.567 [data upload] → 10:20:35.890 [confirmation], 10:30:12.345 [status report] → 10:30:13.678 [confirmation], etc., a total of 80 groups of data transmission; In the authentication association graph (60 nodes), the node attributes include timestamp, source IP (all three IPs), target IP (authentication server), packet size (0.5KB), and typical paths such as 10:05:56.789 [authentication request] → 10:05:57.012 [authentication response], 10:10:45.678 [session request] → 10:10:46.901 [session confirmation], etc., a total of 30 groups of authentication processes; in the instruction delivery association graph (120 nodes), the node attributes include Timestamp, source IP (192.168.1.100), destination IP (192.168.1.101 / 102), data packet size (1.2KB), typical paths such as 10:00:15.023 [control instructions] → 10:00:16.157 [execution confirmation], 10:00:18.234 [task allocation] → 10:00:19.345 [receipt confirmation], etc. A total of 60 groups of instructions were issued, and all related graphs were connected by timestamps to form a complete chain of gang activities.
[0139] S320, extracting suspicious communication records matching the corresponding stages in the preset fraud gang crime pattern template from the event association graphs of each category, and storing them as a first crime record sequence associated with the first suspicious gang, including identity authentication records in the information collection stage, instruction issuance records in the implementation stage, and fund transaction records in the completion stage; Among them, the crime pattern template refers to a template that describes the typical crime process of a fraud gang; the crime stage refers to the different periods in the crime process; the first crime record sequence represents a collection of communication records that reflect the complete crime process; the information collection stage, implementation stage and completion stage correspond to different links of the crime respectively.
[0140] After constructing the event association graph, it is necessary to extract key records that can restore the crime process. Specifically, the system first loads the preset crime pattern template, which defines the typical stages and characteristics of the fraud process, and then searches for communication records that match the characteristics of each stage in the template from various event association graphs, organizes these records in the order of the crime process, and forms a complete crime record sequence.
[0141] In some embodiments, the extraction and organization of crime records can be achieved in a variety of ways: Optionally, a pattern matching method is used to convert the crime pattern template into a feature rule sequence, search for a subgraph structure that meets the rule in the event association graph, extract the corresponding node as the key crime record, and maintain the original time sequence relationship; Optionally, a sequence alignment algorithm is used to align the paths in the event association graph with the crime pattern template, find the best matching path, extract the nodes on the path to form a crime record sequence, and ensure that the sequence fully reflects the crime process.
[0142] It is understandable that other methods may be used to extract and organize crime records, which are not limited here.
[0143] Preferably, in some embodiments, records with a successful communication status in the fund transaction category are identified in the event association graph as key evidence nodes, and each key evidence node corresponds to a fund transaction; based on the associated party address of each key evidence node, the identity authentication records and instruction issuance records corresponding to the address are extracted in other categories of event association graphs, and the extracted records are combined to form a complete set of evidence chains; the records in all evidence chains are merged and arranged in ascending order by timestamp to obtain the first crime record sequence.
[0144] Continuing with the above example, the system extracts key records from various event association graphs: in the information collection phase (10:05-10:15), the identity authentication record includes IP3 sending an authentication request to the authentication server at 10:05:56.789, the authentication server returning authentication success to IP3 at 10:05:57.012, IP2 sending an authentication request to the authentication server at 10:10:45.678, and the authentication server returning authentication success to IP2 at 10:10:46.901; the information transmission record includes IP3 sending a target information report to the data server at 10:06:12.345, and IP2 sending a data analysis result to the data server at 10:11:34.567; in the implementation phase (10:15-10:25), the instruction issuance record includes IP1 issuing a data processing instruction to IP2 at 10:15:15.023, and IP2 returning instruction execution confirmation to IP1 at 10:15:16.157. At 10:15:18.234, IP1 sent the detection task to IP3, and at 10:15:19.345, IP3 returned the task reception confirmation to IP1. The information transmission records include IP2 uploading the processing results to IP1 at 10:20:34.567, and IP1 returning the data reception confirmation to IP2 at 10:20:35.890. In the completion stage (10:25-10:35), the fund transaction records include 10:25:23.456 At 10:25:24.789, the payment gateway returned a successful transaction to IP1. At 10:30:34.567, IP1 sent a transaction request to the payment gateway, and at 10:30:35.890, the payment gateway returned a successful transaction to IP1. The instruction issuance records include IP1 issuing a cleanup instruction to IP2 at 10:35:15.023, and IP2 returning a cleanup completion to IP1 at 10:35:16.157. The system sorts these records by timestamp to form a complete crime record sequence, which clearly shows the entire crime process of the gang from information collection, implementation to completion.
[0145] S321. When the first suspicious gang is marked as a confirmed fraud-related gang, a structured evidence document is generated based on the first crime record sequence. The structured evidence document includes the timestamp, communication content and related party address of each suspicious communication record in the first crime record sequence, as well as the list of addresses of suspicious gang members in the second warning result.
[0146] Among them, the confirmed fraud-involved gang means a gang that has been verified to be involved in fraudulent activities; the structured evidence document refers to the electronic evidence material organized in a standard format; the timestamp indicates the specific time when the communication occurred; the related party address is used to indicate the address information of the parties involved in the communication.
[0147] After a suspicious group is confirmed as a fraud group, a standardized evidence document needs to be generated. Specifically, based on the extracted crime record sequence, the system organizes the key information of each communication record, including the time of occurrence, communication content, and addresses of the participants, into a structured evidence document in a standard format, and records the information of the gang members at the same time to form a complete chain of evidence.
[0148] In some embodiments, the generation of evidence documents can be achieved in a variety of ways: Optionally, a templated document generation method is used to predefine a document template containing various information fields, fill the information in the crime record sequence into the corresponding fields in chronological order, and add gang member information and relationship descriptions to generate a standard format evidence document; Optionally, use a multi-level document structure to divide the evidence document into multiple chapters such as overview, gang information, and the crime process. In the crime process section, organize the communication records according to different stages, and visually display the relationships between gang members through charts and other forms.
[0149] It is understandable that other methods may be used to generate evidence documents, which are not limited here.
[0150] Continuing with the above example, after the suspicious gang is confirmed as a fraud gang, the structured evidence document generated by the system contains the following content: The basic information of the document shows that the case number is CASE-20240115-001, generated at 11:30 on January 15, 2024, and belongs to the category of network communication record evidence, and is related to the case number 2023120001; the members of the gang involved include the main control node (192.168.1.100) responsible for issuing instructions and fund transactions, the data node (192.168.1.101) responsible for data processing and information transmission, and the detection node (192.168.1.102) that performs information collection and status detection, all of which are rated as Level 2 Suspicion level; The key communication records recorded the entire crime process in chronological order, from the information collection stage (10:05:56:789, the detection node sent an authentication request to the authentication server), to the implementation stage (10:15:15:023, the main control node sent a data processing instruction to the data node), and finally to the completion stage (10:25:23:456, the main control node initiated a transaction request to the payment gateway); Evidence analysis showed that the gang had clear division of labor and orderly communication features, adopted a standardized crime process and had anti-reconnaissance awareness, and the complete communication behavior recorded provided reliable electronic evidence support for case investigation.
[0151] In the embodiment of the present application, due to the use of the technical features of combining LSTM-based behavioral feature extraction and dynamic time-warping distance calculation, by performing time series modeling and elastic matching of communication behaviors, the technical problem of the difficulty in accurately depicting the dynamic communication mode of fraud-related gangs in the prior art is effectively solved, thereby achieving the technical effect of improving the accuracy of identifying fraud-related gangs. And due to the use of multi-dimensional event association analysis and evidence chain construction technical features, by performing semantic analysis and association graph construction on communication records, the technical problem of the difficulty in restoring the complete crime process of fraud-related gangs in the prior art is effectively solved, thereby achieving the technical effect of providing reliable electronic evidence support for case investigation.
[0152] The following introduces an exemplary server 100 provided in an embodiment of the present application. Figure 4 It is a schematic diagram of an exemplary hardware structure of the server 100 provided in an embodiment of the present application.
[0153] In some embodiments, the server 100 includes a processor, a memory and a network interface connected by a system bus. Among them, the processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the server is used to store data. The network interface of the server is used to communicate with other external terminals or servers through a network connection. In some embodiments, the network interface can be a wired network interface, and in some embodiments, the network interface can also be a wireless network interface. When the computer program is executed by the processor, the method in the embodiment of the present application is implemented.
[0154] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the server to which the solution of the present application is applied. The specific server may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0155] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying fraudulent network traffic based on LSTM, characterized in that: include: The network traffic data including the communication address, communication time, protocol type and data packet size are collected in real time by the network traffic collection device to obtain the original traffic data to be analyzed; Performing statistical analysis on the data of each communication address in the original traffic data within a preset time window to obtain communication feature data including the number of communications per unit time, average data packet size, number of target addresses, and proportion of protocol types; Using an LSTM network to perform time series modeling on a communication event sequence arranged in chronological order for each communication address in the original traffic data, to obtain behavior feature data that characterizes the time series variation law of the communication behavior; Obtaining a marked fraud gang feature template from a historical feature library, calculating the similarity between the communication feature data and the behavior feature data and the feature template, and obtaining a feature matching degree; Calculating deviation values of the communication feature data and the behavior feature data from corresponding features in a normal behavior feature library to obtain a feature abnormality degree; Based on the feature matching degree and the feature abnormality degree, an adaptive weight coefficient is used for weighted combination to obtain a suspicion score, wherein the adaptive weight coefficient is dynamically adjusted according to the historical warning accuracy rate; For the communication feature data and behavior feature data corresponding to the suspicious communication address whose suspicion score exceeds the preset suspicion threshold, the gang feature description in the fraud gang feature template with the highest matching degree is selected, and the first warning result including the suspicious communication address, the gang suspicion level and the gang feature description is output.
2. The method according to claim 1, characterized in that: The LSTM network is used to perform time series modeling on the communication event sequence arranged in chronological order for each communication address in the original traffic data to obtain behavior feature data that characterizes the time series change law of the communication behavior, specifically including: Encode the communication events of each communication address within a preset time window into an event vector sequence in chronological order, wherein each event vector contains the target address type, communication time interval and data packet size at that moment, and the event vector sequence reflects the behavior change process of the communication address; The LSTM network is used to analyze the changing relationship between adjacent event vectors, extract the communication target switching rules, time interval change trends and packet size change patterns, and obtain the hidden state sequence that characterizes the dynamic changes of communication behavior. The hidden layer state sequence is weighted averaged in the time dimension, with the weight decreasing with the time interval, to obtain the behavior feature data that describes the evolution law of the communication behavior.
3. The method according to claim 1, characterized in that The step of obtaining the marked fraud gang feature template from the historical feature library, calculating the similarity between the communication feature data and the behavior feature data and the feature template, and obtaining the feature matching degree specifically includes: Obtaining a marked fraud gang feature template from the historical feature library, wherein each fraud gang feature template includes a communication feature template and a behavior feature template, wherein the communication feature template includes statistical distribution parameters of the number of communications per unit time, the average data packet size, the number of target addresses, and the proportion of protocol types, and the behavior feature template includes a target address switching rule, a communication time interval change trend, and a feature vector of a data packet size change pattern; The Mahalanobis distance between the communication feature data and the communication feature template, and the dynamic time warping distance between the behavior feature data and the behavior feature template are calculated respectively, wherein the dynamic time warping distance is obtained by constructing a cumulative cost matrix, recording the Euclidean distance between the feature vectors of the corresponding positions of the two sequences at each matrix position, and finding the minimum cumulative cost path by using a dynamic programming method; The comprehensive similarity is calculated based on the Mahalanobis distance and the dynamic time warping distance to obtain the feature matching degree.
4. The method according to claim 1, characterized in that The weighted combination based on the feature matching degree and the feature abnormality degree is performed using an adaptive weight coefficient to obtain a suspiciousness score, which specifically includes: Calculating the distinguishing contribution of the feature matching degree and the feature abnormality degree to the warning result for each record in the historical warning data, wherein the distinguishing contribution is obtained by calculating the distribution difference of the feature matching degree and the feature abnormality degree in the accurate warning sample and the false alarm sample; A feature importance scoring model is constructed based on the distinguishing contribution, and initial weight coefficients of the feature matching degree and the feature abnormality degree are calculated, where the sum of the weight coefficients is 1; After each round of early warning result verification, the newly added accurate early warning samples are added to the feature importance scoring model, the distinction contribution is updated in real time, and the weight coefficient is dynamically adjusted; The updated weight coefficient is used to perform weighted summation on the feature matching degree and the feature abnormality degree to obtain the suspicion score.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: For any two communication addresses in the suspicious communication address set whose suspicion scores exceed a preset suspicion threshold, a first Euclidean distance between their corresponding communication feature data and a second Euclidean distance between their corresponding behavior feature data are calculated; when both the first Euclidean distance and the second Euclidean distance are less than a preset distance threshold, the two communication addresses are classified as the same suspicious group; Calculate the cosine similarity between the average feature vector of each suspicious gang and the feature template of the fraud gang, select the gang feature description in the feature template with the largest cosine similarity, and output the second warning result including the address list of suspicious gang members, the gang suspicion level and the gang feature description.
6. The method according to claim 5, characterized in that The method further comprises: Extracting all suspicious communication records within a preset time range from the original traffic data corresponding to the communication address of the first suspicious group, and sorting the suspicious communication records in chronological order; the first suspicious group is any suspicious group; Perform semantic analysis on each suspicious communication record to identify key event types, including financial transactions, information transmission, identity authentication, and instruction issuance; Classifying the suspicious communication records according to the key event types to obtain an event association graph for each category, wherein nodes represent communication records and edges represent the time sequence of events; Extract suspicious communication records that match the corresponding stages in the preset fraud gang crime pattern template from the event association graphs of each category, and store them as a first crime record sequence associated with the first suspicious gang, including identity authentication records in the information collection stage, instruction issuance records in the implementation stage, and fund transaction records in the completion stage; When the first suspicious group is marked as a confirmed fraud-related group, a structured evidence document is generated based on the first crime record sequence, and the structured evidence document includes the timestamp, communication content and related party address of each suspicious communication record in the first crime record sequence, as well as a list of addresses of suspicious group members in the second warning result.
7. The method according to claim 6, characterized in that The extracting of suspicious communication records matching the corresponding stages in the preset fraud gang crime pattern template from the event association graphs of each category and storing them as the first crime record sequence associated with the first suspicious gang specifically includes: In the event association graph, records with a successful communication status in the fund transaction category are identified as key evidence nodes, each key evidence node corresponding to a fund transaction; Based on the address of the associated party of each key evidence node, the identity authentication record and instruction issuance record corresponding to the address are extracted from other categories of event association graphs, and the extracted records are combined to form a complete chain of evidence; Merge all the records in the evidence chain and arrange them in ascending order according to timestamps to obtain the first crime record sequence.
8. A server, characterized in that: The server includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the server to execute the method described in any one of claims 1-7.
9. A computer program product comprising instructions, characterized in that When the computer program product is run on a server, the server is caused to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on a server, the server is caused to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Early warning method and device for network fraud, computer equipment and storage medium
CN113923011A
Risk data identification method and device and electronic equipment
CN118798626A
Communication network fraud event identification method and device, electronic equipment and storage medium
CN118802223A
Fraud-related group identification method and device, storage medium and electronic equipment
CN119004340A
Network traffic abnormity monitoring method and device based on BiLSTM-Att network
CN119232490A
Cited By
Resource data processing method and device, computer equipment and readable storage medium
CN120408470A
Network gambling behavior identification method and system and medium
CN121486015A
Network gambling behavior identification method, system and medium
CN121486015B