Flow detection method and device, model training method and device, electronic equipment and medium

Through multi-level detection methods, preset detection models and traffic characteristics are used to filter and identify malicious behaviors in network traffic, solving the problem of low detection accuracy in the existing technology and achieving more efficient malicious traffic detection.

CN119995927APending Publication Date: 2025-05-13SHANGHAI DOUXIANG INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411914566.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When detecting malicious behavior in network traffic, the detection accuracy of the prior art is low and it is impossible to effectively identify new attacks.

Method used

Using a multi-level detection method, firstly, the suspicious traffic is filtered through the preset first detection model, then the abnormal type is determined using the second detection model, and the final detection is performed in combination with the traffic characteristics, the first detection result and the second detection result.

Benefits of technology

Improve the accuracy of traffic detection, can more effectively identify malicious traffic, and reduce false alarm rates and missed alarm rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119995927A_ABST
    Figure CN119995927A_ABST
Patent Text Reader

Abstract

The invention provides a traffic detection method and device, a model training method and device, electronic equipment and a medium, and relates to the technical field of computers. The traffic detection method comprises the following steps: acquiring to-be-detected traffic data; using a preset first detection model to detect whether the to-be-detected traffic data is suspicious traffic, and obtaining a first detection result; if the first detection result represents that the to-be-detected traffic data is suspicious traffic, detecting an abnormal type of the to-be-detected traffic data by using a preset second detection model, and obtaining a second detection result; and extracting traffic characteristics from the to-be-detected traffic data, detecting whether the to-be-detected traffic data is malicious traffic by using a preset third detection model based on the traffic characteristics, the first detection result and the second detection result, and outputting a final detection result. Through mutual cooperation of the first detection model, the second detection model and the third detection model, the to-be-detected flow data can be detected more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of computers, and in particular to a flow detection method, a model training method, a device, an electronic device and a medium. Background Art

[0002] In recent years, the Internet industry has flourished, and the Internet has become the main way for people to obtain information. With the emergence of various new websites, network information has grown exponentially, and the security issues that have emerged have become increasingly serious. There are more and more cases of attacking websites through HTTP (Hypertext Transfer Protocol) and other protocols.

[0003] Although a large number of detection and killing solutions have been proposed for web page attacks, the existing detection and killing solutions mainly perform detection through pre-configured detection rules, but this method cannot detect new types of attacks, resulting in low detection accuracy. Summary of the invention

[0004] The present application provides a flow detection method, a model training method, a device, an electronic device and a medium to solve the problem of low detection accuracy in the prior art.

[0005] In a first aspect, the present application provides a traffic detection method, comprising: obtaining traffic data to be detected; using a preset first detection model to detect whether the traffic data to be detected is suspicious traffic, and obtaining a first detection result; if the first detection result characterizes that the traffic data to be detected is suspicious traffic, using a preset second detection model to detect the abnormal type of the traffic data to be detected, and obtaining a second detection result; extracting traffic features from the traffic data to be detected, and using a preset third detection model to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection result, and the second detection result, and outputting a final detection result.

[0006] In the embodiment of the present application, the first detection model is first used to screen the traffic data to be detected, and the data that is clearly non-malicious traffic is filtered out. Then, the traffic data that is malicious according to the first detection result is classified to determine the abnormal type. Finally, the characteristics of the traffic data itself, as well as the first detection result and the second detection result are combined to determine whether it is malicious traffic again, thereby improving the accuracy of detection.

[0007] In combination with the technical solution provided in the first aspect above, in some possible implementations, the traffic data to be detected includes a request message; a preset second detection model is used to detect the abnormal type of the traffic data to be detected, and a second detection result is obtained, including: extracting a first URI (Uniform Resource Identifier) ​​feature, a first header (request header / response header) feature, and a first parameter feature from the request message; inputting the first URI feature, the first header feature, and the first parameter feature into the second detection model respectively to obtain the second detection result, and the second detection result includes a URI detection result, a header detection result, and a parameter detection result.

[0008] In the embodiment of the present application, since different types of attacks may be reflected in different locations in the traffic data to be detected, the first URI feature, the first header feature, and the first parameter feature extracted from the traffic feature to be detected are respectively detected to obtain a second detection result including a URI detection result, a header detection result, and a parameter detection result, thereby facilitating a more accurate determination of the anomaly type in the traffic data to be detected in the subsequent process.

[0009] In combination with the technical solution provided in the first aspect above, in some possible implementations, the traffic feature includes the second URI feature, the second header feature, and the second parameter feature. If the final detection result indicates an abnormality, after obtaining the final detection result, the method further includes: for each sub-feature in the second URI feature, the second header feature, and the second parameter feature, detecting the sub-feature based on a preset feature normal range corresponding to the sub-feature, and obtaining an initial detection result indicating whether the sub-feature is abnormal; counting the number of sub-features in the second URI feature that are characterized as abnormal in the initial detection result and obtaining a URI statistical value, and if the URI statistical value is greater than or equal to the preset number of URIs, outputting that the URI part in the traffic data to be detected is abnormal; if the URI statistical value is less than the preset number of URIs, outputting that the URI part in the traffic data to be detected is abnormal. The method comprises the following steps: first, counting the number of sub-features whose initial detection results are abnormal in the second header feature and obtaining a header statistical value. If the header statistical value is greater than or equal to a preset number of headers, outputting that there is an abnormality in the header part of the traffic data to be detected; second, counting the number of sub-features whose initial detection results are abnormal in the second parameter feature and obtaining a parameter statistical value. If the parameter statistical value is greater than or equal to a preset number of parameters, outputting that there is an abnormality in the parameter part of the traffic data to be detected; and third, determining that there is no abnormality in the parameter part of the traffic data to be detected.

[0010] In the embodiment of the present application, each sub-feature in the second URI feature, the second header feature, and the second parameter feature is analyzed to determine whether there is a sub-feature abnormality. In the case where the number of sub-feature abnormalities is too large, it indicates that there is a high possibility that the part where the sub-feature is located is abnormal. Therefore, in this way, the URI part, the header part, and the parameter part in the flow data to be detected can be analyzed respectively to determine the specific location where the abnormality exists in the flow data to be detected.

[0011] In combination with the technical solution provided in the first aspect above, in some possible implementations, the method further includes: if there is an abnormality in the URI part of the traffic data to be detected, determining that the URI detection result is the abnormal type of the traffic data to be detected; if there is an abnormality in the header part of the traffic data to be detected, determining that the header detection result is the abnormal type of the traffic data to be detected; if there is an abnormality in the parameter part of the traffic data to be detected, determining that the parameter detection result is the abnormal type of the traffic data to be detected.

[0012] In the embodiment of the present application, since the second detection result output by the second detection model includes the URI detection result, the header detection result, and the parameter detection result, after determining the specific location where the anomaly exists in the traffic data to be detected, the final anomaly type can be determined from the second detection result according to the specific location where the anomaly exists, so that the final determined anomaly type can be more accurate.

[0013] In combination with the technical solution provided in the first aspect above, in some possible implementations, before using the preset first detection model to detect whether the traffic data to be detected is suspicious traffic, the method also includes: comparing the traffic data to be detected with a preset whitelist; if the traffic data to be detected is in the whitelist, outputting the traffic data to be detected as normal traffic; if the traffic data to be detected is not in the whitelist, inputting the traffic data to be detected into the pre-trained first detection model.

[0014] In the embodiment of the present application, by screening the traffic data to be detected through a whitelist, the traffic data to be detected that is clearly normal traffic can be directly screened out, thereby reducing the workload of subsequent steps.

[0015] In combination with the technical solution provided in the first aspect above, in some possible implementations, the method also includes: performing quantitative statistics on all target traffic data to be detected within a preset time length; the target traffic data to be detected is the traffic data to be detected that is characterized as malicious traffic by the final detection result; if the number of target traffic data to be detected obtained by counting within the preset time length is greater than a preset number threshold, then all target traffic data to be detected within the preset time length are marked as vulnerability scans.

[0016] In an embodiment of the present application, since vulnerability scanning usually involves a large number of attack behaviors in a short period of time, if the number of target traffic data to be detected obtained by counting within a preset time length is greater than a preset number threshold, it is determined that these target traffic data to be detected may be caused by vulnerability scanning.

[0017] In a second aspect, the present application provides a model training method, comprising: obtaining a training data set, each training data in the training data set comprising a traffic feature corresponding to a traffic data, and a first detection result and a second detection result corresponding to the traffic data; wherein each training data is annotated with a label representing whether it is malicious traffic; the first detection result is a detection result of a first detection model detecting whether the training data is suspicious traffic; the second detection result is a detection result of a second detection model detecting an abnormal type of the training data; and training the initial model based on the training data set to obtain a trained third detection model.

[0018] In a third aspect, the present application provides a flow detection device, comprising: an acquisition module and a processing module, the acquisition module being used to acquire flow data to be detected; the processing module being used to detect whether the flow data to be detected is suspicious flow and obtain a first detection result; if the first detection result indicates that the flow data to be detected is not suspicious flow, the flow data to be detected is output as normal flow; if the first detection result indicates that the flow data to be detected is suspicious flow, a preset first detection model is used to detect the abnormal type of the flow data to be detected and a second detection result is obtained; flow features are extracted from the flow data to be detected, and a preset second detection model is used to detect whether the flow data to be detected is malicious flow based on the flow features, the first detection result and the second detection result, and a final detection result is output.

[0019] In a fourth aspect, the present application provides an electronic device, comprising: a memory and a processor, the memory and the processor being connected; the memory being used to store programs; the processor being used to call the programs stored in the memory to execute the method described in the first aspect above and / or in combination with any possible implementation of the first aspect above, or to execute the method described in the second aspect above.

[0020] In a fifth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a computer, it executes the method described in the first aspect and / or in combination with any possible implementation of the first aspect, or executes the method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 A schematic diagram of a flow chart of a first flow detection method shown in an embodiment of the present application;

[0023] Figure 2 A schematic diagram of a flow chart of a second flow detection method shown in an embodiment of the present application;

[0024] Figure 3 This is a structural block diagram of a flow detection device shown in an embodiment of the present application;

[0025] Figure 4A structural block diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0027] It should be noted that similar numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it is not necessary to further define and explain it in the subsequent drawings. At the same time, in the description of this application, relational terms such as "first", "second", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements limited by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or equipment including the elements.

[0028] The technical solution of the present application will be described in detail below with reference to the accompanying drawings.

[0029] See also Figure 1 , Figure 1 This is a flow chart of a flow detection method shown in an embodiment of the present application. Figure 1 Describe the steps it involves.

[0030] S100: Obtaining traffic data to be detected.

[0031] The flow data to be detected may be flow data obtained in real time, or may be pre-acquired and stored in a local storage medium, and can be directly called when needed.

[0032] S200: Using a preset first detection model to detect whether the traffic data to be detected is suspicious traffic, and obtaining a first detection result.

[0033] In one implementation, the flow data to be detected may be directly used as input data of the first detection model, and then a first detection result output by the first detection model is obtained. The first detection result indicates whether the flow data to be detected is suspicious flow.

[0034] In one implementation, the flow data to be detected may be subjected to a preset standard structured processing to obtain input data in a standard structured format. The input data is then used as input data of a first detection model, and then a first detection result output by the first detection model is obtained. The first detection result indicates whether the flow data to be detected is suspicious flow.

[0035] In one implementation mode, before using a preset first detection model to detect whether the traffic data to be detected is suspicious traffic, the traffic data to be detected may be compared with a preset whitelist.

[0036] If the traffic data to be detected is in the whitelist, the output traffic data to be detected is normal traffic.

[0037] If the traffic data to be detected is not in the whitelist, the traffic data to be detected is input into a pre-trained first detection model.

[0038] By filtering the traffic data to be detected through the whitelist, the traffic data to be detected that is clearly normal traffic can be directly filtered out, thereby reducing the workload of subsequent steps.

[0039] Optionally, the specific rule for comparing the traffic data to be detected with the preset whitelist may be:

[0040] 1. Extract the source IP (Internet Protocol) and destination IP in the traffic data to be detected. If the source IP and the destination IP are in the white list, it is determined that the traffic data to be detected is in the white list.

[0041] 2. Extract the URI in the traffic data to be detected. If the URI is in the whitelist, it is determined that the traffic data to be detected is in the whitelist.

[0042] 3. If the request type of the traffic data to be detected is a get request and the status code is 200, it is determined that the traffic data to be detected is in the whitelist.

[0043] 4. If the request type of the traffic data to be detected is an elasticsearch request (that is, port 9200), since elasticsearch requests are usually intranet requests, it can be determined that the traffic data to be detected is in the whitelist.

[0044] Among them, the above four specific rules for comparing the traffic data to be detected with the preset white list can be selected according to actual needs in actual application. Alternatively, new rules can be added according to actual needs.

[0045] If the first detection result indicates that the traffic data to be detected is normal traffic, the detection process is terminated and it is determined that the traffic data to be detected is sorted traffic.

[0046] In one implementation, the training process of the first detection model may include: first, obtaining a first detection model training data set. Each piece of training data in the first detection model training data set includes a test flow data in a standard structured format and a true label indicating whether it is abnormal flow. Then, the initial first detection model is trained based on the first detection model training data set to obtain a trained first detection model.

[0047] The specific process and principle of training the model are well known to those skilled in the art and will not be described in detail here for the sake of brevity.

[0048] Optionally, the loss function used for training may be a cross-entropy loss function, which is used as a quantitative standard to measure the difference between the probability distribution predicted by the first detection model and the true label.

[0049] Optionally, the first detection model training data set may include a training set and a validation set. The initial first detection model is iteratively trained on the training set and the validation set, and the initial first detection model parameters are regularly evaluated and adjusted to prevent overfitting and ensure the generalization ability of the model. The effectiveness of the model in identifying malicious HTTP traffic is comprehensively evaluated through multi-dimensional indicators such as precision, recall, and F1 score.

[0050] Optionally, the first detection model can be a Transformer model (a deep learning model architecture).

[0051] Optionally, during the training phase, the acquired HTTP traffic data for training can be deeply parsed and structured into a unified and structured format, and the request (request object) and response (response object) traffic packets can be put together. Build a vocabulary tailored for the training task, which is based on all unique tokens (the smallest semantic unit in the text) in the collected HTTP traffic data, and realizes the mapping conversion from text tokens to unique numerical IDs. The first detection model can introduce an input embedding mechanism to convert numerical tokens into dense vectors in a high-dimensional vector space. Construct a multi-layer Transformer encoder, perform deep feature learning through self-attention mechanism and multi-layer feedforward network, and realize efficient capture and abstraction of complex sequence patterns. Design a classification output layer, and use a fully connected layer to map the high-level features of the Transformer encoder to predefined normal / malicious category labels to complete the final classification task.

[0052] Optionally, the first detection model training data set may be obtained by first obtaining normal traffic data and malicious traffic data, and then performing standard formatting on each piece of traffic data obtained, and marking it (ie, the aforementioned true label) to obtain the first detection model training data set.

[0053] The sources of normal traffic data can be: normal traffic of CSIC2010 WEB (a data set used to test WEB attack protection systems); daily traffic of freebuf (a network security industry portal), but it may contain some attack traffic, which needs to be filtered out; capturing the HTTP traffic of ordinary employees of the company who usually access the WEB in large quantities through the network card; simulating WEB traffic in various normal scenarios.

[0054] The sources of malicious traffic data can be: CSIC2010 (a public network intrusion detection data set) WEB malicious traffic; malicious traffic intercepted by waf (Web Application Firewall) (data from github), some of which do not belong to malicious traffic data and need to be filtered out; WEB risk events on the company's prs products, some of which are malicious samples, and the types of malicious samples are: "weak passwords", "threat intelligence domain name detection", "Nginx (high-performance HTTP and reverse proxy web server) directory traversal response", "Apache Tomcat (ordinary server) form authentication user name enumeration request", "threat intelligence-domain name n"; and HTTP protocol generated by leakage scanning tools: mainly some scanning tools scan network traffic; simulate various combined attack WEB traffic.

[0055] S300: If the first detection result indicates that the traffic data to be detected is suspicious traffic, a preset second detection model is used to detect an abnormal type of the traffic data to be detected, and a second detection result is obtained.

[0056] In one implementation, a first detection feature may be extracted from the flow data to be detected, and then the first detection feature is input into a second detection model to obtain a second detection result representing the abnormal type of the flow data to be detected.

[0057] Among them, the first detection feature can be shown in Table 1:

[0058] Table 1

[0059]

[0060]

[0061]

[0062] The first detection feature may include all the features shown in Table 1, or the first detection feature may include only some of the features shown in Table 1. The specific setting method of the first detection feature may be selected according to actual needs.

[0063] In one implementation mode, the traffic data to be detected includes a request message; the specific method of using a preset second detection model to detect the abnormal type of the traffic data to be detected and obtaining the second detection result may be: first extracting a first URI feature, a first header feature, and a first parameter feature from the request message. Then, the first URI feature, the first header feature, and the first parameter feature are respectively input into the second detection model to obtain a second detection result, which includes a URI detection result, a header detection result, and a parameter detection result.

[0064] The first URI feature is a first detection feature extracted from the URI part of the request message. The first header feature is a first detection feature extracted from the header part of the request message. The first parameter feature is a first detection feature extracted from the parameter part of the request message.

[0065] The specific implementation method of the first URI feature, the first header feature, and the first parameter feature is the same as the implementation method of the aforementioned first detection feature, and will not be repeated here for the sake of brief description.

[0066] Since different types of attacks may be reflected in different locations in the traffic data to be detected, the first URI feature, the first header feature, and the first parameter feature extracted from the traffic feature to be detected are detected respectively to obtain a second detection result including a URI detection result, a header detection result, and a parameter detection result, so as to facilitate more accurate determination of the abnormal type in the traffic data to be detected later.

[0067] In one implementation, the second detection model may be a random forest model or other models. The specific type of the second detection model is not limited to the example given here.

[0068] In one implementation, the training process of the second detection model may include: first obtaining a second detection model training data set. Each piece of training data in the second detection model training data set includes any one of a first URI feature, a first header feature, and a first parameter feature, and each piece of training data also includes an abnormality label representing an abnormality type. Then, the initial second detection model is trained based on the second detection model training data set to obtain a trained second detection model.

[0069] The specific process and principle of training the model are well known to those skilled in the art and will not be described in detail here for the sake of brevity.

[0070] Optionally, the second detection model training data set may be obtained by first obtaining normal traffic data and malicious traffic data, then extracting features from each piece of traffic data obtained, respectively obtaining a first URI feature, a first header feature, and a first parameter feature, and marking them (i.e., the aforementioned abnormal label) to obtain a second detection model training data set.

[0071] Among them, in the case that the traffic data is malicious traffic data, it is also possible to first determine which part of the URI part, header part, and parameter part of the traffic data is abnormal, and then when marking, the feature corresponding to the abnormal part is marked as a label representing the abnormality, and the feature corresponding to the other parts is marked as a label representing normality.

[0072] Alternatively, when the traffic data is malicious traffic data, the three features extracted from the traffic data may all be marked as labels representing anomalies.

[0073] Optionally, the method of obtaining normal traffic data and malicious traffic data is the same as the aforementioned method of obtaining normal traffic data and malicious traffic data, and for the sake of brief description, it will not be repeated here.

[0074] Malicious traffic data can also be artificially created. That is, data sets such as SQL injection, XSS attack, and command execution can be collected online first. Then, when creating traffic data, corresponding attacks are randomly inserted to obtain malicious traffic data.

[0075] S400: extracting traffic features from the traffic data to be detected, and using a preset third detection model to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection result, and the second detection result, and outputting a final detection result.

[0076] In one implementation manner, the specific method of extracting traffic features from the traffic data to be detected may be: directly extracting traffic features from the traffic data to be detected.

[0077] Optionally, the specific method of extracting traffic features from the traffic data to be detected may also be: extracting a second URI feature from the URI part of the traffic data to be detected; extracting a second header feature from the header part of the traffic data to be detected; extracting a second parameter feature from the parameter part of the traffic data to be detected. The traffic features include a second URI feature, a second header feature, and a second parameter feature.

[0078] For ease of understanding, the specific implementation of the second URI feature is shown in Table 2:

[0079] Table 2

[0080]

[0081] The second URI feature may include all the features shown in Table 2, or the second URI feature may include only some of the features shown in Table 2. The specific setting method of the second URI feature may be selected according to actual needs.

[0082] For ease of understanding, the specific implementation of the second header feature is shown in Table 3:

[0083] Table 3

[0084]

[0085]

[0086] The second header feature may include all the features shown in Table 3, or the second header feature may include only some of the features shown in Table 3. The specific setting method of the second header feature may be selected according to actual needs.

[0087] For ease of understanding, the specific implementation of the second parameter feature is shown in Table 4:

[0088] Table 4

[0089]

[0090]

[0091] The second parameter feature may include all the features shown in Table 4, or the second parameter feature may include only some of the features shown in Table 4. The specific setting method of the second parameter feature may be selected according to actual needs.

[0092] Optionally, the specific method of extracting traffic features from the traffic data to be detected may also be: extracting global features from the traffic data to be detected; extracting second URI features from the URI part of the traffic data to be detected; extracting second header features from the header part of the traffic data to be detected; extracting second parameter features from the parameter part of the traffic data to be detected. The traffic features include second URI features, second header features, and second parameter features.

[0093] For ease of understanding, the specific implementation of global features is shown in Table 5:

[0094] Table 5

[0095]

[0096] The global features may include all the features shown in Table 5, or the global features may include only some of the features shown in Table 5. The specific setting method of the global features may be selected according to actual needs.

[0097] The specific implementation methods of the second URI feature, the second header feature, and the second parameter feature have been clearly described in the previous text. For the sake of brief description, they will not be repeated here.

[0098] In one implementation, the global feature, the second URI feature, the second header feature, the second parameter feature, the first detection result, and the second detection result may be spliced ​​in a preset order, and then the spliced ​​data may be input into a third detection model to obtain a third detection result output by the third detection model.

[0099] Optionally, since the second detection result may include a URI detection result, a header detection result, and a parameter detection result, in this case, when splicing, the entire second detection result may be spliced ​​as a whole.

[0100] Alternatively, when the global feature, the second URI feature, the second header feature, the second parameter feature, the first detection result, and the second detection result are spliced ​​in a preset order, the URI detection result, the header detection result, and the parameter detection result in the second detection result may be split apart, and the URI detection result may be spliced ​​together with the second URI feature, the header detection result may be spliced ​​together with the second header feature, and the parameter detection result may be spliced ​​together with the second parameter feature.

[0101] For example, the concatenation order may be: global feature, second URI feature, URI detection result, second header feature, header detection result, second parameter feature, parameter detection result, first detection result. Alternatively, the concatenation order may be: global feature, URI detection result, second URI feature, header detection result, second header feature, parameter detection result, second parameter feature, first detection result.

[0102] The examples here are only for ease of understanding. The specific splicing order can be set according to actual needs, as long as in the splicing order, the URI detection result is adjacent to the second URI feature, the header detection result is adjacent to the second header feature, and the parameter detection result is adjacent to the second parameter feature.

[0103] In one implementation mode, when the traffic feature includes a second URI feature, a second header feature, and a second parameter feature, if the final detection result indicates an abnormality, after obtaining the final detection result, each sub-feature in the second URI feature, the second header feature, and the second parameter feature can be further detected based on a preset feature normal range corresponding to the sub-feature to obtain an initial detection result indicating whether the sub-feature is abnormal.

[0104] Then, the number of sub-features characterized as abnormal in the initial detection results in the second URI feature is counted and the URI statistical value (that is, the number of sub-features characterized as abnormal in the initial detection results in the second URI feature) is obtained. If the URI statistical value is greater than or equal to the preset number of URIs, it is output that there is an abnormality in the URI part of the traffic data to be detected; if the URI statistical value is less than the preset number of URIs, it is determined that there is no abnormality in the URI part of the traffic data to be detected.

[0105] And count the number of sub-features whose initial detection results are abnormal in the second header feature and obtain the header statistical value (that is, the number of sub-features whose initial detection results are abnormal in the second header feature). If the header statistical value is greater than or equal to the preset number of headers, it is output that there is an abnormality in the header part of the traffic data to be detected; if the header statistical value is less than the preset number of headers, it is determined that there is no abnormality in the header part of the traffic data to be detected.

[0106] And count the number of sub-features whose initial detection results are abnormal in the second parameter feature and obtain the parameter statistical value (that is, the number of sub-features whose initial detection results are abnormal in the second parameter feature). If the parameter statistical value is greater than or equal to the preset number of parameters, it is output that there is an abnormality in the parameter part of the flow data to be detected; if the parameter statistical value is less than the preset number of parameters, it is determined that there is no abnormality in the parameter part of the flow data to be detected.

[0107] By analyzing each sub-feature in the second URI feature, the second header feature, and the second parameter feature, it is determined whether there is a sub-feature anomaly. If the number of sub-feature anomalies is too large, it indicates that the part where the sub-feature is located is likely to be abnormal. Therefore, in this way, the URI part, the header part, and the parameter part in the traffic data to be detected can be analyzed respectively to determine the specific location of the anomaly in the traffic data to be detected.

[0108] Among them, the specific values ​​of the preset number of URIs, the preset number of headers, and the preset number of parameters can be set according to actual needs, for example, they can be 1, 2, 3, 4, 5, 6, 7, 8, etc., and their specific values ​​are not limited here.

[0109] Optionally, the values ​​of the three preset numbers, namely, the URI preset number, the header preset number, and the parameter preset number, may be equal.

[0110] In one implementation mode, after determining the location where the anomaly exists in the traffic data to be detected, the final anomaly type may be further determined based on the URI detection result, the header detection result, and the parameter detection result included in the second detection result.

[0111] If there is an abnormality in the URI part of the traffic data to be detected, it is determined that the URI detection result is an abnormal type of the traffic data to be detected.

[0112] If there is an abnormality in the header part of the traffic data to be detected, it is determined that the header detection result is the abnormal type of the traffic data to be detected.

[0113] If there is an abnormality in the parameter part of the flow data to be detected, it is determined that the parameter detection result is the abnormal type of the flow data to be detected.

[0114] Since the second detection result output by the second detection model includes a URI detection result, a header detection result, and a parameter detection result, and the URI detection result is obtained based on the first URI feature extracted from the URI part of the traffic data to be detected and the second detection model, the header detection result is obtained based on the first header feature extracted from the header part of the traffic data to be detected and the second detection model, and the parameter detection result is obtained based on the first parameter feature extracted from the parameter part of the traffic data to be detected and the second detection model.

[0115] Therefore, after determining the specific location of the abnormality in the flow data to be detected, the abnormality type determined according to the part where the specific location of the abnormality is located is obviously more accurate. Therefore, the final abnormality type can be determined from the second detection result according to the specific location of the abnormality, so that the final determined abnormality type can be more accurate.

[0116] In one implementation mode, it is also possible to perform quantitative statistics on all target traffic data to be detected within a preset time length; the target traffic data to be detected is the traffic data to be detected that is characterized as malicious traffic by the final detection result.

[0117] If the number of target traffic data to be detected obtained by counting within a preset time length is greater than a preset number threshold, all target traffic data to be detected within the preset time length are marked as vulnerability scanning.

[0118] If the number of target traffic data to be detected obtained within a preset time length is less than or equal to a preset number threshold, no vulnerability scanning mark is performed.

[0119] Among them, the preset time length can be set according to actual needs, for example, it can be 30 seconds, 1 minute, 5 minutes, 10 minutes, half an hour, one hour, half a day, one day, etc., and its specific length is not limited here.

[0120] The preset quantity threshold may also be set to a specific value according to actual needs, for example, it may be 50, 100, 300, etc. The specific value of the preset quantity threshold is not limited to the example given here.

[0121] Since vulnerability scanning usually involves a large number of attacks in a short period of time, if the number of target traffic data to be detected obtained within a preset time period is greater than a preset threshold, it is determined that the target traffic data to be detected may be caused by vulnerability scanning.

[0122] In one implementation, the training method of the third detection model includes: first obtaining a training data set, each training data in the training data set includes a flow feature corresponding to a flow data, and a first detection result and a second detection result corresponding to the flow data. Then, the initial model is trained based on the training data set to obtain a trained third detection model.

[0123] Among them, each training data in the training data set is annotated with a label indicating whether it is malicious traffic; the first detection result is the detection result of the first detection model detecting whether the training data is suspicious traffic; the second detection result is the detection result of the second detection model detecting the abnormal type of the training data.

[0124] Optionally, the specific implementation method of the traffic feature corresponding to the traffic data included in each training data is the same as the implementation method of the aforementioned traffic feature, and for the sake of brief description, it will not be repeated here.

[0125] Optionally, the training data set may be obtained by first obtaining normal traffic data and malicious traffic data, then extracting features from each piece of traffic data obtained to obtain traffic features of each piece of traffic data, and marking the traffic features (i.e., the aforementioned labels indicating whether it is malicious traffic) to obtain the training data set.

[0126] The specific method of obtaining normal traffic data and malicious traffic data has been clearly described in the previous article. For the sake of brief description, it will not be repeated here.

[0127] For ease of understanding, the specific implementation of the flow detection method is described below with examples.

[0128] like Figure 2 As shown, firstly, the HTTP protocol data (that is, the traffic data to be detected) is read from the topic (the basic unit of the Kafka message system) of Kafka (an open source stream processing platform). Then, the traffic data to be detected is compared with the preset whitelist.

[0129] If the traffic data to be detected is in the whitelist, the traffic data to be detected is output as normal traffic, and the detection process is completed.

[0130] If the traffic data to be detected is not in the whitelist, the traffic data to be detected is input into a pre-trained first detection model.

[0131] If the first detection result output by the first detection model indicates that the traffic data to be detected is normal traffic, it is determined that the traffic data to be detected is normal traffic, and the detection process is completed.

[0132] If the first detection result output by the first detection model indicates that the flow data to be detected is suspicious flow, the preset second detection model is used to detect the abnormal type of the flow data to be detected and obtain a second detection result.

[0133] Then, traffic features are extracted from the traffic data to be detected (traffic features include a second URI feature, a second header feature, and a second parameter feature), and a preset third detection model is used to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection results, and the second detection results, and the final detection results are output.

[0134] If the final test result is determined to be normal traffic, the test process is completed.

[0135] If the final detection result is determined to be malicious traffic, then for each sub-feature in the second URI feature, the second header feature, and the second parameter feature, the sub-feature is detected based on the preset feature normal range corresponding to the sub-feature to obtain an initial detection result characterizing whether the sub-feature is abnormal.

[0136] All initial detection results are counted to obtain URI statistics, header statistics, and parameter statistics. The final anomaly type is then determined based on the URI statistics, header statistics, parameter statistics, and the second detection result.

[0137] Specifically, the number of sub-features characterized as abnormal in the initial detection result in the second URI feature is counted and a URI statistical value is obtained. If the URI statistical value is greater than or equal to the preset number of URIs, it is output that there is an abnormality in the URI part of the traffic data to be detected; if the URI statistical value is less than the preset number of URIs, it is determined that there is no abnormality in the URI part of the traffic data to be detected.

[0138] The number of sub-features whose initial detection results are abnormal in the second header feature is counted and the header statistical value is obtained. If the header statistical value is greater than or equal to the preset number of headers, it is output that there is an abnormality in the header part of the traffic data to be detected; if the header statistical value is less than the preset number of headers, it is determined that there is no abnormality in the header part of the traffic data to be detected.

[0139] The number of sub-features whose initial detection results are abnormal in the second parameter feature is counted and the parameter statistical value is obtained. If the parameter statistical value is greater than or equal to the preset number of parameters, it is output that there is an abnormality in the parameter part of the flow data to be detected; if the parameter statistical value is less than the preset number of parameters, it is determined that there is no abnormality in the parameter part of the flow data to be detected.

[0140] If there is an abnormality in the URI part of the traffic data to be detected, it is determined that the URI detection result is an abnormal type of the traffic data to be detected.

[0141] If there is an abnormality in the header part of the traffic data to be detected, it is determined that the header detection result is an abnormal type of the traffic data to be detected.

[0142] If there is an abnormality in the parameter part of the flow data to be detected, it is determined that the parameter detection result is the abnormal type of the flow data to be detected.

[0143] Afterwards, a quantitative count is performed on all the target traffic data to be detected within a preset time length; the target traffic data to be detected is the traffic data to be detected that is characterized as malicious traffic by the final detection result.

[0144] If the number of target traffic data to be detected obtained by counting within a preset time length is greater than a preset number threshold, all target traffic data to be detected within the preset time length are marked as vulnerability scanning.

[0145] Finally, the final detection results of the traffic data to be detected, the type of anomaly, and whether it is a vulnerability scan are written into Kafka to complete the detection process.

[0146] Based on the same technical concept, the present application also provides a flow detection device, such as Figure 3 As shown, the flow detection device 100 includes an acquisition module 110 and a processing module 120 .

[0147] The acquisition module 110 is used to acquire the flow data to be detected.

[0148] The processing module 120 is used to detect whether the traffic data to be detected is suspicious traffic and obtain a first detection result; if the first detection result indicates that the traffic data to be detected is not suspicious traffic, the traffic data to be detected is output as normal traffic; if the first detection result indicates that the traffic data to be detected is suspicious traffic, the abnormal type of the traffic data to be detected is detected using a preset first detection model, and a second detection result is obtained; traffic features are extracted from the traffic data to be detected, and a preset second detection model is used to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection result, and the second detection result, and a final detection result is output.

[0149] The traffic data to be detected includes a request message; the processing module 120 is specifically used to extract a first URI feature, a first header feature, and a first parameter feature from the request message; the first URI feature, the first header feature, and the first parameter feature are respectively input into the second detection model to obtain the second detection result, which includes a URI detection result, a header detection result, and a parameter detection result.

[0150] The traffic feature includes the second URI feature, the second header feature, and the second parameter feature. If the final detection result indicates an abnormality, after obtaining the final detection result, the processing module 120 is further used to detect each sub-feature in the second URI feature, the second header feature, and the second parameter feature based on a preset feature normal range corresponding to the sub-feature to obtain an initial detection result indicating whether the sub-feature is abnormal; count the number of sub-features in the second URI feature that are indicated as abnormal in the initial detection result and obtain a URI statistical value; if the URI statistical value is greater than or equal to the preset number of URIs, output that the URI part in the traffic data to be detected is abnormal; if the URI statistical value is less than the preset number of URIs, determine that the traffic data to be detected is abnormal. There is no abnormality in the URI part of the second header feature; count the number of sub-features whose initial detection results are abnormal in the second header feature and obtain the header statistical value, if the header statistical value is greater than or equal to the preset number of headers, then output that there is an abnormality in the header part of the traffic data to be detected; if the header statistical value is less than the preset number of headers, determine that there is no abnormality in the header part of the traffic data to be detected; count the number of sub-features whose initial detection results are abnormal in the second parameter feature and obtain the parameter statistical value, if the parameter statistical value is greater than or equal to the preset number of parameters, then output that there is an abnormality in the parameter part of the traffic data to be detected; if the parameter statistical value is less than the preset number of parameters, determine that there is no abnormality in the parameter part of the traffic data to be detected.

[0151] The processing module 120 is also used to determine that if there is an abnormality in the URI part of the traffic data to be detected, the URI detection result is the abnormal type of the traffic data to be detected; if there is an abnormality in the header part of the traffic data to be detected, the header detection result is determined to be the abnormal type of the traffic data to be detected; if there is an abnormality in the parameter part of the traffic data to be detected, the parameter detection result is determined to be the abnormal type of the traffic data to be detected.

[0152] Before using the preset first detection model to detect whether the traffic data to be detected is suspicious traffic, the processing module 120 is also used to compare the traffic data to be detected with a preset white list; if the traffic data to be detected is in the white list, the traffic data to be detected is output as normal traffic; if the traffic data to be detected is not in the white list, the traffic data to be detected is input into the pre-trained first detection model.

[0153] The processing module 120 is further used to count the number of all target traffic data to be detected within a preset time length; the target traffic data to be detected is the traffic data to be detected characterized as malicious traffic by the final detection result;

[0154] If the number of target traffic data to be detected obtained by counting within the preset time length is greater than a preset number threshold, all target traffic data to be detected within the preset time length are marked as vulnerability scanning.

[0155] The flow detection device 100 provided in the embodiment of the present application has the same implementation principle and technical effects as those of the aforementioned flow detection method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned flow detection method embodiment.

[0156] Based on the same technical concept, the present application also provides a model training device, which includes a second acquisition module and a second processing module.

[0157] The second acquisition module is used to obtain a training data set, each training data in the training data set includes a traffic feature corresponding to the traffic data, and a first detection result and a second detection result corresponding to the traffic data; wherein each training data is annotated with a label representing whether it is malicious traffic; the first detection result is a detection result of the first detection model detecting whether the training data is suspicious traffic; the second detection result is a detection result of the second detection model detecting the abnormal type of the training data.

[0158] The second processing module is used to train the initial model based on the training data set to obtain a trained third detection model.

[0159] The model training device provided in the embodiment of the present application has the same implementation principle and technical effects as those in the aforementioned model training method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding contents in the aforementioned model training method embodiment.

[0160] See also Figure 4 , which is an electronic device 200 provided in an embodiment of the present application. The electronic device 200 includes: a processor 210 and a memory 220.

[0161] The memory 220 and the processor 210 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory 220 is used to store computer programs, such as storing Figure 3The software function module shown in is the flow detection device 100. Among them, the flow detection device 100 includes at least one software function module that can be stored in the memory 220 in the form of software or firmware or fixed in the operating system (OS) of the electronic device 200. The processor 210 is used to execute the executable module stored in the memory 220, such as the software function module or computer program included in the flow detection device 100. At this time, the processor 210 is used to obtain the flow data to be detected; use the preset first detection model to detect whether the flow data to be detected is suspicious flow, and obtain a first detection result; if the first detection result indicates that the flow data to be detected is suspicious flow, then use the preset second detection model to detect the abnormal type of the flow data to be detected, and obtain a second detection result; extract flow features from the flow data to be detected, and use the preset third detection model based on the flow features, the first detection result, and the second detection result to detect whether the flow data to be detected is malicious flow, and output the final detection result.

[0162] Among them, the memory 220 can be, but is not limited to, RAM (Random Access Memory), ROM (Read Only Memory), PROM (Programmable Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electric Erasable Programmable Read-Only Memory), etc.

[0163] The processor 210 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a CPU (Central Processing Unit), an NP (Network Processor), etc.; it may also be a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. A general-purpose processor may be a microprocessor or the processor 210 may also be any conventional processor, etc.

[0164] The electronic device 200 mentioned above includes but is not limited to a personal computer, a server, etc.

[0165] The embodiment of the present application also provides a computer-readable storage medium (hereinafter referred to as storage medium), on which a computer program is stored, and when the computer program is run by a computer such as the above-mentioned electronic device 200, the above-mentioned flow detection method is executed. The computer-readable storage medium includes: a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and other media that can store program codes.

[0166] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A flow detection method, characterized in that: include: Obtain the traffic data to be detected; Using a preset first detection model to detect whether the traffic data to be detected is suspicious traffic, and obtaining a first detection result; If the first detection result indicates that the traffic data to be detected is suspicious traffic, a preset second detection model is used to detect the abnormal type of the traffic data to be detected, and a second detection result is obtained; Extract traffic features from the traffic data to be detected, and use a preset third detection model to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection result, and the second detection result, and output a final detection result.

2. The method according to claim 1, characterized in that The traffic data to be detected includes a request message; Detecting the abnormal type of the traffic data to be detected by using a preset second detection model, and obtaining a second detection result, including: Extracting a first URI feature, a first header feature, and a first parameter feature from the request message; The first URI feature, the first header feature, and the first parameter feature are respectively input into the second detection model to obtain the second detection result, which includes a URI detection result, a header detection result, and a parameter detection result.

3. The method according to claim 1, characterized in that The traffic feature includes a second URI feature, a second header feature, and a second parameter feature. If the final detection result indicates an abnormality, after obtaining the final detection result, the method further includes: For each sub-feature of the second URI feature, the second header feature, and the second parameter feature, the sub-feature is detected based on a preset feature normal range corresponding to the sub-feature to obtain an initial detection result indicating whether the sub-feature is abnormal; Counting the number of sub-features characterized as abnormal in the initial detection result in the second URI feature and obtaining a URI statistical value, if the URI statistical value is greater than or equal to the preset number of URIs, outputting that the URI part in the traffic data to be detected has an abnormality; if the URI statistical value is less than the preset number of URIs, determining that the URI part in the traffic data to be detected has no abnormality; Counting the number of sub-features of which the initial detection result is abnormal in the second header feature and obtaining a header statistical value, if the header statistical value is greater than or equal to a preset number of headers, outputting that there is an abnormality in the header part of the traffic data to be detected; if the header statistical value is less than the preset number of headers, determining that there is no abnormality in the header part of the traffic data to be detected; The number of sub-features whose initial detection results are abnormal in the second parameter feature is counted and a parameter statistical value is obtained. If the parameter statistical value is greater than or equal to the preset number of parameters, it is output that there is an abnormality in the parameter part of the flow data to be detected; if the parameter statistical value is less than the preset number of parameters, it is determined that there is no abnormality in the parameter part of the flow data to be detected.

4. The method according to claim 3, characterized in that The method further comprises: If there is an abnormality in the URI part of the traffic data to be detected, determining that the URI detection result is an abnormal type of the traffic data to be detected; If there is an abnormality in the header part of the traffic data to be detected, determining the header detection result as the abnormal type of the traffic data to be detected; If there is an abnormality in the parameter part of the flow data to be detected, the parameter detection result is determined to be the abnormal type of the flow data to be detected.

5. The method according to claim 1, characterized in that Before using the preset first detection model to detect whether the traffic data to be detected is suspicious traffic, the method further includes: Comparing the traffic data to be detected with a preset whitelist; If the traffic data to be detected is in the whitelist, outputting the traffic data to be detected as normal traffic; If the traffic data to be detected is not in the whitelist, the traffic data to be detected is input into a pre-trained first detection model.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Performing quantitative statistics on all target traffic data to be detected within a preset time length; the target traffic data to be detected is the traffic data to be detected characterized as malicious traffic by the final detection result; If the number of target traffic data to be detected obtained by counting within the preset time length is greater than a preset number threshold, all target traffic data to be detected within the preset time length are marked as vulnerability scanning.

7. A model training method, characterized in that: include: Obtain a training data set, wherein each piece of training data in the training data set includes a flow feature corresponding to a piece of flow data, and a first detection result and a second detection result corresponding to the flow data; wherein each piece of training data is annotated with a label indicating whether it is malicious flow; the first detection result is a detection result of the first detection model detecting whether the piece of training data is suspicious flow; the second detection result is a detection result of the second detection model detecting the abnormal type of the piece of training data; The initial model is trained based on the training data set to obtain a trained third detection model.

8. A flow detection device, characterized in that: include: An acquisition module is used to acquire the traffic data to be detected; A processing module, used to detect whether the flow data to be detected is suspicious flow, and obtain a first detection result; If the first detection result indicates that the flow data to be detected is not suspicious flow, outputting the flow data to be detected as normal flow; If the first detection result indicates that the traffic data to be detected is suspicious traffic, a preset first detection model is used to detect the abnormal type of the traffic data to be detected, and a second detection result is obtained; Extract traffic features from the traffic data to be detected, and use a preset second detection model to detect whether the traffic data to be detected is malicious traffic based on the traffic features, the first detection result, and the second detection result, and output a final detection result.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are connected; The memory is used to store programs; The processor is used to call the program stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is executed.