Sample threat assessment method, device, electronic device and storage medium

By using quantitative scoring of multiple data sources and the hierarchical analysis method to assess the threat level of samples, the problem of poor evaluation results from a single data source is solved, more efficient threat level assessment and screening is achieved, and more accurate security defense is supported.

CN113901453BActive Publication Date: 2025-09-23QI AN XIN TECHNOLOGY GROUP INC +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111189226.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-09-23
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

The existing method of assessing the threat level of samples relies on a single data source, resulting in poor assessment results, low accuracy, and prone to missed or false detections.

Method used

By obtaining data samples from multiple data sources, determining multiple indicators, and performing quantitative scoring, the indicator weights are obtained by combining the hierarchical analysis method. The threat level of the samples is judged based on the scoring results, and the similarity judgment of high-threat standard samples is used to finally generate a KS curve to adjust the threshold.

Benefits of technology

It achieves unified threat level assessment of samples from multiple data sources, improves assessment effect and accuracy, can more accurately screen out high-threat samples, and support more targeted security defense measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901453B_ABST
    Figure CN113901453B_ABST
Patent Text Reader

Abstract

The present application provides a sample threat level assessment method, device, electronic device and storage medium, which relate to the field of security technology. The method quantitatively scores multiple indicators used to assess the threat level of data samples, and then judges the threat level of the data samples based on the scoring results of each indicator corresponding to the data samples. In this way, the threat level analysis can be performed on data samples from multiple data sources according to unified indicators, thereby measuring the threat level of data samples from different data sources, with better assessment effect and higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of security technology, and in particular to a sample threat assessment method, device, electronic device, and storage medium. Background Art

[0002] As cybersecurity becomes increasingly challenging, data security analysis becomes increasingly important. Existing methods for analyzing the threat level of samples often rely on a single data source for a specific type of detection result. This results in the threat level assessment being dependent on the quality of the data source. When the data source is incomplete, the threat level assessment is poor and inaccurate, leading to missed or false detections. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to provide a sample threat level assessment method, device, electronic device and storage medium to improve the problem of poor sample threat level assessment effect and low accuracy in the prior art.

[0004] In a first aspect, an embodiment of the present application provides a sample threat assessment method, the method comprising: obtaining multiple data samples from multiple data sources; determining multiple indicators for assessing the threat level of each data sample; quantitatively scoring each indicator to obtain a scoring result for each indicator corresponding to each data sample; and judging the threat level of each data sample based on the scoring result for each indicator corresponding to each data sample.

[0005] In the above implementation process, multiple indicators used to evaluate the threat level of data samples are quantitatively scored, and then the threat level of the data samples is judged based on the scoring results of each indicator corresponding to the data samples. In this way, the threat level analysis can be performed on data samples from multiple data sources according to unified indicators, so that the threat level of data samples from different data sources can be measured, with better evaluation results and higher accuracy.

[0006] Optionally, the quantitative scoring of each indicator to obtain the scoring results of each indicator corresponding to each data sample includes:

[0007] Obtaining the weight of each indicator and the evaluation score of each indicator corresponding to each data sample, wherein the weight represents the importance of the indicator for evaluating the threat level of the data sample, and the evaluation score represents the influence of the indicator on evaluating the threat level of the data sample;

[0008] The scoring results of each indicator corresponding to each data sample are obtained according to the weight and the evaluation score.

[0009] In the above implementation process, the scoring results of each indicator are determined by the two dimensions of weight and evaluation score, which can obtain the scoring results of the indicator more reasonably and accurately.

[0010] Optionally, obtaining the weight of each indicator includes:

[0011] Building a hierarchical model based on the multiple indicators, the hierarchical model including a target layer, a criterion layer, and an indicator layer, the target layer quantifies each indicator, the criterion layer includes the multiple data sources, and the indicator layer includes the multiple indicators;

[0012] The weight of each indicator in the hierarchical structure model is obtained by using the hierarchical analysis method.

[0013] In the above implementation process, since the hierarchical analysis method can combine quantitative analysis with qualitative analysis and be used for decision makers' experience judgment to measure the relative importance of various indicators, the weight of each indicator can be obtained more reasonably through the hierarchical analysis method.

[0014] Optionally, the method of using the analytic hierarchy process to obtain the weight of each indicator in the hierarchical structure model includes:

[0015] Constructing judgment matrices corresponding to the criterion layer and the indicator layer respectively according to the importance scale of each element in each layer in the hierarchical structure model;

[0016] The weight of each indicator is obtained according to the judgment matrix.

[0017] In the above implementation process, by constructing a judgment matrix to obtain the weight of the indicator, the importance of each indicator relative to the target can be measured more accurately, and the weight obtained is more accurate.

[0018] Optionally, the importance scale is an average score after multiple expert users score each element, which can balance the different scoring results of multiple expert users and make the scoring more reasonable.

[0019] Optionally, obtain the evaluation scores of various indicators corresponding to each data sample, including:

[0020] Obtain sample feature data for each data sample;

[0021] The evaluation scores of various indicators corresponding to each data sample are obtained based on the sample feature data of each data sample.

[0022] In the above implementation process, the evaluation score of the indicator is determined based on the sample feature data. In this way, the evaluation score can be determined in combination with the specific business scenario, so that the threat level of the data sample can be better assessed in the specific business scenario.

[0023] Optionally, the sample feature data includes: whether it hits the malware family, whether it hits the APT group, the number of visits, the log protocol type, the email protocol type, the link content, whether there is abnormal host behavior, whether there is abnormal network behavior, and whether it is a malicious released file. At least two of the following.

[0024] Optionally, obtaining the scoring results of the indicators corresponding to each data sample according to the weight and the evaluation score includes:

[0025] The weight of the corresponding indicator is multiplied by the evaluation score, and the product obtained is used as the scoring result of the indicator corresponding to each data sample. The scoring result obtained in this way can better evaluate the impact of the indicator corresponding to each data sample on the threat level assessment.

[0026] Optionally, judging the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample includes:

[0027] Obtaining high-threat standard samples and scoring results of various indicators corresponding to the high-threat standard samples;

[0028] Obtaining the similarity between each data sample and the high-threat standard sample based on the scoring results of each indicator corresponding to each data sample and the scoring results of each indicator corresponding to the high-threat standard sample;

[0029] The threat level of each data sample is determined based on the similarity.

[0030] In the above implementation process, the similarity between the data sample and the high-threat standard sample is judged, so that the threat level of the data sample can be assessed more accurately.

[0031] Optionally, judging the threat level of each data sample according to the similarity includes:

[0032] When the similarity is greater than or equal to a set threshold, it is determined that the threat level of the corresponding data sample is greater than or equal to the set threat level, so that high-threat samples can be screened out from multiple data samples.

[0033] Optionally, the set threshold is determined by taking the maximum evaluation score based on the scoring results of the specified indicator among the multiple indicators and taking the minimum evaluation score based on the scoring results of other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample. In this way, the set threshold can be set more reasonably.

[0034] Optionally, the method further includes:

[0035] Determining a first number of samples greater than or equal to the set threat level and a second number of samples less than the set threat level among the plurality of data samples;

[0036] generating a KS curve based on the first sample size and the second sample size;

[0037] The set threshold is adjusted using the KS curve.

[0038] In the above implementation process, the KS curve is used to adjust the set threshold, so that the set threshold can be flexibly adjusted, thereby further improving the accuracy of the assessment of the threat level of the data sample.

[0039] Optionally, the multiple indicators include at least two of the following: malware family, APT group, IP quintuple, protocol type, email information, link, host behavior, network behavior, and released files. This allows multiple data samples to be measured for their threat level using a unified indicator, resulting in a more effective threat assessment.

[0040] Optionally, after determining the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample, the method further includes:

[0041] Target data samples with a threat level greater than or equal to the set threat level are screened and used as samples for threat behavior analysis. This eliminates the need to analyze all data samples, effectively improving analysis efficiency and making it more targeted. This allows for more accurate security measures to be deployed, ensuring a more secure network.

[0042] In a second aspect, an embodiment of the present application provides a sample threat assessment device, the device comprising:

[0043] A sample acquisition module, used to acquire multiple data samples from multiple data sources;

[0044] an indicator determination module, for determining a plurality of indicators for evaluating the threat level of each data sample;

[0045] The quantitative scoring module is used to quantitatively score each indicator and obtain the scoring results of each indicator corresponding to each data sample;

[0046] The threat level assessment module is used to determine the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample.

[0047] Optionally, the quantitative scoring module is used to obtain the weight of each indicator and the evaluation score of each indicator corresponding to each data sample, wherein the weight represents the importance of the indicator for evaluating the threat level of the data sample, and the evaluation score represents the influence of the indicator on evaluating the threat level of the data sample; and the scoring results of each indicator corresponding to each data sample are obtained based on the weight and the evaluation score.

[0048] Optionally, the quantitative scoring module is used to construct a hierarchical model based on the multiple indicators, the hierarchical model includes a target layer, a criterion layer and an indicator layer, the target layer is for quantifying each indicator, the criterion layer includes the multiple data sources, and the indicator layer includes the multiple indicators; using the hierarchical analysis method, the weight of each indicator in the hierarchical model is obtained respectively.

[0049] Optionally, the quantitative scoring module is used to construct a judgment matrix corresponding to the criterion layer and the indicator layer respectively according to the importance scale of each element in each level in the hierarchical structure model; and obtain the weight of each indicator according to the judgment matrix.

[0050] Optionally, the importance scale is an average score given by multiple expert users to each element.

[0051] Optionally, the quantitative scoring module is used to obtain sample feature data of each data sample; and obtain evaluation scores of various indicators corresponding to each data sample based on the sample feature data of each data sample.

[0052] Optionally, the sample feature data includes: whether it hits the malware family, whether it hits the APT group, the number of visits, the log protocol type, the email protocol type, the link content, whether there is abnormal host behavior, whether there is abnormal network behavior, and whether it is a malicious released file. At least two of the following.

[0053] Optionally, the quantitative scoring module is used to multiply the weight of the corresponding indicator by the evaluation score, and the obtained product is used as the scoring result of the indicator corresponding to each data sample.

[0054] Optionally, the threat level assessment module is used to obtain high-threat standard samples and scoring results of various indicators corresponding to the high-threat standard samples; obtain the similarity between each data sample and the high-threat standard sample based on the scoring results of various indicators corresponding to each data sample and the scoring results of various indicators corresponding to the high-threat standard sample; and judge the threat level of each data sample based on the similarity.

[0055] Optionally, the threat level assessment module is configured to determine that the threat level of the corresponding data sample is greater than or equal to a set threat level when the similarity is greater than or equal to a set threshold.

[0056] Optionally, the set threshold is determined by taking the maximum evaluation score of the scoring results of the specified indicator among the multiple indicators and taking the minimum evaluation score of the scoring results of other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample.

[0057] Optionally, the device further comprises:

[0058] A threshold adjustment module is used to determine the number of first samples greater than or equal to the set threat level and the number of second samples less than the set threat level in the multiple data samples; generate a KS curve based on the first sample number and the second sample number; and use the KS curve to adjust the set threshold.

[0059] Optionally, the multiple indicators include at least two of malware families, APT groups, IP five-tuples, protocol types, email information, links, host behavior, network behavior, and released files.

[0060] Optionally, the device further comprises:

[0061] The screening module is used to screen and obtain target data samples whose threat level is greater than or equal to the set threat level, and the target data samples are used as samples for threat behavior analysis.

[0062] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the method provided in the first aspect above are executed.

[0063] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps in the method provided in the first aspect are executed.

[0064] Other features and advantages of the present application will be described in the following description and, in part, will become apparent from the description or be understood by practicing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0066] Figure 1 A flowchart of a sample threat assessment method provided in an embodiment of the present application;

[0067] Figure 2 A schematic diagram of multiple indicators corresponding to multiple data sources provided in an embodiment of the present application;

[0068] Figure 3 A schematic diagram of a hierarchical structure model provided in an embodiment of the present application;

[0069] Figure 4 A schematic diagram of a KS curve provided in an embodiment of the present application;

[0070] Figure 5 A structural block diagram of a sample threat assessment device provided in an embodiment of the present application;

[0071] Figure 6 A schematic diagram of the structure of an electronic device for executing a sample threat level assessment method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present application.

[0073] It should be noted that the terms "system" and "network" in the embodiments of the present invention are used interchangeably. "Multiple" refers to two or more. In view of this, in the embodiments of the present invention, "multiple" can also be understood as "at least two." "And / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the related objects are in an "or" relationship.

[0074] An embodiment of the present application provides a sample threat assessment method, which quantitatively scores multiple indicators used to assess the threat level of data samples, and then determines the threat level of the data samples based on the scoring results of each indicator corresponding to the data samples. In this way, the threat level of data samples from multiple data sources can be analyzed according to unified indicators, thereby measuring the threat level of data samples from different data sources, with better assessment results and higher accuracy.

[0075] Please refer to Figure 1 , Figure 1 A flowchart of a sample threat assessment method provided in an embodiment of the present application, the method comprising the following steps:

[0076] Step S110: Acquire multiple data samples from multiple data sources.

[0077] Among them, multiple data sources can be understood as referring to data sources with security threats. In the embodiment of this application, multiple data sources can include but are not limited to: sample homology analysis results, sample propagation logs, sample dynamic behavior logs, etc. By selecting reasonable data sources to conduct threat assessment on data samples, the problem of poor assessment results caused by incomplete single data sources can be effectively solved. It can be understood that in the embodiment of this application, data samples from these three data sources are used as examples for threat assessment. In actual applications, data samples from more different data sources can also be obtained to conduct threat assessment.

[0078] Step S120: Determine multiple indicators for evaluating the threat level of each data sample.

[0079] Taking the above three data sources as an example, each data source can be divided into different indicators. For example, for the sample homology analysis results, it can include two indicators: malware family, Advanced Persistent Threat (APT) group; for the sample propagation log, it can include four indicators: IP five-tuple, protocol type, email information, link, etc.; for the sample dynamic behavior analysis results, it can include three indicators: host behavior, network behavior, and released files. In this way, a total of nine indicators can be obtained, such as Figure 2 In this way, the threat level of multiple data samples can be measured with a unified indicator, making the threat level assessment more effective.

[0080] It is understandable that, in specific applications, the multiple indicators used to assess the threat level of data samples may include at least two of these nine indicators. Of course, the specific number of indicators can be flexibly selected according to actual business needs, and the specific indicator content can also be flexibly set according to actual business needs. In other words, the corresponding indicators can be determined according to different data sources, and the setting of indicators can be related to the actual business.

[0081] Step S130: quantify and score each indicator to obtain the scoring results of each indicator corresponding to each data sample.

[0082] Since multiple indicators are determined based on various data sources, in order to evaluate the threat level of data samples from multiple data sources according to a unified standard, this application quantitatively scores each indicator and then evaluates the threat level of the data sample based on the scoring results.

[0083] Among them, quantitative scoring refers to quantifying each indicator according to its importance and the degree of influence on the evaluation data sample, so that the threat level of the data sample can be evaluated through the quantified scoring results.

[0084] Step S140: Determine the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample.

[0085] Each data source may contain a large number of data samples. For each data sample, the scoring results of each of the multiple indicators corresponding to each data sample can be obtained. Among them, the scoring results for the same indicator in each data sample may be different. For example, for the indicator "Host Behavior", the score result of this indicator in data sample 1 is 5 points (if the full score is 10 points), while in data sample 2, the score result of this indicator is 3 points. The scoring results are determined based on the importance of each indicator and the degree of influence in each data sample. Therefore, the same indicator may have different scoring results in different data samples.

[0086] In the embodiment of the present application, for each data sample, it corresponds to the scoring results of 9 indicators. Finally, when performing the threat level assessment, one way is to add the scoring results of the 9 indicators corresponding to each data sample or take the average of the 9 scoring results, and use the total scoring result or average score to assess the threat level of the data sample. In some embodiments, the higher the total scoring result or average score, the higher the threat level of its data sample. Conversely, the lower the total scoring result or average score, the lower the threat level of its data sample. Data samples whose total scoring results or average scores are greater than or equal to the set score can be screened out. These data samples can be considered as high-threat samples with a higher threat level. These high-threat samples can be analyzed later to provide a data basis for the deployment of security defense measures.

[0087] In the above implementation process, multiple indicators used to evaluate the threat level of data samples are quantitatively scored, and then the threat level of the data samples is judged based on the scoring results of each indicator corresponding to the data samples. In this way, the threat level analysis can be performed on data samples from multiple data sources according to unified indicators, so that the threat level of data samples from different data sources can be measured, with better evaluation results and higher accuracy.

[0088] On the basis of the above embodiment, in the method of quantitatively scoring each indicator, the weight of each indicator and the evaluation score of each indicator corresponding to each data sample can be obtained, wherein the weight represents the importance of the indicator for evaluating the threat level of the data sample, and the evaluation score represents the influence of the indicator on evaluating the threat level of the data sample. Then, the scoring results of each indicator corresponding to each data sample can be obtained based on the weight and the evaluation score.

[0089] Among them, the weight of the indicators can be obtained by using information concentration method (such as factor analysis method, principal component analysis method), relative size of numbers (such as Analytic Hierarchy Process (AHP), priority diagram method), information volume (such as entropy method), data volatility or correlation (such as independence weight, information volume weight method), etc. These methods can be used to obtain the weight of each indicator. The larger the weight, the more important the indicator is for evaluating the threat level of the data sample. Conversely, the smaller the weight, the less important the indicator is for evaluating the threat level of the data sample.

[0090] The evaluation score can be understood as the degree of influence of each indicator among the nine indicators corresponding to each data sample on the evaluation of the threat level of the data sample. For example, the higher the evaluation score, the greater the impact on the threat level of the evaluated data sample. Conversely, the lower the evaluation score, the smaller the impact on the threat level of the evaluated data sample.

[0091] In the above implementation process, the scoring results of each indicator are determined by the two dimensions of weight and evaluation score, which can obtain the scoring results of the indicator more reasonably and accurately.

[0092] On the basis of the above embodiment, since the importance of the nine indicators is closely related to the business, the analytic hierarchy process (AHP) can be used to obtain the weight of each indicator in the process of quantitative scoring of the indicators. The analytic hierarchy process (AHP) can be understood as a decision-making method that decomposes elements related to decision-making into levels such as goals, criteria, and plans, and performs qualitative and quantitative analysis on this basis. It can decompose the problem into different levels of cohesion and combination according to the nature of the problem and the overall goal to be achieved, forming a multi-level structural model, and reasonably giving the weight of each decision plan.

[0093] In the process of using AHP to obtain the weights of each indicator, we can first build a hierarchical model based on multiple indicators, such as Figure 3 As shown in the figure, the hierarchical model includes a target layer, a criterion layer and an indicator layer. The target layer is used to quantify each indicator, the criterion layer includes multiple data sources, and the indicator layer includes multiple indicators. Then, the hierarchical analysis method is used to obtain the weight of each indicator in the hierarchical model.

[0094] Among them, the target layer in the hierarchical model refers to the target to be solved, the criterion layer refers to the main factors affecting the target, and the indicator layer refers to the specific plan. In the embodiment of the present application, the target to be solved is to quantify each indicator, the main factors affecting the target are multiple data sources, and the specific plan refers to each indicator.

[0095] When building a hierarchical model, the relevant factors can be decomposed into several levels from top to bottom according to their different attributes. Factors in the same level are subordinate to or influence the factors in the previous level, while also dominating or being influenced by the factors in the next level. This concept can be used to build a corresponding hierarchical model in combination with the application scenarios of this application.

[0096] In the above implementation process, since the hierarchical analysis method can combine quantitative analysis with qualitative analysis and be used for decision makers' experience judgment to measure the relative importance of various indicators, the weight of each indicator can be obtained more reasonably through the hierarchical analysis method.

[0097] Based on the above embodiment, using the analytic hierarchy process to obtain indicator weights refers to comparing and ranking elements to obtain a hierarchical ranking. The hierarchical ranking can be understood as the process of calculating the relative importance of all factors at a certain level to the highest level, that is, to the target, and then ranking them. For example, a specific implementation method can be: first, based on the importance scale of each element at each level in the hierarchical structure model, construct a judgment matrix corresponding to the criterion level and the indicator level, respectively, and then obtain the weight of each indicator based on the judgment matrix.

[0098] The importance scale of each element may be determined by expert scoring, and the expert scoring follows certain rules, such as Santy's 1-9 scale method or three-scale method, or the scale value may be determined by the expert.

[0099] In some embodiments, in order to avoid a certain error in the scoring result caused by one expert scoring, the above importance scale can also be the average score of multiple expert users scoring each element. Since different experts consider the importance of each indicator differently, the scoring results may be inconsistent. Therefore, considering the final scoring effect, the average score of multiple expert users scoring each element can be taken as the importance scale. For example, if n expert users are used to score, the scoring result is recorded as x. i , then the average of the scores given by multiple expert users is finally taken as the importance scale, that is,

[0100] Therefore, in the embodiment of the present application, combining the scores of multiple expert users with AHP can make the weight design more in line with actual application scenarios, easier to operate, and more reasonable results.

[0101] When constructing a judgment matrix, the appropriate scale can be determined by comparing each element pairwise. The specific process is to compare different elements (such as element i and element j) to obtain the value X ij (representing the comparison result of element i relative to element j) is filled in the position of row i and column j of the judgment matrix. For example, if X ij =1, it means that element i and element j have the same importance to the elements of the previous level. If X ij =3, indicating that element i is slightly more important than element j.

[0102] In the embodiment of the present application, the criterion layer includes: sample homology analysis results, sample propagation logs and sample dynamic behavior analysis results. At this time, the scoring result X ij Available B ij Indicates that construct n A Order positive reciprocal matrix (i.e. judgment matrix) Where i, j = 1, 2, 3, that is:

[0103]

[0104] In the same way, the judgment matrix of the indicator layer can be constructed, such as the relationship between the indicators corresponding to the sample homology analysis results. Positive reciprocal matrix of order Where i, j = 1, 2; this judgment matrix can represent the importance of two indicators to the sample homology analysis results;

[0105] The relationship between the indicators corresponding to the sample propagation log Positive reciprocal matrix of order Where i, j = 3, 4, 5, 6; this judgment matrix can represent the importance of two indicators to the sample propagation log;

[0106] The relationship between the indicators corresponding to the sample dynamic behavior analysis results Positive reciprocal matrix of order Among them, i, j = 7, 8, 9; this judgment matrix can represent the importance of two indicators to the dynamic behavior analysis results of the sample.

[0107] Therefore, according to the above method, a total of 4 judgment matrices can be obtained. In order to verify the accuracy of the judgment matrix, the judgment matrix can also be checked for consistency. The verification process can be as follows:

[0108] Define getW(M)→(W,λ max ): Indicates obtaining the maximum eigenvalue λ of the judgment matrix M (n rows and n columns) max and eigenvector W, the judgment matrix M is the above four judgment matrices, and the calculation is performed as follows: (1) Normalize each column of the judgment matrix, that is, (2) Add row by row to get the sum vector, that is, (3) The sum vector w is obtained i Regularization, that is Get the eigenvector W=(w1,...,w i ) T ; (4) Calculate the maximum eigenvalue using the sum-product method Then output W and λ max , at this time W is the maximum eigenvalue λ max The corresponding eigenvector, each component in the eigenvector can represent the relative importance of each element.

[0109] Define isConsitst(M,λ max )→true or false: indicates the consistency of the n-order positive reciprocal matrix M. If the input is M and λ max , perform the following calculation process: (1) Calculate the consistency index (2) Calculate the test coefficient Among them, RI is called the average random consistency index, which is related to the matrix order n and can be obtained by looking up the RI index table.

[0110] When CR < 0.10, the consistency check of M is considered to have passed and true is output, otherwise false is output. Therefore, the eigenvectors and maximum eigenvalues ​​of the four matrices A, B1, B2, and B3 can be calculated according to the above process, and then consistency checks can be performed on each matrix.

[0111] If we first call getW(A) on the matrix A, we can get the eigenvector W of A. A and the maximum eigenvalue λ Amax , then call isConsitst(A,λ Amax ), if the return value is true, it means that the eigenvector W of A A Reasonable; for matrix B1, first call getW(B1) to get the eigenvector of B1 and the maximum eigenvalue Then call If the return value is true, it means that the eigenvector of B1 Reasonable; for matrix B2, first call getW(B2) to get the eigenvector of B2 and the maximum eigenvalue Then call If the return value is true, it means that the eigenvector of B2 Reasonable; for matrix B3, first call getW(B3) to get the eigenvector of B3 and the maximum eigenvalue Then call If the return value is true, it means that the eigenvector of B3 Reasonable.

[0112] For example, for the eigenvector corresponding to the criterion layer in, Indicates the importance of sample homology analysis results relative to the target layer, Indicates the importance of the sample propagation log relative to the target layer, Indicates the importance of the sample dynamic behavior analysis results relative to the target layer.

[0113] For the eigenvectors corresponding to the judgment matrix between malware families and APT groups in the indicator layer in, Indicates the importance of the malware family relative to the sample homology analysis results, Indicates the importance of the APT group relative to the sample homology analysis results; the eigenvector corresponding to the judgment matrix between the IP quintuple, protocol type, email information and link in, Indicates the importance of the IP quintuple relative to the sample propagation log. Indicates the importance of the protocol type relative to the sample propagation log, Indicates the importance of the email message relative to the sample propagation log. Indicates the importance of the link relative to the sample propagation log; the eigenvector corresponding to the judgment matrix between host behavior, network behavior and released files in, Indicates the importance of host behavior relative to the sample dynamic behavior analysis results, w C8 Indicates the importance of network behavior relative to the results of sample dynamic behavior analysis. Indicates the importance of the dropped file relative to the sample dynamic behavior analysis results.

[0114] After obtaining the eigenvectors of each matrix in the above process, the weight of each indicator can be the product of the indicator layer weight and the criterion layer weight, such as:

[0115]

[0116]

[0117]

[0118] According to the above formula, when calculating the weight of each indicator, the weight of each indicator can be obtained by the following formula, such as:

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128] In this way, the weights of each of the above 9 indicators can be obtained, and the sum of the weights of the 9 indicators is equal to 1. It should be noted that when quantifying, the type and number of each indicator can vary depending on the business, so the indicator layer can be designed as a separate template, which does not affect the main category quantification. In this way, regardless of the indicator, the above method can be used to quantify the indicator, which is conducive to being applicable to different application scenarios.

[0129] In the above implementation process, by constructing a judgment matrix to obtain the weight of the indicator, the importance of each indicator relative to the target can be measured more accurately, and the weight obtained is more accurate.

[0130] Based on the above embodiment, in the method of obtaining the evaluation scores of each indicator corresponding to each indicator sample, the sample feature data of each data sample can be obtained, and then the evaluation scores of each indicator corresponding to the data sample can be obtained based on the sample feature data of each data sample.

[0131] Among them, in an embodiment of the present application, the sample feature data of the data sample may include: whether it hits the malware family, whether it hits the APT group, the number of visits, the log protocol type, the email protocol type, the link content, whether there is abnormal host behavior, whether there is abnormal network behavior, and whether it is a malicious released file. At least two of the following.

[0132] For example, the maximum score of each indicator is set to be consistent, such as 10 points. Of course, the maximum score can be flexibly set according to actual needs. Then the evaluation scores of the 9 indicators corresponding to a data sample can be obtained as follows, such as for data sample 1:

[0133] (1) Indicator: Malware family. Determine whether the homology analysis result of data sample 1 hits the malware family. If it hits, the output evaluation score is 10. Otherwise, if it does not hit, the output evaluation score is 0. In addition, if the malware family can be classified, if it is divided into multiple categories, it can be determined which category the data sample 1 hits, and the evaluation score can be graded according to the hit result. For example, if it hits the first category, the evaluation score is 2, and if it hits the second category, the evaluation score is 3, etc.

[0134] (2) Indicator: APT group. Determine whether the homology analysis result of data sample 1 hits the APT group. If it hits, the output evaluation score is 10. Otherwise, if it does not hit, the output evaluation score is 0.

[0135] (3) Indicator: IP quintuple. Obtain the IP quintuple of data sample 1, count the number of visits to the IP quintuple in the sample propagation log, and determine the evaluation score based on the number of visits. If the number of visits is less than the first threshold, it can be considered an attack behavior, and the output evaluation score is 10 points. If the number of visits is greater than the second threshold, it can be considered a normal access behavior, and the output evaluation score is 3 points. If the number of visits is between the first and second thresholds, the output evaluation score is 5 points. Of course, the output evaluation scores corresponding to different numbers of visits can also be flexibly set according to actual needs.

[0136] (4) Indicator: Protocol type. The log protocol type of data sample 1 is obtained. If the log protocol type is a common web protocol (such as HTTP), the threat level of data sample 1 is considered low, and the output evaluation score can be low. Conversely, if the log protocol type is an email protocol (such as SMTP), the threat level of data sample 1 is considered high, and the output evaluation score is high. It is understandable that the specific output evaluation score can be flexibly set according to actual needs.

[0137] (5) Indicator: Email information. If the protocol of data sample 1 is an email protocol type, then the specific email information is obtained. If it is judged to be a malicious email based on the email information, the output evaluation score is high, otherwise, the output evaluation score is low. In addition, the email sender's email address and sender's IP address, the recipient's email address and recipient's IP address can also be obtained to determine whether this information comes from key units, that is, whether these key units are on the blacklist. Alternatively, it can be determined whether the suffix of the email attachment is on the blacklist. If so, the output evaluation score is high, otherwise, the output evaluation score is low. It can be understood that the specific output evaluation score can be flexibly set according to actual needs.

[0138] (6) Indicator: Link. Parse the link of data sample 1. If the protocol type of data sample 1 is a Web protocol type, make a judgment based on the parsed content of the link (i.e., the link content). The parsed content may include an IP address or domain name and a file extension. If the IP address is from a key unit (i.e., a blacklist) and / or the file extension is from a blacklist, the output evaluation score is high. Otherwise, the output evaluation score is low. It is understandable that the specific output evaluation score can be flexibly set according to actual needs.

[0139] (7) Indicator: Host behavior. Obtain the host behavior of the dynamic behavior analysis results of data sample 1 and determine whether there is abnormal host behavior (such as disabling the proxy, which may be used for traffic hijacking). If so, the output evaluation score is high, otherwise, the output evaluation score is low. Alternatively, the number of abnormal host behaviors can be obtained. If the number is large, the evaluation score is high, and if the number is small, the evaluation score is low. It is understandable that the specific output evaluation score can be flexibly set according to actual needs.

[0140] (8) Indicator: Network behavior. The network behavior of the dynamic behavior analysis results of data sample 1 is obtained. If there is abnormal network behavior (for example, the IP address obtained by resolving the domain name is from the blacklist), the output evaluation score is high; otherwise, the output evaluation score is low. It is understandable that the specific output evaluation score can be flexibly set according to actual needs.

[0141] (9) Indicator: Released file. Based on the released file of the dynamic behavior analysis result of data sample 1, determine whether the released file is a malicious released file. If so, the output evaluation score is high; otherwise, the output evaluation score is low. It is understandable that the specific output evaluation score can be flexibly set according to actual needs.

[0142] Therefore, according to the above method, the evaluation scores of the 9 indicators corresponding to each data sample can be obtained, such as the evaluation score V = (V1, V2, ..., V9), where V i Represents the evaluation score corresponding to the i-th indicator.

[0143] It should be noted that the two indicators, email information and links, are mutually exclusive (because a data sample is either a web protocol type or an email protocol type). If the evaluation score corresponding to the email information is greater than 0, the evaluation score corresponding to the link is equal to 0.

[0144] In the above implementation process, the evaluation score of the indicator is determined based on the sample feature data. In this way, the evaluation score can be determined in combination with the specific business scenario, so that the threat level of the data sample can be better assessed in the specific business scenario.

[0145] Based on the above example, after obtaining the weight and evaluation score for each indicator, the scoring result for each indicator can be multiplied by the corresponding indicator weight and evaluation score, and the resulting level is used as the scoring result for the indicator corresponding to each data sample. The scoring result obtained in this way can better assess the impact of the indicator corresponding to each data sample on the threat level assessment.

[0146] For example, scoring result = weight * evaluation score = dw i ×V i , where dw i is the weight of the i-th indicator, V i is the evaluation score of the i-th indicator. Of course, the scoring result can also be obtained by other calculation methods using weights and evaluation scores. It is only necessary to follow the principle that the larger the weight, the larger the scoring result, and the larger the evaluation score, the larger the scoring result. That is, the scoring result is positively correlated with the weight and evaluation score.

[0147] On the basis of the above embodiment, in the method of judging the threat level of each data sample, in addition to summing or averaging the scoring results to judge the threat level of the data sample, a high threat standard sample can be pre-set, and the evaluation scores of each indicator corresponding to the high threat standard sample can be full marks, such as the evaluation result is represents the evaluation result of the i-th indicator, dw i The weights obtained according to the AHP method can be Indicates the maximum score, such as 10.

[0148] The high-threat standard samples can be pre-established and stored. When judging the threat level, the high-threat standard samples and the scoring results of each indicator corresponding to the high-threat standard samples can be obtained first. Then, based on the scoring results of each indicator corresponding to each data sample and the scoring results of each indicator corresponding to the high-threat standard sample, the similarity between each data sample and the high-threat standard sample is obtained, and the threat level of each data sample is judged based on the similarity.

[0149] For example, the vector corresponding to the scoring results of each indicator corresponding to the high-threat standard sample is The vector corresponding to the score results for each indicator for each data sample is X(scoreD1, scoreD2, ..., scoreD9). When calculating similarity between two vectors, methods such as Euclidean distance and cosine similarity can be used. In practical applications, different calculation methods can be used based on specific business needs. For example, Euclidean distance determines the similarity between a data sample and a high-threat standard sample based on distance. The smaller the distance, the higher the similarity. Cosine similarity, on the other hand, determines the similarity based on the angle between the data sample and the high-threat standard sample. The larger the cosine value, the higher the similarity.

[0150] If you can choose Euclidean distance to determine the similarity between two samples, you can calculate the Euclidean distance based on the scoring results of each indicator corresponding to each data sample and the scoring results of each indicator corresponding to the high-threat standard sample. The Euclidean distance can be used to characterize the similarity between each data sample and the high-threat standard sample.

[0151] The calculation formula of Euclidean distance is as follows:

[0152]

[0153] Since the value range of Euclidean distance is large, it is generally normalized, that is, Among them, simDis(X,Y)∈[0,1], the larger the value is, the greater the threat level of the data sample is.

[0154] You can also choose cosine similarity to determine the similarity between two samples. This is achieved by calculating the cosine similarity between the score of each data sample and the score of the high-threat standard sample. The calculation method is as follows:

[0155] The value range is (0,1). The larger the value, the greater the threat level of the data sample.

[0156] Therefore, when determining the threat level of a data sample, when it is determined that the similarity is greater than or equal to a set threshold, the threat level of the corresponding data sample is determined to be greater than or equal to the set threat level. For example, when the Euclidean distance or cosine similarity is greater than or equal to a set threshold, the threat level of the data sample is greater than or equal to the set threat level.

[0157] Based on the above embodiment, the threshold value can be flexibly set according to actual circumstances and is related to actual business. For example, the threshold value can be determined by correlation indicators (such as formulation principles), expert consultation and scoring (such as expert scoring in AHP), etc. Taking correlation indicators as an example, the threshold value can be determined by taking the maximum evaluation score of the scoring results of a specified indicator among multiple indicators and taking the minimum evaluation score of the scoring results of other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample.

[0158] For example, the design principles for setting thresholds may include: Principle 1: If a data sample originates from an APT group, the threat level of the data sample is high and it is a high-threat sample; Principle 2: If the dynamic behavior detection results of the data sample are abnormal, the threat level of the data sample is high and it may be a high-threat sample. Therefore, when determining the threshold, it should be ensured that Principles 1 and Principle 2 are met, that is, the highest score is taken for the evaluation score of the indicator (APT group), and the lowest score is taken for the evaluation scores of the other indicators. The evaluation result is According to Principle 1, the initial threshold recomScore1=sim(X1,Y) is set, which represents the similarity with the high-threat standard sample (such as Euclidean distance or cosine distance), where Y represents the score result of the high-threat standard sample.

[0159] Based on principle 2, for the indicators: host behavior, network behavior, and released files, the evaluation scores corresponding to at least two of the three indicators can be full marks. There are 7 combined evaluation results. Finally, the average of these 7 can be taken, and the initial threshold

[0160] The initial threshold is set according to the above two principles, so the final threshold can be threshold=recomScore2<recomScore1, which can make the threshold more reasonable.

[0161] It can be understood that the above-mentioned set threshold can be set manually according to the above-mentioned two principles, and as the system runs, the above-mentioned set threshold can also be flexibly adjusted according to the evaluation of the data sample. In order to achieve more accurate adjustment of the set threshold, a machine learning algorithm can be used for adjustment.

[0162] For example, if a threat level assessment is required for 1,000 data samples, after assessing the threat level of these 1,000 data samples using the above-described assessment method, a first number of samples greater than or equal to the set threat level and a second number of samples less than the set threat level can be determined from the plurality of data samples. For example, the data samples greater than or equal to the set threat level are considered predicted positive samples, and the remaining samples less than the set threat level are considered predicted negative samples. A KS curve can then be generated based on the first number of samples and the second number of samples. For example, the 1,000 data samples can be manually labeled as positive or negative, that is, each data sample is manually confirmed to be a positive or negative sample. A KS curve can then be established based on this data.

[0163] In order to construct the KS curve, we can first calculate the true positive rate (TPR), TPR = TP / (TP+FN), which can represent the proportion of positive samples identified by the evaluation method to all positive samples (i.e., the cumulative proportion of positive samples), and calculate the false positive rate (FPR), FPR = FP / (FP+TN), which can represent the proportion of negative samples mistakenly identified as positive samples by the evaluation method to all negative samples (i.e., the cumulative proportion of negative samples). Among them, TP represents the number of samples manually confirmed as positive and predicted as positive by the evaluation method (i.e., the number of first samples), FN represents the number of samples manually confirmed as positive but predicted as negative by the evaluation method (i.e., the number of second samples), FP represents the number of samples manually confirmed as negative and predicted as positive by the evaluation method, and TN represents the number of samples manually confirmed as negative and predicted as negative by the evaluation method. Therefore, TPR and FPR can be calculated according to the above calculation formula, and the KS curve can be drawn according to TPR and FPR, as shown in the following example. Figure 4 As shown, the horizontal axis of the KS curve can be understood as the threshold. The maximum difference between TPR and FPR indicates that the evaluation method can distinguish positive and negative samples to a greater extent, that is, the greater the prediction accuracy, so the threshold corresponding to the maximum difference can be used as the adjusted set threshold.

[0164] It is understandable that, as the system runs, the KS curve can be generated by continuously evaluating the data samples, and then the threshold corresponding to the maximum difference between TPR and FPR is found to update the current set threshold. That is, each time the set threshold is adjusted, the threshold corresponding to the maximum difference between TPR and FPR is used as the adjusted set threshold. In this way, the set threshold can be flexibly adjusted, thereby further improving the accuracy of the assessment of the threat level of the data samples.

[0165] On the basis of the above embodiments, in order to facilitate the subsequent deployment of security defense measures, it is also possible to obtain the threat level of each data sample and then screen out target data samples whose threat level is greater than or equal to the set threat level. The target data samples can be used as samples for threat behavior analysis, such as analyzing the source of these samples, or analyzing the security risks of these samples (such as analyzing the source IP address, destination IP address, main attack behavior, etc. of these samples). By analyzing these target data samples, it is convenient to provide data basis for the subsequent deployment of related security defense measures.

[0166] Moreover, by screening out target data samples for analysis instead of analyzing all data samples, the analysis efficiency can be effectively improved and more targeted, making the security defense measures deployed subsequently more accurate, thereby ensuring a safer network.

[0167] Alternatively, the threat level corresponding to the target data sample can also be output, and its threat level is the above-mentioned similarity. In this way, by comparing the data sample with the high-threat standard sample, the effectiveness of the recommended score (that is, samples greater than or equal to the set threat level are recommended as high-threat samples) can be ensured, which is conducive to distinguishing high-threat samples from other samples and improving the accuracy of high-threat sample screening.

[0168] Please refer to Figure 5 , Figure 5 This is a structural block diagram of a sample threat assessment device 200 provided in an embodiment of the present application. The device 200 may be a module, program segment or code on an electronic device. It should be understood that the device 200 is similar to the above-mentioned Figure 1 The method embodiment corresponds to the embodiment that can be executed Figure 1 The various steps involved in the method embodiment and the specific functions of the device 200 can be found in the description above. To avoid repetition, detailed description is appropriately omitted here.

[0169] Optionally, the apparatus 200 includes:

[0170] The sample acquisition module 210 is used to acquire multiple data samples from multiple data sources;

[0171] an indicator determination module 220 for determining a plurality of indicators for evaluating the threat level of each data sample;

[0172] The quantitative scoring module 230 is used to perform quantitative scoring on each indicator and obtain the scoring results of each indicator corresponding to each data sample;

[0173] The threat level assessment module 240 is configured to determine the threat level of each data sample based on the scoring results of the various indicators corresponding to each data sample.

[0174] Optionally, the quantitative scoring module 230 is used to obtain the weight of each indicator and the evaluation score of each indicator corresponding to each data sample, wherein the weight represents the importance of the indicator for evaluating the threat level of the data sample, and the evaluation score represents the influence of the indicator on evaluating the threat level of the data sample; and the scoring results of each indicator corresponding to each data sample are obtained based on the weight and the evaluation score.

[0175] Optionally, the quantitative scoring module 230 is used to construct a hierarchical model based on the multiple indicators, the hierarchical model includes a target layer, a criterion layer and an indicator layer, the target layer is for quantifying each indicator, the criterion layer includes the multiple data sources, and the indicator layer includes the multiple indicators; using the hierarchical analysis method, the weight of each indicator in the hierarchical model is obtained respectively.

[0176] Optionally, the quantitative scoring module 230 is used to construct a judgment matrix corresponding to the criterion layer and the indicator layer respectively according to the importance scale of each element in each layer in the hierarchical structure model; and obtain the weight of each indicator according to the judgment matrix.

[0177] Optionally, the importance scale is an average score given by multiple expert users to each element.

[0178] Optionally, the quantitative scoring module 230 is configured to obtain sample feature data of each data sample; and obtain evaluation scores of various indicators corresponding to each data sample based on the sample feature data of each data sample.

[0179] Optionally, the sample feature data includes: whether it hits the malware family, whether it hits the APT group, the number of visits, the log protocol type, the email protocol type, the link content, whether there is abnormal host behavior, whether there is abnormal network behavior, and whether it is a malicious released file. At least two of the following.

[0180] Optionally, the quantitative scoring module 230 is configured to multiply the weight of the corresponding indicator by the evaluation score, and use the obtained product as the scoring result of the indicator corresponding to each data sample.

[0181] Optionally, the threat level assessment module 240 is used to obtain high-threat standard samples and scoring results of various indicators corresponding to the high-threat standard samples; obtain the similarity between each data sample and the high-threat standard sample based on the scoring results of various indicators corresponding to each data sample and the scoring results of various indicators corresponding to the high-threat standard sample; and judge the threat level of each data sample based on the similarity.

[0182] Optionally, the threat level assessment module 240 is configured to determine that the threat level of the corresponding data sample is greater than or equal to a set threat level when the similarity is greater than or equal to a set threshold.

[0183] Optionally, the set threshold is determined by taking the maximum evaluation score of the scoring results of the specified indicator among the multiple indicators and taking the minimum evaluation score of the scoring results of other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample.

[0184] Optionally, the apparatus 200 further includes:

[0185] A threshold adjustment module is used to determine the number of first samples greater than or equal to the set threat level and the number of second samples less than the set threat level in the multiple data samples; generate a KS curve based on the first sample number and the second sample number; and use the KS curve to adjust the set threshold.

[0186] Optionally, the multiple indicators include at least two of malware families, APT groups, IP five-tuples, protocol types, email information, links, host behavior, network behavior, and released files.

[0187] Optionally, the apparatus 200 further includes:

[0188] The screening module is used to screen and obtain target data samples whose threat level is greater than or equal to the set threat level, and the target data samples are used as samples for threat behavior analysis.

[0189] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0190] Please refer to Figure 6 , Figure 6A structural diagram of an electronic device for executing a sample threat assessment method provided in an embodiment of the present application, the electronic device may include: at least one processor 310, such as a CPU, at least one communication interface 320, at least one memory 330 and at least one communication bus 340. Among them, the communication bus 340 is used to realize direct connection and communication between these components. Among them, the communication interface 320 of the device in the embodiment of the present application is used to communicate signaling or data with other node devices. The memory 330 can be a high-speed RAM memory or a non-volatile memory (non-volatile memory), such as at least one disk memory. The memory 330 can optionally also be at least one storage device located away from the aforementioned processor. Computer-readable instructions are stored in the memory 330. When the computer-readable instructions are executed by the processor 310, the electronic device executes the above-mentioned Figure 1 The method process shown.

[0191] I understand. Figure 6 The structure shown is only for illustration, and the electronic device may also include Figure 6 More or fewer components than shown, or with Figure 6 Different configurations shown. Figure 6 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0192] The embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program performs the following operations: Figure 1 The method process in the illustrated method embodiment is performed by the electronic device.

[0193] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can perform the methods provided by the above-mentioned method embodiments, for example, including: obtaining multiple data samples from multiple data sources; determining multiple indicators for evaluating the threat level of each data sample; quantitatively scoring each indicator to obtain a scoring result for each indicator corresponding to each data sample; and judging the threat level of each data sample based on the scoring result for each indicator corresponding to each data sample.

[0194] In summary, the embodiments of the present application provide a sample threat level assessment method, device, electronic device and storage medium, which quantitatively scores multiple indicators used to assess the threat level of data samples, and then judges the threat level of the data samples based on the scoring results of each indicator corresponding to the data samples. In this way, the threat level analysis can be performed on data samples from multiple data sources according to unified indicators, so that the threat level of data samples from different data sources can be measured, with better assessment effect and higher accuracy.

[0195] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0196] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0197] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0198] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0199] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A sample threat assessment method, characterized in that: The method comprises: Obtain multiple data samples from multiple data sources; Identify multiple metrics for assessing the threat level of each data sample; Quantitatively score each indicator and obtain the scoring results of each indicator corresponding to each data sample; Determine the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample; The step of judging the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample includes: Obtaining high-threat standard samples and scoring results of various indicators corresponding to the high-threat standard samples; Obtaining the similarity between each data sample and the high-threat standard sample based on the score results of each indicator corresponding to each data sample and the score results of each indicator corresponding to the high-threat standard sample; Determining the threat level of each data sample based on the similarity; The step of determining the threat level of each data sample based on the similarity includes: When the similarity is greater than or equal to a set threshold, it is determined that the threat level of the corresponding data sample is greater than or equal to the set threat level; The set threshold is determined based on the maximum evaluation score of the scoring results of the specified indicator among the multiple indicators and the minimum evaluation score of the scoring results of the other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample; The method further comprises: determining a first number of samples greater than or equal to the set threat level and a second number of samples less than the set threat level among the plurality of data samples; generating a KS curve based on the first sample size and the second sample size; The set threshold is adjusted using the KS curve.

2. The method according to claim 1, characterized in that The quantitative scoring of each indicator to obtain the scoring results of each indicator corresponding to each data sample includes: Obtaining the weight of each indicator and the evaluation score of each indicator corresponding to each data sample, wherein the weight represents the importance of the indicator for evaluating the threat level of the data sample, and the evaluation score represents the influence of the indicator on evaluating the threat level of the data sample; The scoring results of each indicator corresponding to each data sample are obtained according to the weight and the evaluation score.

3. The method according to claim 2, characterized in that Obtaining the weight of each indicator includes: Building a hierarchical model based on the multiple indicators, the hierarchical model including a target layer, a criterion layer, and an indicator layer, the target layer quantifies each indicator, the criterion layer includes the multiple data sources, and the indicator layer includes the multiple indicators; The weight of each indicator in the hierarchical structure model is obtained by using the hierarchical analysis method.

4. The method according to claim 3, characterized in that The method of using the hierarchical analysis method to obtain the weight of each indicator in the hierarchical structure model includes: Constructing judgment matrices corresponding to the criterion layer and the indicator layer respectively according to the importance scale of each element in each layer in the hierarchical structure model; The weight of each indicator is obtained according to the judgment matrix.

5. The method according to claim 4, characterized in that The importance scale is the average score of each element scored by multiple expert users.

6. The method according to claim 2, characterized in that Get the evaluation scores of each indicator corresponding to each data sample, including: Obtain sample feature data for each data sample; The evaluation scores of various indicators corresponding to each data sample are obtained based on the sample feature data of each data sample.

7. The method according to claim 6, characterized in that The sample feature data includes: whether it hits the malware family, whether it hits the APT group, the number of visits, the log protocol type, the email protocol type, the link content, whether there is abnormal host behavior, whether there is abnormal network behavior, and whether it is a malicious released file. At least two of the following.

8. The method according to any one of claims 1 to 7, characterized in that: The multiple indicators include at least two of malware families, APT groups, IP five-tuples, protocol types, email information, links, host behaviors, network behaviors, and released files.

9. The method according to any one of claims 1 to 7, characterized in that: After determining the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample, the following steps are further included: Target data samples with a threat level greater than or equal to a set threat level are screened and obtained, and the target data samples are used as samples for threat behavior analysis.

10. A sample threat assessment device, characterized in that: The device comprises: A sample acquisition module, used to acquire multiple data samples from multiple data sources; an indicator determination module, for determining a plurality of indicators for evaluating the threat level of each data sample; The quantitative scoring module is used to quantitatively score each indicator and obtain the scoring results of each indicator corresponding to each data sample; The threat level assessment module is used to determine the threat level of each data sample based on the scoring results of each indicator corresponding to each data sample; The threat level assessment module is specifically configured to obtain a high-threat standard sample and the scoring results of each indicator corresponding to the high-threat standard sample; obtain the similarity between each data sample and the high-threat standard sample based on the scoring results of each indicator corresponding to each data sample and the scoring results of each indicator corresponding to the high-threat standard sample; and determine the threat level of each data sample based on the similarity; The threat level assessment module is specifically configured to determine that the threat level of the corresponding data sample is greater than or equal to a set threat level when the similarity is greater than or equal to a set threshold; The set threshold is determined by taking the maximum evaluation score of the scoring results of the specified indicator among the multiple indicators and taking the minimum evaluation score of the scoring results of the other indicators, and the scoring results of each indicator corresponding to the high-threat standard sample; The device further comprises: A threshold adjustment module is used to determine the number of first samples greater than or equal to the set threat level and the number of second samples less than the set threat level among the multiple data samples; generate a KS curve based on the first sample number and the second sample number; and use the KS curve to adjust the set threshold.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 9 is executed.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is executed.

Citation Information

Patent Citations

  • Network threat security detection method and device

    CN111541702A

  • Information source quality analysis method and device, electronic equipment and storage medium

    CN113360739A