A vulnerability exploitation analysis method, system and storage medium based on big data
By building a main information channel and data identification library, a base database based on feature point information and threat coefficients is adopted, and a unified area judgment algorithm is used to solve the speed and scalability problems in big data vulnerability analysis, achieving the effect of quickly identifying threat data.
Patent Information
- Application Number
- CN202510104088.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the analysis of big data vulnerability, it is difficult for the existing technology to improve the running speed and scalability of the algorithm while ensuring accuracy. Choosing suitable data mining and machine learning algorithms is complex and inefficient.
Build a main information channel and data identification library, build a base database based on feature point information and threat coefficients, quickly filter threat data through spider web lines and boundary maps, and use a unified area judgment algorithm to initially judge the threat of data to avoid complex data classification and algorithm selection.
Improve the running speed and scalability of vulnerability analysis, quickly identify potential threat data, and reduce the requirements for algorithm efficiency.
Smart Images

Figure CN119989365B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information security technology, and particularly to a method, system and storage medium for vulnerability exploitation analysis based on big data. Background Art
[0002] The method for vulnerability exploitation analysis based on big data is an advanced means that combines big data technology and vulnerability analysis technology, aiming to mine, analyze and utilize potential vulnerability information from massive data to enhance the security and defense capabilities of the system.
[0003] In the known prior art, after detecting vulnerability data in the system, the collected data is preprocessed, including data cleaning, formatting, deduplication, etc.; a vulnerability database is established to store the preprocessed vulnerability data; big data technology is used to deeply analyze the vulnerability database to discover the correlation and trend between vulnerabilities, and data mining algorithms such as classification, clustering, and scoring models are adopted to identify the patterns and characteristics of vulnerabilities; according to factors such as the type, severity, and impact scope of the vulnerabilities, the possibility of the vulnerabilities being exploited is evaluated; the results of the vulnerability analysis are displayed in a visual manner and a vulnerability analysis report is generated, which details the discovered vulnerabilities, evaluation results, exploitation attempts, and recommended remedial measures.
[0004] However, it is necessary to analyze vulnerability data by means of various data mining and machine learning algorithms. Different algorithms are applicable to different data types and analysis scenarios, and selecting the appropriate algorithm is a complex process; the large amount of data poses high requirements on the efficiency of the algorithm, and it is difficult to improve the running speed and scalability of the algorithm while ensuring accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system and storage medium for vulnerability exploitation analysis based on big data to solve the deficiencies of the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for vulnerability exploitation analysis based on big data, comprising the following steps: constructing a main information channel, with a separation point provided at one end of the main information channel, one end of the conventional channel and one end of the discrimination channel being connected to the separation point, and a data discrimination library being provided at the separation point, the data discrimination library being used to classify existing data into secure data and threat data according to a preset security level and complete the screening of the data;
[0007] Obtaining feature point information based on the existing data, obtaining a feature point threat coefficient A based on preset standard data, constructing a feature chain based on the feature point information and obtaining a feature chain threat coefficient B, and constructing a base database based on the feature point information, the feature point threat coefficient, the feature chain, and the feature chain threat coefficient;
[0008] Set the scanning area based on the discrimination channel, determine the threat coefficient C of new data per unit time based on a simple test, sort according to the value of the threat coefficient C, and determine the parsing order of new data;
[0009] Select a point O on the virtual plane, construct spider web lines with O as the endpoint. There are six spider web lines in total. Mark a maximum characteristic point threat coefficient point on each spider web line and connect them end to end in sequence as the outer boundary line. Mark the characteristic point threat coefficients of the standard data on the spider web lines and mark them as reference points. Connect the reference points end to end in sequence as the standard data boundary map. The outer boundary line is located outside the standard data boundary map. Construct a base map based on the spider web lines, the outer boundary line, and the standard data boundary map;
[0010] Mark the threat coefficients of the unknown characteristic points of the new data obtained based on the simulation test on the spider web lines and mark them as actual points. The actual point and the reference point on the same ray form a control group;
[0011] Based on the free quantity combination of unknown characteristic points and existing characteristic points on the spider web to form unknown characteristic chains, select a point on the base map as the anchor point. Construct a measurement plane based on the actual point, the anchor point, and the point O. Construct a control plane based on the reference point, the anchor point, and the point O. Compare the areas of the measurement plane and the control plane and preliminarily judge the threat of a series of data containing unknown characteristic points and / or unknown characteristic chains. Mark one or a series of data with threats as pending data, conduct comprehensive detection on the pending data, and transmit the classification information to the data discrimination library.
[0012] In a preferred embodiment, the steps of constructing the main information channel and constructing the data discrimination library based on the existing data include:
[0013] Construct a main information channel between the transmission end and the receiving end. The main information channel is used to transmit data, and the data includes: existing data and new data;
[0014] Collect the existing data, divide the existing data into safe data and threat data according to the preset standard data. Construct a data discrimination library based on the safe data and the threat data. The data discrimination library is used to identify the safe information and transmit it to the regular channel. The data discrimination library is used to identify the new information and transmit it to the discrimination channel. The data identification library is used to identify the threat data and intercept the threat data.
[0015] In a preferred embodiment, the steps of obtaining the characteristic point information based on the existing data, obtaining the characteristic point threat coefficient A based on the preset standard data, and constructing the characteristic chain based on the characteristic point information include:
[0016] Obtain the characteristic point information of the existing data. The classification of the characteristic point information includes: the sensitivity of the data, the scale and accuracy of the data, the accessibility and mobility of the data, the timeliness and value of the data, the integrity and accuracy of the data, the ownership and compliance of the data;
[0017] Obtain the threat coefficient of feature points. There are 6 categories of threat coefficients of feature points, namely, the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and mobility, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and the data ownership and compliance a6;
[0018] The formula for the threat coefficient A of feature points is:
[0019] A = ε1a1 + ε2a2 + ε3a3 + ε4a4 + ε5a5 + ε6a6
[0020] Wherein, ε1, ε2, ε3, ε4, ε5, ε6 are the weight coefficients of a1, a2, a3, a4, a5, a6 in the threat coefficient of feature points respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data;
[0021] Form a feature chain through at least two pieces of feature point information.
[0022] In a preferred embodiment, the step of sorting according to the numerical value of the threat coefficient C and determining the parsing order of new data includes:
[0023] Set a scanning area in the identification channel, measure the threat coefficient C of new data per unit time based on VirusTotal. When the number of new data passing through the scanning area per unit time is not less than 1, compare the threat coefficients of all new data and sort them according to the numerical value, and preferentially perform parsing processing on the new data with a large numerical value;
[0024] At least one unknown feature point is included in a new data. A first-level unknown feature chain is formed by combining several unknown feature points, and a second-level unknown feature chain is formed by combining one or more unknown feature points with one or more known feature points. The detection priority of the first-level unknown feature chain is higher than that of the second-level unknown feature chain.
[0025] In a preferred embodiment, the step of constructing a base map based on the spider web line, the outer boundary line, and the standard data boundary map includes:
[0026] Construct the center point O on a plane. Six rays are drawn with O as the common endpoint. The six rays are collectively called spider webs. The included angle between adjacent rays is 60°. Each ray is only marked with the unified threat coefficient of one type of feature point. Mark the unified threat coefficient of the feature point with the largest existing data on the six rays and connect them end to end in sequence to form a hexagon, which is marked as the outer boundary line. Mark the unified threat coefficient of the feature points of the standard data on the six rays and mark them as reference points. Connect all the reference points end to end in sequence to form the standard data boundary map. Construct the base map based on the spider webs, the outer boundary line, and the standard data boundary map.
[0027] In a preferred embodiment, the steps of marking the threat coefficient of the unknown feature points of the new data obtained based on the simulation test on the spider web line and marking them as actual points, and the actual points and reference points on the same ray being a control group include:
[0028] Screen out the unknown feature point information based on the base database, determine the classification of the unknown feature point information, swap the feature point information of this type on the preset standard data with the unknown feature points and combine them into new data. The threat coefficient of the preset standard data is C1, and the threat coefficient C2 of the combined data is obtained based on the simulation test. Then the threat coefficient of the unknown feature point is: C2 - C1. Transmit the threat coefficient of the unknown feature point to the base database, and the unified threat coefficient of the unknown feature point is ε(C2 - C1);
[0029] Mark the unified threat coefficient of the unknown feature point on the spider web line to obtain the actual point. Each actual point and a reference point are on the same ray. The actual point and the reference point on the same ray are a control group.
[0030] In a preferred embodiment, the steps of constructing a measurement surface and a control surface, comparing the areas of the measurement surface and the control surface, and preliminarily judging the threat of a series of data including unknown feature points and / or unknown feature chains include:
[0031] The unknown feature points of the new data and the known feature points of the existing data are freely combined in number to form unknown feature chains. Connect all the points or an anchor point on the unknown feature chain end to end in sequence to form a measurement surface and measure the area s1. In the control group, connect all the reference points or an anchor point on the reference points end to end in sequence to form a control surface and measure the area s2. If s1 is greater than s2, then determine that the data or a series of data containing the unknown feature chain is undetermined data;
[0032] Based on the comprehensive detection, measure the threat coefficient of the undetermined data and classify the data into safe data or threat data according to the preset security level, and transmit the detection result to the data discrimination library.
[0033] A vulnerability exploitation analysis system based on big data, used to implement a vulnerability exploitation analysis method based on big data, is characterized by including:
[0034] A setting module constructs a main information channel. One end of the main information channel is provided with a separation point. One end of a conventional channel and one end of an identification channel are both connected to the separation point. The separation point is provided with a data discrimination library.
[0035] An acquisition module obtains feature point information based on existing data, obtains a feature point threat coefficient based on a preset standard, constructs a feature chain based on the feature point information, and obtains a feature chain threat coefficient based on the preset standard.
[0036] A sorting module sorts according to the value of the threat coefficient C and determines the parsing order of new data.
[0037] A judgment module constructs a measurement surface based on actual points, anchor points, and point O, constructs a control surface based on reference points, anchor points, and point O, compares the areas of the measurement surface and the control surface, and preliminarily judges the threat of a series of data containing unknown feature points and / or unknown feature chains. Mark one or a series of data with threats as pending data, conduct comprehensive detection on the pending data, and transfer the classification information to the data discrimination library.
[0038] A storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned vulnerability exploitation analysis method based on big data are realized.
[0039] In the above technical solution, the technical effects and advantages provided by the present invention are as follows: BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of the method of the present invention.
[0042] Figure 2 It is a system block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0044] Example 1, please refer toFigure 1 As shown in the figure, a vulnerability exploitation analysis method based on big data according to this embodiment includes the following steps:
[0045] S1. Construct a main information channel. One end of the main information channel is provided with a separation point. One end of a conventional channel and one end of an identification channel are both connected to the separation point. The separation point is provided with a data discrimination library, which is used to classify existing data into safe data and threat data according to a preset security level and complete the screening of the data.
[0046] S2. Obtain feature point information based on existing data, obtain a feature point threat coefficient A based on preset standard data, construct a feature chain based on the feature point information and obtain a feature chain threat coefficient B, and construct a base database based on the feature point information, the feature point threat coefficient, the feature chain, and the feature chain threat coefficient.
[0047] S3. Set a scanning area based on the identification channel, determine the threat coefficient C of new data within a unit time based on a simple test, sort according to the numerical value of the threat coefficient C, and determine the parsing order of the new data.
[0048] S4. Select a point O on the virtual plane, construct spider web lines with O as the endpoint. There are a total of six spider web lines. Mark a point with the largest feature point threat coefficient on each spider web line and connect them end to end in sequence as the outer boundary line. Mark the feature point threat coefficient of the standard data on the spider web line and mark it as the reference point. Connect the reference points end to end in sequence to form a standard data boundary map. The outer boundary line is located outside the standard data boundary map. Construct a base map based on the spider web lines, the outer boundary line, and the standard data boundary map; mark the threat coefficient of the unknown feature points of the new data obtained based on the simulation test on the spider web line and mark it as the actual point. The actual point and the reference point on the same ray form a control group; freely combine the unknown feature points and the existing feature points on the spider web into unknown feature chains. Select a point on the base map as the anchor point, construct a measurement surface based on the actual point, the anchor point, and the point O, construct a control surface based on the reference point, the anchor point, and the point O, compare the areas of the measurement surface and the control surface, and preliminarily judge the threat of a series of data containing unknown feature points and / or unknown feature chains. Mark one or a series of data with threats as pending data, conduct comprehensive detection on the pending data, and transmit the classification information to the data discrimination library.
[0049] As described in the above steps S1 - S4, the big - data - based vulnerability exploitation analysis method is an advanced means that combines big - data technology and vulnerability analysis technology, aiming to mine, analyze, and utilize potential vulnerability information from massive data to enhance the security and defense capabilities of the system. However, in actual use, various data - mining and machine - learning algorithms need to be used to analyze vulnerability data. Different algorithms are suitable for different data types and analysis scenarios, and selecting the appropriate algorithm is a complex process; the large amount of data poses high requirements for the efficiency of the algorithm, and it is difficult to improve the running speed and scalability of the algorithm while ensuring accuracy. In this application, six characteristic - point information of data is locked, namely the sensitivity of data, the scale and precision of data, the accessibility and mobility of data, the timeliness and value of data, the integrity and accuracy of data, and the ownership and compliance of data. These six characteristic points are used as the basic components for measuring the threat of data. Only one measurement standard is adopted, and there is no need to precisely classify the data, but the data can be quickly and preliminarily compared and detected; a spider - web line is constructed with the six characteristic - point information as six axes, and a standard - data boundary map is connected by the threat - coefficient of the characteristic points of the standard data. The unknown characteristic points of the new data are marked on the spider - web line, and unknown characteristic chains are formed by the free combination of the unknown characteristic points and the existing characteristic points on the spider - web line. A measurement surface is constructed based on actual points, anchor points, and O - points, and a comparison surface is constructed based on reference points, anchor points, and O - points. Based on the area of the comparison surface as the judgment basis, the data with an area larger than the comparison surface is tentatively judged as pending data, and comprehensive detection is carried out to determine whether the data is threat data. Through the predetermined judgment mode of the area, data with potential threats can be quickly judged, and the running speed and scalability of the algorithm are improved.
[0050] In one embodiment, the construction of the main information channel and the data discrimination library, where the data discrimination library is used to classify the existing data into secure data and threat data according to a preset security level and complete the screening of the data S1, includes:
[0051] S11. Construct a main information channel between the transmission end and the receiving end. The main information channel is used to transmit data, and the data includes: existing data and new data.
[0052] The transmission end is used to transmit external data into the main information channel, and the receiving end is used to receive the data transmitted from the main information channel.
[0053] S12. Collect the existing data, classify the existing data into secure data and threat data according to the preset standard data, construct a data discrimination library based on the secure data and threat data. The data discrimination library is used to identify secure information and transmit it to the regular channel, the data discrimination library is used to identify new information and transmit it to the authentication channel, and the data identification library is used to identify threat data and intercept the threat data.
[0054] A standard data can be selected, and the threat values of all known data are measured according to TDS, SIEM, APT, Green Alliance's next-generation threat detection all-in-one machine UTS-NDR, and Tencent Cloud's advanced threat detection system (NTA). The data with a threat value exceeding the standard data is marked as threat data, and the data with a threat value less than the standard data is marked as secure data.
[0055] The data identification library uses various means such as data verification, access control, firewalls and intrusion detection systems, data encryption, detection of abnormal activities, logging and auditing, data backup and recovery, and the use of professional database monitoring tools to identify data and intercept potential threats.
[0056] As described in the above steps S11 - S12, external data enters the main information channel through the transmission end. The data in the main information channel enters the separation point, and based on the data identification library set at the separation point, the data is screened. Secure data is transmitted to the regular channel, new data is transmitted to the authentication channel, threat data is intercepted, and the data in the regular channel is transmitted to the receiving end. The main information channel, the regular channel, and the authentication channel form a Y-shaped channel, and the separation point is the common intersection of the main information channel, the regular channel, and the authentication channel. Data is screened, diverted, and intercepted through the separation point. This facilitates obtaining new data for subsequent processing and also facilitates timely updating of the data in the data identification library to improve the identification ability of the data identification library.
[0057] In one embodiment, a base database S2 is constructed based on feature point information, feature point threat coefficients, feature chains, and feature chain threat coefficients, including:
[0058] S21. Obtain the feature point information of existing data. The classification of feature point information includes: data sensitivity, data scale and accuracy, data accessibility and mobility, data timeliness and value, data integrity and accuracy, and data ownership and compliance.
[0059] S22. Obtain the feature point threat coefficients. There are 6 types of feature point threat coefficients, namely the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and mobility, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and data ownership and compliance a6.
[0060] S23. The formula for the feature point threat coefficient A is:
[0061] A = ε1a1 + ε2a2 + ε3a3 + ε4a4 + ε5a5 + ε6a6
[0062] Among them, ε1, ε2, ε3, ε4, ε5, and ε6 are the weight coefficients of a1, a2, a3, a4, a5, and a6 in the threat coefficient of feature points respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data.
[0063] For example, the values of ε1, ε2, ε3, ε4, ε5, and ε6 of a piece of data are 0.3, 0.2, 0.1, 0.2, 0.1, and 0.1 respectively, and the values of a1, a2, a3, a4, a5, and a6 are 2, 3, 1, 1, 4, and 2 respectively.
[0064] Then the unified threat coefficient of the sensitivity of this data is ε1a1 = 0.3 * 2 = 0.6;
[0065] The unified threat coefficient of the scale and precision of this data is ε2a2 = 0.2 * 3 = 0.6;
[0066] The unified threat coefficient of the accessibility and mobility of this data is ε3a3 = 0.1 * 3 = 0.3;
[0067] The unified threat coefficient of the timeliness and value of this data is ε4a4 = 0.2 * 1 = 0.2;
[0068] The unified threat coefficient of the integrity and accuracy of this data is ε5a5 = 0.1 * 4 = 0.4;
[0069] The unified threat coefficient of the ownership and compliance of this data is ε6a6 = 0.1 * 2 = 0.2;
[0070] The threat coefficient A of the feature points of this data = 0.6 + 0.6 + 0.3 + 0.2 + 0.4 + 0.2 = 2.3
[0071] S24. Form a feature chain through at least two pieces of feature point information.
[0072] S25. Construct a feature chain based on the feature point information and obtain the threat coefficient B of the feature chain.
[0073] As described in the above steps S21 - S22, using feature points and feature chains as the criteria for data differentiation, it is necessary to determine the existing feature points and feature chains; using the threat values of feature points and the threat values of feature chains as the criteria for judging data threat, it is necessary to determine the existing threat values of feature points and the existing threat values of feature chains. Therefore, it is necessary to collect the existing feature points, the threat coefficient A of feature points, the existing feature chains, the threat coefficient B of feature chains and construct a base database, so as to distinguish existing feature points from unknown feature points, distinguish existing feature chains from unknown feature chains. When obtaining unknown feature points and threat coefficients, unknown feature chains and threat coefficients, it is also convenient to input the newly obtained data information into the base database and update the base database.
[0074] In one embodiment, a scanning area is set based on an identification channel, the threat coefficient C of new data within a unit time is determined based on a simple test, and the new data is sorted according to the value of the threat coefficient C and the parsing order S3 of the new data is determined, including:
[0075] S31. Set a scanning area within the identification channel, measure the threat coefficient C of new data within a unit time based on VirusTotal. When the number of new data passing through the scanning area within a unit time is not less than 1, compare the threat coefficients of all new data and sort them according to the numerical value, and preferentially parse and process the new data with a larger numerical value.
[0076] Through simple measurement by VirusTotal, the threat value of the data can be quickly measured. Although the numerical value is not precise enough, it can simply screen the threat of the data, quickly narrow down the data range with a high threat degree value, and preferentially detect the data with a high threat coefficient, avoiding detecting all data and reducing the workload of the detection system.
[0077] For example, if the threat coefficients of three data, namely A, B, and C, are 2, 3, and 4 respectively, then the threat coefficient of C > the threat coefficient of B > the threat coefficient of A, and then C, B, and A are parsed and processed in sequence.
[0078] S32. At least one unknown feature point is included in a new data. A first-level unknown feature chain is formed by combining several unknown feature points, and a second-level unknown feature chain is formed by one or more unknown feature points and one or more known feature points. The detection priority of the first-level unknown feature chain is higher than that of the second-level unknown feature chain.
[0079] Preferentially detect the first-level feature chain composed of all unknown feature points to determine the harm of subclass data generated only by unknown feature points.
[0080] As described in the above steps S31 - S32, the data is simply detected by VirusTotal to obtain the threat coefficient of the data. By detecting the data within a unit time (which can be set by oneself, usually using 1s as the basic unit of time) and classifying them according to the threat coefficient, the larger the threat coefficient, the greater the possible threat of the data. Preferentially parse and process the data with a larger threat coefficient to determine the threat of the data and the data based on its feature points and feature chains, and perform marking and prevention in advance.
[0081] In one embodiment, a spider web line, an outer boundary line, and a standard data boundary map are constructed. Reference points and actual points are marked on the spider web line. A measurement plane is constructed based on the actual points, anchor points, and point O. A control plane is constructed based on the reference points, anchor points, and point O. The areas of the measurement plane and the control plane are compared, and the threat level of a series of data containing unknown feature points and / or unknown feature chains is preliminarily judged. One or a series of data with threat level is marked as pending data. Comprehensive detection is performed on the pending data, and classification information is transmitted to the data discrimination library S4, including:
[0082] S41. A center point O is constructed on a plane. Six rays are drawn with O as a common endpoint. The six rays are collectively referred to as the spider web line. The included angle between adjacent rays is 60°. Only the unified threat coefficient of one type of feature point is marked on each ray. The unified threat coefficient of the feature point with the largest existing data is marked on the six rays and connected end to end in sequence to form a hexagon, which is marked as the outer boundary line. The unified threat coefficient of the feature points of the standard data is marked on the six rays and marked as reference points. All reference points are connected end to end in sequence to form the standard data boundary map. A base map is constructed based on the spider web line, the outer boundary line, and the standard data boundary map.
[0083] Since the feature point information is divided into six categories, and each category of feature point information needs to be analyzed and counted, a planar composition is adopted. The spider web line is constructed through lines. Each spider web line represents one category of feature point information, and the number on the spider web line represents the unified threat coefficient of one category of feature point information.
[0084] S42. Unknown feature point information is screened out based on the base database, and the classification of the unknown feature point information is determined. The feature point information of this category on the preset standard data is exchanged with the unknown feature point and combined into new data. The threat coefficient of the preset standard data is C1, and the threat coefficient C2 of the combined data is obtained based on simulation testing. Then, the threat coefficient of the unknown feature point is: C2 - C1. The threat coefficient of the unknown feature point is transmitted to the base database, and the unified threat coefficient of the unknown feature point is ε(C2 - C1).
[0085] S43. The unified threat coefficient of the unknown feature point is marked on the spider web line to obtain the actual point. Each actual point and a reference point are on the same ray. The actual point and the reference point on the same ray are a control group.
[0086] Since there are six categories of data feature point information in total, and the weight coefficients of each category in the feature point threat coefficient are different. If the threat coefficient of each feature point is directly marked on the spider web diagram, then there is no specific length standard for the threat coefficient on each spider web line, and the obtained area does not have unity and comparability. In order to eliminate the influence of the weight coefficient, εa (unified threat coefficient) is used to replace a (one category of threat coefficient).
[0087] S44. The unknown feature points of the new data and the known feature points of the existing data are freely combined into an unknown feature chain. Based on all the points or an anchor point on the unknown feature chain, they are connected end to end in sequence to form a measurement surface and the area s1 is measured. In the control group, all the reference points or an anchor point are connected end to end in sequence to form a control surface and the area s2 is measured. If s1 is greater than s2, it is determined that the data or a series of data containing the unknown feature chain is to-be-determined data.
[0088] S45. Based on the comprehensive detection, measure the threat coefficient of the to-be-determined data and classify the data into safe data or threat data according to the preset security level, and transfer the detection result to the data discrimination library.
[0089] As described in the above steps S41 - S45, construct spider webs, external boundaries, standard data edges, and base maps. Based on the standard data edges, construct a control surface. The unknown feature points and the existing feature points on the spider web are freely combined into an unknown feature chain. Based on the unknown feature points, numerous unknown feature chains are derived, thereby increasing the probability of obtaining threatening data and avoiding blind search, learning, and finding new data. By comparing the area of the measurement surface formed by the unknown feature chain with the area of the standard surface, the threat of the data or a series of data containing the unknown feature chain is judged. Without relying on various data mining and machine learning algorithms to analyze vulnerability data, only an algorithm for constructing and comparing areas is used as a unified algorithm, and different algorithms do not need to be adopted for different data types and analysis scenarios.
[0090] Example 2, please refer to Figure 2 , a vulnerability exploitation analysis system based on big data for implementing a vulnerability exploitation analysis method based on big data, including:
[0091] A setting module, constructing a main information channel. One end of the main information channel is provided with a separation point. One end of the regular channel and one end of the identification channel are both connected to the separation point, and a data discrimination library is set at the separation point;
[0092] A collection module, obtaining feature point information based on the existing data, obtaining the feature point threat coefficient based on the preset standard, constructing a feature chain based on the feature point information, and obtaining the feature chain threat coefficient based on the preset standard;
[0093] A sorting module, sorting according to the numerical value of the threat coefficient C and determining the parsing order of the new data;
[0094] A judgment module, constructing a measurement surface based on the actual points, anchor points, and O points, constructing a control surface based on the reference points, anchor points, and O points, comparing the areas of the measurement surface and the control surface, and preliminarily judging the threat of a series of data containing unknown feature points and / or unknown feature chains. Mark one or a series of data with threats as to-be-determined data, conduct comprehensive detection on the to-be-determined data, and transfer the classification information to the data discrimination library.
[0095] Embodiment 3. A storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned big data-based vulnerability exploitation analysis method are implemented.
[0096] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A vulnerability exploitation analysis method based on big data, characterized in that, Including the following steps: Construct a main information channel. One end of the main information channel is provided with a separation point. One end of the conventional channel and one end of the discrimination channel are both connected to the separation point. The separation point is provided with a data discrimination library, which is used to classify the existing data into secure data and threat data according to the preset security level and complete the screening of the data; Obtain feature point information based on the existing data, obtain the threat coefficient A of the feature point based on the preset standard data, construct a feature chain based on the feature point information and obtain the threat coefficient B of the feature chain, and construct a base database based on the feature point information, the threat coefficient of the feature point, the feature chain, and the threat coefficient of the feature chain; Set a scanning area based on the discrimination channel, determine the threat coefficient C of the new data per unit time based on a simple test, sort according to the numerical value of the threat coefficient C and determine the parsing order of the new data; Select a point O on the virtual plane, construct spider web lines with O as the end point. There are six spider web lines in total. Mark a point with the maximum threat coefficient of the feature point on each spider web line and connect them end to end in sequence as the outer boundary line. Mark the threat coefficient of the feature point of the standard data on the spider web line and mark it as the reference point. Connect the reference points end to end in sequence as the standard data boundary map. The outer boundary line is located outside the standard data boundary map. Construct a base map based on the spider web lines, the outer boundary line, and the standard data boundary map; Mark the threat coefficient of the unknown feature point of the new data obtained based on the simulation test on the spider web line and mark it as the actual point. The actual point and the reference point on the same ray are a control group; Based on the free combination of the unknown feature points and the existing feature points on the spider web into unknown feature chains, select a point on the base map as the anchor point. Construct a measurement plane based on the actual point, the anchor point, and the point O. Construct a comparison plane based on the reference point, the anchor point, and the point O. Compare the areas of the measurement plane and the comparison plane and preliminarily judge the threat of a series of data containing unknown feature points and / or unknown feature chains. Mark one or a series of data with threats as pending data, conduct comprehensive detection on the pending data and transfer the classification information to the data discrimination library.
2. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing the main information channel and constructing the data discrimination library based on the existing data include: Construct a main information channel between the transmitting end and the receiving end. The main information channel is used to transmit data, and the data includes: existing data and new data; Collect the existing data, classify the existing data into secure data and threat data according to the preset standard data, construct a data discrimination library based on the secure data and the threat data. The data discrimination library is used to identify the secure information and transmit it to the conventional channel. The data discrimination library is used to identify the new information and transmit it to the discrimination channel. The data identification library is used to identify the threat data and intercept the threat data.
3. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of obtaining the feature point information based on the existing data, obtaining the threat coefficient A of the feature point based on the preset standard data, and constructing a feature chain based on the feature point information include: Obtain the feature point information of the existing data. The classification of the feature point information includes: the sensitivity of the data, the scale and accuracy of the data, the accessibility and mobility of the data, the timeliness and value of the data, the integrity and accuracy of the data, the ownership and compliance of the data; Obtain the threat coefficients of feature points. There are 6 categories of threat coefficients of feature points, namely, the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and mobility, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and the data ownership and compliance a6; The formula for the threat coefficient A of feature points is: A = ε1a1 + ε2a2 + ε3a3 + ε4a4 + ε5a5 + ε6a6 Where ε1, ε2, ε3, ε4, ε5, and ε6 are the weight coefficients of a1, a2, a3, a4, a5, and a6 in the threat coefficient of feature points respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data; Form a feature chain through at least two pieces of feature point information.
4. A vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of sorting according to the numerical value of the threat coefficient C and determining the parsing order of new data include: Set a scanning area in the identification channel, measure the threat coefficient C of new data per unit time based on VirusTotal. When the number of new data passing through the scanning area per unit time is not less than 1, compare the threat coefficients of all new data and sort them according to the numerical value, and preferentially perform parsing processing on new data with a large numerical value; At least one unknown feature point is included in a new data. An unknown feature chain of the first level is formed by combining several unknown feature points, and an unknown feature chain of the second level is formed by combining one or more unknown feature points with one or more known feature points. The detection priority of the unknown feature chain of the first level is higher than that of the unknown feature chain of the second level.
5. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing a base map based on the spider web line, the outer boundary line, and the standard data boundary map include: Construct a center point O on the plane, draw 6 rays with O as the common end point. The 6 rays are collectively called the spider web line, and the included angle between adjacent rays is 60°. Each ray only marks the unified threat coefficient of a type of feature point. Mark the unified threat coefficient of the largest feature point of the existing data on the 6 rays and connect them in sequence head to tail to form a hexagon, which is marked as the outer boundary line. Mark the unified threat coefficient of the feature points of the standard data on the 6 rays and mark them as reference points. Connect all the reference points in sequence head to tail as the standard data boundary map, and construct a base map based on the spider web line, the outer boundary line, and the standard data boundary map.
6. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of marking the threat coefficient of the unknown feature point of the new data obtained based on the simulation test on the spider web line and marking it as the actual point, and the actual point and the reference point on the same ray being a control group include: Screen out the unknown feature point information based on the base database, determine the classification of the unknown feature point information, swap the information of this type of feature point on the preset standard data with the unknown feature point and combine them into new data. The threat coefficient of the preset standard data is C1, and the threat coefficient of the combined data is obtained based on the simulation test as C2. Then the threat coefficient of the unknown feature point is: C2 - C1. Transmit the threat coefficient of the unknown feature point to the base database, and the unified threat coefficient of the unknown feature point is ε(C2 - C1); Mark the unified threat coefficient of the unknown feature points on the cobweb line to obtain actual points. Each actual point and a reference point are on the same ray, and the actual point and the reference point on the same ray form a control group.
7. A vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing a measurement surface and a control surface, comparing the areas of the measurement surface and the control surface, and preliminarily judging the threat level of a series of data including unknown feature points and / or unknown feature chains include: The unknown feature points of the new data and the known feature points of the existing data are freely combined into unknown feature chains. Based on all the points or an anchor point on the unknown feature chain, they are connected end to end in sequence to form a measurement surface and the area s1 is measured. In the control group, all the reference points or an anchor point are connected end to end in sequence to form a control surface and the area s2 is measured. If s1 is greater than s2, it is determined that the data or a series of data including the unknown feature chain is undetermined data. Based on the comprehensive detection of the threat coefficient of the undetermined data, the data is classified into safe data or threat data according to the preset security level, and the detection result is transmitted to the data discrimination library.
8. A vulnerability exploitation analysis system based on big data, which is used to implement a vulnerability exploitation analysis method based on big data according to any one of claims 1-7, characterized in that, Including: A setting module that constructs a main information channel. One end of the main information channel is provided with a separation point. One end of the conventional channel and one end of the identification channel are both connected to the separation point, and the separation point is provided with a data discrimination library. A collection module that obtains feature point information based on the existing data, obtains the feature point threat coefficient based on the preset standard, constructs a feature chain based on the feature point information, and obtains the feature chain threat coefficient based on the preset standard. A sorting module that sorts according to the numerical value of the threat coefficient C and determines the parsing order of the new data. A judgment module that constructs a measurement surface based on the actual point, the anchor point, and the O point, constructs a control surface based on the reference point, the anchor point, and the O point, compares the areas of the measurement surface and the control surface, and preliminarily judges the threat level of a series of data including unknown feature points and / or unknown feature chains. Mark one or a series of data with threat level as undetermined data, conduct comprehensive detection on the undetermined data, and transmit the classification information to the data discrimination library.
9. A storage medium, wherein the medium stores a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for analyzing vulnerability exploitation based on big data according to any one of claims 1-7 above.
Citation Information
Patent Citations
Network security detection method and system
CN118101250A
Vulnerability mining and repairing method and system based on big data and storage medium
CN118364467A