Vulnerability utilization analysis method and system based on big data and storage medium

By building the main information channel and data identification library in vulnerability utilization analysis, and using feature point information and threat coefficients to make preliminary threat judgments, the problem of insufficient running speed and scalability of algorithms in the existing technology is solved, and fast and accurate data threat screening is achieved.

CN119989365AActive Publication Date: 2025-05-13BEIJING WILLBOX TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510104088.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

In the vulnerability analysis, it is difficult for the prior art to improve the running speed and scalability of the algorithm while ensuring accuracy, especially when dealing with large data volumes.

Method used

By building the main information channel and data identification library, using feature point information and threat coefficients to build a base database, conduct preliminary threat judgments based on spider web lines and standard data boundary maps, and quickly filter out the pending data for comprehensive detection.

Benefits of technology

It realizes the rapid judgment of data threats without relying on complex data mining and machine learning algorithms, and improves the running speed and scalability of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989365A_ABST
    Figure CN119989365A_ABST
Patent Text Reader

Abstract

The invention discloses a big data-based vulnerability utilization analysis method and system and a storage medium, and relates to an information security technology, a main information channel is constructed, and a data discrimination library is used for dividing existing data into security data and threat data according to a preset security level and completing data screening; constructing a base database based on the feature point information, the feature point threat coefficient, the feature chain and the feature chain threat coefficient; sorting according to the numerical value of the threat coefficient C and determining the analysis sequence of the new data; constructing a cobweb line, an external side line, a standard data boundary diagram and a base map; an unknown feature chain is formed based on free number combination of unknown feature points on a cobweb and existing feature points, a measurement surface and a contrast surface are constructed, and the threat of data is preliminarily judged. According to the method, a unified analysis algorithm is adopted, and the operation speed and expandability of the algorithm are improved on the basis of ensuring the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information security technology, and in particular to a vulnerability exploitation analysis method, system and storage medium based on big data. Background Art

[0002] The vulnerability exploitation analysis method based on big data is an advanced method that combines big data technology with vulnerability analysis technology. It aims to mine, analyze and utilize potential vulnerability information from massive data to improve the security and defense capabilities of the system.

[0003] In the known prior art, when vulnerability data is detected in the system, the collected data is preprocessed, including data cleaning, formatting, deduplication, etc.; a vulnerability database is established to store the preprocessed vulnerability data; big data technology is used to conduct in-depth analysis of the vulnerability database to discover the correlation and trend between vulnerabilities, and data mining algorithms such as classification, clustering, and scoring models are used to identify the patterns and characteristics of vulnerabilities; the possibility of vulnerability being exploited is evaluated based on factors such as the type, severity, and scope of impact of the vulnerability; the results of the vulnerability analysis are displayed in a visual manner and a vulnerability analysis report is generated, which lists in detail the discovered vulnerabilities, evaluation results, exploitation attempts, and recommended remedial measures.

[0004] However, various data mining and machine learning algorithms are needed to analyze vulnerability data. Different algorithms are suitable for different data types and analysis scenarios. Selecting a suitable algorithm is a complex process. The large amount of data places high demands on the efficiency of the algorithm, and it is difficult to improve the algorithm's running speed and scalability while ensuring accuracy. Summary of the invention

[0005] The purpose of the present invention is to provide a vulnerability exploitation analysis method, system and storage medium based on big data to address the shortcomings of the background technology.

[0006] In order to achieve the above-mentioned object, the present invention provides the following technical solutions: A vulnerability exploitation analysis method based on big data, comprising the following steps: constructing a main information channel, one end of the main information channel is provided with a separation point, one end of the conventional channel and one end of the identification channel are both connected to the separation point, the separation point is provided with a data identification library, the data identification library is used to classify existing data into security data and threat data according to a preset security level and complete data screening;

[0007] Obtain feature point information based on existing data, obtain feature point threat coefficient A based on preset standard data, construct feature chain based on feature point information and obtain feature chain threat coefficient B, and construct a base database based on feature point information, feature point threat coefficient, feature chain, and feature chain threat coefficient;

[0008] Set the scanning area based on the identification channel, determine the threat coefficient C of new data per unit time based on a simple test, sort and determine the parsing order of new data according to the value of the threat coefficient C;

[0009] Select a point O on the virtual plane, and construct a spider web line with point O as the endpoint. There are six spider web lines in total. Mark a feature point with the largest threat coefficient on each spider web line and connect them in sequence from beginning to end to form an external edge line. Mark the feature point threat coefficient of the standard data on the spider web line and mark it as a reference point. Connect the reference points in sequence from beginning to end to form a standard data boundary map. The external edge line is located outside the standard data boundary map. Construct a base map based on the spider web line, the external edge line, and the standard data boundary map.

[0010] The threat coefficients of unknown feature points of new data obtained based on simulation tests are marked on the spider web line and marked as actual points. The actual points and reference points on the same ray form a control group;

[0011] Based on the unknown feature points on the spider web and the free number of existing feature points, an unknown feature chain is formed. A point is selected on the base map as an anchor point. A measurement surface is constructed based on the actual point, anchor point, and point O. A control surface is constructed based on the reference point, anchor point, and point O. The areas of the measurement surface and the control surface are compared, and a preliminarily judged threat of a series of data containing unknown feature points and / or unknown feature chains. One or a series of data containing threats are marked as pending data. The pending data are comprehensively tested and the classification information is transmitted to the data identification library.

[0012] In a preferred embodiment, the steps of constructing a main information channel and constructing a data identification library based on existing data include:

[0013] A main information channel is constructed between the transmission end and the reception end. The main information channel is used to transmit data, including existing data and new data.

[0014] Collect existing data, divide the existing data into security data and threat data according to preset standard data, and build a data identification library based on security data and threat data. The data identification library is used to identify security information and transmit it to regular channels. The data identification library is used to identify new information and transmit it to authentication channels. The data recognition library is used to identify threat data and intercept threat data.

[0015] In a preferred embodiment, the steps of obtaining feature point information based on existing data, obtaining feature point threat coefficient A based on preset standard data, and constructing a feature chain based on the feature point information include:

[0016] Obtain feature point information of existing data. The feature point information classification includes: data sensitivity, data scale and accuracy, data accessibility and liquidity, data timeliness and value, data integrity and accuracy, data ownership and compliance;

[0017] Obtain the threat coefficient of feature points. There are 6 types of threat coefficients of feature points, namely, the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and liquidity, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and the threat coefficient a6 of data ownership and compliance.

[0018] The formula of the feature point threat coefficient A is:

[0019] A=ε1a1+ε2a2+ε3a3+ε4a4+ε5a5+ε6a6

[0020] Among them, ε1, ε2, ε3, ε4, ε5, and ε6 are the weight coefficients of a1, a2, a3, a4, a5, and a6 in the threat coefficient of the feature point respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data;

[0021] A feature chain is formed by at least two feature point information.

[0022] In a preferred embodiment, the step of sorting and determining the parsing order of new data according to the numerical value of the threat coefficient C includes:

[0023] A scanning area is set in the identification channel, and the threat coefficient C of new data per unit time is measured based on VirusTotal. When there is no less than one new data passing through the scanning area per unit time, the threat coefficients of all new data are compared and sorted by value, and the new data with large values ​​are analyzed and processed first;

[0024] A new data includes at least one unknown feature point. Several unknown feature points are combined into a first-level unknown feature chain. One or more unknown feature points and one or more known feature points form a second-level unknown feature chain. The detection priority of the first-level unknown feature chain is greater than the detection priority of the second-level unknown feature chain.

[0025] In a preferred embodiment, the step of constructing a base map based on spider web lines, external edges, and standard data boundary maps includes:

[0026] Construct a center point O on the plane, and draw out 6 rays with O as the common endpoint. The 6 rays are collectively called spider web lines. The angle between adjacent rays is 60°. Each ray only marks the unified threat coefficient of one type of feature point. The unified threat coefficient of the largest feature point in the existing data is marked on the 6 rays and connected end to end in sequence to form a hexagon. The hexagon is marked as the external edge line. The unified threat coefficient of the feature point of the standard data is marked on the 6 rays and marked as a reference point. All reference points are connected end to end in sequence to form a standard data boundary map. The base map is constructed based on the spider web lines, external edge lines, and standard data boundary map.

[0027] In a preferred embodiment, the threat coefficient of the unknown feature point of the new data obtained based on the simulation test is marked on the spider line and marked as the actual point, and the actual point and the reference point on the same ray form a control group, including:

[0028] Based on the base database, the unknown feature point information is screened out, the classification of the unknown feature point information is determined, the feature point information of this type on the preset standard data is exchanged with the unknown feature points and combined into new data. The threat coefficient of the preset standard data is C1, and the threat coefficient of the combined data is C2 based on the simulation test. Then the threat coefficient of the unknown feature point is: C2-C1, and the threat coefficient of the unknown feature point is transferred to the base database. The unified threat coefficient of the unknown feature point is ε(C2-C1);

[0029] The unified threat coefficient of the unknown feature point is marked on the spider web line to obtain the actual point. Each actual point is located on the same ray as a reference point, and the actual points and reference points on the same ray form a control group.

[0030] In a preferred embodiment, the steps of constructing a measurement surface and a comparison surface, comparing the areas of the measurement surface and the comparison surface, and preliminarily judging the threat of a series of data containing unknown feature points and / or unknown feature chains include:

[0031] The unknown feature points of the new data and the known feature points of the existing data are combined into an unknown feature chain. All points on the unknown feature chain or an anchor point are connected end to end in sequence to form a measurement surface and the area s1 is measured. In the control group, all reference points or an anchor point are connected end to end to form a control surface and the area s2 is measured. If s1 is greater than s2, the data or a series of data containing the unknown feature chain is determined to be pending data;

[0032] Based on comprehensive detection, the threat coefficient of the pending data is measured and the data is divided into safe data or threat data according to the preset security level, and the detection results are transmitted to the data identification library.

[0033] A vulnerability exploitation analysis system based on big data, used to implement a vulnerability exploitation analysis method based on big data, characterized by comprising:

[0034] Setting module, constructing main information channel, one end of main information channel is provided with separation point, one end of conventional channel and one end of identification channel are both connected with separation point, and separation point is provided with data identification library;

[0035] The acquisition module obtains feature point information based on existing data, obtains feature point threat coefficients based on preset standards, constructs feature chains based on feature point information, and obtains feature chain threat coefficients based on preset standards;

[0036] The sorting module sorts the threat coefficient C according to its numerical value and determines the parsing order of the new data;

[0037] The judgment module constructs a measurement surface based on the actual point, anchor point, and O point, and constructs a control surface based on the reference point, anchor point, and O point. It compares the areas of the measurement surface and the control surface and preliminarily judges the threat of a series of data containing unknown feature points and / or unknown feature chains, marks one or a series of data containing threats as pending data, conducts comprehensive detection on the pending data, and transmits the classification information to the data identification library.

[0038] A storage medium stores a computer program, which, when executed by a processor, implements the steps of the vulnerability exploitation analysis method based on big data as described above.

[0039] In the above technical solution, the technical effects and advantages provided by the present invention are: BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0041] Figure 1 The present invention is a flow chart of the method.

[0042] Figure 2 It is a system block diagram of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described examples are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0044] Example 1, please refer to Figure 1 As shown, the vulnerability exploitation analysis method based on big data described in this embodiment includes the following steps:

[0045] S1. Construct a main information channel. A separation point is set at one end of the main information channel. One end of the conventional channel and one end of the identification channel are both connected to the separation point. A data identification library is set at the separation point. The data identification library is used to divide the existing data into security data and threat data according to the preset security level and complete the data screening.

[0046] S2. Obtain feature point information based on existing data, obtain feature point threat coefficient A based on preset standard data, construct a feature chain based on the feature point information and obtain feature chain threat coefficient B, and construct a base database based on feature point information, feature point threat coefficient, feature chain, and feature chain threat coefficient.

[0047] S3. Set a scanning area based on the identification channel, determine the threat coefficient C of new data per unit time based on a simple test, sort according to the value of the threat coefficient C and decide the parsing order of the new data.

[0048] S4. Select a point O on the virtual plane, and construct spider web lines with point O as the endpoint. There are six spider web lines in total. Mark a feature point with the largest threat coefficient on each spider web line and connect them in sequence from beginning to end to form an external edge line. Mark the feature point threat coefficient of the standard data on the spider web line and mark it as a reference point. Connect the reference points in sequence from beginning to end to form a standard data boundary map. The external edge line is located outside the standard data boundary map. Construct a base map based on the spider web lines, external edge lines, and the standard data boundary map. The threat coefficient of the unknown feature point of the new data obtained based on the simulation test is marked on the spider web line and marked as The actual point and the reference point on the same ray are a control group; based on the unknown feature points on the spider web and the free number of existing feature points, an unknown feature chain is formed, and a point is selected on the base map as an anchor point. A measurement surface is constructed based on the actual point, anchor point, and point O. A control surface is constructed based on the reference point, anchor point, and point O. The areas of the measurement surface and the control surface are compared, and a preliminary judgment is made on the threat of a series of data containing unknown feature points and / or unknown feature chains. One or a series of data containing threats are marked as pending data, and the pending data are comprehensively tested and the classification information is passed to the data identification library.

[0049] As described in the above steps S1-S4, the vulnerability exploitation analysis method based on big data is an advanced method that combines big data technology with vulnerability analysis technology, and aims to mine, analyze and utilize potential vulnerability information from massive data to improve the security and defense capabilities of the system. However, in actual use, it is necessary to use various data mining and machine learning algorithms to analyze vulnerability data. Different algorithms are suitable for different data types and analysis scenarios. Selecting a suitable algorithm is a complex process; the amount of big data places high demands on the efficiency of the algorithm, and it is difficult to improve the algorithm's running speed and scalability while ensuring accuracy. The present application locks the six characteristic point information of the data, namely, the sensitivity of the data, the scale and accuracy of the data, the accessibility and liquidity of the data, the timeliness and value of the data, the integrity and accuracy of the data, and the ownership and compliance of the data. These six characteristic points are used as the basic components to measure the threat of the data. Only one measurement standard is used, and there is no need to accurately classify the data, but to quickly conduct preliminary comparison and detection of the data; the six characteristic point information is used as six axes to construct spider web lines, and the characteristic point threat coefficients of the standard data are connected to form a standard data boundary map, and the unknown characteristic points of the new data are marked on the spider web lines. The unknown characteristic points on the spider web are combined with the free number of existing characteristic points to form an unknown characteristic chain, and the measurement surface is constructed based on the actual point, anchor point, and O point. The control surface is constructed based on the reference point, anchor point, and O point. The area of ​​the control surface is used as the judgment basis. The data larger than the area of ​​the control surface is to be judged as pending data, and a comprehensive test is performed to determine whether the data is threatening data. Through the predetermined judgment mode of the area, data with potential threats can be quickly judged, and the running speed and scalability of the algorithm can be improved.

[0050] In one embodiment, the main information channel and the data identification library are constructed, and the data identification library is used to classify the existing data into safe data and threat data according to the preset security level and complete the data screening S1, including:

[0051] S11. A main information channel is constructed between the transmission end and the reception end. The main information channel is used to transmit data, including existing data and new data.

[0052] The transmitting end is used to transmit external data to the main information channel, and the receiving end is used to receive data from the main information channel.

[0053] S12. Collect existing data, divide the existing data into security data and threat data according to preset standard data, and build a data identification library based on the security data and threat data. The data identification library is used to identify security information and transmit it to the regular channel. The data identification library is used to identify new information and transmit it to the authentication channel. The data identification library is used to identify threat data and intercept the threat data.

[0054] You can select a standard data and measure the threat value of all known data based on TDS, SIEM, APT, NSFOCUS Next-Generation Threat Detection Appliance UTS-NDR, and Tencent Cloud Advanced Threat Detection System (NTA). Data that exceeds the threat value of the standard data is marked as threat data, and data that is less than the threat value of the standard data is marked as safe data.

[0055] The data identification library identifies data and intercepts potential threats through a variety of means such as data verification, access control, firewalls and intrusion detection systems, data encryption, abnormal activity detection, logging and auditing, data backup and recovery, and the use of professional database monitoring tools.

[0056] As described in the above steps S11-S12, the external data enters the main information channel through the transmission end, and the data in the main information channel enters the separation point. The data is screened based on the data identification library set at the separation point, and the safe data is transmitted to the conventional channel, the new data is transmitted to the identification channel, the threat data is intercepted, and the data of the conventional channel is transmitted to the receiving end. The main information channel, the conventional channel, and the identification channel form a Y-shaped channel. The separation point is the common intersection of the main information channel, the conventional channel, and the identification channel. The data is screened, diverted, and intercepted through the separation point. It is convenient to obtain new data for subsequent processing, and it is also convenient to update the data of the data identification library in a timely manner, and improve the recognition ability of the data identification library.

[0057] In one embodiment, a base database S2 is constructed based on feature point information, feature point threat coefficients, feature chains, and feature chain threat coefficients, including:

[0058] S21. Obtain feature point information of existing data. The classification of feature point information includes: data sensitivity, data scale and accuracy, data accessibility and liquidity, data timeliness and value, data integrity and accuracy, and data ownership and compliance.

[0059] S22. Obtain the threat coefficient of feature points. There are 6 types of threat coefficients of feature points, namely, the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and liquidity, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and the threat coefficient a6 of data ownership and compliance.

[0060] S23. The formula of the feature point threat coefficient A is:

[0061] A=ε1a1+ε2a2+ε3a3+ε4a4+ε5a5+ε6a6

[0062] Among them, ε1, ε2, ε3, ε4, ε5, and ε6 are the weight coefficients of a1, a2, a3, a4, a5, and a6 in the threat coefficient of the feature points respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data.

[0063] For example, the values ​​of ε1, ε2, ε3, ε4, ε5, and ε6 of a data are 0.3, 0.2, 0.1, 0.2, 0.1, and 0.1 respectively, and the values ​​of a1, a2, a3, a4, a5, and a6 are 2, 3, 1, 1, 4, and 2 respectively.

[0064] Then the unified threat coefficient of the sensitivity of the data is ε1a1=0.3*2=0.6;

[0065] The unified threat factor of the scale and accuracy of this data is ε2a2=0.2*3=0.6;

[0066] The unified threat factor of the accessibility and liquidity of this data is ε3a3=0.1*3=0.3;

[0067] The unified threat coefficient of the timeliness and value of the data is ε4a4=0.2*1=0.2;

[0068] The unified threat factor for the integrity and accuracy of the data is ε5a5=0.1*4=0.4;

[0069] The unified threat coefficient of the ownership and compliance of the data is ε6a6=0.1*2=0.2;

[0070] The characteristic point threat coefficient of the data is A = 0.6 + 0.6 + 0.3 + 0.2 + 0.4 + 0.2 = 2.3

[0071] S24. Form a feature chain through at least two feature point information.

[0072] S25. Construct a feature chain based on the feature point information and obtain a feature chain threat coefficient B.

[0073] As described in the above steps S21-S22, using feature points and feature chains as the standard for data differentiation, it is necessary to determine the existing feature points and feature chains; using feature point threat values ​​and feature chain threat values ​​as the judgment standard for data threat, it is necessary to determine the existing feature point threat values ​​and existing feature chain threat values. Therefore, it is necessary to collect existing feature points, feature point threat coefficient A, existing feature chains, feature chain threat coefficient B and construct a base database to distinguish existing feature points from unknown feature points, and distinguish existing feature chains from unknown feature chains. When unknown feature points and threat coefficients, unknown feature chains and threat coefficients are obtained, it is also convenient to input the newly obtained data information into the base database and update the base database.

[0074] In one embodiment, a scanning area is set based on the identification channel, a threat coefficient C of new data per unit time is determined based on a simple test, and the new data is sorted and the parsing order S3 is determined according to the value of the threat coefficient C, including:

[0075] S31. A scanning area is set in the identification channel, and the threat coefficient C of new data per unit time is measured based on VirusTotal. When there is no less than one new data passing through the scanning area per unit time, the threat coefficients of all new data are compared and sorted by numerical value, and the new data with large numerical value are analyzed and processed first.

[0076] Through simple measurement by VirusTotal, the threat value of data can be quickly measured. Although the value is not accurate enough, it can be used to simply screen the threat of data, quickly narrow the range of data with high threat value, and give priority to detecting data with high threat coefficient, avoiding detecting all data and reducing the workload of the detection system.

[0077] For example, the threat coefficients of three data A, B and C are 2, 3 and 4 respectively, then the threat coefficient of C > the threat coefficient of B > the threat coefficient of A, then C, B and A are analyzed and processed in order.

[0078] S32. A new data includes at least one unknown feature point. A first-level unknown feature chain is formed by combining several unknown feature points. A second-level unknown feature chain is formed by one or more unknown feature points and one or more known feature points. The detection priority of the first-level unknown feature chain is greater than the detection priority of the second-level unknown feature chain.

[0079] Prioritize the detection of the first-level feature chain composed of all unknown feature points to determine the hazards of sub-category data generated only by unknown feature points.

[0080] As described in the above steps S31-S32, a simple detection is performed on the data through VirusTotal to obtain the threat coefficient of the data. The data is detected within a unit time (which can be set by yourself, usually 1s is used as the basic unit of time) and classified according to the threat coefficient. The larger the threat coefficient, the greater the threat of the data may be. The data with the larger threat coefficient is analyzed and processed first to determine the threat of the data and the data based on its feature points and feature chains, and mark and prevent them in advance.

[0081] In one embodiment, spider web lines, external edges, and standard data boundary maps are constructed, reference points and actual points are marked on the spider web lines, a measurement surface is constructed based on the actual point, anchor point, and point O, and a control surface is constructed based on the reference point, anchor point, and point O, the areas of the measurement surface and the control surface are compared, and the threat of a series of data containing unknown feature points and / or unknown feature chains is preliminarily determined, one or a series of data containing threat is marked as pending data, and the pending data is comprehensively detected and the classification information is transmitted to the data identification library S4, including:

[0082] S41. Construct a center point O on the plane, and draw out 6 rays with O as the common endpoint. The 6 rays are collectively called spider web lines. The angle between adjacent rays is 60°. Each ray only marks the unified threat coefficient of one type of feature point. Mark the unified threat coefficient of the largest feature point in the existing data on the 6 rays and connect them end to end in sequence to form a hexagon. The hexagon is marked as the external edge line. Mark the unified threat coefficient of the feature point of the standard data on the 6 rays and mark them as reference points. Connect all reference points end to end in sequence to form a standard data boundary map. Construct a base map based on the spider web lines, external edge lines, and standard data boundary map.

[0083] Since the feature point information is divided into 6 categories, each category of feature point information needs to be analyzed and counted. Therefore, a plane composition is used to construct spider web lines through lines. Each spider web line represents a feature point information, and the number on the spider web line represents the unified threat coefficient of a feature point information.

[0084] S42. Filter out unknown feature point information based on the base database, determine the classification of the unknown feature point information, interchange the feature point information of this type on the preset standard data with the unknown feature points and combine them into new data. The threat coefficient of the preset standard data is C1. The threat coefficient of the combined data is C2 based on the simulation test. Then, the threat coefficient of the unknown feature point is: C2-C1. The threat coefficient of the unknown feature point is transferred to the base database. The unified threat coefficient of the unknown feature point is ε(C2-C1).

[0085] S43. Mark the unified threat coefficient of the unknown feature point on the spider web line to obtain the actual point. Each actual point is located on the same ray as a reference point. The actual points and the reference points on the same ray form a control group.

[0086] Since there are six categories of data feature point information, the weight coefficient of each category in the feature point threat coefficient is different. If the threat coefficient of each feature point is directly marked on the spider web diagram, then the threat coefficient on each spider web line does not have a specific length standard, and the obtained area is not uniform and contrasting. In order to eliminate the influence of the weight coefficient, εa (unified threat coefficient) is used instead of a (first-class threat coefficient).

[0087] S44. The unknown feature points of the new data and the known feature points of the existing data are freely combined into an unknown feature chain. All points on the unknown feature chain or one anchor point are connected end to end in sequence to form a measurement surface and the area s1 is measured. In the control group, all reference points or one anchor point are connected end to end in sequence to form a control surface and the area s2 is measured. If s1 is greater than s2, the data or a series of data containing the unknown feature chain is determined to be pending data.

[0088] S45. Based on the comprehensive detection, measure the threat coefficient of the pending data and classify the data into safe data or threat data according to the preset security level, and transmit the detection result to the data identification library.

[0089] As described in the above steps S41-S45, spider web lines, external edges, standard data edges, and base maps are constructed, and a control surface is constructed based on the standard data edge. The unknown feature points on the spider web are combined with the existing feature points in a free number to form an unknown feature chain, and a large number of unknown feature chains are derived based on the unknown feature points, thereby increasing the probability of obtaining threatening data and avoiding blind search and learning, and finding new data. The area of ​​the measured surface formed by the unknown feature chain is compared with the area of ​​the standard surface to determine the threat of data or a series of data containing the unknown feature chain. There is no need to use various data mining and machine learning algorithms to analyze vulnerability data. Only an algorithm for constructing and comparing areas is needed as a unified algorithm, and there is no need to use different algorithms for different data types and analysis scenarios.

[0090] Example 2, please refer to Figure 2 , a vulnerability exploitation analysis system based on big data, used to implement a vulnerability exploitation analysis method based on big data, including:

[0091] Setting module, constructing main information channel, one end of main information channel is provided with separation point, one end of conventional channel and one end of identification channel are both connected with separation point, and separation point is provided with data identification library;

[0092] The acquisition module obtains feature point information based on existing data, obtains feature point threat coefficients based on preset standards, constructs feature chains based on feature point information, and obtains feature chain threat coefficients based on preset standards;

[0093] The sorting module sorts the threat coefficient C according to its numerical value and determines the parsing order of the new data;

[0094] The judgment module constructs a measurement surface based on the actual point, anchor point, and O point, and constructs a control surface based on the reference point, anchor point, and O point. It compares the areas of the measurement surface and the control surface and preliminarily judges the threat of a series of data containing unknown feature points and / or unknown feature chains, marks one or a series of data containing threats as pending data, conducts comprehensive detection on the pending data, and transmits the classification information to the data identification library.

[0095] Embodiment 3, a storage medium, wherein the medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned vulnerability exploitation analysis method based on big data are implemented.

[0096] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A vulnerability exploitation analysis method based on big data, characterized in that: The following steps are involved: Construct a main information channel, one end of which is provided with a separation point, one end of the conventional channel and one end of the identification channel are both connected to the separation point, and the separation point is provided with a data identification library, which is used to classify existing data into safe data and threat data according to preset security levels and complete data screening; Obtain feature point information based on existing data, obtain feature point threat coefficient A based on preset standard data, construct feature chain based on feature point information and obtain feature chain threat coefficient B, and construct a base database based on feature point information, feature point threat coefficient, feature chain, and feature chain threat coefficient; Set the scanning area based on the identification channel, determine the threat coefficient C of new data per unit time based on a simple test, sort and determine the parsing order of new data according to the value of the threat coefficient C; Select a point O on the virtual plane, and construct a spider web line with point O as the endpoint. There are six spider web lines in total. Mark a feature point with the largest threat coefficient on each spider web line and connect them in sequence from beginning to end to form an external edge line. Mark the feature point threat coefficient of the standard data on the spider web line and mark it as a reference point. Connect the reference points in sequence from beginning to end to form a standard data boundary map. The external edge line is located outside the standard data boundary map. Construct a base map based on the spider web line, the external edge line, and the standard data boundary map. The threat coefficients of unknown feature points of new data obtained based on simulation tests are marked on the spider web line and marked as actual points. The actual points and reference points on the same ray form a control group; Based on the unknown feature points on the spider web and the free number of existing feature points, an unknown feature chain is formed. A point is selected on the base map as an anchor point. A measurement surface is constructed based on the actual point, anchor point, and point O. A control surface is constructed based on the reference point, anchor point, and point O. The areas of the measurement surface and the control surface are compared, and a preliminarily judged threat of a series of data containing unknown feature points and / or unknown feature chains. One or a series of data containing threats are marked as pending data. The pending data are comprehensively tested and the classification information is transmitted to the data identification library.

2. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing the main information channel and building a data identification library based on existing data include: A main information channel is constructed between the transmission end and the reception end. The main information channel is used to transmit data, including existing data and new data. Collect existing data, divide the existing data into security data and threat data according to preset standard data, and build a data identification library based on security data and threat data. The data identification library is used to identify security information and transmit it to regular channels. The data identification library is used to identify new information and transmit it to authentication channels. The data recognition library is used to identify threat data and intercept threat data.

3. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of obtaining feature point information based on existing data, obtaining feature point threat coefficient A based on preset standard data, and constructing a feature chain based on the feature point information include: Obtain feature point information of existing data. The feature point information classification includes: data sensitivity, data scale and accuracy, data accessibility and liquidity, data timeliness and value, data integrity and accuracy, data ownership and compliance; Obtain the threat coefficient of feature points. There are 6 types of threat coefficients of feature points, namely, the threat coefficient a1 of data sensitivity, the threat coefficient a2 of data scale and accuracy, the threat coefficient a3 of data accessibility and liquidity, the threat coefficient a4 of data timeliness and value, the threat coefficient a5 of data integrity and accuracy, and the threat coefficient a6 of data ownership and compliance. The formula of the feature point threat coefficient A is: A=ε1a1+ε2a2+ε3a3+ε4a4+ε5a5+ε6a6 Among them, ε1, ε2, ε3, ε4, ε5, and ε6 are the weight coefficients of a1, a2, a3, a4, a5, and a6 in the threat coefficient of the feature point respectively, and εa is the unified threat coefficient of a type of feature point information in the standard data; A feature chain is formed by at least two feature point information.

4. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of sorting the threat coefficient C and determining the parsing order of the new data include: A scanning area is set in the identification channel, and the threat coefficient C of new data per unit time is measured based on VirusTotal. When there is no less than one new data passing through the scanning area per unit time, the threat coefficients of all new data are compared and sorted by value, and the new data with large values ​​are analyzed and processed first; A new data includes at least one unknown feature point. Several unknown feature points are combined into a first-level unknown feature chain. One or more unknown feature points and one or more known feature points form a second-level unknown feature chain. The detection priority of the first-level unknown feature chain is greater than the detection priority of the second-level unknown feature chain.

5. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing a base map based on spider web lines, external edges, and standard data boundary maps include: Construct a center point O on the plane, and draw out 6 rays with O as the common endpoint. The 6 rays are collectively called spider web lines. The angle between adjacent rays is 60°. Each ray only marks the unified threat coefficient of one type of feature point. The unified threat coefficient of the largest feature point in the existing data is marked on the 6 rays and connected end to end in sequence to form a hexagon. The hexagon is marked as the external edge line. The unified threat coefficient of the feature point of the standard data is marked on the 6 rays and marked as a reference point. All reference points are connected end to end in sequence to form a standard data boundary map. The base map is constructed based on the spider web lines, external edge lines, and standard data boundary map.

6. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of obtaining the threat coefficient of the unknown feature point of the new data based on the simulation test and marking it on the spider web line and marking it as the actual point, and the actual point and the reference point on the same ray as a control group include: Based on the base database, the unknown feature point information is screened out, the classification of the unknown feature point information is determined, the feature point information of this type on the preset standard data is exchanged with the unknown feature points and combined into new data. The threat coefficient of the preset standard data is C1, and the threat coefficient of the combined data is C2 based on the simulation test. Then the threat coefficient of the unknown feature point is: C2-C1, and the threat coefficient of the unknown feature point is transferred to the base database. The unified threat coefficient of the unknown feature point is ε(C2-C1); The unified threat coefficient of the unknown feature point is marked on the spider web line to obtain the actual point. Each actual point is located on the same ray as a reference point, and the actual points and reference points on the same ray form a control group.

7. The vulnerability exploitation analysis method based on big data according to claim 1, characterized in that: The steps of constructing a measurement surface and a comparison surface, comparing the areas of the measurement surface and the comparison surface, and preliminarily judging the threat of a series of data containing unknown feature points and / or unknown feature chains include: The unknown feature points of the new data and the known feature points of the existing data are combined into an unknown feature chain. All points on the unknown feature chain or an anchor point are connected end to end in sequence to form a measurement surface and the area s1 is measured. In the control group, all reference points or an anchor point are connected end to end to form a control surface and the area s2 is measured. If s1 is greater than s2, the data or a series of data containing the unknown feature chain is determined to be pending data; Based on comprehensive detection, the threat coefficient of the pending data is measured and the data is divided into safe data or threat data according to the preset security level, and the detection results are transmitted to the data identification library.

8. A vulnerability exploitation analysis system based on big data, used to implement a vulnerability exploitation analysis method based on big data as described in any one of claims 1 to 7, characterized in that: include: Setting module, constructing main information channel, one end of main information channel is provided with separation point, one end of conventional channel and one end of identification channel are both connected with separation point, and separation point is provided with data identification library; The acquisition module obtains feature point information based on existing data, obtains feature point threat coefficients based on preset standards, constructs feature chains based on feature point information, and obtains feature chain threat coefficients based on preset standards; The sorting module sorts the threat coefficient C according to its numerical value and determines the parsing order of the new data; The judgment module constructs a measurement surface based on the actual point, anchor point, and O point, and constructs a control surface based on the reference point, anchor point, and O point. It compares the areas of the measurement surface and the control surface and preliminarily judges the threat of a series of data containing unknown feature points and / or unknown feature chains, marks one or a series of data containing threats as pending data, conducts comprehensive detection on the pending data, and transmits the classification information to the data identification library.

9. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the vulnerability exploitation analysis method based on big data as described in any one of claims 1 to 7 above are implemented.

Citation Information

Patent Citations

  • Network security detection method and system

    CN118101250A

  • Vulnerability mining and repairing method and system based on big data and storage medium

    CN118364467A

  • Systems and methods for threat analysis of computer data

    WO2016061546A1