Security protection method and system based on big data analysis and cloud computing

By building a security situation awareness model on the cloud computing platform and processing and analyzing multi-source heterogeneous data, the problems of inefficiency and insufficient accuracy of traditional security protection methods when processing large-scale data are solved, and more efficient and accurate network security monitoring and protection are achieved.

CN120090856APending Publication Date: 2025-06-03NAVAL UNIV OF ENG PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510303385.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional security protection methods are difficult to deal with complex and changeable cyber attack methods, and are inefficient and insufficient in the processing of large-scale, multi-source heterogeneous data.

Method used

Using security protection methods based on big data analysis and cloud computing, a preset cloud computing platform collects and cleanses multi-source heterogeneous data in real time, builds a security situation awareness model, extracts behavioral feature vectors for risk scores, and uses machine learning algorithms to output dynamic risk estimates, and selects appropriate protection modes based on the degree of risk.

Benefits of technology

Real-time monitoring and accurate judgment of network security is achieved, real-time and accuracy of security protection is improved, and targeted defense measures can be taken according to different security threats and attack scenarios to ensure the safe and stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090856A_ABST
    Figure CN120090856A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of security protection, and relates to a security protection method and system based on big data analysis and cloud computing, and the method comprises the following steps: collecting multi-source heterogeneous data in real time, and carrying out cleaning and standardization processing on the multi-source heterogeneous data; creating a cloud native data warehouse, encrypting the multi-source heterogeneous data and loading the cloud native data warehouse; constructing a security situation awareness model, extracting key behavior feature vectors from the preprocessed multi-source heterogeneous data, inputting the behavior feature vectors into the security situation awareness model, and performing risk scoring on each behavior feature vector; outputting a dynamic risk estimation value based on a machine learning algorithm; and setting a plurality of risk pre-estimation threshold values, comparing the dynamic risk pre-estimation value with the plurality of risk pre-estimation threshold values, and selecting a proper protection mode. The method can solve the problems of low efficiency and insufficient accuracy when a traditional method is used for processing large-scale and multi-source heterogeneous data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of security protection, and relates to a security protection method and system based on big data analysis and cloud computing. Background Art

[0002] With the rapid development of information technology, network security issues have become increasingly prominent and have become an important factor restricting the informatization process. Traditional security protection methods often rely on static rule matching and signature detection, and it is difficult to cope with increasingly complex and changeable network attack means. At the same time, these methods have problems such as low efficiency and insufficient accuracy when dealing with large-scale, multi-source heterogeneous data, and cannot meet the urgent needs of current network security protection.

[0003] With the rise of cloud computing and big data technologies, new ideas and means have been provided for network security protection. Cloud computing platforms, with their powerful computing capabilities and elastic scalability, can efficiently process and analyze large-scale data. They can mine potential associations and rules in the data and provide more accurate decision-making support for security protection.

[0004] Based on the above problems, traditional methods have problems of low efficiency and insufficient accuracy when dealing with large-scale, multi-source heterogeneous data. Summary of the Invention

[0005] In order to solve the above problems, the present invention provides a security protection method and system based on big data analysis and cloud computing.

[0006] In the first aspect, the present invention provides a security protection method based on big data analysis and cloud computing, adopting the following technical solutions:

[0007] A security protection method based on big data analysis and cloud computing includes the following steps:

[0008] S1. Combine a preset cloud computing platform to collect multi-source heterogeneous data in real time, and perform cleaning and standardization processing on the multi-source heterogeneous data;

[0009] S2. Based on the preset cloud computing platform, create a cloud-native data warehouse, encrypt the preprocessed multi-source heterogeneous data and load it into the cloud-native data warehouse;

[0010] S3. Build a security situation awareness model, extract key behavior feature vectors from the preprocessed multi-source heterogeneous data, input the behavior feature vectors into the security situation awareness model, the security situation awareness model generates corresponding scoring criteria for each behavior feature vector, and perform risk scoring on each behavior feature vector;

[0011] S4. Combine the risk score values of the key behavior feature vectors and output a dynamic risk prediction value based on a machine learning algorithm;

[0012] S5. Set several risk prediction thresholds, compare the dynamic risk prediction value with the several risk prediction thresholds, and select an appropriate protection mode.

[0013] In a further solution of the present invention, step S1 includes the following steps:

[0014] The server to be protected is deployed with a lightweight collection plugin or a lightweight agent program;

[0015] In combination with a preset cloud computing platform, multi-source heterogeneous data is collected in real time, and the multi-source heterogeneous data is cleaned and standardized. The multi-source heterogeneous data includes network data, behavior data, and system data;

[0016] For network data, collect the time of network connection, IP address, and the amount of data transmitted;

[0017] For behavior data, collect user login logs and operation records;

[0018] For system data, collect the CPU instantaneous peak value and memory occupancy alarm.

[0019] In a further solution of the present invention, step S2 includes the following steps:

[0020] Perform encrypted hashing on the fields containing personal identity information to convert the fields containing personal identity information into irreversible hash values. The preprocessed multi-source heterogeneous data forms a complete data packet;

[0021] The server calls the TPM2.0 security chip to generate an RSA-3072 key pair, an RSA public key and an RSA private key. The server to be protected uses the RSA private key to sign the data packet, and the RSA public key is pre-stored in the preset cloud computing platform;

[0022] The preset cloud computing platform creates a cloud-native data warehouse, loads the data packet into the cloud-native data warehouse, and the preset cloud computing platform uses the RSA public key to verify the signature. If the signature does not match, the data packet will be discarded.

[0023] In a further solution of the present invention, constructing a security situation awareness model includes the following steps:

[0024] For the data packets with historical records in the cloud-native data warehouse, extract the key behavior feature vectors from the preprocessed multi-source heterogeneous data in the data packets. Experts will score and label each key behavior feature vector according to their own experience, and the labeled behavior feature vectors form the training set of the security situation awareness model;

[0025] Among them, the key behavioral feature vectors include CPU utilization rate, memory usage rate, number of system vulnerabilities, number of behavioral anomalies, and abnormal changes in network traffic;

[0026] Construct a security situation awareness model based on deep learning, and use the training set of the security situation awareness model to train the security situation awareness model, so that the security situation awareness model will formulate a quantitative scoring standard for each key behavioral feature vector.

[0027] A further solution of the present invention, step S3, includes the following steps:

[0028] Combined with historical record data and expert experience, preset the low threshold U of the abnormal degree of CPU utilization rate U 1 and the high threshold U 2 , satisfying

[0029] When the abnormal degree U of CPU utilization rate < U 1 , it indicates that the CPU utilization rate is in a normal state;

[0030] When the abnormal degree U of CPU utilization rate 1 ≤U < U 2 , it indicates that the CPU utilization rate is slightly abnormal;

[0031] When the abnormal degree U of CPU utilization rate ≥ U 2 , it indicates that the CPU utilization rate is severely abnormal.

[0032] A further solution of the present invention, step S3, further includes the following steps:

[0033] Combined with historical record data and expert experience, preset the low threshold M of the abnormal degree of memory usage rate M 1 and the high threshold M 2 , satisfying

[0034] When the abnormal degree M of memory usage rate < M 1 , it indicates that the memory usage rate is in a normal state;

[0035] When the abnormal degree M of memory usage rate 1 ≤M < M 2 , it indicates that the memory usage rate is slightly abnormal;

[0036] When the abnormal degree M of memory usage rate ≥ M 2 , it indicates that the memory usage rate is severely abnormal;

[0037] Combined with historical record data and expert experience, preset the low threshold NTA of the abnormal change degree of network traffic NTA 1 and the high threshold NTA 2 , satisfying

[0038] When the degree of abnormal change in network traffic |NTA| < NTA 1 , it belongs to the normal state;

[0039] When the degree of abnormal change in network traffic NTA 1 ≤|NTA| < NTA 2 , it belongs to a mild abnormality;

[0040] When the degree of abnormal change in network traffic |NTA| ≥ NTA 2 , it belongs to a serious abnormality.

[0041] A further solution of the present invention, step S3, further includes the following steps:

[0042] Combining the data of historical records and expert experience, preset the limit threshold V of the number of system vulnerabilities V max ;

[0043] If there are no known vulnerabilities or all known vulnerabilities have been repaired, it belongs to the normal state; if the number of system vulnerabilities V < V max , it belongs to a mild abnormality; if the number of system vulnerabilities V ≥ V max , it belongs to a severe abnormality;

[0044] Combining the data of historical records and expert experience, preset the limit threshold AB of the number of abnormal behaviors AB max ;

[0045] If there are no known abnormal behaviors, it belongs to the normal state; if the number of abnormal behaviors AB < AB max , it belongs to a mild abnormality; if the number of abnormal behaviors AB ≥ AB max , it belongs to a severe abnormality.

[0046] A further solution of the present invention, step S4, includes the following steps:

[0047] Calculate the dynamic risk prediction value through the already scored CPU utilization rate, memory usage rate, number of system vulnerabilities, number of abnormal behaviors, and degree of abnormal change in network traffic, satisfying the following formula,

[0048] S = W 1 ×S U +W 2 ×S M +W 3 ×S V +W 4 ×S AB +W 5 ×S NTA

[0049] Among them, S represents the risk prediction value; S URepresents the CPU utilization score value; S M Represents the memory usage score value; S V Represents the system vulnerability count score value; S AB Represents the number of abnormal behavior score value; S NTA Represents the score value of abnormal change in network traffic; W 1 、W 2 、W 3 、W 4 、W 5 Represents the weight coefficient of each behavior feature vector, determined according to expert experience.

[0050] A further solution of the present invention, step S5, includes the following steps:

[0051] Combining historical record data and expert experience, set the low-level risk threshold S 1 and the high-level risk threshold S 2 ;

[0052] If S < S 1 , select the low-risk protection mode, trigger the traffic mirroring function, and perform in-depth packet detection, such as implementing access control or traffic filtering;

[0053] If S 1 ≤S < S 2 , select the medium-risk protection mode, and on the basis of the low-risk protection mode, superimpose a dynamic behavior verification mechanism, such as two-factor authentication or behavior trace tracking;

[0054] If S ≥ S 2 , select the high-risk protection mode, immediately activate the fusing mechanism, and forcibly isolate high-risk nodes and block suspicious paths.

[0055] In a second aspect, the present invention provides a security protection system based on big data analysis and cloud computing, adopting the following technical solution:

[0056] A security protection system based on big data analysis and cloud computing, characterized by including the following modules:

[0057] A multi-source heterogeneous data acquisition module, in combination with a preset cloud computing platform, to collect multi-source heterogeneous data in real time;

[0058] A cloud-native data warehouse creation module, used to create a cloud-native data warehouse, encrypt the preprocessed multi-source heterogeneous data and load it into the cloud-native data warehouse;

[0059] A security situation awareness model construction module is used to construct a security situation awareness model. Key behavioral feature vectors are extracted from the preprocessed multi-source heterogeneous data, and the behavioral feature vectors are input into the security situation awareness model to generate corresponding scoring criteria for each behavioral feature vector and perform risk scoring on each behavioral feature vector;

[0060] A risk prediction value calculation module combines the risk score values of the key behavioral feature vectors and outputs a dynamic risk prediction value based on a machine learning algorithm;

[0061] A protection mode selection module sets several risk prediction thresholds, compares the size between the dynamic risk prediction value and the several risk prediction thresholds, and selects an appropriate protection mode.

[0062] In summary, the present invention includes the following beneficial technical effects:

[0063] 1. By combining a preset cloud computing platform, multi-source heterogeneous data is collected in real time and subjected to cleaning and standardization processing. This method and system can quickly obtain and analyze a large amount of security-related data. With the powerful computing power of cloud computing and the accuracy of big data analysis, real-time monitoring and accurate judgment of network security can be achieved, improving the real-time performance and accuracy of security protection;

[0064] 2. Before the preprocessed multi-source heterogeneous data is loaded into the cloud-native data warehouse, this method and system perform encrypted hash processing on the fields containing personal identity information and convert them into irreversible hash values, thereby ensuring the privacy and security of the data. By the server calling the TPM2.0 security chip to generate an RSA-3072 key pair for data packet signature and verification, the security of data transmission and storage is further enhanced, effectively preventing the risks of data leakage and tampering;

[0065] 3. By constructing a security situation awareness model, extracting key behavioral feature vectors and performing risk scoring, and then combining a machine learning algorithm to output a dynamic risk prediction value. By setting different risk prediction thresholds, an appropriate protection mode can be selected according to the risk level. This flexible risk assessment and protection mode selection mechanism enables this method and system to take targeted defense measures according to different security threats and attack scenarios, thus more effectively ensuring the safe and stable operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 Disclosed is a flow schematic diagram of a security protection method based on big data analysis and cloud computing.

[0067] Figure 2 Disclosed is a structural schematic diagram of a security protection system based on big data analysis and cloud computing. DETAILED DESCRIPTION OF THE INVENTION

[0068] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0069] The following will make a preferred and detailed description of the present invention in conjunction with the attached Figure 1-2 drawings.

[0070] Referring to the attached Figure 1 drawings, the present invention proposes a security protection method based on big data analysis and cloud computing, including the following steps:

[0071] S1. In combination with a preset cloud computing platform, collect multi-source heterogeneous data in real time, and perform cleaning and standardization processing on the multi-source heterogeneous data;

[0072] S2. Based on the preset cloud computing platform, create a cloud-native data warehouse, encrypt the preprocessed multi-source heterogeneous data, and load it into the cloud-native data warehouse;

[0073] S3. Build a security situation awareness model, extract key behavior feature vectors from the preprocessed multi-source heterogeneous data, input the behavior feature vectors into the security situation awareness model, and the security situation awareness model generates corresponding scoring criteria for each behavior feature vector and performs risk scoring on each behavior feature vector;

[0074] S4. Combine the risk score values of the key behavior feature vectors and output a dynamic risk prediction value based on a machine learning algorithm;

[0075] S5. Set several risk prediction thresholds, compare the size between the dynamic risk prediction value and the several risk prediction thresholds, and select an appropriate protection mode.

[0076] In one embodiment of the present invention, step S1 includes the following steps:

[0077] Deploy lightweight collection plugins or lightweight proxy programs on the servers to be protected;

[0078] In combination with a preset cloud computing platform, collect multi-source heterogeneous data in real time, and perform cleaning and standardization processing on the multi-source heterogeneous data. The multi-source heterogeneous data includes network data, behavior data, and system data.

[0079] 1. Network data: Collect the time of network connection, IP address, and data volume transmitted;

[0080] Exemplarily, the IP address of the attacker's login attempt is collected and recorded: 192.168.1.100, and the size of a single data packet exceeds 1500 bytes.

[0081] 2. Behavioral data, collecting user login logs and operation records;

[0082] Exemplarily, it is collected and recorded that the user attempted to log in 5 times unsuccessfully at 02:15, and the user deleted an important log file at 17:30 after successful login.

[0083] 3. System data, collecting CPU instant peak values and memory occupancy alarms;

[0084] Exemplarily, it is detected and recorded that during a 5-second period at a certain moment, the CPU usage rate soars from 20% to 98%, and the unknown process "miner.exe" occupies 50% of the CPU's memory.

[0085] Clean and standardize multi-source heterogeneous data, converting multi-source heterogeneous data in different formats and standards into a unified format and standard.

[0086] In one embodiment of the present invention, step S2 includes the following steps:

[0087] Perform encryption hashing (SHA-256) processing on the fields containing personal identity information to convert the fields containing personal identity information into irreversible hash values, and the preprocessed multi-source heterogeneous data forms a complete data packet;

[0088] The server calls the TPM2.0 security chip to generate an RSA-3072 key pair: RSA public key and RSA private key. The server to be protected uses the RSA private key to sign the data packet, and the RSA public key is pre-stored in a preset cloud computing platform;

[0089] The preset cloud computing platform creates a cloud-native data warehouse, loads the data packet into the cloud-native data warehouse, and the preset cloud computing platform uses the RSA public key to verify the signature. If the signature does not match, the data packet will be discarded.

[0090] In one embodiment of the present invention, constructing a security situation awareness model includes the following steps:

[0091] For the data packets with historical records in the cloud-native data warehouse, extract the key behavioral feature vectors from the preprocessed multi-source heterogeneous data in the data packets. Experts will score and label each key behavioral feature vector based on their own experience, and the labeled behavioral feature vectors form the training set of the security situation awareness model;

[0092] Among them, the key behavioral feature vectors include CPU utilization rate, memory usage rate, number of system vulnerabilities, number of behavioral anomalies, and abnormal changes in network traffic;

[0093] Build a security situation awareness model based on deep learning, and use the training set of the security situation awareness model to train the security situation awareness model, so that the security situation awareness model will formulate a quantitative scoring standard for each key behavioral feature vector.

[0094] In one embodiment of the present invention, step S3 includes the following steps:

[0095] For each key behavioral feature vector, the security situation awareness model will formulate a quantitative scoring standard for each key behavioral feature vector, facilitating the conversion of each key behavioral feature vector into a comparable numerical value;

[0096] 1. Combine the historical record data and expert experience to preset the low threshold U of the abnormal degree of CPU utilization rate U 1 and the high threshold U 2 , satisfying:

[0097] When the abnormal degree U of CPU utilization rate < U 1 , it indicates that the CPU utilization rate is in a normal state;

[0098] When the abnormal degree U of CPU utilization rate 1 ≤U < U 2 , it indicates that the CPU utilization rate is mildly abnormal;

[0099] When the abnormal degree U of CPU utilization rate ≥ U 2 , it indicates that the CPU utilization rate is severely abnormal.

[0100] Exemplarily, combine the historical record data and expert experience to preset the low threshold U of the abnormal degree of CPU utilization rate 1 = 70% and the high threshold U 2 = 85%;

[0101] When the abnormal degree U of CPU utilization rate < 70%, it is in a normal state and the score is 1 point;

[0102] When the abnormal degree U of CPU utilization rate 70% ≤ U < 85%, it is mildly abnormal and the score is 3 points;

[0103] When the abnormal degree u of CPU utilization rate ≥ 85%, it is severely abnormal and the score is 5 points.

[0104] 2. Combine the historical record data and expert experience to preset the low threshold M of the abnormal degree of memory usage rate M 1 and the high threshold M 2 , satisfying:

[0105] When the abnormal degree of memory usage M < M 1 , it indicates that the memory usage is in a normal state;

[0106] When the abnormal degree of memory usage M 1 ≤ M < M 2 , it indicates that the memory usage is slightly abnormal;

[0107] When the abnormal degree of memory usage M ≥ M 2 , it indicates that the memory usage is severely abnormal.

[0108] Exemplarily, combining the historical record data and expert experience, the low threshold M 1 of the abnormal degree of memory usage is preset as 80%; the high threshold M 2 is 90%;

[0109] When the abnormal degree of memory usage M < 80%, it belongs to the normal state and is rated 1 point;

[0110] When the abnormal degree of memory usage 80% ≤ M < 90%, it belongs to the slightly abnormal state and is rated 3 points;

[0111] When the abnormal degree of memory usage M ≥ 90%, it belongs to the severely abnormal state and is rated 5 points.

[0112] 3. Combining the historical record data and expert experience, the limit threshold V max of the number of system vulnerabilities V is preset;

[0113] If there are no known vulnerabilities or all known vulnerabilities have been repaired, it belongs to the normal state; if the number of system vulnerabilities V < V max , it belongs to the slightly abnormal state; if the number of system vulnerabilities V ≥ V max , it belongs to the severely abnormal state.

[0114] Exemplarily, combining the historical record data and expert experience, the limit threshold V max of the number of system vulnerabilities V is preset as 3;

[0115] If there are no known vulnerabilities or all known vulnerabilities have been repaired, it belongs to the normal state and is rated 0 points;

[0116] If the number of system vulnerabilities V < 3, it belongs to the slightly abnormal state and is rated 2 points;

[0117] If the number of system vulnerabilities V ≥ 3, it belongs to the severely abnormal state and is rated 5 points.

[0118] 4. Combining the historical record data and expert experience, the limit threshold AB max;

[0119] If there is no known abnormal behavior, it belongs to the normal state; if the number of abnormal behaviors AB < AB max , it belongs to mild abnormality; if the number of abnormal behaviors AB ≥ AB max , it belongs to severe abnormality.

[0120] Exemplarily, combining the data of historical records and expert experience, the limit threshold V of the number of system vulnerabilities V is preset max = 3;

[0121] There is no known abnormal behavior, belonging to the normal state, and the score is 0 points;

[0122] If the number of abnormal behaviors AB < 3, it belongs to mild abnormality, and the score is 2 points;

[0123] If the number of abnormal behaviors AB ≥ 3, it belongs to severe abnormality, and the score is 5 points.

[0124] 5. Combining the data of historical records and expert experience, preset the low threshold NTA and high threshold NTA of the abnormal change degree NTA of network traffic 1 , satisfying: 2 When the abnormal change degree of network traffic |NTA| < NTA

[0125] 1 , it belongs to the normal state;

[0126] When the abnormal change degree of network traffic NTA 1 ≤|NTA| < NTA 2 <000040>, it belongs to mild abnormality;

[0127] When the abnormal change degree of network traffic |NTA| ≥ NTA 2 , it belongs to severe abnormality.

[0128] Exemplarily, combining the data of historical records and expert experience, preset the low threshold NTA of the abnormal change degree of network traffic 1 = 20% and high threshold |NTA| 2 = 50%;

[0129] The abnormal change of network traffic |NTA| < 20%, belonging to the normal state, and the score is 1 point;

[0130] The abnormal change of network traffic 20% ≤ |NTA| < 50%, belonging to mild abnormality, and the score is 3 points;

[0131] The abnormal change of network traffic |NTA| ≥ 50%, belonging to severe abnormality, and the score is 5 points.

[0132] In one embodiment of the present invention, step S4 includes the following steps:

[0133] Calculate a dynamic risk prediction value based on the already scored CPU utilization rate, memory usage rate, number of system vulnerabilities, number of behavioral anomalies, and abnormal changes in network traffic, which satisfies the following formula:

[0134] S = W 1 × S U + W 2 × S M + W 3 × S V + W 4 × S AB + W 5 × S NTA

[0135] where S represents the risk prediction value; S U represents the scored value of CPU utilization rate; S M represents the scored value of memory usage rate; S V represents the scored value of the number of system vulnerabilities; S AB represents the scored value of the number of behavioral anomalies; S NTA represents the scored value of abnormal changes in network traffic; W 1 、W 2 、W 3 、W 4 、W 5 represent the weight coefficients of each behavioral feature vector, which are judged according to expert experience.

[0136] Exemplarily, the CPU utilization rate score S U = 5, the memory usage rate score S M = 3, the number of system vulnerabilities score S V = 2, the number of behavioral anomalies score S AB = 1, the abnormal changes in network traffic score S NTA = 4; according to expert experience judgment, W 1 = 0.3, W 2 = 0.2, W 3 = 0.2, W 4 = 0.1, W 5 = 0.1, and it can be calculated that:

[0137] S = 0.3×5 + 0.2×3 + 0.2×2 + 0.1×1 + 0.1×4 = 3

[0138] In one embodiment of the present invention, step S5 includes the following steps:

[0139] Combined with the data of historical records and expert experience, set the low-level risk threshold S 1and a high-level risk threshold S 2 ;

[0140] If S < S 1 , select a low-risk protection mode, trigger the traffic mirroring function, and perform in-depth packet detection, such as implementing access control or traffic filtering;

[0141] If S 1 ≤ S < S 2 , select a medium-risk protection mode, and on the basis of the low-risk protection mode, superimpose a dynamic behavior verification mechanism, such as two-factor authentication or behavior trace tracking;

[0142] If S ≥ S 2 , select a high-risk protection mode, immediately activate the fusing mechanism, forcibly isolate high-risk nodes and block suspicious paths to ensure system security.

[0143] Exemplarily, combining the data of historical records and expert experience, set a low-level risk threshold S 1 = 2 and a high-level risk threshold S 2 = 4;

[0144] The risk prediction value S calculated in step S4 is 3, 2 ≤ S < 4, it is judged that the current network security situation is a high risk, immediately activate the fusing mechanism, forcibly isolate high-risk nodes and block suspicious paths to ensure system security.

[0145] See Appendix Figure 2 , the present invention also proposes a security protection system based on big data analysis and cloud computing, including the following modules:

[0146] A multi-source heterogeneous data collection module, combined with a preset cloud computing platform, to collect multi-source heterogeneous data in real time;

[0147] A cloud-native data warehouse creation module, used to create a cloud-native data warehouse, encrypt the preprocessed multi-source heterogeneous data and load it into the cloud-native data warehouse;

[0148] A security situation awareness model construction module, used to construct a security situation awareness model, extract key behavior feature vectors from the preprocessed multi-source heterogeneous data, input the behavior feature vectors into the security situation awareness model, generate corresponding scoring criteria for each behavior feature vector, and perform risk scoring on each behavior feature vector;

[0149] A risk prediction value calculation module, combined with the risk score values of the key behavior feature vectors, outputs a dynamic risk prediction value based on a machine learning algorithm;

[0150] The protection mode selection module sets several risk prediction thresholds, compares the dynamic risk prediction value with the several risk prediction thresholds, and selects an appropriate protection mode.

[0151] Each of the above-mentioned modules can be implemented in whole or in part by software, hardware, and their combinations, and supports being embedded in the processor of the computer device in hardware form or being independent of it. At the same time, it also supports being stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above-mentioned modules.

[0152] It should be noted that the user information (including but not limited to user device information and personal information, etc.) and data (including but not limited to data for analysis, stored data, and displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data require relevant legal standards.

[0153] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A security protection method based on big data analysis and cloud computing, characterized in that: The following steps are involved: S1. Combined with the preset cloud computing platform, multi-source heterogeneous data is collected in real time, and the multi-source heterogeneous data is cleaned and standardized; S2. Based on the preset cloud computing platform, a cloud-native data warehouse is created, and the pre-processed multi-source heterogeneous data is encrypted and loaded into the cloud-native data warehouse; S3. Construct a security situation awareness model, extract key behavior feature vectors from preprocessed multi-source heterogeneous data, input the behavior feature vectors into the security situation awareness model, and the security situation awareness model generates corresponding scoring criteria for each behavior feature vector, and performs risk scoring on each behavior feature vector; S4. Combine the risk score values ​​of key behavioral feature vectors and output a dynamic risk estimate based on the machine learning algorithm; S5. Set several risk estimation thresholds, compare the dynamic risk estimation value with several risk estimation thresholds, and select a suitable protection mode.

2. According to claim 1, a security protection method based on big data analysis and cloud computing is characterized in that: Step S1 includes the following steps: The server that needs to be protected is deployed with a lightweight collection plug-in or a lightweight agent program; Combined with the preset cloud computing platform, multi-source heterogeneous data is collected in real time, and the multi-source heterogeneous data is cleaned and standardized. The multi-source heterogeneous data includes network data, behavior data, and system data; Network data, including network connection time, IP address, and amount of data transmitted; Behavioral data, collecting user login logs and operation records; System data, collects CPU instantaneous peak value and memory usage alarm.

3. According to the security protection method based on big data analysis and cloud computing in claim 1, it is characterized in that: Step S2 includes the following steps: Perform cryptographic hashing on the fields containing personal identity information to convert the fields containing personal identity information into irreversible hash values. The pre-processed multi-source heterogeneous data form a complete data package. The server calls the TPM2.0 security chip to generate an RSA-3072 key pair, RSA public key, and RSA private key. The server to be protected uses the RSA private key to sign the data packet. The RSA public key is pre-stored in the preset cloud computing platform. The preset cloud computing platform creates a cloud-native data warehouse, and the data packet is loaded into the cloud-native data warehouse. The preset cloud computing platform uses the RSA public key to verify the signature. If the signature does not match, the data packet will be discarded.

4. A security protection method based on big data analysis and cloud computing according to claim 3, characterized in that: Building a security situation awareness model includes the following steps: The cloud-native data warehouse contains historical data packets. The pre-processed multi-source heterogeneous data in the data packets extract key behavioral feature vectors. Experts will score and label each key behavioral feature vector based on their own experience. The labeled behavioral feature vectors form the training set of the security situation awareness model. Among them, the key behavioral feature vectors include CPU utilization, memory usage, number of system vulnerabilities, number of behavioral anomalies, and abnormal changes in network traffic; A security situation awareness model based on deep learning is constructed, and the security situation awareness model is trained using a training set of the security situation awareness model, so that the security situation awareness model can formulate a quantitative scoring standard for each key behavioral feature vector.

5. A security protection method based on big data analysis and cloud computing according to claim 4, characterized in that: Step S3 includes the following steps: Combined with historical data and expert experience, preset the low threshold U1 and high threshold U2 for the abnormal degree of CPU utilization U, satisfying When the abnormal degree of CPU utilization U < U1, it indicates that the CPU utilization is in a normal state; When the abnormal degree of CPU utilization U1 ≤ U < U2, it indicates that the CPU utilization is in a mild abnormal state; When the abnormal degree of CPU utilization U ≥ U2, it indicates that the CPU utilization is in a severe abnormal state.

6. A security protection method based on big data analysis and cloud computing according to claim 5, characterized in that: Step S3 also includes the following steps: Combined with historical data and expert experience, preset the low threshold M1 and high threshold M2 for the abnormal degree of memory usage M, satisfying When the abnormal degree of memory usage M < M1, it indicates that the memory usage is in a normal state; When the abnormal degree of memory usage M1 ≤ M < M2, it indicates that the memory usage is in a mild abnormal state; When the abnormal degree of memory usage M ≥ M2, it indicates that the memory usage is in a severe abnormal state; Combined with historical data and expert experience, preset the low threshold NTA1 and high threshold NTA2 for the abnormal change degree NTA of network traffic, satisfying When the abnormal change degree of network traffic |NTA| < NTA1, it belongs to the normal state; When the abnormal change degree of network traffic NTA1 ≤ |NTA| < NTA2, it belongs to the mild abnormal state; When the abnormal change degree of network traffic |NTA| ≥ NTA2, it belongs to the severe abnormal state.

7. A security protection method based on big data analysis and cloud computing according to claim 6, characterized in that: Step S3 also includes the following steps: Combine historical data and expert experience to pre-set the limit threshold V of the number of system vulnerabilities V max ; If there are no known vulnerabilities or all known vulnerabilities have been fixed, it is a normal state; if there are a number of system vulnerabilities V <V max , which is a mild anomaly; if the number of system vulnerabilities V ≥ V max , which is a severe abnormality; Combine historical data and expert experience to pre-set the limit threshold AB of the number of abnormal behaviors AB max ; If there is no known abnormal behavior, it is a normal state; if there is an abnormal number of behaviors AB <AB max , which is a mild abnormality; if there are abnormal behaviors, the number of AB≥AB max , which is a severe abnormality.

8. A security protection method based on big data analysis and cloud computing according to claim 7, characterized in that: Step S4 includes the following steps: Calculate the dynamic risk prediction value through the already scored CPU utilization, memory usage, number of system vulnerabilities, number of behavioral anomalies, and abnormal change of network traffic, satisfying the following formula S=W1×S U +W2×S M +W3×S V +W4×S AB +W5×S NTA Where S represents the risk estimate; S U Indicates the CPU utilization score value; S M Indicates the memory usage score value; S V Indicates the number of system vulnerabilities. AB Indicates the score value of the number of abnormal behaviors; S NTA Indicates the score value of abnormal changes in network traffic; W1, W2, W3, W4, and W5 represent the weight coefficients of each behavior feature vector, which are determined based on expert experience.

9. A security protection method based on big data analysis and cloud computing according to claim 8, characterized in that: Step S5 includes the following steps: Combined with historical data and expert experience, set the low-level risk threshold S1 and high-level risk threshold S2; If S < S1, select the low-risk protection mode, trigger the traffic mirroring function, and perform in-depth packet detection, such as implementing access control or traffic filtering; If S1 ≤ S < S2, select the medium-risk protection mode, and on the basis of the low-risk protection mode, superimpose a dynamic behavior verification mechanism, such as two-factor authentication or behavior track tracing; If S ≥ S2, select the high-risk protection mode, immediately start the fusing mechanism, and forcibly isolate high-risk nodes and block suspicious paths.

10. A security protection system based on big data analysis and cloud computing, characterized in that: It includes the following modules: Multi-source heterogeneous data collection module, combined with a preset cloud computing platform, to collect multi-source heterogeneous data in real time; Cloud-native data warehouse creation module, used to create a cloud-native data warehouse, encrypt the preprocessed multi-source heterogeneous data and load it into the cloud-native data warehouse; Security situation awareness model construction module, used to construct a security situation awareness model, extract key behavioral feature vectors from the preprocessed multi-source heterogeneous data, input the behavioral feature vectors into the security situation awareness model, generate corresponding scoring criteria for each behavioral feature vector, and perform risk scoring on each behavioral feature vector; Risk prediction value calculation module, combined with the risk score values of key behavioral feature vectors, and output a dynamic risk prediction value based on machine learning algorithms; The protection mode selection module sets several risk estimation thresholds, compares the dynamic risk estimation value with several risk estimation thresholds, and selects an appropriate protection mode.