Data security protection method based on big data analysis

Through big data analysis and gated recurrent unit networks, combined with historical database logs, the target attacker is identified and locked, solving the problem of inaccurate attacker identity identification in knowledge graph analysis and improving data security.

CN120805155APending Publication Date: 2025-10-17CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510855539.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

In IoT technology, errors are prone to occur during knowledge graph analysis, resulting in inaccurate identification of attackers and affecting data security.

Method used

Through a method based on big data analysis, the access logs and modification logs of the historical database are used for clustering processing, combined with gated recurrent unit network analysis, to identify and lock the target attacker, upgrade the firewall and restore the tampered data.

Benefits of technology

It improves the accuracy of attacker identification, enhances data security, reduces the risk of misjudgment, and ensures the data security of industrial equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805155A_ABST
    Figure CN120805155A_ABST
Patent Text Reader

Abstract

The invention provides a data security protection method based on big data analysis, and relates to the technical field of data security. The method comprises the following steps: clustering first historical node data in a historical database according to determined tampered reference historical node data to obtain a second historical node data set and a third historical node data set; determining a target node corresponding to each piece of data in the second historical node data set as a reference target node to obtain a reference target node set; and taking the access log and the modification log corresponding to each reference target node as fourth historical node data, and analyzing the fourth historical node data by using a gating loop unit network to determine a target attacker. Misjudgment caused by knowledge extraction errors when the knowledge graph is constructed is avoided, the target attacker can be accurately determined, and the security of data in the database can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data security technology, and specifically, to a data security protection method based on big data analysis. Background Art

[0002] The widespread adoption of IoT technology has enabled widespread interconnection of industrial equipment. The deep integration of industrial control systems with the IoT and the internet has significantly improved production efficiency and information technology, but it has also introduced serious cybersecurity risks. Related technologies use knowledge graphs to analyze data from historical databases or remote monitoring systems, identifying data leaks or tampering and tracing attackers to improve data security. However, when constructing knowledge graphs, errors can easily occur during the process of extracting knowledge from vast amounts of data. This can lead to inaccurate knowledge graphs, which in turn affects attacker identification, leading to misjudgment of the true attack path and threatening data security. Summary of the Invention

[0003] The embodiments of the present application provide a data security protection method based on big data analysis, aiming to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0004] A first aspect of the embodiments of the present application provides a data security protection method based on big data analysis, comprising: Obtaining a first historical node data set based on first historical node data corresponding to each target node in the historical database, wherein the first historical node data includes device information corresponding to a device monitored by the target node; According to the access log and the modification log in the historical database, first historical node data that has been illegally modified and illegally accessed in the first historical node data set is determined as reference historical node data to obtain a reference historical node data set; Based on the reference historical node data set, clustering the first historical node data in the first historical node data set to obtain a second historical node data set and a third historical node data set, wherein the second historical node data set is a set of the tampered first historical node data, and the third historical node data set is a set of the untampered first historical node data; Determine the target node corresponding to each second historical node data in the second historical node data set as a reference target node to obtain a reference target node set; Determine the access log and modification log corresponding to each reference target node in the reference target node set as fourth historical node data to obtain a fourth historical node data set; The fourth historical node data in the fourth historical node data set is analyzed through a gated recurrent unit network to obtain a target attacker.

[0005] In an optional implementation, the method further comprises: According to the access log and the modification log corresponding to each reference target node in the set of reference target nodes, obtaining the attack means of the target attacker; According to the attack means of the target attacker, upgrading the firewall of the historical database and performing data recovery on each second historical node data in the set of second historical node data.

[0006] In an optional implementation, the method further comprises: Each target node respectively encrypts the collected initial historical node data through the respective encryption unit carried by each target node to obtain the first historical node data corresponding to each target node and stores the first historical node data in the historical database; According to the key corresponding to the encryption unit carried by each target node, obtaining the first historical node data corresponding to each target node in the historical database.

[0007] In an optional implementation, based on the set of reference historical node data, the first historical node data in the set of first historical node data is subjected to clustering processing to obtain a set of second historical node data and a set of third historical node data, comprising: Preprocessing the reference historical node data in the set of reference historical node data to obtain a vector of each reference historical node data, and preprocessing each first historical node data in the set of first historical node data to obtain a vector of each first historical node data; According to the vector of each reference historical node data, determining a first clustering center; Calculating the relative similarity between the vector of each first historical node data and the first clustering center; Taking the vector of the first historical node data with the smallest relative similarity as a second clustering center; According to the first clustering center and the second clustering center, the first historical node data in the set of first historical node data is subjected to clustering processing to obtain the set of second historical node data and the set of third historical node data.

[0008] In an optional implementation, preprocessing the reference historical node data in the set of reference historical node data to obtain a vector of each reference historical node data, and preprocessing each first historical node data in the set of first historical node data to obtain a vector of each first historical node data, comprising: The feature information of each reference historical node data is obtained by feature extraction on each reference historical node data through a hash algorithm. The feature information of each reference historical node data is taken as a vector of each reference historical node data. The feature information of each first historical node data is obtained by feature extraction on each first historical node data through a hash algorithm. The feature information of each first historical node data is taken as a vector of each first historical node data. The feature information is a hash value obtained by encrypting data through a hash algorithm.

[0009] In an optional embodiment, the relative similarity of the vector of each first historical node data and the first clustering center is calculated, including: The relative similarity of the vector of each first historical node data and the first clustering center is calculated according to the following formula:

[0010] Wherein, α i represents the relative similarity of the vector of each first historical node data and the first clustering center; δ i represents the correlation factor; Q i represents the number of components of the vector of the first historical node data in the first historical node data set; || represents absolute value calculation; a i represents the i-th component in the first clustering center; b i represents the i-th component in the vector of each first historical node data; represents the module length of the vector of each first historical node data; represents the module length of the first clustering center.

[0011] In an optional embodiment, according to the first clustering center and the second clustering center, the first historical node data in the first historical node data set is clustered to obtain the second historical node data set and the third historical node data set, including: The first distance of each first historical node data and the first clustering center is calculated, and the second distance of each first historical node data and the second clustering center is calculated; It is judged whether the first distance corresponding to each first historical node data is greater than the second distance; In the case that the first distance corresponding to the first historical node data is less than or equal to the second distance, the first historical node data is determined as the second historical node data in the second historical node data set; In a case where the first distance corresponding to the first historical node data is greater than the second distance, the first historical node data is determined as third historical node data in the third historical node data set.

[0012] In an optional implementation, the first distance of each first historical node data to the first cluster center and the second distance of each first historical node data to the second cluster center are calculated, including: mapping the vector of each first historical node data, the first cluster center and the second cluster center into a same-dimensional multi-dimensional space to obtain a distribution of the vector of each first historical node data, the first cluster center and the second cluster center in the multi-dimensional space respectively; determining a distance parameter corresponding to the first historical node data set and the reference historical node data set according to a first similarity between the first historical node data set and the reference historical node data set; calculating the first distance of each first historical node data to the first cluster center and the second distance of each first historical node data to the second cluster center according to the distribution and the distance parameter.

[0013] In an optional implementation, the first distance of each first historical node data to the first cluster center and the second distance of each first historical node data to the second cluster center are calculated according to the distribution and the distance parameter, including: the first distance or the second distance is calculated by the following formula:

[0014] wherein, S represents the Euclidean distance between each first historical node data and the first cluster center or the second cluster center; n represents the number of components of the vector of the first historical node data; ω i represents the distance parameter; a i represents the i-th component in the first cluster center or the second cluster center; b i represents the i-th component in each first historical node data; β i represents the first similarity between the first historical node data set and the reference historical node data set; γ represents a self-defined parameter for controlling the speed of change of the first weight information size.

[0015] In an optional implementation, the fourth historical node data in the fourth historical node data set is analyzed by a gated recurrent unit network to obtain a target attacker, including: construct a supply chain topology graph according to fourth historical node data in the fourth historical node data set, the supply chain topology graph comprising nodes and edges, the nodes representing the reference target nodes, and the edges representing interaction relationships between the nodes; process time information of the abnormal event in the fourth historical node data corresponding to each node through the gated recurrent unit network to obtain abnormal event development information corresponding to each node; analyze connection strength between each node and other nodes according to the abnormal event development information corresponding to each node to obtain a risk propagation mode of each node; determine a risk level of each node according to the risk propagation mode of each node and a hierarchical relationship between nodes in the supply chain topology graph; determine a user corresponding to the fourth historical node data of a node with a risk level greater than a preset risk level threshold as a pending attacker to obtain a pending attacker set; determine a rating of each pending attacker according to a dependency relationship between the node corresponding to each pending attacker and other nodes; determine a pending attacker with a target rating as the target attacker.

[0016] The second aspect of the embodiments of the application provides a data security protection device based on big data analysis, the device comprising: an acquisition module configured to obtain a first historical node data set according to first historical node data corresponding to each target node in a historical database, the first historical node data comprising device information corresponding to a device monitored by the target node; a first determination module configured to determine, according to access logs and modification logs in the historical database, first historical node data with illegal modification and illegal access in the first historical node data set as reference historical node data to obtain a reference historical node data set; a processing module configured to perform clustering processing on the first historical node data in the first historical node data set based on the reference historical node data set to obtain a second historical node data set and a third historical node data set, the second historical node data set being a set of tampered first historical node data, and the third historical node data set being a set of un-tampered first historical node data; a second determination module configured to determine a target node corresponding to each second historical node data in the second historical node data set as a reference target node to obtain a reference target node set; a third determining module, configured to determine the access log and the modification log corresponding to each reference target node in the reference target node set as fourth historical node data, to obtain a fourth historical node data set; an analyzing module, configured to analyze the fourth historical node data in the fourth historical node data set through a gated recurrent unit network, to obtain a target attacker.

[0017] The third aspect of the embodiments of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the computer program is executed by the processor to implement the data security protection method based on big data analysis according to the first aspect of the present application.

[0018] The fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the data security protection method based on big data analysis according to the first aspect of the present application.

[0019] In the data security protection method based on big data analysis provided by the present application, the first historical node data in the historical database is subjected to clustering processing according to the determined tampered reference historical node data, so that the tampered first historical node data (second historical node data set) in the historical database can be quickly and accurately identified, and misjudgment caused by knowledge extraction error when constructing a knowledge graph is avoided. Each data in the second historical node data set is determined as a reference target node to form a reference target node set, which directly points to the target node of the device that may be attacked, and provides a key clue for tracking the attacker. The access log and the modification log corresponding to the reference target node are collected as fourth historical node data, and the gated recurrent unit network is used to analyze these data to determine the target attacker. The gated recurrent unit network can effectively learn the time sequence dependency in the historical access log and the modification log, identify the behavior pattern and characteristics of the attacker, and thus more accurately lock the target attacker. Accurate determination of the target attacker helps to upgrade the firewall of the database according to the attack means of the target attacker, thereby improving the security of the data in the database. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1is a step flow chart of a data security protection method based on big data analysis according to an embodiment of the present application; Figure 2 is a step flow chart of topology driving and gated recurrent unit network analysis data of the data security protection method based on big data analysis according to an embodiment of the present application; Figure 3 is a structural schematic diagram of the data security protection device based on big data analysis according to an embodiment of the present application; Figure 4 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0023] In the drawings, sometimes the size of the constituent elements, the thickness of the layers or the regions is exaggerated for the sake of clarity, and therefore, any one of the implementations of the present disclosure is not necessarily limited to the size shown in the drawings, and the shape and size of the components in the drawings do not reflect the actual scale. In addition, the drawings schematically show ideal examples, and any one of the implementations of the present disclosure is not limited to the shape or value shown in the drawings.

[0024] Referring to Figure 1 , Figure 1 is a step flow chart of a data security protection method based on big data analysis according to an embodiment of the present application. As shown in Figure 1 , the method comprises steps S11-S16.

[0025] Step S11: obtaining a first historical node data set according to first historical node data corresponding to each target node in a historical database, wherein the first historical node data comprises device information corresponding to a device monitored by the target node.

[0026] In this embodiment, the target nodes are used to monitor the running status of the equipment in real time and collect the equipment information of the equipment in real time, which can be various sensors (such as temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.) used for monitoring the equipment on the production line. The equipment information includes the running parameters (such as rotation speed, current, voltage) of the equipment and the production environment data (such as temperature, humidity, pressure) and the like. The first historical node data corresponding to each target node is the equipment information corresponding to the equipment monitored by each target node. The target nodes are interconnected through the Internet of Things technology, and the collected node data is uploaded to the historical database or the remote monitoring system through the Internet of Things technology. Users can access the historical database or the remote monitoring system to view the running status of each device on the production line and obtain the first historical node data corresponding to each target node. According to the first historical node data corresponding to each target node in the historical database, a first historical node data set is obtained.

[0027] Step S12: According to the access log and the modification log in the historical database, the first historical node data in the first historical node data set that has illegal modification and illegal access is determined as reference historical node data, and a reference historical node data set is obtained.

[0028] In this embodiment, the access log and the modification log are obtained from the historical database or the remote monitoring system. These logs usually record the operation information of the target node data of the user, including operation time, operation type (access or modification), operation user, operation content and the like. According to the access log and the modification log in the historical database or the remote monitoring system, the first historical node data set is traversed, and the first historical node data in which there is illegal modification and illegal access (i.e. tampered) is determined as reference historical node data, and a reference historical node data set is obtained. The reference historical node data in this application can also determine some tampered historical node data as reference historical node data according to the historical access log and the historical modification log obtained from the historical database or the remote monitoring system in the past. Through the access log and the modification log, the recorded tampered first historical node data is obtained, which provides a key reference for subsequent classification processing.

[0029] Step S13: Based on the reference historical node data set, the first historical node data in the first historical node data set is processed to obtain a second historical node data set and a third historical node data set, the second historical node data set is a set of tampered first historical node data, and the third historical node data set is a set of non-tampered first historical node data.

[0030] In this embodiment, it is considered that not all tampered first historical node data is recorded in the access log and the modification log in the historical database or the remote detection system. Therefore, the first historical node data in the first historical node data set can be clustered based on the reference historical node data set to obtain a second historical node data set and a third historical node data set. Each second historical node data in the second historical node data set is tampered first historical node data in the first historical node data set; and each third historical node data in the third historical node data set is first historical node data in the first historical node data set that is not tampered. The data of the reference historical node data set is used as a reference to make the clustering result more consistent with the actual situation and improve the accuracy of tampered data identification. Through clustering, the tampered first historical node data (second historical node data set) in the historical database can be quickly and accurately identified, providing data basis for subsequent identification of target attackers.

[0031] Step S14: determining each second historical node data in the second historical node data set as a reference target node to obtain a reference target node set.

[0032] In this embodiment, each second historical node data in the second historical node data set can be determined as a reference target node according to the correspondence between each second historical node data and the target node to obtain a reference target node set. The determined reference target node can provide a direction for subsequent log analysis and key node information for subsequent analysis of target attackers.

[0033] Step S15: determining the access log and the modification log corresponding to each reference target node in the reference target node set as fourth historical node data to obtain a fourth historical node data set.

[0034] In this embodiment, the access log and the modification log corresponding to each reference target node in the reference target node set are queried from the historical database or the remote detection system, and the access log and the modification log corresponding to each reference target node are determined as fourth historical node data to obtain a fourth historical node data set. The log information (fourth historical node data) of the reference target node corresponding to the tampered data is obtained to provide data support for subsequent identification of target attackers.

[0035] Step S16: analyzing the fourth historical node data in the fourth historical node data set through a gated recurrent unit network to obtain a target attacker.

[0036] In this embodiment, the gated recurrent unit network is an improved recurrent neural network for processing sequence data (such as time series, log streams) to capture the time dependence between events. By analyzing the fourth historical node data in the fourth historical node data set through the gated recurrent unit network, the target attacker is obtained. By using the powerful sequence data processing capability of the gated recurrent unit network, the behavior characteristics of the attacker can be mined from the complex log data, the target attacker can be accurately identified, and then the attack means of the attacker can be understood, so that the security protection can be targeted.

[0037] In the data security protection method based on big data analysis provided in the application, the first historical node data in the historical database is clustered according to the determined tampered reference historical node data, which can quickly and accurately identify the tampered first historical node data (second historical node data set) in the historical database, and avoid misjudgment caused by knowledge extraction error when building a knowledge graph. Each data in the second historical node data set corresponds to a target node, which is determined as a reference target node to form a reference target node set, which directly points to the target node of the device that may be attacked, providing a key clue for tracking the attacker. Collect the access log and modification log corresponding to the reference target node as the fourth historical node data, and analyze these data by using the gated recurrent unit network to determine the target attacker. The gated recurrent unit network can effectively learn the time sequence dependence in the historical access log and modification log, identify the behavior pattern and characteristics of the attacker, and thus more accurately lock the target attacker. Accurate determination of the target attacker helps to upgrade the firewall of the database according to the attack means of the target attacker, thereby improving the security of the data in the database.

[0038] In one embodiment, the application further provides a data security protection method based on big data analysis, which further comprises steps S21 and S22.

[0039] Step S21: obtaining the attack means of the target attacker according to the access log and modification log corresponding to each reference target node in the reference target node set; Step S22: upgrading the firewall of the historical database according to the attack means of the target attacker, and recovering the data of each second historical node data in the second historical node data set.

[0040] In this embodiment, the access log and the modification log corresponding to the reference target node related to the target attacker are analyzed to obtain key information such as operation steps of the attacker, tools used, and time regularity of the attack. Through the arrangement and analysis of the key information, the attack means (such as SQL injection, cross-site scripting attack, brute force cracking, etc.) of the target attacker are obtained. According to the attack means of the target attacker, the firewall of the historical database is upgraded, for example, for the attack means of the SQL injection attack, a regular expression rule can be configured in the firewall to strictly verify all data entering the database. Then, the backup file related to the second historical node data is searched from the data backup library, and according to the backup file, data recovery is performed on each second historical node data in the second historical node data set. According to the corresponding protection means corresponding to the attack means of the target attacker, the security of the data in the historical database can be improved, and the tampered data in the historical database can be recovered, which can reduce the impact of the tampered data on the production line.

[0041] In an embodiment, the application further provides a data security protection method based on big data analysis, which further comprises steps S31 and S32.

[0042] Step S31: Each target node respectively encrypts the collected initial historical node data through the encryption unit carried by each target node to obtain the first historical node data corresponding to each target node, and stores the first historical node data in the historical database. Step S32: According to the key corresponding to the encryption unit carried by each target node, the first historical node data corresponding to each target node in the historical database is obtained.

[0043] In this embodiment, when each target node uploads the initial historical node data collected by itself to the historical database or the remote monitoring system, a suitable encryption algorithm can be selected according to the security requirements, such as a symmetric encryption algorithm or an asymmetric encryption algorithm. After selecting the encryption algorithm, the initial historical node data collected by each target node is encrypted through the encryption unit carried by each target node to obtain the first historical node data corresponding to each target node, and the first historical node data is stored in the historical database or the remote monitoring system. By encrypting the data collected by the target node, even if the data is intercepted during transmission or illegally accessed in the historical database or the remote monitoring system, the original device information cannot be obtained, which can prevent data leakage and ensure the safety of the data on the production line.

[0044] In this embodiment, when the user accesses the history database or the remote monitoring system, the identity of the user needs to be authenticated, such as verifying the username and password, using digital certificates, etc. If the user is an authorized user, the system will assign the corresponding key to the user according to the key information corresponding to the encryption unit of each target node. Then the user accesses the data in the history database or the remote monitoring system according to the key corresponding to the encryption unit carried by each target node respectively, so as to obtain the first historical node data corresponding to each target node in the history database or the remote monitoring system. Only the user carrying the corresponding key can view the corresponding node data, ensuring the safe access of the data on the production line.

[0045] In an embodiment, the present application also provides a data security protection method based on big data analysis. In the method, the step S13 of "performing clustering processing on the first historical node data in the first historical node data set based on the reference historical node data set to obtain a second historical node data set and a third historical node data set" can specifically include the following steps S41 to S45: Step S41: pre-processing the reference historical node data in the reference historical node data set to obtain a vector of each reference historical node data, and pre-processing each first historical node data in the first historical node data set to obtain a vector of each first historical node data; Step S42: determining a first clustering center according to the vector of each reference historical node data.

[0046] Step S43: calculating the relative similarity of the vector of each first historical node data and the first clustering center; Step S44: taking the vector of the first historical node data with the smallest relative similarity as a second clustering center; Step S45: performing clustering processing on the first historical node data in the first historical node data set according to the first clustering center and the second clustering center to obtain the second historical node data set and the third historical node data set.

[0047] In this embodiment, the reference historical node data in the reference historical node data set and each first historical node data in the first historical node data set are preprocessed. The purpose of preprocessing is to convert the original data (reference historical node data and first historical node data) into a structured vector form that can be used for clustering. The method of preprocessing can be feature extraction of the data to obtain respective corresponding feature information, and the respective corresponding feature information is used as a vector corresponding to the respective data, so as to obtain a vector of each reference historical node data and a vector of each first historical node data. By preprocessing to convert the data into a unified vector form, it is convenient for subsequent clustering processing of the first historical node data; and the preprocessed data can better reflect the differences and similarities between the data, which helps to improve the accuracy of subsequent clustering of the first historical node data.

[0048] In this embodiment, the vector of each historical node data contains multiple components, and correspondingly, the vector of each historical node data also contains multiple components. Since the first historical node data includes device information corresponding to the device monitored by the target node, the device information includes running parameters (such as speed, current, voltage) and production environment data (such as temperature, humidity, pressure) of the device, etc., each component of the vector of each historical node data is a feature (feature information) of each parameter of the extracted device information. For example, a certain first historical data specifically includes current, voltage and temperature, and the feature information obtained after preprocessing includes current feature, voltage feature and temperature feature; these features are the components of the vector of the first historical data. In this embodiment, the mean vector of all reference historical node data vectors can be calculated, and the obtained mean vector is used as the first clustering center. Then, the relative similarity of each first historical node data vector and the first clustering center is calculated, and the first historical node data vector with the smallest relative similarity is used as the second clustering center. According to the first clustering center and the second clustering center, the first historical node data in the first historical node data set is clustered to obtain a second historical node data set and a third historical node data set. The clustering center corresponding to the second historical node data in the second historical node data set is the first clustering center; and the clustering center corresponding to the third historical node data in the third historical node data set is the second clustering center. According to the first clustering center and the second clustering center, the clustering processing is performed, which avoids the classification deviation caused by random initialization in common clustering, and can accurately distinguish the first historical node data that is tampered and the first historical node data that is not tampered.

[0049] In an implementation, the present application further provides a data security protection method based on big data analysis, in which the "preprocessing the reference historical node data in the reference historical node data set to obtain a vector of each reference historical node data, and preprocessing each first historical node data in the first historical node data set to obtain a vector of each first historical node data" in the step S41 can specifically include the following steps S51-S54: Step S51: performing feature extraction on each reference historical node data by a hash algorithm to obtain feature information of each reference historical node data; Step S52: taking the feature information of each reference historical node data as the vector of each reference historical node data; Step S53: performing feature extraction on each first historical node data by a hash algorithm to obtain feature information of each first historical node data; Step S54: taking the feature information of each first historical node data as the vector of each first historical node data; Wherein, the feature information is a hash value obtained by encrypting the data by a hash algorithm.

[0050] In the embodiment, the feature information of each reference historical node data obtained by performing feature extraction on each reference historical node data by a hash algorithm is a hash value obtained by encrypting the data by a hash algorithm, and the feature information of each reference historical node data is taken as the vector of each reference historical node data. Similarly, the vector of each first historical node data can be obtained. The feature information of each reference historical node data can also be obtained by performing feature extraction on each reference historical node data by other encryption algorithms (such as a check algorithm), and the feature information of the data can be a check value obtained by using other encryption algorithms. The data is vectorized by a hash algorithm, which realizes the unified representation of the data and facilitates subsequent data processing and analysis, and can protect the privacy and security of the data to a certain extent.

[0051] In an implementation, the present application further provides a data security protection method based on big data analysis, in which the "calculating the relative similarity of each first historical node data vector and the first clustering center" in the step S43 can be calculated by the following formula: The relative similarity of each first historical node data vector and the first clustering center is calculated according to the following formula:

[0052] Wherein, α i represents the relative similarity of each first historical node data vector and the first clustering center; δ i represents the correlation factor; Qi denotes the number of components of the vector of the first historical node data in the first historical node data set; || denotes absolute value calculation; a i denotes the i-th component in the first cluster center; b i denotes the i-th component in the vector of each first historical node data; denotes the length of the vector of each first historical node data; denotes the length of the first cluster center.

[0053] In an embodiment, the present application also provides a data security protection method based on big data analysis, in which the step S41 of "performing clustering processing on the first historical node data in the first historical node data set according to the first cluster center and the second cluster center to obtain the second historical node data set and the third historical node data set" can specifically include the following steps S61 to S64: Step S61: calculating a first distance of each first historical node data from the first cluster center, and a second distance of each first historical node data from the second cluster center; Step S62: judging whether the first distance corresponding to each first historical node data is greater than the second distance; Step S63: in the case that the first distance corresponding to the first historical node data is less than or equal to the second distance, determining that the first historical node data is a second historical node data in the second historical node data set; Step S64: in the case that the first distance corresponding to the first historical node data is greater than the second distance, determining that the first historical node data is a third historical node data in the third historical node data set.

[0054] In the present embodiment, a first distance of each first historical node data from the first cluster center is calculated, and a second distance of each first historical node data from the second cluster center is calculated. By judging whether the first distance corresponding to each first historical node data is greater than the second distance, if it is greater, it means that the first historical node data is a third historical node data in the third historical node data set, otherwise, it means that the first historical node data is a second historical node data in the second historical node data set, so as to obtain the second historical node data set and the third historical node data set. The cluster center corresponding to the second historical node data set is the first cluster center, and the cluster center corresponding to the third historical node data set is the second cluster center.

[0055] In an implementation, the application further provides a data security protection method based on big data analysis, in which the "calculating a first distance of each first historical node data from the first cluster center and a second distance of each first historical node data from the second cluster center" in the step S61 can specifically include the following steps S71-S73: Step S71: mapping the vector of each first historical node data, the first cluster center and the second cluster center into a same-dimensional multi-dimensional space to obtain a distribution of the vector of each first historical node data, the first cluster center and the second cluster center in the multi-dimensional space respectively; Step S72: determining a distance parameter corresponding to the first historical node data set and the reference historical node data set according to a first similarity between the first historical node data set and the reference historical node data set; Step S73: calculating the first distance of each first historical node data from the first cluster center and the second distance of each first historical node data from the second cluster center according to the distribution and the distance parameter.

[0056] In the embodiment, the vector of each first historical node data, the first cluster center and the second cluster center are mapped into a same-dimensional multi-dimensional space to obtain a distribution of the vector of each first historical node data, the first cluster center and the second cluster center in the multi-dimensional space respectively. The data of different dimensions are processed by dimension increasing to map all the data into a same-dimensional multi-dimensional space. For example, the components of a general vector of first historical node data are 5, and the components of a certain first historical node data are only 4, so a preset value (such as 0) is taken as the missing component of the data to complete the dimension increasing processing of the data. By mapping all the data into a same-dimensional multi-dimensional space, distance calculation and visualization can be performed in a unified space, which provides a space basis for subsequent clustering.

[0057] In the embodiment, the first similarity between the first historical node data set and the reference historical node data set can be calculated by set similarity measurement (such as Jaccard similarity, cosine similarity, Hellinger distance, etc.). The distance parameter is adjusted according to the first similarity, the lower the first similarity, the larger the distance parameter, and vice versa. The distance parameter is used to weight the components of the Euclidean distance in combination with the custom parameter. Then, the first distance of each first historical node data from the first cluster center and the second distance of each first historical node data from the second cluster center are calculated according to the distribution and the distance parameter.

[0058] In an implementation, the present application also provides a data security protection method based on big data analysis, in which the step S73 of calculating the first distance between each first historical node data and the first clustering center and the second distance between each first historical node data and the second clustering center according to the distribution and the distance parameter can be calculated by the following formula: The first distance or the second distance is calculated by the following formula:

[0059] wherein S represents the Euclidean distance between each first historical node data and the first clustering center or the second clustering center; n represents the number of components of the vector of the first historical node data; ω i represents the distance parameter; a i represents the i-th component in the first clustering center or the second clustering center; b i represents the i-th component in each first historical node data; β i represents the first similarity between the first historical node data set and the reference historical node data set; γ represents a custom parameter for controlling the speed of change of the first weight information size.

[0060] In an implementation, referring to Figure 2 , Figure 2 is a step flowchart of the topology-driven and gated recurrent unit network analysis data of the data security protection method based on big data analysis according to an embodiment of the present application. As shown in Figure 2 , the step S16 of analyzing the fourth historical node data in the fourth historical node data set by the gated recurrent unit network to obtain the target attacker can specifically include the following steps S71 to S77: Step S71: constructing a supply chain topology graph according to the fourth historical node data in the fourth historical node data set, wherein the supply chain topology graph includes nodes and edges, the nodes represent the reference target nodes, and the edges represent the interaction relationship between the nodes; Step S72: processing the time information of the abnormal events in the fourth historical node data corresponding to each node by the gated recurrent unit network to obtain the abnormal event development information corresponding to each node; Step S73: analyzing the connection strength between each node and other nodes according to the abnormal event development information corresponding to each node to obtain the risk propagation mode of each node; Step S74: determining the risk level of each node according to the risk propagation mode of each node and the hierarchical relationship between the nodes in the supply chain topology graph; Step S75: determining the user who performs illegal access and illegal modification in the fourth historical node data corresponding to the node with a risk level greater than the preset risk level threshold as a pending attacker, to obtain a pending attacker set; Step S76: determining the rating of each pending attacker according to the dependency relationship between the node corresponding to each pending attacker in the pending attacker set and other nodes; Step S77: determining the pending attacker with a target rating as the target attacker.

[0061] In this embodiment, the fourth historical node data in the fourth historical node data set is the access log and modification log corresponding to each reference target node. Each reference target node (with data tampering) is taken as each node of the supply chain topology graph, and if two nodes have interaction records in the logs, an edge is established, and finally a constructed supply chain topology graph is obtained to clearly show the association relationship between nodes. The time information of abnormal events in the fourth historical node data corresponding to each node is processed by the gating recurrent unit network, the long-term dependency relationship in the abnormal event time information is captured, and the abnormal event development information (trend, periodicity, etc.) corresponding to each node is output. According to the abnormal event development information corresponding to each node, the connection strength between each node and other nodes is analyzed in combination with the supply chain topology graph, so as to identify the risk propagation mode of each node. The risk propagation mode can include linear propagation, star propagation, mesh propagation, etc.

[0062] In this embodiment, the risk assessment index of each node is determined according to the risk propagation mode of each node and the hierarchical relationship between nodes in the supply chain topology graph. For example, the position of the node in the risk propagation path (such as whether it is on the critical path), the connection strength with other nodes, the level, the importance of the node (such as whether it is a core business node), etc. Each risk assessment index is assigned a corresponding weight to reflect its influence on the risk level. Based on the weight, the risk score of each node can be calculated using weighted summation or other methods. According to the risk score, the risk level of each node is determined.

[0063] In this embodiment, the user who performs illegal access and illegal modification (for example, frequently attempts illegal login, modifies key data, accesses sensitive information, etc.) in the fourth historical node data corresponding to the node with a risk level greater than the preset risk level threshold is determined as a pending attacker, and a pending attacker set is obtained. Since illegal operations may be misoperations, exploratory attacks or accidental behaviors, it is necessary to further determine the target attacker according to the pending attacker set. By analyzing the supply chain topology graph, the dependency relationship between the node corresponding to the pending attacker in the pending attacker set and other nodes is obtained, and each pending attacker in the pending attacker set is rated according to the strength of the dependency relationship. The pending attacker rated as a target level (a preset highest level) is determined as a target attacker, and the target attacker is one or more.

[0064] Since the production line will make corresponding adjustments to the production process or material ratio on the production line according to the historical node data collected by each target node, thereby improving the quality of the finished products produced by the production line. And competitors may tamper with the data in the historical database, thereby causing incorrect adjustments to the production process or material ratio on the production line, resulting in a decrease in the quality of the finished products produced by the production line. Therefore, it is necessary to analyze the historical data in the historical database, determine the real identity of the attacker, obtain the target attacker, and according to the attack means of the target attacker, the firewall of the database is upgraded, thereby improving the security of the data in the database.

[0065] The present application can accurately distinguish between tampered and non-tampered first historical node data by calculating the similarity of each first historical node data in the first historical node data set and the nodes in the first cluster, and determining the first historical node data with the smallest similarity as the second cluster, and performing clustering processing according to the first cluster center and the second cluster center. The fourth historical node data is the access log and modification log corresponding to the tampered first historical node data determined by clustering. According to the principle of graph theory, a supply chain topology graph is constructed according to each fourth historical node data in the fourth historical node information set, and the connection strength of each node (corresponding to the reference target node) in the supply chain topology graph is analyzed according to the development time information of the fourth historical node information, the mode of risk propagation in the supply chain topology graph is identified, and the target attacker set is determined according to the mode of risk propagation and the hierarchical relationship between each node in the topology graph. The target attacker is determined by rating the pending attacker according to the strength of the dependency relationship between the reference target nodes. Through dynamic risk mode identification, double determination mechanism (preliminary determination according to risk level, and final determination according to strength of dependency relationship), the identity of the attacker can be accurately identified. Therefore, the firewall of the database can be upgraded according to the attack means of the target attacker, and the security of the data in the database can be improved.

[0066] Based on the same inventive concept, one embodiment of the present application provides a data security protection device based on big data analysis. Figure 3 is a structural schematic diagram of the data security protection device based on big data analysis provided by one embodiment of the present application, as Figure 3 shown, the device comprises: An acquisition module, configured to obtain a first historical node data set according to first historical node data corresponding to each target node in a historical database, wherein the first historical node data comprises device information corresponding to a device monitored by the target node; A first determination module, configured to determine, according to access logs and modification logs in the historical database, first historical node data in the first historical node data set that has illegal modification and illegal access as reference historical node data, and obtain a reference historical node data set; A processing module, configured to perform clustering processing on the first historical node data in the first historical node data set based on the reference historical node data set, and obtain a second historical node data set and a third historical node data set, wherein the second historical node data set is a set of tampered first historical node data, and the third historical node data set is a set of un-tampered first historical node data; A second determination module, configured to determine, as a reference target node, a target node corresponding to each second historical node data in the second historical node data set, and obtain a reference target node set; A third determination module, configured to determine, as fourth historical node data, access logs and modification logs corresponding to each reference target node in the reference target node set, and obtain a fourth historical node data set; An analysis module, configured to analyze, through a gated recurrent unit network, fourth historical node data in the fourth historical node data set, and obtain a target attacker.

[0067] Based on the same inventive concept, one embodiment of the present application provides an electronic device, referring to Figure 4 , Figure 4 is a schematic diagram of an electronic device according to one embodiment of the present application. As Figure 4 shown, the electronic device 100 comprises a memory 110 and a processor 120, the memory 110 and the processor 120 are communicatively connected through a bus, the memory 110 stores a computer program, the computer program can run on the processor 120, and then implement steps in the data security protection method based on big data analysis described in any of the embodiments of the present application.

[0068] Based on the same inventive concept, the disclosure further provides a computer readable storage medium, when instructions in the computer readable storage medium are executed by a processor of a computer device, the computer device is enabled to perform the steps in the data security protection method based on big data analysis in any of the embodiments of the present application.

[0069] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0070] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, device or computer program product. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0071] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the method, terminal device (system) and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be realized by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The device for implementing the functions specified in one block or multiple blocks.

[0072] These computer program instructions can also be stored in a computer readable storage medium which can guide the computer or other programmable data processing terminal device to work in a specific way, so that the instructions stored in the computer readable storage medium produce a product including instruction devices, which implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one flow or multiple flows and / or blocks Figure 1 The device for implementing the functions specified in one block or multiple blocks.

[0073] These computer program instructions can also be loaded into a computer or other programmable data processing terminal device, so that a series of operation steps are performed on the computer or other programmable terminal device to produce a computer implemented process, so that the instructions executed on the computer or other programmable terminal device provide a product for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1one or more processes and / or blocks Figure 1 the steps of a function specified in one or more blocks.

[0074] While preferred embodiments of the application have been described, those skilled in the art will appreciate that other modifications than those specifically described can be made within the scope of the application. Accordingly, the appended claims are intended to embrace all such alternatives as well.

[0075] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are not intended to denote a limitation (such as "first" versus "second"). Also, the terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0076] The above provides a kind of data security protection method based on big data analysis provided by the present application, has carried out detailed introduction, the principle and implementation mode of the present application are described in this paper by specific example, the above example is only for helping to understand the method of the present application and its core idea;For the person skilled in the art, according to the idea of the present application, there will be changes in specific implementation mode and application range, as described above, the content of the specification should not be understood as the limitation of the present application.

Claims

1. A data security protection method based on big data analysis, characterized in that: The method comprises: Obtaining a first historical node data set based on first historical node data corresponding to each target node in the historical database, wherein the first historical node data includes device information corresponding to a device monitored by the target node; According to the access log and the modification log in the historical database, first historical node data that has been illegally modified and illegally accessed in the first historical node data set is determined as reference historical node data to obtain a reference historical node data set; Based on the reference historical node data set, clustering the first historical node data in the first historical node data set to obtain a second historical node data set and a third historical node data set, wherein the second historical node data set is a set of the tampered first historical node data, and the third historical node data set is a set of the untampered first historical node data; Determine the target node corresponding to each second historical node data in the second historical node data set as a reference target node to obtain a reference target node set; Determine the access log and modification log corresponding to each reference target node in the reference target node set as fourth historical node data to obtain a fourth historical node data set; The fourth historical node data in the fourth historical node data set is analyzed through a gated recurrent unit network to obtain a target attacker.

2. The data security protection method based on big data analysis according to claim 1 is characterized in that: The method further comprises: Obtaining the attack means of the target attacker according to the access log and modification log corresponding to each reference target node in the reference target node set; According to the attack method of the target attacker, the firewall of the historical database is upgraded, and data recovery is performed on each second historical node data in the second historical node data set.

3. The data security protection method based on big data analysis according to claim 1 is characterized in that: The method further comprises: Each target node encrypts the collected initial historical node data through its own encryption unit to obtain the first historical node data corresponding to each target node and stores it in the historical database; The first historical node data corresponding to each target node in the historical database is obtained according to the key corresponding to the encryption unit carried by each target node.

4. The data security protection method based on big data analysis according to claim 1 is characterized in that: Based on the reference historical node data set, clustering processing is performed on the first historical node data in the first historical node data set to obtain a second historical node data set and a third historical node data set, including: Preprocessing the reference historical node data in the reference historical node data set to obtain a vector of each reference historical node data, and preprocessing each first historical node data in the first historical node data set to obtain a vector of each first historical node data; Determine the first cluster center according to the vector of each reference historical node data; Calculating the relative similarity between the vector of each first historical node data and the first cluster center; The vector of the first historical node data with the smallest relative similarity is used as the second cluster center; The first historical node data in the first historical node data set is clustered according to the first cluster center and the second cluster center to obtain the second historical node data set and the third historical node data set.

5. The data security protection method based on big data analysis according to claim 4 is characterized in that: Preprocessing the reference historical node data in the reference historical node data set to obtain a vector of each reference historical node data, and preprocessing each first historical node data in the first historical node data set to obtain a vector of each first historical node data, including: The feature information of each reference historical node data is obtained by extracting the feature of each reference historical node data through the hash algorithm; Taking the characteristic information of each reference historical node data as the vector of each reference historical node data; Extract features from each first historical node data using a hash algorithm to obtain feature information of each first historical node data; Taking the characteristic information of each first historical node data as the vector of each first historical node data; The characteristic information is a hash value obtained by encrypting the data using a hash algorithm.

6. The data security protection method based on big data analysis according to claim 4 is characterized in that: Calculating the relative similarity between the vector of each first historical node data and the first cluster center includes: The relative similarity between the vector of each first historical node data and the first cluster center is calculated according to the following formula: Among them, α i Represents the relative similarity between the vector of each first historical node data and the first cluster center; δ i represents the correlation factor; Q i represents the number of components of the vector of the first historical node data in the first historical node data set; || represents absolute value calculation; a i represents the i-th component in the first cluster center; b i The i-th component in the vector representing each first historical node data; The modulus of the vector representing each first historical node data; Indicates the modulus of the first cluster center.

7. The data security protection method based on big data analysis according to claim 4 is characterized in that: Clustering the first historical node data in the first historical node data set according to the first cluster center and the second cluster center to obtain the second historical node data set and the third historical node data set includes: Calculating a first distance between each first historical node data and the first cluster center, and calculating a second distance between each first historical node data and the second cluster center; Determine whether the first distance corresponding to each first historical node data is greater than the second distance; When the first distance corresponding to the first historical node data is less than or equal to the second distance, determining that the first historical node data is the second historical node data in the second historical node data set; When the first distance corresponding to the first historical node data is greater than the second distance, the first historical node data is determined to be the third historical node data in the third historical node data set.

8. The data security protection method based on big data analysis according to claim 7 is characterized in that: Calculating a first distance between each first historical node data and the first cluster center, and calculating a second distance between each first historical node data and the second cluster center, including: Mapping the vector of each first historical node data, the first cluster center, and the second cluster center to a multidimensional space of the same dimension, and obtaining the distribution of the vector of each first historical node data, the first cluster center, and the second cluster center in the multidimensional space; Determining a distance parameter corresponding to the first historical node data set and the reference historical node data set according to a first similarity between the first historical node data set and the reference historical node data set; According to the distribution and the distance parameter, a first distance between each first historical node data and the first cluster center is calculated, and a second distance between each first historical node data and the second cluster center is calculated.

9. The data security protection method based on big data analysis according to claim 8 is characterized in that: Calculating a first distance between each first historical node data and the first cluster center, and calculating a second distance between each first historical node data and the second cluster center according to the distribution and the distance parameter, including: The first distance or the second distance is calculated by the following formula: Where S represents the Euclidean distance between each first historical node data and the first cluster center or the second cluster center; n represents the number of components of the vector of the first historical node data; ω i represents the distance parameter; a i represents the i-th component in the first cluster center or the second cluster center; b i represents the i-th component in each first historical node data; β i represents the first similarity between the first historical node data set and the reference historical node data set; γ represents a custom parameter used to control the speed of change of the size of the first weight information.

10. The data security protection method based on big data analysis according to claim 1 is characterized in that: Analyzing the fourth historical node data in the fourth historical node data set through a gated recurrent unit network to obtain a target attacker includes: constructing a supply chain topology graph based on the fourth historical node data in the fourth historical node data set, the supply chain topology graph including nodes and edges, the nodes representing the reference target nodes, and the edges representing the interaction relationships between the nodes; The time information of the abnormal event in the fourth historical node data corresponding to each node is processed by the gated recurrent unit network to obtain the abnormal event development information corresponding to each node; According to the abnormal event development information corresponding to each node, the connection strength between each node and other nodes is analyzed to obtain the risk propagation model of each node; Determining the risk level of each node based on the risk propagation model of each node and the hierarchical relationship between the nodes in the supply chain topology; Determine the user who illegally accesses and modifies the fourth historical node data corresponding to the node whose risk level is greater than the preset risk level threshold as a pending attacker, and obtain a set of pending attackers; Determining a rating of each pending attacker based on a dependency relationship between a node corresponding to each pending attacker in the pending attacker set and other nodes; The pending attacker rated as the target level is determined as the target attacker.