Method and related products for predicting data security risks

By constructing a topology graph of nodes and using machine learning to predict data security events, the method addresses the limitations of single-dimensional risk prediction, enhancing accuracy and enabling proactive risk mitigation.

CN119135406BActive Publication Date: 2025-07-15HANGZHOU MINGSHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411241534.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-07-15
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

The existing data security risk prediction methods mainly focus on a single dimension, ignore the influence of factors such as personnel and application systems, and have a limited prediction scope. It requires full inspection after an attack. The workload is large and time-consuming and labor-intensive, making it difficult to accurately predict the next data security incident.

Method used

Build a topology diagram, including personnel, data, application systems and cloud resource nodes, predict data security risks through machine learning models, and combine the risk characteristics and association relationships of nodes to predict the node type and specific events of the next data security event.

Benefits of technology

It realizes efficient and accurate prediction of data security events, improves the depth and breadth of prediction, reduces unnecessary waste of resources, and improves the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119135406B_ABST
    Figure CN119135406B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method for predicting data security risks and related products thereof. The method includes: constructing a topology graph among the nodes of a target object according to the node types; determining a security event sequence and a candidate node set of the topology graph according to the data security events that have occurred at each node in the topology graph, where the candidate nodes are the nodes in the topology graph that are directly adjacent to the end node of the security event sequence and have not yet had data security events occur; and inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of the candidate nodes in the candidate node set. The solution of the present disclosure can accurately predict the node types and specific data security events of the next data security event, so as to accurately predict the data security risk values of the candidate nodes directly adjacent to the current attack chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of data security technologies. More specifically, this disclosure relates to a method, an electronic device, and a computer-readable storage medium for predicting data security risks. Background Art

[0002] With the rapid development of the Internet and cloud technologies, data has become the most important asset in all industries. Each organization and enterprise has started to use more and more technologies and methods to process, store, and use data, and the resulting data security issues have become increasingly important. Security incidents such as data leakage, hacking attacks, and unauthorized operations may all cause huge losses to organizations or enterprises.

[0003] Advanced Persistent Threat (APT) attacks have diverse threat channels, a long concealment time, and a high degree of correlation with various factors in the environment where the target of the attack is located, making it difficult to predict what kind of attack event will occur in what environment next time.

[0004] In view of this, there is an urgent need to provide a solution for predicting data security risks in order to predict the node types of the next data security incident and the specific data security events, so as to achieve efficient and accurate prediction of data security risks related to the current attack chain. Summary of the Invention

[0005] In order to solve at least one or more of the above-mentioned technical problems, this disclosure proposes a solution for predicting data security risks in the following aspects.

[0006] In a first aspect, this disclosure provides a method for predicting data security risks, including: constructing a topology graph between each node of a target object according to node types, where the topology graph includes horizontal associations between nodes of the same node type and vertical associations between nodes of different node types; determining a security event sequence and a candidate node set of the topology graph according to the data security events that have occurred at each node in the topology graph, where the candidate node is a node in the topology graph that is directly adjacent to the end node of the security event sequence and has not yet had a data security event; and inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of each candidate node in the candidate node set.

[0007] In some embodiments, the node types include personnel, data, application systems, and cloud resources.

[0008] In some embodiments, constructing a topology graph between nodes of a target object according to node types includes: detecting risk characteristics of each node in the topology graph by using a risk characteristic detection tool; and determining the possibility of a data security event occurring for each node based on the risk characteristics.

[0009] In some embodiments, after obtaining the data security risk values of the candidate nodes in the candidate node set, the method further includes: determining the value-weighted risk of the candidate nodes based on the data security value and the data security risk value of the candidate nodes; and outputting event warnings and prevention tips outward according to the value-weighted risk.

[0010] In some embodiments, the data security value is determined according to value factors of the candidate node; the value factors include: whether it is a data node or a node directly storing data, whether the amount of stored data is greater than 1000, whether the classification of the stored data is at or above a preset classification, whether the stored data is generated or updated within a preset time limit, and whether the stored data is correlated with other data sets.

[0011] In some embodiments, inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of the candidate nodes in the candidate node set includes: inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the possibility of each node type being attacked and the probability of a data security event occurring when a node of each node type is attacked after the security event sequence occurs; and determining the data security risk value of the candidate node according to the node type of the candidate node, the risk characteristics of the candidate node, the possibility of a data security event occurring for the candidate node, the possibility of each node type being attacked, and the probability of a data security event occurring when a node of each node type is attacked.

[0012] In some embodiments, after determining the possibility of a data security event occurring for each node, the method further includes: sorting the possibility of a data security event occurring for each node to determine the data security event corresponding to the maximum possibility value of each node; and outputting a possibility event prompt outward for the data security event corresponding to the maximum possibility value.

[0013] In some embodiments, the preset security event prediction model is a model obtained by training using a machine learning model. Training the preset security event prediction model using a machine learning model includes: collecting and analyzing historical security event data of a training object to obtain a security event sequence of the training object and information on induced security events, where the information on induced security events includes the node type attacked after the security event sequence and the data security events that occurred at the attacked node; and inputting the security event sequence and the information on induced security events as training data into the machine learning model to train the machine learning model and obtain the preset security event prediction model.

[0014] In a second aspect, the present disclosure provides an electronic device, including: a processor; and a memory that stores program instructions for predicting data security risks. When the program instructions are run by the processor, the methods and their various embodiments described in the foregoing first aspect are implemented.

[0015] In a third aspect, the present disclosure provides a computer-readable storage medium that stores program instructions for predicting data security risks. When the program instructions are executed by a processor, the methods and their various embodiments described in the foregoing first aspect are implemented.

[0016] Through the solution for predicting data security risks provided as above, the embodiments of the present disclosure can determine the security event sequence of the topology graph and the candidate node set based on the data security events that have occurred at each node in the topology graph. Then, by inputting the security event sequence into the preset security event prediction model for security event prediction, the data security risk values of the candidate nodes in the candidate node set can be obtained. Thus, the solution of the present disclosure can accurately predict the node type of the next data security event and the specific data security event, thereby accurately predicting the data security risk values of the candidate nodes directly adjacent to the current attack chain.

[0017] Further, in some embodiments, the node types in the embodiments of the present disclosure may include personnel, data, application systems, and cloud resources, taking into account various factors associated with data security events, not only increasing the depth and breadth of data security risk prediction, but also being able to simulate a multi-dimensional and complex environment, which helps to improve the accuracy of data security risk prediction.

[0018] Furthermore, in some embodiments, the embodiments of the present disclosure can determine the data security value of each node according to the value factors of each node. Furthermore, the predicted data security risk value can be risk-reduced based on the data security value of the candidate nodes to determine the value-weighted risk of the candidate nodes. Based on this, the value-weighted risk of the embodiments of the present disclosure is the data security risk value obtained by comprehensively considering the node type and data value factors of the nodes. In other words, for nodes of the same node type, different data security risk values may be obtained under the influence of data value factors, which helps to improve the refinement degree of data security risk prediction. At the same time, it is possible to avoid over-prevention of nodes with low data security value, resulting in waste of resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of example and not limitation, and the same or corresponding reference numerals denote the same or corresponding parts, wherein:

[0020] Figure 1 Shows an exemplary flowchart of a method for predicting data security risks according to an embodiment of the present disclosure;

[0021] Figure 2 Shows an exemplary schematic diagram of the association relationship topology of four types of nodes according to an embodiment of the present disclosure;

[0022] Figure 3 Shows an exemplary instance of the association relationship topology of four types of nodes according to an embodiment of the present disclosure;

[0023] Figure 4 Shows an exemplary schematic diagram of a security event sequence according to an embodiment of the present disclosure;

[0024] Figure 5 Shows an exemplary schematic diagram of a candidate node according to an embodiment of the present disclosure;

[0025] Figure 6 Shows Figure 1 An exemplary flowchart of one step;

[0026] Figure 7 Shows Figure 1 An exemplary flowchart of another step;

[0027] Figure 8 Shows an exemplary structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0029] It should be understood that when terms such as "first", "second", "third", and "fourth" are used in the claims, the description, and the drawings of the present application, they are only used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" used in the description and claims of the present disclosure indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0030] It should also be understood that the terms used in the description of the present disclosure herein are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. As used in the description and claims of the present disclosure, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms. It should also be further understood that the term "and / or" used in the description and claims of the present disclosure refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0031] As used in this specification and the claims, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0032] The following will describe in detail the specific implementation manners of the present disclosure with reference to the accompanying drawings.

[0033] Overview of the Present Disclosure

[0034] In the process of obtaining the solution of the present disclosure, the applicant found that data security is built on top of network security, and without network security, there is no data security. However, when predicting data security risks, if only network security is considered, the expected attack protection effect cannot be achieved. Therefore, in addition to predicting data security risks and protecting against attacks on the host, factors such as the people, application systems, and data that the host may be associated with also need to be considered. For example, a person being socially engineered may lead to the leakage of the application system password. Another example is that a security vulnerability in a cloud resource host may lead to intrusion and the acquisition of sensitive data, etc.

[0035] However, most of the existing data security risk prediction methods focus on relatively single dimensions. Usually, they only perform data security risk prediction and attack protection for host-type assets and / or the data itself, ignoring the influence of factors such as people and application systems. The prediction scope is limited, and the backtracking sequence is relatively fixed. At the same time, once an attack occurs, all assets need to be investigated, which is a large amount of work and time-consuming.

[0036] In order to more accurately predict the upcoming data security events, the present disclosure collects and analyzes the occurred data security events, classifies all the occurred data security events into nodes of four node types: people, data, application systems, and cloud resources, and analyzes the risk characteristics of the nodes of different node types and their correlation with data security events. Subsequently, the risk characteristics and correlation are combined with the security event sequence, and a multi-dimensional machine learning model is used to predict the node type of the next data security event and the specific data security event, so as to achieve efficient and accurate prediction of the security risks related to the current attack chain and reduce risks according to the node value.

[0037] For the convenience of understanding, before introducing the solution of the present disclosure in detail, the technical terms involved in the present disclosure will be introduced first.

[0038] Graph database: A type of database used to store the relationships between entities. It uses a graph structure to represent data, where entities are represented as nodes and the relationships between entities are represented as edges. Graph databases are particularly suitable for processing complex relationship networks. In a graph database, the relationships between data are as important as the data itself and are stored as part of the data. Such an architecture gives graph databases advantages such as high performance, scalability, and flexibility, and can quickly respond to complex association queries because the relationships between entities have been pre-stored in the database.

[0039] Machine learning model: An algorithm that automatically learns and generates prediction results based on the given input data (also called features). These models can be divided into two major categories: supervised learning models and unsupervised learning models.

[0040] Exemplary method

[0041] Next, in combination with Figure 1 an exemplary introduction will be given to method 100 for predicting data security risks in some embodiments of the present disclosure. It can be understood that method 100 can be executed by any suitable device with data processing capabilities, such as, for example, but not limited to, terminal devices and servers, etc.

[0042] As Figure 1 shown, at step S101, method 100 can construct a topology graph between the nodes of the target object according to the node type. Here, the target object is used to identify any object that is taken as the target when predicting data security risks, and this object can be any organization or institution.

[0043] In the solution of the present disclosure, the entities of the target object can be represented by the nodes in the topology graph, and the mutual relationships between the entities can be represented by the association relationships between the nodes. The types of nodes in the topology graph (or can also be called node types) can include, but are not limited to, personnel, data, application systems, and cloud resources. The association relationships between the nodes in the topology graph can include horizontal associations between nodes of the same node type and vertical associations between nodes of different node types.

[0044] Based on this, when constructing the topology graph between the nodes of the target object in the solution of the present disclosure, the nodes corresponding to the entities of the target object can be constructed first, and the nodes can be classified according to the node type. Then, the association relationships between the nodes can be constructed according to the mutual relationships between the entities, so as to obtain the topology graph of the target object. From the implementation level, the present disclosure can use a graph database to store the topology graph of the target object. Through the relational storage method of the graph database, the association relationships between the nodes can be better shown, and it can provide convenience for complex data management and query operations.

[0045] For the convenience of understanding, here in combination with Figure 2 and Figure 3 a more detailed introduction will be given to the topology graph between the nodes. As Figure 2 shown, it shows an exemplary schematic diagram of the association relationship topology graph of four types of nodes. In Figure 2Among them, the domain is the value range of four types of nodes: personnel, data, application systems, and cloud resources. In this topology diagram, a node can be connected to other nodes of the same type through horizontal association, or to nodes of different types through vertical association. For example, a cloud resource can be connected to and access other cloud resources within a local area network, which is horizontal association. An application system is running on this cloud resource, which is vertical association. Here, the node type of cloud resource will be used as an example to illustrate the association relationship between nodes. By analogy, the association relationships of other node types such as personnel, data, and application systems can be understood.

[0046] Specifically, the curve connecting to itself on the right side of the cloud resource represents the horizontal association between different nodes of the same cloud resource type; the straight line connecting the cloud resource to the data on the left represents the vertical association between the node of the cloud resource type and the node of the data type; the straight line connecting the cloud resource to the application system on the upper side represents the vertical association between the node of the cloud resource type and the node of the application system type; the straight line connecting the cloud resource to the personnel on the upper left represents the vertical association between the node of the cloud resource type and the node of the personnel type.

[0047] Furthermore, Figure 3 an exemplary instance of the association relationship topology diagram of the four types of nodes is shown. As Figure 3 shown, an entity of the target object corresponds to a node, and the mutual relationships between the entities determine the association relationships between the nodes. In this instance, the nodes are grouped into four types: personnel, application systems, cloud resources, and data according to the node type, and the number of nodes included in each type varies.

[0048] Specifically, the personnel type includes nodes such as Zhang San, Li Si, Wang Wu, and Personnel XXX. Here, Personnel XXX represents other personnel type nodes. The application system type includes nodes such as the financial system, performance system, attendance system, and Application System XXX. Here, Application System XXX represents other application system type nodes. The cloud resource type includes ECS running the financial system and performance system, ECS running the attendance system, RDS storing employee identity information, OSS storing unstructured information related to employees, and Cloud Resource XXX. Here, Cloud Resource XXX represents other cloud resource type nodes. The data type nodes include the employee identity information table, employee face pictures, fingerprint information files, and Data XXX. Here, Data XXX represents other data type nodes. Here, ECS is the abbreviation of Elastic Compute Service, RDS is the abbreviation of Relational Database Service, and OSS is the abbreviation of Object Storage Service.

[0049] AsFigure 3 As shown, the nodes of the personnel type can be horizontally associated through co - working relationships, and the nodes of each personnel type and the nodes of the application system type can be vertically associated through development, operation and maintenance, testing or design relationships. The nodes of the application system type can be horizontally associated through data interaction, call interfaces or data transmission relationships, and in addition to the vertical association with the nodes of the personnel type, each node can also be vertically associated with the nodes of the cloud resource type through the running - on relationship. The nodes of the cloud resource type can be horizontally associated through intranet interconnection, data acquisition or data storage, and in addition to the vertical association with the nodes of the application relationship type, each node can also be vertically associated with the nodes of the data type through the storage relationship. The nodes of the data type can be horizontally associated through consanguinity or logical association relationships. It should be understood that Figure 3 The node types, node quantities and node relationships shown in are merely exemplary and illustrative. Those skilled in the art can set the node types, node quantities and node relationships according to actual needs, and the present disclosure does not make specific limitations in this regard.

[0050] Continuing from the foregoing step S101, at step S102, method 100 can determine the security event sequence and candidate node set of the topology graph according to the data security events that have occurred for each node in the topology graph.

[0051] In the solution of the present disclosure, the data security event indicates the specific situation of a node being attacked or negatively affected. The data security event may include, but is not limited to, password leakage, being socially engineered, being injected with a shell reverse instruction, data tampering, password cracking, being loaded with an attack payload, unauthorized access, and being subjected to a distributed denial - of - service (DDOS) attack. It should be understood that the data security events shown here are merely exemplary and illustrative, and those skilled in the art can expand them according to actual needs, and the present disclosure does not make specific limitations in this regard.

[0052] From an implementation perspective, the present disclosure can use a data security event set to store the data security events to be used. As an example, the data security event set can be expressed as {password leakage, being socially engineered, being injected with a shell reverse instruction, data tampering, password cracking, being loaded with an attack payload, unauthorized access, being subjected to a DDOS attack}.

[0053] The security event sequence disclosed herein indicates a set of associated nodes that have been attacked successively over a period of time. The security event sequence may specifically include each node that has been attacked, the order in which each node has been attacked, the type of each node, and the data security events that have occurred at each node. It is easy to understand that the node that is directly adjacent to the last node of the security event sequence and has not yet experienced a data security event has the highest probability of experiencing a data security event in the next attack and should be guarded against. Therefore, in the embodiments disclosed herein, the node that is directly adjacent to the last node of the security event sequence and has not yet experienced a data security event is determined as a candidate node, and its data security risk value is predicted so that the staff can formulate corresponding prevention strategies based on the data security risk value to ensure the data security of the candidate node.

[0054] For the sake of easy understanding, continue to combine the foregoing Figure 3 to introduce in detail how to determine the security event sequence and the candidate node set of the topology graph. Before that, first introduce the data security events that have occurred at each node in the topology graph.

[0055] The order of occurrence of data security events: 1. Li Si obtained the password of Zhang San's financial system stored in plain text in the enterprise shared document. 2. When Li Si browsed the financial system, it was found that the financial system would call the interface of the performance system to obtain employee information, and the interface provided by the performance system did not implement an authentication mechanism and could be accessed without restriction. Therefore, Li Si injected a shell reverse connection instruction into the interface. 3. Li Si successfully invaded the ECS running the performance system through the shell reverse connection instruction and achieved full control of the ECS.

[0056] Analyzing the data security events that have occurred at each node in the topology graph, it can be seen that at data security event 1, the attacked node is the financial system, the order of attack is 1, the type of the node is an application system, and the data security event that occurred is the password being leaked; at data security event 2, the attacked node is the performance system, the order of attack is 2, the type of the node is an application system, and the data security event that occurred is being injected with a shell reverse connection instruction; at data security event 3, the attacked node is the ECS running the performance system, the order of attack is 3, the type of the node is cloud resources, and the data security event that occurred is unauthorized access. Thus, a security event sequence of length 3 as Figure 4 shown can be obtained.

[0057] After determining the security event sequence, the node that is directly adjacent to the last node of the security event sequence and has not yet experienced a data security event can be determined as a candidate node. As Figure 3As shown, the end node of the security event sequence is the ECS of the operation performance system. The nodes that are directly adjacent to this end node and have not yet experienced a data security event are the RDS storing employee identity information and the ECS running the attendance system. The types of both of these nodes are cloud resource types. At the same time, the ECS of the operation performance system is horizontally associated with the RDS storing employee identity information through a data acquisition relationship, and the ECS of the operation performance system is horizontally associated with the ECS running the attendance system through an intranet interconnection relationship. Thus, the candidate nodes as shown in Figure 5 can be obtained.

[0058] Continuing from the aforementioned step S102, at step S103, method 100 can input the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of each candidate node in the candidate node set.

[0059] In the solution disclosed herein, the preset security event prediction model is a model obtained by training using a machine learning model, which is used to predict the likelihood of a certain type of candidate node being attacked and the probability of a specific data security event occurring when the type of node is attacked after the security event sequence occurs. In actual operation, the machine learning model can adopt any one of a recurrent neural network (RNN), support vector regression (SVR), long short-term memory network (LSTM), and transformer network.

[0060] In the solution disclosed herein, when using a machine learning model to train and obtain a preset security event prediction model, the historical security event data of the training object can be first collected, analyzed, and sorted to obtain the security event sequence of the training object and the information on the induced security events. The information on the induced security events here includes the type of node attacked after the security event sequence and the data security event that occurred to the attacked node. Then, the security event sequence and the information on the induced security events can be used as training data to input into the machine learning model to train the machine learning model and obtain the preset security event prediction model.

[0061] In the solution disclosed herein, the training object is similar to the aforementioned target object and can be any organization or institution. The historical security event data of the training object can be regularly collected, processed, analyzed, sorted, and obtained from multiple channels such as open source intelligence, various organizations in cooperation, the long-term operation logs of enterprise intranets, and data from attack and defense drills.

[0062] The data security risk values disclosed herein can identify the risks of candidate nodes being attacked. After the preset security event prediction model predicts the likelihood of a certain type of candidate node being attacked and the probability of specific data security events occurring when that type of node is attacked, the likelihood of each candidate node being attacked and the probability of specific data security events occurring when it is attacked can be determined by combining the node types of the candidate nodes, thereby determining the data security risk values of each candidate node.

[0063] Further, after method 100 determines the data security risk values of each candidate node, it can also output the data security risk values of each candidate node externally. Additionally or optionally, method 100 can output the data security risk values of each candidate node in a visual or audible manner.

[0064] The above combination Figure 1 describes method 100 for predicting data security risks. This method 100 can determine the security event sequence and candidate node set of the topology graph based on the data security events that have occurred for each node in the topology graph. Then, by inputting the security event sequence into the preset security event prediction model for security event prediction, the data security risk values of each candidate node in the candidate node set can be obtained. Thus, the solution disclosed herein can accurately predict the node type of the next data security event and the specific data security event, and thereby predict the data security risk values of each candidate node directly adjacent to the current attack chain.

[0065] Figure 6 Shows Figure 1 An exemplary flowchart of step S101 in the method. It can be understood that the description below in combination with Figure 1 is a specific implementation of the aforementioned step S101. Therefore, the features described in the foregoing in combination with Figure 1 can be similarly applied herein.

[0066] As Figure 6 shown, at step S1011, method 100 can use a risk feature detection tool to detect the risk features of each node in the topology graph.

[0067] In the solution disclosed herein, the risk feature detection tool is a set of tools for checking the daily status of nodes within the domain. In one embodiment, the risk feature detection tool can include a vulnerability scanning tool, a code auditing tool, and a weak password scanning tool. The risk feature detection tool can update the risk features of each node in the graph database regularly according to the daily detection results and in real-time combination with the data security event records. In actual operation, there may be risk features that the risk feature detection tool cannot detect. In such cases, the risk features of the nodes can be determined through manual review and analysis.

[0068] In the solution disclosed herein, the risk characteristics of a node are the characteristics that make the node a target of an attacker and subject to attacks, and the risk characteristics of nodes of different node types can be different. As shown in Table 1 below, the risk characteristics of nodes of the personnel type can include not signing a confidentiality agreement, not receiving security training, and not participating in an admission investigation. The risk characteristics of nodes of the application system type can include passwords being stored in plain text, no multi-factor authentication, and no authentication for interface calls. The risk characteristics of nodes of the data type can include not being desensitized, no encryption, and no watermark. The risk characteristics of nodes of the cloud resource type can include having vulnerabilities, having weak passwords, and unrestricted opening of high-risk ports.

[0069] It should be understood that the risk characteristics shown in Table 1 are only exemplary and illustrative, and those skilled in the art can expand the risk characteristics of each node type according to actual needs, and the present disclosure does not make specific limitations in this regard. From an implementation perspective, the present disclosure can use a risk characteristic set to store the risk characteristics to be used. As an example, the risk characteristic set can be represented as {passwords are stored in plain text, not receiving security training, no authentication for interfaces, data not encrypted, having vulnerabilities, having weak passwords, unrestricted opening of high-risk ports}.

[0070] Table 1 Example Table of Risk Characteristics of Nodes of Each Type

[0071]

[0072] Next, in combination with an example, an exemplary introduction is given to determining the risk characteristics of each node in the topology diagram by using a risk characteristic detection tool and manual review and analysis.

[0073] First, after regular automated checks, the personnel management system issues a warning: Zhang San, an operations and maintenance personnel, has not received data security training organized by the enterprise after joining the company, which may lead to the risk of data leakage due to improper behavior. Therefore, the risk characteristic (not receiving security training) of the node (Zhang San, the personnel) in the graph database can be updated.

[0074] Second, since Zhang San has not received security training, he is unaware of the harm of storing passwords in plain text and mistakenly saves the operation and maintenance account password of the financial system in plain text in a shared document within the enterprise. Therefore, the risk characteristic (passwords are stored in plain text) of the node (application system financial system) in the graph database can be updated.

[0075] Then, the code audit tool scans and discovers that the performance system does not provide authentication capabilities for some interfaces, thus leaving a security vulnerability backdoor. Therefore, the risk characteristic (no authentication for interfaces) of the node (application system performance system) in the graph database can be updated.

[0076] Next, the weak password scanning tool issues a warning after periodic scanning: the ECS running the attendance system has a weak password that can be easily cracked, which may lead to the risk of unauthorized access to cloud resources. Therefore, the risk feature (presence of a weak password) of the node (ECS of the cloud resource running the attendance system) in the graph database can be updated.

[0077] Then, due to Zhang San mistakenly storing the account password of the RDS storing employee identity information in plain text in a shared document within the enterprise. Therefore, the risk feature (password stored in plain text) of the node (RDS of the cloud resource storing employee identity information) in the graph database can be updated.

[0078] Finally, the vulnerability scanning tool scans and discovers that most of the RDSs within the organization (including the RDS storing employee identity information) have security vulnerabilities due to the old version of the database engine, resulting in the risk of being exploited. Therefore, the risk feature (presence of vulnerabilities) of the node (RDS of the cloud resource storing employee identity information) in the graph database can be updated.

[0079] Subsequently, when it is necessary to analyze the risk features of nodes in the graph database, the following form can be used to determine whether a node has a certain risk feature:

[0080] (1)

[0081] As an example, has(Zhang San, not receiving security training) = 1, indicating that Zhang San has the risk feature of not receiving security training. As another example, has(ECS running the attendance system, password stored in plain text) = 0, indicating that the ECS running the attendance system does not have the risk feature of password stored in plain text.

[0082] Furthermore, when it is necessary to perform statistical analysis on the risk features of multiple nodes in the graph database, the quantitative form shown in Table 2 below can be used to represent whether multiple nodes have a certain risk feature. As can be seen from Table 2, the person Zhang San has the risk feature of not receiving security training; the application system, the financial system, has the risk feature of password stored in plain text; the application system, the performance system, has the risk feature of no authentication for the interface; the RDS of the cloud resource storing employee identity information has the risk features of password stored in plain text and presence of vulnerabilities; the ECS of the cloud resource running the attendance system has the risk feature of presence of a weak password.

[0083] Table 2 Quantification Table of Risk Features of Various Types of Nodes

[0084]

[0085] Continuing from the previous step S1011, at step S1012, method 100 can determine the possibility of a data security event occurring for each node based on the risk feature of each node.

[0086] It is understandable that there is an inevitable correlation between the risk characteristics of a node and the data security events that will occur at the node. In actual operation, the correlation between risk characteristics and data security events can be obtained through statistical analysis of historical data security event data, and a numerical value within the range of [0, 1] can be used to represent the degree of correlation between the two. In one embodiment, the correlation between risk characteristics and data security events can be represented in the following form:

[0087] (2)

[0088] As an example, , it indicates that the degree of correlation between the risk characteristic that the password is stored in plain text and the data security event that the password is leaked is 1. In other words, if a node has the risk characteristic that the password is stored in plain text, the possibility that the node will have a data security event of password leakage is 1, that is, the possibility of password leakage is very high.

[0089] As another example, , it indicates that the degree of correlation between the risk characteristic of having a vulnerability and the data security event of unauthorized access is 0.3. In other words, if a node has the risk characteristic of having a vulnerability, the possibility that the node will have a data security event of unauthorized access is 0.3, that is, the possibility of unauthorized access is not high.

[0090] As yet another example, , it indicates that the degree of correlation between the risk characteristic of not undergoing an admission investigation and the data security event of data tampering is 0. In other words, if a node has the risk characteristic of not undergoing an admission investigation, the possibility that the node will have a data security event of data tampering is 0, that is, there is no possibility of data tampering.

[0091] Subsequently, when it is necessary to conduct statistical analysis on the correlation between different risk characteristics and different data security events, the following quantitative form as shown in Table 3 can be used to represent the correlation between different risk characteristics and different data security events. From the implementation level, this disclosure can use a matrix to store the correlation between risk characteristics and data security events.

[0092] Table 3 Quantification Table of the Correlation between Risk Characteristics and Data Security Events

[0093]

[0094] Based on the above-mentioned correlation between the risk characteristics and data security events obtained through statistical analysis of historical data security event data, the possibility of each node having different data security events can be determined based on the risk characteristics of each node.

[0095] Furthermore, the method 100 can sort the probabilities of different data security events occurring at each node to determine the data security event corresponding to the maximum probability for each node. Thereafter, the method 100 can also output a likelihood event prompt in a visual or audible manner for the data security event corresponding to the maximum probability. Here, the likelihood event prompt is used to prompt the user of the data security event most likely to occur at each node, so that the user can make corresponding preventive preparations.

[0096] Figure 7 shows Figure 1 An exemplary flowchart of step S103 in the method. It can be understood that the description below in conjunction with Figure 1 is a specific implementation of the foregoing step S103. Therefore, the features described in conjunction with Figure 1 above can be similarly applied here.

[0097] As Figure 7 shown, at step S1031, the method 100 can input the security event sequence into a preset security event prediction model for security event prediction to obtain the probability of each node type being attacked and the probability of a data security event occurring when a node of each node type is attacked after the security event sequence occurs.

[0098] Taking the cloud resource type node as an example, the prediction results output by the preset security event prediction model are as follows:

[0099] P(cloud resource type node is attacked | data security event sequence occurs) = 0.35, indicating that the probability of the cloud resource type node being attacked is 0.35;

[0100] P(password is leaked | data security event sequence occurs, cloud resource type node is attacked) = 0.5, indicating that the probability of the password being leaked when the cloud resource type node is attacked is 0.5;

[0101] P(password is cracked | data security event sequence occurs, cloud resource type node is attacked) = 0.1, indicating that the probability of the password being cracked when the cloud resource type node is attacked is 0.1;

[0102] P(attack payload is deployed | data security event sequence occurs, cloud resource type node is attacked) = 0.35, indicating that the probability of an attack payload being deployed when the cloud resource type node is attacked is 0.35;

[0103] P(unauthorized access | data security event sequence occurs, cloud resource type node is attacked) = 0.05, indicating that the probability of unauthorized access occurring when the cloud resource type node is attacked is 0.05;

[0104] P(Other data security incidents occur | Data security incident sequence occurs, cloud resource type node is attacked) = 0, indicating that when a cloud resource type node is attacked, the probability of other data security incidents occurring in the cloud resource type node is 0.

[0105] It can be seen that for nodes of the same node type, the prediction results of the preset security event prediction model are the same. Then, for the two candidate nodes (the RDS node storing employee identity information and the ECS node running the attendance system) determined in step S102 above, since both candidate nodes are of the cloud resource type, the prediction results of the two nodes are the same.

[0106] In step S1032, method 100 can determine the data security risk value of the candidate node according to the node type of the candidate node, the risk characteristics of the candidate node, the possibility of data security incidents occurring in the candidate node, the possibility of each node type being attacked, and the probability of data security incidents occurring when each node type of node is attacked.

[0107] In actual operation, the following formula can be used to calculate the data security risk value of the candidate node directly adjacent to the end of the security event sequence when a security event sequence occurs:

[0108]

[0109] Among them, Indicates whether the candidate node has a certain risk characteristic; Indicates the relevance between a certain risk characteristic of the candidate node and a certain data security incident; Indicates the probability that the candidate node is attacked after the security event sequence occurs; Indicates the probability that the candidate node is attacked and a certain data security incident occurs after the security event sequence occurs, Indicates the data security risk value of the candidate node.

[0110] Next, continue to refer to Figure 1 , in the solution disclosed in this disclosure, after method 100 determines the data security risk value of the candidate node in step S103, steps S104 and S105 can also be executed.

[0111] Then, in step S104, method 100 can determine the value-weighted risk of each candidate node based on the data security value and the data security risk value of each candidate node. Specifically, the product of the data security value and the data security risk value can be used as the value-weighted risk of the candidate node. In actual operation, the following formula can be used to calculate the value-weighted risk of the candidate node.

[0112]

[0113] Among them, represents the value-weighted risk of the candidate node, represents the data security risk value of the candidate node, represents the data security value of the candidate node.

[0114] In the solution disclosed herein, the data security value of a candidate node can be a numerical value within the range of [0, 1]. Method 100 can determine the data security value of each candidate node according to the value factors of each candidate node. Specifically, the data security value of a candidate node can be calculated based on whether the candidate node meets the value factors. The value factors disclosed herein can include, but are not limited to: whether it is a data node or a node directly storing data, whether the amount of stored data is greater than a preset number of bytes, whether the classification of the stored data is at or above a preset classification, whether the stored data is generated or updated within a preset time limit, and whether the stored data is correlated with other data sets. It can be understood that those skilled in the art can set the specific values of the preset number of bytes, preset classification, and preset time limit according to actual needs, and the present disclosure does not make specific limitations thereto.

[0115] In one implementation scenario, the more value factors a candidate node meets, the higher the data security value of the candidate node can be determined. Conversely, the fewer value factors a candidate node meets, the lower the data security value of the candidate node can be determined. In this scenario, the data security value corresponding to each value factor can be preset, and the sum of the data security values corresponding to the multiple value factors met by the candidate node can be used as the final data security value of the candidate node. In one embodiment, the data security value of the candidate node can be represented in the following form:

[0116]

[0117] As an example, at the aforementioned step S102, the data security values of the two determined candidate nodes (the RDS node storing employee identity information and the ECS node running the attendance system) are respectively: V(RDS node storing employee identity information) = 1.0; V(ECS node running the attendance system) = 0.5. To facilitate understanding of how to determine the data security risk value and value-weighted risk of the candidate node, the calculation processes of the data security risk value and value-weighted risk of the aforementioned two candidate nodes (the RDS node storing employee identity information and the ECS node running the attendance system) will be introduced next.

[0118] From the foregoing Figure 3It can be known that the type of the RDS node storing employee identity information is cloud resources. As can be seen from the aforementioned Table 2, the RDS node storing employee identity information has risk characteristics of the password being stored in plain text and having vulnerabilities. As can be seen from the aforementioned Table 3, the relevance of the password being stored in plain text to each data security event in the data security event sequence, and the relevance of the storage vulnerability to each data security event in the data security event sequence. In addition, the data security key value of the RDS node storing employee identity information is V(RDS node storing employee identity information)=1.0.

[0119] Furthermore, from the aforementioned Figure 3 It can be known that the type of the ECS node running the attendance system is cloud resources. As can be seen from the aforementioned Table 2, the ECS node running the attendance system has the risk characteristic of having a weak password. As can be seen from the aforementioned Table 3, the relevance of having a weak password to each data security event in the data security event sequence. In addition, V(ECS node running the attendance system)=0.5.

[0120] Based on this, the calculation processes of the data security risk value and the value-weighted risk of the RDS node storing employee identity information are as follows:

[0121]

[0122] The calculation processes of the data security risk value and the value-weighted risk of the ECS node running the attendance system are as follows:

[0123]

[0124] Based on the above calculation results, it can be concluded that the risk of the RDS node storing employee identity information is higher than that of the ECS node running the attendance system. For the RDS node storing employee identity information, due to its risk characteristic of the password being stored in plain text, attention should be paid to preventing data security events of the password being leaked; due to its risk characteristic of having vulnerabilities, attention should be paid to preventing data security events of being loaded with attack payloads and unauthorized access; for the EC node running the attendance system, due to its risk characteristic of having a weak password, attention should be paid to preventing data security events of the password being cracked.

[0125] Finally, at step S105, method 100 can output event warnings and prevention tips outward according to the value-weighted risk. Here, the event warnings and prevention tips are used to prompt the user of the data security events that the candidate nodes need to prevent. Specifically, method 100 can output the event warnings and prevention tips in a visual or audible manner.

[0126] Furthermore, method 100 can also determine the early warning and prevention levels according to the value-weighted risk. In actual operation, method 100 can preset multiple early warning and prevention levels, and set the value range of the value-weighted risk corresponding to each early warning and prevention level. Thus, method 100 can determine the early warning and prevention level of the candidate node according to the value range where the value-weighted risk of the candidate node is located. It can be understood that those skilled in the art can set multiple early warning and prevention levels and the value range of the value-weighted risk corresponding to each early warning and prevention level according to actual needs, and this disclosure does not make specific limitations.

[0127] Next, in combination with Figure 8 an electronic device 800 provided by an embodiment of the present disclosure will be introduced exemplarily. As Figure 8 shown, the electronic device 800 of the embodiment of the present disclosure may include a processor 801, a memory 802, and a communication bus 803.

[0128] In the process of a specific embodiment, the above-mentioned processor 801 may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing image processing device (DSPD), a programmable logic image processing device (PLD), a field programmable gate array (FPGA), a CPU, a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above-mentioned processor functions may also be others, and this embodiment does not make specific limitations.

[0129] In the embodiment of the present disclosure, the above-mentioned communication bus 803 is used to realize the connection and communication between the processor 801 and the memory 802; the memory 802 stores program instructions for predicting data security risks; when the above-mentioned processor 801 executes the program instructions stored in the memory 802, it realizes the method for predicting data security risks described in the present disclosure in combination with the attached Figures 1 to 7 drawings.

[0130] The above in combination with Figure 8Devices that can be used to perform the prediction of data security risks for the present disclosure are described. It should be understood that the device structure or architecture here is only exemplary, and the implementation manner and implementation entity of the present application are not limited by it, but can be changed without departing from the spirit of the present application. It can be understood that the description of each embodiment in the present disclosure emphasizes the differences between the embodiments, and the same or corresponding parts can be referred to each other. For the sake of brevity, the present disclosure will not elaborate one by one.

[0131] According to the above description in conjunction with the drawings, those skilled in the art can also understand that the embodiments of the present application can also be implemented through software programs. Accordingly, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores program instructions for predicting data security risks, and the program instructions can be used to implement the method for predicting data security risks described in the present disclosure in conjunction with the attached Figures 1 to 7 drawings.

[0132] It should be noted that although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0133] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, changes, and alternative ways can be conceived by those skilled in the art without departing from the spirit and idea of the present disclosure. It should be understood that various alternative solutions to the embodiments of the present disclosure described herein can be adopted in the practice of the present disclosure. The appended claims are intended to define the scope of protection of the present disclosure and thus cover equivalents or alternatives within the scope of these claims.

[0134] The collection and acquisition of various data in the present application comply with relevant laws and regulations and are authorized by the data providers. Any organization or individual that needs to obtain external data shall obtain authorization in accordance with the law and ensure data security, and shall not illegally collect, use, process, or transmit unauthorized or unprotected data, and shall not illegally buy, sell, provide, or disclose unauthorized or unprotected data.

Claims

1. A method for predicting data security risks, comprising: Constructing a topology graph among the nodes of the target object according to the node types, wherein the topology graph includes horizontal associations among nodes of the same node type and vertical associations among nodes of different node types; Determining a security event sequence and a candidate node set of the topology graph according to the data security events that have occurred for each node in the topology graph, wherein the candidate node is a node in the topology graph that is directly adjacent to the end node of the security event sequence and has not had a data security event occur; and Inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of the candidate nodes in the candidate node set; Wherein, inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the data security risk values of the candidate nodes in the candidate node set includes: Inputting the security event sequence into a preset security event prediction model for security event prediction to obtain the possibility of each node type being attacked and the probability of a data security event occurring when a node of each node type is attacked after the security event sequence occurs; and Determining the data security risk value of the candidate node according to the node type of the candidate node, the risk characteristics of the candidate node, the possibility of the candidate node having a data security event occur, the possibility of each node type being attacked, and the probability of a data security event occurring when a node of each node type is attacked.

2. The method according to claim 1, wherein The node types include personnel, data, application systems, and cloud resources.

3. The method according to claim 2, wherein Constructing a topology graph among the nodes of the target object according to the node types includes: Detecting the risk characteristics of each node in the topology graph by using a risk characteristic detection tool; and Based on the risk characteristics, determining the possibility of each node having a data security event occur.

4. According to the method of claim 1, after obtaining the data security risk values of the candidate nodes in the candidate node set, the method further includes: Determining the value-weighted risk of the candidate nodes based on the data security value and the data security risk value of the candidate nodes; And Outputting an event warning and a prevention prompt outward according to the value-weighted risk.

5. The method according to claim 4, wherein, The data security value is determined according to the value factors of the candidate node; the value factors include: whether it is a data node or a node directly storing data, whether the amount of stored data is greater than 1000, whether the classification of the stored data is at or above a preset classification, whether the stored data is generated or updated within a preset time limit, and whether the stored data is mutually associated with other data sets.

6. According to the method of claim 3, after determining the possibility of each node having a data security event occur, the method further includes: Sorting the possibility of each node having a data security event occur to determine the data security event corresponding to the maximum possibility value of each node; And Outputting a possibility event prompt outward for the data security event corresponding to the maximum possibility value.

7. According to the method according to any one of claims 1-6, wherein The preset security event prediction model is a model obtained by training using a machine learning model. Training the preset security event prediction model using a machine learning model includes: Collecting and analyzing historical security event data of the training object to obtain the security event sequence of the training object and the information on induced security events. Among them, the information on induced security events includes the node type attacked after the security event sequence and the data security event that occurred to the attacked node; and Using the security event sequence and the information on induced security events as training data and inputting them into the machine learning model to train the machine learning model to obtain the preset security event prediction model.

8. An electronic device, comprising: A processor; And A memory storing program instructions for predicting data security risks. When the program instructions are run by the processor, the electronic device is caused to execute the method according to any one of claims 1-7.

9. A computer-readable storage medium storing program instructions for predicting data security risks. When the program instructions are executed by a processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Log detection method, system, device and medium

    CN112395159A

  • Intelligent economic risk identification method and system

    CN118096392A