Network data asset ownership identification method and system based on Bayesian causal reasoning

By applying Bayesian causal reasoning method in network data asset attribution recognition, a causal Bayesian network is constructed, and the problem of machine learning and deep learning models focusing on correlation rather than causality is solved, achieving efficient and highly interpretable network data asset attribution recognition.

CN119496645BActive Publication Date: 2025-05-16JINQICHUANG (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411610797.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-05-16
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

Machine learning and deep learning models focus on the correlation between features and target variables in network data asset attribution recognition, rather than causality, making the model difficult to interpret and perform weakly in dynamically changing network environments, requiring frequent retraining, resulting in high computational and time costs.

Method used

Using Bayesian causal reasoning method, by obtaining network data asset data sets, extracting asset characteristics, building Bayesian networks, using chi-square test and point-to-point conditional causality, dynamically update network structure and parameters, and reduce the frequency of retraining.

Benefits of technology

It realizes the causal relationship between features and targets in the attribution identification of network data assets. The model output has a clear causal chain, which is easy to explain and track, and can adapt to changes in the data environment, reduce calculation and time costs, and improve the efficiency and accuracy of network data asset management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119496645B_ABST
    Figure CN119496645B_ABST
Patent Text Reader

Abstract

The present invention provides a network data asset attribution identification method and system based on Bayesian causal reasoning, and relates to the field of data processing technology. The method includes: obtaining a network data asset data set; using heuristic rules, preferentially selecting asset features with high significance and high confidence as preliminary nodes of the Bayesian network; constructing an undirected preliminary Bayesian network based on the preliminary nodes; adding the remaining asset features as new nodes to the undirected preliminary Bayesian network to form an undirected intermediate Bayesian network; based on the point-to-point conditional causality of the node pairs composed of two parent nodes in the potential causal structure, determining the direction of the undirected edge, and converting the undirected intermediate Bayesian network into a directed final Bayesian network; obtaining the network data asset to be identified; based on the Bayesian theorem, calculating the posterior probability of each category label under a given feature combination; taking the category with the largest posterior probability as the final category of the network data asset to be identified, and outputting it.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a network data asset attribution identification method and system based on Bayesian causal reasoning. Background Art

[0002] The significance of network data asset ownership identification is to improve asset management, security and compliance in the network environment. Accurately identifying the ownership of assets in the network can help organizations discover and protect key assets and prevent them from network attacks. By identifying the ownership of data assets, you can more clearly understand the relationship between different assets, which helps to identify potential security threats and take timely protective measures.

[0003] With the rapid development of science and technology and the urgent need for network data asset attribution identification, more and more modern technologies are being applied to network data asset attribution identification. Currently, network data asset attribution identification mainly includes attribution identification based on machine learning and attribution identification based on deep learning.

[0004] However, machine learning and deep learning models mainly focus on the correlation between features and target variables, rather than causal relationships. These models are often "black box" in nature, and it is difficult to explain the specific reasons for the model output. At the same time, they are weak in dealing with dynamically changing network environments, especially when network assets are constantly changing. Models need to be retrained frequently to adapt to new data, which results in high computational and time costs. Summary of the invention

[0005] In order to solve the technical problem that machine learning and deep learning models mainly focus on the correlation between features and target variables rather than causal relationships, these models are often "black box" in nature, it is difficult to explain the specific reasons for the model output, and they perform poorly when dealing with dynamically changing network environments, especially when network assets are constantly changing, and the models need to be frequently retrained to adapt to new data, which results in high computational and time costs, the present invention provides a network data asset attribution identification method and system based on Bayesian causal reasoning.

[0006] The technical solution provided by the embodiment of the present invention is as follows:

[0007] First aspect:

[0008] An embodiment of the present invention provides a network data asset attribution identification method based on Bayesian causal reasoning, comprising:

[0009] S1: Obtain network data asset dataset;

[0010] S2: extracting multiple asset features of the network data asset;

[0011] S3: Using heuristic rules, asset features with high significance and confidence are selected as the initial nodes of the Bayesian network;

[0012] S4: constructing an undirected preliminary Bayesian network according to the preliminary nodes based on the selection principle of maximizing the log-likelihood increment;

[0013] S5: adding the remaining asset features as new nodes one by one to the undirected preliminary Bayesian network;

[0014] S6: Using the chi-square test, the position of the new node in the network is determined according to the causal relationship between the new node and the existing nodes, forming an undirected intermediate Bayesian network;

[0015] S7: Extracting a latent causal structure consisting of two parent nodes and a common child node from an undirected intermediate Bayesian network;

[0016] S8: based on the point-to-point conditional causality of the node pairs consisting of two parent nodes in the potential causal structure, determining the direction of the undirected edge, and converting the undirected intermediate Bayesian network into a directed final Bayesian network;

[0017] S9: Obtain the network data assets to be identified;

[0018] S10: extracting multiple asset features of the network data asset to be identified;

[0019] S11: using the chain rule, according to the topological structure of the final Bayesian network, calculating the joint probability under the combination of asset features;

[0020] S12: Based on Bayes’ theorem, the posterior probability of each category label under a given feature combination is calculated;

[0021] S13: The category with the largest a posteriori probability is taken as the final category of the network data asset to be identified and outputted.

[0022] Second aspect:

[0023] An embodiment of the present invention provides a network data asset attribution identification system based on Bayesian causal reasoning, comprising:

[0024] processor;

[0025] A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the network data asset ownership identification method based on Bayesian causal reasoning as described in the first aspect is implemented.

[0026] The third aspect:

[0027] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method for identifying network data asset ownership based on Bayesian causal reasoning as described in the first aspect is implemented.

[0028] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0029] In the present invention, the posterior probability of each category label under a given feature combination is calculated based on the Bayesian theorem, and the category with the largest posterior probability is taken as the final category of the network data asset to be identified. The network data asset attribution identification is completed through Bayesian causal reasoning. Bayesian causal reasoning can not only capture the correlation between features and targets, but also pay more attention to the causal relationship between features. The results output by the model have a clear causal chain, which is easy to explain and track. The Bayesian causal reasoning model can dynamically update the network structure and parameters without frequently retraining the entire model. It can adapt to changes in the data environment, reduce computing and time costs, and help improve the overall efficiency and accuracy of network data asset management. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0031] Figure 1 A flowchart of a network data asset ownership identification method based on Bayesian causal reasoning provided by an embodiment of the present invention;

[0032] Figure 2 A schematic diagram of the structure of a network data asset attribution identification system based on Bayesian causal reasoning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0033] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0034] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0035] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0036] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.

[0037] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0038] Reference Manual Attached Figure 1 , showing a flow chart of a network data asset ownership identification method based on Bayesian causal reasoning provided by an embodiment of the present invention.

[0039] The embodiment of the present invention provides a network data asset attribution identification method based on Bayesian causal reasoning, which can be implemented by a network data asset attribution identification device based on Bayesian causal reasoning, and the network data asset attribution identification device based on Bayesian causal reasoning can be a terminal or a server. The processing flow of the network data asset attribution identification method based on Bayesian causal reasoning can include the following steps:

[0040] S1: Get the network data asset dataset.

[0041] S2: Extract multiple asset features of network data assets.

[0042] Among them, asset characteristics include: IP address, MAC address, host name, port information, packet size, packet frequency, access frequency, common access sources, transmission protocol, number of connections, running services, access control list, authentication method, encryption method, last update date, active period and last activity time.

[0043] Among them, the IP address is one of the unique identifiers of assets in the network, which can help identify the location of assets in the network and divide different network areas (such as internal network, external network). Through the analysis of IP addresses, it is possible to infer the network segment or geographical location to which the asset belongs.

[0044] Among them, the MAC address is the physical address of the device, which can uniquely identify the network interface. The MAC address is usually associated with the device manufacturer and helps to determine the brand or device type of the asset. It is particularly important for the identification of intranet devices and is suitable for the location and tracking of assets in the local area network.

[0045] Among them, the host name can usually provide basic usage or role information of the asset (such as web server, database server), helping to identify the function and department of the asset. Especially in a well-managed network, the host name can directly reflect the usage and environment of the device.

[0046] Among them, the open port reflects the type of service provided by the device (such as port 80 for HTTP service and port 443 for HTTPS). By analyzing the port, you can infer the business type and service function of the device and help distinguish different types of servers or application devices.

[0047] The size of the data packet can reflect the communication characteristics of the asset. For example, a large data packet may be used for file transfer services, while a small data packet may be used for status updates. This feature helps to determine the communication mode of the device and thus infer its business type.

[0048] The packet frequency reflects the activity frequency and traffic characteristics of the device, and a high frequency usually indicates an active device or service. By analyzing the packet frequency, the usage intensity of the device can be identified to help determine its role (such as an active database server or an idle device).

[0049] Among them, access frequency indicates how often the device is accessed by other network assets. A high access frequency usually means that the device is a critical resource (such as a file server or database server), which helps to identify its importance and business role in the network.

[0050] Among them, common access sources can help identify the user groups or dependencies of assets. For example, frequent access from a specific IP address can indicate the department or business system using the asset, and access source analysis can help with attribution identification and security control.

[0051] The transmission protocol (such as TCP, UDP, ICMP, etc.) determines the communication mode of the device. The use of a specific protocol often corresponds to a specific service or device type. For example, UDP is mostly used for streaming media transmission. The service function of the device can be further inferred by the protocol type.

[0052] The number of connections indicates the number of connections to a device within a specific period of time, which can reflect the device's workload and network interaction frequency. A higher number of connections usually indicates that the device is a core resource with high traffic, such as a web server or load balancer.

[0053] Among them, the services running on the device (such as HTTP service, database service, FTP service) directly determine its business functions. By identifying the running services of the device, its role in the network can be clarified, which is an important basis for attribution identification.

[0054] The access control list defines the IP range or user group that is allowed or denied access. ACL can help determine the security level of assets and the user population, thereby inferring their business category and application environment (such as whether it is a publicly available device).

[0055] Among them, the authentication method adopted by the device (such as username / password, two-factor authentication) can reflect its security requirements and protection level. Advanced authentication methods are usually applicable to important assets to help identify key systems or sensitive equipment.

[0056] Among them, data encryption methods (such as SSL / TLS, IPSec) reflect the data protection strategy of the device. Important business assets usually enable encrypted transmission, and the sensitivity and importance of the device can be inferred through the encryption method.

[0057] The most recent update date of a device can reflect its maintenance and security status. Frequently updated devices are usually active or valued resources, which helps determine their usage and importance in the network.

[0058] The active cycle represents the typical active time of the device (e.g., active on weekdays and idle at night). Specific active patterns help identify the business scenarios of assets (e.g., office equipment, server equipment), making it easier to categorize equipment.

[0059] Among them, the last activity time of the device can help determine whether it is online or idle. For ownership identification, it can distinguish between assets currently in use and assets that have been inactive for a long time, which facilitates resource cleanup and management.

[0060] S3: Using heuristic rules, asset features with high significance and confidence are preferentially selected as the initial nodes of the Bayesian network.

[0061] In a possible implementation, S3 specifically includes sub-steps S301 to S304:

[0062] S301: Based on information gain, calculate the significance of each asset feature:

[0063] s(X)=H(Y)-H(Y|X)

[0064] Where s represents significance, X represents the target asset feature, s(X) represents the significance of the target asset feature X, Y represents the category label, H(Y) represents the information entropy of the category label Y, and H(Y|X) represents the conditional entropy of the category label Y given the target asset feature X.

[0065] It should be noted that significance is used to measure the correlation or association between a feature and the target variable (category label). The higher the significance of a feature, the stronger its predictive ability for the target variable and the greater its contribution to classification or recognition.

[0066] S302: Based on Bayesian probability, calculate the confidence level of each asset feature:

[0067]

[0068] Among them, conf represents confidence, conf(X) represents the confidence of the target asset feature X, a represents correctness, ei represents the i-th network data asset sample, P(a|X,ei) represents the correctness probability of the target asset feature X under the i-th network data asset sample, and n represents the total number of network data asset samples.

[0069] Alternatively, if a feature is correct, it means that the feature contributes positively to the prediction or classification task, and conversely, if a feature is correct, it means that the feature does not contribute positively to the prediction or classification task.

[0070] Furthermore, how to judge whether the feature is correct? Use statistical methods (such as information gain, mutual information, correlation coefficient, etc.) to judge the correlation between the feature and the target variable. If the significance of the feature is high, the feature is considered to be correct. The contribution of the feature can also be judged through model experiments. For example, by adding the feature to the model, observe whether the accuracy or prediction performance of the model is improved. If the addition of the feature improves the model effect, it can be considered that the feature is correct.

[0071] It should be noted that confidence indicates the reliability of a feature in a prediction or classification task, and is usually used to evaluate the accuracy or stability of the feature in actual data. Features with higher confidence have more stable performance in different data samples, which helps improve classification consistency.

[0072] S303: Calculate the priority selection index based on the significance and confidence of each asset feature:

[0073] b(X)=βs(X)+(1-β)conf(X)

[0074] Where b represents the priority selection index, b(X) represents the priority selection index of the target asset feature X, s(X) represents the significance of the target asset feature X, conf(X) represents the confidence of the target asset feature X, and β represents the significance weight coefficient.

[0075] Among them, those skilled in the art can set the size of the significance weight coefficient β according to actual conditions, and the present invention does not limit this.

[0076] S304: Sort each asset feature in descending order of priority index, and select a preset number of asset features with the highest ranking as preliminary nodes of the Bayesian network.

[0077] Among them, those skilled in the art can set the size of the preset number according to actual conditions, and the present invention does not limit it. Optionally, the preset number is 5.

[0078] In the present invention, the heuristic rules are used to select features with high significance and confidence as preliminary nodes, which not only improves the accuracy and stability of the model, but also reduces the computational cost and enhances the interpretability and generalization ability of the model. This method takes into account both efficiency and effect, making the Bayesian network more efficient, accurate and robust when processing network data asset attribution identification tasks.

[0079] S4: Based on the preliminary nodes, an undirected preliminary Bayesian network is constructed based on the selection principle of maximizing the log-likelihood increment.

[0080] In a possible implementation, S4 specifically includes:

[0081] S401: Combine preliminary nodes to form multiple groups of node pairs.

[0082] S402: Calculate the independent log-likelihood of the joint probability of each group of node pairs when there is no independent undirected edge between two nodes.

[0083] Among them, log-likelihood is a very important concept in statistics, especially in maximum likelihood estimation (MLE) and Bayesian network construction. The log-likelihood function is usually used to evaluate the probability of a certain model parameter under given data. In Bayesian networks, log-likelihood is used to evaluate the strength of association between nodes. It should be noted that the calculation method of log-likelihood is a mature prior art, and the present invention will not repeat it in detail.

[0084] S403: Calculate the associated log-likelihood of the joint probability of each group of node pairs when there is an associated undirected edge between two nodes.

[0085] S404: Subtract the associated log-likelihood from the independent log-likelihood to calculate the log-likelihood increment of each group of node pairs.

[0086] S405: Determine whether the log-likelihood increment of each group of node pairs is greater than the increment threshold. If so, establish undirected edges between the node pairs to construct an undirected preliminary Bayesian network.

[0087] Among them, those skilled in the art can set the size of the increment threshold according to actual conditions, and the present invention does not limit it.

[0088] In the present invention, by calculating the log-likelihood increment, those connections that really contribute to the model can be effectively screened out. If the log-likelihood increment between two nodes is not large, it means that the association between the two nodes is weak and there is no need to establish a connection. This helps prevent the model from being too complicated and avoids the problem of overfitting. At the same time, reducing unnecessary connections can make the model more concise, easy to understand and explain. This is very beneficial for model maintenance and debugging in practical applications.

[0089] S5: Add the remaining asset features as new nodes one by one to the undirected preliminary Bayesian network.

[0090] S6: Using the chi-square test, the position of the new node in the network is determined according to the causal relationship between the new node and the existing nodes, forming an undirected intermediate Bayesian network.

[0091] Among them, the Chi-squared test is a statistical method used to test whether there is a significant association between two categorical variables.

[0092] In a possible implementation, S6 specifically includes sub-steps S601 to S604:

[0093] S601: When a new node is added to the Bayesian network, the chi-square statistic between the new node and other nodes is calculated:

[0094]

[0095] Among them, X 2 represents the chi-square statistic, Q ab represents the actual observation data collected between node a and node b, T ab represents the theoretical data between node a and node b, r a represents the state number of node a, r b Indicates the state number of node b.

[0096] S602: According to the set significance level and degrees of freedom, query the chi-square distribution table to determine the chi-square threshold.

[0097] The degrees of freedom are:

[0098] f=(r a -1)(r b -1)

[0099] Where f represents the degree of freedom, r a represents the state number of node a, r b Indicates the state number of node b.

[0100] It should be noted that setting the significance level can control the false positive rate in the chi-square test and ensure that the correct judgment is made within an acceptable probability range. This can effectively reduce misjudgments, that is, avoid mistaking the existence of a relationship between nodes when there is actually no causal relationship. At the same time, by using different degrees of freedom, the sensitivity of the test can be adjusted dynamically. Nodes with more states can capture the differences in data distribution more finely, while nodes with fewer states can avoid overfitting.

[0101] S603: Compare the chi-square statistic with the chi-square threshold. When the chi-square statistic is greater than or equal to the chi-square threshold, it indicates that there is a causal relationship between the nodes. When the chi-square statistic is less than the chi-square threshold, it indicates that there is no causal relationship between the nodes.

[0102] Among them, those skilled in the art can set the size of the chi-square threshold according to actual conditions, and the present invention does not limit it.

[0103] S604: Determine the position of the new node in the network according to the causal relationship between the new node and the existing nodes. When there is no causal relationship between the new node and any existing node, refuse to add the node to the undirected preliminary Bayesian network.

[0104] In the present invention, the risk of erroneously adding irrelevant nodes to the network can be effectively reduced by using the chi-square test. Only statistically significant causal relationships are included in the model, thereby improving the accuracy and reliability of the model.

[0105] S7: Extracting a latent causal structure consisting of two parent nodes and a common child node from an undirected intermediate Bayesian network.

[0106] S8: Based on the point-to-point conditional causality of node pairs consisting of two parent nodes in the potential causal structure, the direction of the undirected edge is determined, and the undirected intermediate Bayesian network is converted into a directed final Bayesian network.

[0107] Point-to-Point Conditional Causality is a method for evaluating the causal relationship between two variables given other variables. In Bayesian networks, point-to-point conditional causality is often used to determine the causal direction between nodes, thereby converting an undirected Bayesian network into a directed Bayesian network.

[0108] It should be noted that by calculating point-to-point conditional causality, the causal strength between two parent nodes can be determined, thereby determining who is the cause and who is the effect. Conditional causality calculation is based on the degree of information change of the node under given conditions, and can accurately identify the causal relationship between node pairs.

[0109] In a possible implementation, S8 specifically includes:

[0110] S801: Calculate the point-to-point conditional causality of a node pair consisting of two parent nodes in the potential causal structure:

[0111] PC(X i →X j |Y)=H(X j |Y)-H(X j |X i ,Y)

[0112] Among them, X i represents the i-th node, X j represents the jth node, Y represents the category label, PC(X i →X j |Y) means that the i-th node X is given a class label Y. i For the jth node X j The conditional causality of H(X j |Y) indicates that the jth node X is given a class label Y. j The conditional entropy, H(X j |X i ,Y) means that given the category label Y and the i-th node X i In the case of the jth node X j The conditional entropy of .

[0113] PC(X j →X i |Y)=H(X i |Y)-H(X i |X j ,Y)

[0114] Among them, PC(X j →X i |Y) indicates that the jth node X is given a class label Y. j For the i-th node X i The conditional causality of H(X i |Y) means that the i-th node X is given a class label Y. i The conditional entropy, H(X i |X j ,Y) means that given the category label Y and the jth node X j In the case of the i-th node X i The conditional entropy of .

[0115] S802: Determine the i-th node X given the category label Y i For the jth node X jThe conditional causality of the j-th node X j For the i-th node X i Are the conditional causality of all nodes less than the conditional causality threshold? If so, delete the i-th node X i For the jth node X j Otherwise, go to the next step.

[0116] S803: Determine the i-th node X given the category label Y i For the jth node X j Is the conditional causality of greater than the jth node X j For the i-th node X i If so, add the conditional causality from the i-th node X i To the jth node X j Otherwise, add a directed edge from the jth node X i To the i-th node X j The directed edges of are added to the intermediate Bayesian network, converting the undirected intermediate Bayesian network into a directed final Bayesian network.

[0117] In the present invention, the direction of the undirected edge is determined by point-to-point conditional causality, and the undirected network is converted into a directed final Bayesian network, which not only improves the accuracy of causal relationship identification, but also optimizes the network structure, improves the interpretability and reasoning efficiency of the model. This method is based on statistical causal reasoning, ensuring that each edge in the Bayesian network has a clear causal directionality, making the model more robust and efficient in a changing network environment.

[0118] In a possible implementation, after S8 and before S11, the process further includes: using the attribution identification result of the historical network data assets as evidence, and constructing a conditional probability table through an optimization algorithm.

[0119] It should be noted that after the conditional probability table is constructed, the conditional probability value can be directly obtained from the conditional probability table when the chain rule and Bayes' theorem are used later.

[0120] Optionally, the optimization algorithm may specifically be an expectation maximization (EM) algorithm, a genetic algorithm, a particle swarm optimization algorithm, a simulated annealing algorithm, or the like.

[0121] In the present invention, historical data reflects the asset types and behavior patterns in the network environment, and the optimization algorithm can extract representative probability distributions from them, thereby reducing the model's dependence on current data and avoiding overfitting. This makes the model more adaptable when processing new data or unseen asset types. As new historical data is continuously added, the conditional probability table can be regularly updated through the optimization algorithm to ensure that the conditional probability reflects the latest network data asset characteristics. This dynamic update mechanism helps the model automatically adapt to new trends when the network environment changes.

[0122] S9: Obtain the network data assets to be identified.

[0123] S10: Extract multiple asset features of the network data assets to be identified.

[0124] S11: Use the chain rule to calculate the joint probability under the combination of asset features based on the topological structure of the final Bayesian network.

[0125] In a possible implementation, S11 specifically includes: using the chain rule to calculate the joint probability under the asset feature combination according to the following formula:

[0126]

[0127] Among them, x represents the asset feature combination. There are n asset features in the asset feature combination, namely x1, x2, …, x n , y represents the category label, P(x,y) represents the joint probability distribution of the asset feature combination x and the category label y, P(y) represents the prior probability of the category label y, P(x i |x1,x2,…,x i-1 ,y) means that given the category label y and all previous features x1,x2,…,x i-1 In the case of x i The conditional probability of occurrence.

[0128] S12: Based on Bayes’ theorem, the posterior probability of each category label under a given feature combination is calculated.

[0129] Among them, Bayes' Theorem is a basic formula in probability theory, which is used to describe the probability of an event under known conditions, especially how to update the probability of an event when the conditions change. Bayes' Theorem is widely used in machine learning, statistical inference and causal reasoning, especially for reasoning and decision-making under uncertain conditions.

[0130] In a possible implementation, S12 specifically includes: based on Bayes' theorem, according to the following formula, calculating the posterior probability of each category label under a given feature combination:

[0131]

[0132] Among them, P(y|x) represents the posterior probability of category label y under asset feature combination x, and P(x) represents the marginal probability of asset feature combination x.

[0133] S13: The category with the largest a posteriori probability is taken as the final category of the network data asset to be identified and output.

[0134] In the present invention, the chain rule is used to decompose the joint probability and the Bayesian theorem is used to calculate the posterior probability. This method not only improves the accuracy of classification, but also enhances the efficiency, interpretability and robustness of the model. The category with the largest posterior probability is selected as the final classification result, which makes this method perform well in complex network data asset attribution identification and provides an efficient and reliable solution for practical applications.

[0135] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0136] In the present invention, the posterior probability of each category label under a given feature combination is calculated based on the Bayesian theorem, and the category with the largest posterior probability is taken as the final category of the network data asset to be identified. The network data asset attribution identification is completed through Bayesian causal reasoning. Bayesian causal reasoning can not only capture the correlation between features and targets, but also pay more attention to the causal relationship between features. The results output by the model have a clear causal chain, which is easy to explain and track. The Bayesian causal reasoning model can dynamically update the network structure and parameters without frequently retraining the entire model. It can adapt to changes in the data environment, reduce computing and time costs, and help improve the overall efficiency and accuracy of network data asset management.

[0137] Reference Manual Attached Figure 2 , showing a structural schematic diagram of a network data asset attribution identification system based on Bayesian causal reasoning provided by the present invention.

[0138] The present invention also provides a network data asset attribution identification system 20 based on Bayesian causal reasoning, which is applied to the above-mentioned network data asset attribution identification method based on Bayesian causal reasoning, and includes:

[0139] Processor 201.

[0140] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the network data asset attribution identification method based on Bayesian causal reasoning as in the method embodiment is implemented.

[0141] The network data asset attribution identification system 20 based on Bayesian causal reasoning provided by the present invention can execute the above-mentioned network data asset attribution identification method based on Bayesian causal reasoning and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.

[0142] In the present invention, the posterior probability of each category label under a given feature combination is calculated based on the Bayesian theorem, and the category with the largest posterior probability is taken as the final category of the network data asset to be identified. The network data asset attribution identification is completed through Bayesian causal reasoning. Bayesian causal reasoning can not only capture the correlation between features and targets, but also pay more attention to the causal relationship between features. The results output by the model have a clear causal chain, which is easy to explain and track. The Bayesian causal reasoning model can dynamically update the network structure and parameters without frequently retraining the entire model. It can adapt to changes in the data environment, reduce computing and time costs, and help improve the overall efficiency and accuracy of network data asset management.

[0143] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0144] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0145] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0146] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0147] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0148] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0149] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0150] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0151] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0152] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0153] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0154] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0155] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for identifying the ownership of network data assets based on Bayesian causal reasoning as described in the method embodiment is implemented.

[0156] A computer-readable storage medium provided by the present invention can implement the steps and effects of the network data asset ownership identification method based on Bayesian causal reasoning in the above method embodiment. To avoid repetition, the present invention will not go into details.

[0157] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0158] There are a few points to note:

[0159] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.

[0160] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.

[0161] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.

[0162] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A network data asset ownership identification method based on Bayesian causal reasoning, characterized in that: include: S1: Obtain network data asset dataset; S2: extracting multiple asset features of the network data asset; S3: Using heuristic rules, asset features with high significance and confidence are selected as the initial nodes of the Bayesian network; S4: constructing an undirected preliminary Bayesian network according to the preliminary nodes based on the selection principle of maximizing the log-likelihood increment; S5: adding the remaining asset features as new nodes one by one to the undirected preliminary Bayesian network; S6: Using the chi-square test, the position of the new node in the network is determined according to the causal relationship between the new node and the existing nodes, forming an undirected intermediate Bayesian network; S7: Extracting a latent causal structure consisting of two parent nodes and a common child node from an undirected intermediate Bayesian network; S8: based on the point-to-point conditional causality of the node pairs consisting of two parent nodes in the potential causal structure, determining the direction of the undirected edge, and converting the undirected intermediate Bayesian network into a directed final Bayesian network; S9: Obtain the network data assets to be identified; S10: extracting multiple asset features of the network data asset to be identified; S11: using the chain rule, according to the topological structure of the final Bayesian network, calculating the joint probability under the combination of asset features; S12: Based on Bayes’ theorem, the posterior probability of each category label under a given feature combination is calculated; S13: The category with the largest a posteriori probability is taken as the final category of the network data asset to be identified and outputted.

2. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 1 is characterized in that: The asset characteristics include: IP address, MAC address, host name, port information, data packet size, data packet frequency, access frequency, common access sources, transmission protocol, number of connections, running services, access control list, authentication method, encryption method, last update date, active period and last activity time.

3. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 1 is characterized in that: The S3 specifically includes: S301: Calculate the significance of each asset feature based on information gain; S302: Calculate the confidence level of each asset feature based on Bayesian probability; S303: Calculate a priority selection index based on the significance and confidence of each asset feature; S304: Sort the asset features in descending order according to the priority index, and select a preset number of asset features with the highest ranking as preliminary nodes of the Bayesian network.

4. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 1 is characterized in that: The S4 specifically includes: S401: Combining the preliminary nodes to form multiple groups of node pairs; S402: Calculate the independent log-likelihood of the joint probability of each group of node pairs when there is no independent undirected edge between the two nodes; S403: Calculate the associated log-likelihood of the joint probability of each group of node pairs when there is an associated undirected edge between two nodes; S404: subtract the associated log-likelihood from the independent log-likelihood to calculate the log-likelihood increment of each group of node pairs; S405: Determine whether the log-likelihood increment of each group of node pairs is greater than the increment threshold; if so, establish undirected edges between the node pairs to construct an undirected preliminary Bayesian network.

5. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 1 is characterized in that: The S6 specifically includes: S601: When a new node is added to the Bayesian network, a chi-square statistic between the new node and other nodes is calculated; S602: According to the set significance level and degree of freedom, query the chi-square distribution table to determine the chi-square threshold; S603: Compare the chi-square statistic with the chi-square threshold. When the chi-square statistic is greater than or equal to the chi-square threshold, it indicates that there is a causal relationship between the nodes. When the chi-square statistic is less than the chi-square threshold, it indicates that there is no causal relationship between the nodes. S604: Determine the position of the new node in the network according to the causal relationship between the new node and the existing nodes, and refuse to add the node to the undirected preliminary Bayesian network when there is no causal relationship between the new node and any existing node.

6. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 5 is characterized in that: The S8 specifically includes: S801: Calculate the point-to-point conditional causality of a node pair consisting of two parent nodes in the potential causal structure: PC(X i →X j |Y)=H(X j |Y)-H(X j |X i ,Y) Among them, X i represents the i-th node, X j represents the jth node, Y represents the category label, PC(X i →X j |Y) means that the i-th node X is given a class label Y. i For the jth node X j The conditional causality of H(X j |Y) means that the jth node X is given a class label Y. j The conditional entropy, H(X j |X i ,Y) means that given the category label Y and the i-th node X i In the case of the jth node X j The conditional entropy of PC(X j →X i |Y)=H(X i |Y)-H(X i |X j ,Y) Among them, PC(X j →X i |Y) means that the jth node X is given a class label Y. j For the i-th node X i The conditional causality of H(X i |Y) means that the i-th node X is given a class label Y. i The conditional entropy, H(X i |X j ,Y) means that given the category label Y and the jth node X j In the case of the i-th node X i The conditional entropy of S802: Determine the i-th node X given the category label Y i For the jth node X j The conditional causality of the j-th node X j For the i-th node X i Are the conditional causality of all nodes less than the conditional causality threshold? If so, delete the i-th node X i For the jth node X j Undirected edges between them; otherwise, go to the next step; S803: Determine the i-th node X given the category label Y i For the jth node X j Is the conditional causality of greater than the jth node X j For the i-th node X i Conditional causality of; if so, add the node X from the i-th node i To the jth node X j to the intermediate Bayesian network; otherwise, add a directed edge from the jth node X i To the i-th node X j The directed edges are added to the intermediate Bayesian network, converting the undirected intermediate Bayesian network into a directed final Bayesian network.

7. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 5 is characterized in that: The S11 is specifically: Using the chain rule, we can calculate the joint probability under the asset feature combination according to the following formula: Among them, x represents the asset feature combination. There are n asset features in the asset feature combination, namely x1, x2, …, x n , y represents the category label, P(x,y) represents the joint probability distribution of the asset feature combination x and the category label y, P(y) represents the prior probability of the category label y, P(x i |x1,x2,…,x i-1 ,y) means that given the category label y and all previous features x1,x2,…,x i-1 In the case of x i The conditional probability of occurrence.

8. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 7 is characterized in that: The S12 is specifically: Based on Bayes' theorem, the posterior probability of each category label under a given feature combination is calculated according to the following formula: Among them, P(y|x) represents the posterior probability of category label y under asset feature combination x, and P(x) represents the marginal probability of asset feature combination x.

9. The network data asset attribution identification method based on Bayesian causal reasoning according to claim 1 is characterized in that: After S8 and before S11, the method further includes: Taking the attribution identification results of historical network data assets as evidence, a conditional probability table is constructed through optimization algorithm.

10. A network data asset ownership identification system based on Bayesian causal reasoning, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the network data asset attribution identification method based on Bayesian causal reasoning as described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Recommendation technology based on correlation rules and Bayesian network integration

    CN107103000A

  • Intelligent decision-making method and system for enabling enterprise based on Bayesian network data

    CN114548709A