A case reasoning method and device based on incomplete information of data cross-border transmission
Patent Information
- Application Number
- CN202410450014.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-04-15
AI Technical Summary
[0005]本发明提供了一种基于数据跨境传输不完全信息的案例推理方法及装置,为解决现有技术存在的数据属性的缺失问题、分类算法的类别概率相近无法区分问题等做出贡献
Smart Images

Figure CN118296509B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data security and relates to a case-based reasoning method and apparatus based on incomplete information from cross-border data transmission. Background Technology
[0002] With the deepening development of informatization and digitalization, global data is growing rapidly, and large-scale, multi-source cross-border transmission of important data has gradually become the norm. However, the unauthorized export of important data poses a significant threat to data sovereignty and national security. In recent years, the national data security situation has been far from optimistic. Regulatory methods based on manual assessment are limited in both accuracy and efficiency. On the one hand, the recipients of cross-border data often involve multiple countries and regions, and the characteristics of cross-border transmission networks are uncertain and unstable, making it difficult to effectively identify abnormal cross-border data transmission paths or behaviors. On the other hand, regulatory agencies face problems such as incomplete access to relevant information, diverse data risk categories, and limited samples of cross-border data transmission paths when determining the risk categories of cross-border data transmission, making it difficult for regulatory agencies to make accurate judgments based on incomplete information. Therefore, designing an intelligent identification and reasoning method for cross-border data path risks under uncertain information is an urgent problem to be solved.
[0003] While existing technologies have provided some quantitative analysis of cross-border data risks from the perspectives of complex networks and link prediction, they fail to consider the impact of deviations in attributes such as economic environment, policy preferences, institutional characteristics, and transmission frequency, as well as the influence of incomplete path information on cross-border risk assessment. This can lead to data violations and irreparable harm. Furthermore, current regulatory authorities use intelligent identification and classification algorithms to predict the risk categories of cross-border paths. However, these algorithms often make classification decisions based on the higher probability of the category in the analysis results, rarely considering situations with similar category probabilities. This results in the inability to accurately classify categories in scenarios with multiple similar characteristics, and to some extent, the accuracy of classification algorithm results cannot be guaranteed, especially given the practical need for precise classification of cross-border data risk categories.
[0004] Traditional cross-border risk identification methods analyze attributes such as the characteristics of cross-border data transmission paths. Classification techniques are used to train and test historical transmission paths to obtain a reasonable classifier for cross-border data risk categories. However, during classifier training, the attributes of cross-border data are obtained by regulators in a short period through observation and processing, resulting in attribute values that deviate from reality. The timeliness required for assessment results prevents industry regulators from efficiently and accurately confirming and refining incomplete information, thus compromising the accuracy of classification results. Furthermore, general classification algorithms often output the category with the highest probability as the final result. However, cross-border risk categories are often similar, leading to similar classification probability values. Simply outputting the category with the highest probability lacks rigor and scientific validity, and may even result in numerous misjudgments, preventing regulators from taking appropriate regulatory actions. Summary of the Invention
[0005] This invention provides a case-based reasoning method and apparatus based on incomplete information from cross-border data transmission, contributing to solving problems such as missing data attributes and inability to distinguish between categories with similar probabilities in existing technologies.
[0006] This invention provides a case-based reasoning method based on incomplete information from cross-border data transmission, including:
[0007] Step S1: Obtain the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, which is an unknown risk cross-border path;
[0008] Step S2: Divide the attributes of cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each historical path under each type;
[0009] Step S3: Establish a distributed robust classification method based on uncertain cross-border transmission information, classify the target path into risk categories, and output the category confidence of the target path and the weight matrix of cross-border transmission data attributes;
[0010] Step S4: Calculate the case similarity based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data;
[0011] Step S5: Calculate the mixed similarity based on case similarity and category confidence;
[0012] Step S6: Filter out the historical paths with the highest mixing similarity to the target path, and output the similar historical paths of the target path and their corresponding risk categories.
[0013] Optionally, the cross-border data information has a total of M attributes, including data transmission time (Time_trans), data transmission volume (Vol_trans), data transmission frequency (Freq_trans), whether it is a newly added path (Add_path_flag), institution type (Type_org), institution IP address zone (Zone_org), whether cross-border data transmission has been declared according to the process (Rep_flag), the time of promulgation of the latest cross-border policy when the path is generated (Time_pol), and the policy orientation of cross-border data transmission when the path is generated (Deg_att).
[0014] Optionally, the risk category information for historical and target paths includes K types of risks, such as abnormal cross-border traffic, warnings of unknown threats, large-scale outbound encrypted traffic, and frequent cross-border communication.
[0015] Optionally, step S2: The attributes of the cross-border data are categorized into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and the attribute similarity between the target path and each historical path is calculated, including:
[0016] The formula for calculating attribute similarity when the attribute is a clear symbol is as follows:
[0017]
[0018] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ik The similarity of attributes j between them; and These are the paths P that belong to category k for attribute j. ij The attribute values of the target path P0; n represents the n types of this attribute; i represents a path in the historical path, i∈[N], and [N] represents the set of indexes of historical paths in the case library; [M I [] represents the set of indices for clearly symbolic type attributes in the case library;
[0019] The similarity calculation formula for attributes with clear numbers is as follows:
[0020]
[0021] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ik The similarity of attributes j between them; i represents a path in the history path, i∈[N], [N] represents the set of indexes of the history paths in the case library; [M II [] represents the set of indices for clear number type attributes in the case library; Represents attribute value and The distance between them;
[0022] The formula for calculating attribute similarity when the attribute is a fuzzy linguistic variable is as follows:
[0023]
[0024] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ik The similarity of attributes j between them; i represents a path in the history path, i∈[N], [N] represents the set of indexes of the history paths in the case library; [M III ] represents the set of indices for the type attributes of fuzzy linguistic variables in the case library, M = [M I ]∪[M II ]∪[M III According to triangular fuzziness, fuzzy linguistic variables can be represented as follows: and Represents attribute value and The distance between them.
[0025] Optionally, step S3: Establish a distributed robust classification method based on uncertain cross-border transmission information, classify the target path into risk categories, and output the category confidence score and cross-border transmission data attribute weight matrix of the target path, including:
[0026] By utilizing incomplete and uncertain cross-border data information from historical paths, reference distributions of each attribute are obtained, and fuzzy sets of each attribute are constructed based on Wasserstein distance distribution and reference distribution to capture real and unknown attribute distribution information.
[0027] Based on the fuzzy sets of each attribute and the error function of the classification algorithm, we construct an objective function that minimizes the fitting error and a distributed robust optimization model based on cross-border information of incomplete data.
[0028] Based on the reference distribution and Wasserstein fuzzy set, the optimization model is reconstructed into a model form that can be directly solved;
[0029] Solve the reconstructed model and output the attribute weight matrix of cross-border transmission data based on historical path cross-border data information;
[0030] Incomplete and uncertain cross-border data information is input into an optimization model based on incomplete cross-border data information, and the final output is the confidence score of each category.
[0031] Optionally, case similarity is calculated based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data. The calculation method for case similarity is as follows:
[0032]
[0033] Among them, Sim j (P0,P ik Sim(P0, P) represents the attribute similarity, w represents the attribute weight matrix, and Sim(P0, P) represents the attribute similarity. ik () represents the case similarity.
[0034] Optionally, the mixed similarity can be calculated as follows:
[0035] Mixed similarity = Max{Class confidence * Case similarity}
[0036] Optionally, the historical paths with the highest mixed similarity to the target path are selected, and the similar historical paths to the target path and their corresponding risk categories are output, including:
[0037] Each calculated mixed similarity is sorted from largest to smallest, and the historical path with the largest mixed similarity to the target path is selected.
[0038] Determine if the target path and its corresponding historical path have the same risk category, and output the similar historical paths of the target path and their corresponding risk categories.
[0039] The present invention also provides a case-based reasoning device based on incomplete information from cross-border data transmission, comprising:
[0040] The acquisition module is used to acquire the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, which is an unknown risk cross-border path.
[0041] The attribute similarity calculation module is used to classify the attributes of cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each historical path under each type.
[0042] The distributed robust classification module is used to output the attribute weight matrix of cross-border transmission data based on a distributed robust classification method for uncertain cross-border transmission information; and to classify the target path into risk categories and output the category confidence of the target path.
[0043] The case similarity calculation module is used to calculate the case similarity based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data.
[0044] A hybrid similarity calculation module, configured to calculate hybrid similarity based on case similarity and category confidence;
[0045] An output module, configured to screen out the historical path with the maximum hybrid similarity to the target path, and output the similar historical path of the target path and the corresponding risk category.
[0046] Beneficial effects of the present invention:
[0047] The present invention aims to provide a technology that considers the uncertainty of cross-border attributes to reduce the influence of incomplete cross-border path information on path risk type classification, and establishes a distributed robust classification model by integrating the risk preference, tolerable prediction error and current supervision intensity of data cross-border supervision authorities, so as to effectively identify the risk category and probability of cross-border paths. Meanwhile, the present invention is further combined with the case-based reasoning method to design a hybrid case-based reasoning technology based on "optimization + classification" technology, which provides an efficient and robust case analysis and research and judgment scheme for data supervision authorities in practice. The present invention considers the extreme scenario of cross-border data transmission paths, and performs distributed robust classification on cross-border paths based on incomplete cross-border transmission information, thereby ensuring the stability of classification results. Then, through high-dimensional features such as device levels, institution attributes and transmission behaviors, a dynamically updated case-based reasoning library for cross-border data security risks is implemented based on the calculated attribute similarity between cases. The case similarity is calculated based on the attribute similarity and the class credibility to obtain similar cases of cross-border paths with unknown risks in the reasoning library, so as to provide suggestions for industry supervision authorities to take relevant measures. The present invention effectively overcomes the disadvantages of traditional machine learning such as low classification accuracy, poor robustness and excessive dependence on complete data, improves the stability of identifying the risk types of cross-border paths in different cross-border scenarios, and reduces the degree of manual participation in identifying the risk types of cross-border paths. Description of Drawings
[0048] Figure 1 is a schematic flow diagram of a case-based reasoning method based on incomplete information of cross-border data transmission provided by the present invention;
[0049] Figure 2 is an overall framework diagram of a case-based reasoning method based on incomplete information of cross-border data transmission provided by the present invention;
[0050] Figure 3 is a schematic step diagram of a distributed robust classification method based on uncertain cross-border transmission information provided by the present invention;
[0051] Figure 4 is a schematic diagram of a case-based reasoning device based on incomplete information of cross-border data transmission provided by the present invention. Detailed Description of the Embodiments
[0052] This patent employs a case-based reasoning method to provide early warning and intelligent control of data security risks associated with incomplete information in cross-border data transmission. First, in the case-based reasoning model, the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network are stored in a case library. When a new, unknown-risk cross-border path appears in the network, the model can obtain its cross-border data transmission attributes; this path is the target path, and its relevant cross-border transmission attributes constitute the target case description. Second, using a hybrid similarity ranking method, the model retrieves the historical path and its corresponding risk category that are most similar to the current target path's cross-border transmission attributes from the case library. Further, it determines whether the search results match the target path's cross-border data transmission attributes. If they match, the most similar historical path case is reused, and the risk category of the target path is output as the risk category corresponding to that historical path. If they do not match, the adopted historical path is adjusted and modified according to the actual situation, and the matching of the search results is judged again to finally obtain the target path's risk category. Finally, once a new target path and its risk category are identified, if necessary, the new target path and risk category can be stored as a new case in the case library for further retrieval and application. In this cyclical mode, by continuously increasing the number of cases in the historical case library, the reasoning ability of the case reasoning system can be enhanced, while realizing the self-learning ability of case reasoning.
[0053] To address the technical challenges in existing cross-border risk path assessment methods, this invention proposes a case-based reasoning method based on incomplete information. This method identifies the categories of attributes within a path, finely classifying them into three types: clear symbols, clear numbers, and fuzzy linguistic variables. Firstly, for each type of attribute feature, different calculation methods are used to compare it with attribute values from other cases, obtaining the similarity of each attribute in the path to be processed. Secondly, based on observed incomplete cross-border information, this invention employs distributed robust optimization techniques to efficiently and robustly identify the probability of cross-border risk categories and the weights of corresponding attributes, while ensuring the accuracy of classification predictions. This is achieved by combining information such as current policy trends, regulatory authorities' risk preferences, and regulatory intensity, further reducing the probability of misjudgment due to classification uncertainty. Finally, by simultaneously considering case attribute similarity and case category confidence, the method accurately identifies the case in the database most similar to the target path and the corresponding regulatory response, ensuring the accuracy and reliability of the assessment.
[0054] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0055] This invention provides a schematic diagram of a case-based reasoning method based on incomplete information from cross-border data transmission, as shown below. Figure 1 and Figure 2 As shown, the process includes:
[0056] Step S1: Obtain the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, which is an unknown risk cross-border path;
[0057] In one embodiment, the cross-border data information for historical paths and unknown risk cross-border paths includes M attributes, such as: data transmission time (Time_trans), data transmission volume (Vol_trans), data transmission frequency (Freq_trans), and whether it is a newly added path (Add_path_flag); organization information such as organization type (Type_org), organization IP address zone (Zone_org), and whether cross-border data transmission declaration has been carried out according to the process (Rep_flag); and cross-border policy information such as the promulgation time of the latest cross-border policy when the path was generated (Time_pol) and the policy orientation for cross-border data transmission at the time of path generation (Deg_att). The above cross-border data information corresponds to different attributes of the cross-border path Z = {Z1, Z2, ..., Z...} m}, where m represents the number of attributes, Z j This represents attribute j.
[0058] It obtains information on the risk categories of historical and unknown cross-border paths, totaling K risk categories, such as: abnormal cross-border, unknown threat warning, large-scale outbound encrypted traffic, and frequent cross-border communication.
[0059] Let P = {P1, P2, ..., P} n Let P represent a finite set of historical paths. i = i ,R i > represents path i, S i R represents the data transfer attribute value of path i. i This represents the risk category for path i. P0 =<S0,R0> S0 represents the target path, which describes the attributes of cross-border data transfer along the target path, and R0 represents the risk category of the target path.
[0060] Step S2: Divide the attributes of cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each historical path under each type;
[0061] In one embodiment, it is assumed that the cross-border data information has a total of 9 attributes, which can be divided into 3 types according to data type: clear symbols, clear numbers, and fuzzy linguistic variables. Let Z... I Z II Z III Z represents the attribute set representing clear numbers, clear symbols, and fuzzy linguistic variables, respectively.I Z II Z III They are independent and Z = Z I ∪Z II ∪Z III Based on the above assumptions, the feasibility of this invention in a case study is verified. Specifically, data transmission time (Time_trans), data transmission volume (Vol_trans), data transmission frequency (Freq_trans), and the time of promulgation of the latest cross-border policy when the path is generated (Time_pol) can be classified as clear numbers; whether it is a new path (Add_path_flag), institution type (Type_org), institution IP address zone (Zone_org), and whether cross-border data transmission declaration has been carried out according to the procedure (Rep_flag) can be classified as clear symbols; and the policy guidance for cross-border data transmission when the path is generated (Deg_att) can be classified as a fuzzy linguistic variable.
[0062] Depending on the data type, the target path and the path P belonging to class k are calculated in different ways. ik Sim similarity between them for each attribute j (P0,P ik ).
[0063] The formula for calculating attribute similarity when the attribute is a clear symbol is as follows:
[0064]
[0065] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ik The similarity of attributes j between them; and These are the paths P that belong to category k for attribute j. ij The attribute values of the target path P0; n represents the n types of this attribute; i represents a path in the historical path, i∈[N], and [N] represents the set of indexes of historical paths in the case library; [M I [] represents the set of indexes for clearly symbolic type attributes in the case library.
[0066] The similarity calculation formula for attributes with clear numbers is as follows:
[0067]
[0068] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ikThe similarity of attributes j between them; i represents a path in the history path, i∈[N], [N] represents the set of indexes of the history paths in the case library; [M II [] represents the set of indexes for clear number type attributes in the case library. Represents attribute value and The distance between them is calculated using the following formula:
[0069]
[0070] The formula for calculating attribute similarity when the attribute is a fuzzy linguistic variable is as follows:
[0071]
[0072] Among them, Sim j (P0,P ik ) represents the path P0 to the target path P0 and the path P belonging to category k. ik The similarity of attributes j between them; i represents a path in the history path, i∈[N], [N] represents the set of indexes of the history paths in the case library; [M III ] represents the set of indices for the type attributes of fuzzy linguistic variables in the case library, M = [M I ]∪[M II ]∪[M III According to triangular fuzziness, fuzzy linguistic variables can be represented as follows: and Represents attribute value and The distance between them is calculated using the following formula:
[0073]
[0074] in,
[0075] Step S3: Establish a distributed robust classification method based on uncertain cross-border transmission information, classify the target path into risk categories, and output the category confidence of the target path and the weight matrix of cross-border transmission data attributes;
[0076] It should be noted that the cross-border data transmission attributes and risk categories of historical paths need to be input into the aforementioned distributed robust classification model to train the model, outputting the distributed robust classifier parameters w, where w is the attribute weight matrix corresponding to all cross-border transmission data. The cross-border data transmission attributes of the target path are then input into the trained distributed robust classification model, outputting the class confidence score of the target path. This method effectively overcomes the uncertainty and noise disturbance of cross-border data, exhibiting strong resilience and classification accuracy. Specifically, the class confidence score of the target path and the cross-border transmission data attribute weight matrix are obtained according to the following steps.
[0077] Optionally, a distributed robust classification method based on uncertain cross-border transmission information, such as... Figure 3 As shown, it includes:
[0078] Step S31: Use incomplete and uncertain cross-border data information from historical paths to obtain the reference distribution of each attribute, and construct fuzzy sets of each attribute based on Wasserstein distance and reference distribution to capture real and unknown distribution information;
[0079] Step S32: Based on the fuzzy sets of each attribute and the error function of the classification algorithm, construct the objective function that minimizes the fitting error and the distributed robust optimization model based on incomplete cross-border information of data;
[0080] Step S33: Based on the reference distribution and Wasserstein fuzzy set, reconstruct the optimization model into a form that can be directly solved;
[0081] Step S34: Solve the reconstructed model and output the attribute weight matrix of cross-border transmission data based on historical path cross-border data information;
[0082] Step S35: Input the incomplete and uncertain cross-border data information into the optimization model based on incomplete cross-border data information, and finally output the confidence scores of each category.
[0083] In one specific embodiment, incomplete and uncertain cross-border data information from historical paths is used to obtain the reference distribution of each attribute, and a fuzzy set of each attribute is constructed based on the reference distribution to capture real and unknown distribution information.
[0084] Based on the fuzzy sets of each attribute and the error function of the classification algorithm, we construct an objective function that minimizes the fitting error and a distributed robust optimization model based on cross-border information of incomplete data.
[0085] Based on the reference distribution and the Wasserstein fuzzy set, the optimization model is reconstructed into a form that can be directly solved;
[0086] Solve the reconstructed optimization model and output the attribute weight matrix of cross-border transmission data based on historical path cross-border data information;
[0087] Incomplete and uncertain cross-border data information is input into an optimization model based on incomplete cross-border data information, and the final output is the confidence score of each category.
[0088] Specifically, in step S31, a fuzzy set is constructed based on the Wasserstein distribution. In the risk classification problem of cross-border data paths, cross-border data is often incomplete and uncertain due to various factors, making it difficult to accurately know the true distribution of the data. To accurately identify the probability values of cross-border risk categories and the weight values of each attribute under the uncertainty and incompleteness of cross-border attributes, this embodiment considers using distributed robust optimization based on Wasserstein distance to capture the true distribution of cross-border attributes. Although the true distribution cannot be obtained and expressed in a finite amount of time, the fuzzy set based on the distribution distance can characterize the uncertainty and fuzziness between the true distribution and the observed distribution.
[0089] Therefore, this embodiment constructs a distributed fuzzy set based on Wasserstein distance, using information such as the risk preferences of regulatory authorities. Let θ be the radius of the Wasserstein sphere, representing the reference distribution of the transmitted feature attributes. For the center of the Wasserstein sphere, construct the following fuzzy set for each cross-boundary attribute:
[0090]
[0091] Among them, the risk appetite θ of the regulatory authorities is a measure of uncertain future risks and returns.
[0092] In step S32, based on the fuzzy sets of each attribute and the error function of the classification algorithm, an objective function that minimizes the fitting error and a distributed robust optimization model based on incomplete cross-border information of data are constructed.
[0093] To ensure the optimization model maintains its performance under the greatest possible uncertainty, considering the worst-case scenario helps build a robust optimization model to improve the overall system's robustness. Therefore, the objective function should minimize the loss function value under the "worst-case" condition. To address the multi-classification problem of cross-border risk, an optimization model considering incomplete sample attributes for identifying multiple risk categories is established as follows:
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] Where [N] = {1,2,…,n,…,|N|} represents the set of cross-border path sample indexes in the case library.
[0100] [K] = {1,2,…,n,…,|K|} represents the set of cross-border risk category indices. The objective function is the minimum risk function of the optimal classifier obtained by training with partial historical samples containing features s and classification labels r. The parameter values of multiple linear classifiers are represented by w.
[0101] Constraint 1 requires that the risk function on the left side of the inequality sign, representing the worst-case scenario in the fuzzy set D of cross-border path attributes, be greater than or equal to the sum of the errors of the linear classifiers of the K classes on N samples. This means the classifier's classification error should satisfy the worst-case scenario for the cross-border path attributes. It's worth noting that Constraint 1 cannot be solved directly and requires a certain equivalent transformation. Constraint 2 defines the loss function for the nth sample under the classifier of the kth class, used to characterize the difference between the classifier's classification result and the true result. Constraint 3 defines the general form of the linear function, used to represent the linear classifier function of the kth class.
[0102] In step S33, the optimization model is reconstructed into a form that can be directly solved based on the reference distribution and the Wasserstein fuzzy set.
[0103] In this embodiment, since constraint 1 needs to solve how to capture unknown distribution information based on existing distribution information, some necessary transformations are shown below.
[0104] Because the s-value in the real-world scenario differs from the actually observed property... There is an error between the values, so let Q = (q ls ) l∈[L],n∈[N] The s-value represents the value in the real-world scenario and the attribute actually observed. The joint distribution of values, where the real-world s values are {s1,...,s} L According to the reference distribution and fuzzy set functions The reconstructed model is as follows:
[0105]
[0106]
[0107]
[0108]
[0109] When the above model has an optimal solution, its strong dual problem also has an optimal solution, and the optimal objective function values of the two problems are the same. Through strong dual transformation, the goal-oriented distributed robust optimization is reconstructed into equivalent constraints as follows:
[0110] mimw6(w,s,r)
[0111] st
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118] Solving the multi-class probability model based on incomplete cross-border information yields multi-class probabilities. These probabilities not only reflect the likelihood of each sample belonging to different categories, but also, through robust optimization model training, maintain high stability and accuracy even in the face of data uncertainty. The model ultimately outputs class confidence scores and attribute weight matrices w.
[0119] Step S4: Calculate the case similarity based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data;
[0120] In one embodiment, case similarity is calculated based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data. Specifically, the case similarity is calculated as follows:
[0121]
[0122] Among them, Sim j (P0,P ik Sim(P0, P) represents the attribute similarity, w represents the attribute weight matrix, and Sim(P0, P) represents the attribute similarity. ik () represents the case similarity.
[0123] Step S5: Calculate the mixed similarity based on the category confidence and case similarity.
[0124] The category confidence score reflects the degree of certainty that the model in this embodiment belongs to the category of a certain path, and is output by the distributed robust classification model in step S3. Case similarity refers to the proximity of two cross-border data paths in the feature space, and is obtained by the formula in step S4. Considering both category confidence score and case similarity score together helps improve the performance of hybrid similarity in applications. The formula for calculating hybrid similarity is as follows:
[0125] Mixed similarity = Max{Class confidence * Case similarity}
[0126] Step S6: Filter out the historical paths with the highest mixing similarity to the target path, and output the similar historical paths of the target path and their corresponding risk categories.
[0127] Optionally, the historical paths with the highest mixed similarity to the target path are selected, and the similar historical paths to the target path and their corresponding risk categories are output, including:
[0128] Each calculated mixture similarity is sorted from largest to smallest, and the target path and historical path with the largest mixture similarity are selected.
[0129] Determine if the target path and its corresponding historical path have the same risk category, and output the similar historical paths of the target path and their corresponding risk categories.
[0130] In one embodiment, step S6 is implemented by sorting the paths according to their mixed similarity from highest to lowest. Path P7 has the highest mixed similarity with the target path P0, and its risk category is "abnormal cross-border." The output path is P7, which is a similar historical path to the target path P0, and its risk category is "abnormal cross-border." The output result of this embodiment provides a management basis for industry regulatory departments.
[0131] It should be noted that this invention is the first to propose a case-based reasoning method in cross-border path assessment that addresses the issues of incomplete cross-border attribute features and inaccurate classification probabilities. Unlike traditional case-based reasoning methods, this study combines robust optimization to obtain robust results for cross-border feature weights and integrates individual case similarity with case category probability, proposing a hybrid similarity calculation method. This contributes to the further implementation and application of case-based reasoning methods in other fields.
[0132] This invention also provides a case-based reasoning device based on incomplete information from cross-border data transmission, such as... Figure 4 As shown, it includes:
[0133] The acquisition module is used to acquire the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, which is an unknown risk cross-border path.
[0134] The attribute similarity calculation module is used to classify the attributes of cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each historical path under each type.
[0135] The distributed robust classification module is used to output the attribute weight matrix of cross-border transmission data based on a distributed robust classification method for uncertain cross-border transmission information; and to classify the target path into risk categories and output the category confidence of the target path.
[0136] The case similarity calculation module is used to calculate the case similarity based on the attribute similarity between the target path and each historical path and the attribute weight matrix of cross-border transmission data.
[0137] The hybrid similarity calculation module is used to calculate hybrid similarity based on case similarity and category confidence.
[0138] The output module is used to filter out the historical paths with the highest mixing similarity to the target path, and output the similar historical paths of the target path and their corresponding risk categories.
[0139] It should be noted that the case reasoning device based on incomplete information from cross-border data transmission provided by the present invention can implement the same method steps as the above-described method embodiments, and will not be repeated here.
Claims
1. A case reasoning method based on incomplete information of cross-border data transmission, characterized in that, include: Step S1: Obtain the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, wherein the target path is an unknown risk cross-border path; Step S2: Divide the attributes of the cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each historical path under each type; Step S3: Establish a distributed robust classification method based on uncertain cross-border transmission information to classify the target path into risk categories, and output the category confidence score of the target path and the attribute weight matrix of the cross-border transmission data; establishing a distributed robust classification method based on uncertain cross-border transmission information to classify the target path into risk categories, and outputting the category confidence score of the target path and the attribute weight matrix of the cross-border transmission data includes: Based on the incomplete and uncertain cross-border data information in the historical path, the reference distribution of each attribute is obtained, and the fuzzy set of each attribute is constructed based on the Wasserstein distance and the reference distribution to capture the real and unknown attribute distribution information. Based on the fuzzy sets of the attributes and the error function of the classification algorithm, an objective function that minimizes the fitting error and a distributed robust optimization model based on incomplete cross-border information of data are constructed. Based on the reference distribution and the Wasserstein fuzzy set, the optimization model is reconstructed into a model form that can be directly solved; The reconstructed model is solved, and the attribute weight matrix of the cross-border transmission data is output based on the historical path cross-border data information. The incomplete and uncertain cross-border data information of the target path is input into the optimization model based on incomplete cross-border data information, and finally the confidence scores of each category are output. Step S4: Calculate the case similarity based on the attribute similarity between the target path and each of the historical paths and the attribute weight matrix of the cross-border transmission data; Step S5: Calculate the mixed similarity based on the case similarity and the category confidence. Step S6: Filter out the historical paths with the highest mixing similarity to the target path, and output the similar historical paths of the target path and their corresponding risk categories.
2. The case reasoning method based on incomplete information of data cross-border transmission according to claim 1, characterized in that, Step S1: The cross-border data information has M attributes, including data transmission time (Time_trans), data transmission volume (Vol_trans), data transmission frequency (Freq_trans), whether it is a new path (Add_path_flag), institution type (Type_org), institution IP address zone (Zone_org), whether cross-border data transmission declaration has been carried out according to the process (Rep_flag), the promulgation time of the latest cross-border policy when the path is generated (Time_pol), and the policy orientation of cross-border data transmission when the path is generated (Deg_att).
3. The case-based reasoning method based on incomplete information from cross-border data transmission as described in claim 1, characterized in that, Step S1: The risk category information of the historical path and the target path includes K types of risks, such as abnormal cross-border, unknown threat warning, large-scale outbound encrypted traffic, and frequent cross-border communication.
4. The case-based reasoning method based on incomplete information from cross-border data transmission as described in claim 1, characterized in that, Step S2: Divide the attributes of the cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and calculate the attribute similarity between the target path and each of the historical paths, including: The formula for calculating attribute similarity when the attribute is a clear symbol is as follows: in, Indicates the target path Paths belonging to category k Attribute similarity between attributes j These are the paths belonging to category k for attribute j. and target path The attribute value; n represents the n possible types of the attribute; i represents the path in the history path. , [ This represents the set of indexes for historical paths in the case library; [] represents the set of indices for clearly symbolic type attributes in the case library; The similarity calculation formula for attributes with clear numbers is as follows: in, Indicates the target path Paths belonging to category k Attribute similarity between attributes j i represents a path in the historical path. , [ This represents the set of indexes for historical paths in the case library; Represents the set of indices for clear number type attributes in the case library; Represents attribute value The distance between them; The formula for calculating attribute similarity when the attribute is a fuzzy linguistic variable is as follows: in, Indicates the target path Paths belonging to category k Attribute similarity between attributes j i represents a path in the historical path. , [ This represents the set of indexes for historical paths in the case library; [] represents the set of indices for the type attributes of fuzzy linguistic variables in the case library. Based on triangular fuzziness, fuzzy linguistic variables are represented as follows: and Represents attribute value The distance between them.
5. The case-based reasoning method based on incomplete information from cross-border data transmission as described in claim 1, characterized in that, Step S4: Calculate the case similarity based on the attribute similarity between the target path and each of the historical paths and the attribute weight matrix of the cross-border transmission data. The case similarity is calculated as follows: in, For attribute similarity, This is the attribute weight matrix. This represents the similarity between cases.
6. The case-based reasoning method based on incomplete information from cross-border data transmission as described in claim 1, characterized in that, Step S5: The calculation method for the hybrid similarity is as follows: Mixed similarity = Max{Class confidence} Case similarity}.
7. The case-based reasoning method based on incomplete information from cross-border data transmission as described in claim 1, characterized in that, Step S6: Filter out the historical paths with the highest mixed similarity, and output the similar historical paths of the target path and their corresponding risk categories, including: Each calculated mixed similarity is sorted from largest to smallest, and the historical path with the largest mixed similarity to the target path is selected. Determine that the target path and the corresponding historical path have the same risk category, and output the similar historical paths of the target path and their corresponding risk categories.
8. A case-based reasoning device based on incomplete information from cross-border data transmission, characterized in that, include: The acquisition module is used to acquire the cross-border data transmission attributes and risk categories of historical paths in the cross-border data transmission network, as well as the cross-border data transmission attributes of the target path, wherein the target path is an unknown risk cross-border path. The attribute similarity calculation module is used to classify the attributes of the cross-border data into three types: clear symbols, clear numbers, and fuzzy linguistic variables, and to calculate the attribute similarity between the target path and each of the historical paths under each type. The distributed robust classification module is used to output the attribute weight matrix of the cross-border transmission data according to the distributed robust classification method based on uncertain cross-border transmission information; to establish a distributed robust classification method based on uncertain cross-border transmission information, to classify the target path into risk categories, and to output the category confidence of the target path and the attribute weight matrix of the cross-border transmission data, including: Based on the incomplete and uncertain cross-border data information in the historical path, the reference distribution of each attribute is obtained, and the fuzzy set of each attribute is constructed based on the Wasserstein distance and the reference distribution to capture the real and unknown attribute distribution information. Based on the fuzzy sets of the attributes and the error function of the classification algorithm, an objective function that minimizes the fitting error and a distributed robust optimization model based on incomplete cross-border information of data are constructed. Based on the reference distribution and the Wasserstein fuzzy set, the optimization model is reconstructed into a model form that can be directly solved; The reconstructed model is solved, and the attribute weight matrix of the cross-border transmission data is output based on the historical path cross-border data information. The incomplete and uncertain cross-border data information of the target path is input into the optimization model based on incomplete cross-border data information, and the confidence scores of each category are finally output; and the target path is classified into risk categories, and the category confidence scores of the target path are output. The case similarity calculation module is used to calculate the case similarity based on the attribute similarity between the target path and each of the historical paths and the attribute weight matrix of the cross-border transmission data; A hybrid similarity calculation module is used to calculate hybrid similarity based on the case similarity and the category confidence. The output module is used to filter out the historical paths with the highest mixing similarity to the target path, and output the similar historical paths of the target path and their corresponding risk categories.
Citation Information
Patent Citations
A robust active and reactive power coordination optimization method for active distribution network based on time series scenario analysis
CN109274134A
Knowledge inference and fault diagnosis method based on knowledge graph
CN114756686A