Cross-domain multi-dimensional data trusted sharing processing method and platform
By performing random trusted desensitization and correlation analysis on cross-domain multidimensional data sets, the optimal solution is optimized, which solves the problems of privacy protection and data utility maintenance in cross-domain data sharing and realizes safe and efficient data sharing.
Patent Information
- Application Number
- CN202510781937.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies make it difficult to balance privacy protection and data utility maintenance in cross-domain data sharing, and ignore data correlation, resulting in damage to the structure and meaning of desensitized data.
By obtaining multidimensional data sets for random trusted desensitization processing, the desensitization utility and data utility loss are analyzed in combination with data correlation, the trusted processing fitness is calculated, and the optimal desensitization solution is obtained through optimization.
It achieves the goal of maintaining high data utility while protecting privacy, promoting trusted data sharing, balancing data security and utility, and maintaining data relevance.
Smart Images

Figure CN120705123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a cross-domain multi-dimensional data trusted sharing processing method and platform. Background Art
[0002] In today's digital age, data has become a critical resource driving socioeconomic development and scientific and technological progress. With the rapid development of information technology and the widespread adoption of internet applications, vast amounts of multidimensional data have been generated across diverse fields and industries. This data holds immense value and is crucial for scientific research, business decision-making, public services, and other fields. However, cross-domain data sharing continues to face numerous challenges due to issues such as data privacy, data security, and data ownership.
[0003] Traditional data sharing methods often focus on the simple exchange or copying of data, while neglecting the need to protect privacy and maintain data utility during the sharing process. Especially in the context of cross-domain data sharing, data from different domains have varying levels of sensitivity and sharing value, and traditional desensitization methods often struggle to balance data security and utility. On the one hand, excessive desensitization can cause data to lose its original value and fail to meet sharing needs; on the other hand, insufficient desensitization can lead to data leakage risks, threatening personal privacy and corporate security. Existing desensitization technologies are mostly designed for a single domain or type of data, lacking comprehensive considerations for cross-domain, multi-dimensional data. In the process of cross-domain data sharing, complex relationships often exist between data from different domains, and these relationships are crucial for understanding and applying the data. However, existing desensitization methods often overlook this correlation, making it difficult for desensitized data to maintain its original structure and meaning in cross-domain applications. Summary of the Invention
[0004] The present invention addresses the technical problems in the existing technology that it is difficult to balance privacy protection and data utility maintenance when sharing cross-domain multidimensional data, and ignores data relevance, resulting in damage to the structure and meaning of data after desensitization. The present invention provides a cross-domain multidimensional data trusted sharing processing method and platform to solve the problem.
[0005] The technical solution of the present invention to solve the above technical problems is as follows:
[0006] In a first aspect, the present invention provides a cross-domain multidimensional data trusted sharing processing method, the method comprising: obtaining multiple data sets of multidimensional data to be shared across domains, performing random trusted desensitization processing on the multiple data sets to obtain a first trusted processing scheme; performing desensitization utility analysis on the first trusted processing scheme based on the correlation between the multidimensional data to obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes desensitization utility loss analysis; performing data utility loss analysis on the first trusted processing scheme based on the multidimensional data to obtain a first data utility loss parameter, and calculating a first trusted processing fitness in combination with the first desensitization utility parameter; optimizing the first trusted processing scheme based on the first trusted processing fitness to obtain an optimal trusted processing scheme, performing trusted desensitization processing on the multiple data sets, and performing cross-domain data trusted sharing.
[0007] In a second aspect, the present invention provides a cross-domain multidimensional data trusted sharing processing platform, which includes: a data desensitization module, which is used to obtain multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing scheme; a utility analysis module, which is used to perform desensitization utility analysis on the first trusted processing scheme based on the correlation between the multidimensional data, and obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes desensitization utility loss analysis; a loss analysis module, which is used to perform data utility loss analysis on the first trusted processing scheme based on the multidimensional data, obtain a first data utility loss parameter, and calculate a first trusted processing fitness in combination with the first desensitization utility parameter; a trusted sharing module, which is used to optimize the first trusted processing scheme based on the first trusted processing fitness, obtain an optimal trusted processing scheme, perform trusted desensitization processing on the multiple data sets, and perform cross-domain data trusted sharing.
[0008] The beneficial effects of the present invention are: by obtaining cross-domain multidimensional data sets and performing random trusted desensitization processing, combining data correlation to analyze the desensitization utility and data utility loss, calculating the trusted processing adaptability, and optimizing to obtain the optimal desensitization solution, cross-domain data is achieved while maintaining high data utility while protecting privacy, promoting trusted data sharing, balancing data security and utility, and maintaining the technical effect of data correlation. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of a cross-domain multi-dimensional data trusted sharing processing method provided by the present invention.
[0010] Figure 2 This is a structural diagram of a cross-domain multi-dimensional data trusted sharing processing platform provided by the present invention.
[0011] Explanation of the accompanying drawings: data desensitization module 11, utility analysis module 12, loss analysis module 13, trusted sharing module 14. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0014] In the description of the present invention, the term "for example" is used to mean "used as an example, illustration or illustration". Any embodiment of the present invention described as "for example" is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present invention can be implemented without using these specific details. In other examples, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed herein.
[0015] Example 1:
[0016] like Figure 1 As shown, an embodiment of the present invention provides a cross-domain multi-dimensional data trusted sharing processing method, the method comprising:
[0017] S10: Acquire multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing solution.
[0018] For example, during trusted sharing, multiple datasets of multidimensional data to be shared across domains are first obtained. These datasets may cover diverse fields, such as healthcare, finance, and education, and each dataset contains multiple types of data, such as sensitive information like age, salary, and educational background. Subsequently, these datasets are subjected to random trusted desensitization. This process relies on the randomness and trustworthiness of the desensitization parameters. Specifically, the desensitization parameters can be set to a floating range for sensitive fields such as age and salary, for example, using different desensitization ranges such as ±10% and ±20%. This ensures that the data maintains a certain level of authenticity while effectively preventing direct leakage of sensitive information. By randomly generating multiple first desensitization parameters for the multidimensional data within this parameter space, a first trusted processing scheme is constructed. This scheme aims to balance the needs of data privacy protection and data utility maintenance, laying the foundation for subsequent data sharing. For example, in a medical data sharing scenario, the age field may be randomly adjusted to ±10% of the actual age, protecting patient privacy while maintaining the data's usability for disease analysis.
[0019] S20: Performing a desensitization utility analysis on the first trusted processing solution according to the correlation between the multi-dimensional data to obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes a desensitization utility loss analysis.
[0020] Preferably, a desensitization effectiveness analysis is performed on the initially generated first trusted processing solution, focusing on the correlation between multidimensional data, in order to ensure that the related data remains consistent during the desensitization process and to prevent the risk of inferring the original sensitive information through partial data. The correlation reflects the closeness of the intrinsic connection between data and is a factor that cannot be ignored when analyzing the desensitization effectiveness. Specifically, the desensitization effectiveness analysis includes a desensitization effectiveness loss analysis, which is to evaluate the impact of the desensitization process on the data correlation. For example, in a data set containing two sensitive fields, age and salary, if the age field is desensitized with a ±5% margin, while the salary field is desensitized by 20%, this differentiated desensitization strategy may lead to the destruction of the correlation between the data, making it possible for attackers to infer specific individuals through slight changes in the age field, thereby threatening data privacy. Therefore, by calculating the desensitization utility parameter, the degree of this correlation destruction can be quantified, providing a basis for subsequent optimization of the desensitization strategy. In actual operation, the calculation of the desensitization utility parameters will comprehensively consider the correlation between each data field and the selection of desensitization processing parameters to ensure that the related data can still maintain the necessary ambiguity after desensitization, while maximizing the shared value of the data.
[0021] S30: Performing a data utility loss analysis on the first trusted processing solution based on the multidimensional data to obtain a first data utility loss parameter, and calculating a first trusted processing fitness in combination with the first desensitization utility parameter.
[0022] Furthermore, performing a data utility loss analysis on the generated first trusted processing solution is a key step in optimizing the desensitization strategy. This step focuses on assessing the impact of desensitization on the reference value of the multidimensional data, quantifying the degree of utility loss after desensitization. The data utility loss parameter represents the core metric of this evaluation process, reflecting the degradation of information content and analytical value of the desensitized data compared to the original data. Specifically, this analysis simulates the post-desensitization state of the data for each data category in the multidimensional data, combining the desensitization parameters set in the first trusted processing solution, such as a ±5% fluctuation for the age field or a 20% adjustment for the salary field. This simulated data is then fed into a pre-built data utility loss analyzer, which outputs corresponding data utility loss parameters. These parameters not only account for the utility changes of individual data fields but also integrate the impact of inter-field correlations on the overall data utility.
[0023] After obtaining the first data utility loss parameter, it is necessary to combine it with the first desensitization utility parameter calculated previously to comprehensively evaluate the adaptability of the first trusted processing scheme. As the result of this comprehensive evaluation, the first trusted processing fitness measures the desensitization scheme's ability to strike a balance between protecting data privacy and maintaining data utility. For example, if a desensitization scheme performs well in protecting age privacy (high desensitization utility parameter), but at the same time causes a significant decrease in the analytical value of salary data (high data utility loss parameter), its overall fitness may be low. Through this calculation process, deficiencies in the desensitization scheme can be identified and provide direction for subsequent optimization, such as adjusting the desensitization processing parameters to reduce data utility loss, or improving the desensitization strategy to better balance privacy protection and data utility.
[0024] S40: Optimize the first trusted processing solution according to the first trusted processing adaptability to obtain an optimal trusted processing solution, perform trusted desensitization processing on the multiple data sets, and perform trusted cross-domain data sharing.
[0025] Specifically, systematic optimization of the initially generated trusted processing scheme is a core step in achieving the goal of efficient data sharing. This step focuses on iteratively adjusting the desensitization processing parameters to explore and determine the optimal trusted processing scheme that strikes the best balance between protecting data privacy and maintaining data utility. The optimization aims to reduce the loss of data utility while ensuring that the desensitized data can still meet the needs of cross-domain sharing. Specifically, the optimization process involves applying different combinations of desensitization processing parameters to multiple data sets, such as adjusting the desensitization range of the age field to ±3%, ±7%, etc., or changing the desensitization ratio of the salary field to 15%, 25%, etc., to evaluate the comprehensive impact of different parameter settings on data utility and desensitization utility.
[0026] By constructing an evaluation model and combining the obtained data utility loss parameters and desensitization utility parameters, the fitness of each trusted processing solution is calculated, and the solution with the highest fitness is selected as the optimal trusted processing solution. For example, in data sharing scenarios in the medical and financial fields, if a solution only slightly affects the analytical value of medical expense data while protecting the patient's age privacy, and its overall fitness is better than other solutions, then this solution will be selected as the optimal solution. Subsequently, this optimal solution is used to perform trusted desensitization processing on multiple data sets, ensuring that the desensitized data can effectively protect personal privacy during cross-domain sharing while providing valuable information support to the recipient, ultimately achieving secure and efficient cross-domain data sharing.
[0027] In a preferred embodiment, multiple data sets of multidimensional data to be shared across domains are obtained, and the multiple data sets are randomly and credibly desensitized to obtain a first credible processing solution, including: obtaining multiple data sets of multidimensional data to be shared across domains; respectively obtaining desensitizing processing parameter intervals for desensitizing the multidimensional data to form a credible desensitizing processing solution space; and randomly generating multiple first desensitizing processing parameters for the multidimensional data within the credible desensitizing processing solution space to obtain a first credible processing solution.
[0028] Optionally, multiple data sets of multidimensional data to be shared across domains are obtained. These data sets may come from different domains, such as healthcare, financial services, and education and research, and each data set contains multiple types of data, such as sensitive information such as age, income, and education level. Subsequently, for each type of data, the parameter range for its desensitization processing is determined. These parameter ranges form the basis of the space of trusted desensitization processing solutions. The desensitization processing parameter range refers to the acceptable desensitization range set for different data types to protect data privacy. For example, the age field may be set to a floating range of ±5%, ±10%, etc., and the income field may be scaled or fuzzified.
[0029] After defining the desensitization parameter ranges for each data type, they are combined to form a multi-dimensional space of trusted desensitization solutions, encompassing all possible combinations of desensitization parameters. Next, multiple first desensitization parameter combinations for multidimensional data are randomly generated within this solution space, each representing a potential desensitization strategy. Through this random generation process, a first trusted solution is obtained, which provides the basis for subsequent data utility analysis, desensitization utility analysis, and optimization. For example, in a cross-domain sharing scenario for medical and financial data, a set of desensitization parameters might be randomly generated, where the age field is desensitized by ±7% and the income field is scaled to within 80%-90% of its original value. This serves as a preliminary trusted solution for evaluation and optimization.
[0030] In a preferred embodiment, based on the correlation between the multidimensional data, a desensitization utility analysis is performed on the first trusted processing scheme to obtain a first desensitization utility parameter, including: analyzing the correlation between the data categories of the multidimensional data based on the cross-domain data sharing records in the historical time to obtain multiple correlations; based on the multiple first desensitization processing parameters in the first trusted processing scheme, calculating the absolute difference amplitude of the first desensitization processing parameters between each two data categories to obtain multiple first desensitization utility loss parameters; based on the multiple correlations, performing weighted calculation on the multiple first desensitization utility loss parameters to obtain a first fused desensitization utility loss parameter, and calculating the first desensitization utility parameter.
[0031] Specifically, during trusted sharing, a desensitization utility analysis is performed on the generated first trusted solution to assess its effectiveness. This process first analyzes the correlations between different data categories in multidimensional data based on historical cross-domain data sharing records. These correlations reflect the inherent connections and mutual influence between data. For example, significant correlations exist between age and disease type in medical data, and between income and credit scores in financial data.
[0032] After obtaining multiple correlation degrees, the absolute difference amplitude of the desensitization processing parameters between each two data categories is further calculated in combination with multiple first desensitization processing parameters in the first trusted processing scheme. The absolute difference amplitude refers to the quantitative expression of the difference in the desensitization processing parameters of two data categories. For example, if the desensitization amplitude of one data category is 20% and the other is 5%, then their absolute difference amplitude is calculated as (20%-5%) / 20%=75%, which is used as a first desensitization utility loss parameter between the two data categories, reflecting the potential degree of damage of the desensitization processing to the correlation between the data categories.
[0033] Subsequently, a weighted calculation is performed on the first desensitization utility loss parameter based on the multiple correlations obtained to comprehensively consider the impact of the strength of the correlation between different data categories on the overall desensitization utility, thereby obtaining the first fused desensitization utility loss parameter. Finally, the first desensitization utility parameter is obtained by calculating 1 minus the fused desensitization utility loss parameter. This parameter intuitively reflects the ability of the first trusted processing solution to maintain the correlation between data categories and the overall data utility while protecting data privacy. For example, in the medical and financial data sharing scenario, if the first desensitization utility parameter is higher through the above analysis, it means that the current desensitization processing solution has better preserved the correlation and analytical value between data while protecting privacy.
[0034] In a preferred embodiment, based on the cross-domain data sharing records in the historical time, the correlation between the data categories of the multidimensional data is analyzed to obtain multiple correlation degrees, including: traversing the multiple data categories of the multidimensional data and combining them in pairs to obtain multiple data category groups; based on the cross-domain data sharing records in the historical time, collecting multiple historical data sharing records; collecting the number of times multiple data category groups appear in the multiple historical data sharing records to obtain multiple association occurrence times; and based on the multiple association occurrence times, allocating and calculating to obtain multiple correlation degrees.
[0035] Furthermore, to assess the correlation between multidimensional data categories, we first traverse the multiple data categories in the multidimensional data and combine them pairwise to form multiple data category groups. For example, if a dataset contains three data categories: age, income, and education level, three data category groups might be formed: "age-income," "age-education level," and "income-education level." Subsequently, based on historical cross-domain data sharing records, we collect a large number of historical data sharing records as a basis for analysis. Within these historical records, we count the number of times each data category group appears in all shared records, thereby obtaining multiple associated occurrence counts. For example, if the "age-income" combination appears 30 times in 100 shared records, its associated occurrence count is 30. The associated occurrence count reflects the frequency of co-occurrence of the data category group in historical sharing. Finally, based on these associated occurrence counts, we calculate the ratio of each associated occurrence count to the total associated occurrence count, and then calculate multiple degrees of association. For example, if the total number of associations across all data category groups is 100, and the "age-income" combination only occurs 30 times, then the correlation is 30 / 100 = 0.3, reflecting the close relationship between the "age" and "income" data categories in historical sharing. This process systematically quantifies the correlations between multidimensional data categories, providing an important basis for subsequent desensitization effectiveness analysis and processing solution optimization.
[0036] In a preferred embodiment, based on the multidimensional data, a data utility loss analysis is performed on the first trusted processing scheme to obtain a first data utility loss parameter, including: combining multiple data categories of the multidimensional data with multiple first desensitizing processing parameters in the first trusted processing scheme to obtain multiple data utility loss analysis input data; inputting the multiple data utility loss analysis input data into a pre-built data utility loss analyzer to output multiple data utility loss parameters; and calculating and obtaining a first data utility loss parameter based on the multiple data utility loss parameters.
[0037] In detail, in order to evaluate the impact of the first trusted processing scheme on the data utility, a data utility loss analysis needs to be conducted. Specifically, multiple data categories contained in the multidimensional data, such as age, income, education level, etc., are respectively combined with multiple first desensitization processing parameters in the first trusted processing scheme. The input data for data utility loss analysis refers to the data categories combined with specific desensitization processing parameters, which form the basis for subsequent analysis. For example, if a certain data category is age, and the desensitization processing parameter for age in the first trusted processing scheme is ±5%, then the input data for the data utility loss analysis is the age field and its corresponding ±5% desensitization processing parameter.
[0038] Subsequently, the multiple data utility loss analysis input data are fed into a pre-built data utility loss analyzer. This analyzer, based on a preset algorithm or model, can quantify the impact of the desensitization process on the utility of each data category and output multiple data utility loss parameters. For example, the analyzer may assess that the age field loses 10% of its data utility under ±5% desensitization. Finally, based on these multiple data utility loss parameters, a first data utility loss parameter is calculated through weighted averaging or other statistical methods. This parameter comprehensively reflects the impact of the first trusted processing scheme on the overall utility of the multidimensional data, providing an important basis for optimizing subsequent processing schemes.
[0039] In a preferred embodiment, the construction process of the data utility loss analyzer includes: collecting a sample data category set and a sample desensitization processing parameter set based on data sharing records of different data categories, and desensitizing different sample data categories and sample desensitization processing parameters, marking the data utility loss ratio after desensitization, and obtaining a sample data utility loss parameter set; using machine learning to construct a data utility loss analyzer; using the sample data category set, sample desensitization processing parameter set and sample data utility loss parameter set as input training data and supervised training data, and performing iterative supervised training on the data utility loss analyzer until convergence.
[0040] For example, in the process of building a data utility loss analyzer, it is necessary to systematically collect a set of sample data categories and a set of sample desensitization parameters based on historical data sharing records of different data categories. The sample data category set covers various sensitive data categories such as age, income, and medical records, while the sample desensitization parameter set includes specific desensitization parameters applied to these data categories, such as a ±5% floating range for the age field or a 20% scaling for the income field.
[0041] Then, for each combination of sample data category and sample desensitization parameters, desensitization is performed and the loss ratio of the data utility after desensitization is annotated, thus forming a set of sample data utility loss parameters. For example, for the age field, if the data utility is lost by 10% after applying ±5% desensitization, this loss ratio is annotated as the corresponding sample data utility loss parameter.
[0042] Next, using machine learning technology, a data utility loss analyzer is constructed based on these labeled sample data. The data utility loss analyzer is essentially a predictive model that aims to learn the intrinsic relationship between data categories, desensitization processing parameters and data utility loss. Finally, the data utility loss analyzer is iteratively supervised and trained using a set of sample data categories, a set of sample desensitization processing parameters and a set of sample data utility loss parameters as input training data and supervised training data. During the training process, the analyzer continuously adjusts its internal parameters to minimize the error between the predicted value and the true value until the model converges, that is, reaches the preset prediction accuracy or the upper limit of the number of iterations. For example, in medical and financial data sharing scenarios, through training with a large amount of sample data, the data utility loss analyzer can accurately predict the degree of utility loss of each data category under different desensitization processing parameters, providing strong support for the optimization of subsequent desensitization processing solutions.
[0043] In a preferred embodiment, the first trusted processing fitness is calculated in combination with the first desensitizing utility parameter, including: calculating the first data utility parameter based on the first data utility loss parameter; calculating the first trusted processing fitness based on the first desensitizing utility parameter and the first data utility parameter.
[0044] Preferably, in the process of evaluating the adaptability of the first trusted processing scheme, a first data utility parameter is calculated based on the obtained first data utility loss parameter. The first data utility parameter reflects the extent to which the dataset maintains its original analytical value and practicality under a specific desensitization process. Its calculation generally involves an inverse operation or transformation of the data utility loss parameter. For example, if the first data utility loss parameter indicates a 30% reduction in data utility, the first data utility parameter may be set to 70% to quantify the retained data utility.
[0045] Subsequently, combined with the calculated first desensitization utility parameter, which measures the effectiveness of the desensitization process in protecting data privacy, the first trusted processing fitness is calculated by comprehensively considering the first desensitization utility parameter and the first data utility parameter. As the result of this comprehensive evaluation, the first trusted processing fitness reflects the desensitization solution's ability to balance privacy protection and data utility maintenance. For example, if the first desensitization utility parameter indicates that the desensitization process performs well in privacy protection, and the first data utility parameter shows that the degree of data utility retention is also high, the calculated first trusted processing fitness will be at a high level, indicating that the desensitization process has high applicability and effectiveness in cross-domain data sharing. This fitness indicator provides an important basis for the optimization and selection of subsequent processing solutions.
[0046] In a preferred embodiment, the first trusted processing scheme is optimized to obtain the optimal trusted processing scheme, the multiple data sets are trusted desensitized, and cross-domain data trusted sharing is performed, including: continuing to perform random trusted desensitization on the multiple data sets to obtain a second trusted processing scheme, calculating the second trusted processing fitness, and optimizing the second trusted processing scheme; after the optimization converges, outputting the optimal trusted processing scheme with the largest trusted processing fitness, performing trusted desensitization on the multiple data sets, and cross-domain data trusted sharing is performed.
[0047] Specifically, to obtain the optimal data processing strategy, randomized trusted desensitization must be continuously applied to multiple datasets to generate a diverse set of trusted desensitization solutions. Randomized trusted desensitization means that desensitization parameters (such as the degree of data obfuscation and replacement rules) are randomly generated during each process, thereby exploring a diverse space of desensitization strategies. For each newly generated trusted desensitization solution, its corresponding trusted desensitization fitness is calculated. This fitness comprehensively considers the desensitization solution's performance in protecting data privacy and maintaining data utility. Subsequently, based on the calculated trusted desensitization fitness, an optimization algorithm (such as a genetic algorithm or simulated annealing) is used to iteratively optimize the trusted desensitization solution, aiming to gradually improve its fitness and identify a more optimal desensitization strategy. This optimization process may involve adjusting desensitization parameters and combining multiple desensitization techniques to balance the needs of privacy protection and data utility. For example, in a cross-domain sharing scenario involving medical and financial data, an optimization algorithm might gradually adjust the degree of desensitization of patient age and the degree of obfuscation of income data to ensure that the data remains usable for effective risk assessment and analysis while protecting individual privacy. When the optimization process converges, meaning fitness no longer increases significantly, the optimal trusted processing solution with the highest fitness is output. Finally, this optimal solution is used to perform trusted desensitization on multiple datasets, ensuring that the processed data meets privacy protection requirements while retaining sufficient data utility, thereby supporting trusted data sharing across domains.
[0048] The embodiment of the present invention provides a cross-domain multidimensional data trusted sharing processing method, which has at least the following technical effects:
[0049] 1. By analyzing historical cross-domain data sharing records, quantifying the correlation between multidimensional data categories, and using this as a basis for utility analysis of desensitization solutions, we no longer view desensitization in isolation, but instead take into account the inherent connections between data. This ensures that desensitization can maximize the correlation and analytical value between data while protecting privacy, thereby improving the overall effectiveness of desensitization.
[0050] 2. Construct a data utility loss analyzer based on machine learning. By training on a large amount of sample data, it can accurately predict the degree of data utility loss under different desensitization processing parameters, realize the quantitative assessment of data utility loss, and provide a scientific basis for the optimization of desensitization processing solutions, so that desensitization processing can find a better balance between privacy protection and data utility.
[0051] 3. By calculating the trusted processing fitness, the desensitizing processing scheme is iteratively optimized until the optimal trusted processing scheme with the highest fitness is found. The fitness function is introduced as the optimization target, and the privacy protection effect and data utility loss of the desensitizing processing are comprehensively considered to ensure that the final output desensitizing processing scheme not only meets the privacy protection requirements but also maximizes the practicality and analytical value of the data, thereby supporting cross-domain trusted data sharing.
[0052] Example 2:
[0053] like Figure 2 As shown, based on the same inventive concept as the cross-domain multidimensional data trusted sharing processing method provided in the first embodiment, the embodiment of the present invention further provides a cross-domain multidimensional data trusted sharing processing platform, the platform comprising:
[0054] The data desensitization module 11 is used to obtain multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing solution.
[0055] The utility analysis module 12 is configured to perform a desensitization utility analysis on the first trusted processing solution according to the correlation between the multi-dimensional data to obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes a desensitization utility loss analysis.
[0056] The loss analysis module 13 is used to perform data utility loss analysis on the first trusted processing scheme based on the multidimensional data, obtain a first data utility loss parameter, and calculate a first trusted processing fitness in combination with the first desensitization utility parameter.
[0057] The trusted sharing module 14 is used to optimize the first trusted processing solution according to the first trusted processing adaptability to obtain the optimal trusted processing solution, perform trusted desensitization processing on the multiple data sets, and perform trusted sharing of cross-domain data.
[0058] Furthermore, the data desensitization module 11 is further configured to perform the following steps:
[0059] Acquire multiple data sets of multidimensional data to be shared across domains; obtain desensitizing processing parameter intervals for desensitizing the multidimensional data respectively to form a trusted desensitizing processing solution space; randomly generate multiple first desensitizing processing parameters for the multidimensional data within the trusted desensitizing processing solution space to obtain a first trusted processing solution.
[0060] Furthermore, the utility analysis module 12 is further configured to perform the following steps:
[0061] According to the cross-domain data sharing records in the historical time, the correlation between the data categories of the multidimensional data is analyzed to obtain multiple correlations; according to the multiple first desensitization processing parameters in the first trusted processing scheme, the absolute difference amplitude of the first desensitization processing parameters between each two data categories is calculated to obtain multiple first desensitization utility loss parameters; according to the multiple correlations, the multiple first desensitization utility loss parameters are weightedly calculated to obtain the first fused desensitization utility loss parameter, and the first desensitization utility parameter is calculated.
[0062] Furthermore, the utility analysis module 12 is further configured to perform the following steps:
[0063] The multiple data categories of the multidimensional data are traversed and combined in pairs to obtain multiple data category groups; based on the cross-domain data sharing records in the historical time, multiple historical data sharing records are collected; the number of times the multiple data category groups appear in the multiple historical data sharing records is collected to obtain multiple associated occurrence times; based on the multiple associated occurrence times, multiple correlation degrees are obtained by distribution calculation.
[0064] Furthermore, the loss analysis module 13 is further configured to perform the following steps:
[0065] Combine the multiple data categories of the multidimensional data with the multiple first desensitizing processing parameters in the first trusted processing scheme to obtain multiple data utility loss analysis input data; input the multiple data utility loss analysis input data into a pre-built data utility loss analyzer to obtain multiple data utility loss parameters as output; and calculate and obtain a first data utility loss parameter based on the multiple data utility loss parameters.
[0066] Furthermore, the loss analysis module 13 is further configured to perform the following steps:
[0067] According to the data sharing records of different data categories, a sample data category set and a sample desensitization processing parameter set are collected, and different sample data categories and sample desensitization processing parameters are desensitized, and the utility loss ratio of the data after desensitization is marked to obtain a sample data utility loss parameter set; machine learning is used to construct a data utility loss analyzer; the sample data category set, sample desensitization processing parameter set and sample data utility loss parameter set are used as input training data and supervised training data, and the data utility loss analyzer is iteratively supervised trained until convergence.
[0068] Furthermore, the loss analysis module 13 is further configured to perform the following steps:
[0069] A first data utility parameter is calculated based on the first data utility loss parameter; and a first trusted processing fitness is calculated based on the first desensitization utility parameter and the first data utility parameter.
[0070] Furthermore, the trusted sharing module 14 is further configured to perform the following steps:
[0071] Continue to perform random trusted desensitization processing on the multiple data sets, obtain a second trusted processing scheme, calculate the second trusted processing fitness, and optimize the second trusted processing scheme; after the optimization converges, output the optimal trusted processing scheme with the largest trusted processing fitness, perform trusted desensitization processing on the multiple data sets, and perform trusted sharing of cross-domain data.
[0072] Through the above detailed description of a cross-domain multi-dimensional data trusted sharing processing method in this specification, those skilled in the art can clearly understand a cross-domain multi-dimensional data trusted sharing processing platform in this embodiment. For the platform disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0073] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cross-domain multi-dimensional data trusted sharing processing method, characterized in that: The method comprises: Acquire multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing solution; Performing a desensitization utility analysis on the first trusted processing solution according to the correlation between the multidimensional data to obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes a desensitization utility loss analysis; Performing a data utility loss analysis on the first trusted processing solution based on the multidimensional data to obtain a first data utility loss parameter, and calculating a first trusted processing fitness based on the first desensitization utility parameter; The first trusted processing solution is optimized according to the first trusted processing adaptability to obtain an optimal trusted processing solution, and the multiple data sets are trusted desensitized to perform trusted cross-domain data sharing.
2. The cross-domain multi-dimensional data trusted sharing processing method according to claim 1 is characterized in that: Acquire multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing solution, including: Acquire multiple data sets of multidimensional data to be shared across domains; Obtaining desensitization processing parameter intervals for performing desensitization processing on the multidimensional data respectively to form a credible desensitization processing solution space; A plurality of first desensitization processing parameters of multidimensional data are randomly generated in the credible desensitization processing solution space to obtain a first credible processing solution.
3. The cross-domain multi-dimensional data trusted sharing processing method according to claim 1 is characterized in that: Performing a desensitization effectiveness analysis on the first trusted processing solution based on the correlation between the multi-dimensional data to obtain a first desensitization effectiveness parameter includes: Analyzing the correlation between data categories of the multidimensional data based on cross-domain data sharing records in historical time to obtain multiple correlations; Calculating the absolute difference of the first desensitization processing parameters between each two data categories according to the plurality of first desensitization processing parameters in the first trusted processing scheme to obtain a plurality of first desensitization utility loss parameters; According to the multiple correlation degrees, a weighted calculation is performed on the multiple first desensitization utility loss parameters to obtain a first fused desensitization utility loss parameter, and then the first desensitization utility parameter is obtained by calculation.
4. The cross-domain multi-dimensional data trusted sharing processing method according to claim 3 is characterized in that: Based on the cross-domain data sharing records in the historical time, the correlation between the data categories of the multidimensional data is analyzed to obtain multiple correlations, including: Traversing and combining multiple data categories of the multidimensional data in pairs to obtain multiple data category groups; Collect multiple historical data sharing records based on cross-domain data sharing records within historical time; Collect the number of occurrences of multiple data category groups in multiple historical data sharing records to obtain multiple associated occurrence counts; According to the plurality of association occurrence times, a plurality of association degrees are obtained by allocation calculation.
5. The cross-domain multi-dimensional data trusted sharing processing method according to claim 1 is characterized in that: Performing a data utility loss analysis on the first trusted processing solution according to the multi-dimensional data to obtain a first data utility loss parameter includes: Combining the multiple data categories of the multidimensional data with the multiple first desensitization processing parameters in the first trusted processing solution to obtain multiple data utility loss analysis input data; Inputting the plurality of data utility loss analysis input data into a pre-built data utility loss analyzer, and outputting a plurality of data utility loss parameters; A first data utility loss parameter is calculated based on the multiple data utility loss parameters.
6. The cross-domain multi-dimensional data trusted sharing processing method according to claim 5 is characterized in that: The construction process of the data utility loss analyzer includes: According to the data sharing records of different data categories, a sample data category set and a sample desensitization processing parameter set are collected, and different sample data categories and sample desensitization processing parameters are desensitized, and the utility loss ratio of the data after desensitization is marked to obtain a sample data utility loss parameter set; Use machine learning to build a data utility loss analyzer; The sample data category set, the sample desensitization processing parameter set and the sample data utility loss parameter set are used as input training data and supervised training data to perform iterative supervised training on the data utility loss analyzer until convergence.
7. The cross-domain multi-dimensional data trusted sharing processing method according to claim 1 is characterized in that: Calculating a first trusted processing fitness based on the first desensitization utility parameter includes: Calculating a first data utility parameter according to the first data utility loss parameter; A first trusted processing fitness is calculated based on the first desensitization utility parameter and the first data utility parameter.
8. The cross-domain multi-dimensional data trusted sharing processing method according to claim 1 is characterized in that: Optimize the first trusted processing solution to obtain the optimal trusted processing solution, perform trusted desensitization processing on the multiple data sets, and perform trusted cross-domain data sharing, including: Continue to perform random trusted desensitization processing on the multiple data sets to obtain a second trusted processing solution, calculate the second trusted processing fitness, and optimize the second trusted processing solution; After the optimization converges, the optimal trusted processing solution with the maximum trusted processing adaptability is output, the multiple data sets are trusted desensitized, and cross-domain data trusted sharing is performed.
9. A cross-domain multi-dimensional data trusted sharing and processing platform, characterized by: The platform is used to implement the cross-domain multi-dimensional data trusted sharing processing method according to any one of claims 1 to 8, comprising: A data desensitization module is used to obtain multiple data sets of multidimensional data to be shared across domains, perform random trusted desensitization processing on the multiple data sets, and obtain a first trusted processing solution; a utility analysis module, configured to perform a desensitization utility analysis on the first trusted processing solution based on the correlation between the multidimensional data to obtain a first desensitization utility parameter, wherein the desensitization utility analysis includes a desensitization utility loss analysis; a loss analysis module, configured to perform a data utility loss analysis on the first trusted processing solution based on the multidimensional data, obtain a first data utility loss parameter, and calculate a first trusted processing fitness based on the first desensitization utility parameter; A trusted sharing module is used to optimize the first trusted processing solution according to the first trusted processing adaptability to obtain the optimal trusted processing solution, perform trusted desensitization processing on the multiple data sets, and perform trusted sharing of cross-domain data.