A big data privacy aggregation encryption method based on polymorphic fuzzy logic
The dynamic encryption mechanism is constructed through polymorphic fuzzy logic and Kuyu optimization algorithm, which solves the problem of balance between data privacy protection and processing efficiency in big data scenarios, and realizes efficient, secure encryption and sharing of heterogeneous data.
Patent Information
- Application Number
- CN202510104896.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The existing technology is difficult to achieve the balance between data privacy protection and processing efficiency in big data scenarios. The traditional static encryption method has high computational complexity, the simple design of desensitization processing rules can easily lead to the leakage of sensitive information, and the fuzzy logic method is difficult to dynamically adapt to heterogeneous data and multi-scenario needs, and lacks a unified privacy protection mechanism.
The fuzzy logic rule library is constructed using polymorphic fuzzy logic, combined with the Kuyu optimization algorithm to optimize the fuzzy membership function and rules, and through the dynamic encryption key generation strategy, the data is fuzzed and aggregated and encrypted, stored in a distributed storage system and controlled access rights.
It realizes flexible processing of heterogeneous data, improves privacy protection strength and encryption efficiency, ensures data security and availability, and meets the real-time needs of big data scenarios.
Smart Images

Figure CN119538315B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of encryption technology, and in particular to a big data privacy aggregation encryption method based on polymorphic fuzzy logic. Background Art
[0002] With the rapid development of big data technology, data privacy protection has gradually become a focus of public and technical attention. In big data scenarios involving sensitive information in medicine, finance, and social networking, how to protect privacy and security during data sharing and processing has become an important issue that needs to be addressed urgently. Data privacy protection not only concerns personal rights and interests, but is also directly related to the popularization and application of data-driven technologies. However, the current technical system faces many challenges in achieving a balance between privacy protection and data utilization.
[0003] At present, traditional data privacy protection methods mainly use static encryption technology or simple desensitization processing. Static encryption technology prevents unauthorized access by encrypting the entire data, but its disadvantage is that the encryption strength is linearly related to the computational complexity. In big data scenarios, due to the huge amount of data and strong heterogeneity, static encryption often leads to low processing efficiency and difficulty in meeting real-time requirements. At the same time, although desensitization processing can reduce computational complexity, overly simple rule design can easily weaken the availability of data and even lead to the leakage of sensitive information.
[0004] In recent years, fuzzy logic and intelligent optimization algorithms have been gradually introduced into the field of data privacy protection to address the above-mentioned issues. Fuzzy logic utilizes its uncertainty processing capabilities to achieve multi-level protection of complex data, while optimization algorithms can dynamically adjust encryption parameters to adapt to different scenarios. However, there are still deficiencies in practical applications: on the one hand, the rule design of existing fuzzy logic methods mostly relies on human experience, making it difficult to dynamically adapt to heterogeneous data and multi-scenario requirements; on the other hand, traditional optimization algorithms are prone to falling into local optimality, resulting in privacy protection strength and computational efficiency failing to reach ideal levels. In addition, existing technologies lack a unified privacy protection mechanism in data aggregation and distributed storage, making it difficult to simultaneously meet the requirements of data processing efficiency and privacy security. Summary of the Invention
[0005] One purpose of the present invention is to propose a big data privacy aggregation encryption method based on polymorphic fuzzy logic. The present invention effectively improves the strength and efficiency of privacy protection, and also provides a more secure and reliable solution for big data privacy protection scenarios.
[0006] According to an embodiment of the present invention, a big data privacy aggregation encryption method based on polymorphic fuzzy logic includes the following steps:
[0007] S1. Preprocess the target dataset to generate a dataset to be encrypted containing privacy-sensitive information;
[0008] S2, building a fuzzy logic rule base based on polymorphic fuzzy logic;
[0009] S3, using the bitter fish optimization algorithm to optimize the fuzzy membership function parameters and fuzzification rules in the fuzzy logic rule base to generate dynamically optimized fuzzy logic rules;
[0010] S4. Use the optimized fuzzy logic rules to perform fuzzification processing on each data element in the encrypted data set, convert the privacy-sensitive data into a fuzzy set, and form a fuzzy data set;
[0011] S5. An aggregate encryption mechanism is introduced based on the fuzzified dataset to aggregate privacy-sensitive data with similar characteristics in the fuzzified dataset according to their similarity. The data block partitioning strategy is determined according to the optimization goal to form an optimized aggregated data block.
[0012] S6. Encrypt the optimized aggregated data block using a dynamic encryption key generation strategy to generate an encrypted data block, and simultaneously generate index information corresponding to the encrypted data block;
[0013] S7. Storing the encrypted data block and index information in a distributed storage system, and dynamically controlling access to the distributed storage system based on data access permission rules, so that users with different permissions can only obtain data in the encrypted data block that matches their permissions.
[0014] Optionally, S1 includes the following specific processes:
[0015] S11. Clean the target data set to detect and remove redundant data, incomplete data, and abnormal data in the target data set;
[0016] S12. Standardize the format of the target data set after data cleaning, and standardize the numerical data, categorical data, and time data respectively;
[0017] S13, label the privacy-sensitive attributes of the standardized target dataset, predefine the sensitive attribute set S based on domain knowledge, and label each column of data in the target dataset. Classify the attributes and determine whether they belong to the sensitive attribute set S. If the data column If it belongs to the sensitive attribute set S, then the column is marked as a privacy-sensitive attribute;
[0018] S14. Generate a dataset to be encrypted containing privacy-sensitive information based on the target dataset with data cleaning, format standardization, and privacy-sensitive attribute annotation. .
[0019] Optionally, S2 includes the following specific processes:
[0020] S21. Construct a polymorphic fuzzy membership function for fuzzily describing each attribute data in the dataset to be encrypted according to different data types, privacy sensitivity levels, and scenario requirements in the target dataset The polymorphic fuzzy membership function is realized by integrating Gaussian fuzzy membership functions, triangular fuzzy membership functions, and discrete-type fuzzy membership functions, and uniformly expresses the fuzzy descriptions of the same attribute in different scenarios:
[0021] ;
[0022] Among them, represents the attribute with the attribute serial number in the dataset to be encrypted, represents the privacy requirement for the th application scenario, is the value of the attribute in the data record, , and are weight parameters used to balance the contribution degrees of Gaussian fuzzy membership functions, triangular fuzzy membership functions, and discrete-type fuzzy membership functions in the scenario . The weight parameters determine the influence strength of different types of data in the fuzzification process in big data privacy aggregation encryption, and are the center and standard deviation parameters of the Gaussian membership function, , ,
[0024] ;
[0025] Where, For the kth rule in the scene The matching degree of the encrypted data set is as follows: is a set of privacy-sensitive attributes, corresponding to the labeled sensitive attributes, For rules For attributes The weight parameter, Adjust the parameters for rule matching, is the number of records in the dataset to be encrypted, To accumulate the fuzzy membership of the attribute after traversing all records in the dataset to be encrypted, Indicates that for any data record d in the data set;
[0026] S23. Perform polymorphic initialization on the fuzzified parameter set P to meet the dynamic adaptability of big data privacy aggregation encryption to diverse data types and privacy requirements;
[0027] S24. Construct a polymorphic fuzzy logic rule base F based on the polymorphic fuzzy membership function, the fuzzy rule set R, and the fuzzy parameter set P:
[0028] .
[0029] Optionally, S3 includes the following specific processes:
[0030] S31, fuzzy logic rule base F defines the comprehensive optimization objective function for the three goals of privacy protection strength, encryption efficiency and decryption accuracy :
[0031] ;
[0032] in, It is a measure of the degree of privacy protection after fuzzification and is defined as the inverse of the deviation between the fuzzy membership and the target threshold. It reflects the encryption efficiency and is defined as the inverse of the encryption time and data size. Characterizes the similarity between the decrypted data and the original data, , , is the weight parameter;
[0033] S32. Construct and initialize the bitter fish population , each individual Represents a set of specific parameter configurations in the fuzzy logic rule base F, setting the population size , the proportion of escaped individuals , Bitterfish activation threshold and the maximum number of iterations ;
[0034] S33. For each individual in the population Calculate the fitness value based on the fuzzy logic rule base and optimization objective function:
[0035] ;
[0036] in, Reflecting individual populations The comprehensive degree of satisfaction with the optimization goal is used to select the optimal solution;
[0037] S34. The update rule of Bitterfish optimization includes the population individuals whose fitness has not triggered the activation threshold. , update the population individual parameters:
[0038] ;
[0039] in, is the updated population individual parameter, is the individual parameter of the current generation population, is the learning step size, is the optimal individual of the current generation, is a random disturbance;
[0040] For fitness below the activation threshold The population individuals trigger the bitter fish activation mechanism and generate new population individuals to replace them:
[0041] ;
[0042] in, It is a new population individual generated by the bitter fish activation mechanism. is the step size, used to control the generation range, is a uniform random number;
[0043] After multiple rounds of iteration, the parameter set converges to the optimal solution and generates an optimized parameter set. ;
[0044] S35. After all iterations are completed or the optimization target converges, the dynamic optimization fuzzy logic rule set corresponding to the optimal solution is output. :
[0045] ;
[0046] in, , is the optimized fuzzy set, corresponding to the updated membership function parameters, Output fuzzy sets for optimized rules, is the fuzzy rule generated by optimization;
[0047] S36, based on the optimized fuzzy rule set and optimized fuzzy membership function and the optimized parameter set , build a dynamically optimized fuzzy logic rule base:
[0048] .
[0049] Optionally, S4 includes the following specific processes:
[0050] S41. Treat each data element in the encrypted data set based on the optimized fuzzy logic rule base. Perform fuzzy processing to obtain all data elements after fuzzy processing , and organized into fuzzy sets ;
[0051] S42. Synthesize the fuzzy sets of all privacy-sensitive attributes to form a complete fuzzy data set :
[0052] ;
[0053] in, A collection of application scenarios;
[0054] S43, the complete fuzzy data set According to the distributed storage requirements, the data is stored in shards and scene index information is added to each data slice to obtain the fuzzy data set stored in shards. .
[0055] Optionally, S5 includes the following specific processes:
[0056] S51. For fuzzy data sets Define the similarity function between data elements , similarity function is used to measure the fuzzy data elements and fuzzy data elements similarity of characteristics;
[0057] S52. Preliminary aggregation of data elements with high similarity in the fuzzified data set is performed according to the following conditions:
[0058] ;
[0059] in, is the similarity threshold, Represents fuzzy data elements and fuzzy data elements Belong to the same aggregate data block;
[0060] Initial aggregation generates a set of data blocks ,in is the mth aggregate data block.
[0061] Optionally, S6 includes the following specific processes:
[0062] S61, for each data block initially generated , dynamically define data blocks based on their characteristics Aggregate encryption key :
[0063] ;
[0064] in, is a hash function, It is the cumulative value of the fuzzy membership of all elements in the aggregated data block, reflecting the fuzzy characteristics of the data block. Is the unique identifier of the data block;
[0065] Leveraging dynamically generated keys Data Block Encrypt to obtain the encrypted aggregate data block ;
[0066] S62. Using the optimization objective function Encrypted aggregate data blocks Perform partition optimization to generate an optimized set of aggregated data blocks :
[0067] ;
[0068] in, To measure the privacy protection strength of the aggregated data block The blurring effect, For encryption efficiency, measure the aggregation of data blocks during encryption The relationship between size and computation time, To measure the decryption accuracy, aggregate data blocks The similarity of the decrypted data to the original data.
[0069] The beneficial effects of the present invention are:
[0070] (1) The present invention adopts polymorphic fuzzy logic to dynamically construct fuzzy membership functions and fuzzification rules according to the data type, privacy sensitivity level and application scenario, which solves the limitation of traditional static fuzzy logic methods that are difficult to adapt to complex scenarios. By organically combining three different fuzzy membership functions, it achieves the unified processing capability of heterogeneous data, enabling it to flexibly cope with diverse big data environments.
[0071] (2) The present invention uses the bitter fish optimization algorithm to globally optimize the parameters and rules of the fuzzy logic rule base, ensuring a balance between privacy protection strength and encryption efficiency. Traditional optimization algorithms are prone to fall into the problem of local optimality, while the bitter fish algorithm improves the global search capability of the algorithm by introducing an escape mechanism and a random perturbation factor, making the optimization of the fuzzy logic rule base more efficient.
[0072] (3) The aggregate encryption mechanism proposed in this invention defines dynamic encryption keys for aggregated data blocks, enabling the encryption process to dynamically adjust key generation rules based on the characteristics of the data blocks. This solves the problem of poor adaptability of traditional static key mechanisms in heterogeneous data scenarios. Dynamic key generation based on the cumulative value of the fuzzy membership of the data blocks and the unique identifier ensures that the encryption process of each data block is independent and more secure. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0074] Figure 1 This is a flowchart of a big data privacy aggregation encryption method based on polymorphic fuzzy logic proposed by the present invention;
[0075] Figure 2 This is a flowchart of optimizing the fuzzy logic rule base based on the bitter fish optimization algorithm in a big data privacy aggregation encryption method based on polymorphic fuzzy logic proposed by the present invention. DETAILED DESCRIPTION
[0076] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0077] refer to Figure 1-Figure 2 , a big data privacy aggregation encryption method based on polymorphic fuzzy logic, including the following steps:
[0078] S1. Preprocess the target dataset to generate a dataset to be encrypted containing privacy-sensitive information;
[0079] S2, building a fuzzy logic rule base based on polymorphic fuzzy logic;
[0080] S3, using the bitter fish optimization algorithm to optimize the fuzzy membership function parameters and fuzzification rules in the fuzzy logic rule base to generate dynamically optimized fuzzy logic rules;
[0081] S4. Use the optimized fuzzy logic rules to perform fuzzification processing on each data element in the encrypted data set, convert the privacy-sensitive data into a fuzzy set, and form a fuzzy data set;
[0082] S5. An aggregate encryption mechanism is introduced based on the fuzzified dataset to aggregate privacy-sensitive data with similar characteristics in the fuzzified dataset according to their similarity. The data block partitioning strategy is determined according to the optimization goal to form an optimized aggregated data block.
[0083] S6. Encrypt the optimized aggregated data block using a dynamic encryption key generation strategy to generate an encrypted data block, and simultaneously generate index information corresponding to the encrypted data block;
[0084] S7. Storing the encrypted data block and index information in a distributed storage system, and dynamically controlling access to the distributed storage system based on data access permission rules, so that users with different permissions can only obtain data in the encrypted data block that matches their permissions.
[0085] In this embodiment, S1 includes the following specific processes:
[0086] S11. Clean the target data set to detect and remove redundant data, incomplete data, and abnormal data in the target data set;
[0087] S12. Standardize the format of the target data set after data cleaning, and standardize the numerical data, categorical data, and time data respectively;
[0088] S13, label the privacy-sensitive attributes of the standardized target dataset, predefine the sensitive attribute set S based on domain knowledge, and label each column of data in the target dataset. Classify the attributes and determine whether they belong to the sensitive attribute set S. If the data column If it belongs to the sensitive attribute set S, then the column is marked as a privacy-sensitive attribute;
[0089] S14. Generate a dataset to be encrypted containing privacy-sensitive information based on the target dataset with data cleaning, format standardization, and privacy-sensitive attribute annotation. .
[0090] In this embodiment, S2 includes the following specific processes:
[0091] S21. Construct a polymorphic fuzzy membership function for fuzzily describing each attribute data in the dataset to be encrypted according to different data types, privacy sensitivity levels, and scenario requirements in the target dataset The polymorphic fuzzy membership function is realized by integrating Gaussian fuzzy membership functions, triangular fuzzy membership functions, and discrete-type fuzzy membership functions, and uniformly expresses the fuzzy descriptions of the same attribute in different scenarios:
[0092] ;
[0093] Among them, represents the attribute with the attribute number in the dataset to be encrypted, represents the privacy requirement for the th application scenario, and is the value of the attribute in the data record, 丶 and are weight parameters used to balance the contribution degrees of Gaussian fuzzy membership functions, triangular fuzzy membership functions, and discrete-type fuzzy membership functions in the scenario . The weight parameters determine the influence intensity of different types of data in the fuzzification process in big data privacy aggregation encryption, and are the center and standard deviation parameters of the Gaussian membership function, , , are the lower bound, peak, and upper bound parameters of the triangular membership function, used to fuzzify the attribute data in the numerical domain to meet the requirements of different privacy sensitivity levels, is the v-th possible value in the discrete value set of categorical data, is an indicator function used to fuzzily describe specific value matching for categorical data, and further perform granularity processing on the privacy characteristics of categorical attributes in the big data privacy aggregation encryption process, is the number of possible values of categorical data in the scenario . g represents the Gaussian membership function, t represents the triangular membership function, cat represents the category characteristics of discrete data, and v represents the index of the discrete value currently being processed;
[0094] S22. Based on the polymorphic fuzzy membership function, construct a fuzzy rule set R to perform fuzzy decision mapping on the privacy-sensitive attributes and application requirements in the scenario 3] in the dataset to be encrypted, and define the matching degree of the fuzzy rule as:
[0095] ;
[0096] Where, For the kth rule in the scene The matching degree of the encrypted data set is as follows: is a set of privacy-sensitive attributes, corresponding to the labeled sensitive attributes, For rules For attributes The weight parameter, Adjust the parameters for rule matching, is the number of records in the dataset to be encrypted, To accumulate the fuzzy membership of the attribute after traversing all records in the dataset to be encrypted, Indicates that for any data record d in the data set;
[0097] S23. Perform polymorphic initialization on the fuzzified parameter set P to meet the dynamic adaptability of big data privacy aggregation encryption to diverse data types and privacy requirements;
[0098] S24. Construct a polymorphic fuzzy logic rule base F based on the polymorphic fuzzy membership function, the fuzzy rule set R, and the fuzzy parameter set P:
[0099] .
[0100] In this embodiment, S3 includes the following specific processes:
[0101] S31, fuzzy logic rule base F defines the comprehensive optimization objective function for the three goals of privacy protection strength, encryption efficiency and decryption accuracy :
[0102] ;
[0103] in, It is a measure of the degree of privacy protection after fuzzification and is defined as the inverse of the deviation between the fuzzy membership and the target threshold. It reflects the encryption efficiency and is defined as the inverse of the encryption time and data size. Characterizes the similarity between the decrypted data and the original data, , , is the weight parameter;
[0104] S32. Construct and initialize the bitter fish population , each individual Represents a set of specific parameter configurations in the fuzzy logic rule base F, setting the population size , the proportion of escaped individuals , Bitterfish activation threshold and the maximum number of iterations ;
[0105] S33. For each individual in the population Calculate the fitness value based on the fuzzy logic rule base and optimization objective function:
[0106] ;
[0107] in, Reflecting individual populations The comprehensive degree of satisfaction with the optimization goal is used to select the optimal solution;
[0108] S34. The update rule of Bitterfish optimization includes the population individuals whose fitness has not triggered the activation threshold. , update the population individual parameters:
[0109] ;
[0110] in, is the updated population individual parameter, is the individual parameter of the current generation population, is the learning step size, is the optimal individual of the current generation, is a random disturbance;
[0111] For fitness below the activation threshold The population individuals trigger the bitter fish activation mechanism and generate new population individuals to replace them:
[0112] ;
[0113] in, It is a new population individual generated by the bitter fish activation mechanism. is the step size, used to control the generation range, is a uniform random number;
[0114] After multiple rounds of iteration, the parameter set converges to the optimal solution and generates an optimized parameter set. ;
[0115] S35. After all iterations are completed or the optimization target converges, the dynamic optimization fuzzy logic rule set corresponding to the optimal solution is output. :
[0116] ;
[0117] in, , is the optimized fuzzy set, corresponding to the updated membership function parameters, Output fuzzy sets for optimized rules, is the fuzzy rule generated by optimization;
[0118] S36, based on the optimized fuzzy rule set and optimized fuzzy membership function and the optimized parameter set , build a dynamically optimized fuzzy logic rule base:
[0119] .
[0120] In this embodiment, S4 includes the following specific processes:
[0121] S41. Treat each data element in the encrypted data set based on the optimized fuzzy logic rule base. Perform fuzzy processing to obtain all data elements after fuzzy processing , and organized into fuzzy sets ;
[0122] S42. Synthesize the fuzzy sets of all privacy-sensitive attributes to form a complete fuzzy data set :
[0123] ;
[0124] in, A collection of application scenarios;
[0125] S43, the complete fuzzy data set According to the distributed storage requirements, the data is stored in shards and scene index information is added to each data slice to obtain the fuzzy data set stored in shards. .
[0126] In this embodiment, S5 includes the following specific processes:
[0127] S51. For fuzzy data sets Define the similarity function between data elements , similarity function is used to measure the fuzzy data elements and fuzzy data elements similarity of characteristics;
[0128] S52. Preliminary aggregation of data elements with high similarity in the fuzzified data set is performed according to the following conditions:
[0129] ;
[0130] in, is the similarity threshold, Represents fuzzy data elements and fuzzy data elements Belong to the same aggregate data block;
[0131] Initial aggregation generates a set of data blocks ,in is the mth aggregate data block.
[0132] In this embodiment, S6 includes the following specific processes:
[0133] S61, for each data block initially generated , dynamically define data blocks based on their characteristics Aggregate encryption key :
[0134] ;
[0135] in, is a hash function, It is the cumulative value of the fuzzy membership of all elements in the aggregated data block, reflecting the fuzzy characteristics of the data block. Is the unique identifier of the data block;
[0136] Leveraging dynamically generated keys Data Block Encrypt to obtain the encrypted aggregate data block ;
[0137] S62. Using the optimization objective function Encrypted aggregate data blocks Perform partition optimization to generate an optimized set of aggregated data blocks :
[0138] ;
[0139] in, To measure the privacy protection strength of the aggregated data block The blurring effect, For encryption efficiency, measure the aggregation of data blocks during encryption The relationship between size and computation time, To measure the decryption accuracy, aggregate data blocks The similarity of the decrypted data to the original data.
[0140] Example 1:
[0141] To verify the feasibility and effectiveness of this invention, the experimenters selected two public medical datasets, MIMIC-III and HCUP, as the experimental data sources. Both datasets contain a large amount of electronic health record data, which are highly heterogeneous and privacy-sensitive.
[0142] The MIMIC-III dataset contains information on 40,000 patients hospitalized at Medical Center A, covering structured and unstructured data on clinical diagnosis, laboratory tests, and treatment plans.
[0143] The HCUP dataset is a collection of inpatient data provided by the B Medical Data Center, containing 50,000 patient records, including personal identification information (name, gender, address) and medical history (disease diagnosis, surgical records).
[0144] The experimenters divided the two data sets into training and test sets in a ratio of 8:2. The goal of the experiment was to verify the performance of the invention in terms of privacy protection strength, encryption efficiency, decryption accuracy and data utilization.
[0145] In the experiment, the experimenters first preprocessed the dataset to clean invalid records and standardized the field format. For the numerical fields of patient age, diagnosis records and laboratory test values, the experimenters performed normalization. For text fields (doctors' clinical instructions), the BERT model was used to extract semantic embedding features and truncate the text length to unify the data format.
[0146] For the privacy-sensitive data in the experiment (names, diagnostic records, laboratory test values), an initial fuzzy logic rule base was constructed based on the data type and privacy sensitivity level, including:
[0147] Gaussian membership function: used for numerical data (age), fuzzifying "30 years old" into a membership interval of 0.7 to [20-40];
[0148] Discrete membership function: used for categorical data (diagnosis codes), mapping similar diagnoses to fuzzy sets, with membership values of 0.8 and 0.9 for "hypertension stage I" and "hypertension stage II" respectively;
[0149] Triangular membership function: used for laboratory test values, defining the "normal range" as the peak of the fuzzy interval.
[0150] Based on the initial rule base, the experimenters used the bitter fish optimization algorithm to optimize the parameters and weights in the fuzzy logic rule base. The training process included 200 rounds of iterations, with a population size of 50 individuals per generation, a learning rate set to 0.01, an activation threshold of 0.05, and a random perturbation factor of 0.1. The optimization objective function was a comprehensive trade-off between privacy protection strength, encryption efficiency, and decryption accuracy.
[0151] During the experiment, the experimenters performed similarity aggregation on the fuzzified data set, aggregating similar laboratory test records (blood sugar value difference <10%) into the same data block. The encryption key of each data block was dynamically generated based on the characteristics of the aggregated data, and the SHA-256 algorithm was used to generate a unique encryption key to ensure data security and independence.
[0152] To demonstrate the effectiveness of this invention, the experimenters compared the performance of traditional static encryption methods and the method of this invention on two data sets. The test results are as follows:
[0153] Table 1 Experimental results on the MIMIC-III dataset
[0154]
[0155] Table 2 Experimental results on the HCUP dataset
[0156]
[0157] As can be seen from Tables 1 and 2 above, the fuzzy membership deviation of the method of the present invention on the two datasets is significantly lower than that of traditional methods in terms of privacy protection strength, with improvements of 41.7% and 50% for the MIMIC-III and HCUP datasets, respectively, effectively enhancing privacy protection capabilities. In terms of encryption and decryption efficiency, compared with traditional methods, the fuzzification time and encryption time of the method of the present invention are reduced by approximately 40%, and the decryption accuracy is improved by approximately 10%, which can meet the real-time and data availability requirements of medical big data scenarios. In terms of data utilization, because the present invention adopts dynamically optimized fuzzification rules and aggregate encryption mechanisms, it retains the main characteristic information of the data while protecting privacy, which increases data utilization to over 90%.
[0158] Through the examples, it can be seen that the method of the present invention is not only superior to traditional methods in terms of privacy protection strength and efficiency, but can also achieve efficient and secure data encryption and sharing in heterogeneous medical data scenarios, and has high practical application value.
[0159] This invention adopts polymorphic fuzzy logic to dynamically construct fuzzy membership functions and fuzzification rules according to the data type, privacy sensitivity level and application scenario, which solves the limitation of traditional static fuzzy logic methods that are difficult to adapt to complex scenarios. By organically combining three different fuzzy membership functions, it realizes the unified processing capability of heterogeneous data, enabling it to flexibly respond to diverse big data environments. Experimental verification shows that in multi-scenario testing, the privacy protection sensitivity of this method is improved by 17%. While the sensitive information is blurred, the analysis utilization rate of the data is maintained at more than 90%, effectively taking into account both privacy protection and data availability.
[0160] This paper uses the bitter fish optimization algorithm to globally optimize the parameters and rules of the fuzzy logic rule base, ensuring a balance between privacy protection strength and encryption efficiency. Traditional optimization algorithms are prone to falling into the problem of local optimality, while the bitter fish algorithm improves the global search capability of the algorithm by introducing an escape mechanism and a random perturbation factor, making the optimization of the fuzzy logic rule base more efficient.
[0161] The aggregate encryption mechanism proposed in the present invention defines dynamic encryption keys for aggregated data blocks, enabling the encryption process to dynamically adjust key generation rules according to the characteristics of the data blocks, thereby solving the problem of poor adaptability of traditional static key mechanisms in heterogeneous data scenarios. The dynamic generation of keys based on the cumulative value of the fuzzy membership of the data blocks and the unique identifier ensures that the encryption process of each data block is independent and more secure.
[0162] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A big data privacy aggregation encryption method based on polymorphic fuzzy logic, characterized by: The steps include: S1. Preprocess the target dataset to generate a dataset to be encrypted containing privacy-sensitive information; S2, building a fuzzy logic rule base based on polymorphic fuzzy logic; The S2 includes the following specific processes: S21. Construct a method for treating encrypted datasets based on different data types, privacy sensitivity levels, and scenario requirements in the target dataset. Polymorphic fuzzy membership function for fuzzy description of each attribute data The polymorphic fuzzy membership function is realized by integrating Gaussian fuzzy membership function, triangular fuzzy membership function and discrete fuzzy membership function, and uniformly expresses the fuzzy description of the same attribute in different scenarios: ; in, Indicates the attribute with attribute number j in the data set to be encrypted. Represents the privacy requirement for the lth application scenario, x is the attribute The value in the data record, and is a weight parameter used to balance the Gaussian fuzzy membership function, triangle fuzzy membership function and discrete fuzzy membership function in the scene The contribution of the weight parameter in big data privacy aggregation encryption determines the influence of different types of data in the fuzzification process. and are the center and standard deviation parameters of the Gaussian membership function, are the lower bound, peak value, and upper bound parameters of the triangle membership function, which are used to fuzzify the attribute data in the numerical domain to meet the needs of different privacy sensitivity levels. is the vth possible value in the discrete value set of categorical data, It is an indicator function that is used to fuzzily describe the matching of specific values for categorized data, and then granularize the privacy characteristics of categorized attributes in the process of big data privacy aggregation encryption. For categorized data in the scene The number of possible values under g, g represents the Gaussian membership function, t represents the triangular membership function, cat represents the category characteristics of discrete data, and v represents the discrete value index currently being processed; S22, based on the polymorphic fuzzy membership function, construct the fuzzy rule set R to treat the encrypted data set Privacy-sensitive attributes and scenarios in The fuzzy decision mapping is performed based on the application requirements below, and the matching degree of the fuzzy rules is defined as: ; Where, For the kth rule in the scene The matching degree of the encrypted data set is as follows: is a set of privacy-sensitive attributes, corresponding to the labeled sensitive attributes, For rules For attributes The weight parameter, Adjust the parameters for rule matching, is the number of records in the dataset to be encrypted, After traversing all records in the dataset to be encrypted, The accumulation of fuzzy membership, Indicates that for any data record d in the data set; S23. Perform polymorphic initialization on the fuzzified parameter set P to meet the dynamic adaptability of big data privacy aggregation encryption to diverse data types and privacy requirements; S24. Construct a polymorphic fuzzy logic rule base F based on the polymorphic fuzzy membership function, the fuzzy rule set R, and the fuzzy parameter set P: ; S3, using the bitter fish optimization algorithm to optimize the fuzzy membership function parameters and fuzzification rules in the fuzzy logic rule base to generate dynamically optimized fuzzy logic rules; S4. Use the optimized fuzzy logic rules to perform fuzzification processing on each data element in the encrypted data set, convert the privacy-sensitive data into a fuzzy set, and form a fuzzy data set; S5. An aggregate encryption mechanism is introduced based on the fuzzified dataset to aggregate privacy-sensitive data with similar characteristics in the fuzzified dataset according to their similarity. The data block partitioning strategy is determined according to the optimization goal to form an optimized aggregated data block. S6. Encrypt the optimized aggregated data block using a dynamic encryption key generation strategy to generate an encrypted data block, and simultaneously generate index information corresponding to the encrypted data block; S7. Storing the encrypted data block and index information in a distributed storage system, and dynamically controlling access to the distributed storage system based on data access permission rules, so that users with different permissions can only obtain data in the encrypted data block that matches their permissions.
2. The big data privacy aggregation encryption method based on polymorphic fuzzy logic according to claim 1 is characterized in that: The S1 includes the following specific processes: S11. Clean the target data set to detect and remove redundant data, incomplete data, and abnormal data in the target data set; S12. Standardize the format of the target data set after data cleaning, and standardize the numerical data, categorical data, and time data respectively; S13, label the privacy-sensitive attributes of the standardized target dataset, predefine the sensitive attribute set S based on domain knowledge, and label each column of data in the target dataset. Classify the attributes and determine whether they belong to the sensitive attribute set S. If the data column If it belongs to the sensitive attribute set S, then the column is marked as a privacy-sensitive attribute; S14. Generate a dataset to be encrypted containing privacy-sensitive information based on the target dataset with data cleaning, format standardization, and privacy-sensitive attribute annotation. .
3. The big data privacy aggregation encryption method based on polymorphic fuzzy logic according to claim 1 is characterized in that: The S3 includes the following specific processes: S31, fuzzy logic rule base F defines the comprehensive optimization objective function for the three goals of privacy protection strength, encryption efficiency and decryption accuracy : ; in, It is a measure of the degree of privacy protection after fuzzification and is defined as the inverse of the deviation between the fuzzy membership and the target threshold. It reflects the encryption efficiency and is defined as the inverse of the encryption time and data size. Characterizes the similarity between the decrypted data and the original data, is the weight parameter; S32. Construct and initialize the bitter fish population , each individual Represents a set of specific parameter configurations in the fuzzy logic rule base F, setting the population size , the proportion of escaped individuals , Bitterfish activation threshold and the maximum number of iterations ; S33. For each individual in the population Calculate the fitness value based on the fuzzy logic rule base and optimization objective function: ; in, Reflecting individual populations The comprehensive degree of satisfaction with the optimization goal is used to select the optimal solution; S34. The update rule of Bitterfish optimization includes the population individuals whose fitness has not triggered the activation threshold. , update the population individual parameters: ; in, is the updated population individual parameter, is the individual parameter of the current generation population, is the learning step size, is the optimal individual of the current generation, is a random disturbance; For fitness below the activation threshold The population individuals trigger the bitter fish activation mechanism and generate new population individuals to replace them: ; in, It is a new population individual generated by the bitter fish activation mechanism. is the step size, used to control the generation range, is a uniform random number; After multiple rounds of iteration, the parameter set converges to the optimal solution and generates an optimized parameter set. ; S35. After all iterations are completed or the optimization target converges, the dynamic optimization fuzzy logic rule set corresponding to the optimal solution is output. : ; in, is the optimized fuzzy set, corresponding to the updated membership function parameters, Output fuzzy sets for optimized rules, is the fuzzy rule generated by optimization; S36, based on the optimized fuzzy rule set and optimized fuzzy membership function and the optimized parameter set , build a dynamically optimized fuzzy logic rule base: 。 4. The big data privacy aggregation encryption method based on polymorphic fuzzy logic according to claim 1 is characterized in that: The S4 includes the following specific processes: S41. Treat each data element in the encrypted data set based on the optimized fuzzy logic rule base. Perform fuzzy processing to obtain all data elements after fuzzy processing , and organized into fuzzy sets ; S42. Synthesize the fuzzy sets of all privacy-sensitive attributes to form a complete fuzzy data set : ; in, A collection of application scenarios; S43, the complete fuzzy data set According to the distributed storage requirements, the data is stored in shards and scene index information is added to each data slice to obtain the fuzzy data set stored in shards. .
5. The big data privacy aggregation encryption method based on polymorphic fuzzy logic according to claim 1 is characterized in that: The S5 includes the following specific processes: S51. For fuzzy data sets Define the similarity function between data elements , similarity function is used to measure the fuzzy data elements and fuzzy data elements similarity of characteristics; S52. Preliminary aggregation of data elements with high similarity in the fuzzified data set is performed according to the following conditions: ; in, is the similarity threshold, Represents fuzzy data elements and fuzzy data elements Belong to the same aggregate data block; Initial aggregation generates a set of data blocks ,in is the mth aggregate data block.
6. The big data privacy aggregation encryption method based on polymorphic fuzzy logic according to claim 5 is characterized in that: The S6 includes the following specific processes: S61, for each data block initially generated , dynamically define data blocks based on their characteristics Aggregate encryption key : ; in, is a hash function, It is the cumulative value of the fuzzy membership of all elements in the aggregated data block, reflecting the fuzzy characteristics of the data block. Is the unique identifier of the data block; Leveraging dynamically generated keys Data Block Encrypt to obtain the encrypted aggregate data block ; S62. Using the optimization objective function Encrypted aggregate data blocks Perform partition optimization to generate an optimized set of aggregated data blocks : ; in, To measure the privacy protection strength of the aggregated data block The blurring effect, For encryption efficiency, measure the aggregation of data blocks during encryption The relationship between size and computation time, To measure the decryption accuracy, aggregate data blocks The similarity of the decrypted data to the original data.
Citation Information
Patent Citations
Establishment and using method of reliability model having fuzzy polymorphism characteristic
CN102024084A
Fuzzy polymorphic manufacturing system task reliability evaluation method based on extended random flow network
CN110210531A