An Artificial Intelligence-Based Privacy-Enhanced Data Security Aggregation Method and System

Through differential privacy algorithms, homomorphic encryption and federated learning frameworks based on artificial intelligence, the independence and aggregation degree of data sharding are dynamically adjusted, the contradiction between data security and business needs in privacy computing is solved, and safe and efficient data aggregation is achieved.

CN119961980BActive Publication Date: 2025-07-11PINGFU INFORMATION TECH HEBEI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510421071.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In the data security aggregation scenario of privacy computing, the independence and aggregation degree of data sharding are difficult to take into account, which makes it difficult to resolve the contradiction between data security and business needs. Especially in scenarios such as finance and medical care, or in scenarios with high real-time recommendation and risk control, it is difficult for the existing technology to flexibly adjust the independence and aggregation degree of data sharding to meet different business needs while ensuring data security.

Method used

Using an artificial intelligence-based method, data sharding is processed through differential privacy algorithms and homomorphic encryption technology, and distributed aggregation calculation is performed using the federated learning framework. The independence and aggregation degree of data sharding are adjusted in combination with machine learning algorithms, data leakage risks are monitored in real time, and security protection measures are activated when necessary to achieve dynamic balance.

Benefits of technology

On the premise of ensuring data security, the independence and aggregation degree of data sharding is flexibly adjusted according to the needs of different business scenarios, improve the efficiency and accuracy of data analysis, reduce the risk of data leakage, and meet the dynamic balance of business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961980B_ABST
    Figure CN119961980B_ABST
Patent Text Reader

Abstract

The present application provides an artificial intelligence-based privacy computing data security aggregation method and system, including: obtaining preset data security levels and accuracy requirements according to business scenario requirements, and determining the independence threshold and aggregation degree range of data sharding; using the differential privacy algorithm to add noise to data shards, reducing the leakage risk while ensuring data independence; encrypting data shards through homomorphic encryption technology to ensure data security during transmission and calculation; if the accuracy of the preliminary aggregation result is lower than the preset threshold, adjust the aggregation degree, increase the interaction times of data shards, and re-perform aggregation calculation; if the adjusted aggregation result meets the accuracy requirements, decrypt the aggregated data to obtain the final result; through system architecture design, embed the adjustment mechanism of data shard independence and aggregation degree into the calculation process to achieve dynamic balance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and in particular to a privacy computing data security aggregation method and system based on artificial intelligence. Background Art

[0002] In the data security aggregation scenario of privacy computing, there is a contradiction between data independence and aggregation degree. On the one hand, to ensure data security, each data shard needs to be able to independently complete part of the aggregation task to avoid cross-leakage of data during transmission and calculation. On the other hand, different business scenarios have different requirements for data accuracy and calculation efficiency, and it is necessary to balance data security and business needs by controlling the aggregation degree.

[0003] However, in actual business scenarios, it is often difficult to balance the independence and aggregation degree of data shards. When the data shards are too independent, although data security can be maximally guaranteed, it may lead to a decrease in the accuracy of the aggregation result and fail to meet business requirements. When the aggregation degree is too high, although data accuracy and calculation efficiency can be improved, it may increase the risk of data leakage.

[0004] In addition, the trade-off between data security and business needs also varies in different business scenarios. In some scenarios with extremely high requirements for data security, such as the financial and medical fields, certain data accuracy and calculation efficiency may need to be sacrificed to ensure absolute data security. In some scenarios with high requirements for real-time performance, such as real-time recommendation and risk control fields, the aggregation degree may need to be appropriately increased to meet the timeliness requirements of the business.

[0005] Therefore, how to flexibly adjust the independence and aggregation degree of data shards according to the requirements of different business scenarios while ensuring data security is a key technical challenge faced by privacy computing data security aggregation. This requires in-depth research and innovation at multiple levels such as algorithm design, system architecture, and security mechanism to find an optimal balance point that takes into account both data security and business needs. Summary of the Invention

[0006] The present invention provides a privacy computing data security aggregation method based on artificial intelligence, mainly including:

[0007] According to the requirements of the business scenario, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation degree range of the data shards;

[0008] Adopt the differential privacy algorithm to perform noise addition processing on the data shards, while ensuring data independence, reducing the leakage risk;

[0009] Through homomorphic encryption technology, encrypt the data shards to ensure the security of data during transmission and calculation;

[0010] According to the preset aggregation degree range, adopt the federated learning framework to perform distributed aggregation calculation on the encrypted data shards to obtain a preliminary aggregation result;

[0011] If the accuracy of the preliminary aggregation result is lower than the preset threshold, adjust the aggregation degree, increase the interaction times of the data shards, and re-perform the aggregation calculation;

[0012] If the adjusted aggregation result meets the accuracy requirements, decrypt the aggregated data to obtain the final result;

[0013] According to the real-time requirement, judge whether it is necessary to optimize the aggregation process, adopt a lightweight encryption algorithm or reduce the shard interaction times to improve the calculation efficiency;

[0014] Through system architecture design, embed the independence of data shards and the adjustment mechanism of aggregation degree into the calculation process to achieve dynamic balance;

[0015] According to the security mechanism design, monitor the data leakage risk during the aggregation process in real time. If an abnormality is found, terminate the calculation and start the security protection measures.

[0016] The present invention provides a privacy computing data security aggregation system based on artificial intelligence, mainly including:

[0017] A data security level and accuracy acquisition module, which is used to obtain the preset data security level and accuracy requirements according to the business scenario requirements, and determine the independence threshold of data shards and the aggregation degree range;

[0018] A differential privacy noise addition module, which is used to add noise to the data shards by using the differential privacy algorithm to reduce the leakage risk while ensuring the independence of the data;

[0019] A homomorphic encryption processing module, which is used to encrypt the data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation;

[0020] A federated learning aggregation calculation module, which is used to perform distributed aggregation calculation on the encrypted data shards according to the preset aggregation degree range by adopting the federated learning framework to obtain a preliminary aggregation result;

[0021] An aggregation result accuracy adjustment module, which is used to adjust the aggregation degree, increase the interaction times of the data shards, and re-perform the aggregation calculation if the accuracy of the preliminary aggregation result is lower than the preset threshold;

[0022] The aggregated data decryption module is used to decrypt the aggregated data to obtain the final result if the adjusted aggregated result meets the accuracy requirements;

[0023] The real-time optimization module is used to determine whether the aggregation process needs to be optimized according to real-time requirements, adopt a lightweight encryption algorithm or reduce the number of shard interactions to improve the calculation efficiency;

[0024] The dynamic balance embedding module is used to embed the independence of data shards and the adjustment mechanism of the aggregation degree into the calculation process through system architecture design to achieve dynamic balance;

[0025] The data leakage monitoring module is used to monitor the data leakage risk during the aggregation process in real time according to the security mechanism design. If an anomaly is found, the calculation is terminated and security protection measures are initiated.

[0026] The technical solution provided by the embodiments of the present invention may include the following beneficial effects:

[0027] The present invention discloses a privacy computing data security aggregation method based on artificial intelligence. This method proposes a dynamic balance solution for the contradiction between data privacy protection and aggregation effect. First, according to the preset data security level and accuracy requirements, the independence threshold of data shards and the aggregation degree range are determined. Then, differential privacy and homomorphic encryption technologies are used to process the data to reduce the leakage risk while ensuring data independence. Next, a federated learning framework is used to perform distributed aggregation calculation on the encrypted data shards. If the accuracy of the aggregated result is insufficient, it is optimized by adjusting the aggregation degree and increasing the number of interactions. Finally, according to real-time requirements, a lightweight encryption algorithm or reducing the number of shard interactions is adopted to improve the calculation efficiency. The present invention also includes a real-time risk monitoring mechanism to initiate security protection measures when an anomaly is found. This method effectively solves the conflict between data privacy protection and aggregation effect and realizes secure and efficient data aggregation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flowchart of a privacy computing data security aggregation method based on artificial intelligence of the present invention;

[0029] Figure 2 is a schematic diagram of a specific embodiment of a privacy computing data security aggregation method based on artificial intelligence of the present invention;

[0030] Figure 3 is another schematic diagram of a specific embodiment of a privacy computing data security aggregation method based on artificial intelligence of the present invention;

[0031] Figure 4 is a schematic diagram of the structure of a privacy computing data security aggregation system based on artificial intelligence of the present invention;

[0032] Figure 5 This is a structural diagram of an electronic device for secure aggregation of privacy computing data based on artificial intelligence provided by an embodiment of the present invention. Specific embodiments

[0033] The technical solutions of the present invention will be described clearly and completely below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] As Figures 1-3 shown, a method for secure aggregation of privacy computing data based on artificial intelligence in this embodiment may specifically include:

[0035] S101. According to the requirements of the service scenario, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation degree range of data sharding.

[0036] Obtain the preset data security level and accuracy requirements; according to the data security level and accuracy requirements, use preset conditions to calculate the independence threshold of data sharding to obtain the first threshold range; for the first threshold range, combine the upper and lower limits of the preset aggregation degree to delimit the first aggregation degree range of data sharding; if the independence threshold of the data sharding is higher than the preset conditions, use the preset logical association rules to adjust the first aggregation degree range to obtain the second aggregation degree range; according to the second aggregation degree range, recalculate the independence threshold of the data sharding to obtain the second threshold range; use the clustering analysis method in the machine learning algorithm to perform independence verification on the data sharding to determine whether it meets the preset conditions; if the independence threshold and aggregation degree range of the data sharding meet the preset service scenario requirements, determine the final data sharding scheme.

[0037] For the acquisition of data security levels and precision requirements, it can be understood that the business needs for security and precision should be clarified before data sharding. Exemplarily, in the financial risk control scenario, assume the data security level is defined as "high" because it involves user privacy data such as transaction records; the precision requirement is "medium" because only risk patterns need to be identified rather than precise numerical predictions. Based on this, the preset condition may be that after data sharding, it is necessary to ensure that a single piece cannot be reversely deduced to obtain the complete information, while maintaining the accuracy rate of risk analysis above 85%. The calculation of the independence threshold involves multiple factors. For example, indicators such as the correlation between data and information entropy can be considered. Suppose there is a set of user browsing records. If the browsing patterns of different users are highly similar, then the independence is low; conversely, if the browsing patterns are significantly different, then the independence is high. The determination of the first threshold range may consider the weighted average of these factors. The delineation of the aggregation degree range needs to balance data privacy and analysis utility. Taking the retail industry as an example, too high an aggregation degree may lead to the inability to identify the purchase tendencies of individual consumers, while too low an aggregation degree may expose sensitive personal information. Therefore, it is necessary to determine an appropriate range according to business needs and regulatory requirements. When the independence threshold is higher than expected, the aggregation degree may need to be adjusted. For example, when analyzing social network data, if it is found that user groups are highly independent, the aggregation degree may need to be reduced to retain more valuable information. Such adjustments may involve redefining data grouping or adopting a more fine-grained analysis method. Recalculating the independence threshold is an iterative process. For example, when analyzing telecom user data, the initial threshold may be set based on call duration. After adjustment, multi-dimensional information such as SMS usage frequency and data traffic may need to be incorporated to obtain a more comprehensive second threshold range. Cluster analysis is an effective method for verifying the independence of data sharding. For example, when conducting customer segmentation, the K-means algorithm can be used. If customers in different shards show obvious clusters in dimensions such as consumption habits and age distribution, it indicates that the shards have good independence. The determination of the final data sharding scheme needs to comprehensively consider multiple factors. Taking the customer data of an insurance company as an example, it may be necessary to find a balance among customer privacy protection, risk assessment accuracy, and market segmentation effect. Only when the sharding scheme can simultaneously meet data security, analysis precision, and business needs can it be recognized as the final scheme. This process of data sharding and verification can not only improve the efficiency and accuracy of data analysis but also play an important role in protecting privacy and meeting regulatory requirements. Through refined data processing, enterprises can better understand customer needs, optimize business processes, and ensure compliance operations at the same time.

[0038] S102. Apply the differential privacy algorithm to add noise to the data shards, while ensuring data independence and reducing the leakage risk.

[0039] Obtain the preset data security level and accuracy requirements, and determine the initial independence threshold for data sharding. For each data shard, add an appropriate amount of noise using the differential privacy algorithm to obtain the data shard after noise addition. If the independence threshold is lower than the preset threshold, adjust the noise addition parameter and perform the noise addition process again until the independence threshold meets the preset condition. Use the clustering analysis method in machine learning to verify the independence of the adjusted data shards. If the independence threshold of the data shard does not meet the preset condition, return to the noise addition process; if it meets the preset condition, adjust the aggregation degree range according to the verification result until the business scenario requirements are met.

[0040] Exemplarily, obtain the preset data security level and accuracy requirements, and determine the initial independence threshold for data sharding in combination with the business scenario requirements. Use the differential privacy algorithm to add an appropriate amount of noise to each data shard to obtain the data shard after noise addition. Calculate the independence threshold based on the data shard after noise addition. If the independence threshold is lower than the preset threshold, adjust the noise addition parameter and perform the noise addition process again until the independence threshold meets the preset condition. For the adjusted data shards, obtain the upper and lower limits of the preset aggregation degree, and determine the aggregation degree range of the data shards. Fine-tune the aggregation degree range in combination with the logical association rules to obtain the adjusted aggregation degree range. Use the clustering analysis method in machine learning to verify the independence of the adjusted data shards. If the independence threshold of the data shard does not meet the preset condition, return to the noise addition process. If the independence threshold of the data shard meets the preset condition, further determine whether its aggregation degree range meets the business scenario requirements. If not, adjust the aggregation degree range according to the verification result until the business scenario requirements are met. Determine the final data sharding scheme including the noise addition parameter, independence threshold, and aggregation degree range. Apply the final scheme to the data sharding process to obtain the processed data shards. Persistently store the processed data shards and combine them with the data access control mechanism. Regularly re-evaluate and adjust the data shards.

[0041] Data security levels and precision requirements are key factors in determining the initial independence threshold for data sharding. For example, in the financial industry, customer transaction data is typically classified as highly sensitive information, requiring the highest level of protection and precision. Suppose a bank sets the initial independence threshold for transaction data at 0.8 (range 0-1), which means that the transaction behaviors of different customer groups should have a high degree of differentiation. The differential privacy algorithm plays an important role in data sharding. It protects individual privacy by adding carefully designed random noise while maintaining the overall statistical characteristics of the data. Taking bank transaction data as an example, noise conforming to the Laplace distribution may be added to the monthly total transaction amount of each customer. The initial noise parameter ε may be set to 0.1, indicating a strong degree of privacy protection. After adding the noise, it is necessary to evaluate whether the independence of the data shards reaches the preset threshold. If the independence threshold is lower than 0.8, for example, only 0.6, it means that the differentiation of the transaction behaviors of different customer groups is insufficient. At this time, it is necessary to adjust the noise addition parameter, and ε may be adjusted to 0.2 to reduce the impact of the noise on the data characteristics. Repeat this process until the preset threshold is reached. Cluster analysis is an effective method for verifying the independence of data shards. For example, the K-means algorithm can be used to cluster the adjusted customer transaction data. If customers can be clearly divided into different groups such as high-frequency small-amount transactions and low-frequency large-amount transactions, it means that the data shards have good independence. If the clustering results show that the data shards still do not meet the independence requirements, it is necessary to return to the noise addition link. This process may require multiple iterations. For example, it may be found that simply adjusting the ε parameter is not enough to improve independence, and at this time, it may be necessary to consider adding noise in other dimensions (such as transaction frequency, transaction type). Finally, adjust the aggregation degree range according to the verification results. In the bank case, it may be found that the original planned aggregation method of 100 customers in a group cannot meet the business requirements because this will obscure important customer behavior patterns. Through adjustment, it may finally be determined that the aggregation method of 50 customers in a group can not only protect personal privacy but also provide sufficient insights for precision marketing. This data processing method can not only improve the accuracy of data analysis but also play an important role in protecting customer privacy and meeting financial regulatory requirements. Through refined data sharding and verification, banks can better understand customer needs, optimize product design, and ensure compliance operations.

[0042] S103. Encrypt the data shards through homomorphic encryption technology to ensure the security of the data during transmission and calculation.

[0043] Obtain data shards to be processed; for each of the data shards, perform encryption processing using homomorphic encryption technology to obtain encrypted data shards; use the clustering analysis method in machine learning to perform independence verification on the encrypted data shards to obtain the independence threshold of the encrypted data shards; determine whether the independence threshold is lower than a preset threshold, if so, adjust the encryption parameters, and return to perform the encryption processing until the independence threshold meets the preset conditions; determine the aggregation degree range according to the independence verification result; determine whether the aggregation degree range meets the requirements of the business scenario, if not, adjust the aggregation degree range, and return to perform the independence verification until the requirements of the business scenario are met.

[0044] Exemplarily, such as Figure 3As shown, for each data shard, homomorphic encryption technology is used for encryption processing to obtain the encrypted data shard. If the independence threshold is lower than the preset threshold, the encryption parameters are adjusted and the encryption processing is redone until the independence threshold meets the preset conditions. The clustering analysis method in machine learning is used to verify the independence of the encrypted data shards. If the independence threshold of the data shards does not meet the preset conditions, the encryption processing link is returned; if it meets the preset conditions, according to the verification results, the aggregation degree range is adjusted until the business scenario requirements are met. Homomorphic encryption is an advanced encryption technology that allows calculations to be directly performed on encrypted data without decryption. This feature makes it widely used in the fields of data analysis and machine learning. Data sharding processing is an important means to protect sensitive information. In the financial field, such as when a bank processes customer transaction records, a balance needs to be struck between protecting privacy and data availability. After obtaining the data shards to be processed, homomorphic encryption technology is used for encryption processing. Homomorphic encryption allows calculations to be directly performed on ciphertext without decryption, greatly improving data security. For example, a bank may homomorphically encrypt information such as the total monthly transaction amount and transaction frequency of customers. The initial encryption parameters may be set at a high security level, such as using a 2048-bit key. After encryption, the independence of the data shards needs to be verified. This step uses the clustering analysis method in machine learning, such as the K-means algorithm. Clustering on encrypted data can evaluate the distinguishability of different customer groups without exposing the original data. Suppose the bank sets the independence threshold at 0.75, indicating a high expected distinguishability between different customer groups. If the clustering result shows that the independence threshold is only 0.6, it means that the encrypted data shards may be too ambiguous to distinguish different customer groups. At this time, the encryption parameters need to be adjusted. The bank may reduce the encryption intensity, such as changing to a 1024-bit key, to achieve a new balance between protecting privacy and retaining data characteristics. After adjustment, re-encryption and independence verification are carried out until the preset threshold of 0.75 is reached. This process may require multiple iterations, fine-tuning the encryption parameters each time until the optimal configuration is found. Based on the independence verification results, the aggregation degree range is determined. This step aims to find an appropriate data aggregation level that can both protect individual privacy and meet business needs. For example, the bank may initially set every 500 customers as an aggregation group. However, through analysis, it is found that this aggregation method may mask important customer behavior patterns and is not conducive to precision marketing. Therefore, the aggregation degree range needs to be adjusted. The bank may try to narrow the aggregation group to every 200 customers and then re-verify the independence. If the new aggregation method meets both the independence requirements (the threshold is still greater than or equal to 0.75) and can provide sufficient insights for business decisions, this aggregation degree range can be adopted. This process may require multiple adjustments and verifications until a balance point is found. Through this refined data processing method, the bank can both protect customer privacy and optimize business operations.For example, in this smaller aggregation group, the bank may identify a newly emerging high-value customer segment, which provides valuable insights for precision marketing and product development. At the same time, since the data is always encrypted, even if a data breach occurs during the analysis process, the personal information of customers will not be directly exposed, greatly reducing the privacy risk.

[0045] S104. According to the preset aggregation degree range, use the federated learning framework to perform distributed aggregation calculation on the encrypted data shards to obtain a preliminary aggregation result.

[0046] Obtain the encrypted data shards, and according to the preset aggregation degree range, use the federated learning framework to perform distributed aggregation calculation on the encrypted data to obtain a preliminary aggregation result; for the preliminary aggregation result, use the clustering analysis method to calculate the independence value of the encrypted data. If the independence value is lower than the preset threshold, adjust the encryption parameters and re-encrypt the data shards to obtain new encrypted data; use the federated learning framework to perform distributed aggregation calculation on the encrypted data within the adjusted aggregation degree range to obtain the final preliminary aggregation result.

[0047] Exemplarily, using a federated learning framework, encrypted data shards are obtained, and distributed aggregation calculations are performed on the encrypted data according to a preset aggregation degree range to obtain a preliminary aggregation result. For the preliminary aggregation result, a clustering analysis method is used to calculate the independence value of the encrypted data, and it is judged whether the independence value is lower than a preset threshold. If it is lower than the preset threshold, the encryption parameters are adjusted and the encryption process is re-executed. According to the adjusted encryption parameters, the data shards are re-encrypted to obtain new encrypted data, and the distributed aggregation calculations are performed on the new encrypted data using the federated learning framework to obtain an updated preliminary aggregation result. For the updated preliminary aggregation result, the independence value is calculated again, and it is judged whether the independence value meets the preset conditions. If it meets, the final aggregation degree range is determined. According to the final aggregation degree range, it is judged whether the range meets the business requirements. If it does not meet, the aggregation degree range is adjusted and the independence value calculation is re-executed. Using the federated learning framework, distributed aggregation calculations are performed on the encrypted data within the adjusted aggregation degree range to obtain a final aggregation result. According to the final aggregation result, a regression analysis method in machine learning is used to optimize the aggregated data to obtain an optimized aggregation result. Federated learning is a distributed machine learning framework that can perform multi-party collaborative calculations while protecting data privacy. In this scenario, federated learning is used for distributed aggregation calculations on encrypted data, which not only protects the confidentiality of the data but also improves the calculation efficiency. Taking the financial field as an example, assume that multiple banks hope to jointly build a credit scoring model but are not willing to directly share customer data. Each bank first encrypts and shards its own customer data. Then, according to a preset aggregation degree range (such as a group of every 1000 customers), each bank uses the federated learning framework for distributed calculations. During this process, only the model parameters are exchanged among the parties, rather than the original data, thus protecting the respective data privacy in collaboration. After the preliminary aggregation result is obtained, independence verification is required. Here, a clustering analysis method, such as the hierarchical clustering algorithm, is used to evaluate the independence of the encrypted data. The independence value reflects the distinguishability between different customer groups. Assume that the preset threshold is 0.8, indicating that a high distinguishability between different customer groups is expected. If the calculated independence value is only 0.6, it means that the current encryption scheme may be too strict, resulting in excessive masking of data features. At this time, the encryption parameters need to be adjusted and the data needs to be reprocessed. For example, the security level of homomorphic encryption can be tried to be reduced from a 2048-bit key to a 1024-bit key. This adjustment aims to find a balance between data availability and privacy protection. After adjustment, the encryption process is re-executed to obtain a new encrypted data set. Next, the distributed aggregation calculations are performed on the new encrypted data using the adjusted aggregation degree range (such as changing to a group of every 800 customers). The purpose of this step is to find the best aggregation level that can both protect privacy and provide valuable information.Through multiple iterations and adjustments, an aggregated result that meets the independence requirements and can provide sufficient insights for business decision-making is finally obtained. The advantage of this method is that it can provide valuable group insights for financial institutions while protecting individual privacy. For example, a bank may find that certain customer groups in specific regions or age brackets have higher credit scores, which can guide them in formulating more precise marketing strategies and risk management policies. At the same time, since the data remains encrypted throughout the process, even in the event of a data breach, customers' personal information will not be directly exposed, greatly reducing the privacy risk.

[0048] S105. If the accuracy of the preliminary aggregated result is lower than the preset threshold, adjust the aggregation degree, increase the number of interactions of the data shards, and re-perform the aggregation calculation.

[0049] The accuracy of the aggregated result is used to characterize the closeness of the aggregated result (such as mean, variance, distribution characteristics) to the true value of the original data, and can be calculated in various ways, such as:

[0050] Absolute Error: |aggregated value - true value|;

[0051] Relative Error: |aggregated value - true value| / true value;

[0052] Confidence Interval: In differential privacy, the range of fluctuations in the result after noise addition (such as the error boundary at an 85% confidence level).

[0053] Exemplarily, in the application of federated learning in the field of financial risk control, the confidence interval is used to characterize. Assuming that the accuracy of the preliminary aggregated result is 75%, lower than the preset threshold of 85%, at this time, the aggregation degree needs to be adjusted to improve the accuracy. First, increase the number of interactions of the sharded data from the default 5 times to 10 times to enhance the ability to extract data features. Adopt a distributed optimization algorithm based on gradient descent, set the learning rate to 01, and the number of iterations to 100, and re-perform the aggregation calculation on the encrypted sharded data. During the calculation process, use differential privacy technology to add Laplace noise, and set the noise parameter ε to 1 to further protect data privacy. After recalculation, an updated aggregated result is obtained, and the accuracy is improved to 82%. Then, use the principal component analysis (PCA) method to perform dimensionality reduction on the updated aggregated result, retaining 95% of the variance information to reduce the data dimension and improve the calculation efficiency.

[0054] Through the K-means clustering algorithm, set the number of clusters to 5, perform clustering analysis on the dimensionality-reduced data, and calculate the ratio of the within-cluster distance to the between-cluster distance to evaluate the rationality of the data distribution.

[0055] If the ratio is lower than the preset threshold 7, further adjust the aggregation degree, increase the interaction times of the sharded data to 15 times, and re-perform the aggregation calculation.

[0056] Finally, through multiple iterations and optimizations, an aggregation result with an accuracy of 88% is obtained, meeting the business requirements.

[0057] S106. If the adjusted aggregation result meets the accuracy requirement, decrypt the aggregated data to obtain the final result.

[0058] Obtain the aggregation result and calculate the accuracy value of the aggregation result; compare the accuracy value with the preset accuracy threshold to determine whether the accuracy value meets the preset accuracy threshold; if the accuracy value meets the preset accuracy threshold, use the preset decryption algorithm to decrypt the aggregation result to obtain decrypted data; for the decrypted data, perform format conversion and standardization processing to obtain standardized data; according to the standardized data, use the preset result generation algorithm to obtain the final result; store the final result in the preset database and mark it as the processed state; according to the storage result, generate a processing log, and the processing log records the key parameters and status information of the aggregation, decryption, and generation links.

[0059] Exemplarily, for the aggregation result, calculate its accuracy value, compare it with the preset accuracy threshold, and determine whether it meets the accuracy requirement. If the accuracy value meets the preset threshold, use the preset decryption algorithm to decrypt the aggregation result to generate decrypted data. For the decrypted data, perform format conversion and standardization processing to generate standardized data. According to the standardized data, use the preset result generation algorithm to generate the final result. Store the final result in the preset database and mark it as the processed state. According to the storage result, generate a processing log, recording the key parameters and status information of the aggregation, decryption, generation, etc.

[0060] S107. According to the real-time requirement, determine whether it is necessary to optimize the aggregation process, adopt a lightweight encryption algorithm or reduce the shard interaction times to improve the calculation efficiency.

[0061] Obtain real-time requirement parameters, calculate the processing time of the current aggregation process, and determine whether the processing time meets a preset time threshold; if the processing time does not meet the preset time threshold, obtain the encryption algorithm used in the aggregation process, and select a target lightweight encryption algorithm from a preset set of lightweight encryption algorithms according to the complexity of the encryption algorithm, and use the target lightweight encryption algorithm to replace the encryption algorithm; obtain the number of shards and interaction times of the aggregated data, and obtain an optimized shard interaction strategy by reducing the number of shards or optimizing the shard interaction logic; perform an aggregation process simulation according to the target lightweight encryption algorithm and the optimized shard interaction strategy to obtain an optimized aggregation time; determine whether the optimized aggregation time meets the preset time threshold, if not, adjust the parameters of the target lightweight encryption algorithm or the shard interaction strategy until the optimized aggregation time meets the preset time threshold; solidify the aggregation process parameters that meet the preset time threshold to obtain an optimized aggregation processing module; obtain the key performance indicators of the optimized aggregation processing module and generate an aggregation efficiency log.

[0062] Exemplarily, obtain real-time requirement parameters, calculate the processing time of the current aggregation process, and determine whether the aggregation efficiency meets a preset time threshold. If not, enter the optimization phase. Analyze the complexity of the encryption algorithm used in the aggregation process, and select a lightweight encryption algorithm to replace the existing algorithm to reduce the consumption of computing resources. Evaluate the number of shards and interaction times of the aggregated data, and reduce the data transmission delay by reducing the number of shards or optimizing the shard interaction logic. Combine the lightweight encryption algorithm and the optimized shard interaction strategy, re-perform the aggregation process simulation, record the aggregation time, and verify the efficiency improvement effect. If the optimized aggregation time still does not meet the real-time requirement, further adjust the encryption algorithm parameters or the shard strategy until the preset time threshold is met. Solidify the optimized aggregation process parameters, update the aggregation processing module, and ensure that subsequent aggregation operations are executed according to the optimized process. Monitor the optimized aggregation process, record the key performance indicators in real time, and generate an aggregation efficiency log for subsequent analysis and further optimization reference.

[0063] The real-time requirement parameter is a key indicator for measuring the efficiency of the aggregation process. In a financial trading system, assuming that the preset time threshold is 100 milliseconds, the current aggregation process takes 150 milliseconds, exceeding the threshold by 50%. At this time, the aggregation algorithm needs to be optimized to improve efficiency. The choice of encryption algorithm directly affects the processing time. For example, the original RSA algorithm has high computational complexity, and it can be considered to be replaced with a lightweight elliptic curve encryption algorithm. While ensuring security, elliptic curve encryption greatly reduces computational overhead and is expected to shorten the processing time to about 120 milliseconds. The sharding strategy of aggregated data is also an important factor affecting efficiency. Assume that 100 shards are originally used, and each shard requires 3 interactions. By adjusting the number of shards to 50 and optimizing the interaction logic to reduce the number of interactions to 2, the processing time can be further shortened. This optimization not only reduces the data transmission overhead, but also reduces the complexity of parallel processing. Simulating the optimized aggregation process is a key step in verifying the effect. Using historical data for simulation testing can estimate the performance of the new solution. If the optimized aggregation time still does not reach the target of 100 milliseconds, further adjustment of parameters is required. For example, you can try to reduce the security parameters of elliptic curve encryption to further improve efficiency while ensuring security. Solidification of aggregation process parameters is an important measure to ensure long-term stability. Write the optimized parameters, such as encryption algorithm type, key length, number of shards, etc., to the configuration file or database to ensure that the system can still run efficiently after restart. This method is also convenient for subsequent version control and rollback operations. Generating aggregation efficiency logs is crucial for system monitoring and continuous optimization. The logs should contain key performance indicators, such as average processing time, peak processing time, resource utilization, etc. By analyzing the trends of these indicators, performance bottlenecks can be discovered in a timely manner, providing a basis for further optimization. For example, if it is found that the processing time in certain periods is significantly extended, it may mean that hardware resources need to be increased or the load balancing strategy needs to be optimized. The entire optimization process embodies an iterative performance tuning method. From identifying problems, proposing optimization solutions, verifying effects to solidifying parameters, a closed loop is formed. This method is not only applicable to the optimization of the aggregation process, but can also be extended to other high-performance computing scenarios, such as real-time data analysis, large-scale parallel computing, and other fields. Through continuous optimization and adjustment, system performance can be continuously improved to meet growing business needs.

[0064] S108. Through system architecture design, the independence of data shards and the adjustment mechanism of the degree of aggregation are embedded in the computing process to achieve dynamic balance.

[0065] Obtain the current data sharding status information of the system, where the data sharding status information includes the number of data shards and the independence index of each shard; judge whether the aggregation degree corresponding to the current data sharding status information meets the standard according to a preset aggregation degree threshold; if the aggregation degree does not meet the standard, start the shard adjustment mechanism, and the shard adjustment mechanism adopts an embedded architecture design and is integrated into the calculation process of data processing; through the shard adjustment mechanism, optimize the number of data shards and the independence index of each shard to obtain the adjusted data sharding status information; adopt a dynamic balance strategy to monitor the aggregation effect corresponding to the adjusted data sharding status information in real time and generate a dynamic balance log; according to the dynamic balance log, use a parameter optimization algorithm to optimize the adjustment parameters in the shard adjustment mechanism to obtain the optimized adjustment parameters; if the aggregation degree corresponding to the adjusted data sharding status information reaches the preset aggregation degree threshold, solidify the adjusted data sharding status information and the optimized adjustment parameters to form an optimized calculation process module.

[0066] Exemplarily, the acquisition of data sharding status information is the key starting point for optimizing the aggregation process. The system collects shard quantity and independence metrics through real-time monitoring, and these metrics reflect the balance and independence degree of data distribution. For example, in a distributed database system, there may be 100 data shards, and the independence metric range of each shard is between 0.6 and 0.9. The higher the independence metric, the less data overlap between shards, which is beneficial to improving parallel processing efficiency. The aggregation degree threshold is an important criterion for measuring whether the data sharding status meets the system requirements. Suppose the set aggregation degree threshold of the system is 0.8, and the current aggregation degree of the system is 0.75, which means that the shard adjustment mechanism needs to be started. The shard adjustment mechanism adopts an embedded architecture and is directly integrated into the data processing flow, which can achieve real-time and dynamic adjustment and reduce additional system overhead. The core of the shard adjustment mechanism is to optimize the data shard quantity and independence metrics. In the above example, the system may decide to reduce the shard quantity from 100 to 80 while increasing the independence metric of each shard. This adjustment can be achieved by reallocating data, merging small shards, or splitting large shards. After the adjustment, assuming the average independence metric increases to 0.85, the overall aggregation degree of the system also increases. The application of the dynamic balance strategy ensures that the system can maintain the best performance after adjustment. Through real-time monitoring, the system generates dynamic balance logs, recording key metrics such as shard size, query response time, resource utilization rate, etc. For example, the log may show that within the first 30 minutes after the adjustment, the query response time has decreased by an average of 20%, but the resource utilization rate of some shards has increased to more than 90%. Based on the dynamic balance logs, the parameter optimization algorithm will further adjust the parameters of the shard mechanism. This may include adjusting the threshold of the shard size, the calculation weight of the independence metric, etc. For example, the system may find that increasing the upper limit of the shard size by 10% can control the resource utilization rate within the ideal range without significantly increasing the query time. When the adjusted aggregation degree reaches the preset threshold (such as 0.8), the system will solidify the optimized configuration. This includes the new data sharding status information (such as 80 shards, average independence metric 0.85) and the optimized adjustment parameters. These information are integrated into the optimized computing process module to ensure that the system can maintain an efficient state during subsequent operations. This dynamic optimization process not only improves the overall performance of the system, but also enhances its adaptability and scalability. Through continuous monitoring and adjustment, the system can cope with challenges such as data volume growth and query pattern changes and always maintain the best operating state.

[0067] S109. According to the security mechanism design, monitor the data leakage risk during the aggregation process in real time. If an anomaly is found, terminate the calculation and start the security protection measures.

[0068] Obtain the real-time data stream during the aggregation process, where the real-time data stream includes the data transmission rate and data content characteristics; determine whether the data leakage risk indicator in the real-time data stream exceeds a preset security threshold; if the data leakage risk indicator exceeds the standard, terminate the current aggregation calculation process and record the abnormal data segment; according to the security protection log, use a machine learning algorithm to optimize the parameters in the security protection mechanism to obtain optimized security protection parameters; if the optimized security protection parameters can effectively reduce the data leakage risk, solidify and form an optimized security protection module.

[0069] Exemplarily, as Figure 2 shown, obtain the real-time data stream during the aggregation process, where the real-time data stream includes the data transmission rate and data content characteristics; use a preset security threshold to determine whether the data leakage risk indicator in the real-time data stream exceeds the standard; if the data leakage risk indicator exceeds the standard, immediately terminate the current aggregation calculation process and record the abnormal data segment; start the security protection mechanism, where the security protection mechanism includes a data encryption module and an access control module; through the data encryption module, encrypt the abnormal data segment to generate an encrypted data packet; use the access control module to restrict the access rights to the encrypted data packet and only allow authorized users to access; monitor the implementation effect of the security protection mechanism in real time to generate a security protection log; according to the security protection log, use a machine learning algorithm to optimize the parameters in the security protection mechanism to obtain optimized security protection parameters; if the optimized security protection parameters can effectively reduce the data leakage risk, solidify the optimized security protection parameters to form an optimized security protection module; continuously monitor the data stream during the aggregation process to ensure the real-time effectiveness of the security protection module; regularly update the preset security threshold to adapt to the changing data environment and security threats; through a preset security audit mechanism, regularly review the security protection log, identify potential security vulnerabilities, and repair them.

[0070] During the data aggregation process, real-time monitoring of the data stream is crucial for ensuring information security. The system constructs a comprehensive portrait of the real-time data stream by collecting metrics such as data transmission rate and content characteristics. For example, in a financial data processing system, under normal circumstances, the amount of data transmitted per second is approximately 500KB, and it mainly contains digitized transaction records. If suddenly the data transmission rate surges to 2MB per second and a large amount of text data appears in the content, this may indicate a potential data leakage risk. The system continuously evaluates the data leakage risk metrics and compares them with the preset security thresholds. Suppose the risk threshold set by the system is 0.7 (with a full score of 1). When it detects that the abnormal data pattern causes the risk metric to rise to 0.8, the system will immediately trigger the security protection mechanism. The design of this mechanism aims to respond quickly to potential threats and minimize the possibility of data leakage to the greatest extent. Once the risk metric exceeds the threshold, the system will immediately interrupt the current aggregation calculation process. Although this measure may temporarily affect the data processing efficiency, it is necessary from the perspective of information security. At the same time, the system will detailedly record the data fragments that lead to the abnormality, including information such as timestamps, data characteristics, and risk scores. These records not only contribute to subsequent security analysis but also provide valuable samples for optimizing the security protection mechanism. The accumulation of security protection logs provides rich training data for machine learning algorithms. The system may adopt algorithms such as support vector machines (SVM) or deep learning networks. By analyzing the patterns of historical abnormal data, it continuously optimizes the security protection parameters. For example, the algorithm may find that certain specific data patterns are highly correlated with high-risk events, thus adjusting the weights of the corresponding features in the risk assessment model. The optimized security protection parameters need to be strictly verified. The system may use historical data for backtesting or conduct simulation tests in an isolated environment. If the new parameters perform well in the tests, reducing the false positive rate by 30% while maintaining a true positive detection rate of over 99%, then this set of parameters is considered effective. Finally, the verified optimized parameters will be integrated into the security protection module of the system. This process not only improves the security of the system but also enhances its adaptability to new threats. Through this continuous optimization cycle, the data aggregation system can maintain a high level of security protection while ensuring efficiency.

[0071] As Figure 4 shown, another embodiment of the present invention provides an artificial intelligence-based privacy computing data security aggregation system, mainly including:

[0072] A data security level and accuracy acquisition module, configured to obtain preset data security levels and accuracy requirements according to business scenario needs, and determine the independence threshold and aggregation degree range of data sharding;

[0073] Differential privacy noise addition module, which is used to add noise to data shards using differential privacy algorithms, reducing the leakage risk while ensuring data independence;

[0074] Homomorphic encryption processing module, which is used to encrypt data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation;

[0075] Federated learning aggregation calculation module, which is used to perform distributed aggregation calculation on encrypted data shards according to the preset aggregation degree range using the federated learning framework to obtain a preliminary aggregation result;

[0076] Aggregation result accuracy adjustment module, which is used to adjust the aggregation degree and increase the interaction times of data shards to re-perform aggregation calculation if the accuracy of the preliminary aggregation result is lower than the preset threshold;

[0077] Aggregated data decryption module, which is used to decrypt the aggregated data to obtain the final result if the adjusted aggregation result meets the accuracy requirements;

[0078] Real-time optimization module, which is used to determine whether it is necessary to optimize the aggregation process according to real-time requirements, and adopt lightweight encryption algorithms or reduce the shard interaction times to improve the calculation efficiency;

[0079] Dynamic balance embedding module, which is used to embed the independence of data shards and the adjustment mechanism of aggregation degree into the calculation process through system architecture design to achieve dynamic balance;

[0080] Data leakage monitoring module, which is used to monitor the data leakage risk during the aggregation process in real time according to the security mechanism design, and terminate the calculation and start security protection measures if anomalies are found.

[0081] In summary, the embodiments of the present invention disclose a privacy computing data security aggregation method and system based on artificial intelligence. This method proposes a dynamic balance solution for the contradiction between data privacy protection and aggregation effect. First, according to the preset data security level and accuracy requirements, determine the independence threshold of data shards and the aggregation degree range. Then, use differential privacy and homomorphic encryption technologies to process the data, reducing the leakage risk while ensuring data independence. Next, use the federated learning framework to perform distributed aggregation calculation on the encrypted data shards. If the accuracy of the aggregation result is insufficient, optimize it by adjusting the aggregation degree and increasing the interaction times. Finally, according to real-time requirements, adopt lightweight encryption algorithms or reduce the shard interaction times to improve the calculation efficiency. The present invention also includes a real-time risk monitoring mechanism, which starts security protection measures when anomalies are found. This method effectively solves the conflict between data privacy protection and aggregation effect and realizes secure and efficient data aggregation.

[0082] As Figure 5 shown, it is a structural diagram of an electronic device for secure aggregation of privacy computing data based on artificial intelligence provided by an embodiment of the present invention.

[0083] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a smart city big data fusion analysis cloud program. Among them, the processor 10 may be composed of integrated circuits in some embodiments. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and executing various functions of the electronic device and processing data by running or executing programs or modules stored in the memory 11, and calling data stored in the memory 11.

[0084] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. The memory 11 may be an internal storage unit of the electronic device in some embodiments, such as the mobile hard disk of the electronic device. The memory 11 may also be an external storage device of the electronic device in other embodiments, such as a plug-in mobile hard disk, a smart media card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of the smart city big data fusion analysis cloud program, but also to temporarily store data that has been output or will be output.

[0085] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to implement the connection and communication between the memory 11 and at least one processor 10, etc.

[0086] Only the electronic device with components is shown in the figure. Those skilled in the art can understand that the structure shown in the figure does not constitute a limitation on the electronic device, and it may include fewer or more components than shown in the figure, or combine some components, or have different component arrangements.

[0087] It should also be understood that in the embodiments herein, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0088] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this article.

[0089] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0090] In several embodiments provided in this document, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.

[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the objectives of the embodiments of this document.

[0092] In addition, each functional unit in the various embodiments of this document can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0093] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this document, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this document. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, and other media that can store program codes.

[0094] The above is only the preferred embodiment of one or more embodiments of this specification, and it is not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the protection scope of one or more embodiments of this specification.

Claims

1. A privacy computing data security aggregation method based on artificial intelligence, characterized in that, The method includes: According to the requirements of the business scenario, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation degree range of data sharding; Use the differential privacy algorithm to add noise to the data shards; Through homomorphic encryption technology, encrypt the data shards; According to the preset aggregation degree range, use the federated learning framework to perform distributed aggregation calculation on the encrypted data shards to obtain a preliminary aggregation result; If the accuracy of the preliminary aggregation result is lower than the preset threshold, adjust the aggregation degree, increase the interaction times of the data shards, and re-perform the aggregation calculation; If the adjusted aggregation result meets the accuracy requirements, decrypt the aggregated data to obtain the final result; The step of obtaining the preset data security level and accuracy requirements according to the requirements of the business scenario and determining the independence threshold and aggregation degree range of data sharding includes: Obtain the preset data security level and accuracy requirements; According to the data security level and accuracy requirements, use preset conditions to calculate the independence threshold of data sharding to obtain a first threshold range; For the first threshold range, combine the upper and lower limits of the preset aggregation degree to delimit the first aggregation degree range of data sharding; If the independence threshold of the data sharding is higher than the preset conditions, use the preset logical association rules to adjust the first aggregation degree range to obtain a second aggregation degree range; According to the second aggregation degree range, recalculate the independence threshold of the data sharding to obtain a second threshold range; Use the clustering analysis method in the machine learning algorithm to perform independence verification on the data sharding to determine whether it meets the preset conditions; If the independence threshold and aggregation degree range of the data sharding meet the preset business scenario requirements, determine the final data sharding scheme for data sharding; After obtaining the final result, the method further includes: according to the real-time requirement, determine whether it is necessary to optimize the aggregation process, use a lightweight encryption algorithm or reduce the shard interaction times to improve the calculation efficiency; Through system architecture design, embed the adjustment mechanism of the independence and aggregation degree of data sharding into the calculation process to achieve dynamic balance; According to the security mechanism design, monitor the data leakage risk during the aggregation process in real time. If an anomaly is found, terminate the calculation and start the security protection measures.

2. The method according to claim 1, wherein The step of using the differential privacy algorithm to add noise to the data shards includes: Obtain the preset data security level and accuracy requirements, and determine the initial independence threshold of data sharding; For each data shard, use the differential privacy algorithm to add an appropriate amount of noise to obtain the data shard after adding noise; If the independence threshold is lower than the preset threshold, adjust the noise addition parameter and re-perform the noise addition process until the independence threshold meets the preset conditions.

3. The method according to claim 1, characterized in that, The step of encrypting the data shards through homomorphic encryption technology includes: Obtain the data shards to be processed; For each of the data shards, use homomorphic encryption technology to perform encryption processing to obtain the encrypted data shards; Using the clustering analysis method in machine learning, the independence verification of the encrypted data shards is carried out to obtain the independence threshold of the encrypted data shards; Judge whether the independence threshold is lower than the preset threshold. If so, adjust the encryption parameters and return to perform the encryption process until the independence threshold meets the preset conditions.

4. The method according to claim 1, wherein According to the preset aggregation degree range, using the federated learning framework, the distributed aggregation calculation is performed on the encrypted data shards to obtain the preliminary aggregation result, including: Obtain the encrypted data shards, and according to the preset aggregation degree range, use the federated learning framework to perform distributed aggregation calculation on the encrypted data to obtain the preliminary aggregation result; For the preliminary aggregation result, use the clustering analysis method to calculate the independence value of the encrypted data. If the independence value is lower than the preset threshold, adjust the encryption parameters and re-encrypt the data shards to obtain the new encrypted data; Use the federated learning framework to perform distributed aggregation calculation on the encrypted data within the adjusted aggregation degree range to obtain the final preliminary aggregation result.

5. The method according to claim 4, wherein If the adjusted aggregation result meets the accuracy requirement, decrypt the aggregated data to obtain the final result, including: Obtain the final preliminary aggregation result and calculate the accuracy value of the final preliminary aggregation result; Compare the accuracy value with the preset accuracy threshold to judge whether the accuracy value meets the preset accuracy threshold; If the accuracy value meets the preset accuracy threshold, use the preset decryption algorithm to decrypt the final preliminary aggregation result to obtain the decrypted data; For the decrypted data, perform format conversion and standardization processing to obtain the standardized data; According to the standardized data, use the preset result generation algorithm to obtain the final result; Store the final result in the preset database and mark it as the completed processing status; Generate a processing log according to the storage result, and the key parameters and status information of the aggregation, decryption, and generation links are recorded in the processing log.

6. The method according to claim 1, characterized in that, According to the real-time requirement, judge whether it is necessary to optimize the aggregation process, adopt a lightweight encryption algorithm or reduce the number of shard interactions to improve the calculation efficiency, including: Obtain the real-time requirement parameters, calculate the processing time of the current aggregation process, and judge whether the processing time meets the preset time threshold; If the processing time does not meet the preset time threshold, obtain the encryption algorithm used in the aggregation process, and select the target lightweight encryption algorithm from the preset set of lightweight encryption algorithms according to the complexity of the encryption algorithm, and use the target lightweight encryption algorithm to replace the encryption algorithm; Obtain the number of shards and the number of interactions of the aggregated data, and obtain the optimized shard interaction strategy by reducing the number of shards or optimizing the shard interaction logic; According to the target lightweight encryption algorithm and the optimized shard interaction strategy, perform aggregation process simulation to obtain the optimized aggregation time; Determine whether the optimized aggregation time meets the preset time threshold. If not, adjust the parameters of the target lightweight encryption algorithm or the shard interaction strategy until the optimized aggregation time meets the preset time threshold; Solidify the aggregation process parameters that meet the preset time threshold to obtain an optimized aggregation processing module; Obtain the key performance indicators of the optimized aggregation processing module and generate an aggregation efficiency log.

7. The method according to claim 1, wherein Through system architecture design, embed the adjustment mechanism of the independence of data shards and the degree of aggregation into the calculation process to achieve dynamic balance, including: Obtain the current data shard status information of the system, where the data shard status information includes the number of data shards and the independence index of each shard; Judge whether the degree of aggregation corresponding to the current data shard status information meets the standard according to the preset aggregation degree threshold; If the degree of aggregation does not meet the standard, start the shard adjustment mechanism, and the shard adjustment mechanism adopts an embedded architecture design and is integrated into the calculation process of data processing; Through the shard adjustment mechanism, optimize the number of data shards and the independence index of each shard to obtain the adjusted data shard status information; Adopt a dynamic balance strategy to monitor the aggregation effect corresponding to the adjusted data shard status information in real time and generate a dynamic balance log; According to the dynamic balance log, use a parameter optimization algorithm to optimize the adjustment parameters in the shard adjustment mechanism to obtain optimized adjustment parameters; If the degree of aggregation corresponding to the adjusted data shard status information reaches the preset aggregation degree threshold, solidify the adjusted data shard status information and the optimized adjustment parameters to form an optimized calculation process module.

8. A data security aggregation system based on privacy computing, characterized in that, The system includes: A data security level and accuracy acquisition module, which is used to obtain the preset data security level and accuracy requirements according to the business scenario requirements, and determine the independence threshold of data shards and the aggregation degree range; A differential privacy noise addition module, which is used to add noise to data shards using the differential privacy algorithm to reduce the leakage risk while ensuring data independence; A homomorphic encryption processing module, which is used to encrypt data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation; A federated learning aggregation calculation module, which is used to perform distributed aggregation calculation on encrypted data shards using a federated learning framework according to the preset aggregation degree range to obtain a preliminary aggregation result; An aggregation result accuracy adjustment module, which is used to adjust the aggregation degree and increase the number of interactions of data shards and re-perform aggregation calculation if the accuracy of the preliminary aggregation result is lower than the preset threshold; An aggregated data decryption module, which is used to decrypt the aggregated data to obtain the final result if the adjusted aggregation result meets the accuracy requirements; A real-time optimization module, which is used to judge whether it is necessary to optimize the aggregation process according to real-time requirements, and adopt a lightweight encryption algorithm or reduce the number of shard interactions to improve the calculation efficiency; A dynamic balance embedding module, which is used to embed the adjustment mechanism of the independence of data shards and the degree of aggregation into the calculation process through system architecture design to achieve dynamic balance; A data leakage monitoring module is used to monitor the data leakage risk during the aggregation process in real time according to the security mechanism design. If an anomaly is detected, the calculation is terminated and security protection measures are initiated. According to the requirements of the business scenario, obtaining the preset data security level and accuracy requirements, and determining the independence threshold and aggregation degree range of data sharding, including: Obtaining the preset data security level and accuracy requirements; According to the data security level and accuracy requirements, using preset conditions, calculating the independence threshold of data sharding to obtain the first threshold range; For the first threshold range, combining the upper and lower limits of the preset aggregation degree, defining the first aggregation degree range of data sharding; If the independence threshold of the data sharding is higher than the preset conditions, using the preset logical association rules to adjust the first aggregation degree range to obtain the second aggregation degree range; According to the second aggregation degree range, recalculating the independence threshold of the data sharding to obtain the second threshold range; Using the clustering analysis method in the machine learning algorithm to verify the independence of the data sharding and determine whether it meets the preset conditions; If the independence threshold and aggregation degree range of the data sharding meet the preset business scenario requirements, determine the final data sharding scheme for data sharding.

Citation Information

Patent Citations

  • Federal learning optimization method based on neural network model privacy protection

    CN118643511A

  • Communication content security encryption method

    CN119071074A

  • Private data protection method and system based on homomorphic encryption and federated learning

    CN119513919A

  • Big data secure storage method and system based on cloud computing

    CN119720300A