Privacy computing data security aggregation method and system based on artificial intelligence
By adopting technologies such as differential privacy, homomorphic encryption and federated learning in privacy computing, the independence and aggregation degree of data sharding is dynamically adjusted, the contradiction between data security and business needs is solved, and efficient and secure data aggregation is achieved.
Patent Information
- Application Number
- CN202510421071.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the data security aggregation scenario of privacy computing, it is difficult to balance the independence and aggregation degree of data sharding, resulting in a contradiction between data security and business needs.
Using an artificial intelligence-based approach, data sharding is processed through differential privacy algorithms and homomorphic encryption technology to ensure data independence and security. Use the federated learning framework to perform distributed aggregation calculations, and dynamically adjust the degree of aggregation and the number of interactions of data shards according to business needs.
It realizes that the independence and aggregation degree of data sharding can be flexibly adjusted according to the needs of different business scenarios while ensuring data security, improves the accuracy and computing efficiency of data aggregation, and reduces the risk of data leakage.
Smart Images

Figure CN119961980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and in particular to a privacy computing data security aggregation method and system based on artificial intelligence. Background Art
[0002] In the data security aggregation scenario of privacy computing, there is a contradiction between data independence and aggregation degree. On the one hand, in order to ensure data security, each data shard needs to be able to independently complete part of the aggregation task to avoid cross-leakage of data during transmission and calculation. On the other hand, different business scenarios have different requirements for data accuracy and computing efficiency, and it is necessary to balance data security and business needs by controlling the degree of aggregation.
[0003] However, in actual business scenarios, it is often difficult to strike a balance between the independence and aggregation of data shards. When data shards are too independent, although data security can be maximized, the accuracy of the aggregation results may be reduced and fail to meet business needs. When the aggregation level is too high, although data accuracy and computing efficiency can be improved, the risk of data leakage may increase.
[0004] In addition, different business scenarios have different trade-offs between data security and business needs. In some scenarios with extremely high data security requirements, such as finance and medical fields, it may be necessary to sacrifice a certain amount of data accuracy and computing efficiency to ensure absolute data security. In some scenarios with high real-time requirements, such as real-time recommendations and risk control, it may be necessary to appropriately increase the degree of aggregation to meet the timeliness requirements of the business.
[0005] Therefore, how to flexibly adjust the independence and aggregation of data shards according to the needs of different business scenarios while ensuring data security is a key technical challenge facing privacy computing data security aggregation. This requires in-depth research and innovation in multiple aspects such as algorithm design, system architecture, and security mechanisms to find an optimal balance between data security and business needs. Summary of the invention
[0006] The present invention provides a privacy computing data security aggregation method based on artificial intelligence, which mainly includes: According to the business scenario requirements, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation range of data shards; Using differential privacy algorithms, noise is added to data shards to ensure data independence while reducing the risk of leakage. Through homomorphic encryption technology, data shards are encrypted to ensure the security of data during transmission and calculation; According to the preset aggregation degree range, the federated learning framework is used to perform distributed aggregation calculations on the encrypted data shards to obtain preliminary aggregation results; If the accuracy of the preliminary aggregation result is lower than the preset threshold, the aggregation degree is adjusted, the number of interactions of the data shards is increased, and the aggregation calculation is performed again; If the adjusted aggregation result meets the accuracy requirement, the aggregated data is decrypted to obtain the final result; According to the real-time requirements, determine whether the aggregation process needs to be optimized, adopt lightweight encryption algorithms or reduce the number of shard interactions to improve computing efficiency; Through system architecture design, the independence of data shards and the adjustment mechanism of aggregation degree are embedded in the computing process to achieve dynamic balance; According to the design of the security mechanism, the risk of data leakage in the aggregation process is monitored in real time. If an abnormality is found, the calculation is terminated and security protection measures are initiated.
[0007] The present invention provides a privacy computing data security aggregation system based on artificial intelligence, which mainly includes: The data security level and accuracy acquisition module is used to obtain the preset data security level and accuracy requirements according to the business scenario requirements, and determine the independence threshold and aggregation range of the data shards; The differential privacy noise adding module is used to add noise to data shards using the differential privacy algorithm, thereby reducing the risk of leakage while ensuring data independence; Homomorphic encryption processing module, used to encrypt data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation; The federated learning aggregation calculation module is used to perform distributed aggregation calculation on the encrypted data shards according to the preset aggregation degree range using the federated learning framework to obtain preliminary aggregation results; The aggregation result precision adjustment module is used to adjust the aggregation degree, increase the number of interactions of data shards, and re-perform the aggregation calculation if the precision of the preliminary aggregation result is lower than the preset threshold; The aggregated data decryption module is used to decrypt the aggregated data to obtain the final result if the adjusted aggregated result meets the accuracy requirement; The real-time optimization module is used to determine whether the aggregation process needs to be optimized according to real-time requirements, adopt lightweight encryption algorithms or reduce the number of shard interactions to improve computing efficiency; Dynamic balance embedding module, which is used to embed the independence and aggregation adjustment mechanism of data shards into the computing process through system architecture design to achieve dynamic balance; The data leakage monitoring module is used to monitor the data leakage risk in the aggregation process in real time according to the security mechanism design. If any abnormality is found, the calculation is terminated and security protection measures are initiated.
[0008] The technical solution provided by the embodiment of the present invention may have the following beneficial effects: The present invention discloses a privacy computing data security aggregation method based on artificial intelligence, which proposes a dynamic balance solution to the contradiction between data privacy protection and aggregation effect. First, according to the preset data security level and accuracy requirements, the independence threshold and aggregation degree range of the data shards are determined. Then, differential privacy and homomorphic encryption technology are used to process the data to reduce the risk of leakage while ensuring data independence. Next, the encrypted data shards are subjected to distributed aggregation calculations using a federated learning framework. If the aggregation result is not accurate enough, it is optimized by adjusting the aggregation degree and increasing the number of interactions. Finally, according to the real-time requirements, a lightweight encryption algorithm is used or the number of shard interactions is reduced to improve the computing efficiency. The present invention also includes a real-time risk monitoring mechanism to initiate security protection measures when anomalies are found. The method effectively solves the conflict between data privacy protection and aggregation effect, and realizes safe and efficient data aggregation. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 This is a flow chart of a privacy computing data security aggregation method based on artificial intelligence of the present invention; Figure 2 It is a schematic diagram of a specific embodiment of a privacy computing data security aggregation method based on artificial intelligence of the present invention; Figure 3 This is another schematic diagram of a specific embodiment of a privacy computing data security aggregation method based on artificial intelligence of the present invention; Figure 4 This is a schematic diagram of the structure of a privacy computing data security aggregation system based on artificial intelligence of the present invention; Figure 5 A structural diagram of an electronic device for privacy computing data security aggregation based on artificial intelligence provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0010] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0011] like Figure 1-3As shown, a privacy computing data security aggregation method based on artificial intelligence in this embodiment may specifically include: S101. According to the business scenario requirements, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation range of the data shards.
[0012] Obtain a preset data security level and precision requirement; based on the data security level and precision requirement, use preset conditions to calculate the independence threshold of the data shard to obtain a first threshold range; for the first threshold range, combine the preset upper and lower limits of the degree of aggregation to define a first degree of aggregation range for the data shard; if the independence threshold of the data shard is higher than the preset condition, use a preset logical association rule to adjust the first degree of aggregation range to obtain a second degree of aggregation range; based on the second degree of aggregation range, recalculate the independence threshold of the data shard to obtain a second threshold range; use a clustering analysis method in a machine learning algorithm to verify the independence of the data shard to determine whether it meets the preset conditions; if the independence threshold and degree of aggregation range of the data shard meet the preset business scenario requirements, determine the final data sharding solution.
[0013] Regarding the acquisition of data security level and accuracy requirements, it can be understood that the business needs for security and accuracy must be clarified before data sharding. For example, in the financial risk control scenario, it is assumed that the data security level is defined as "high" because it involves user privacy data, such as transaction records; the accuracy requirement is "medium" because it only needs to identify risk patterns rather than accurate numerical predictions. Based on this, the preset condition may be that after data sharding, it is necessary to ensure that a single piece cannot reversely derive complete information, while maintaining the accuracy of risk analysis above 85%. The calculation of the independence threshold involves multiple factors. For example, indicators such as correlation between data and information entropy can be considered. Assuming that there is a set of user browsing records, if the browsing patterns of different users are highly similar, then the independence is low; conversely, if the browsing patterns are significantly different, then the independence is high. The determination of the first threshold range may consider the weighted average of these factors. The delineation of the aggregation degree range requires a balance between data privacy and analytical utility. Taking the retail industry as an example, too high an aggregation degree may make it impossible to identify the purchasing tendencies of individual consumers, while too low an aggregation degree may expose sensitive personal information. Therefore, it is necessary to determine the appropriate range based on business needs and regulatory requirements. When the independence threshold is higher than expected, the degree of aggregation may need to be adjusted. For example, when analyzing social network data, if it is found that the user groups are highly independent, the degree of aggregation may need to be reduced to retain more valuable information. This adjustment may involve redefining data grouping or adopting a more fine-grained analysis method. Recalculating the independence threshold is an iterative process. For example, when analyzing telecommunications user data, the initial threshold may be set based on call duration. After adjustment, it may be necessary to include multi-dimensional information such as SMS usage frequency and data traffic to obtain a more comprehensive second threshold range. Cluster analysis is an effective way to verify the independence of data shards. For example, when performing customer grouping, the K-means algorithm can be used. If customers in different shards show obvious clusters in dimensions such as consumption habits and age distribution, it means that the shards have good independence. The determination of the final data sharding plan requires comprehensive consideration of multiple factors. Taking the customer data of an insurance company as an example, it may be necessary to find a balance between customer privacy protection, risk assessment accuracy, and market segmentation effect. Only when the sharding plan can meet data security, analysis accuracy, and business needs at the same time can it be identified as the final plan. This data segmentation and verification process not only improves the efficiency and accuracy of data analysis, but also plays an important role in protecting privacy and meeting regulatory requirements. Through refined data processing, companies can better understand customer needs, optimize business processes, and ensure compliance operations.
[0014] S102. Use a differential privacy algorithm to add noise to data shards to reduce the risk of leakage while ensuring data independence.
[0015] Obtain the preset data security level and accuracy requirements, and determine the initial independence threshold of the data shards. For each data shard, use the differential privacy algorithm to add an appropriate amount of noise to obtain the data shard after adding noise. If the independence threshold is lower than the preset threshold, adjust the noise addition parameters and re-add noise until the independence threshold meets the preset conditions. Use the clustering analysis method in machine learning to verify the independence of the adjusted data shards. If the independence threshold of the data shard does not meet the preset conditions, return to the noise addition process; if it meets the preset conditions, adjust the aggregation range according to the verification results until the business scenario requirements are met.
[0016] Exemplarily, the preset data security level and accuracy requirements are obtained, and the initial independence threshold of the data shard is determined in combination with the business scenario requirements. A differential privacy algorithm is used to add an appropriate amount of noise to each data shard to obtain the data shard after adding noise. According to the data shard after adding noise, its independence threshold is calculated. If the independence threshold is lower than the preset threshold, the noise addition parameter is adjusted, and the noise addition process is performed again until the independence threshold meets the preset conditions. For the adjusted data shard, the preset upper and lower limits of the degree of aggregation are obtained to determine the degree of aggregation range of the data shard. The degree of aggregation range is fine-tuned in combination with the logical association rule to obtain the adjusted degree of aggregation range. The cluster analysis method in machine learning is used to verify the independence of the adjusted data shard. If the independence threshold of the data shard does not meet the preset conditions, the noise addition process is returned to the noise addition process. If the independence threshold of the data shard meets the preset conditions, it is further determined whether its degree of aggregation range meets the business scenario requirements. If not, the degree of aggregation range is adjusted according to the verification result until the business scenario requirements are met. Determine the final data sharding scheme including noise addition parameters, independence threshold and aggregation degree range. Apply the final scheme to the data sharding process to obtain processed data shards. Persistently store the processed data shards and combine them with data access control mechanisms. Regularly re-evaluate and adjust the data shards.
[0017] Data security level and accuracy requirements are key factors in determining the initial independence threshold of data shards. For example, in the financial industry, customer transaction data is often classified as highly sensitive information, requiring the highest level of protection and accuracy. Suppose a bank sets the initial independence threshold of transaction data to 0.8 (range 0-1), which means that the transaction behaviors of different customer groups should be highly distinguishable. Differential privacy algorithms play an important role in data sharding. It protects individual privacy by adding carefully designed random noise while maintaining the overall statistical properties of the data. Taking bank transaction data as an example, noise that conforms to the Laplace distribution may be added to the monthly transaction total of each customer. The initial noise parameter ε may be set to 0.1, indicating a strong degree of privacy protection. After adding noise, it is necessary to evaluate whether the independence of the data shards reaches the preset threshold. If the independence threshold is lower than 0.8, such as only 0.6, it means that the transaction behaviors of different customer groups are not distinguishable enough. At this time, it is necessary to adjust the noise addition parameter, perhaps adjusting ε to 0.2, to reduce the impact of noise on data characteristics. Repeat this process until the preset threshold is reached. Cluster analysis is an effective way to verify the independence of data shards. For example, the K-means algorithm can be used to cluster the adjusted customer transaction data. If customers can be clearly divided into different groups such as high-frequency small transactions and low-frequency large transactions, it means that the data shards have good independence. If the clustering results show that the data shards still do not meet the independence requirements, it is necessary to return to the noise addition link. This process may require multiple iterations. For example, it may be found that simply adjusting the ε parameter is not enough to improve independence. At this time, it may be necessary to consider adding noise to other dimensions (such as transaction frequency and transaction type). Finally, adjust the aggregation range based on the verification results. In the case of a bank, it may be found that the originally planned aggregation method of 100 customers in a group does not meet business needs because it will cover up important customer behavior patterns. Through adjustments, it may eventually be determined that the aggregation method of 50 customers in a group can protect personal privacy and provide sufficient insights for precision marketing. This data processing method can not only improve the accuracy of data analysis, but also play an important role in protecting customer privacy and meeting financial regulatory requirements. Through refined data segmentation and verification, banks can better understand customer needs, optimize product design, and ensure compliance operations.
[0018] S103. Encrypt data shards through homomorphic encryption technology to ensure data security during transmission and calculation.
[0019] Obtain the data shards to be processed; for each of the data shards, perform encryption processing using homomorphic encryption technology to obtain encrypted data shards; use the clustering analysis method in machine learning to verify the independence of the encrypted data shards to obtain the independence threshold of the encrypted data shards; determine whether the independence threshold is lower than the preset threshold, and if so, adjust the encryption parameters and return to perform encryption processing until the independence threshold meets the preset conditions; determine the range of aggregation degree according to the independence verification result; determine whether the range of aggregation degree meets the business scenario requirements, and if not, adjust the range of aggregation degree and return to perform independence verification until the business scenario requirements are met.
[0020] For example, Figure 3As shown, for each data shard, homomorphic encryption technology is used to perform encryption processing to obtain the encrypted data shard. If the independence threshold is lower than the preset threshold, the encryption parameters are adjusted and the encryption processing is performed again until the independence threshold meets the preset conditions. The clustering analysis method in machine learning is used to verify the independence of the encrypted data shard. If the independence threshold of the data shard does not meet the preset conditions, the encryption processing link is returned; if the preset conditions are met, the aggregation degree range is adjusted according to the verification results until the business scenario requirements are met. Homomorphic encryption is an advanced encryption technology that allows calculations to be performed directly on encrypted data without decryption. This feature makes it widely used in data analysis and machine learning. Data sharding processing is an important means to protect sensitive information. In the financial field, when banks process customer transaction records, they need to strike a balance between privacy protection and data availability. After obtaining the data shard to be processed, homomorphic encryption technology is used for encryption processing. Homomorphic encryption allows calculations to be performed directly on ciphertext without decryption, which greatly improves data security. For example, a bank may homomorphically encrypt information such as a customer's monthly transaction total and transaction frequency. Initial encryption parameters may be set to a high security level, such as using a 2048-bit key. After encryption, the independence of the data shards needs to be verified. This step uses cluster analysis methods in machine learning, such as the K-means algorithm. Clustering on encrypted data can evaluate the differentiation of different customer groups without exposing the original data. Suppose the bank sets the independence threshold to 0.75, indicating that it expects a high degree of differentiation between different customer groups. If the clustering results show that the independence threshold is only 0.6, it means that the encrypted data shards may be too vague to distinguish different customer groups. At this time, the encryption parameters need to be adjusted. The bank may reduce the encryption strength, such as using a 1024-bit key, to achieve a new balance between protecting privacy and retaining data characteristics. After the adjustment, encryption and independence verification are performed again until the preset threshold of 0.75 is reached. This process may require multiple iterations, each time fine-tuning the encryption parameters until the optimal configuration is found. Based on the independence verification results, the aggregation level range is determined. This step aims to find an appropriate level of data aggregation that can protect individual privacy while meeting business needs. For example, the bank may initially set every 500 customers as an aggregation group. However, through analysis, it was found that this aggregation method may mask important customer behavior patterns and is not conducive to precision marketing. Therefore, the range of aggregation needs to be adjusted. Banks may try to reduce the aggregation group to one group of 200 customers and then re-verify independence. If the new aggregation method meets the independence requirements (the threshold is still greater than or equal to 0.75) and can provide sufficient insights for business decisions, this range of aggregation can be adopted. This process may require multiple adjustments and verifications until a balance is found. Through this refined data processing method, banks can both protect customer privacy and optimize business operations.For example, in this smaller aggregate group, the bank may have discovered an emerging high-value customer group, which provides valuable insights for precision marketing and product development. At the same time, because the data is always encrypted, even if a data leak occurs during the analysis process, the customer's personal information will not be directly exposed, greatly reducing privacy risks.
[0021] S104. According to a preset aggregation degree range, a federated learning framework is used to perform distributed aggregation calculations on the encrypted data shards to obtain preliminary aggregation results.
[0022] Obtain the encrypted data shards, and use a federated learning framework to perform distributed aggregation calculations on the encrypted data according to a preset aggregation degree range to obtain preliminary aggregation results; for the preliminary aggregation results, use a clustering analysis method to calculate the independence value of the encrypted data; if the independence value is lower than a preset threshold, adjust the encryption parameters, re-encrypt the data shards, and obtain new encrypted data; use a federated learning framework to perform distributed aggregation calculations on the encrypted data within the adjusted aggregation degree range to obtain the final preliminary aggregation results.
[0023] Exemplarily, a federated learning framework is used to obtain encrypted data shards, and distributed aggregation calculation is performed on the encrypted data according to a preset aggregation degree range to obtain a preliminary aggregation result. For the preliminary aggregation result, a cluster analysis method is used to calculate the independence value of the encrypted data, and it is determined whether the independence value is lower than a preset threshold. If it is lower than the preset threshold, the encryption parameters are adjusted and the encryption process is re-executed. According to the adjusted encryption parameters, the data shards are re-encrypted to obtain new encrypted data, and the new encrypted data is distributedly aggregated using the federated learning framework to obtain an updated preliminary aggregation result. For the updated preliminary aggregation result, the independence value is calculated again to determine whether the independence value meets the preset conditions. If it does, the final aggregation degree range is determined. According to the final aggregation degree range, it is determined whether the range meets the business requirements. If not, the aggregation degree range is adjusted and the independence value calculation is re-executed. The federated learning framework is used to perform distributed aggregation calculation on the encrypted data within the adjusted aggregation degree range to obtain a final aggregation result. According to the final aggregation result, the regression analysis method in machine learning is used to optimize the aggregated data to obtain an optimized aggregation result. Federated learning is a distributed machine learning framework that can perform multi-party collaborative computing while protecting data privacy. In this scenario, federated learning is used to perform distributed aggregation computing on encrypted data, which not only protects the confidentiality of the data but also improves computing efficiency. Taking the financial field as an example, suppose that several banks want to jointly build a credit scoring model, but are unwilling to directly share customer data. Each bank first encrypts and shards its own customer data. Then, according to the preset aggregation degree range (such as every 1,000 customers as a group), each bank uses the federated learning framework for distributed computing. In this process, the parties only exchange model parameters instead of raw data, thereby protecting their respective data privacy in collaboration. After the preliminary aggregation results are obtained, independence verification is required. Cluster analysis methods, such as hierarchical clustering algorithms, are used here to evaluate the independence of encrypted data. The independence value reflects the degree of distinction between different customer groups. Assume that the preset threshold is 0.8, indicating that different customer groups are expected to have high distinguishability. If the calculated independence value is only 0.6, it means that the current encryption scheme may be too strict, resulting in excessive masking of data features. At this time, it is necessary to adjust the encryption parameters and reprocess the data. For example, you can try to reduce the security level of homomorphic encryption from a 2048-bit key to a 1024-bit key. This adjustment aims to find a balance between data availability and privacy protection. After the adjustment, re-encrypt the data to obtain a new encrypted data set. Next, use the adjusted aggregation range (such as one group for every 800 customers) to perform distributed aggregation calculations on the new encrypted data. The purpose of this step is to find the best aggregation level that can protect privacy while providing valuable information.After multiple iterations and adjustments, we finally get an aggregated result that meets the independence requirements and can provide sufficient insights for business decisions. The advantage of this method is that it can provide financial institutions with valuable group insights while protecting individual privacy. For example, banks may find that customer groups in certain regions or age groups have higher credit scores, which can guide them to formulate more precise marketing strategies and risk management policies. At the same time, since the data is always encrypted throughout the process, even if a data leak occurs, the customer's personal information will not be directly exposed, greatly reducing privacy risks.
[0024] S105. If the accuracy of the preliminary aggregation result is lower than the preset threshold, adjust the aggregation degree, increase the number of interactions of the data shards, and re-perform the aggregation calculation.
[0025] The accuracy of the aggregation result is used to characterize the closeness of the aggregation result (such as mean, variance, distribution characteristics) to the true value of the original data. It can be calculated in many ways, such as: Absolute Error: |aggregate value − true value|; Relative Error: |Aggregate value − True value | / True value; Confidence Interval: In differential privacy, the fluctuation range of the result after noise is added (such as the error boundary at an 85% confidence level).
[0026] For example, in the application of federated learning in the field of financial risk control, confidence intervals are used to characterize. Assuming that the accuracy of the preliminary aggregation result is 75%, which is lower than the preset threshold of 85%, the degree of aggregation needs to be adjusted to improve the accuracy. First, the number of interactions of sharded data is increased from the default 5 to 10 times to enhance the ability to extract data features. A distributed optimization algorithm based on gradient descent is used, the learning rate is set to 01, the number of iterations is 100, and the encrypted sharded data is re-aggregated. During the calculation process, differential privacy technology is used, Laplace noise is added, and the noise parameter ε is set to 1 to further protect data privacy. After recalculation, the updated aggregation result is obtained, and the accuracy is improved to 82%. Then, the principal component analysis (PCA) method is used to reduce the dimension of the updated aggregation result, retaining 95% of the variance information to reduce the data dimension and improve the calculation efficiency.
[0027] The K-means clustering algorithm was used to set the number of clusters to 5, and cluster analysis was performed on the reduced-dimensional data. The ratio of the intra-cluster distance to the inter-cluster distance was calculated to evaluate the rationality of the data distribution.
[0028] If the ratio is lower than the preset threshold of 7, the aggregation degree is further adjusted, the number of interactions of the shard data is increased to 15 times, and the aggregation calculation is performed again.
[0029] Finally, after multiple iterations and optimizations, an aggregation result with an accuracy of 88% was obtained, which met business requirements.
[0030] S106. If the adjusted aggregation result meets the accuracy requirement, the aggregated data is decrypted to obtain the final result.
[0031] Obtain the aggregation result and calculate the precision value of the aggregation result; compare the precision value with a preset precision threshold to determine whether the precision value meets the preset precision threshold; if the precision value meets the preset precision threshold, use a preset decryption algorithm to decrypt the aggregation result to obtain decrypted data; perform format conversion and standardization on the decrypted data to obtain standardized data; use a preset result generation algorithm based on the standardized data to obtain a final result; store the final result in a preset database and mark it as completed; generate a processing log based on the stored result, wherein the processing log records key parameters and status information of the aggregation, decryption, and generation links.
[0032] Exemplarily, for the aggregation result, calculate its precision value, compare it with the preset precision threshold, and determine whether it meets the precision requirement. If the precision value meets the preset threshold, use the preset decryption algorithm to decrypt the aggregation result to generate decrypted data. Perform format conversion and standardization on the decrypted data to generate standardized data. Based on the standardized data, use the preset result generation algorithm to generate the final result. Store the final result in the preset database and mark it as completed. Generate a processing log based on the stored results to record key parameters and status information of aggregation, decryption, generation, and other links.
[0033] S107. According to the real-time requirements, determine whether the aggregation process needs to be optimized, adopt a lightweight encryption algorithm or reduce the number of shard interactions to improve computing efficiency.
[0034] Acquire real-time requirement parameters, calculate the processing time of the current aggregation process, and determine whether the processing time meets a preset time threshold; if the processing time does not meet the preset time threshold, obtain the encryption algorithm used in the aggregation process, select a target lightweight encryption algorithm from a preset set of lightweight encryption algorithms according to the complexity of the encryption algorithm, and replace the encryption algorithm with the target lightweight encryption algorithm; obtain the number of shards and the number of interactions of the aggregated data, and obtain an optimized shard interaction strategy by reducing the number of shards or optimizing the shard interaction logic; simulate the aggregation process according to the target lightweight encryption algorithm and the optimized shard interaction strategy to obtain the optimized aggregation time; determine whether the optimized aggregation time meets the preset time threshold, and if not, adjust the parameters of the target lightweight encryption algorithm or the shard interaction strategy until the optimized aggregation time meets the preset time threshold; solidify the aggregation process parameters that meet the preset time threshold to obtain an optimized aggregation processing module; obtain key performance indicators of the optimized aggregation processing module, and generate an aggregation efficiency log.
[0035] Exemplarily, obtain the real-time requirement parameters, calculate the processing time of the current aggregation process, and determine whether the aggregation efficiency meets the preset time threshold. If not, enter the optimization phase. Analyze the complexity of the encryption algorithm used in the aggregation process, and select a lightweight encryption algorithm to replace the existing algorithm to reduce computing resource consumption. Evaluate the number of shards and the number of interactions of the aggregated data, and reduce the data transmission delay by reducing the number of shards or optimizing the shard interaction logic. Combine the lightweight encryption algorithm and the optimized shard interaction strategy, re-simulate the aggregation process, record the aggregation time, and verify the efficiency improvement effect. If the optimized aggregation time still does not meet the real-time requirements, further adjust the encryption algorithm parameters or the shard strategy until the preset time threshold is met. Solidify the optimized aggregation process parameters, update the aggregation processing module, and ensure that subsequent aggregation operations are performed according to the optimized process. Monitor the optimized aggregation process, record key performance indicators in real time, and generate aggregation efficiency logs for subsequent analysis and further optimization reference.
[0036] The real-time requirement parameter is a key indicator for measuring the efficiency of the aggregation process. In a financial trading system, assuming that the preset time threshold is 100 milliseconds, the current aggregation process takes 150 milliseconds, exceeding the threshold by 50%. At this time, the aggregation algorithm needs to be optimized to improve efficiency. The choice of encryption algorithm directly affects the processing time. For example, the original RSA algorithm has high computational complexity, and it can be considered to be replaced with a lightweight elliptic curve encryption algorithm. While ensuring security, elliptic curve encryption greatly reduces computational overhead and is expected to shorten the processing time to about 120 milliseconds. The sharding strategy of aggregated data is also an important factor affecting efficiency. Assume that 100 shards are originally used, and each shard requires 3 interactions. By adjusting the number of shards to 50 and optimizing the interaction logic to reduce the number of interactions to 2, the processing time can be further shortened. This optimization not only reduces the data transmission overhead, but also reduces the complexity of parallel processing. Simulating the optimized aggregation process is a key step in verifying the effect. Using historical data for simulation testing can estimate the performance of the new solution. If the optimized aggregation time still does not reach the target of 100 milliseconds, further adjustment of parameters is required. For example, you can try to reduce the security parameters of elliptic curve encryption to further improve efficiency while ensuring security. Solidification of aggregation process parameters is an important measure to ensure long-term stability. Write the optimized parameters, such as encryption algorithm type, key length, number of shards, etc., to the configuration file or database to ensure that the system can still run efficiently after restart. This method is also convenient for subsequent version control and rollback operations. Generating aggregation efficiency logs is crucial for system monitoring and continuous optimization. The logs should contain key performance indicators, such as average processing time, peak processing time, resource utilization, etc. By analyzing the trends of these indicators, performance bottlenecks can be discovered in a timely manner, providing a basis for further optimization. For example, if it is found that the processing time in certain periods is significantly extended, it may mean that hardware resources need to be increased or the load balancing strategy needs to be optimized. The entire optimization process embodies an iterative performance tuning method. From identifying problems, proposing optimization solutions, verifying effects to solidifying parameters, a closed loop is formed. This method is not only applicable to the optimization of the aggregation process, but can also be extended to other high-performance computing scenarios, such as real-time data analysis, large-scale parallel computing, and other fields. Through continuous optimization and adjustment, system performance can be continuously improved to meet growing business needs.
[0037] S108. Through system architecture design, the independence of data shards and the adjustment mechanism of the degree of aggregation are embedded in the computing process to achieve dynamic balance.
[0038] The current data shard status information of the system is obtained, and the data shard status information includes the number of data shards and the independence index of each shard; according to the preset aggregation degree threshold, it is judged whether the aggregation degree corresponding to the current data shard status information meets the standard; if the aggregation degree does not meet the standard, the shard adjustment mechanism is started, and the shard adjustment mechanism adopts an embedded architecture design and is integrated into the calculation process of data processing; through the shard adjustment mechanism, the number of data shards and the independence index of each shard are optimized to obtain the adjusted data shard status information; a dynamic balance strategy is adopted to monitor the aggregation effect corresponding to the adjusted data shard status information in real time, and a dynamic balance log is generated; according to the dynamic balance log, a parameter optimization algorithm is adopted to optimize the adjustment parameters in the shard adjustment mechanism to obtain the optimized adjustment parameters; if the aggregation degree corresponding to the adjusted data shard status information reaches the preset aggregation degree threshold, the adjusted data shard status information and the optimized adjustment parameters are solidified to form an optimized calculation process module.
[0039] Exemplarily, the acquisition of data shard status information is the key starting point for optimizing the aggregation process. The system collects the number of shards and independence indicators through real-time monitoring, which reflect the balance and independence of data distribution. For example, in a distributed database system, there may be 100 data shards, and the independence index of each shard ranges from 0.6 to 0.9. The higher the independence index, the less data overlap between shards, which is conducive to improving parallel processing efficiency. The aggregation threshold is an important criterion for measuring whether the data shard status meets the system requirements. Assume that the aggregation threshold set by the system is 0.8, and the aggregation degree of the current system is 0.75, which means that the shard adjustment mechanism needs to be started. The shard adjustment mechanism adopts an embedded architecture and is directly integrated into the data processing process. It can realize real-time and dynamic adjustment and reduce additional system overhead. The core of the shard adjustment mechanism is to optimize the number of data shards and independence indicators. In the above example, the system may decide to reduce the number of shards from 100 to 80, while improving the independence index of each shard. This adjustment can be achieved by redistributing data, merging small shards, or splitting large shards. After the adjustment, assuming that the average independence index is increased to 0.85, the overall convergence of the system will also increase. The application of the dynamic balancing strategy ensures that the system can maintain optimal performance after the adjustment. Through real-time monitoring, the system generates a dynamic balancing log to record key indicators such as shard size, query response time, and resource utilization. For example, the log may show that in the first 30 minutes after the adjustment, the query response time decreased by 20% on average, but the resource utilization of some shards increased to more than 90%. Based on the dynamic balancing log, the parameter optimization algorithm will further adjust the parameters of the sharding mechanism. This may include adjusting the threshold of the shard size, the calculation weight of the independence index, etc. For example, the system may find that increasing the upper limit of the shard size by 10% can control the resource utilization within the ideal range without significantly increasing the query time. When the adjusted convergence reaches the preset threshold (such as 0.8), the system will solidify the optimized configuration. This includes the new shard status information (such as 80 shards, average independence index 0.85) and the optimized adjustment parameters. This information is integrated into the optimized computing process module to ensure that the system can maintain high efficiency in subsequent operations. This dynamic optimization process not only improves the overall performance of the system, but also enhances its adaptability and scalability. Through continuous monitoring and adjustment, the system can cope with challenges such as data volume growth and query pattern changes, and always maintain optimal operation.
[0040] S109. According to the security mechanism design, the data leakage risk in the aggregation process is monitored in real time. If an abnormality is found, the calculation is terminated and security protection measures are initiated.
[0041] Acquire the real-time data stream in the aggregation process, wherein the real-time data stream includes the data transmission rate and data content characteristics; determine whether the data leakage risk index in the real-time data stream exceeds the preset security threshold; if the data leakage risk index exceeds the threshold, terminate the current aggregation calculation process and record the abnormal data fragment; according to the security protection log, use the machine learning algorithm to optimize the parameters in the security protection mechanism to obtain the optimized security protection parameters; if the optimized security protection parameters can effectively reduce the data leakage risk, solidify to form an optimized security protection module.
[0042] For example, Figure 2 As shown, a real-time data stream in an aggregation process is obtained, and the real-time data stream includes a data transmission rate and data content characteristics; a preset security threshold is used to determine whether a data leakage risk index in the real-time data stream exceeds a standard; if the data leakage risk index exceeds a standard, the current aggregation calculation process is immediately terminated and the abnormal data fragment is recorded; a security protection mechanism is started, and the security protection mechanism includes a data encryption module and an access control module; the abnormal data fragment is encrypted by the data encryption module to generate an encrypted data packet; the access control module is used to limit the access rights to the encrypted data packet to allow only authorized users to access; the implementation effect of the security protection mechanism is monitored in real time to generate a security protection log; according to the security protection log, a machine learning algorithm is used to optimize the parameters in the security protection mechanism to obtain optimized security protection parameters; if the optimized security protection parameters can effectively reduce the risk of data leakage, the optimized security protection parameters are solidified to form an optimized security protection module; the data stream in the aggregation process is continuously monitored to ensure the real-time effectiveness of the security protection module; the preset security threshold is regularly updated to adapt to the ever-changing data environment and security threats; the security protection log is regularly reviewed by a preset security audit mechanism to identify potential security vulnerabilities and repair them.
[0043] During the data aggregation process, real-time monitoring of data flows is crucial to ensure information security. The system builds a comprehensive portrait of real-time data flows by collecting indicators such as data transmission rate and content characteristics. For example, in a financial data processing system, the amount of data transmitted per second is normally about 500KB, and it mainly contains digital transaction records. If the data transmission rate suddenly surges to 2MB per second and a large amount of text data appears in the content, this may indicate a potential risk of data leakage. The system continuously evaluates the data leakage risk indicator and compares it with the preset security threshold. Assuming that the risk threshold set by the system is 0.7 (the full score is 1), when an abnormal data pattern is detected that causes the risk indicator to rise to 0.8, the system will immediately trigger the security protection mechanism. This mechanism is designed to respond quickly to potential threats and minimize the possibility of data leakage. Once the risk indicator exceeds the threshold, the system will immediately interrupt the current aggregation calculation process. Although this measure may temporarily affect the efficiency of data processing, it is necessary from the perspective of information security. At the same time, the system will record the data fragments that cause the anomaly in detail, including information such as timestamps, data characteristics, and risk scores. These records not only help subsequent security analysis, but also provide valuable samples for optimizing security protection mechanisms. The accumulation of security protection logs provides a wealth of training data for machine learning algorithms. The system may use algorithms such as support vector machines (SVM) or deep learning networks to continuously optimize security protection parameters by analyzing the patterns of historical abnormal data. For example, the algorithm may find that certain specific data patterns are highly correlated with high-risk events, thereby adjusting the weights of the corresponding features in the risk assessment model. The optimized security protection parameters need to be strictly verified. The system may use historical data for backtesting or conduct simulation tests in an isolated environment. If the new parameters perform well in the test and can reduce the false positive rate by 30% while maintaining a true positive detection rate of more than 99%, then the set of parameters is considered effective. Finally, the verified optimized parameters are integrated into the system's security protection module. This process not only improves the security of the system, but also enhances its adaptability to new threats. Through this continuous optimization cycle, the data aggregation system can always maintain a high level of security protection capabilities while ensuring efficiency.
[0044] like Figure 4 As shown, another embodiment of the present invention provides a privacy computing data security aggregation system based on artificial intelligence, which mainly includes: The data security level and accuracy acquisition module is used to obtain the preset data security level and accuracy requirements according to the business scenario requirements, and determine the independence threshold and aggregation range of the data shards; The differential privacy noise adding module is used to add noise to data shards using the differential privacy algorithm, thereby reducing the risk of leakage while ensuring data independence; Homomorphic encryption processing module, used to encrypt data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation; The federated learning aggregation calculation module is used to perform distributed aggregation calculation on the encrypted data shards according to the preset aggregation degree range using the federated learning framework to obtain preliminary aggregation results; The aggregation result precision adjustment module is used to adjust the aggregation degree, increase the number of interactions of data shards, and re-perform the aggregation calculation if the precision of the preliminary aggregation result is lower than the preset threshold; The aggregated data decryption module is used to decrypt the aggregated data to obtain the final result if the adjusted aggregated result meets the accuracy requirement; The real-time optimization module is used to determine whether the aggregation process needs to be optimized according to real-time requirements, adopt lightweight encryption algorithms or reduce the number of shard interactions to improve computing efficiency; Dynamic balance embedding module, which is used to embed the independence and aggregation adjustment mechanism of data shards into the computing process through system architecture design to achieve dynamic balance; The data leakage monitoring module is used to monitor the data leakage risk in the aggregation process in real time according to the security mechanism design. If any abnormality is found, the calculation is terminated and security protection measures are initiated.
[0045] In summary, the embodiment of the present invention discloses a privacy computing data security aggregation method and system based on artificial intelligence. The method proposes a dynamic balance solution to the contradiction between data privacy protection and aggregation effect. First, according to the preset data security level and accuracy requirements, the independence threshold and aggregation degree range of the data shards are determined. Then, differential privacy and homomorphic encryption technology are used to process the data to reduce the risk of leakage while ensuring data independence. Next, the encrypted data shards are distributedly aggregated using a federated learning framework. If the aggregation result is not accurate enough, it is optimized by adjusting the aggregation degree and increasing the number of interactions. Finally, according to the real-time requirements, a lightweight encryption algorithm is used or the number of shard interactions is reduced to improve the computing efficiency. The present invention also includes a real-time risk monitoring mechanism to initiate security protection measures when an anomaly is found. This method effectively solves the conflict between data privacy protection and aggregation effect, and achieves safe and efficient data aggregation.
[0046] like Figure 5 As shown, it is a structural diagram of an electronic device for privacy computing data security aggregation based on artificial intelligence provided by one embodiment of the present invention.
[0047] The electronic device may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13, and may also include a computer program stored in the memory 11 and run on the processor 10, such as a smart city big data fusion analysis cloud program. In some embodiments, the processor 10 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, which connects the various components of the entire electronic device using various interfaces and lines, and executes various functions of the electronic device and processes data by running or executing programs or modules stored in the memory 11, and calling data stored in the memory 11.
[0048] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (SmartMediaCard, SMC), a secure digital (SecureDigital, SD) card, a flash card (FlashCard), etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit of the electronic device and an external storage device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the smart city big data fusion analysis cloud program, but also can be used to temporarily store data that has been output or is to be output.
[0049] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize connection and communication between the memory 11 and at least one processor 10, etc.
[0050] The figure only shows an electronic device with components. Those skilled in the art will understand that the structure shown in the figure does not constitute a limitation on the electronic device, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0051] It should also be understood that in the embodiments of this article, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0052] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this article.
[0053] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0054] In the several embodiments provided herein, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0055] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this article.
[0056] In addition, each functional unit in each embodiment of this invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional unit.
[0057] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this article is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of this article. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0058] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A privacy computing data security aggregation method based on artificial intelligence, characterized in that: The method comprises: According to the business scenario requirements, obtain the preset data security level and accuracy requirements, and determine the independence threshold and aggregation range of data shards; Use differential privacy algorithm to add noise to data shards; Encrypt data shards through homomorphic encryption technology; According to the preset aggregation degree range, the federated learning framework is used to perform distributed aggregation calculations on the encrypted data shards to obtain preliminary aggregation results; If the accuracy of the preliminary aggregation result is lower than the preset threshold, the aggregation degree is adjusted, the number of interactions of the data shards is increased, and the aggregation calculation is performed again; If the adjusted aggregation result meets the accuracy requirement, the aggregated data is decrypted to obtain the final result.
2. The method according to claim 1, characterized in that According to the business scenario requirements, the preset data security level and accuracy requirements are obtained, and the independence threshold and aggregation range of the data shards are determined, including: Obtain the preset data security level and accuracy requirements; According to the data security level and accuracy requirements, using preset conditions, calculating the independence threshold of the data shards to obtain a first threshold range; For the first threshold range, combined with the preset upper and lower limits of the aggregation degree, a first aggregation degree range of the data shards is defined; If the independence threshold of the data shards is higher than a preset condition, a preset logical association rule is used to adjust the first aggregation degree range to obtain a second aggregation degree range; Recalculating the independence threshold of the data shards according to the second aggregation degree range to obtain a second threshold range; Using the cluster analysis method in the machine learning algorithm to verify the independence of the data shards to determine whether they meet the preset conditions; If the independence threshold and aggregation degree range of the data shards meet the preset business scenario requirements, the final data sharding plan is determined to perform data sharding.
3. The method according to claim 1, characterized in that The differential privacy algorithm is used to add noise to the data shards, including: Obtain the preset data security level and accuracy requirements, and determine the initial independence threshold of data shards; For each data shard, a differential privacy algorithm is used to add an appropriate amount of noise to obtain the data shard after adding noise; If the independence threshold is lower than a preset threshold, the noise adding parameters are adjusted and the noise adding process is performed again until the independence threshold meets the preset condition.
4. The method according to claim 1, characterized in that: The data shards are encrypted using homomorphic encryption technology, including: Get the data shards to be processed; For each of the data shards, homomorphic encryption technology is used to perform encryption processing to obtain encrypted data shards; Using a cluster analysis method in machine learning to verify the independence of the encrypted data shards, and obtaining an independence threshold of the encrypted data shards; Determine whether the independence threshold is lower than a preset threshold. If so, adjust the encryption parameters and return to perform encryption processing until the independence threshold meets the preset condition.
5. The method according to claim 1, characterized in that According to the preset aggregation degree range, the federated learning framework is used to perform distributed aggregation calculation on the encrypted data shards to obtain preliminary aggregation results, including: Obtain the encrypted data shards, and use the federated learning framework to perform distributed aggregation calculations on the encrypted data according to the preset aggregation degree range to obtain preliminary aggregation results; Based on the preliminary aggregation result, a cluster analysis method is used to calculate the independence value of the encrypted data. If the independence value is lower than a preset threshold, the encryption parameters are adjusted, and the data shards are re-encrypted to obtain new encrypted data; The federated learning framework is used to perform distributed aggregation calculations on the encrypted data within the adjusted aggregation degree range to obtain the final preliminary aggregation results.
6. The method according to claim 5, characterized in that If the adjusted aggregation result meets the accuracy requirement, the aggregated data is decrypted to obtain the final result, including: Obtaining a final preliminary aggregation result, and calculating a precision value of the final preliminary aggregation result; Compare the accuracy value with a preset accuracy threshold to determine whether the accuracy value meets the preset accuracy threshold; If the precision value meets the preset precision threshold, the final preliminary aggregation result is decrypted using a preset decryption algorithm to obtain decrypted data; Performing format conversion and standardization processing on the decrypted data to obtain standardized data; According to the standardized data, a preset result generation algorithm is used to obtain a final result; The final result is stored in a preset database and marked as a completed processing state; A processing log is generated according to the storage result, in which key parameters and status information of the aggregation, decryption and generation links are recorded.
7. The method according to claim 1, characterized in that After obtaining the final result, the method further comprises: According to the real-time requirements, determine whether the aggregation process needs to be optimized, adopt lightweight encryption algorithms or reduce the number of shard interactions to improve computing efficiency; Through system architecture design, the independence of data shards and the adjustment mechanism of aggregation degree are embedded in the computing process to achieve dynamic balance; According to the design of the security mechanism, the risk of data leakage in the aggregation process is monitored in real time. If an abnormality is found, the calculation is terminated and security protection measures are initiated.
8. The method according to claim 7, characterized in that According to the real-time requirements, it is determined whether the aggregation process needs to be optimized, a lightweight encryption algorithm is used, or the number of shard interactions is reduced to improve computing efficiency, including: Obtain real-time demand parameters, calculate the processing time of the current aggregation process, and determine whether the processing time meets a preset time threshold; If the processing time does not meet the preset time threshold, the encryption algorithm used in the aggregation process is obtained, and according to the complexity of the encryption algorithm, a target lightweight encryption algorithm is selected from a preset set of lightweight encryption algorithms, and the encryption algorithm is replaced by the target lightweight encryption algorithm; Obtain the number of shards and the number of interactions of the aggregated data, and obtain an optimized shard interaction strategy by reducing the number of shards or optimizing the shard interaction logic; According to the target lightweight encryption algorithm and the optimized sharding interaction strategy, the aggregation process is simulated to obtain the optimized aggregation time; Determine whether the optimized aggregation time meets the preset time threshold; if not, adjust the parameters of the target lightweight encryption algorithm or the shard interaction strategy until the optimized aggregation time meets the preset time threshold; Solidify the aggregation process parameters that meet the preset time threshold to obtain an optimized aggregation processing module; The key performance indicators of the optimized aggregation processing module are obtained, and an aggregation efficiency log is generated.
9. The method according to claim 7, characterized in that: Through the system architecture design, the independence of data shards and the adjustment mechanism of aggregation degree are embedded in the computing process to achieve dynamic balance, including: Obtaining the current data sharding status information of the system, wherein the data sharding status information includes the number of data shards and the independence index of each shard; According to a preset aggregation degree threshold, determine whether the aggregation degree corresponding to the current data shard status information meets the standard; If the aggregation level does not meet the standard, a shard adjustment mechanism is started, wherein the shard adjustment mechanism adopts an embedded architecture design and is integrated into the computing flow of data processing; By means of the shard adjustment mechanism, the number of data shards and the independence index of each shard are optimized to obtain adjusted data shard status information; Adopt a dynamic balancing strategy to monitor the aggregation effect corresponding to the adjusted data sharding status information in real time and generate a dynamic balancing log; According to the dynamic balance log, a parameter optimization algorithm is used to optimize the adjustment parameters in the shard adjustment mechanism to obtain optimized adjustment parameters; If the degree of aggregation corresponding to the adjusted data shard status information reaches a preset degree of aggregation threshold, the adjusted data shard status information and the optimized adjustment parameters are solidified to form an optimized calculation process module.
10. An artificial intelligence-based privacy computing data security aggregation system, characterized in that: The system comprises: The data security level and accuracy acquisition module is used to obtain the preset data security level and accuracy requirements according to the business scenario requirements, and determine the independence threshold and aggregation range of the data shards; The differential privacy noise adding module is used to add noise to data shards using the differential privacy algorithm, thereby reducing the risk of leakage while ensuring data independence; Homomorphic encryption processing module, used to encrypt data shards through homomorphic encryption technology to ensure the security of data during transmission and calculation; The federated learning aggregation calculation module is used to perform distributed aggregation calculation on the encrypted data shards according to the preset aggregation degree range and adopt the federated learning framework to obtain preliminary aggregation results; The aggregation result precision adjustment module is used to adjust the aggregation degree, increase the number of interactions of data shards, and re-perform the aggregation calculation if the precision of the preliminary aggregation result is lower than the preset threshold; The aggregated data decryption module is used to decrypt the aggregated data to obtain the final result if the adjusted aggregated result meets the accuracy requirement; The real-time optimization module is used to determine whether the aggregation process needs to be optimized according to real-time requirements, adopt lightweight encryption algorithms or reduce the number of shard interactions to improve computing efficiency; Dynamic balance embedding module, which is used to embed the independence and aggregation adjustment mechanism of data shards into the computing process through system architecture design to achieve dynamic balance; The data leakage monitoring module is used to monitor the data leakage risk in the aggregation process in real time according to the security mechanism design. If any abnormality is found, the calculation is terminated and security protection measures are initiated.
Citation Information
Patent Citations
Machine learning method based on federated learning, electronic device and storage medium
CN112101579A
Federal learning security aggregation method based on QoS gain, medium and device
CN117973565A
Federal learning-oriented privacy protection method, system, device and medium
CN118590332A
Federal learning optimization method based on neural network model privacy protection
CN118643511A
Communication content security encryption method
CN119071074A
Cited By
Cluster communication method and system based on artificial intelligence
CN121334170A