Federal learning-based privacy computing service method and system
Through the coordinated regulation of hierarchical masking of sensitive fields and dynamic allocation of encryption computing power, the imbalance between privacy protection and computing efficiency in federated learning is solved, high security and low communication overhead of cross-institutional joint modeling are achieved, and model training efficiency and privacy protection capabilities are improved.
Patent Information
- Application Number
- CN202510759424.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Existing federated learning technologies find it difficult to achieve a balance between high security and low communication overhead while protecting data privacy in cross-institutional and multi-party joint modeling scenarios. In particular, in heterogeneous data distribution scenarios, model training efficiency is low and the amount of parameter transmission during communication is large, resulting in waste of computing resources and insufficient privacy protection.
Through the coordinated optimization of hierarchical masking of sensitive fields and dynamic allocation of encryption computing power, a closed-loop control system of privacy classification-parallel encryption-feedback correction is constructed. Local masking and dynamic noise superposition are used to generate differentiated desensitized data. The number of data block segmentation and noise injection intensity are dynamically adjusted through cross-validation of encryption time and privacy strength to achieve decoupled control of encryption efficiency and privacy strength.
It achieves a dynamic balance between high security and low communication overhead in federated learning, ensures fine-grained information hiding of data with multiple privacy levels and computing power matching of heterogeneous devices, and improves the training efficiency and privacy protection capabilities of the model.
Smart Images

Figure CN120671182A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of privacy computing services, and in particular to a privacy computing service method and system based on federated learning. Background Art
[0002] In cross-institutional, multi-party joint modeling scenarios, participants must implement joint modeling while protecting data privacy, while also meeting the requirements of high security and low communication overhead. Specifically, data from participating institutions must be strictly stored locally, and raw data transmission must be prohibited to avoid violating privacy regulations. At the same time, model training must maintain efficient convergence in heterogeneous data distribution scenarios, and parameter transmission must be significantly compressed during communication to reduce network bandwidth pressure and avoid wasting computing resources due to the high complexity of encryption algorithms.
[0003] The current mainstream solution is a federated learning secure aggregation protocol based on homomorphic encryption and model compression. This solution encrypts local model parameters through homomorphic encryption technology, ensuring that the gradients or parameters uploaded by participants are aggregated in ciphertext, thereby preventing the server or third party from obtaining the original data privacy through reverse deduction. Simultaneously, model parameter compression technology is combined to lightweight communication content, dynamically adjusting the compression rate to balance model accuracy and transmission efficiency. In addition, some solutions introduce differential privacy mechanisms, adding controllable noise to gradients during the local model training phase to defend against attacks based on parameter inference data characteristics. The noise intensity is adjusted through a privacy budget to balance model usability. Summary of the Invention
[0004] This application provides a privacy computing service method and system based on federated learning to solve the problem of imbalance between privacy protection and computing efficiency in federated learning in the existing technology.
[0005] In a first aspect, this application provides a privacy computing service method based on federated learning, including: Obtaining raw data from federated learning participants, identifying the types of sensitive fields in the raw data, classifying the sensitive fields into different privacy levels based on the types, and partially masking the contents of the sensitive fields based on the privacy levels and superimposing dynamic noise on the partially masked areas to generate desensitized data; The desensitized data is divided into multiple data blocks according to the preset encryption algorithm complexity, and the data blocks are allocated to independent operation units for parallel encryption processing. The number of the operation units is proportional to the complexity of the encryption algorithm to form an operation unit allocation strategy; Transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregate calculation result fed back by the federated learning collaboration node, and simultaneously record the encryption operation time and the privacy strength value corresponding to the aggregate calculation result; Calculating the difference between the encryption operation time and a preset threshold, adjusting the number of segments of the desensitized data block according to the difference to match the operation unit allocation strategy, and cross-validating the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value; Based on the associated parameters, when the privacy strength value is lower than the preset strength, the dynamic noise of the corresponding privacy level is increased and the number of segments of the desensitized data block is adjusted synchronously, so that the desensitized data block after dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism.
[0006] Optionally, obtaining raw data from federated learning participants, identifying the types of sensitive fields in the raw data, classifying the sensitive fields into different privacy levels according to the types, and partially masking the contents of the sensitive fields according to the privacy levels and superimposing dynamic noise on the partially masked areas to generate desensitized data, including: Obtain the original data provided by the federated learning participants, match the fields of the original data with the preset sensitive field type list one by one, the preset sensitive field type list includes the types of identity identification and location trajectory, and mark the successfully matched fields as sensitive field types; Obtaining a privacy risk weight table associated with the type of the sensitive field, and classifying the field containing the identity identifier into the first privacy level based on the privacy risk weight table. Calculating a density value for the field containing the location trajectory based on the number of trajectory points in the location trajectory, and classifying the field into the second privacy level when the density value exceeds a trajectory point threshold; retaining the first three characters of the field of the first privacy level and replacing the subsequent characters with masking symbols, and superimposing dynamic noise on the masking symbols one by one according to the positions of the first three characters; Performing a segmented masking operation on the trajectory field of the second privacy level, wherein the segmented masking operation divides continuous trajectory points into segments according to time intervals, masks the trajectory points at the end of each segment, obtains the fluctuation range of the values of the trajectory points, and superimposes dynamic noise matching the fluctuation range on the masked area; The sensitive fields that have been masked and superimposed with dynamic noise are reorganized in the format of the original data to generate desensitized data.
[0007] Optionally, the desensitized data is divided into multiple data blocks according to a preset encryption algorithm complexity, and the data blocks are assigned to independent computing units for parallel encryption processing. The number of computing units is proportional to the complexity of the encryption algorithm to form a computing unit allocation strategy, including: Determining the number of segments of the desensitized data according to a preset encryption algorithm complexity, wherein the encryption algorithm complexity is defined by the hierarchical depth of the encryption operation and the length of the data block; Divide the desensitized data into sequentially labeled data blocks according to the number of divisions corresponding to the current encryption algorithm complexity level, where the length of the data block is inversely proportional to the encryption algorithm complexity; Allocating the data block to independent computing units according to the number of segments, wherein the computing units are bound to one data block and trigger encryption processing, wherein the total number of computing units is strictly consistent with the number of segments, and the processing capacity of the computing units matches the length of the data block; Monitor the real-time rate at which the operation unit processes the data block. If the real-time rate is lower than the preset benchmark rate corresponding to the complexity of the current encryption algorithm, split the data block into sub-data blocks according to a preset splitting ratio, and allocate newly added operation units to synchronously process the sub-data blocks to form an operation unit allocation strategy.
[0008] Optionally, monitoring the real-time rate at which the computing unit processes the data block; if the real-time rate is lower than a preset reference rate corresponding to the complexity of the current encryption algorithm, dividing the data block into sub-data blocks according to a preset splitting ratio, and allocating newly added computing units to synchronously process the sub-data blocks, so as to form a computing unit allocation strategy, including: monitoring in real time the real-time rate at which the computing unit processes the data block, and comparing the real-time rate with a preset reference rate bound to the complexity of the current encryption algorithm on a unit-by-unit basis to generate a rate difference status mark; When the rate difference state is detected to be lower than the preset reference rate, a data block splitting operation is triggered, wherein the splitting operation divides the current data block into multiple sub-data blocks according to a preset splitting ratio, and the length of each sub-data block is inversely proportional to the splitting ratio; Assign sub-data blocks to be processed to the newly added computing units, and bind each sub-data block to an independent computing unit for encryption processing to form a computing unit allocation strategy.
[0009] Optionally, the encrypted data block is transmitted to the federated learning collaboration node through an encrypted transmission channel, and the aggregated calculation result fed back by the federated learning collaboration node is obtained, and the time consumption of the encryption operation and the privacy strength value corresponding to the aggregated calculation result are simultaneously recorded, including: Attach a sequence identifier and a channel key to the encrypted data blocks, and send the encrypted data blocks in batches to the federated learning collaboration nodes through the encrypted transmission channel according to the sequence identifier. When sending, record the start transmission timestamp and end transmission timestamp of each batch of data blocks to record the time consumed by the encryption operation; Verifying the continuity of the sequence identifier of the encrypted data block and the validity of the channel key at the receiving end of the federated learning collaboration node, and performing an aggregate calculation on the verified data block to generate an aggregate calculation result; A privacy strength value corresponding to the aggregate calculation result is calculated based on the encryption feature of the data block.
[0010] Optionally, calculating the difference between the encryption operation time and a preset threshold, adjusting the number of segments of the desensitized data block based on the difference to match the operation unit allocation strategy, and cross-validating the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value, including: Comparing the encryption operation time with a preset threshold point by point to calculate the percentage of the encryption operation time to the preset threshold as a difference value; Extract the deviation direction of the difference value, and adjust the number of segments of the desensitized data block according to the deviation direction and the percentage value. If the difference value is a positive deviation, the number of segments is increased proportionally to match the upper limit of the number of units of the operation unit allocation strategy. If the difference value is a negative deviation, the number of segments is reduced proportionally to release redundant operation units of the operation unit allocation strategy. Extracting the privacy strength value, and binding the difference value of each data block to the corresponding privacy strength value one by one to generate a verification pair of the difference value and the privacy strength value; Analyze the correlation between the difference value and the privacy strength value in the verification pair. When the difference value deviates positively and the privacy strength value is lower than a preset strength, increase the number of segmentations and improve the positive correlation parameter of the privacy strength value. When both the difference value and the privacy strength value meet the preset standards, generate a balanced correlation parameter to maintain the current number of segmentations and privacy strength. The forward correlation parameter is integrated with the equilibrium correlation parameter to generate the correlation parameter.
[0011] Optionally, based on the association parameter, when the privacy strength value is lower than a preset strength, the dynamic noise of the corresponding privacy level is increased and the number of segments of the desensitized data block is simultaneously adjusted, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism, including: Checking whether the privacy strength value marked in the associated parameter is lower than a preset strength. When it is detected that the privacy strength value is lower than the preset strength, dynamic noise is increased for the fields belonging to the corresponding privacy level in the desensitized data block to generate a noise-enhanced desensitized data block; Calculate the length change of the data block after the dynamic noise is increased, and adjust the number of segments of the desensitized data block according to the length change according to the preset segmentation rules to ensure that the length of the data block is inversely proportional to the number of segments; The desensitized data blocks after the adjusted number of splits are redistributed to independent computing units for encryption processing. Each computing unit is bound to the corresponding data block according to the updated number of splits and performs encryption processing to form a dynamic linkage mechanism.
[0012] Secondly, this application provides a privacy computing service system based on federated learning, including: An identification module, configured to obtain raw data from federated learning participants, identify the types of sensitive fields in the raw data, classify the sensitive fields into different privacy levels based on the types, and partially mask the contents of the sensitive fields based on the privacy levels and superimpose dynamic noise on the partially masked areas to generate desensitized data; An allocation module is used to divide the desensitized data into multiple data blocks according to a preset encryption algorithm complexity, and allocate the data blocks to independent computing units for parallel encryption processing, wherein the number of computing units is proportional to the complexity of the encryption algorithm to form a computing unit allocation strategy; A transmission module, configured to transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregated calculation results fed back by the federated learning collaboration node, and synchronously record the encryption operation time and the privacy strength value corresponding to the aggregated calculation result; a calculation module, configured to calculate a difference between the encryption operation time and a preset threshold, adjust the number of segments of the desensitized data block according to the difference to match the operation unit allocation strategy, and cross-validate the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value; A module is formed, which is used to increase the dynamic noise of the corresponding privacy level and synchronously adjust the number of segments of the desensitized data block based on the associated parameters when the privacy strength value is lower than the preset strength, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism.
[0013] In a third aspect, an embodiment of the present application provides a computing device comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a privacy computing service method based on federated learning as described in the first aspect above.
[0014] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program. When the computer program is executed by a computer, it implements a privacy computing service method based on federated learning as described in the first aspect.
[0015] In the technical solution of this application, an adaptive framework for dynamic linkage between privacy strength quantitative assessment and encryption strategy is constructed through a multi-layer protection mechanism of hierarchical desensitization of sensitive fields, elastic data encryption segmentation, dynamic resource scheduling, and encrypted transmission verification. Based on the difference value-driven data block segmentation adjustment and noise enhancement compensation technology, it breaks through the conflict between privacy protection and model effectiveness in the traditional static desensitization mode, and achieves precise prevention and control of privacy leakage risks while ensuring data availability. The cross-validation mechanism of encryption operation efficiency and privacy strength forms a dynamic balance of real-time feedback optimization, which solves the difficult problem of balancing security and computing resource utilization in multi-node collaboration of federated learning, and provides full-link closed-loop protection and elastic expansion capabilities for cross-domain data security sharing.
[0016] Furthermore, through the precise identification of preset sensitive fields and the dynamic division of privacy levels, a differentiated management and control system for core privacy elements such as identity identification and location trajectory has been established. A dual strategy of preserving leading characters and masking segments of trajectories is adopted to accurately hide sensitive information while ensuring the availability of some data. Combined with dynamic noise superposition technology, the noise intensity is matched according to the characteristics of the masked area and the range of trajectory fluctuations, effectively resolving the inherent contradiction between data distortion and privacy leakage in traditional desensitization methods. Through sensitive field format reorganization technology, the compatibility of desensitized data with the federated learning framework is ensured, providing a new protection paradigm that combines privacy strength and model effectiveness for multi-party data security collaboration.
[0017] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 A flowchart of a privacy computing service method based on federated learning provided by this application is shown; Figure 2 A scenario diagram showing a privacy computing service method based on federated learning provided by this application is shown; Figure 3 A scenario diagram showing a privacy computing service method based on federated learning provided by this application is shown; Figure 4 A schematic diagram of the structure of a privacy computing service system based on federated learning provided by this application is shown; Figure 5A schematic structural diagram of a computing device provided by the present application is shown. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0021] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0022] Research has found that while current secure aggregation frameworks for cross-institutional federated learning achieve preliminary privacy protection and communication optimization through homomorphic encryption and parameter compression techniques, their core bottlenecks lie in the inadequate adaptability of static privacy protection mechanisms to dynamic computational resource requirements, and a lack of quantitative basis for the coordinated regulation of privacy strength and communication efficiency. While homomorphic encryption can securely aggregate data in an encrypted state, its inherent computational complexity and communication overhead increase exponentially (especially in scenarios with high-dimensional model parameters), exacerbating the computing power imbalance between heterogeneous devices. Furthermore, existing solutions employ a globally unified differential privacy noise injection strategy, which struggles to adapt to the differentiated protection requirements of sensitive fields at multiple privacy levels. For example, applying the same noise intensity to diagnostic records (high privacy level) and demographic information (low privacy level) in medical data can lead to excessive perturbations, resulting in reduced model accuracy or insufficient privacy protection. Furthermore, while model compression techniques can reduce the single-transaction load, the fixed compression ratio setting cannot dynamically adapt to the changing correlation between encryption time and privacy strength for different data blocks, potentially resulting in loss of key parameters or reduced convergence stability.
[0023] To address these challenges, this paper proposes a method for dynamic encryption and noise control in federated learning for multiple privacy levels. Its innovation lies in the collaborative optimization of hierarchical masking of sensitive fields and dynamic allocation of encryption computing power, building a closed-loop control system of "privacy grading, parallel encryption, and feedback correction." Specifically, privacy levels are divided based on sensitive field type, and differentiated desensitized data is generated by combining local masking with dynamic noise superposition. Desensitized data is further segmented into parallel processing blocks based on the complexity of the encryption algorithm, and computing power resources are dynamically allocated based on the number of computing units, achieving decoupled control of encryption efficiency and privacy strength. By monitoring cross-validation metrics of encryption time and privacy strength in real time, the number of data block segmentations and the intensity of noise injection are dynamically adjusted, forming a multi-objective balance mechanism between computing power consumption, privacy protection, and communication load. This method overcomes the static privacy-efficiency trade-off inherent in traditional federated learning. On the one hand, local masking achieves fine-grained information hiding for high-privacy-level fields, mitigating the negative impact of global noise injection on model accuracy. On the other hand, a parallel encryption strategy based on data block segmentation significantly reduces the computational latency of homomorphic encryption by dynamically adjusting the segmentation granularity to match the computing power distribution of heterogeneous devices. More importantly, by modeling the correlation parameters between encryption time and privacy strength, and establishing a linkage feedback mechanism between noise intensity and number of segmentations, the system can adaptively optimize privacy protection strength and communication efficiency according to real-time performance indicators, achieving a technological leap from "fixed compression rate" to "dynamic resource adaptation" and from "global noise perturbation" to "multi-level masking enhancement", providing a solution for cross-institutional joint modeling that takes into account high security, low communication overhead and high model availability.
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0025] Figure 1 A flowchart of a privacy computing service method based on federated learning is provided for the embodiment of this application, such as Figure 1 As shown, the method includes: 101. Obtain the original data of the federated learning participants, identify the types of sensitive fields in the original data, classify the sensitive fields into different privacy levels according to the types, and partially mask the contents of the sensitive fields according to the privacy levels and superimpose dynamic noise on the partially masked areas to generate desensitized data.
[0026] Optionally, step 101 may specifically include the following steps: 1011. Obtain the original data provided by the federated learning participants, match the fields of the original data with a preset sensitive field type list one by one, the preset sensitive field type list including identity identifiers and location track types, and mark the successfully matched fields as sensitive field types; 1012. Obtain a privacy risk weight table associated with the type of the sensitive field, and classify the field containing the identity identifier into the first privacy level according to the privacy risk weight table. Calculate a density value for the field containing the location trajectory based on the number of trajectory points in the location trajectory, and classify it into the second privacy level when the density value exceeds a trajectory point threshold. 1013. Retain the first three characters of the field of the first privacy level and replace the subsequent characters with masking symbols, and superimpose dynamic noise on the masking symbols one by one according to the positions of the first three characters; 1014. Perform a segmented masking operation on the trajectory field of the second privacy level. The segmented masking operation divides the continuous trajectory points into segments according to time intervals, masks the trajectory points at the end of each segment, obtains the fluctuation range of the values of the trajectory points, and superimposes dynamic noise matching the fluctuation range on the masked area. 1015. The sensitive fields after masking and superimposing dynamic noise are reorganized according to the format of the original data to generate desensitized data.
[0027] In the above scheme, federated learning participants are data holders or individuals participating in collaborative machine learning. Raw data is a collection of unprocessed user information. Sensitive field types are categorized as information within the data that may impact personal privacy. The preset sensitive field type list is a predefined list of sensitive data types, such as identity identifiers and location trajectories. Identity identifiers are information elements that can directly or indirectly identify an individual. Location trajectories are time-series data that record an individual's movement paths. The privacy risk weight table is a quantitative indicator table based on the degree of privacy leakage risk associated with sensitive field types. The first privacy level is the high-risk privacy protection level for identity identifier fields. The second privacy level is the medium-risk privacy protection level for location trajectories, based on the density of trajectory points. The trajectory point threshold is the critical value for determining whether the density of location trajectories is too high. Masking symbols are placeholders used to replace the original characters in sensitive fields. Dynamic noise is randomly perturbed data superimposed on the masked area, whose value changes over time or context. Segmented masking is the process of segmenting continuous trajectory points by time interval and masking out portions of the data. Segments are continuous subsequences of trajectory data divided by time intervals. The fluctuation range is the range of change in the trajectory point value over time. Reorganization is the process of reorganizing the desensitized fields into their original data structure. Desensitized data is a privacy-preserving dataset that has been masked and noise-processed.
[0028] In this embodiment of the present application, first, in step 1011, the system obtains the raw data provided by the federated learning participants and compares each field of the raw data with a preset list of sensitive field types (including types such as identity identifiers and location trajectories) using a field matching algorithm (such as regular expression or keyword matching). Fields that successfully match are marked as sensitive field types. For example, fields named "ID number" or "GPS trajectory" are classified as identity identifiers or location trajectories, respectively.
[0029] Subsequently, in step 1012, a predefined privacy risk weight table (containing risk level mapping rules for different field types) is queried based on the type of the marked sensitive field. Identity fields are directly classified as the first privacy level. For location trajectory fields, a density calculation algorithm is used to count the number of trajectory points (e.g., the number of trajectory points per unit time or space) and compare it with a preset trajectory point threshold. If the density value exceeds the threshold, the data is classified as the second privacy level, ensuring stronger privacy protection for high-density trajectory data.
[0030] Next, in step 1013, a partial masking operation is performed on the fields of the first privacy level: the first three characters of the field content (such as the first three digits of the administrative region code of an ID number) are retained, and the subsequent characters are replaced with a fixed number of masking symbols (such as "*"). In the place of the masking symbols, dynamic noise is superimposed using a dynamic noise generation algorithm (such as random perturbation based on a hash function or a differential privacy noise mechanism), ensuring that the masked data is irreversible and retains some statistical characteristics. The noise amplitude is dynamically adjusted based on the field type and privacy level. For example, a low-amplitude random character replacement is used for the identity field.
[0031] Then, in step 1014, a segmented masking operation is performed on the location trajectory field of the second privacy level: continuous trajectory points are divided into multiple segments using a time window partitioning algorithm (such as fixed interval or dynamic clustering). The trajectory points at the end of each segment (such as the last few points in each time window) are masked, and the masking range is adaptively determined based on the distribution density of the trajectory points. At the same time, through statistical analysis of the numerical fluctuation range of the trajectory points (such as latitude and longitude offset or speed variation range), dynamic noise matching the fluctuation range (such as noise based on the Laplace mechanism or Gaussian noise) is superimposed on the masked area to ensure that the noise amplitude is consistent with the variation trend of the original data, thereby avoiding distortion of the data distribution characteristics.
[0032] Finally, in step 1015, the masked and noise-superimposed sensitive fields are reorganized according to the original data structure. The reorganization process retains the original content of non-sensitive fields and replaces only the processed sensitive fields. This generates desensitized data that complies with federated learning data specifications, ensuring a balance between data availability and privacy protection.
[0033] In a practical application, in a cross-hospital electronic medical record joint modeling scenario, the raw data uploaded by the federated learning participant (a tertiary hospital) includes patient IDs and department visit trajectories (step 1011). The system matches the pre-set sensitive field type list, identifying the patient ID as an identity identifier and the department visit time series as a location trajectory (step 1011). Based on the privacy risk weighting table, the patient ID is classified as privacy level 1, while the patient visit trajectories are classified as privacy level 2 because the frequency of cross-department visits per day exceeds the trajectory point threshold (step 1012). A partial masking is performed on the patient ID field: the "PAT" prefix is retained and subsequent characters are replaced with "". Dynamic noise is also superimposed on the masked area, for example, the original field "PAT2023-045" is desensitized to "PAT#k7%9" (step 1013). For each department visit trajectory, the last department name is masked on an hourly basis, and noise is added based on the historical trajectory fluctuation range. For example, the last segment of "Cardiology → Radiology → Laboratory" is masked and replaced with "Cardiology → Pharmacy" (step 1014). When reassembling the desensitized data, the patient ID and the noisy visit trajectory are encapsulated in the original JSON format (step 1015) for use in the federated learning model training task of liver disease prediction. This processing ensures that each hospital's local data, when participating in the joint modeling, protects patient privacy while maintaining the effectiveness of cross-institutional feature alignment, achieving a balance between privacy and model performance.
[0034] The overall solution in step 101 above implements a data desensitization mechanism with multi-level privacy protection and dynamic noise adaptation. By identifying sensitive field types and categorizing privacy levels, a differentiated management system based on core sensitive elements such as identity and location trajectory is established. A dual protection strategy of local masking and dynamic noise overlay is employed to achieve hierarchical control of privacy intensity while retaining some data availability. Refined processing such as leading character masking and segmented trajectory masking is implemented for fields of different privacy levels. Combined with the correlation and matching of dynamic noise range and trajectory fluctuation characteristics, this effectively addresses the difficult balance between data availability and privacy protection in traditional desensitization methods, laying the foundation for secure sharing of multi-source heterogeneous data in federated learning scenarios.
[0035] 102. Divide the desensitized data into multiple data blocks according to the preset encryption algorithm complexity, and allocate the data blocks to independent operation units for parallel encryption processing. The number of the operation units is proportional to the complexity of the encryption algorithm to form an operation unit allocation strategy.
[0036] Optionally, step 102 may specifically include the following steps: 1021. Determine the number of segments of the desensitized data according to a preset encryption algorithm complexity, where the encryption algorithm complexity is defined by the level depth of the encryption operation and the length of the data block; 1022. Divide the desensitized data into sequentially labeled data blocks according to the number of divisions corresponding to the current encryption algorithm complexity level, wherein the length of the data block is inversely proportional to the complexity of the encryption algorithm; 1023. Allocate the data block to independent computing units according to the number of segments. Each computing unit is bound to one data block and triggers encryption processing. The total number of computing units is strictly consistent with the number of segments, and the processing capacity of the computing units matches the length of the data block. 1024. Monitor the real-time rate at which the operation unit processes the data block. If the real-time rate is lower than the preset benchmark rate corresponding to the complexity of the current encryption algorithm, split the data block into sub-data blocks according to a preset splitting ratio, and allocate newly added operation units to synchronously process the sub-data blocks to form an operation unit allocation strategy.
[0037] In the above scheme, encryption algorithm complexity is a measure of the difficulty of the encryption process. The number of splits is the number of blocks into which the desensitized data is divided, determined by the encryption algorithm complexity. The hierarchical depth is the number of recursive or iterative processing steps in the encryption algorithm. The data block length is the number of bytes in a single data segment. Sequence-marked data blocks are encryption processing units numbered in the order of the original data. Independent computing units are computing resource entities that perform encryption tasks. Processing capacity is the number of bytes processed per second by a computing unit. Real-time rate is the actual speed at which a computing unit processes data blocks. The preset baseline rate is the minimum processing speed requirement corresponding to the encryption algorithm complexity. The split ratio is the ratio of the number of sub-blocks to the number of original blocks when a data block is split twice. Sub-data blocks are smaller data processing units resulting from the secondary split of a data block. The computing unit allocation strategy is a processing solution that dynamically adjusts the matching relationship between computing units and data blocks.
[0038] In this embodiment of the present application, first, at step 1021, the system determines the number of segments to be used for the desensitized data based on the preset encryption algorithm complexity (defined by the encryption algorithm's layer depth and data block length). The complexity assessment module uses a dynamic programming algorithm to analyze the computational resources required for encryption. For example, multi-layer encryption or long key algorithms correspond to higher complexity, and thus generates a number of segments proportional to the complexity (e.g., more data blocks are divided when the complexity is high).
[0039] Then, in step 1022, a data sharding algorithm is used to segment the masked data into a plurality of sequentially labeled data blocks, according to the number of segments corresponding to the current complexity level. The sharding rule is: the higher the complexity, the shorter the length of each data block. For example, fixed-length sharding or dynamic adaptive sharding is used to ensure that the computational load of each data block matches the processing power of the computing unit.
[0040] Next, in step 1023, independent computing units (such as containerized instances or threads) are created based on the number of partitions. Each unit is bound to a data block and triggers encryption processing. The resource scheduling module uses a load balancing algorithm to ensure that the total number of computing units is strictly consistent with the number of partitions and that the computing power of the units (such as the number of CPU cores and memory) is adapted to the data block length. For example, longer data blocks are assigned to high-spec computing units, while shorter data blocks are assigned to low-spec units, maximizing parallel efficiency.
[0041] Finally, in step 1024, the real-time rate at which each computing unit processes the data block is monitored. If the rate falls below a preset baseline rate corresponding to the current complexity (e.g., a bytes-per-second threshold), the dynamic adjustment module further splits the original data block into sub-blocks according to a preset split ratio (e.g., splitting it in half) and automatically adds new computing units to process the sub-blocks simultaneously. During this adjustment process, the task scheduler uses dynamic resource allocation algorithms (e.g., elastic scaling) to ensure that the new units are uniquely associated with the sub-blocks. Ultimately, a computing unit allocation strategy is dynamically optimized based on real-time performance, ensuring the efficient completion of encryption tasks.
[0042] In practical applications, in a joint training scenario for an interbank credit scoring model, the system performs encryption preprocessing on the desensitized customer occupation and income range data (output from step 101). Based on the encryption algorithm complexity of the AES-256 algorithm (step 1021), the desensitized data is segmented into eight sequentially labeled data blocks, with each block length adapted to the complexity level. The occupation field is segmented into fine-grained blocks due to its short character length, while the income range is segmented into coarse-grained blocks due to its wide numerical range (step 1022). Eight independent computing units are assigned to each data block (step 1023). The computing unit corresponding to the occupation field block activates a high-speed cache mechanism to accommodate short character processing requirements. During operation, monitoring revealed that the processing rate of a computing unit for the income range block fell below the baseline (step 1024), triggering a dynamic splitting strategy: the data block is split into four sub-data blocks based on the horizontal split ratio, and four additional computing units are activated for parallel processing. After this adjustment, the total number of computing units increases to 12, and the sub-data block lengths are adapted to the unit's processing capacity, restoring the overall encryption task rate to the preset threshold. This allocation strategy ensures that occupation and income data meet real-time requirements when jointly modeling across banks while maintaining encryption strength and privacy security boundaries.
[0043] The overall solution in step 102 above achieves flexible data encryption and dynamic resource scheduling optimization based on algorithm complexity. Through a dynamic mapping mechanism between encryption algorithm complexity and the number of data block splits, a collaborative matching model is established between data block length, number of computing units, and processing power. Real-time rate monitoring and secondary data block splitting techniques are used to dynamically regulate the efficiency of encryption task execution, breaking through the processing delay bottleneck caused by sudden changes in algorithm complexity under fixed resource allocation. The strong binding strategy between computing units and data blocks, combined with an elastic expansion mechanism, improves resource utilization for large-scale data parallel processing while ensuring encryption security, providing high-throughput support for federated learning encrypted communications.
[0044] Among them, step 1024 may specifically include the following processes: real-time monitoring of the real-time rate at which the operation unit processes the data block, and comparing the real-time rate with the preset reference rate bound to the complexity of the current encryption algorithm unit by unit to generate a rate difference status mark; when the rate difference status mark is detected to be lower than the preset reference rate, triggering the data block splitting operation, the splitting operation divides the current data block into multiple sub-data blocks according to a preset splitting ratio, and the length of each sub-data block is inversely proportional to the splitting ratio; allocating sub-data blocks to be processed to the newly added operation unit, and binding each sub-data block to an independent operation unit for encryption processing to form an operation unit allocation strategy.
[0045] 103. Transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregate calculation result fed back by the federated learning collaboration node, and synchronously record the time consumption of the encryption operation and the privacy strength value corresponding to the aggregate calculation result.
[0046] Optionally, step 103 may specifically include the following steps: 1031. Attach a sequence identifier and a channel key to the encrypted data block, and send the encrypted data blocks in batches to the federated learning collaboration node through the encrypted transmission channel according to the sequence identifier. When sending, record the start transmission timestamp and end transmission timestamp of each batch of data blocks to record the time consumed by the encryption operation. 1032. Verify the continuity of the sequence identifier of the encrypted data block and the validity of the channel key at the receiving end of the federated learning collaboration node, and perform aggregate calculation on the verified data blocks to generate an aggregate calculation result. 1033. Calculate a privacy strength value corresponding to the aggregate calculation result based on the encryption feature of the data block.
[0047] In the above scheme, the encrypted transmission channel is a secure communication link that uses an encryption protocol to protect the data transmission process. A federated learning collaboration node is a server or terminal device that performs multi-party data aggregation calculations. The aggregated calculation result is the joint output value of the data from multiple participants processed by the federated learning model. The encryption operation duration is the total length of time from the start of data block transmission to the completion of encryption processing. The privacy strength value is a quantitative indicator of the degree of privacy protection of the aggregated calculation result. The sequence identifier is a unique number that marks the order in which data blocks are transmitted. The channel key is a key parameter used to authenticate the identities of the communicating parties in the encrypted transmission channel. The start transmission timestamp is the system clock value recorded when the data block transmission begins. The end timestamp is the system clock value recorded when the data block transmission completes. The receiving end is the module in the federated learning collaboration node responsible for receiving and verifying the data block. The encryption feature is a set of mathematical properties formed by the combination of the data block encryption algorithm and the key parameters.
[0048] In this embodiment of the present application, first, through step 1031, the system attaches a unique sequence identifier (such as an increasing sequence number or hash value) and a channel key (such as a session key based on the TLS protocol) to the encrypted data block, ensuring that the order of the data blocks during transmission is traceable and the channel encryption is secure. The data blocks are sent to the federated learning collaboration node in batches according to the sequence identifier through an encrypted transmission channel (such as an SSL / TLS or IPSec tunnel). During transmission, the start and end transmission timestamps of each batch of data blocks are recorded through a time synchronization protocol to calculate the total time consumed by the encryption operation (such as the end timestamp minus the start timestamp).
[0049] Subsequently, in step 1032, the federated learning collaboration node at the receiving end verifies the continuity of the data block's sequence identifiers (e.g., checking whether the sequence number is increasing) using a sequence verification algorithm. It also confirms the validity of the channel key using a key verification algorithm (e.g., digital signature or hash check). Verified data blocks are input into the aggregate calculation module (e.g., the FedAvg algorithm), which performs weighted averaging or gradient fusion on multi-node data to generate an aggregate calculation result. If sequence or key verification fails, a data retransmission mechanism is triggered.
[0050] Finally, in step 1033, a privacy strength value is calculated using a privacy strength assessment algorithm based on the encryption characteristics of the data block (such as encryption algorithm type, key length, and noise intensity). For example, a differential privacy budget (ε value) or the security level of homomorphic encryption is used to quantify the degree of privacy protection, ensuring that the privacy leakage risk of the aggregation result is controllable. This value is used to subsequently optimize the encryption strategy or adjust the privacy protection strength of the federated learning model.
[0051] In a practical application, in a federated learning scenario for cross-e-commerce platform user behavior analysis, Platform A appends an ascending sequence identifier and a dynamically generated channel key to the encrypted user browsing duration and product click sequence data blocks (output from Step 102) and sends them sequentially to the federated collaboration nodes via a TLS encrypted transmission channel. After verifying the continuity of the data block sequence identifiers and the validity of the channel key (Step 1032), the collaboration node performs a weighted average aggregation calculation on the multi-platform data to generate a global user behavior feature vector. The system calculates the privacy strength value of the aggregation result based on the homomorphic properties of the encrypted data blocks (Step 1033), confirming that the updated model parameters do not leak individual browsing preferences. Simultaneously recorded transmission time indicates that the higher encryption level of the product click sequence data blocks results in increased transmission time, triggering subsequent optimization of the channel key rotation frequency to balance security and timeliness. This process ensures that when jointly optimizing recommendation models between e-commerce platforms, user sensitive behavior data meets privacy strength thresholds throughout the entire encrypted transmission and aggregation calculation process, while also supporting dynamic tuning of model iteration efficiency.
[0052] The overall solution in step 103 above implements a closed-loop verification system for end-to-end encrypted transmission and quantitative privacy strength assessment. Through a combined verification mechanism of sequence identifiers and channel keys, dual guarantees are established for the transmission integrity of encrypted data blocks and channel security. Combined with transmission timestamp recording and aggregated computation result feedback, full-link monitoring of encryption time consumption and privacy strength is achieved. A privacy strength value calculation model based on encryption feature analysis maps block-level encryption quality to a quantifiable privacy protection metric, resolving the difficulty of objectively assessing data security status in federated learning scenarios and providing a verifiable privacy protection benchmark for cross-node collaboration.
[0053] 104. Calculate the difference between the time consumption of the encryption operation and a preset threshold, and adjust the number of segments of the desensitized data block according to the difference to match the operation unit allocation strategy. At the same time, cross-validate the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value.
[0054] Optionally, step 104 may specifically include the following steps: 1041. Compare the encryption operation time with the preset threshold point by point to calculate the percentage of the encryption operation time to the preset threshold as a difference value; 1042. Extract the deviation direction of the difference value, and adjust the number of segments of the desensitized data block according to the deviation direction and the percentage value. If the difference value is a positive deviation, the number of segments is proportionally increased to match the upper limit of the number of units of the operation unit allocation strategy. If the difference value is a negative deviation, the number of segments is proportionally reduced to release redundant operation units of the operation unit allocation strategy. 1043. Extract the privacy strength value, and bind the difference value of each data block to the corresponding privacy strength value one by one to generate a verification pair of the difference value and the privacy strength value; 1044. Analyze the correlation between the difference value and the privacy strength value in the verification pair. When the difference value deviates positively and the privacy strength value is lower than a preset strength, increase the number of segments and increase the positive correlation parameter of the privacy strength value. When both the difference value and the privacy strength value meet the preset standards, generate a balanced correlation parameter to maintain the current number of segments and the privacy strength. 1045. Integrate the forward correlation parameter and the equilibrium correlation parameter to generate a correlation parameter.
[0055] In the above scheme, the preset threshold is a pre-set upper or lower reference value for encryption processing time. The difference value is the percentage difference between the encryption processing time and the preset threshold. The deviation direction is the positive or negative trend of the difference value relative to the preset threshold. The upper limit of the number of units in the computing unit allocation strategy is the maximum number of parallel processing units allowed by the computing unit allocation strategy. Redundant computing units are idle computing units that exceed the requirements of the current encryption task. The verification pair is a verification data set formed by binding the difference value and the privacy strength value. The correlation is the degree of statistical correlation between the difference value and the privacy strength value. The preset strength is the minimum protection standard threshold that the privacy strength value must meet. The positive correlation parameter is the positive impact coefficient for adjusting the privacy strength value when the difference value deviates positively. The balance correlation parameter is the stability parameter for maintaining the current strategy when both the difference value and the privacy strength value meet the standards. The correlation parameter is a dynamic adjustment coefficient that combines the positive correlation parameter and the balance correlation parameter.
[0056] In this embodiment of the present application, first, in step 1041, the system compares the encryption operation's duration against a preset baseline threshold (e.g., the maximum allowable processing time) at each point in time using the real-time monitoring module. The system then calculates the percentage by which the duration exceeds or falls below the threshold as a difference value. The comparison algorithm uses a sliding window technique, dynamically capturing the duration data for each time period and calculating the percentage difference against the threshold. For example, if the duration exceeds the threshold by 20%, the difference value can be +20%.
[0057] Subsequently, step 1042 extracts the deviation direction of the difference value (a positive deviation indicates that the time consumption exceeded the limit, and a negative deviation indicates that the time consumption was lower than expected). The dynamic adjustment module then adjusts the number of segments of the desensitized data block based on the deviation direction and percentage value. For example, if the difference value deviates positively by 30%, the number of segments is increased proportionally (e.g., the number of segments increases by 5% for every 10% difference) until the upper limit of the number of units in the computing unit allocation strategy is reached. If the difference value deviates negatively, the number of segments is reduced proportionally to release redundant computing units to optimize resource usage.
[0058] Next, in step 1043, the privacy strength value (e.g., the ε value for differential privacy or the encryption security level) for each data block is extracted. Using a data binding algorithm, the difference values are associated with the corresponding privacy strength values, generating a verification pair that reflects the relationship between the two. For example, if a data block has a difference value of +15% and a privacy strength value of medium, a verification pair (+15%, medium) is generated for subsequent analysis.
[0059] Then, in step 1044, a correlation analysis is performed on the difference value and the privacy strength value in the verification pair. If the difference value is positively deviated and the privacy strength value is lower than the preset security standard (for example, the privacy strength value does not reach a high level), the parameter adjustment algorithm is used to simultaneously increase the number of segments and improve the positive correlation parameter of the privacy strength value (for example, by increasing the privacy strength weight coefficient). If both the difference value and the privacy strength value meet the preset standards (for example, the time consumption is within the threshold and the privacy strength meets the standard), a balance correlation parameter is generated to maintain the current stable state of the number of segments and privacy strength.
[0060] Finally, in step 1045, the forward correlation parameter (used to optimize performance and privacy) and the balance correlation parameter (used to maintain a stable state) are integrated into a unified correlation parameter through a parameter fusion algorithm. This parameter is fed back to the system's dynamic adjustment module to guide the subsequent coordinated optimization of data segmentation, encryption operations, and privacy protection, forming a closed-loop adaptive control mechanism.
[0061] In a practical application, during a joint risk assessment modeling scenario involving insurance institutions, the system detected that encryption computation time exceeded a preset threshold (step 1041) and that the difference value exhibited a positive deviation. Based on the deviation direction (step 1042), the number of segments of the customer income masked data block was increased from 12 to 18, freeing up more computing units for parallel processing to meet the policy upper limit. Simultaneously verifying the correlation between the difference value and the privacy strength value (step 1043), it was found that the privacy strength value of the newly segmented block was slightly lower than the preset standard due to a decrease in the noise overlay density. After system analysis and verification (step 1044), a positive correlation parameter adjustment was triggered: while maintaining the number of segments, the dynamic noise intensity of the newly segmented data block was increased to enhance privacy protection, while also optimizing the task scheduling logic of the computing units. The resulting correlation parameters (step 1045) balanced encryption time and privacy strength, ensuring that user income data, when training a cross-institutional premium prediction model, meets the real-time encrypted transmission requirements while also achieving the privacy security baseline for joint modeling through a dynamic noise compensation mechanism, thereby improving both risk assessment accuracy and data protection effectiveness.
[0062] The overall solution in step 104 above achieves a dynamic balance between encryption efficiency and privacy protection strength. Through a cross-validation mechanism between encryption time differences and privacy strength values, a bidirectional feedback loop is established for adjusting the number of data splits and optimizing privacy parameters. The identification of deviations in the difference values and the dynamic correction of the split ratios enable elastic scaling of the computing unit allocation strategy, ensuring the timeliness of encryption tasks while avoiding redundant resource consumption. The generation of parameters linking privacy strength values to processing efficiency integrates encryption performance and privacy protection into a unified evaluation framework, providing a dynamic decision-making basis for the coordinated optimization of security and efficiency in federated learning scenarios.
[0063] 105. Based on the associated parameters, when the privacy strength value is lower than the preset strength, the dynamic noise corresponding to the privacy level is increased and the number of segments of the desensitized data block is adjusted synchronously, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism. Optionally, step 105 may specifically include the following steps: 1051. Check whether the privacy strength value marked in the associated parameter is lower than a preset strength. When it is detected that the privacy strength value is lower than the preset strength, enhance dynamic noise for the fields of the corresponding privacy level in the desensitized data block to generate a noise-enhanced desensitized data block. 1052. Calculate the length change of the data block after the dynamic noise is increased, and adjust the number of segments of the desensitized data block according to the length change and the preset segmentation rule to ensure that the length of the data block is inversely proportional to the number of segments; 1053. The desensitized data blocks after the adjusted number of segments are reallocated to independent operation units for encryption processing. Each operation unit is bound to the corresponding data block according to the updated number of segments and performs encryption processing to form a dynamic linkage mechanism.
[0064] In the above solution, the dynamic linkage mechanism is the automated control logic that coordinates the privacy strength value with the number of data segments. The preset segmentation rule is a block adjustment criterion defined based on the relationship between data block length and the number of segments. The desensitized data block after noise enhancement is the privacy-preserving dataset generated by reprocessing after increasing the dynamic noise strength. The length change is the difference in the number of bytes in the data block after adjusting for dynamic noise.
[0065] In this embodiment of the present application, first, at step 1051, the system checks whether the privacy strength value marked in the associated parameter is lower than a preset security strength threshold. If the privacy strength value is detected to be insufficient (e.g., due to insufficient data masking or too low noise intensity), a dynamic noise enhancement algorithm is used to increase the superposition intensity of dynamic noise (e.g., by increasing the density or range of random noise) for fields with low privacy levels (e.g., identity identifiers or location trajectories) in the desensitized data block, generating a noise-enhanced desensitized data block to ensure that the privacy protection level meets the preset standard.
[0066] Subsequently, in step 1052, the change in the length of the data block after noise enhancement (e.g., the increase in the number of bytes due to noise superposition) is calculated. The data block length monitoring module collects the change statistics and dynamically adjusts the number of segments based on a preset segmentation rule (e.g., the inverse relationship between data block length and the number of segments). For example, if the data block length increases due to increased noise, the number of segments is proportionally increased to shorten the length of each data block, ensuring a balance between encryption processing efficiency and resource utilization.
[0067] Finally, in step 1053, the adjusted desensitized data blocks are redistributed to independent computing units according to the new number of partitions. The resource scheduling module creates or releases computing units through a dynamic expansion or contraction mechanism, ensuring that each unit is bound to a data block and performs encryption processing. During the encryption process, the processing power of the computing units is matched in real time to the length of the data blocks (for example, long data blocks are allocated to high-performance units), forming a dynamic linkage mechanism that coordinates privacy strength and the number of partitions, achieving continuous optimization of privacy protection and computational efficiency.
[0068] In a practical application, in a federated learning scenario for urban traffic flow prediction, the system detects that the privacy strength of a masked vehicle type data block falls below a preset level (step 1051). This triggers a dynamic noise enhancement mechanism: a random noise mask is applied to the vehicle brand field, which belongs to the second privacy level, masking "new energy - sedan" to "new energy *# - sedan@car." The noise-enhanced data block increases in length due to character expansion. Based on the preset segmentation rule (step 1052), the system adjusts the number of segments from 12 to 18, ensuring that the data block length is inversely proportional to the number of segments. After this adjustment, six newly added independent computing units (step 1053) bind fine-grained data blocks and prioritize high-frequency fluctuations in the noise-enhanced field during AES encryption. This dynamic linkage mechanism aligns encryption time and privacy strength in real time. When jointly modeling vehicle type data across cities, noise superposition mitigates the risk of brand information leakage while optimizing the number of segments to maintain encryption efficiency. This allows the traffic flow prediction model to continuously incorporate the dynamic characteristics of multi-source road network data while protecting driver privacy, improving the accuracy of traffic light control during peak hours.
[0069] The overall solution in step 105 above implements a closed-loop enhancement mechanism that integrates privacy strength self-repair with encryption strategies. Through privacy strength threshold monitoring and dynamic noise enhancement technology, an adaptive compensation system is established for scenarios where privacy protection is insufficient. A dynamic segmentation rule adjustment strategy based on changes in data block length ensures resource adaptability during encryption processing of desensitized data after noise enhancement. A real-time reallocation mechanism for computing units and updated data blocks creates a dynamic balance between improving privacy protection strength and maintaining encryption processing efficiency. This effectively addresses the dynamic changes in privacy leakage risks during federated learning collaboration and achieves continuous coordinated optimization of security protection capabilities and system stability.
[0070] The following is a complete embodiment of steps 101 to 105: In a cross-institutional student behavior analysis scenario, the raw data uploaded by federated learning participants (multiple universities) includes student IDs and classroom activity trajectories. Using a pre-set list of sensitive field types, the system identifies the student ID as an identity field and the classroom activity trajectories as location fields. Based on a privacy risk weighting table, the student ID is classified as privacy level 1, and high-frequency classroom trajectories are classified as privacy level 2. The first three characters of the student ID field are retained and overlaid with dynamic noise to generate a desensitized identifier. The classroom trajectories are segmented by course section, masking the end points. Dynamic noise matching the activity frequency is injected into the masked areas. Desensitized data is segmented into multiple data blocks based on the complexity of the encryption algorithm. The number of computation units is proportional to the complexity level to ensure that the data block length is adapted to the unit's processing capacity. If the speed of a computation unit is found to be below the baseline during real-time monitoring, the data block is split and additional computation units are added for parallel processing. The encrypted data blocks are appended with a sequence identifier and channel key and sent to the federated collaboration nodes via an encrypted transmission channel. The nodes verify the identifier continuity and key validity, aggregate them, and generate a global learning behavior signature. The privacy strength value is calculated and the encryption time is recorded simultaneously. When the encryption time difference shows a positive deviation and the privacy strength is insufficient, the system increases the noise intensity of the student ID field, dynamically adjusts the number of data block segments, and expands the computing unit capacity to balance processing efficiency. The noise-enhanced data blocks are redistributed to the computing unit for encryption, forming a dynamic linkage mechanism. This solution enables educational institutions to jointly optimize teaching models through privacy level classification, dynamic noise superposition, and flexible scheduling of computing resources. This ensures that student identity and behavior data meet privacy protection requirements when shared across institutions, while also supporting model iteration for teaching effectiveness analysis.
[0071] like Figure 2 、 Figure 3 As shown in the figure, the dynamic coordination mechanism of privacy protection and computational efficiency in federated learning is explained. Figure 2Focusing on single-scenario privacy control, the diagram presents a three-tier architecture: data provider, parallel computing cluster, and federated learning node. The data source on the left, represented by a building icon, represents multi-domain data input. The parallel cluster in the center utilizes distributed nodes to perform encrypted computations and gradient aggregation. The federated nodes on the right receive the aggregated results and update the global model. A dynamic monitoring system at the bottom tracks privacy strength (such as the differential privacy noise level) and computational time in real time, dynamically adjusting parameters through a feedback loop to ensure a balance between privacy protection and training efficiency. Figure 3 Expanding to a multi-institutional collaboration scenario, the medical and financial data sources on the left undergo pre-processing by anonymizing sensitive fields and segmenting data. The parallel computing layer in the middle integrates encryption and data desensitization technologies, and the federated nodes on the right complete secure multi-party aggregation. In this scenario, the dynamic monitoring system focuses on regulating noise intensity parameters, quantitatively analyzing the relationship between privacy leakage risk and model accuracy to achieve adaptive protection in cross-institutional data collaboration. Together, these two diagrams form a closed-loop system of "data preprocessing-secure computation-dynamic optimization." Through layered encryption, gradient perturbation, and real-time monitoring, this system enhances the practicality and cross-domain collaboration capabilities of the federated learning system while protecting data privacy.
[0072] Figure 4 A structural diagram of a privacy computing service system based on federated learning is provided for the embodiment of this application, such as Figure 2 As shown, the system includes: Identification module 41, configured to obtain raw data from federated learning participants, identify the types of sensitive fields in the raw data, classify the sensitive fields into different privacy levels based on the types, and partially mask the contents of the sensitive fields based on the privacy levels and superimpose dynamic noise on the partially masked areas to generate desensitized data; an allocation module 42 for dividing the desensitized data into a plurality of data blocks according to a preset encryption algorithm complexity, and allocating the data blocks to independent computing units for parallel encryption processing, wherein the number of computing units is proportional to the encryption algorithm complexity, so as to form a computing unit allocation strategy; Transmission module 43, configured to transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregated calculation result fed back by the federated learning collaboration node, and simultaneously record the encryption operation time and the privacy strength value corresponding to the aggregated calculation result; a calculation module 44 configured to calculate a difference between the encryption operation time and a preset threshold, adjust the number of segments of the desensitized data block based on the difference to match the operation unit allocation strategy, and cross-validate the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value; A module 45 is formed, which is used to increase the dynamic noise of the corresponding privacy level and synchronously adjust the number of segments of the desensitized data block based on the associated parameters when the privacy strength value is lower than the preset strength, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism.
[0073] Figure 4 The privacy computing service system based on federated learning can execute Figure 1 The implementation principles and technical effects of the federated learning-based privacy computing service method described in the illustrated embodiment will not be elaborated on here. The specific manner in which each module and unit performs operations in the federated learning-based privacy computing service system described in the above embodiment has been described in detail in the relevant embodiments of the method and will not be elaborated on here.
[0074] In one possible design, Figure 4 A privacy computing service system based on federated learning in the embodiment shown can be implemented as a computing device, such as Figure 5 As shown, the computing device may include a storage component 51 and a processing component 52; The storage component 51 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 52 .
[0075] The processing component 52 is used for the above Figure 1 The embodiment provides a privacy computing service method based on federated learning.
[0076] The processing component 52 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0077] The storage component 51 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0078] Of course, a computing device may also include other components, such as input / output interfaces, display components, communication components, etc.
[0079] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0080] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0081] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0082] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 A privacy computing service method based on federated learning in the illustrated embodiment.
[0083] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0084] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0085] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A privacy computing service method based on federated learning, characterized in that: include: Obtaining raw data from federated learning participants, identifying the types of sensitive fields in the raw data, classifying the sensitive fields into different privacy levels based on the types, and partially masking the contents of the sensitive fields based on the privacy levels and superimposing dynamic noise on the partially masked areas to generate desensitized data; The desensitized data is divided into multiple data blocks according to the preset encryption algorithm complexity, and the data blocks are allocated to independent operation units for parallel encryption processing. The number of the operation units is proportional to the complexity of the encryption algorithm to form an operation unit allocation strategy; Transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregate calculation result fed back by the federated learning collaboration node, and simultaneously record the encryption operation time and the privacy strength value corresponding to the aggregate calculation result; Calculating the difference between the encryption operation time and a preset threshold, adjusting the number of segments of the desensitized data block according to the difference to match the operation unit allocation strategy, and cross-validating the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value; Based on the associated parameters, when the privacy strength value is lower than the preset strength, the dynamic noise of the corresponding privacy level is increased and the number of segments of the desensitized data block is adjusted synchronously, so that the desensitized data block after dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism.
2. The method according to claim 1, characterized in that Obtaining raw data from federated learning participants, identifying the types of sensitive fields in the raw data, classifying the sensitive fields into different privacy levels based on the types, and partially masking the contents of the sensitive fields based on the privacy levels and superimposing dynamic noise on the partially masked areas to generate desensitized data, including: Obtain the original data provided by the federated learning participants, match the fields of the original data with the preset sensitive field type list one by one, the preset sensitive field type list includes the types of identity identification and location trajectory, and mark the successfully matched fields as sensitive field types; Obtaining a privacy risk weight table associated with the type of the sensitive field, and classifying the field containing the identity identifier into the first privacy level based on the privacy risk weight table. Calculating a density value for the field containing the location trajectory based on the number of trajectory points in the location trajectory, and classifying the field into the second privacy level when the density value exceeds a trajectory point threshold; retaining the first three characters of the field of the first privacy level and replacing the subsequent characters with masking symbols, and superimposing dynamic noise on the masking symbols one by one according to the positions of the first three characters; Performing a segmented masking operation on the trajectory field of the second privacy level, wherein the segmented masking operation divides continuous trajectory points into segments according to time intervals, masks the trajectory points at the end of each segment, obtains the fluctuation range of the values of the trajectory points, and superimposes dynamic noise matching the fluctuation range on the masked area; The sensitive fields that have been masked and superimposed with dynamic noise are reorganized in the format of the original data to generate desensitized data.
3. The method according to claim 1, characterized in that The desensitized data is divided into multiple data blocks according to the preset encryption algorithm complexity, and the data blocks are allocated to independent operation units for parallel encryption processing. The number of operation units is proportional to the complexity of the encryption algorithm to form an operation unit allocation strategy, including: Determining the number of segments of the desensitized data according to a preset encryption algorithm complexity, wherein the encryption algorithm complexity is defined by the hierarchical depth of the encryption operation and the length of the data block; Divide the desensitized data into sequentially labeled data blocks according to the number of divisions corresponding to the current encryption algorithm complexity level, where the length of the data block is inversely proportional to the encryption algorithm complexity; Allocating the data block to independent computing units according to the number of segments, wherein the computing units are bound to one data block and trigger encryption processing, wherein the total number of computing units is strictly consistent with the number of segments, and the processing capacity of the computing units matches the length of the data block; Monitor the real-time rate at which the operation unit processes the data block. If the real-time rate is lower than the preset benchmark rate corresponding to the complexity of the current encryption algorithm, split the data block into sub-data blocks according to a preset splitting ratio, and allocate newly added operation units to synchronously process the sub-data blocks to form an operation unit allocation strategy.
4. The method according to claim 3, wherein Monitoring the real-time rate at which the computing unit processes the data block; if the real-time rate is lower than a preset reference rate corresponding to the complexity of the current encryption algorithm, dividing the data block into sub-data blocks according to a preset splitting ratio, and allocating newly added computing units to synchronously process the sub-data blocks, thereby forming a computing unit allocation strategy, including: monitoring in real time the real-time rate at which the computing unit processes the data block, and comparing the real-time rate with a preset reference rate bound to the complexity of the current encryption algorithm on a unit-by-unit basis to generate a rate difference status mark; When the rate difference state is detected to be lower than the preset reference rate, a data block splitting operation is triggered, wherein the splitting operation divides the current data block into multiple sub-data blocks according to a preset splitting ratio, and the length of each sub-data block is inversely proportional to the splitting ratio; Assign sub-data blocks to be processed to the newly added computing units, and bind each sub-data block to an independent computing unit for encryption processing to form a computing unit allocation strategy.
5. The method according to claim 1, wherein Transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregated calculation results fed back by the federated learning collaboration node, and simultaneously record the encryption operation time and the privacy strength value corresponding to the aggregated calculation result, including: Attach a sequence identifier and a channel key to the encrypted data blocks, and send the encrypted data blocks in batches to the federated learning collaboration nodes through the encrypted transmission channel according to the sequence identifier. When sending, record the start transmission timestamp and end transmission timestamp of each batch of data blocks to record the time consumed by the encryption operation; Verifying the continuity of the sequence identifier of the encrypted data block and the validity of the channel key at the receiving end of the federated learning collaboration node, and performing an aggregate calculation on the verified data block to generate an aggregate calculation result; A privacy strength value corresponding to the aggregate calculation result is calculated based on the encryption feature of the data block.
6. The method according to claim 1, wherein Calculating the difference between the encryption operation time and a preset threshold, adjusting the number of segments of the desensitized data block based on the difference to match the operation unit allocation strategy, and cross-validating the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value, including: Comparing the encryption operation time with a preset threshold point by point to calculate the percentage of the encryption operation time to the preset threshold as a difference value; Extract the deviation direction of the difference value, and adjust the number of segments of the desensitized data block according to the deviation direction and the percentage value. If the difference value is a positive deviation, the number of segments is increased proportionally to match the upper limit of the number of units of the operation unit allocation strategy. If the difference value is a negative deviation, the number of segments is reduced proportionally to release redundant operation units of the operation unit allocation strategy. Extracting the privacy strength value, and binding the difference value of each data block to the corresponding privacy strength value one by one to generate a verification pair of the difference value and the privacy strength value; Analyze the correlation between the difference value and the privacy strength value in the verification pair. When the difference value deviates positively and the privacy strength value is lower than a preset strength, increase the number of segmentations and improve the positive correlation parameter of the privacy strength value. When both the difference value and the privacy strength value meet the preset standards, generate a balanced correlation parameter to maintain the current number of segmentations and privacy strength. The forward correlation parameter is integrated with the equilibrium correlation parameter to generate the correlation parameter.
7. The method according to claim 1, wherein Based on the associated parameters, when the privacy strength value is lower than the preset strength, the dynamic noise of the corresponding privacy level is increased and the number of segments of the desensitized data block is adjusted synchronously, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism, including: Checking whether the privacy strength value marked in the associated parameter is lower than a preset strength. When it is detected that the privacy strength value is lower than the preset strength, dynamic noise is increased for the fields belonging to the corresponding privacy level in the desensitized data block to generate a noise-enhanced desensitized data block; Calculate the length change of the data block after the dynamic noise is increased, and adjust the number of segments of the desensitized data block according to the length change according to the preset segmentation rules to ensure that the length of the data block is inversely proportional to the number of segments; The desensitized data blocks after the adjusted number of splits are redistributed to independent computing units for encryption processing. Each computing unit is bound to the corresponding data block according to the updated number of splits and performs encryption processing to form a dynamic linkage mechanism.
8. A privacy computing service system based on federated learning, characterized in that: include: An identification module, configured to obtain raw data from federated learning participants, identify the types of sensitive fields in the raw data, classify the sensitive fields into different privacy levels based on the types, and partially mask the contents of the sensitive fields based on the privacy levels and superimpose dynamic noise on the partially masked areas to generate desensitized data; An allocation module is used to divide the desensitized data into multiple data blocks according to a preset encryption algorithm complexity, and allocate the data blocks to independent computing units for parallel encryption processing, wherein the number of computing units is proportional to the complexity of the encryption algorithm to form a computing unit allocation strategy; A transmission module, configured to transmit the encrypted data block to the federated learning collaboration node through an encrypted transmission channel, obtain the aggregated calculation results fed back by the federated learning collaboration node, and synchronously record the encryption operation time and the privacy strength value corresponding to the aggregated calculation result; a calculation module, configured to calculate a difference between the encryption operation time and a preset threshold, adjust the number of segments of the desensitized data block according to the difference to match the operation unit allocation strategy, and cross-validate the difference with the privacy strength value to generate an association parameter between the operation unit allocation strategy and the privacy strength value; A module is formed, which is used to increase the dynamic noise of the corresponding privacy level and synchronously adjust the number of segments of the desensitized data block based on the associated parameters when the privacy strength value is lower than the preset strength, so that the desensitized data block after the dynamic noise enhancement is re-segmented and encrypted to form a dynamic linkage mechanism.
9. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a privacy computing service method based on federated learning as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, a privacy computing service method based on federated learning is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Real-time stream data processing method, device, computer apparatus, and storage medium
CA3148075A1
MePC-F model-based real-time federal learning data privacy security strengthening method in Internet of Vehicles
CN115310121A
Federal learning-based privacy protection type large-scale model training and deployment method
CN118734360A
Self-adaptive task scheduling execution unit management method and system
CN119376903A
Multi-private-domain visitor portrait sharing and privacy protection routing method based on federal learning
CN119383014A
Cited By
Intelligent federal learning cross-component privacy query method, related device and storage medium
CN120995504A
Longitudinal federal data privacy calculation method and system based on artificial intelligence
CN121261970A