Model permission processing method, device and medium applying federated learning
Patent Information
- Application Number
- CN202511807541.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-12-03
AI Technical Summary
目前,模型权限处理通常采用基于参与方身份的权限配置方式,这种做法由于未能考虑农业生产中作物生长周期不同阶段对模型参数敏感程度的影响,导致权限配置单一,无法根据参数的敏感程度实现精细化的安全防护,使得高敏感参数的安全性不足或低敏感参数的可用性降低,影响联邦学习系统的安全性和合规性
[0006] This invention acquires crop growth stage characteristic data of target farmland within a preset growth cycle. The feature data, collected through a distributed sensor network and aligned with timestamps, provides a precise and dynamic basis for subsequent permission allocation. Based on this crop growth stage characteristic data, model permission levels are divided, and the generated stage-based permission configuration set enables dynamic adjustment of permissions at different growth stages. This ensures that permission management is closely integrated with the actual needs of crop growth. The stage-based permission configuration set is associated and mapped with the local model parameters of each participant in the federated learning system. The resulting local model parameter set with permission identifiers achieves precise binding between permissions and specific model parameters, avoiding coarse-grained permission management. Differential encryption is performed on the local model parameter set with permission identifiers. By adding a noise perturbation factor dynamically adjusted based on permission level, a permission-encrypted model parameter set is generated, ensuring that the encryption strength matches the permission level of the parameters. This guarantees the security of high-permission parameters while also considering the usability of low-permission parameters. Based on the permission-encrypted model parameter set, model access permission verification is performed at the federated learning model call interface. The generated cross-institutional model call permission verification results effectively control the access scope of model parameters, ensuring the compliance and security of model parameter calls during the federated learning process.
Smart Images

Figure CN121682863B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a model permission processing method, device and medium using federated learning. Background Technology
[0002] With the rapid development of intelligent agriculture, federated learning has gained increasing attention in the agricultural field. Federated learning offers the advantage of enabling multi-agency collaborative model training while protecting data privacy. Model access control is a key technical aspect of federated learning, primarily used to control the access scope and operational permissions of different participants to model parameters, ensuring security and compliance during model training and invocation. Currently, model access control typically employs a participant-based access configuration approach. This approach fails to consider the impact of different stages of crop growth cycles on the sensitivity of model parameters, resulting in a simplistic access configuration that cannot achieve refined security protection based on parameter sensitivity. This leads to insufficient security for highly sensitive parameters or reduced availability for less sensitive parameters, impacting the security and compliance of the federated learning system. Summary of the Invention
[0003] In view of this, the present invention provides a model permission processing method, device, and medium applying federated learning. The technical solution of the embodiments of the present invention is implemented as follows: On one hand, this invention provides a model permission processing method using federated learning. The method includes: acquiring crop growth stage feature data of a target farmland within a preset growth cycle; the crop growth stage feature data is collected via a distributed sensor network at a preset sampling frequency and timestamped; performing model permission level classification processing based on the crop growth stage feature data to generate a stage-based permission configuration set, the stage-based permission configuration set containing model parameter access permission identifiers and permission effective time intervals corresponding to different growth stages; and associating and mapping the stage-based permission configuration set with the local model parameters of each participant in the federated learning system to generate a local permission configuration set with permission identifiers. A local model parameter set, wherein the local model parameter set with permission identifiers includes local model parameters and their corresponding permission level labels and access control conditions; differential encryption processing is performed on the local model parameter set with permission identifiers, and a permission-encrypted model parameter set is generated by adding a noise perturbation factor that is dynamically adjusted based on the permission level. The noise perturbation intensity of the permission-encrypted model parameter set is positively correlated with the permission level label; based on the permission-encrypted model parameter set, a model access permission verification operation is performed at the federated learning model call interface to generate a cross-institution model call permission verification result. The permission verification result is used to indicate whether the model parameter call request meets the preset permission configuration requirements.
[0004] On the other hand, the present invention provides a computer device including a memory and a processor, the memory storing a computer program executable on the processor, the processor executing the program to implement the steps in the method described above.
[0005] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0006] This invention acquires crop growth stage characteristic data of target farmland within a preset growth cycle. The feature data, collected through a distributed sensor network and aligned with timestamps, provides a precise and dynamic basis for subsequent permission allocation. Based on this crop growth stage characteristic data, model permission levels are divided, and the generated stage-based permission configuration set enables dynamic adjustment of permissions at different growth stages. This ensures that permission management is closely integrated with the actual needs of crop growth. The stage-based permission configuration set is associated and mapped with the local model parameters of each participant in the federated learning system. The resulting local model parameter set with permission identifiers achieves precise binding between permissions and specific model parameters, avoiding coarse-grained permission management. Differential encryption is performed on the local model parameter set with permission identifiers. By adding a noise perturbation factor dynamically adjusted based on permission level, a permission-encrypted model parameter set is generated, ensuring that the encryption strength matches the permission level of the parameters. This guarantees the security of high-permission parameters while also considering the usability of low-permission parameters. Based on the permission-encrypted model parameter set, model access permission verification is performed at the federated learning model call interface. The generated cross-institutional model call permission verification results effectively control the access scope of model parameters, ensuring the compliance and security of model parameter calls during the federated learning process. Attached Figure Description
[0007] Figure 1 This is a schematic diagram illustrating the implementation process of a model permission processing method using federated learning, provided in an embodiment of the present invention.
[0008] Figure 2 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0009] This invention provides a model permission processing method using federated learning, which can be executed by the processor of a computer device. The computer device can refer to a server, laptop, tablet, desktop computer, or other device with data processing capabilities.
[0010] Figure 1 This is a schematic diagram illustrating the implementation process of a model permission processing method using federated learning, as provided in an embodiment of the present invention. Figure 1As shown, the method includes: Step S100: Obtain crop growth stage characteristic data of the target farmland within a preset growth cycle. The crop growth stage characteristic data is collected by a distributed sensor network at a preset sampling frequency and aligned with timestamps.
[0011] A preset growth cycle is a complete time period from the start to the end of crop growth, predetermined based on the crop's biological characteristics and growth patterns. For example, for wheat, the preset growth cycle might be the entire process from sowing to harvest. Crop growth stage characteristic data is a collection of information describing various characteristics exhibited by the crop at different growth stages. These characteristics cover multiple aspects such as the crop's physiological state and growth environment, including plant height, leaf number, chlorophyll content, soil moisture, and nutrient content. A distributed sensor network is a network system composed of multiple sensors distributed at different locations in the target farmland. These sensors have data acquisition capabilities and can collect data related to crop growth in real time. A preset sampling frequency refers to the frequency at which the sensors collect data at preset time intervals, such as once every hour.
[0012] For example, to obtain crop growth stage characteristic data of a target farmland within a preset growth cycle, a distributed sensor network needs to be scientifically and rationally deployed in the target farmland. For instance, based on factors such as farmland topography and crop planting layout, soil moisture sensors, light sensors, temperature sensors, and crop physiological index sensors are installed in different areas of the farmland. Soil moisture sensors can be capacitance-based sensors, indirectly obtaining soil moisture information by measuring the dielectric constant of the soil; light sensors can be photodiode-type sensors, using the photoelectric effect to convert light intensity into an electrical signal for measurement; temperature sensors can be thermocouple sensors, measuring ambient temperature based on the thermoelectric effect of thermocouples. These sensors collect data at a preset sampling frequency, for example, setting data to be collected every hour. The collected data carries its own timestamp information and is then sent to a data processing center via a data transmission module. At the data processing center, a timestamp sorting algorithm is used to process the data, arranging the data collected by different sensors in chronological order to achieve timestamp alignment. For example, the bubble sort algorithm can be used to sort the data by timestamps, placing data with earlier timestamps at the beginning and data with later timestamps at the end.
[0013] Step S200: Based on the crop growth stage characteristic data, perform model permission level division processing to generate a stage permission configuration set. The stage permission configuration set includes the model parameter access permission identifier and permission effective time interval corresponding to different growth stages.
[0014] Model permission level classification refers to the process of dividing model usage permissions into different levels based on the characteristics and needs of crop growth stage data. A stage-based permission configuration set is a collection of multiple permission configuration units, each corresponding to a different crop growth stage. Each permission configuration unit records the model parameter access permission identifier and the permission's effective time interval for that growth stage. The model parameter access permission identifier is a code or symbol used to uniquely identify the access permission of model parameters, clearly defining the specific scope of access to model parameters, such as whether read or modification operations are allowed. The permission's effective time interval refers to the period during which the corresponding model parameter access permission is valid; outside this period, the permission will expire.
[0015] In one implementation, step S200 may specifically include the following steps S210 to S260: Step S210: Perform time series segmentation processing on the crop growth stage feature data, identify the time boundaries of each growth stage based on the mutation points of crop physiological characteristics, and generate a growth stage time axis containing multiple continuous growth stages. Each stage of the growth stage time axis is marked with a start timestamp and an end timestamp.
[0016] When performing time series segmentation, a sliding window algorithm can be used to process crop growth stage characteristic data. First, a fixed-size sliding window is set and moved across the time series data. Within each window, statistical indicators of crop physiological characteristics, such as mean and variance, are calculated. When a statistical indicator within a window changes significantly, it is considered that there may be a mutation point in the crop physiological characteristic within that window. For example, when the variance of crop plant height suddenly increases within a window, it indicates that the crop plant height has changed drastically during that time period, and there may be a mutation point in the physiological characteristic. Then, by further analyzing the specific changes in crop physiological characteristics, the precise location of the mutation point is determined. Based on these mutation points, the crop growth stage characteristic data is segmented into different time periods, thus generating a growth stage time axis. After generating the growth stage time axis, a start timestamp and an end timestamp are marked for each growth stage. For example, the time of the first data point of each growth stage can be extracted from the crop growth stage characteristic data as the start timestamp, and the time of the last data point as the end timestamp.
[0017] Step S220: Extract the crop growth feature vector corresponding to each growth stage in the growth stage time axis. The crop growth feature vector contains multiple physiological indicators that reflect the physiological state of the crop. The dimension of the feature vector is consistent with the number of physiological indicators collected.
[0018] A crop growth feature vector is a vector composed of multiple physiological indicators reflecting the crop's physiological state. It integrates multiple physiological characteristics of the crop at a certain growth stage and represents them in vector form. Physiological indicators are specific parameters used to describe the crop's physiological state, such as plant height, leaf area, photosynthetic rate, and respiration rate. The dimension of the feature vector refers to the number of elements contained in the vector, which is equal to the number of physiological indicators collected.
[0019] When extracting crop growth feature vectors, for each growth stage on the growth timeline, all physiological indicator data corresponding to that growth stage are selected from the crop growth stage feature data. For example, for a certain growth stage, physiological indicator data such as crop height, leaf area, photosynthetic rate, and respiration rate are extracted from the crop growth stage feature data. Then, these physiological indicator data are arranged in a certain order to form a vector, which is the crop growth feature vector corresponding to that growth stage.
[0020] Step S230: Analyze the fluctuation range of crop growth feature vectors at each growth stage, and calculate the feature fluctuation coefficient. The feature fluctuation coefficient is the average difference between the maximum and minimum values of each feature dimension within the stage divided by the average feature vector magnitude of the stage, which is used to characterize the degree of dynamic change of the feature at that stage.
[0021] When calculating the characteristic fluctuation coefficient, firstly, for each growth stage on the growth timeline, the difference between the maximum and minimum values of each dimension in the crop growth characteristic vector is calculated. For example, for a 5-dimensional crop growth characteristic vector, the difference between the maximum and minimum values of the data in each dimension is calculated. Then, these differences are averaged to obtain the average difference between the maximum and minimum values of each characteristic dimension within the stage. Next, the average value of the crop growth characteristic vector within that growth stage is calculated to form the stage average characteristic vector. Then, the magnitude of the stage average characteristic vector is calculated. Finally, the average difference between the maximum and minimum values of each characteristic dimension within the stage is divided by the magnitude of the stage average characteristic vector to obtain the characteristic fluctuation coefficient for that growth stage.
[0022] Step S240: Based on the feature fluctuation coefficient and the prediction error sensitivity of the federated learning model parameters, establish a correlation model between the growth stage and the permission level. The prediction error sensitivity is obtained through the correlation analysis between historical feature data and model prediction error. The larger the feature fluctuation coefficient and the higher the sensitivity, the higher the corresponding permission level.
[0023] The prediction error sensitivity of federated learning model parameters refers to the degree to which the prediction error of the federated learning model is sensitive to changes in crop growth characteristic data. It reflects the extent to which small changes in crop growth characteristic data can affect the prediction error of the federated learning model. Historical characteristic data refers to crop growth stage characteristic data collected in the past crop growth cycle.
[0024] In one implementation, step S240 may specifically include the following steps S241 to S246: Step S241: Collect the crop growth feature vectors and the prediction error rate data of the federated learning model parameters for each growth stage in the historical crop growth cycle, and generate a feature-error association dataset. The feature-error association dataset contains the prediction error rate sequence of model parameters under different feature fluctuation coefficients.
[0025] The prediction error rate of federated learning model parameters refers to the proportion of error between the predicted result and the actual result when the federated learning model predicts crop growth-related data. The feature-error association dataset contains information on the correlation between crop growth feature vectors and model parameter prediction error rates. It organizes and records the model parameter prediction error rates under different feature fluctuation coefficients, forming a sequence. During data collection, crop growth feature vector data for each growth stage in multiple past crop growth cycles is extracted from historical data storage systems, along with the corresponding federated learning model parameter prediction error rates. For example, for the past five wheat growth cycles, crop growth feature vectors and model prediction error rates for each growth stage in each cycle are extracted. Then, the feature fluctuation coefficient for each growth stage is calculated, and the feature fluctuation coefficient is correlated with the corresponding model parameter prediction error rate. These correlated data are sorted according to the feature fluctuation coefficient to generate a sequence of model parameter prediction error rates under different feature fluctuation coefficients, thus obtaining the feature-error association dataset. A database can be used to store the feature-error association dataset, with each record containing the feature fluctuation coefficient and the corresponding model parameter prediction error rate.
[0026] Step S242: Perform bivariate correlation analysis on the feature-error association dataset, calculate the correlation index between the feature fluctuation coefficient and the prediction error rate, and generate sensitivity weight values. The sensitivity weight values are used to characterize the degree of influence of feature fluctuation on the prediction accuracy of model parameters.
[0027] When performing bivariate correlation analysis, the Pearson correlation coefficient method can be used to calculate the correlation index between the feature fluctuation coefficient and the prediction error rate. First, extract the feature fluctuation coefficient and the prediction error rate from the feature-error association dataset. Then, calculate the Pearson correlation coefficient using the formula, and use it as the correlation index between the feature fluctuation coefficient and the prediction error rate. The strength and direction of the correlation between the feature fluctuation coefficient and the prediction error rate can be determined based on the magnitude and sign of the correlation index. Finally, the correlation index is standardized to generate a sensitivity weight value. For example, the correlation index can be mapped to the interval [0,1] to obtain the sensitivity weight value; the larger the value, the greater the impact of feature fluctuation on the prediction accuracy of the model parameters.
[0028] Step S243: Multiply the sensitivity weight value with the feature fluctuation coefficient to obtain the comprehensive sensitivity index. The comprehensive sensitivity index is used to quantify the comprehensive influence of feature fluctuations during the growth stage on the model parameters.
[0029] When calculating the comprehensive sensitivity index, the sensitivity weight value obtained in step S242 is directly multiplied by the feature fluctuation coefficient calculated in step S230. For example, if the sensitivity weight value is w and the feature fluctuation coefficient is k, then the comprehensive sensitivity index S = w × k. Through this product operation, the degree of influence of feature fluctuation on the prediction accuracy of model parameters (represented by the sensitivity weight value) is combined with the magnitude of feature fluctuation (represented by the feature fluctuation coefficient), resulting in an index that can more comprehensively reflect the comprehensive influence of feature fluctuation on model parameters during the growth stage.
[0030] Step S244: Based on the preset permission level division standard of the federated learning system, the range of values of the comprehensive sensitivity index is divided into multiple continuous intervals, each interval corresponding to a permission level. The interval boundary values are determined by cluster analysis of the system's historical permission configuration data.
[0031] When dividing the data into intervals, the first step is to obtain the default permission level classification criteria of the federated learning system, such as dividing permission levels into low, medium, and high. Then, cluster analysis is performed on the system's historical permission configuration data. The K-means clustering algorithm can be used to cluster the comprehensive sensitivity index in the historical permission configuration data. Based on the clustering results, the center and boundaries of each cluster are determined, and these boundaries are used as the dividing boundaries for the comprehensive sensitivity index value range, thus dividing the comprehensive sensitivity index value range into multiple continuous intervals. Each interval corresponds to a permission level; for example, intervals with lower comprehensive sensitivity indices correspond to low permission levels, and intervals with higher comprehensive sensitivity indices correspond to high permission levels.
[0032] Step S245: Using the start and end timestamps of the growth stage and the comprehensive sensitivity index as inputs, construct the input layer of the association model. Map the input to the permission level output layer through a fully connected network layer. The activation function of the network layer adopts a non-linear mapping function to ensure the probability distribution characteristics of the output permission level.
[0033] When constructing the association model, an input layer is first created, using the start and end timestamps of the growth phase and the comprehensive sensitivity index as input nodes. For example, the input layer can have three nodes, corresponding to the start, end, and comprehensive sensitivity indices, respectively. Next, a fully connected network layer is constructed, the number of nodes of which can be adjusted according to the actual situation. The nodes in the input layer are fully connected to the nodes in the fully connected network layer, with each connection having a corresponding weight. In the fully connected network layer, the input data is linearly combined and transformed. Then, a non-linear mapping function, such as the Softmax function, is applied between the fully connected network layer and the permission level output layer. The Softmax function converts the output of the fully connected network layer into a probability distribution, giving each permission level a corresponding probability value. Finally, the permission level output layer outputs the corresponding permission level according to the probability distribution; the permission level with the highest probability is the final output permission level.
[0034] Step S246: Train the association model based on the feature-error association dataset, optimize the model parameters by minimizing the loss value between the predicted permission level and the actual permission level, until the permission level prediction accuracy of the model on the validation set reaches the preset threshold, and complete the construction of the association model between the growth stage and the permission level.
[0035] When training the association model, the feature-error association dataset is divided into a training set and a validation set, for example, in an 8:2 ratio. The association model is trained using the training set, with each input sample (containing feature fluctuation coefficients and prediction error sensitivity). The association model outputs a predicted permission level, and the loss between the predicted and actual permission levels is calculated, for example, using the cross-entropy loss function. Then, optimization algorithms (such as stochastic gradient descent) are used to adjust the parameters of the association model, gradually reducing the loss. After each training iteration, the model is evaluated using the validation set, and the permission level prediction accuracy on the validation set is calculated. When the model's permission level prediction accuracy on the validation set reaches a preset threshold, training stops, completing the growth phase and the construction of the permission level association model.
[0036] Step S250: Combine the start timestamp, end timestamp, and permission level output by the associated model for each growth stage to generate a permission configuration unit containing permission level labels, permission effective start time, and permission effective end time. Each growth stage corresponds to one permission configuration unit.
[0037] When generating permission configuration units, for each growth stage in the growth stage timeline, the start and end timestamps of that growth stage are obtained. The start timestamp is then used as the permission's effective start time, and the end timestamp as the permission's effective end time. Simultaneously, the permission level corresponding to that growth stage output by the association model is converted into a permission level label. For example, if the permission level output by the association model is a high permission level, its permission level label can be marked as "high". The permission level label, the permission's effective start time, and the permission's effective end time are combined to form the permission configuration unit. Each growth stage corresponds to one such permission configuration unit, serving as the foundational information for subsequent permission management.
[0038] Step S260: Arrange all permission configuration units in the order of the growth stage timeline, merge configuration units with the same permission level in adjacent stages, and add permission priority identifiers to resolve permission conflicts when time intervals overlap, thus obtaining a complete set of staged permission configurations.
[0039] For example, all permission configuration units are first arranged in chronological order according to their growth stages. Then, the arranged permission configuration units are traversed, checking if adjacent permission configuration units have the same permission level. If they do, the permission effective time intervals of these two permission configuration units are merged. For example, if the permission effective time interval of the previous permission configuration unit is [t1, t2] and the permission effective time interval of the next permission configuration unit is [t2, t3], and their permission levels are the same, they are merged into one permission configuration unit with a permission effective time interval of [t1, t3]. During the merging process, the permission level label of one of the permission configuration units is retained. Next, a permission priority identifier is added to each permission configuration unit. The permission priority identifier can be determined based on factors such as the importance of the growth stage and the level of permission. For example, it can be stipulated that the more critical the growth stage and the higher the permission level, the higher the priority of the permission configuration unit. When time intervals overlap, the permission settings of the permission configuration unit with higher permission priority are adopted first. Finally, all processed and identifiable permission configuration units are integrated together to obtain a complete set of staged permission configurations.
[0040] Step S300: Associate and map the set of phased permission configurations with the local model parameters of each participant in the federated learning system to generate a set of local model parameters with permission identifiers. The set of local model parameters with permission identifiers includes local model parameters and their corresponding permission level labels, permission effective time intervals, and access control conditions.
[0041] Association mapping is the process of mapping and associating permission information in the phased permission configuration set with the local model parameters of each participant in the federated learning system. The purpose is to assign corresponding permission information to each local model parameter. The local model parameters of each participant in the federated learning system refer to the parameters of the local models possessed by each institution or node participating in the federated learning; these parameters are key data for model computation and prediction. The set of local model parameters with permission identifiers is a collection containing local model parameters and their corresponding permission level labels, permission effective time intervals, and access control conditions. Permission level labels specify the access permission level for the local model parameter. The permission effective time interval specifies the time period during which the permission for the local model parameter is valid. Access control conditions are a series of conditions used to restrict access to local model parameters, such as the types of participants allowed to access, encryption requirements for data transmission, and frequency limits for parameter calls.
[0042] In one implementation, step S300 may specifically include the following steps S310 to S360: Step S310: Obtain the local model metadata of each participant from the participant node registration center of the federated learning system. The local model metadata includes model architecture description, parameter name list, parameter dimension information and parameter storage path. The model architecture description records the functional positioning and data flow relationship of each layer of parameters.
[0043] The participant node registry center in a federated learning system is a module or database used to manage and record participant node information. Participant local model metadata refers to the relevant metadata of the local models possessed by each institution or node participating in the federated learning process. This metadata helps in understanding the structure and parameters of the local models. The model architecture description is a detailed description of the overall structure and hierarchical relationships of the local model, recording the functional positioning of parameters at each layer and the flow and relationships of data within the model. The parameter name list is a list containing all parameter names in the local model, allowing for easy identification and referencing of each parameter. Parameter dimension information refers to the dimensionality of each parameter; for example, for a matrix parameter, parameter dimension information can represent the number of rows and columns of the matrix. The parameter storage path refers to the storage location of the local model parameters in the storage system, allowing for convenient access and retrieval of model parameters.
[0044] Step S320: Perform functional classification processing on the parameter name list in the local model metadata. Based on the role of the parameters in the federated learning process, divide the local model parameters into parameter groups of different functional types, with each parameter group corresponding to a different data interaction range.
[0045] Functional classification is the process of categorizing local model parameters according to their function in the federated learning process. The role of a parameter in the federated learning process refers to the specific task and function it undertakes at each stage, such as data transmission, model training, and parameter aggregation. Functional parameter groups are sets of parameters with the same or similar functions. The scope of data interaction refers to the range within which parameters participate in data exchange and sharing within the federated learning system. Different functional parameter groups correspond to different scopes of data interaction; for example, some parameters may only interact within the local model, while others may require cross-institutional interaction with other participants.
[0046] In one implementation, step S320 may specifically include the following steps S321 to S326: Step S321: Parse the model architecture description of the local model metadata, extract the input source identifier and output target identifier of each layer parameter, generate the parameter data flow graph, the parameter data flow graph uses directed edges to represent the parameter transmission path, and nodes represent the model layer or external interaction interface.
[0047] When parsing the model architecture description, natural language processing techniques and regular expressions can be used to extract input source identifiers and output target identifiers. For example, for the text "The output of layer A serves as the input of layer B" in the model architecture description, regular expressions can be used to match "layer A" as the output target identifier and "layer B" as the input source identifier. Then, based on the extracted input source identifiers and output target identifiers, a parameter data flow graph can be constructed. The parameter data flow graph can be implemented using graph data structures from graph theory. Nodes can be represented by objects, and directed edges can be used to connect nodes. For example, one node object can be created to represent model layer A, another node object can be created to represent model layer B, and then a directed edge object can be created from node A to node B, indicating that parameters are passed from layer A to layer B.
[0048] Step S322: Identify parameters whose output target identifiers in the parameter data flow diagram contain cross-participant aggregation interfaces, and classify such parameters into the first functional type parameter group. The cross-participant aggregation interface is the interaction interface responsible for cross-participant parameter aggregation in the federated learning system.
[0049] The cross-participant aggregation interface is used in federated learning systems to aggregate parameters across multiple participants, summarizing and integrating model parameters from multiple participants. The first functional type parameter group is a group of parameters whose output target identifier includes the cross-participant aggregation interface; these parameters need to participate in the cross-participant parameter aggregation operation during federated learning.
[0050] When identifying parameters, all nodes and edges in the parameter data flow graph are traversed, and each parameter node's output target identifier is checked to see if it contains a cross-participant aggregation interface. This check can be performed by comparing the output target identifier with a predefined cross-participant aggregation interface identifier, and parameters that meet the criteria are grouped into the first function type parameter group.
[0051] In one implementation, step S322 may specifically include the following steps S3221 to S3226: Step S3221: Extract the output target identifier set of all parameter nodes from the parameter data flow graph. The output target identifier set is a list of interface identifiers of all target nodes pointed to by the parameter nodes.
[0052] Parameter nodes are nodes in the parameter data flow graph, representing parameters in the model. The output target identifier set is a collection containing the output target identifiers of all parameter nodes; it is a list of interface identifiers of all target nodes pointed to by the parameter nodes. When extracting the output target identifier set, all parameter nodes in the parameter data flow graph are traversed. For each parameter node, the interface identifiers of all target nodes it points to are obtained. For example, using a graph data structure traversal algorithm, such as a depth-first search algorithm, starting from a parameter node, all its target nodes are visited along the directed edges, and the interface identifiers of the target nodes are recorded. The output target identifiers of all parameter nodes are then aggregated into a list, forming the output target identifier set. A list data structure can be used to store the output target identifier set.
[0053] Step S3222: Traverse the output target identifier set and check whether it contains the cross-participant aggregation interface identifier preset by the federated learning system. The identifier is the standard interface identifier defined in the system configuration file.
[0054] During the check, a preset cross-participant aggregation interface identifier is read from the system configuration file. Then, each interface identifier in the output target identifier set is compared sequentially with the preset cross-participant aggregation interface identifier. String comparison functions can be used for this comparison. If an interface identifier in the output target identifier set is identical to the preset cross-participant aggregation interface identifier, then there exists an output target identifier that contains parameters of the cross-participant aggregation interface.
[0055] Step S3223: For parameter nodes containing cross-participant aggregation interface identifiers, extract their corresponding parameter names and parameter dimension information, and record them in the first function type candidate list.
[0056] For example, when a parameter node containing a cross-participant aggregation interface identifier is found in the output target identifier set, the parameter name and parameter dimension information corresponding to that parameter node are obtained from the parameter data flow graph. This information can be obtained through the properties of the parameter node object. The extracted parameter name and parameter dimension information are recorded in the first function type candidate list. The first function type candidate list can be stored using a list-nested dictionary approach, with each dictionary containing parameter name and parameter dimension information.
[0057] Step S3224: Query the aggregation parameter registry of the federated learning system to obtain the list of parameter names that have participated in the federated aggregation in the past. Compare the candidate list of the first function type with the list of historical parameter names, retain the parameters in the intersection of the two, and remove the parameters that were only added in the current period but have not actually participated in the aggregation.
[0058] The aggregation parameter registry of the federated learning system is a database or data structure used to record parameter information that has historically participated in federated aggregation. The list of parameter names that historically participated in federated aggregation refers to the list of parameter names that actually participated in parameter aggregation operations during past federated learning processes, obtained from the aggregation parameter registry. The first function type candidate list is a list generated in step S3223 containing parameter names and parameter dimension information that may belong to the first function type parameter group. The intersection refers to the parameters corresponding to parameter names commonly included in both the first function type candidate list and the historical parameter name list. Parameters added only in the current cycle but not actually participating in aggregation refer to parameters that newly appear in the current federated learning cycle, but which have not participated in parameter aggregation operations during past federated learning processes.
[0059] Step S3225: Perform aggregation frequency statistics on the retained parameters, calculate the average number of aggregations in the most recent multiple federated learning cycles, and if the average number of aggregations is greater than zero, then it is confirmed as a first function type parameter.
[0060] When performing aggregation frequency statistics, the aggregation count records for each retained parameter within the most recent multiple federated learning cycles are retrieved from the aggregation parameter registry of the federated learning system. For example, an SQL statement is used to query the aggregation count of a parameter within a specified cycle range. The aggregation counts for each parameter within the most recent multiple federated learning cycles are summed, and then divided by the number of federated learning cycles to obtain the average aggregation count. If the average aggregation count is greater than zero, the parameter is considered to have actually participated in parameter aggregation operations during the federated learning process and is identified as a first-function type parameter.
[0061] Step S3226: Add a first function type label and an aggregation frequency attribute to the confirmed first function type parameter, and generate a list of first function type parameters. The aggregation frequency attribute is used for weight adjustment during subsequent permission level mapping.
[0062] The first function type label is used to identify that a parameter belongs to the first function type parameter group. This label makes it easy to identify and distinguish parameters of different function types. The aggregation frequency attribute records the average number of aggregations of a parameter over the most recent federated learning cycles, which can be used as a weighting adjustment factor when mapping permission levels. The first function type parameter list is a list containing all confirmed first function type parameters and their related information, including parameter name, first function type label, and aggregation frequency attribute.
[0063] Step S323: Identify parameters in the parameter data flow diagram where both the input source identifier and the output target identifier contain only the local model internal interface. Classify such parameters into the second function type parameter group. The local model internal interface refers to the interface in the participant's local model that has not established data interaction with external nodes.
[0064] The local model internal interface is the interface used for internal data interaction within the participant's local model; these interfaces do not exchange data with external nodes. The second function type parameter group is formed by grouping parameters whose input source identifier and output target identifier both contain only the local model internal interface. These parameters only interact internally within the local model during federated learning. When identifying parameters, all parameter nodes in the parameter data flow graph are traversed. For each parameter node, it is checked whether its input source identifier and output target identifier both contain only the local model internal interface. This can be checked by comparing the input source identifier and output target identifier with predefined local model internal interface identifiers. If both the input source identifier and output target identifier contain only the local model internal interface identifier, the parameter is classified into the second function type parameter group.
[0065] Step S324: Identify parameters whose output target identifiers in the parameter data flow diagram contain other participant interaction interfaces but not cross-participant aggregation interfaces. Classify such parameters into the third function type parameter group. Other participant interaction interfaces refer to the interaction interfaces of other institutional nodes in the federated learning system, excluding local nodes.
[0066] When identifying parameters, all parameter nodes in the parameter data flow graph are traversed. For each parameter node, its output target identifier is checked to see if it contains other participant interaction interfaces but not cross-participant aggregation interfaces. This check can be performed by comparing the output target identifier with predefined other participant interaction interface identifiers and cross-participant aggregation interface identifiers. If the output target identifier contains other participant interaction interface identifiers but not cross-participant aggregation interface identifiers, the parameter is classified into the third function type parameter group. A list can be used to store the parameter names in the third function type parameter group.
[0067] Step S325: Perform cross-validation on the grouped parameters to check if there are any parameters that simultaneously meet the conditions of multiple functional types. If so, determine the final parameter group according to the priority of the output target in the parameter data flow diagram. The priority is sorted from largest to smallest according to the data interaction range.
[0068] During cross-validation, all grouped parameters are iterated through. For each parameter, it is checked whether it appears in multiple function type parameter groups simultaneously. If such a parameter exists, its final parameter group is determined based on the priority of the output target in the parameter data flow graph. For example, the data interaction scope of the first function type parameter group is cross-participant aggregation, the data interaction scope of the second function type parameter group is within the local model, and the data interaction scope of the third function type parameter group is interaction with other participants but not aggregation. The first function type parameter group has the highest priority and the second function type parameter group has the lowest priority, sorted by data interaction scope from largest to smallest. If a parameter satisfies the conditions of both the first and third function types, it is assigned to the first function type parameter group.
[0069] Step S326: Generate a parameter classification result table, which includes parameter name, function type label, data flow description and classification basis.
[0070] The parameter classification result table records the parameter classification results, clearly displaying the classification information for each parameter. The parameter name is the name of the parameter in the local model, used to uniquely identify a parameter. The function type label is a label used to identify the function type parameter group to which the parameter belongs, such as first function type label, second function type label, third function type label, etc. The data flow description describes the direction and path of data flow of the parameter in the model, helping to understand the parameter's function and role. The classification basis refers to the reasons and basis for classifying a parameter into a certain function type parameter group, such as classification based on the output target identifier in the parameter data flow diagram.
[0071] Step S330: Extract the permission level tags, permission effective time intervals, and access control rules from the phased permission configuration set, and generate a permission configuration element set. The access control rules include the types of participants allowed to access, data transmission encryption requirements, and parameter call frequency limits.
[0072] The permission configuration element set is a collection of key permission configuration information from the phased permission configuration set, integrating permission level labels, permission effective time intervals, and access control rules. Permission level labels identify the permission level; the permission effective time interval specifies the valid period of the permission; and access control rules are a series of rules used to restrict access to model parameters. Allowed participant types refer to the types of participants permitted to access model parameters according to the permission configuration; for example, some parameters may only allow access from participants in specific domains. Data transmission encryption requirements refer to the encryption conditions that must be met when transmitting parameter data, such as using a preset encryption algorithm to encrypt the data. Parameter call frequency limits refer to the limits on the number of times model parameters can be called; for example, specifying that each participant can only call a parameter a certain number of times within a certain period.
[0073] Step S340: Establish basic mapping rules between parameter function types and permission levels. Parameter groups of different function types are mapped to the permission levels required for the corresponding data interaction range in the phased permission configuration set, and the priority order of the mapping rules is recorded.
[0074] When establishing basic mapping rules, the data interaction scope of different functional type parameter groups is analyzed. For example, the data interaction scope of the first functional type parameter group is cross-participant aggregation, which typically requires a higher permission level; the data interaction scope of the second functional type parameter group is within the local model, requiring a relatively lower permission level. Based on the phased permission configuration set, the permission level required for the corresponding data interaction scope of each functional type parameter group is found, and a mapping relationship is established. For example, the first functional type parameter group is mapped to a high permission level, and the second functional type parameter group is mapped to a low permission level. Simultaneously, the priority order of the mapping rules is recorded. A dictionary can be used to store the basic mapping rules, where the key is the functional type label and the value is the corresponding permission level; a list can be used to store the priority order of the mapping rules.
[0075] Step S350: Match the local model parameters of each participant with the permission configuration element set according to the parameter name, add the corresponding permission level label, permission effective time range and access control rules to each parameter, and generate parameter units with permission attributes.
[0076] A parameter unit with permission attributes is a unit that contains local model parameters and their corresponding permission level labels, permission effective time intervals, and access control rules. In this step, the local model parameters of each participant are matched against the permission configuration element set. This involves finding the corresponding permission information in the permission configuration element set based on the parameter name and adding this permission information to the parameter. During the matching process, the local model parameters of each participant are traversed. For each parameter, the corresponding permission configuration element is searched in the permission configuration element set based on its parameter name. This can be achieved using the key-value lookup function of a dictionary, using the parameter name as the key to find the corresponding value in the permission configuration element set. If a corresponding permission configuration element is found, its permission level label, permission effective time interval, and access control rules are added to the parameter information. All parameters with added permission attributes are summarized into a list to generate parameter units with permission attributes.
[0077] Step S360: Group all parameter units with permission attributes according to their parameter storage paths, add participant organization identifiers and parameter version numbers to obtain a local model parameter set with permission identifiers. The parameter version number is used to track the update and iteration of parameters.
[0078] During grouping, all parameter units with permission attributes are traversed. For each parameter unit, it is assigned to the appropriate group based on its parameter storage path. Grouping can be implemented using a dictionary, where the key is the parameter storage path and the value is a list of parameter units with permission attributes under that path. Then, a participant organization identifier and parameter version number are added to each parameter unit with permission attributes. The participant organization identifier can be obtained from the participant node registry center, and the parameter version number can be automatically generated or manually set based on parameter updates. Finally, all grouped parameter units with permission attributes are aggregated to obtain a local model parameter set with permission identifiers. The local model parameter set with permission identifiers can be stored using a list nested into dictionaries, where each dictionary contains parameter unit information, participant organization identifier, and parameter version number.
[0079] Step S400: Perform differential encryption processing on the local model parameter set with permission identifiers. Generate a permission-encrypted model parameter set by adding a noise perturbation factor that is dynamically adjusted based on the permission level. The noise perturbation intensity of the permission-encrypted model parameter set is positively correlated with the permission level label.
[0080] In one implementation, step S400 may specifically include the following steps S410 to S470: Step S410: Extract the permission level labels and parameter dimension information of all parameters from the local model parameter set with permission identifiers, and generate a permission-dimensional lookup table. The permission-dimensional lookup table includes parameter name, permission level label, number of parameter rows and number of parameter columns. The number of parameter rows and columns are used to determine the dimension of the noise disturbance factor.
[0081] During the operation, the local model parameter set with permission identifiers is traversed. For each parameter, its permission level label, row number, and column number are extracted. This information can be obtained, for example, by accessing the parameter object's attributes. The extracted parameter names, permission level labels, row numbers, and column numbers are then compiled into a table to generate a permission-dimension lookup table. This lookup table facilitates the subsequent generation of appropriate noise perturbation factors based on the permission level and dimension information of different parameters.
[0082] Step S420: Assign an independent noise generation source to each permission level label. The noise generation source is an instance of a cryptographically secure pseudo-random number generator. Different permission level labels correspond to different initial seed values. The initial seed value is generated by combining the hash value of the permission level label with the system timestamp.
[0083] The noise generation source is the origin of the noise. In this step, an instance of a cryptographically secure pseudo-random number generator is used as the noise generation source. The cryptographically secure pseudo-random number generator can generate random number sequences with high randomness and security, ensuring the unpredictability of the generated noise. Different permission level labels correspond to different initial seed values. The initial seed value is a key parameter for starting the pseudo-random number generator; different values will result in different generated random number sequences. The initial seed value is generated by combining the hash value of the permission level label with the system timestamp, ensuring the uniqueness of the noise generation source for each permission level. In the actual allocation process, for each permission level label in the permission-dimension lookup table, its hash value is calculated. Simultaneously, the current system timestamp is obtained; the system timestamp is a numerical value representing the current time. The hash value of the permission level label is combined with the system timestamp, for example, through concatenation or other operations, to generate the initial seed value. Then, this initial seed value is used to initialize an instance of the cryptographically secure pseudo-random number generator, which serves as the noise generation source for the corresponding permission level label. Based on this, each permission level has an independent and unique noise generation source.
[0084] Step S430: Call the noise generation source to generate a basic noise matrix that matches the parameter dimension. The elements of the basic noise matrix follow a preset probability distribution, and the number of rows and columns of the matrix are strictly consistent with the dimension of the corresponding parameter.
[0085] The base noise matrix is a noise matrix added to the local model parameter set with access control. Its elements follow a preset probability distribution, such as a Gaussian or uniform distribution. This preset probability distribution determines the characteristics of the noise elements; different distributions produce noise with different characteristics. The number of rows and columns of the matrix strictly matches the dimensions of the corresponding parameters. This ensures that the noise matrix can be accurately added to the parameters, achieving encrypted parameter processing.
[0086] In one implementation, step S430 may specifically include the following steps S431 to S436: Step S431: Initialize the cryptographically secure pseudo-random number generator. Input the initial seed value corresponding to the permission level label into the generator, and set the random number sequence generation mode to continuous non-repeating mode to ensure that each call generates an independent random number sequence.
[0087] For example, first, an instance of a cryptographically secure pseudo-random number generator is obtained. Then, the initial seed value corresponding to the permission level label is passed as a parameter to the generator's initialization function to complete the generator's initialization. Next, the generator's generation mode is set to continuous, non-repeating mode, for example, by calling the generator's setting method or setting relevant parameters. Based on this, each time the generator is called, an independent and non-repeating random number sequence will be generated based on the input initial seed value.
[0088] Step S432: Read the number of rows and columns of the current parameter from the permission-dimension lookup table, and calculate the total number of random numbers to be generated. The total number is the product of the number of rows and columns.
[0089] For example, the parameter to be processed is found from the permission-dimension lookup table, and its corresponding row and column numbers are obtained. Then, the row and column numbers are multiplied to obtain the total number of random numbers to be generated. This total number will serve as the basis for subsequent random number sequence generation, ensuring that the number of generated random numbers is sufficient to fill the underlying noise matrix.
[0090] Step S433: Start the pseudo-random number generator and generate a specified number of random numbers in row priority order. Each random number is converted into a value that follows a preset probability distribution through a probability distribution transformation algorithm. During the transformation process, a preset transformation method is used to map uniformly distributed random numbers to target probability distribution random numbers.
[0091] A probability distribution transformation algorithm is an algorithm that converts uniformly distributed random numbers into random numbers that follow a preset probability distribution. The preset transformation method is the specific transformation method in the algorithm, such as using the inverse transformation method or the accept-reject method to map uniformly distributed random numbers to random numbers with a target probability distribution, such as Gaussian or Poisson distribution. In actual execution, a pseudo-random number generator that has been initialized and configured is started. Based on the total number of random numbers to be generated calculated in step S432, a corresponding number of random numbers are generated in row priority order. For each generated random number, the preset probability distribution transformation algorithm and transformation method are used to convert it from a uniform distribution to a value of the preset probability distribution. For example, if the preset probability distribution is a Gaussian distribution, the inverse transformation method can be used to convert uniformly distributed random numbers into Gaussian distributed random numbers.
[0092] Step S434: Fill the generated probability distribution values into a matrix of preset dimensions in row-major order to generate a basic noise matrix. The row and column elements of the matrix are arranged in the same order as the dimensions of the corresponding parameters.
[0093] During the filling process, an empty matrix with preset dimensions is created. Then, random numbers generated in step S433, following a preset probability distribution, are sequentially filled into each position of the matrix in row-major order. For example, the first row of the matrix is filled first, with random numbers being filled from left to right. After the first row is filled, the second row is filled, and so on, until all random numbers are filled into the matrix, forming the basic noise matrix. The row and column element order of this matrix is exactly the same as the dimension order of the corresponding parameters, ensuring the correctness of the encryption operation.
[0094] Step S435: Perform a randomness test on the basic noise matrix, calculate the statistical eigenvalues of all elements of the matrix, and if the statistical eigenvalues meet the characteristics of the preset probability distribution, then the matrix is determined to meet the randomness requirements; otherwise, regenerate the random number sequence and construct the matrix.
[0095] Randomness testing is an operation that evaluates the randomness of the underlying noise matrix. Its purpose is to ensure that the elements in the matrix are truly random and conform to the characteristics of a predefined probability distribution. Statistical eigenvalues are numerical values that describe the distributional characteristics of the matrix elements, such as mean, variance, skewness, and kurtosis. The predefined probability distribution's characteristics refer to the statistical features it possesses; for example, a Gaussian distribution has a set mean and variance. If the statistical eigenvalues of the underlying noise matrix meet the predefined probability distribution's characteristics, the matrix is considered to satisfy the randomness requirement and can be used for subsequent encryption operations; otherwise, a new random number sequence needs to be generated and the matrix reconstructed to ensure the randomness and effectiveness of the noise.
[0096] In one implementation, step S435 may specifically include the following steps S4351 to S4358: Step S4351: Convert the basic noise matrix into a one-dimensional array, and count the total number of elements in the array. The total number is the product of the number of rows and columns of the matrix.
[0097] For example, iterate through each row and column of the underlying noise matrix, adding the elements of the matrix sequentially to a one-dimensional array in row-major order. For instance, first add the elements of the first row to the array, then the second row, and so on, until all elements of the matrix have been added to the array. Finally, count the total number of elements in this one-dimensional array; this total number is the product of the number of rows and columns of the matrix, providing the basis for subsequent calculations of statistical eigenvalues.
[0098] Step S4352: Calculate the arithmetic mean of all elements in the one-dimensional array as the overall mean of the basic noise matrix. The overall mean reflects whether the central trend of the noise sequence conforms to the preset probability distribution.
[0099] When calculating the overall mean, all elements in the one-dimensional array are summed to obtain the sum of the elements. Then, the sum is divided by the total number of elements (the result obtained in step S4351) to obtain the overall mean. The overall mean is compared with the center value of the preset probability distribution. If the two are close, it means that the central trend of the noise sequence conforms to the preset probability distribution; if the difference is large, it may be necessary to regenerate the noise matrix.
[0100] Step S4353: Calculate the sum of squares of the differences between each element in the one-dimensional array and the overall mean to obtain the sum of squares of deviations from the mean. The sum of squares of deviations from the mean is used to measure the degree to which an element deviates from the central trend.
[0101] To calculate the sum of squared deviations from the mean, each element in the one-dimensional array is iterated through, and the difference between that element and the overall mean is calculated. This difference is then squared. The sum of the squared differences of all elements is then obtained. The larger this value, the greater the deviation of the element from the central tendency.
[0102] Step S4354: Divide the sum of squared deviations from the mean by the total number minus one to obtain the sample variance. The sample variance reflects whether the dispersion of the noise sequence meets the preset probability distribution requirements.
[0103] When calculating the sample variance, the sum of squared deviations from the mean obtained in step S4353 is divided by the total number of elements minus one (i.e., n-1, where n is the total number of elements). The calculated sample variance is compared with the discrete values of the preset probability distribution. If the two are close, it indicates that the dispersion of the noise sequence meets the requirements of the preset probability distribution; if the difference is large, it may be necessary to regenerate the noise matrix.
[0104] Step S4355: Read the preset mean tolerance threshold and variance tolerance threshold from the system security configuration file. The mean tolerance threshold is the maximum range that the overall mean is allowed to deviate from the preset probability distribution center value, and the variance tolerance threshold is the maximum range that the sample variance is allowed to deviate from the preset probability distribution discrete value.
[0105] The system security configuration file stores system security-related configuration information, including preset mean tolerance thresholds and variance tolerance thresholds. The mean tolerance threshold is a pre-set value that specifies the maximum range by which the overall mean can deviate from the preset probability distribution center value. Similarly, the variance tolerance threshold is a pre-set value that specifies the maximum range by which the sample variance can deviate from the preset probability distribution discrete value.
[0106] For example, the system security configuration file is opened, and preset mean tolerance thresholds and variance tolerance thresholds are read from it. These thresholds are preset based on the system's security requirements and the characteristics of the preset probability distribution. For instance, if the preset probability distribution is a Gaussian distribution, the system may set appropriate mean tolerance thresholds and variance tolerance thresholds based on the theoretical mean and variance of the Gaussian distribution, combined with security requirements.
[0107] Step S4356: Calculate the mean deviation and variance deviation. The mean deviation is the absolute value of the difference between the overall mean and the center value of the preset probability distribution, and the variance deviation is the absolute value of the difference between the sample variance and the discrete value of the preset probability distribution.
[0108] The mean deviation value measures the degree of difference between the overall mean and the central value of a preset probability distribution. By taking the absolute value of the difference, the influence of positive and negative values can be avoided, providing a more intuitive reflection of the degree of deviation. Similarly, the variance deviation value measures the degree of difference between the sample variance and the discrete values of the preset probability distribution.
[0109] When calculating the mean deviation, the overall mean calculated in step S4352 is subtracted from the center value of the preset probability distribution, and the absolute value of the difference is taken. When calculating the variance deviation, the sample variance calculated in step S4354 is subtracted from the discrete value of the preset probability distribution, and the absolute value of the difference is taken. By calculating these two deviation values, the degree of conformity between the statistical characteristics of the basic noise matrix and the preset probability distribution can be determined more accurately.
[0110] Step S4357: If the mean deviation is less than the mean tolerance threshold and the variance deviation is less than the variance tolerance threshold, then the basic noise matrix is determined to meet the randomness requirement.
[0111] Based on the mean tolerance threshold and variance tolerance threshold read in step S4355, and the mean deviation and variance deviation values calculated in step S4356, a judgment is made. If the mean deviation value is less than the mean tolerance threshold, it indicates that the overall mean is within the allowable deviation range; if the variance deviation value is less than the variance tolerance threshold, it indicates that the sample variance is also within the allowable deviation range. Only when both conditions are met simultaneously is the basic noise matrix deemed to meet the randomness requirement and can be used for subsequent encryption operations. During the judgment process, the mean deviation value is compared with the mean tolerance threshold, and the variance deviation value is also compared with the variance tolerance threshold. If both comparison results meet the conditions, i.e., the mean deviation value is less than the mean tolerance threshold and the variance deviation value is less than the variance tolerance threshold, then the statistical characteristics of the basic noise matrix are considered to meet the requirements of the preset probability distribution and satisfy the randomness requirement; otherwise, the basic noise matrix is considered not to meet the randomness requirement and further processing is required.
[0112] Step S4358: If not satisfied, update the initial seed value of the pseudo-random number generator. The new seed value is a combination of the original seed value and the number of times the current matrix generation failed. Repeat the step of generating the basic noise matrix until the randomness requirement is met.
[0113] When the basic noise matrix does not meet the randomness requirement, the initial seed value of the pseudo-random number generator needs to be updated to generate different random number sequences, thereby constructing a new basic noise matrix. The new seed value is a combination of the original seed value and the number of times the current matrix generation failed. This ensures that the seed value is different after each update, thus generating different random number sequences. The steps for generating the basic noise matrix are then repeated, including initialization, random number generation, matrix construction, and randomness checks starting from step S431, until the generated basic noise matrix meets the randomness requirement.
[0114] In one implementation, step S4358 may specifically include the following steps S43581 to S43587: Step S43581: Record the number of times the current basic noise matrix generation fails. The initial value is one, and it increments by one after each failed test.
[0115] For example, a variable is used to store the number of times the current base noise matrix generation has failed, initially set to one. When the randomness test result in step S4357 indicates that the base noise matrix does not meet the requirements, the value of this variable is incremented by one. In this way, the number of generation failures is accurately recorded, providing a basis for subsequent updates to the initial seed value.
[0116] Step S43582: Calculate a new seed value. The new seed value is generated by combining the original seed value with the current number of failures, ensuring that the seed value is unique and increments after each failure.
[0117] The calculation of the new seed value ensures that the seed value is different after each update, thus enabling the pseudo-random number generator to generate different random number sequences. A new seed value is generated by combining the original seed value with the current number of failures through operations such as concatenation, addition, and multiplication. This ensures that the seed value is unique and increments after each failure, avoiding the generation of identical random number sequences. When calculating the new seed value, the original seed value and the current number of failures are obtained. These are then processed according to pre-defined combination rules. For example, if concatenation is used, the original seed value is concatenated with the string representation of the current number of failures to obtain the new seed value. In this way, a new, unique, and incrementing seed value is obtained, which is used to reinitialize the pseudo-random number generator.
[0118] Step S43583: Input the new seed value into the pseudo-random number generator, reset the generator's internal state, and clear the previous random number sequence cache.
[0119] The new seed value calculated in step S43582 is passed as a parameter to the initialization function of the pseudo-random number generator to complete the generator's re-initialization. Simultaneously, relevant methods of the generator are called or parameters are set to reset the generator's internal state and clear the previously generated random number sequence cache. Based on this, the pseudo-random number generator can restart generating random number sequences based on the new seed value, preparing for the construction of a new basic noise matrix.
[0120] Step S43584: Regenerate a random number sequence with a specified number of rows and columns in row priority order, perform probability distribution transformation, and construct a new basic noise matrix.
[0121] Regenerating a random number sequence with a specified number of rows and columns in row-major order is to obtain different random numbers to construct a new base noise matrix. Performing a probability distribution transformation is to convert the generated random numbers into values that follow a preset probability distribution, ensuring that the elements of the new base noise matrix have the required distribution characteristics.
[0122] When regenerating the random number sequence, the length of the random number sequence to be generated is determined based on the number of rows and columns of the parameters in the permission-dimension lookup table. The re-initialized pseudo-random number generator is started, generating a specified number of random numbers in row-major order. For each generated random number, a preset probability distribution transformation algorithm and method are used to convert it into a value that follows a preset probability distribution. Then, these values are filled into a matrix of preset dimensions in row-major order to construct a new basic noise matrix.
[0123] Step S43585: Re-perform the randomness test on the newly constructed basic noise matrix, calculate the overall mean and sample variance, and compare the relationship between the mean deviation and the mean tolerance threshold, and the variance deviation and the variance tolerance threshold.
[0124] Re-perform the randomness test on the newly constructed basic noise matrix to assess whether the new matrix meets the randomness requirements of the preset probability distribution. Calculating the overall mean and sample variance is to obtain the statistical characteristic values of the matrix for subsequent comparisons. Comparing the mean deviation with the mean tolerance threshold and the variance deviation with the variance tolerance threshold is a key step in determining whether the matrix meets the randomness requirements.
[0125] For example, following steps S4351-S4354, the overall mean and sample variance of the newly constructed basic noise matrix are calculated. Then, the mean deviation and variance deviation are calculated, i.e., the absolute value of the difference between the overall mean and the central value of the preset probability distribution, and the absolute value of the difference between the sample variance and the discrete value of the preset probability distribution. The mean deviation is compared with the mean tolerance threshold, and the variance deviation is compared with the variance tolerance threshold. Based on the comparison results, it is determined whether the newly constructed basic noise matrix meets the randomness requirement.
[0126] Step S43586: If the new matrix passes the test, stop the iteration and record the current seed value and the number of failures.
[0127] When the newly constructed base noise matrix passes the randomness test—that is, the mean deviation is less than the mean tolerance threshold and the variance deviation is less than the variance tolerance threshold—it indicates that the new matrix meets the randomness requirements of the preset probability distribution and can be used for subsequent encryption operations. At this point, the iteration process stops, and the base noise matrix is no longer regenerated. Simultaneously, the currently used seed value and the number of times the base noise matrix generation failed are recorded; this information can be used for subsequent traceability and analysis.
[0128] Step S43587: If it still fails, repeat the steps of updating the seed value and generating the matrix until the number of failures reaches the preset maximum number of failures. At this time, trigger a system alarm and use the backup noise generation algorithm to generate the basic noise matrix.
[0129] If the newly constructed base noise matrix still fails the randomness test, it means that the current generation process has failed to obtain a noise matrix that meets the requirements. In this case, it is necessary to repeat the steps of updating the seed value and generating the matrix, and continue to attempt to generate a base noise matrix that meets the randomness requirements. The preset maximum number of failures is a pre-set threshold. When the number of failures reaches this threshold, it indicates that there may be other problems. At this time, a system alarm is triggered, and a backup noise generation algorithm is used to generate the base noise matrix.
[0130] Step S436: The basic noise matrix that has passed the test is associated with the parameter name and stored. The generation time and the state parameters of the random number generator are recorded. The state parameters are used to trace the generation process of the noise disturbance factor.
[0131] For example, a data structure (such as a dictionary) can be used to store the association between the base noise matrix and the parameter names, with the parameter name as the key and the corresponding base noise matrix as the value. Simultaneously, the generation time of the base noise matrix is recorded, for example, using a system timestamp. Furthermore, the state parameters of the random number generator are recorded, and this information is stored in the same data structure or in a separate storage location.
[0132] Step S440: Query the preset noise intensity adjustment coefficient table according to the permission level label to obtain the corresponding intensity coefficient. The noise intensity adjustment coefficient table contains the mapping relationship between permission level and intensity coefficient. The higher the permission level, the larger the intensity coefficient, ensuring that higher permission parameters add stronger noise disturbance.
[0133] The preset noise intensity adjustment coefficient table is a pre-defined table that records the mapping relationship between permission levels and intensity coefficients. Permission level labels identify the permission level of a parameter. By querying the noise intensity adjustment coefficient table, the intensity coefficient corresponding to that permission level label can be obtained. The intensity coefficient is used to adjust the noise intensity; the higher the permission level, the larger the corresponding intensity coefficient. This ensures that when encrypting parameters, higher-permission parameters receive stronger noise perturbations, better protecting the privacy of higher-permission parameters.
[0134] Step S450: Multiply the basic noise matrix and intensity coefficients element by element to generate a noise perturbation factor that is dynamically adjusted based on the permission level. The dimension of the noise perturbation factor is consistent with that of the basic noise matrix, and the element values increase as the permission level increases.
[0135] Element-wise multiplication of the base noise matrix with the intensity coefficients is performed to adjust the base noise matrix according to the access level, generating a dynamically adjusted noise perturbation factor. The dimension of the noise perturbation factor remains consistent with the base noise matrix to ensure accurate addition with parameters in subsequent encryption operations. The element values increase with increasing access level because higher access levels correspond to larger intensity coefficients, resulting in larger element values for the multiplied noise perturbation factor.
[0136] Step S460: For each parameter in the local model parameter set with permission identifier, perform matrix addition with the corresponding noise perturbation factor to generate encrypted model parameters, and record the generation source information and intensity coefficient of the noise perturbation factor to obtain the parameter unit with encrypted trajectory.
[0137] For each parameter in the local model parameter set with access permissions, a matrix addition operation is performed with its corresponding noise perturbation factor to add noise to the parameter, thus encrypting it. After generating the encrypted model parameters, the source information and intensity coefficient of the noise perturbation factor are recorded. This information constitutes the encrypted trajectory, used for subsequent auditing and traceability. The parameter unit with the encrypted trajectory is a unit that contains the encrypted model parameters and their encrypted trajectory information.
[0138] For example, the local model parameter set with permission identifiers is traversed. For each parameter, its corresponding noise perturbation factor (associated with the parameter name) is found. The parameter and the noise perturbation factor are then added together using a matrix addition operation. For example, if the parameter is P and the noise perturbation factor is N', then the encrypted model parameter P' = P + N'. Simultaneously, the generation source information of the noise perturbation factor (such as the seed value of the random number generator, state parameters, etc.) and its intensity coefficient are recorded. This information is then combined with the encrypted model parameters to form a parameter unit with an encrypted trajectory.
[0139] Step S470: Sort all parameter units with encrypted traces according to permission level labels, add encryption timestamps and checksums, and combine them into a permission encryption model parameter set. The checksum is used to verify the integrity of the parameters during transmission.
[0140] For example, a sorting algorithm (such as quicksort, bubble sort, etc.) is used to sort all parameter units with encrypted traces according to their permission level labels, for example, sorting them from highest to lowest permission level. The current system timestamp is obtained and added as an encrypted timestamp to each parameter unit with an encrypted trace. A checksum is calculated for each parameter unit, for example, using a hash algorithm (such as MD5, SHA-256, etc.) to calculate the checksum. The encrypted timestamp and checksum are added to the parameter unit. Finally, all parameter units with encrypted traces are combined to form the permission encryption model parameter set. During subsequent transmission, the integrity of the parameters can be verified by comparing the checksums.
[0141] Step S500: Based on the permission-encrypted model parameter set, perform a model access permission verification operation at the federated learning model call interface to generate a permission verification result for cross-institutional model calls. The permission verification result is used to indicate whether the model parameter call request meets the preset permission configuration requirements.
[0142] In one implementation, step S500 may specifically include the following steps S510 to S570: Step S510: Receive model parameter call requests sent by participants in the federated learning system, parse the caller identifier, target parameter name, call timestamp, and call purpose description in the request, and generate request parsing results. The caller identifier is the unique institutional identifier of the participant in the system, and the call purpose description includes the parameter usage scenario and expected application method.
[0143] In a federated learning system, participants refer to the various institutions or nodes involved in the process. A model parameter access request is a request sent by a participant to access model parameters. Request parsing is the process of extracting and processing the information within the request. The caller identifier is a unique identifier for each participant in the federated learning system, used to distinguish different participants. The target parameter name is the name of the model parameter the caller requests access to. The call timestamp records the time the request was sent. The purpose description includes the parameter's usage scenario and expected application, such as for model training or prediction. The request parsing result is the result obtained after parsing the request information, containing the caller identifier, target parameter name, call timestamp, and purpose description.
[0144] Step S520: Query the permission level label, permission effective time range and access control rules corresponding to the target parameter name from the permission encryption model parameter set, and generate the target parameter permission file. The target parameter permission file records the complete permission constraints of the parameter.
[0145] The permission encryption model parameter set is a collection containing model parameters and their permission information. Retrieving the permission level label, permission effective time interval, and access control rules corresponding to the target parameter name from this set is to obtain the specific permission information of the target parameter. The permission level label identifies the permission level of the target parameter, the permission effective time interval specifies the time period during which the parameter can be accessed, and the access control rules are a series of restrictions on parameter access, such as the types of participants allowed to access the parameter and data transmission encryption requirements. The target parameter permission file is a file that records the complete permission constraints of the target parameter, including information such as permission level labels, permission effective time intervals, and access control rules.
[0146] Step S530: Verify whether the call timestamp is within the permission effective time interval. If the call timestamp is earlier than the permission effective start time or later than the permission effective end time, a permission verification failure result will be generated directly. The reason for the failure is that the permission has not taken effect or has expired.
[0147] During the verification process, the call timestamp is compared with the start and end times of the permissions in the target parameter permission file. A time comparison algorithm can be used to determine if the call timestamp is within the permission's effective time range. If the call timestamp is not within this range, a permission verification failure result is generated, such as returning a result object containing a "verification failed" flag and the reason "permission not effective or expired". In this way, it is ensured that the call request is made within the valid time range of the permissions.
[0148] Step S540: Query the participant permission database of the federated learning system to obtain the institution attribute information corresponding to the caller identifier. The institution attribute information includes the domain classification, data contribution rating, security certification level and historical cooperation records. The data contribution rating reflects the data provision of the participant in federated learning.
[0149] During the query, a database query statement or data access interface is used to search for the corresponding record in the participant's permission database based on the caller's identifier. From the found record, information such as the domain category, data contribution rating, security certification level, and historical cooperation records are extracted and organized into organizational attribute information.
[0150] Step S550: Match and verify the organization attribute information with the access control rules in the target parameter permission file, check whether the domain category is in the allowed access list, whether the data contribution rating meets the minimum requirements, and whether the security authentication level meets the encryption requirements, and generate attribute matching results.
[0151] Matching and verifying organizational attribute information against access control rules in the target parameter permission file is a crucial step in determining whether a call request meets the permission conditions, comprehensively considering the organizational attributes of the participants and the access control requirements of the target parameters. The allowed access list is a list of the domain categories of participants permitted to access the target parameters, as specified in the access control rules. The minimum data contribution rating requirement is the minimum standard for data contribution rating set by the access control rules. The security authentication level encryption requirement is the security authentication level requirement for participants in the access control rules to ensure the security of data transmission. The attribute matching result is the judgment result of the matching between the organizational attribute information and the access control rules, indicating whether the call request meets the attribute-related permission requirements.
[0152] In one implementation, step S550 may specifically include the following steps S551 to S556: Step S551: Extract the list of allowed domain categories, minimum data contribution rating standards, and security authentication level standards from the access control rules of the target parameter permission file, and generate an attribute verification standard set. The domain category list is a set of category identifiers of the domain to which the participants belong, preset by the system.
[0153] Extracting the list of allowed domain categories, minimum data contribution rating standards, and security authentication level standards from the access control rules of the target parameter permission file is to clarify the specific attribute requirements that participants need to meet. The list of allowed domain categories is a set of domain category identifiers for participants allowed to access the target parameter. The minimum data contribution rating standard is a set of rating standards that participants must meet to access the target parameter. The security authentication level standard specifies the security authentication level requirements for participants. The attribute verification standard set is a collection of these attribute verification standards, providing clear criteria for subsequent attribute matching verification. During the extraction process, a parsing algorithm is used to extract the list of allowed domain categories, minimum data contribution rating standards, and security authentication level standards from the access control rules of the target parameter permission file. This information is then organized into an attribute verification standard set. For example, regular expressions or data structure search methods can be used for extraction. In this way, the standard information required for attribute verification is accurately obtained.
[0154] Step S552: Extract the domain classification identifier, data contribution rating index, and security certification level identifier of the caller from the organization attribute information. The data contribution rating index reflects the quantity and quality of the data provided by the participant, and the security certification level identifier reflects the level of the participant's security protection capability.
[0155] During the extraction process, the domain classification identifier, data contribution rating index, and security certification level identifier of the caller are obtained from the organization attribute information. This information can be retrieved by accessing the attributes or fields of the organization attribute information object. For example, if the organization attribute information is stored in a dictionary, the corresponding value can be retrieved using the dictionary key. In this way, the caller's attribute information is accurately obtained, providing data support for subsequent matching and verification.
[0156] Step S553: Check if the domain category identifier of the caller exists in the list of allowed domain categories. If it does not exist, mark the domain category verification as failed.
[0157] Checking whether the caller's domain category identifier exists in the list of allowed domain categories is a crucial step in determining whether a participant has the necessary domain permissions to access the target parameter. If the caller's domain category identifier is not in the list of allowed domain categories, it means that the participant's domain does not meet the access requirements, and the domain category verification fails. During the check, a list search algorithm, such as linear search or binary search, is used to search for the caller's domain category identifier in the list of allowed domain categories. If the identifier is not found, the domain category verification fails, for example, by setting a flag to "failed." This method provides an initial verification of the participant's domain permissions.
[0158] Step S554: If it exists, compare the data contribution rating index with the minimum data contribution rating standard. If the index does not meet the standard, mark the data contribution verification as failed.
[0159] If the domain category identifier of the caller exists in the list of allowed domain categories, it means that the participant has passed the initial verification of domain permissions. At this point, it is necessary to further compare the data contribution rating metric with the minimum data contribution rating standard. If the data contribution rating metric does not meet the minimum data contribution rating standard, it means that the participant's data contribution does not meet the requirements, and the data contribution verification is marked as failed.
[0160] In one implementation, step S554 may specifically include the following steps S5541 to S5546: Step S5541: Extract data contribution rating indicators from the institutional attribute information. The data contribution rating indicators are calculated by comprehensively considering the amount of training data, data annotation quality, and data update frequency provided by the caller within a preset time range. The calculation method is to first normalize the indicators of each dimension to a unified dimensionless score and then sum them according to preset weights.
[0161] Extracting data contribution rating indicators from institutional attribute information aims to obtain specific quantitative metrics of participants' data contributions. These indicators are calculated by comprehensively considering factors such as the amount of training data provided by the caller within a preset timeframe, the quality of data annotation, and the frequency of data updates. The amount of training data reflects the quantity of data provided, the quality of data annotation reflects the accuracy and usability of the data, and the frequency of data updates reflects the timeliness of the data. The calculation method involves first normalizing each dimension of the indicators to a unified dimensionless score, eliminating dimensional differences between different dimensions, and then weighting and summing them according to preset weights to obtain a comprehensive score as the data contribution rating indicator. During the extraction process, information such as the amount of training data, the quality of data annotation, and the frequency of data updates from the caller's institutional attribute information is obtained. This information is then normalized, for example, using a min-max normalization method to convert each dimension of the indicators to the range [0,1]. Finally, the normalized indicators are weighted and summed according to preset weights.
[0162] Step S5542: Read the minimum data contribution rating standard from the attribute verification standard set. The standard is determined according to the permission level of the target parameter. The higher the permission level, the stricter the standard.
[0163] Retrieving the minimum data contribution rating standard from the attribute validation standard set is to obtain the data contribution rating threshold that participants need to meet. The minimum data contribution rating standard is determined based on the permission level of the target parameter; the higher the permission level, the stricter the data contribution requirements for participants, meaning the higher the minimum data contribution rating standard. During the retrieval process, the minimum data contribution rating standard is obtained from the attribute validation standard set. This standard can be obtained by accessing the attributes or fields of the attribute validation standard set object. For example, if the attribute validation standard set is stored in a dictionary, the minimum data contribution rating standard can be obtained using the dictionary key.
[0164] Step S5543: Compare the data contribution rating index with the minimum data contribution rating standard. If the rating index is greater than or equal to the standard, the data contribution is deemed to have met the standard.
[0165] Comparing the data contribution rating metric with the minimum data contribution rating standard is a crucial step in determining whether a participant's data contribution meets the requirements. If the data contribution rating metric is greater than or equal to the minimum data contribution rating standard, it indicates that the participant's data contribution meets the requirements, and the data contribution is deemed satisfactory. The comparison process directly compares the data contribution rating metric with the minimum data contribution rating standard. If the data contribution rating metric is greater than or equal to the standard, the data contribution is deemed satisfactory, and the data contribution verification is marked as passed.
[0166] Step S5544: If the rating index is less than the standard, query the recent new data contribution records preset by the caller, calculate the contribution increment value of the new data. The contribution increment value is the sum of the new data after the various dimensions of the new data are normalized to a unified dimensionless score and then weighted with the same weight.
[0167] If the data contribution rating index is lower than the minimum data contribution rating standard, it indicates that the participant's historical data contributions have not met the requirements. In this case, the caller's recent preset new data contribution records are queried, and the contribution increment value of the new data is calculated. The contribution increment value of the new data is obtained by normalizing the dimensions of the new data, such as the amount of training data, data annotation quality, and data update frequency, and then summing them with the same weights as those used to calculate the data contribution rating index. During the query process, the caller's recent preset new data contribution records are retrieved from the participant's permission database or related data storage. The various dimensions of the new data are normalized, for example, using a minimum-maximum normalization method to convert them to the range [0,1]. Then, the normalized indicators are summed with the same weights to obtain the contribution increment value of the new data.
[0168] Step S5545: Compare the contribution increment value with the difference between the standard and the rating indicator. If the contribution increment value is greater than or equal to the difference, the data contribution potential is determined to meet the standard and is marked as needing to supplement contribution proof.
[0169] Comparing the incremental contribution value with the difference between the standard and rating indicators is to determine whether the contribution of recently added data by the participant can make up for the lack of historical data contribution. If the incremental contribution value is greater than or equal to the difference between the standard and rating indicators, it indicates that the participant has the potential to meet the data contribution requirements through the added data, and the data contribution potential is determined to be met. However, the participant needs to provide supplementary contribution proof, such as providing supporting materials such as detailed source and processing procedures of the added data.
[0170] During the comparison process, the difference between the standard and the rating indicator is calculated as D = SR (where S is the minimum data contribution rating standard and R is the data contribution rating indicator). Then, the contribution increment value I is compared with the difference D. If I ≥ D, the data contribution is deemed to have potentially met the standard and is marked as requiring supplementary contribution proof.
[0171] Step S5546: If the contribution increment value is still less than the difference, the data contribution verification is determined to have failed, and the reason for failure is recorded as insufficient historical data contribution and recent increment not meeting the standard.
[0172] If the contribution increment value is still less than the difference between the standard and the rating indicator, it indicates that the contributor's recently added data contribution cannot make up for the insufficiency of historical data contribution, and the data contribution verification is determined to be failed. During the determination process, when I<D, the data contribution verification is determined to be failed. The failure cause is recorded, for example, a failure cause field containing "insufficient historical data contribution and recent increment failing to meet the standard" is added to the verification result object.
[0173] Step S555: if the data contribution rating meets the standard, checking whether the security certification level identifier is higher than or equal to the security certification level standard, wherein the security certification levels are sorted in ascending order of protection capability; if the current level is lower than the required level, marking that the security certification verification fails.
[0174] If the data contribution rating meets the standard, it indicates that the contributor meets the requirement in terms of data contribution. At this time, it is necessary to further check whether the security certification level identifier is higher than or equal to the security certification level standard. The security certification levels are sorted in ascending order of protection capability, for example, they can be divided into Level 1, Level 2, Level 3, etc. If the contributor's security certification level identifier is lower than the security certification level standard, it indicates that the contributor's security protection capability does not meet the requirement, and the security certification verification is marked as failed.
[0175] In the checking process, the caller's security certification level identifier is compared with the security certification level standard. It can be determined whether the current level is higher than or equal to the required level according to the sorting rule of security certification levels. If the current level is lower than the required level, the security certification verification is marked as failed, for example, a flag bit is set to "failed". In this way, the security certification permission of the contributor is verified.
[0176] Step S556: if all verification items are passed, generating an attribute matching pass result and recording the specific matching condition of each verification item; if there is at least one verification item that fails, generating an attribute matching failure result.
[0177] It is determined whether all verification items are passed according to the results of domain classification verification, data contribution verification and security certification verification. If all verification items are passed, it indicates that the contributor's organization attribute information fully conforms to the access control rule of the target parameter, an attribute matching pass result is generated, and the specific matching condition of each verification item is recorded, such as that the affiliated domain classification is matched, the data contribution meets the standard, the security certification level meets the requirement, etc. If there is at least one verification item that fails, it indicates that the contributor's organization attribute information does not satisfy the access control rule of the target parameter, an attribute matching failure result is generated, and the specific failed verification item and cause are recorded.
[0178] During the judgment process, the status of the domain classification verification flag, data contribution verification flag, and security authentication verification flag is checked. If all flags are "passed," an attribute matching pass result is generated; if any flag is "failed," an attribute matching failure result is generated.
[0179] Step S560: Combine the description of the purpose of the call with the purpose restriction clauses in the access control rules to verify whether the parameter call scenario conforms to the preset purpose scope. If the purpose description contains unauthorized scenario keywords, the purpose is determined to be non-compliant.
[0180] During the verification process, the usage restriction clauses in the access control rules are first parsed to extract a list of unauthorized scenario keywords. Then, the usage description is analyzed to check for unauthorized scenario keywords. String matching algorithms, such as regular expression matching, can be used to search for unauthorized scenario keywords in the usage description. If any unauthorized scenario keywords are found in the usage description, the parameter usage scenario is deemed non-compliant.
[0181] Step S570: Combine the time validity verification results, attribute matching results, and usage compliance verification results to generate the permission verification results for cross-organization model calls. The permission verification results include verification pass or fail indicators and detailed verification item results.
[0182] The time validity verification result is the verification conclusion in step S530 regarding whether the call timestamp is within the permission effective time interval; the attribute matching result is the verification conclusion in step S550 regarding the matching of the caller organization attribute information with the target parameter access control rules; and the usage compliance verification result is the verification conclusion in step S560 regarding whether the parameter call scenario conforms to the preset usage scope. Combining the verification results from these three aspects, it is possible to comprehensively and accurately determine whether cross-organization model call requests meet the preset permission configuration requirements.
[0183] When generating permission verification results, the time validity verification result is checked first. If the call timestamp is not within the permission's effective time range, the permission verification is directly deemed to have failed, and the permission verification result includes a "Verification Failed" flag, with a detailed record of the reason for failure as "Permission not effective or expired". If the time validity verification passes, the attribute matching result is checked next. If the attribute matching result fails, the specific failed verification item and reason are recorded, such as "Domain classification mismatch", "Data contribution not meeting standards", "Insufficient security certification level", etc., and the permission verification result is also marked as "Verification Failed". If the attribute matching result passes, the usage compliance verification result is checked next. If the usage is non-compliant, the details of "Usage non-compliance" are recorded, and the permission verification result is marked as "Verification Failed". Only when the time validity verification, attribute matching verification, and usage compliance verification all pass, is the permission verification result marked as "Verification Passed", and the specific passing status of each verification item is recorded in the detailed verification item results, such as "Time valid", "Attribute matching passed", "Usage compliant", etc.
[0184] Figure 2 A hardware entity diagram of a computer device provided in an embodiment of the present invention, such as... Figure 2 As shown, the hardware entity of the computer device 1000 includes a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can run on the processor 1001, and the processor 1001 executes the program to implement the steps in the method of any of the above embodiments.
[0185] This invention also provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the model permission processing method using federated learning as described in any of the above embodiments.
[0186] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A model permission processing method applying federated learning, characterized in that, The method includes: Acquire crop growth stage characteristic data of the target farmland within a preset growth cycle. The crop growth stage characteristic data is collected by a distributed sensor network at a preset sampling frequency and aligned with timestamps. Based on the crop growth stage characteristic data, the model permission level is divided to generate a stage permission configuration set, which includes the model parameter access permission identifier and permission effective time interval corresponding to different growth stages. The local model metadata of each participant is obtained from the participant node registration center of the federated learning system. This local model metadata includes a model architecture description, a list of parameter names, parameter dimension information, and parameter storage paths. The model architecture description records the functional positioning and data flow relationships of parameters at each layer. The parameter name list in the local model metadata is functionally categorized, and based on the parameter's role in the federated learning process, local model parameters are divided into parameter groups of different functional types, each corresponding to a different data interaction range. Permission level tags, permission effective time intervals, and access control rules are extracted from the phased permission configuration set to generate a permission configuration element set. The access control rules include the types of participants allowed to access, data transmission encryption requirements, and parameter adjustment... Frequency limits are used; basic mapping rules are established between parameter function types and permission levels. Parameter groups of different function types are mapped to the permission levels required for the corresponding data interaction range in the phased permission configuration set, and the priority order of the mapping rules is recorded; the local model parameters of each participant are matched with the permission configuration element set according to the parameter name, and a corresponding permission level label, permission effective time interval, and access control rule are added to each parameter to generate parameter units with permission attributes; all parameter units with permission attributes are grouped according to the parameter storage path, and participant organization identifiers and parameter version numbers are added to obtain a local model parameter set with permission identifiers; the local model parameter set with permission identifiers includes local model parameters and their corresponding permission level labels and access control conditions; Differential encryption processing is performed on the local model parameter set with permission identifier. A permission-encrypted model parameter set is generated by adding a noise perturbation factor that is dynamically adjusted based on the permission level. The noise perturbation intensity of the permission-encrypted model parameter set is positively correlated with the permission level label. Based on the permission-encrypted model parameter set, a model access permission verification operation is performed at the federated learning model call interface to generate a permission verification result for cross-institutional model calls. The permission verification result is used to indicate whether the model parameter call request meets the preset permission configuration requirements.
2. The method according to claim 1, characterized in that, The step of performing model permission level classification based on the crop growth stage feature data to generate a stage permission configuration set includes: The crop growth stage feature data is processed by time series segmentation. The time boundaries of each growth stage are identified based on the mutation points of crop physiological characteristics. A growth stage time axis containing multiple consecutive growth stages is generated. Each stage of the growth stage time axis is marked with a start time stamp and an end time stamp. Extract the crop growth feature vector corresponding to each growth stage in the growth stage time axis. The crop growth feature vector contains multiple physiological indicators that reflect the physiological state of the crop. The dimension of the feature vector is consistent with the number of physiological indicators collected. The fluctuation range of the crop growth feature vector at each growth stage is analyzed, and the feature fluctuation coefficient is calculated. The feature fluctuation coefficient is the average of the difference between the maximum and minimum values of each feature dimension within the stage divided by the average feature vector magnitude of the stage, which is used to characterize the degree of dynamic change of the feature at that stage. Based on the feature fluctuation coefficient and the prediction error sensitivity of the federated learning model parameters, a correlation model between growth stage and permission level is established. The prediction error sensitivity is obtained through correlation analysis between historical feature data and model prediction error. The larger the feature fluctuation coefficient and the higher the sensitivity, the higher the corresponding permission level. By combining the start timestamp, end timestamp, and permission level output by the associated model for each growth stage, a permission configuration unit is generated that includes a permission level label, permission effective start time, and permission effective end time. Each growth stage corresponds to one permission configuration unit. All permission configuration units are arranged in the order of the growth stage timeline. Configuration units with the same permission level in adjacent stages are merged, and permission priority identifiers are added to resolve permission conflicts when time intervals overlap, resulting in a complete set of staged permission configurations.
3. The method according to claim 2, characterized in that, The establishment of a correlation model between growth stage and permission level based on the prediction error sensitivity of the characteristic fluctuation coefficient and federated learning model parameters includes: Collect crop growth feature vectors and prediction error rate data of federated learning model parameters at each growth stage in the historical crop growth cycle to generate a feature-error association dataset. The feature-error association dataset contains a sequence of model parameter prediction error rates under different feature fluctuation coefficients. A bivariate correlation analysis was performed on the feature-error association dataset to calculate the correlation index between the feature fluctuation coefficient and the prediction error rate, and a sensitivity weight value was generated. The sensitivity weight value is used to characterize the degree of influence of feature fluctuation on the prediction accuracy of model parameters. The sensitivity weight value is multiplied by the feature fluctuation coefficient to obtain the comprehensive sensitivity index, which is used to quantify the comprehensive influence of feature fluctuations during the growth stage on the model parameters. Based on the default permission level classification standard of the federated learning system, the range of values for the comprehensive sensitivity index is divided into multiple continuous intervals, each interval corresponding to a permission level. The interval boundary values are determined by cluster analysis of the system's historical permission configuration data. Using the start and end timestamps of the growth stage and the comprehensive sensitivity index as inputs, the input layer of the association model is constructed. The input is mapped to the permission level output layer through a fully connected network layer. The activation function of the network layer adopts a non-linear mapping function to ensure the probability distribution characteristics of the output permission level. The association model is trained based on the feature-error association dataset. The model parameters are optimized by minimizing the loss value between the predicted permission level and the actual permission level until the permission level prediction accuracy of the model on the validation set reaches a preset threshold, thus completing the construction of the association model between the growth stage and the permission level.
4. The method according to claim 1, characterized in that, The parameter name list in the local model metadata is functionally categorized, and the local model parameters are divided into parameter groups of different functional types based on their role in the federated learning process, including: The model architecture description of the local model metadata is parsed, the input source identifier and output target identifier of each layer parameter are extracted, and a parameter data flow graph is generated. The parameter data flow graph uses directed edges to represent the parameter transmission path and nodes to represent model layers or external interaction interfaces. The parameters whose output target identifiers in the parameter data flow diagram contain cross-participant aggregation interfaces are identified and classified into the first functional type parameter group. The cross-participant aggregation interface is the interaction interface responsible for cross-participant parameter aggregation in the federated learning system. The parameters whose input source identifier and output target identifier in the parameter data flow diagram both contain only the local model internal interface are classified into the second function type parameter group. The local model internal interface refers to the interface in the participant's local model that has not established data interaction with external nodes. The parameters whose output target identifiers in the parameter data flow diagram include other participant interaction interfaces but do not include cross-participant aggregation interfaces are classified into the third function type parameter group. The other participant interaction interfaces refer to the interaction interfaces of other institutional nodes in the federated learning system other than the local node. Cross-validate the grouped parameters to check if there are any parameters that simultaneously meet the conditions of multiple functional types. If so, determine the final parameter group based on the priority of the output target in the parameter data flow diagram. The priority is sorted from largest to smallest according to the data interaction range. Generate a parameter classification result table, which includes parameter name, function type label, data flow description and classification basis.
5. The method according to claim 4, characterized in that, The parameter data flow diagram outputs a target identifier containing parameters from a cross-participant aggregation interface. These parameters are categorized into a first functional type parameter group, including: Extract the output target identifier set of all parameter nodes from the parameter data flow graph. The output target identifier set is a list of interface identifiers of all target nodes pointed to by the parameter nodes. Traverse the set of output target identifiers and check whether it contains a cross-participant aggregation interface identifier preset by the federated learning system. The identifier is a standard interface identifier defined in the system configuration file. For parameter nodes that contain cross-participant aggregation interface identifiers, extract their corresponding parameter names and parameter dimension information and record them in the first function type candidate list; Query the registry of aggregation parameters of the federated learning system to obtain the list of parameter names that have participated in the federated aggregation in the past. Compare the candidate list of the first function type with the list of historical parameter names, retain the parameters in the intersection of the two, and remove the parameters that were only added in the current period but have not actually participated in the aggregation. The parameters to be retained are aggregated and the frequency is statistically analyzed. The average number of aggregations in the most recent federated learning cycles is calculated. If the average number of aggregations is greater than zero, it is confirmed as a first-function type parameter. Add a first function type label and an aggregation frequency attribute to the confirmed first function type parameter to generate a list of first function type parameters.
6. The method according to claim 1, characterized in that, The step of performing differential encryption processing on the local model parameter set with permission identifiers, and generating a permission-encrypted model parameter set by adding a noise perturbation factor dynamically adjusted based on permission level, includes: Extract the permission level labels and parameter dimension information of all parameters from the local model parameter set with permission identifiers, and generate a permission-dimensional lookup table. The permission-dimensional lookup table includes parameter name, permission level label, number of parameter rows and number of parameter columns. The number of parameter rows and number of columns are used to determine the dimension of the noise disturbance factor. Each permission level label is assigned an independent noise generation source, which is an instance of a cryptographically secure pseudo-random number generator. Different permission level labels correspond to different initial seed values, which are generated by combining the hash value of the permission level label with the system timestamp. The noise generation source is invoked to generate a basic noise matrix that matches the parameter dimension. The elements of the basic noise matrix follow a preset probability distribution, and the number of rows and columns of the matrix are strictly consistent with the dimension of the corresponding parameter. The noise intensity adjustment coefficient table is queried according to the permission level label to obtain the corresponding intensity coefficient. The noise intensity adjustment coefficient table contains the mapping relationship between permission level and intensity coefficient. The higher the permission level, the larger the intensity coefficient. The basic noise matrix is multiplied element-wise with the intensity coefficients to generate a noise perturbation factor that is dynamically adjusted based on the permission level. The dimension of the noise perturbation factor is consistent with that of the basic noise matrix, and the element values increase as the permission level increases. For each parameter in the local model parameter set with permission identifier, perform matrix addition with the corresponding noise perturbation factor to generate encrypted model parameters, and record the generation source information and intensity coefficient of the noise perturbation factor to obtain parameter units with encrypted trajectories. Sort all parameter units with encrypted trajectories according to permission level labels, add encrypted timestamps and checksums, and combine them into a permission encryption model parameter set.
7. The method according to claim 6, characterized in that, The method involves calling a noise generation source to generate a basic noise matrix that matches the parameter dimensions. The elements of the basic noise matrix follow a preset probability distribution, including: Initialize the cryptographically secure pseudo-random number generator by inputting the initial seed value corresponding to the permission level label into the generator and setting the random number sequence generation mode to continuous and non-repeating mode to ensure that each call generates an independent random number sequence. Read the number of rows and columns of the current parameter from the permission-dimension lookup table, and calculate the total number of random numbers to be generated, where the total number is the product of the number of rows and columns; Start the pseudo-random number generator to generate a specified number of random numbers in row priority order. Each random number is converted into a value that follows a preset probability distribution through a probability distribution transformation algorithm. During the transformation process, a preset transformation method is used to map uniformly distributed random numbers to target probability distribution random numbers. The generated probability distribution values are filled into a matrix of a preset dimension in row-major order to generate a basic noise matrix. The row and column elements of the matrix are arranged in the same order as the dimensions of the corresponding parameters. The randomness of the basic noise matrix is tested by calculating the statistical characteristic values of all elements of the matrix. If the statistical characteristic values meet the characteristic requirements of the preset probability distribution, the matrix is determined to meet the randomness requirements; otherwise, a random number sequence is regenerated and the matrix is constructed. The basic noise matrix that passes the test is stored in association with the parameter name, and the generation time and state parameters of the random number generator are recorded.
8. The method according to claim 7, characterized in that, The randomness test of the basic noise matrix involves calculating the statistical eigenvalues of all elements of the matrix. If the statistical eigenvalues meet the characteristics of a preset probability distribution, the matrix is determined to satisfy the randomness requirement; otherwise, a random number sequence is regenerated and the matrix is reconstructed. This includes: The basic noise matrix is converted into a one-dimensional array, and the total number of elements in the array is counted. The total number is the product of the number of rows and columns of the matrix. Calculate the arithmetic mean of all elements in the one-dimensional array as the overall mean of the basic noise matrix. The overall mean reflects whether the central trend of the noise sequence conforms to a preset probability distribution. The sum of squares of the differences between each element in the one-dimensional array and the overall mean is calculated to obtain the sum of squares of deviations from the mean, which is used to measure the degree to which an element deviates from the central trend. The sample variance is obtained by dividing the sum of squared deviations from the mean by the total number of samples minus one. The sample variance reflects whether the dispersion of the noise sequence meets the requirements of the preset probability distribution. Read the preset mean tolerance threshold and variance tolerance threshold from the system security configuration file. The mean tolerance threshold is the maximum range that the overall mean is allowed to deviate from the center value of the preset probability distribution, and the variance tolerance threshold is the maximum range that the sample variance is allowed to deviate from the discrete value of the preset probability distribution. Calculate the mean deviation and variance deviation. The mean deviation is the absolute value of the difference between the overall mean and the center value of the preset probability distribution, and the variance deviation is the absolute value of the difference between the sample variance and the discrete value of the preset probability distribution. If the mean deviation is less than the mean tolerance threshold and the variance deviation is less than the variance tolerance threshold, then the basic noise matrix is determined to meet the randomness requirement. If the conditions are not met, the initial seed value of the pseudo-random number generator is updated. The new seed value is a combination of the original seed value and the number of times the current matrix generation has failed. The step of generating the basic noise matrix is then repeated until the randomness requirement is met.
9. A computer device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Personalized differential privacy federal learning method and system
CN117556459A
Cultivated land quality monitoring method
CN120891177A