Federal learning subset adaptive isolation method and system supporting multiple tenants
By separating private and ordinary data in multi-tenant federated learning, constructing independent models and performing real-time accuracy analysis, the problems of data privacy protection and model accuracy are solved, and data isolation and training efficiency are improved.
Patent Information
- Application Number
- CN202510974263.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In multi-tenant federated learning scenarios, data privacy protection, heterogeneous data processing, and model training efficiency become key challenges. Existing technologies are difficult to adapt to the diversity of data types and needs in multi-tenant environments, and may lead to a decrease in model accuracy or an increase in computational overhead when processing privacy-sensitive data.
By establishing a database that maps tasks to data requirements, separating privacy-required data from general-required data, constructing independent models and performing real-time accuracy analysis, and dynamically adjusting the fictitious scale of the training data output model, data isolation and privacy protection can be achieved.
It effectively protects data privacy, adapts to the needs of multi-tenant heterogeneous data, improves model accuracy and training efficiency, and is suitable for multi-tenant federated learning scenarios.
Smart Images

Figure CN120851242A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of federated learning technology, and in particular to a method and system for adaptive isolation of federated learning subsets that supports multi-tenancy. Background Art
[0002] With the rapid development of cloud computing and edge computing, federated learning, as a distributed machine learning technology, is widely used in multi-device collaborative training scenarios, especially in multi-tenant environments, enabling model training without sharing raw data. However, in multi-tenant federated learning scenarios, data privacy protection, heterogeneous data processing, and model training efficiency have become key challenges. Traditional federated learning methods typically assume that all participants have similar data distributions and have relatively uniform methods for handling privacy data, making it difficult to adapt to the diversity of data types and requirements in multi-tenant environments. Furthermore, existing methods often protect privacy-sensitive data through simple data encryption or differential privacy techniques, but these methods may lead to decreased model accuracy or increased computational overhead. On the other hand, in multi-tenant scenarios, the dynamic changes in task requirements demand that the system be able to quickly match data resources and make adaptive adjustments, while existing technologies lack fine-grained isolation and dynamic optimization mechanisms for data subsets. Summary of the Invention
[0003] The purpose of this invention is to provide a method and system for adaptive isolation of subsets in federated learning that supports multi-tenancy, which can dynamically divide privacy data and ordinary data according to task requirements and build independent models to protect privacy.
[0004] This invention discloses a multi-tenant federated learning subset adaptive isolation method, comprising: Step S100: Establish a federated learning training task, use the federated learning training task as a retrieval condition, find the corresponding data requirement information in the preset task-data requirement correspondence database, and determine the required data that the edge server needs to provide based on the data requirement information. Based on privacy requirements, divide the required data into privacy requirement data and ordinary requirement data. Step S200: Independent model construction is performed on the privacy requirement data to obtain the training data output model. The cloud server trains the model based on the ordinary requirement data and the simulated requirement data output by the training parameter output model to build a federated training model. Step S300: Perform real-time accuracy analysis on the output of the federated training model to determine the accuracy level. Based on the accuracy level, adjust the fictitious scale of the training data output model.
[0005] In some embodiments disclosed in this invention, the method for constructing a task-data requirement mapping database includes: Step S101: Construct several standard tasks. Each standard task corresponds to several standard task keywords and requirement data information. The requirement data information includes the type of requirement data, the volume of requirement data, and the requirement data template. The requirement data template contains several data factor windows. Each data factor window is used to input data factors. Data factors include text information, numerical information, and image information. Step S102: The combination of standard task keywords is identified as the search tag, and the corresponding search tag records include the explanation of the training task content and the training objective.
[0006] In some embodiments disclosed in this invention, the method for independently constructing models from privacy-required data to obtain training data output models includes: Step S201: Determine several types of variable parameters in the privacy requirement data, construct a variable parameter performance model for each variable parameter, analyze the change characteristics of each variable parameter performance model, obtain several change characteristic change factors, and combine the change characteristic factors to generate a change characteristic factor group. Step S202: Determine the variation range of each variation characteristic factor in the variation characteristic factor group, and based on the variation range, perform several dynamic adjustments on the variation characteristic factor group to obtain several adjusted variation characteristic factor groups. Step S203: Based on the adjusted change feature factor group, adjust the change parameter performance model to obtain the adjusted change parameter performance model, and identify the adjusted change parameter performance model as the simulated demand data.
[0007] In some embodiments disclosed in this invention, the method for determining the variation fluctuation range of each variation characteristic factor in the variation characteristic factor group includes: Step S2021: Obtain historical demand data of the same type as privacy demand data and demand data constructed by expert experience, and uniformly record them as reference demand data. Determine several types of reference change parameters for the reference demand data, construct a reference change parameter performance model for each reference change parameter, and determine the reference change feature factor group. Step S2022: Classify the change feature factor groups corresponding to the privacy requirement data into equivalent categories to form several change feature factor sets, and classify different reference change feature factor groups into different change feature factor sets. Step S2023: Determine the range of variation of each variable characteristic factor in the set of variable characteristic factors, and associate the range of variation of each variable characteristic factor with the group of variable characteristic factors in the set of variable characteristic factors.
[0008] In some embodiments disclosed in this invention, the method for equivalent classification of change feature factor groups includes: Step S20221: Set an equal weight coefficient for each variable characteristic factor in the variable characteristic factor group, and set several factor difference intervals for each characteristic factor, with equal parameters set for each factor difference interval. Step S20222: Calculate the factor difference between the corresponding change feature factors, determine the factor difference interval to which the factor difference belongs, determine the equivalent parameters between the change feature factors, determine the degree of equivalence between change feature factor groups based on the equivalent parameters corresponding to all change feature factors, and determine whether change feature factor groups belong to the same category based on the degree of equivalence. The expression for calculating the degree of equality is: ; Where D represents the same degree, d i K is the equivalent parameter corresponding to the i-th characteristic factor of change. i is the equivalent weight coefficient corresponding to the i-th variable characteristic factor, n is the total number of variable characteristic factors in the variable characteristic factor group, L is the equivalent parameter influence adjustment coefficient, and b is the equivalent parameter influence adjustment constant.
[0009] In some embodiments disclosed in this invention, the method for classifying different reference variation feature factor groups into different variation feature factor sets includes: Step S20223: Average the variable characteristic factor group in the variable characteristic factor set to obtain the average variable characteristic factor group. Compare the reference variable characteristic factor group and the average variable characteristic factor group for equivalence, and determine whether to include the reference variable characteristic factor group in the variable characteristic factor set based on the degree of equivalence.
[0010] In some embodiments disclosed in this invention, the method for performing real-time accuracy analysis on the output of the federated training model to determine the degree of accuracy includes: Step 301: Determine the change feature factor group of the simulated demand data output by the federated training model, and denote it as the first change feature factor group. Denote the set of several first change feature factor groups as the first change feature factor set. Determine the change feature factor group of the privacy demand data, and denote it as the second change feature factor group. Denote the set of several second change feature factor groups as the second change feature factor set. Step S302: Determine the change fluctuation characteristics of the first set of change feature factors, denoted as the first change fluctuation characteristic; determine the change fluctuation characteristics of the second set of change feature factors, denoted as the second change fluctuation characteristic; compare the first change fluctuation characteristic and the second change fluctuation characteristic to determine the degree of agreement between them, and thus determine the accuracy.
[0011] In some embodiments disclosed in this invention, the method for determining the matching of the first change fluctuation feature and the second change fluctuation feature includes: Step S3021: Determine the characteristic factor change sequence of each change characteristic factor in the first change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the first change fluctuation feature as the first characteristic factor change sequence group; determine the characteristic factor change sequence of each change characteristic factor in the second change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the second change fluctuation feature as the second characteristic factor change sequence group. Step S3022: Compare the first feature factor change sequence group and the second feature factor change sequence group to determine the matching ratio of the matching feature factors among all feature factors in each relative feature factor change sequence. The average of the matching ratios corresponding to all feature factor change sequences is taken as the accuracy of the output of the federated training model.
[0012] In some embodiments disclosed in this invention, the method for adjusting the fictitious scale of the output model of the training data output model based on the degree of accuracy includes: Step S303: Based on the accuracy, the range of variation fluctuations of the variable feature factors when constructing the training data model is narrowed. The ratio of the narrowed range of variation fluctuations to the original range of variation fluctuations is considered a fictitious scale. Step S304: Based on the fictional scale, determine the range of variation fluctuations of the variation feature factors after narrowing, and use this as a benchmark to filter the data output by the training data output model.
[0013] Some embodiments of the present invention also disclose a multi-tenant federated learning subset adaptive isolation system, comprising: The first module is used to establish federated learning training tasks. The federated learning training tasks are used as search conditions to find the corresponding data requirement information in the preset task-data requirement correspondence database. Based on the data requirement information, the required data that the edge server needs to provide is determined. Based on privacy requirements, the required data is divided into privacy requirement data and ordinary requirement data. The second module is used to independently build models for privacy requirement data to obtain training data output models. The cloud server trains the model based on ordinary requirement data and simulated requirement data output by the training parameter output model to build a federated training model. The third module is used to perform real-time accuracy analysis on the output of the federated training model, determine the level of accuracy, and adjust the fictitious scale of the training data output model based on the level of accuracy.
[0014] This invention discloses a method and system for adaptive subset isolation in multi-tenant federated learning, relating to the field of federated learning technology. The method includes establishing a federated learning training task; determining the required data for edge servers based on a pre-defined task-data requirement correspondence database, and dividing it into privacy-requirement data and ordinary requirement data according to privacy requirements; constructing an independent model for the privacy-requirement data to generate a training data output model; training the federated training model on a cloud server using ordinary requirement data and simulated requirement data; and analyzing the accuracy of the federated training model output in real time, dynamically adjusting the fictitious scale of the training data output model based on the accuracy. This invention effectively protects data privacy, adapts to the heterogeneous data requirements of multi-tenant systems, and improves model accuracy and training efficiency through data subset isolation, independent model construction, and dynamic adjustment mechanisms, making it suitable for multi-tenant federated learning scenarios.
[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the steps of a federated learning subset adaptive isolation method that supports multi-tenancy, as disclosed in an embodiment of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific embodiments. It should be understood that the preferred embodiments described herein are only for illustration and explanation of the present invention and should not be construed as limiting the scope of protection of the present invention. Those skilled in the art can make some non-essential improvements and adjustments based on the following content of the present invention. In the present invention, unless otherwise expressly specified and limited, the technical terms used in the present invention should have the ordinary meaning understood by those skilled in the art.
[0019] Example:
[0020] This invention discloses a multi-tenant federated learning subset adaptive isolation method, see reference. Figure 1 ,include: Step S100: Establish a federated learning training task, use the federated learning training task as a retrieval condition, find the corresponding data requirement information in the preset task-data requirement correspondence database, and determine the required data that the edge server needs to provide based on the data requirement information. Based on privacy requirements, divide the required data into privacy requirement data and ordinary requirement data.
[0021] Step S100 works by establishing a federated learning training task and using a pre-defined task-data requirement mapping database to accurately match the data resources needed for the task. Simultaneously, the data is categorized according to privacy requirements to achieve data isolation and privacy protection in a multi-tenant environment. First, the system generates a training task based on task characteristics (e.g., image classification, text prediction) and uses this as a retrieval condition to search for corresponding data requirement information in the task-data requirement mapping database, such as data type (e.g., image, text), data volume, and format requirements. Based on this information, the system determines the data content that the edge server needs to provide. To meet privacy protection requirements, data is divided into privacy-required data (e.g., user personal information, medical records) and general-required data (e.g., publicly available weather data). For example, in a medical federated learning task, hospital patient medical records are labeled as privacy-required data, while publicly available medical statistics are general-required data. This categorization ensures that privacy-sensitive data receives special protection in subsequent processing, while general-required data can be directly used for cloud training, thereby achieving efficient data utilization and privacy security in a multi-tenant environment.
[0022] Step S200: Independent model construction is performed on the privacy requirement data to obtain the training data output model. The cloud server trains the model based on the ordinary requirement data and the simulated requirement data output by the training parameter output model to build a federated training model.
[0023] The principle behind step S200 is to build an independent model for privacy-required data to generate simulated data to protect privacy, while simultaneously combining it with ordinary requirement data for federated training in the cloud to construct an efficient federated training model. Specifically, the edge server builds an independent model for privacy-required data (such as user behavior data), and generates a training data output model by analyzing data change characteristics (such as time series fluctuations and image edge features). This model outputs simulated requirement data instead of the original data. For example, in a multi-tenant recommendation system task, the edge device might build a local model based on the user's historical purchase records (privacy data), outputting simulated purchase preference data. The cloud server then uses ordinary requirement data (such as product catalogs) and these simulated requirement data for joint training to generate a federated training model. This method effectively avoids the leakage of original privacy data by processing privacy data at the edge and only uploading simulated data, while ensuring that the cloud can optimize the model based on diverse data subsets, adapting to the data heterogeneity of multi-tenants.
[0024] Step S300: Perform real-time accuracy analysis on the output of the federated training model to determine the accuracy level. Based on the accuracy level, adjust the fictitious scale of the training data output model.
[0025] Step S300 works by dynamically adjusting the fictitious scale of the training data output model based on real-time analysis of the model's output accuracy, thereby optimizing model performance and adapting to task requirements. The system first performs an accuracy analysis on the federated training model's output (i.e., prediction results based on simulated demand data and ordinary demand data), comparing its variation characteristics with real demand data (such as the fluctuation range of feature factors and the proportion of agreement) to determine its accuracy. For example, in a traffic flow prediction task, if the model's predicted traffic pattern has a low degree of agreement with actual traffic data, its accuracy will be deemed insufficient. Based on this, the system adjusts the fictitious scale of the training data output model, i.e., by limiting the fluctuation range of variable feature factors in the simulated data (such as reducing the fluctuation amplitude of traffic peaks), making the output data closer to the real distribution. The adjusted model regenerates simulated data and participates in the next round of federated training, thereby improving the model's accuracy and stability. This dynamic adjustment mechanism ensures that the federated learning system can continuously optimize in a multi-tenant environment and adapt to the accuracy requirements of different tasks.
[0026] In some embodiments disclosed in this invention, the method for constructing a task-data requirement mapping database includes: Step S101: Construct several standard tasks. Each standard task corresponds to several standard task keywords and requirement data information. The requirement data information includes the type of requirement data, the volume of requirement data, and the requirement data template. The requirement data template contains several data factor windows. Each data factor window is used to input data factors, which include text information, numerical information, and image information.
[0027] The principle of step S101 is to establish a structured task-data requirement mapping database by constructing standard tasks and their corresponding data requirements, providing a foundation for subsequent task retrieval and data matching. The system first defines several standard tasks (such as image classification and speech recognition), each task associated with a set of standard task keywords (such as "image, classification, object detection") and detailed data requirements. This information includes the required data type (such as image, audio), data volume (such as 1000 images), and a required data template. The template includes multiple data factor windows to standardize the input data factors, such as text information (label description), numerical information (numerical features), or image information (pixel data). For example, in a face recognition task, the template might require an input factor window containing a face image (image information), age (numerical information), and identity label (text information). This structured task definition and template design ensures the standardization and retrieval of data requirements, providing a unified framework for data matching across different edge servers in multi-tenant federated learning, while also supporting flexible processing of heterogeneous data.
[0028] Step S102: The combination of standard task keywords is identified as the search tag, and the corresponding search tag records include the explanation of the training task content and the training objective.
[0029] The principle of step S102 is to transform standard task keyword combinations into search tags and associate them with detailed task explanations and objectives to achieve rapid mapping between tasks and data requirements.
[0030] In some embodiments disclosed in this invention, the method for independently constructing models from privacy-required data to obtain training data output models includes: Step S201: Determine several types of variable parameters in the privacy requirement data, construct a variable parameter performance model for each variable parameter, analyze the change characteristics of each variable parameter performance model, and obtain several change characteristic change factors (such as the curvature and amplitude of the change curve at a time node, or the block size and edge shape of a certain block on an image). Combine the change characteristic factors to generate a change characteristic factor group.
[0031] The principle of step S201 is to analyze the changing parameters of the privacy-required data, construct independent models representing these changing parameters, and extract key change feature factors to generate feature combinations that can be used for generating simulated data. First, the system identifies the changing parameters in the privacy-required data (such as flow rates in a time series or pixel distribution in an image) and constructs a model for each parameter to describe its dynamic characteristics. Next, by analyzing the changing features of the models (such as the curvature and amplitude of time series curves, or the edge shape and size of image blocks), several change feature factors are extracted. For example, in a medical federated learning task, patient blood pressure data (privacy-required data) might be analyzed as a time series, and the extracted feature factors include the curvature and peak amplitude of blood pressure fluctuations. These factors are combined into a set of change feature factors, representing the core dynamic characteristics of the data. This method abstracts the key information of the privacy data through feature extraction and combination, providing a foundation for subsequent generation of simulated data, while avoiding the direct use of raw privacy data and ensuring privacy protection.
[0032] Step S202: Determine the variation range of each variation characteristic factor in the variation characteristic factor group, and based on the variation range, perform several dynamic adjustments on the variation characteristic factor group to obtain several adjusted variation characteristic factor groups.
[0033] The principle of step S202 is to quantify the fluctuation range of each feature factor in the changing feature factor group and generate diverse feature groups through dynamic adjustment, providing flexibility for the generation of simulation data. The system first determines the fluctuation range of each changing feature factor; for example, the range of the curvature factor in a time series might be [0.1, 0.5], and the range of the image edge shape factor might be an interval of a specific angle. Based on these ranges, the system dynamically adjusts the factor group multiple times, generating several adjusted changing feature factor groups. For example, in a multi-tenant traffic prediction task, the system might generate multiple different feature groups by adjusting the amplitude range of the traffic peak factor to simulate traffic patterns under different scenarios. This adjustment process increases the diversity of simulation data by controlling the randomization or regularization of the fluctuation range, while preserving the statistical characteristics of the original data, providing rich input for subsequent model adjustment and federated training.
[0034] Step S203: Based on the adjusted change feature factor group, adjust the change parameter performance model to obtain the adjusted change parameter performance model, and identify the adjusted change parameter performance model as the simulated demand data.
[0035] The principle of step S203 is to optimize the variable parameter performance model using the adjusted set of variable feature factors to generate simulated demand data to replace the original privacy data in federated training. Based on the adjusted set of feature factors (such as the adjusted curvature and amplitude combination), the system updates the parameters or optimizes the structure of the original variable parameter performance model to generate the adjusted performance model.
[0036] In some embodiments disclosed in this invention, the method for determining the variation fluctuation range of each variation characteristic factor in the variation characteristic factor group includes: Step S2021: Obtain historical demand data of the same type as privacy demand data and demand data constructed by expert experience, and uniformly refer to them as reference demand data. Determine several types of reference change parameters for the reference demand data, construct a reference change parameter performance model for each reference change parameter, and determine the reference change feature factor group.
[0037] The principle of step S2021 is to construct reference requirement data by collecting historical data and expert experience data of the same type as the privacy requirement data, and to analyze its changing parameters and characteristic factors to provide a benchmark for determining the fluctuation range of the changing characteristic factors. The system first acquires historical requirement data (such as past user behavior records) and expert-constructed experience data (such as typical behavior patterns), collectively referred to as reference requirement data. Next, it identifies the reference changing parameters (such as click frequency) in this data, constructs a reference changing parameter performance model for each parameter, analyzes its changing characteristics (such as the amplitude of frequency fluctuations), and extracts a group of reference changing characteristic factors.
[0038] Step S2022: Classify the change feature factor groups corresponding to the privacy requirement data into equivalent categories to form several change feature factor sets, and classify different reference change feature factor groups into different change feature factor sets.
[0039] The principle of step S2022 is to classify the change feature factor groups of privacy requirement data with reference change feature factor groups through equivalence classification, forming change feature factor sets for subsequent determination of fluctuation range. The system first performs equivalence analysis on the change feature factor groups of privacy requirement data (such as the curvature of time series and the edge angle of images), classifying them into several change feature factor sets based on the similarity between factors (such as numerical distribution and change patterns). Simultaneously, the reference change feature factor groups are also assigned to corresponding factor sets.
[0040] Step S2023: Determine the range of variation of each variable characteristic factor in the set of variable characteristic factors, and associate the range of variation of each variable characteristic factor with the group of variable characteristic factors in the set of variable characteristic factors.
[0041] The principle of step S2023 is to determine the fluctuation range of each variable characteristic factor in the set of variable characteristic factors and establish a correlation with the factor groups in the factor set to support the generation of simulation data and model optimization. The system analyzes the characteristic factors in each factor set and calculates their fluctuation range (e.g., the range of the curvature factor [0.1, 0.5]), typically based on statistical methods or the distribution characteristics of reference data. Then, these fluctuation ranges are correlated with the variable characteristic factor groups in the factor set to ensure that the characteristic factors in each factor group have clear range constraints.
[0042] In some embodiments disclosed in this invention, the method for equivalent classification of change feature factor groups includes: Step S20221: Each variable characteristic factor in the variable characteristic factor group is assigned an equal weight coefficient, and several factor difference intervals are set for each characteristic factor, with equal parameters set for each factor difference interval.
[0043] The principle of step S20221 is to establish a quantitative comparison standard by assigning equal weight coefficients and factor difference intervals to each feature factor in the variable feature factor group, providing a basis for subsequent equivalence classification. The system first sets an equal weight coefficient for each variable feature factor (such as the curvature of a time series or the edge angle of an image) to reflect its relative importance in the factor group. For example, the curvature factor may be given a higher weight due to its greater influence on data patterns. Next, several factor difference intervals (such as curvature difference values [0, 0.1], [0.1, 0.2]) are defined for each factor, with each interval associated with an equal parameter to quantify the similarity between factors.
[0044] Step S20222: Calculate the factor difference between the corresponding change feature factors, determine the factor difference interval to which the factor difference belongs, determine the equivalent parameters between the change feature factors, determine the degree of equivalence between change feature factor groups based on the equivalent parameters corresponding to all change feature factors, and determine whether change feature factor groups belong to the same category based on the degree of equivalence.
[0045] The expression for calculating the degree of equality is: ; Where D represents the same degree, d i K is the equivalent parameter corresponding to the i-th characteristic factor of change. i is the equivalent weight coefficient corresponding to the i-th variable characteristic factor, n is the total number of variable characteristic factors in the variable characteristic factor group, L is the equivalent parameter influence adjustment coefficient, and b is the equivalent parameter influence adjustment constant.
[0046] D: Equivalence, representing the similarity between two groups of changing characteristic factors, output in exponential form.
[0047] exp: Exponential function (base e) is used to convert weighted sums to a non-linear scale, amplifying or smoothing the results.
[0048] L: Equivalent parameter influence adjustment coefficient, which adjusts the overall influence of the weighted sum.
[0049] K i : The weight coefficient of the i-th characteristic factor of change, reflecting its importance.
[0050] d i : The difference or similarity index of the i-th characteristic factor of change.
[0051] b: Equivalent parameter influence adjustment constant, used to correct the reference value.
[0052] n: The total number of variable characteristic factors in the variable characteristic factor group.
[0053] In some embodiments disclosed in this invention, the method for classifying different reference variation feature factor groups into different variation feature factor sets includes: Step S20223: Average the variable characteristic factor group in the variable characteristic factor set to obtain the average variable characteristic factor group. Compare the reference variable characteristic factor group and the average variable characteristic factor group for equivalence, and determine whether to include the reference variable characteristic factor group in the variable characteristic factor set based on the degree of equivalence.
[0054] In some embodiments disclosed in this invention, the method for performing real-time accuracy analysis on the output of the federated training model to determine the degree of accuracy includes: Step 301: Determine the change feature factor group of the simulated demand data output by the federated training model, and denote it as the first change feature factor group. Denote the set of several first change feature factor groups as the first change feature factor set. Determine the change feature factor group of the privacy demand data, and denote it as the second change feature factor group. Denote the set of several second change feature factor groups as the second change feature factor set.
[0055] Step S302: Determine the change fluctuation characteristics of the first set of change feature factors, denoted as the first change fluctuation characteristic; determine the change fluctuation characteristics of the second set of change feature factors, denoted as the second change fluctuation characteristic; compare the first change fluctuation characteristic and the second change fluctuation characteristic to determine the degree of agreement between them, and thus determine the accuracy.
[0056] In some embodiments disclosed in this invention, the method for determining the matching of the first change fluctuation feature and the second change fluctuation feature includes: Step S3021: Determine the characteristic factor change sequence of each characteristic factor in the first change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the first change fluctuation feature as the first characteristic factor change sequence group; determine the characteristic factor change sequence of each characteristic factor in the second change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the second change fluctuation feature as the second characteristic factor change sequence group.
[0057] Step S3022: Compare the first feature factor change sequence group and the second feature factor change sequence group to determine the matching ratio of the matching feature factors among all feature factors in each relative feature factor change sequence. The average of the matching ratios corresponding to all feature factor change sequences is taken as the accuracy of the output of the federated training model.
[0058] In some embodiments disclosed in this invention, the method for adjusting the fictitious scale of the output model of the training data output model based on the degree of accuracy includes: Step S303: Based on the accuracy, the range of variation fluctuations of the variable feature factors when constructing the training data model is narrowed. The ratio of the narrowed range of variation fluctuations to the original range of variation fluctuations is considered a fictitious scale.
[0059] Step S304: Based on the fictional scale, determine the range of variation fluctuations of the variation feature factors after narrowing, and use this as a benchmark to filter the data output by the training data output model.
[0060] Some embodiments of the present invention also disclose a multi-tenant federated learning subset adaptive isolation system, comprising: The first module is used to establish federated learning training tasks. It uses the federated learning training tasks as search conditions to find the corresponding data requirement information in the preset task-data requirement correspondence database. Based on the data requirement information, it determines the required data that the edge server needs to provide. Based on privacy requirements, the required data is divided into privacy requirement data and ordinary requirement data.
[0061] The second module is used to independently build models for privacy-required data to obtain training data output models. The cloud server trains the model based on ordinary requirement data and simulated requirement data output by the training parameter output model to build a federated training model.
[0062] The third module is used to perform real-time accuracy analysis on the output of the federated training model, determine the level of accuracy, and adjust the fictitious scale of the training data output model based on the level of accuracy.
[0063] This invention discloses a method and system for adaptive subset isolation in multi-tenant federated learning, relating to the field of federated learning technology. The method includes establishing a federated learning training task; determining the required data for edge servers based on a pre-defined task-data requirement correspondence database, and dividing it into privacy-requirement data and ordinary requirement data according to privacy requirements; constructing an independent model for the privacy-requirement data to generate a training data output model; training the federated training model on a cloud server using ordinary requirement data and simulated requirement data; and analyzing the accuracy of the federated training model output in real time, dynamically adjusting the fictitious scale of the training data output model based on the accuracy. This invention effectively protects data privacy, adapts to the heterogeneous data requirements of multi-tenant systems, and improves model accuracy and training efficiency through data subset isolation, independent model construction, and dynamic adjustment mechanisms, making it suitable for multi-tenant federated learning scenarios.
[0064] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented in hardware or by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0065] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A subset adaptive isolation method for federated learning that supports multi-tenancy, characterized in that, include: Step S100: Establish a federated learning training task, use the federated learning training task as a retrieval condition, find the corresponding data requirement information in the preset task-data requirement correspondence database, and determine the required data that the edge server needs to provide based on the data requirement information. Based on privacy requirements, divide the required data into privacy requirement data and ordinary requirement data. Step S200: Independent model construction is performed on the privacy requirement data to obtain the training data output model. The cloud server trains the model based on the ordinary requirement data and the simulated requirement data output by the training parameter output model to build a federated training model. Step S300: Perform real-time accuracy analysis on the output of the federated training model to determine the accuracy level. Based on the accuracy level, adjust the fictitious scale of the training data output model.
2. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 1, characterized in that, Methods for constructing a task-data requirement mapping database include: Step S101: Construct several standard tasks. Each standard task corresponds to several standard task keywords and requirement data information. The requirement data information includes the type of requirement data, the volume of requirement data, and the requirement data template. The requirement data template contains several data factor windows. Each data factor window is used to input data factors. Data factors include text information, numerical information, and image information. Step S102: The combination of standard task keywords is identified as the search tag, and the corresponding search tag records include the explanation of the training task content and the training objective.
3. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 2, characterized in that, Methods for building independent models from privacy-required data to obtain training data output models include: Step S201: Determine several types of variable parameters in the privacy requirement data, construct a variable parameter performance model for each variable parameter, analyze the change characteristics of each variable parameter performance model, obtain several change characteristic change factors, and combine the change characteristic factors to generate a change characteristic factor group. Step S202: Determine the variation range of each variation characteristic factor in the variation characteristic factor group, and based on the variation range, perform several dynamic adjustments on the variation characteristic factor group to obtain several adjusted variation characteristic factor groups. Step S203: Based on the adjusted change feature factor group, adjust the change parameter performance model to obtain the adjusted change parameter performance model, and identify the adjusted change parameter performance model as the simulated demand data.
4. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 3, characterized in that, Methods for determining the range of variation fluctuations for each variation characteristic factor in a group of variation characteristic factors include: Step S2021: Obtain historical demand data of the same type as privacy demand data and demand data constructed by expert experience, and uniformly record them as reference demand data. Determine several types of reference change parameters for the reference demand data, construct a reference change parameter performance model for each reference change parameter, and determine the reference change feature factor group. Step S2022: Classify the change feature factor groups corresponding to the privacy requirement data into equivalent categories to form several change feature factor sets, and classify different reference change feature factor groups into different change feature factor sets. Step S2023: Determine the range of variation of each variable characteristic factor in the set of variable characteristic factors, and associate the range of variation of each variable characteristic factor with the group of variable characteristic factors in the set of variable characteristic factors.
5. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 4, characterized in that, Methods for classifying change characteristic factor groups into equivalent categories include: Step S20221: Set an equal weight coefficient for each variable characteristic factor in the variable characteristic factor group, and set several factor difference intervals for each characteristic factor, with equal parameters set for each factor difference interval. Step S20222: Calculate the factor difference between the corresponding change feature factors, determine the factor difference interval to which the factor difference belongs, determine the equivalent parameters between the change feature factors, determine the degree of equivalence between change feature factor groups based on the equivalent parameters corresponding to all change feature factors, and determine whether change feature factor groups belong to the same category based on the degree of equivalence. The expression for calculating the degree of equality is: ; Where D represents the same degree, d i K is the equivalent parameter corresponding to the i-th characteristic factor of change. i is the equivalent weight coefficient corresponding to the i-th variable characteristic factor, n is the total number of variable characteristic factors in the variable characteristic factor group, L is the equivalent parameter influence adjustment coefficient, and b is the equivalent parameter influence adjustment constant.
6. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 5, characterized in that, Methods for classifying different sets of reference variable feature factors into different sets of variable feature factors include: Step S20223: Average the variable characteristic factor group in the variable characteristic factor set to obtain the average variable characteristic factor group. Compare the reference variable characteristic factor group and the average variable characteristic factor group for equivalence, and determine whether to include the reference variable characteristic factor group in the variable characteristic factor set based on the degree of equivalence.
7. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 4, characterized in that, Methods for performing real-time accuracy analysis on the output of federated training models to determine the degree of accuracy include: Step 301: Determine the change feature factor group of the simulated demand data output by the federated training model, and denote it as the first change feature factor group. Denote the set of several first change feature factor groups as the first change feature factor set. Determine the change feature factor group of the privacy demand data, and denote it as the second change feature factor group. Denote the set of several second change feature factor groups as the second change feature factor set. Step S302: Determine the change fluctuation characteristics of the first set of change feature factors, denoted as the first change fluctuation characteristic; determine the change fluctuation characteristics of the second set of change feature factors, denoted as the second change fluctuation characteristic; compare the first change fluctuation characteristic and the second change fluctuation characteristic to determine the degree of agreement between them, and thus determine the accuracy.
8. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 7, characterized in that, Methods for determining the agreement between the first and second fluctuation characteristics include: Step S3021: Determine the characteristic factor change sequence of each change characteristic factor in the first change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the first change fluctuation feature as the first characteristic factor change sequence group; determine the characteristic factor change sequence of each change characteristic factor in the second change fluctuation feature, and denote the set of all characteristic factor change sequences corresponding to the second change fluctuation feature as the second characteristic factor change sequence group. Step S3022: Compare the first feature factor change sequence group and the second feature factor change sequence group to determine the matching ratio of the matching feature factors among all feature factors in each relative feature factor change sequence. The average of the matching ratios corresponding to all feature factor change sequences is taken as the accuracy of the output of the federated training model.
9. The method for adaptive isolation of subsets in federated learning supporting multi-tenancy as described in claim 4, characterized in that, Methods for adjusting the fictitious scale of the training data output model based on accuracy include: Step S303: Based on the accuracy, the range of variation fluctuations of the variable feature factors when constructing the training data model is narrowed. The ratio of the narrowed range of variation fluctuations to the original range of variation fluctuations is considered a fictitious scale. Step S304: Based on the fictional scale, determine the range of variation fluctuations of the variation feature factors after narrowing, and use this as a benchmark to filter the data output by the training data output model.
10. A federated learning subset adaptive isolation system supporting multi-tenancy, characterized in that, The federated learning subset adaptive isolation method for implementing any one of claims 1-9 includes: The first module is used to establish federated learning training tasks. The federated learning training tasks are used as search conditions to find the corresponding data requirement information in the preset task-data requirement correspondence database. Based on the data requirement information, the required data that the edge server needs to provide is determined. Based on privacy requirements, the required data is divided into privacy requirement data and ordinary requirement data. The second module is used to independently build models for privacy requirement data to obtain training data output models. The cloud server trains the model based on ordinary requirement data and simulated requirement data output by the training parameter output model to build a federated training model. The third module is used to perform real-time accuracy analysis on the output of the federated training model, determine the level of accuracy, and adjust the fictitious scale of the training data output model based on the level of accuracy.