Federal large model learning method based on confrontation feature aggregation and multi-party computing system
By using a federated large model learning method based on adversarial feature aggregation, adversarial features with small storage space and containing noise are generated and transmitted, solving the communication efficiency and privacy protection problems of federated large models, and achieving efficient model training and privacy protection.
Patent Information
- Application Number
- CN202511119764.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-11-25
AI Technical Summary
Existing technologies cannot effectively address the challenges of communication efficiency and privacy protection in federated large models, especially when facing complex adversarial attacks, as they pose risks of high communication overhead and privacy breaches.
A federated large model learning method based on adversarial feature aggregation is adopted. Adversarial features are generated through a noise generator and uploaded to the central service provider for model training, which reduces communication overhead and improves privacy protection.
It improves the communication efficiency and privacy protection capabilities of the federated large model, reduces the risk of data leakage, and maintains the effectiveness of model training.
Smart Images

Figure CN121009952A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to a federated large model learning method and device based on adversarial feature aggregation. BACKGROUND
[0002] Currently, there are various types of attacks on federated large models, including text replacement, gradient attacks, prompt injection, etc. As these attack methods continue to evolve, defense becomes more difficult. And as the application of large models in various fields continues to expand, the risk of adversarial attacks also increases. For example, in the financial field, attackers may use large models to generate false transaction information.
[0003] Currently, there are various types of attacks on federated large models, including text replacement, gradient attacks, prompt injection, etc. As these attack methods continue to evolve, defense becomes more difficult. And as the application of large models in various fields continues to expand, the risk of adversarial attacks also increases. For example, in the financial field, attackers may use large models to generate false transaction information.
[0004] However, in terms of federated large model defense against adversarial attacks, existing technologies cannot well solve the two major problems of communication efficiency and privacy protection. SUMMARY
[0005] The present application provides a federated large model learning method based on adversarial feature aggregation and a multi-party computing system to solve the problems of low communication efficiency and privacy protection in federated large models.
[0006] In a first aspect, the present application provides a federated large model learning method based on adversarial feature aggregation, applied to a multi-party computing system, the multi-party computing system comprising a plurality of participants and a central service party, the plurality of participants being in communication connection with the central service party, the method comprising:
[0007] For any one of the plurality of participants, the participant, based on a noise generator, performs adversarial feature generation processing on local data to obtain adversarial features, the noise generator being used to generate noise data;
[0008] uploading the adversarial features to the central service party;
[0009] The central service party trains a model for a specific task through an adversarial feature cluster to obtain a target model, the adversarial feature cluster being obtained by aggregating the adversarial features issued by the plurality of participants.
[0010] Secondly, this application provides a multi-party computing system, which includes multiple participating parties and a central service provider. The multiple participating parties are all communicatively connected to the central service provider. Each participating party includes a processing module and an uploading module, and the central service provider includes a training module.
[0011] For any one of the multiple participants, the processing module of that participant is used to perform adversarial feature generation processing on local data based on a noise generator to obtain adversarial features, wherein the noise generator is used to generate noise data.
[0012] An upload module is used to upload the adversarial features to the central service provider;
[0013] The training module of the central service provider is used to train a model for a specific task through an adversarial feature cluster to obtain a target model. The adversarial feature cluster is obtained by aggregating adversarial features issued by the multiple participating parties.
[0014] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0015] The memory stores computer-executed instructions;
[0016] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0019] The federated large model learning method and multi-party computation system based on adversarial feature aggregation provided in this application involve participating parties generating adversarial features from their local data using a noise generator. These adversarial features are then uploaded to the central service provider. Due to the small storage space required for adversarial features and the addition of noisy data, the communication overhead between the participating parties and the server is reduced. The addition of adversarial noise also improves the privacy protection of local data. The central service provider trains a model for a specific task using the adversarial feature cluster to obtain the target model, thus solving the two major problems of communication efficiency and privacy protection in traditional federated large model algorithms. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] Figure 1 A schematic diagram of the architecture of a multi-party computation system for implementing a federated large model learning method based on adversarial feature aggregation, provided for embodiments of this application;
[0022] Figure 2 A flowchart illustrating the federated large model learning method based on adversarial feature aggregation provided in this application embodiment. Figure One ;
[0023] Figure 3 A flowchart illustrating a federated large model learning method based on adversarial feature aggregation provided in this application embodiment. Figure Two ;
[0024] Figure 4 A schematic diagram of the architecture of a federated large model learning method based on adversarial feature aggregation provided in an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the structure of a multi-party computing system provided in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0030] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0031] It should be noted that the federated large model learning method and multi-party computation system based on adversarial feature aggregation provided in this application can be used in the field of artificial intelligence, or in any field other than artificial intelligence. The application field of the federated large model learning method and multi-party computation system based on adversarial feature aggregation in this application is not limited.
[0032] Currently, attack methods against federated large models are becoming increasingly complex and diverse, including text substitution, gradient attacks, and hint injection. As these attack methods continue to evolve, defense becomes more difficult. Furthermore, the expanding applications of large models across various fields increase the risk of attacks. For example, in the financial sector, attackers might use large models to generate fake transaction information.
[0033] Currently, defense against adversarial attacks on large models involves several layers: the data layer (including data augmentation, data cleaning, and filtering), the model layer, and others. The model layer includes model distillation and model ensemble; other layers include security frameworks and tools, and red team testing. Regarding collaborative modeling of large models, federated learning frameworks such as FATE have introduced federated learning modules specifically for large models, supporting fine-tuning of large models and providing various pre-defined learning modes and model algorithms.
[0034] However, when it comes to defending against adversarial attacks on large-scale federated models, existing technologies cannot effectively address the two major issues of communication efficiency and privacy protection.
[0035] The federated large model learning method based on adversarial feature aggregation provided in this application replaces the traditional model parameter passing with adversarial feature passing, which has a smaller storage space and incorporates noisy data, aiming to solve the above-mentioned technical problems of the prior art.
[0036] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0037] Figure 1 This application provides an embodiment of an architecture diagram of a multi-party computation system for implementing a federated large model learning method based on adversarial feature aggregation; as shown below. Figure 1 As shown, the multi-party computing system 10 includes multiple participants 101 and a central service provider 102, and all participants 101 are communicatively connected to the central service provider 102.
[0038] Figure 2 A flowchart illustrating the federated large model learning method based on adversarial feature aggregation provided in this application embodiment. Figure One The method provided in this embodiment is applied to, for example, Figure 1 In the multi-party computation system shown, such as Figure 2 As shown, the method includes:
[0039] S201. For any one of the plurality of participants, the participant performs adversarial feature generation processing on the local data based on the noise generator to obtain adversarial features.
[0040] Among them, the participants refer to the data holders who own local data and participate in the federated model; the noise generator is used to generate noisy data, and the adversarial features are used to improve the privacy of local data. The adversarial features have a small storage space, for example, they can be a 4096-dimensional vector.
[0041] Federated learning, as a distributed machine learning method, enables multiple data holders to jointly train a model without directly sharing data. Federated large-scale models can be trained without data leaving the data domain, thus meeting the cross-institutional collaborative training needs in the field of large-scale models. However, existing federated large-scale model methods also face challenges such as privacy risks, low training efficiency, high overhead, and vulnerability to adversarial attacks.
[0042] Participants input their local data into a noise generator, which then performs adversarial feature generation on the local data. This generates adversarial features that improve the privacy of the local data while minimizing storage space; for example, these can be 4096-dimensional vectors. These small-storage adversarial features replace traditional model parameter passing, reducing communication overhead between participants and the server, and improving the model training efficiency of the central service provider.
[0043] In one possible implementation, the adversarial feature generation process based on a noise generator on local data is described in detail to obtain the adversarial features, including:
[0044] The local data is used to extract features using a feature extractor issued by the central service provider to obtain intermediate features; the intermediate features are then processed by a noise generator to generate adversarial features, resulting in noise data; the intermediate features and the noise data are then fused to obtain the adversarial features.
[0045] Among them, the feature extractor is obtained through the large model parameters uniformly distributed by the central service provider. This large model is fine-tuned by federated learning using data uploaded by multiple participants and can adapt to data from multiple parties; intermediate features are used to indicate vector feature data with small storage space; noisy data is adversarial data generated by the attack decoder. The attack decoder is used to simulate attackers. Understandably, attackers want to recover the original data through shared data.
[0046] The local data is used to extract features through the feature extractor issued by the central service provider to obtain intermediate features with small storage space. The intermediate features are then processed by the noise generator to generate adversarial features, resulting in adversarial data, i.e., noise data, which is generated against the attack decoder. The intermediate features and the noise data are then fused to generate adversarial features, thereby improving the privacy protection of the participants' local data.
[0047] By generating intermediate features and then generating adversarial features based on these intermediate features, communication efficiency and privacy protection of local data from multiple participants are improved.
[0048] In one possible implementation, the local data includes sample data and label data; the process of generating adversarial features from the intermediate features using a noise generator to obtain noisy data is described in detail, including:
[0049] The sample data and the label data are input into the noise generator. The loss is determined based on the loss function and the label data, and the gradient is determined based on the loss and the sample data. The gradient is then designed using a sign function, and the product of the designed gradient and the control parameter is used as the noise data. The noise data is then superimposed on the sample data to determine the adversarial features.
[0050] Among them, the control parameters are used to control the noise amplitude.
[0051] The sample data is input into the noise generator for prediction, yielding the generator's prediction output. The loss and gradient are then determined based on the predicted data, label data, and the loss function. The formulas are as follows:
[0052]
[0053] in, For loss, For gradient, Used to indicate the impact of each sample data on the loss.
[0054] The gradient is signified using a sign function, which is as follows:
[0055]
[0056] The product of the sign-reduced gradient and the control parameter is taken as the noise data, as shown in the following formula:
[0057]
[0058] in, This is noise data; These are control parameters.
[0059] The adversarial features are determined by overlaying the noise data with the sample data, as shown in the following formula:
[0060]
[0061] in, To counteract the characteristics.
[0062] By generating adversarial features, the privacy protection of local data of multiple participants is improved, avoiding consequences such as data leakage, privacy violations, and malicious code generation.
[0063] S202. Upload the adversarial features to the central service provider;
[0064] S203. The central service provider trains a model for a specific task through an adversarial feature cluster to obtain a target model. The adversarial feature cluster is obtained by aggregating the adversarial features issued by the multiple participating parties.
[0065] The central service provider uses an adversarial feature cluster obtained by aggregating adversarial features issued by multiple participants to train a model for a specific task to obtain a target model. The target model is trained in a supervised manner using adversarial features aggregated from multiple participants. The target model can be a classification model for a specific classification task or a prediction model for a specific prediction task. This application does not impose any restrictions on this.
[0066] The federated large model learning method based on adversarial feature aggregation provided in this application involves participating parties generating adversarial features from their local data using a noise generator. These adversarial features are then uploaded to the central service provider. Due to the small storage space required for adversarial features and the addition of noisy data, the communication overhead between the participating parties and the server is reduced. The addition of adversarial noise also improves the privacy protection of local data. The central service provider trains a model for a specific task using the adversarial feature cluster to obtain the target model, thus solving the two major problems of communication efficiency and privacy protection in traditional federated large model algorithms.
[0067] Figure 3 A flowchart illustrating a federated large model learning method based on adversarial feature aggregation provided in this application embodiment. Figure Two ,like Figure 3 As shown, in this embodiment... Figure 2 Based on the examples, a federated large model learning method based on adversarial feature aggregation is described in detail. This method also includes:
[0068] S301. The adversarial feature and the intermediate feature are respectively processed by the attack decoder to obtain the first restored data corresponding to the adversarial feature and the second restored data corresponding to the intermediate feature;
[0069] The attack decoder is used to simulate an attacker, who, understandably, wants to recover the original data through shared data.
[0070] To prevent attackers from retrieving participants' original private data from shared data, the privacy protection of local data is improved by adding noisy data. At the same time, privacy quantification is performed on the recovered data before and after adding noisy data to achieve a balance between security and usability in adversarial features.
[0071] Before receiving the shared data, participating parties do not need the adversarial features of other participants. Since the feature extractor for extracting intermediate features from local data is public, participating parties can use their own local data to train the attack decoder. Participating parties input the intermediate features and adversarial features into the attack decoder respectively, perform data recovery processing, and obtain the first recovered data corresponding to the adversarial features and the second recovered data corresponding to the intermediate features.
[0072] S302. Perform privacy measurement on the first recovered data and the second recovered data to obtain the privacy measurement result;
[0073] The privacy metric results are used to measure the validity of noisy data.
[0074] Privacy measures are applied to the first and second recovered data to obtain privacy measurement results, which measure the effectiveness of the noisy data. Privacy measures may include, for example, comparative analysis or quantitative privacy.
[0075] In one possible implementation, the privacy measurement results include KL divergence and similarity; a detailed explanation is provided regarding the privacy measurement results obtained by performing privacy measurements on the first recovered data and the second recovered data, including:
[0076] Determine the first probability distribution corresponding to the first recovered data and the second probability distribution corresponding to the second recovered data, and determine the KL divergence based on the first probability distribution and the second probability distribution;
[0077] Among them, KL divergence is used to indicate the degree of deviation in recovering local data and to explain the extent of privacy data leakage.
[0078] Based on the first and second recovered data, the sample space is determined. Based on the sample space, the first probability distribution of the random variables in the first recovered data and the second probability distribution of the random variables in the second recovered data are determined respectively, and the KL divergence is determined using the following formula.
[0079]
[0080] in, Let KL divergence be the KL divergence. It follows the first probability distribution; Second probability distribution; The number of random variables.
[0081] By minimizing the KL divergence between the real data distribution and the privacy-preserving data distribution, it is ensured that the privacy-preserving output does not deviate excessively from the real distribution of the original data, thus maintaining the usefulness of the data while protecting privacy.
[0082] Determine the similarity evaluation vectors corresponding to the first recovered data and the second recovered data, and determine the similarity between the local data and the recovered local data based on the similarity evaluation vectors.
[0083] Similarity is used to indicate the risk of privacy leakage; for example, similarity can be cosine similarity. The similarity evaluation vector indicates the similarity between two vectors in one direction. For example, in text analysis, a document can be represented as a term frequency-inverse document frequency (TF-IDF) weighted vector, and cosine similarity is used to measure the similarity between these document vectors. If the similarity value is close to 1, it indicates a highly consistent distribution and a high risk of privacy leakage; conversely, if the similarity value is close to -1 or 0, it indicates a significant difference in distribution and a low risk of privacy leakage.
[0084] A similarity evaluation vector is determined for evaluating the first and second recovered data in the same direction, and the similarity between the local data and the recovered local data is determined using the following formula:
[0085]
[0086] in, Cosine similarity; This is the first similarity evaluation vector corresponding to the first recovered data; This is the second similarity evaluation vector corresponding to the second recovered data; This represents the number of similarity evaluation vectors.
[0087] By performing privacy quantification on the recovered data before and after adding noise, a balance can be struck between security and usability for adversarial features.
[0088] In one possible implementation, before performing data recovery processing on the adversarial feature and the intermediate feature respectively through the attack decoder to obtain the first recovered data corresponding to the adversarial feature and the second recovered data corresponding to the intermediate feature, the method further includes:
[0089] The adversarial features are input into the initial attack decoder to obtain the recovered data; the loss function is determined based on the difference between the recovered data and the local data; the initial attack decoder is optimized according to the loss function, and the recovered data is re-determined based on the adversarial features until the loss function converges or the number of training steps meets the preset threshold.
[0090] By training an attack decoder to recover the input data, the effectiveness of adversarial features can be quantitatively analyzed to measure the effectiveness of noisy data.
[0091] Figure 4 A schematic diagram of the architecture of a federated large model learning method based on adversarial feature aggregation provided in this application embodiment is shown below. Figure 4As shown, each participant extracts features from local data using a feature extractor to obtain intermediate features, and then processes these intermediate features using a noise generator to obtain noisy data. Adversarial features are then derived from the intermediate and noisy data and uploaded to the central service provider. An attack decoder performs data recovery on the intermediate and adversarial features respectively, and privacy metrics are performed based on the recovered data. The central service provider trains a target model for a specific task using the adversarial features uploaded by K participants.
[0092] The federated large model learning method based on adversarial feature aggregation provided in this application performs data recovery processing on intermediate features and adversarial features through an attack decoder, obtaining first recovered data corresponding to the adversarial features and second recovered data corresponding to the intermediate features. This provides a data foundation for subsequent quantification of the privacy protection capability of adversarial features. Privacy measurement is performed on the first and second recovered data to obtain privacy measurement results. The effectiveness of the noise is measured by comparing the data recovered by the attack decoder before adding noise and the data recovered by the attack decoder without adding noise, thereby quantifying the privacy protection capability of adversarial features.
[0093] Figure 5 This is a schematic diagram of the structure of a multi-party computing system provided in an embodiment of this application, such as... Figure 5 As shown, the multi-party computing system 10 provided in this embodiment includes: multiple participants 101 and a central service provider 102. The multiple participants 101 are all communicatively connected to the central service provider 102. Each participant includes a processing module 1011 and an uploading module 1012. The central service provider includes a training module 1021.
[0094] For any one of the plurality of participants, the participant's processing module 1011 is used to perform adversarial feature generation processing on local data based on a noise generator to obtain adversarial features, wherein the noise generator is used to generate noise data.
[0095] Upload module 1012 is used to upload the adversarial features to the central service provider;
[0096] The training module 1021 of the central service provider is used to train a model for a specific task through an adversarial feature cluster to obtain a target model. The adversarial feature cluster is obtained by aggregating adversarial features issued by the multiple participating parties.
[0097] In one possible implementation, the processing module 1011 is further configured to extract features from local data using a feature extractor issued by the central service provider to obtain intermediate features; perform adversarial feature generation processing on the intermediate features using a noise generator to obtain noise data; and fuse the intermediate features and the noise data to obtain the adversarial features.
[0098] In one possible implementation, the local data includes sample data and label data; the processing module 1011 is further configured to input the sample data and the label data into the noise generator, determine the loss based on the loss function and the label data, and determine the gradient based on the loss and the sample data;
[0099] The gradient is signified by a sign function, and the product of the signified gradient and the control parameter is used as the noise data. The control parameter is used to control the noise amplitude.
[0100] The noise data is superimposed on the sample data to determine the adversarial features.
[0101] In one possible implementation, the processing module 1011 is further configured to perform data recovery processing on the adversarial feature and the intermediate feature respectively through the attack decoder to obtain the first recovered data corresponding to the adversarial feature and the second recovered data corresponding to the intermediate feature;
[0102] Privacy metrics are performed on the first recovered data and the second recovered data to obtain privacy metric results, which are used to measure the validity of noisy data.
[0103] In one possible implementation, the privacy measurement results include KL divergence and similarity; the processing module 1011 is further configured to determine a first probability distribution corresponding to the first recovered data and a second probability distribution corresponding to the second recovered data, and determine KL divergence based on the first probability distribution and the second probability distribution, wherein the KL divergence is used to indicate the degree of deviation in recovering local data;
[0104] A similarity assessment vector is determined for the first recovered data and the second recovered data, and based on the similarity assessment vector, the similarity between the local data and the recovered local data is determined, wherein the similarity is used to indicate the risk of leakage.
[0105] In one possible implementation, the processing module 1011 is further configured to input the adversarial features into the initial attack decoder to obtain the recovered data;
[0106] The loss function is determined based on the degree of difference between the recovered data and the local data;
[0107] The initial attack decoder is optimized according to the loss function, and the recovery data is re-determined based on the adversarial features until the loss function converges or the number of training steps meets a preset threshold.
[0108] The multi-party computation system provided in this embodiment can execute the methods provided in the above-described method embodiments. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0109] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.
[0110] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.
[0111] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0112] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0113] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0114] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0115] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0116] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0117] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0118] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.
[0119] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0120] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0121] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0122] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0123] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0124] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0125] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A federated large model learning method based on adversarial feature aggregation, characterized in that, The method, applied to a multi-party computation system, includes multiple participating parties and a central service provider, wherein all participating parties are communicatively connected to the central service provider. For any one of the plurality of participants, the participant performs adversarial feature generation processing on local data based on a noise generator to obtain adversarial features, wherein the noise generator is used to generate noise data; The adversarial features are uploaded to the central service provider; The central service provider trains a model for a specific task using an adversarial feature cluster to obtain a target model. The adversarial feature cluster is obtained by aggregating adversarial features issued by the multiple participating parties.
2. The method according to claim 1, characterized in that, The adversarial feature generation process based on the noise generator on local data yields adversarial features, including: Intermediate features are obtained by extracting features from local data using a feature extractor provided by the central service provider. The intermediate features are subjected to adversarial feature generation processing by a noise generator to obtain noise data. The intermediate features and the noise data are fused together to obtain the adversarial features.
3. The method according to claim 2, characterized in that, The local data includes sample data and label data; the adversarial feature generation process performed on the intermediate features using a noise generator to obtain noise data includes: The sample data and the label data are input into the noise generator, the loss is determined based on the loss function and the label data, and the gradient is determined based on the loss and the sample data. The gradient is signified by a sign function, and the product of the signified gradient and the control parameter is used as the noise data. The control parameter is used to control the noise amplitude. The noise data is superimposed on the sample data to determine the adversarial features.
4. The method according to claim 2, characterized in that, The method further includes: The attack decoder performs data recovery processing on the adversarial feature and the intermediate feature respectively to obtain the first recovered data corresponding to the adversarial feature and the second recovered data corresponding to the intermediate feature; Privacy metrics are performed on the first recovered data and the second recovered data to obtain privacy metric results, which are used to measure the validity of noisy data.
5. The method according to claim 4, characterized in that, The privacy measurement results include KL divergence and similarity; the privacy measurement of the first recovered data and the second recovered data to obtain privacy measurement results includes: A first probability distribution corresponding to the first recovered data and a second probability distribution corresponding to the second recovered data are determined, and the KL divergence is determined based on the first probability distribution and the second probability distribution, wherein the KL divergence is used to indicate the degree of deviation in the recovery of local data; A similarity assessment vector is determined for the first recovered data and the second recovered data, and based on the similarity assessment vector, the similarity between the local data and the recovered local data is determined, wherein the similarity is used to indicate the risk of leakage.
6. The method according to claim 4, characterized in that, Before performing data recovery processing on the adversarial feature and the intermediate feature respectively through the attack decoder to obtain the first recovered data corresponding to the adversarial feature and the second recovered data corresponding to the intermediate feature, the method further includes: The adversarial features are input into the initial attack decoder to obtain the recovered data; The loss function is determined based on the degree of difference between the recovered data and the local data; The initial attack decoder is optimized according to the loss function, and the recovery data is re-determined based on the adversarial features until the loss function converges or the number of training steps meets a preset threshold.
7. A multi-party computation system, characterized in that, The multi-party computing system includes multiple participating parties and a central service provider. The multiple participating parties are all communicatively connected to the central service provider. Each participating party includes a processing module and an upload module. The central service provider includes a training module. For any one of the multiple participants, the processing module of that participant is used to perform adversarial feature generation processing on local data based on a noise generator to obtain adversarial features, wherein the noise generator is used to generate noise data. An upload module is used to upload the adversarial features to the central service provider; The training module of the central service provider is used to train a model for a specific task through an adversarial feature cluster to obtain a target model. The adversarial feature cluster is obtained by aggregating adversarial features issued by the multiple participating parties.
8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.