Model training method and device, electronic equipment and storage medium

By screening trusted participants and effective local model parameters, the problems of poor data quality and malicious model upload in federated learning are solved, efficient and accurate training of the global model is achieved, and the reliability of risk control decisions is improved.

CN120579657APending Publication Date: 2025-09-02AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510706903.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In federated learning, there are problems such as poor data quality, uneven data distribution and degradation in global risk control model performance caused by malicious model upload, which affects the risk control effect and is difficult to identify malicious participants.

Method used

By obtaining the local model parameters, number of parameter uploads, local sample information and historical usage information of the participants, using preset clustering methods and data processing to determine the confidence results, and filter out trusted participants and their effective local model parameters for global model training.

Benefits of technology

It improves the efficiency and accuracy of model training, ensures the effectiveness of the global model, prevents the influence of malicious models, and improves the reliability of risk control decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579657A_ABST
    Figure CN120579657A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring local model parameters, parameter uploading times, local sample information and historical use information corresponding to each first participant in a current communication round; determining a first confidence result based on a preset clustering mode and the local model parameters; determining a second confidence result based on the parameter uploading times and the local sample information; determining a third confidence result based on the number of communication rounds corresponding to the current communication round and the historical use information; determining a target confidence result based on the first confidence result, the second confidence result and the third confidence result; and determining a second participant based on a preset screening mode and the target confidence result, and training the global model by using local model parameters corresponding to the second participant. Through the technical scheme of the embodiment of the invention, effective training of the global model can be realized, and the efficiency and accuracy of model training are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of model training technology, and in particular to a model training method, device, electronic device and storage medium. Background Art

[0002] In recent years, risk control capabilities have become an invisible barrier to entry in the financial industry. Financial institutions face significant challenges in risk control due to information asymmetry, the lack of personal and corporate credit records, high manual review costs, and the difficulty in identifying overdue customer risks. For example, risk control capabilities can be reflected in the ability to utilize risk control models to assess the creditworthiness of borrowers in bank lending operations.

[0003] However, there is a lack of sufficient training data to improve the robustness of risk control models. Furthermore, there is a conflict between data privacy protection and data sharing. Federated learning technology is often introduced to enhance data sharing and utilization without leaking the original data. This allows risk control models to maintain good performance in the face of varying data distributions, outliers, or adversarial attacks, thereby improving their robustness.

[0004] Currently, model training in federated learning often relies on the local models of all participants, which are then integrated to form a global model. However, this model training approach, which relies on the local models of all participants, can lead to poor training results for some participants if their data quality is poor or unevenly distributed. These local weaknesses can be transmitted to the global model through the model aggregation process in federated learning, thereby affecting overall risk control effectiveness. Furthermore, in federated learning, without effective security mechanisms and regulatory oversight, malicious participants could upload models designed to undermine the global model. However, because each participant in federated learning encrypts their local model parameters when uploading them, the validity of the model parameters cannot be directly verified. Ultimately, such malicious models can degrade the performance of the global model and even lead to erroneous risk control decisions, a phenomenon known as model poisoning. Summary of the Invention

[0005] The embodiments of the present invention provide a model training method, device, electronic device and storage medium to screen participants, so as to accurately and conveniently determine the second participant that can participate in the global model training and its corresponding valid local model parameters, so as to realize the effective training of the global model and improve the efficiency and accuracy of model training.

[0006] In a first aspect, an embodiment of the present invention provides a model training method, comprising:

[0007] Obtain the local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round;

[0008] Performing clustering processing based on a preset clustering method and the local model parameters to determine a first confidence result corresponding to each first participant;

[0009] Performing data processing based on the number of parameter uploads and the local sample information to determine a second confidence result corresponding to each first participant;

[0010] performing data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant;

[0011] Performing data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant;

[0012] Based on a preset screening method and the target confidence result, each first participant is screened to determine the second participant, and the global model is trained using the local model parameters corresponding to the second participant.

[0013] Optionally, the method also includes: clustering based on a preset clustering method and local model parameters corresponding to each first participant, determining multiple model parameter sets and the number of participants corresponding to each model parameter set; performing data processing based on the number of participants and the total number of participants corresponding to the first participant, and determining a first confidence result corresponding to each first participant.

[0014] Optionally, the method also includes: for each first participant, dividing the number of parameter uploads corresponding to the current participant by the local sample size in the local sample information to obtain a first division result, and using the first division result as the second confidence result corresponding to the current participant.

[0015] Optionally, the method also includes: for each first participant, dividing the number of historical usages in the historical usage information corresponding to the current participant by the number of communication rounds corresponding to the current communication round to obtain a second division result, and using the second division result as the third confidence result corresponding to the current participant.

[0016] Optionally, the method also includes: performing data fusion based on the conflict factor between the first confidence result and the second confidence result, the first confidence result and the second confidence result to determine the fourth confidence result corresponding to each first participant; performing data fusion based on the conflict factor between the third confidence result and the fourth confidence result, the third confidence result and the fourth confidence result to determine the target confidence result corresponding to each first participant.

[0017] Optionally, the method also includes: performing data processing on the target confidence results corresponding to each first participant to determine the average confidence results and median confidence results corresponding to each first participant; screening each first participant based on the average confidence result and the median confidence result to determine the second participant that can participate in the global model training in the current communication round.

[0018] In a second aspect, an embodiment of the present invention further provides a model training device, the device comprising:

[0019] An information acquisition module is used to obtain the local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round;

[0020] a first confidence result determination module, configured to perform clustering processing based on a preset clustering method and the local model parameters to determine a first confidence result corresponding to each first participant;

[0021] a second confidence result determination module, configured to perform data processing based on the number of parameter uploads and the local sample information to determine a second confidence result corresponding to each first participant;

[0022] a third confidence result determination module, configured to perform data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant;

[0023] a target confidence result determination module, configured to perform data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant;

[0024] The second participant determination module is used to screen each first participant based on a preset screening method and the target confidence result to determine the second participant, and use the local model parameters corresponding to the second participant to train the global model.

[0025] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:

[0026] one or more processors;

[0027] a memory for storing one or more programs;

[0028] When the one or more programs are executed by the one or more processors, the one or more processors implement the model training method provided by any embodiment of the present invention.

[0029] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a model training method as provided in any embodiment of the present invention.

[0030] In a fifth aspect, an embodiment of the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements a model training method as provided in any embodiment of the present invention.

[0031] The technical solution of the embodiment of the present invention obtains the local model parameters, parameter upload times, local sample information and historical usage information corresponding to each first participant in the current communication round; performs clustering processing based on a preset clustering method and the local model parameters to determine a first confidence result corresponding to each first participant; performs data processing based on the parameter upload times and the local sample information to determine a second confidence result corresponding to each first participant; performs data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant; performs data fusion based on the first confidence result, the second confidence result and the third confidence result to determine a target confidence result corresponding to each first participant; screens each first participant based on a preset screening method and the target confidence result to determine a second participant, and uses the local model parameters corresponding to the second participant to train the global model, so that the participants can be screened, and then the second participants that can participate in the global model training and their corresponding valid local model parameters are accurately and conveniently determined, so as to achieve effective training of the global model and improve the efficiency and accuracy of model training.

[0032] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0034] Figure 1 This is a flow chart of a model training method provided in Example 1 of the present invention;

[0035] Figure 2 This is a flow chart of a model training method provided in Example 2 of the present invention;

[0036] Figure 3 This is a schematic diagram of the structure of a model training device provided in Example 3 of the present invention;

[0037] Figure 4 It is a structural diagram of an electronic device for implementing the model training method of an embodiment of the present invention. DETAILED DESCRIPTION

[0038] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0039] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0040] Example 1

[0041] Figure 1A flowchart of a model training method is provided for the first embodiment of the present invention. This embodiment is applicable to the case where it is determined in each communication round whether the first participant can participate in the global model training. The method can be executed by a model training device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0042] S110: Obtain local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round.

[0043] For example, in the model training process of federated learning, there are local models, a central server, and multiple first participants. All first participants perform machine learning with the help of the server. The model training process of federated learning can be divided into multiple communication rounds, each of which mainly consists of four steps. In the first step, each first participant will perform model training locally to obtain the model gradient, and use an encryption algorithm to ensure that the model parameter information is not leaked. In the second step, the server will adopt a model aggregation algorithm based on homomorphic encryption and use the local model parameters uploaded by the first participant to perform model aggregation. In the third step, the server will send the aggregated global model to each first participant. In the fourth step, each first participant will decrypt the received global model and use the decrypted parameters to update their respective local model parameters, so that the model gradient decreases. Repeat the above training process and perform continuous iterations until the loss function remains stable.

[0044] The current communication round may refer to any communication round during the model training process of the global model in federated learning on the server. Communication rounds are divided based on the model parameters received by the participants. The first participant may refer to the participant that uploaded the local model parameters in the current communication round. In this embodiment, federated learning may adopt an asynchronous training strategy, such as dynamic participant selection: the server may only select a certain number of participants in each training round (e.g., through a heuristic algorithm), and unselected participants do not participate in the upload in this communication round. Another example is resource-sensitive scheduling: participants with weak computing power or high network latency may be skipped, reducing their upload frequency. Local model parameters may refer to the local model parameters obtained by the first participant using local training samples to train the local model in the current communication round. Local model parameters need to be uploaded to the server. The number of parameter uploads may refer to the number of times the first participant uploaded local model parameters to the server in the current communication round and all previous historical communication rounds. Local sample information may refer to the volume of training samples used by the first participant to train the local model. In this embodiment, the local sample information does not include the specific values ​​of the local samples, thereby avoiding the leakage of local samples and further realizing the privacy protection of local data.

[0045] The historical usage information may refer to information about the first participant's participation in global model training in historical communication rounds. The historical usage information may include, but is not limited to, historical usage counts and historically used model parameters. The historical usage counts may refer to the number of times the first participant's local model parameters were used for global model training in historical communication rounds. The historically used model parameters may refer to the local model parameters used for each historical usage count. In this embodiment, the historical usage count may be understood as the number of times the first participant was determined as the second participant through validity screening in historical communication rounds.

[0046] Specifically, based on the pre-set program code for realizing the information acquisition function, when all first participants have uploaded their corresponding local model parameters in each communication round, the number of parameter uploads, local sample information and historical usage information corresponding to each first participant in the current communication round are obtained.

[0047] S120: Perform clustering processing based on a preset clustering method and local model parameters to determine a first confidence result corresponding to each first participant.

[0048] The preset clustering method may refer to a pre-set parameter clustering method. For example, the preset clustering method may be, but is not limited to, a K-means clustering method. The first confidence result may refer to a preliminary determination of whether the first participant is trustworthy based on the clustering results. For example, a higher first confidence result corresponds to a higher degree of trustworthiness of the first participant.

[0049] Specifically, the local model parameters are clustered using K-means clustering to form multiple clusters. The smaller the size of a cluster (i.e., the number of parameters), the higher the parameter outliers of the cluster, and the lower the credibility. Assume that 8 of the 10 nodes have similar parameters (large clusters), and 2 have abnormal parameters due to data contamination (small clusters). The small clusters represent low-credibility nodes. The size of the cluster containing the local model parameters corresponding to the first participant is quantified to determine the first confidence result corresponding to each first participant.

[0050] In this embodiment, cluster credibility can also be judged in combination with multi-dimensional indicators, such as the proportion of samples within the cluster, the consistency of parameters within the cluster, and the distance between clusters. Among them, the proportion of samples within the cluster: the lower the proportion, the lower the credibility. The consistency of parameters within the cluster: the compactness is measured by SSE (intra-cluster squared error), and the smaller the SSE, the higher the credibility. The distance between clusters: the farther away from other clusters, the more likely it is an anomaly. Therefore, the first confidence result can be calculated using multi-dimensional indicators and the preset weights corresponding to each dimensional indicator.

[0051] On the basis of the above technical solution, "clustering based on the preset clustering method and local model parameters to determine the first confidence result corresponding to each first participant" may include: clustering based on the preset clustering method and the local model parameters corresponding to each first participant to determine multiple model parameter sets and the number of participants corresponding to each model parameter set; performing data processing based on the number of participants and the total number of participants corresponding to the first participant to determine the first confidence result corresponding to each first participant.

[0052] The model parameter set may refer to a cluster obtained through clustering. The number of participants may refer to the number of first participants to which the local model parameters contained in the model parameter set belong. The total number of participants may refer to the number of all first participants in the current communication round.

[0053] Specifically, the local model parameters corresponding to each first participant are clustered using a K-means clustering method to determine multiple model parameter sets and the number of participants corresponding to each model parameter set. The number of participants corresponding to each model parameter set is divided by the total number of participants corresponding to the first participant to obtain a division result, and the division result is used as the first confidence result corresponding to the first participant to which the local model parameters contained in the model parameter set belong.

[0054] S130: Perform data processing based on the number of parameter uploads and local sample information to determine a second confidence result corresponding to each first participant.

[0055] The second confidence result can refer to a preliminary determination of the trustworthiness of the first participant based on data processing results related to the participant's own size. For example, a higher second confidence result corresponds to a lower trustworthiness of the first participant. Considering that federated learning involves participants continuously uploading local model parameters, when local model parameters are uploaded, it indicates that the local model parameters have been updated. If a participant intentionally uploads harmful parameters to affect the global model, then they will have to continuously upload parameters, i.e., upload at a high frequency. For a normal participant with insufficient data, the frequency of local model updates will be low.

[0056] Specifically, the number of parameter upload times and local sample information are quantified to determine the second confidence result corresponding to each first participant.

[0057] Based on the above technical solution, "data processing based on the number of parameter uploads and local sample information to determine the second confidence result corresponding to each first participant" may include: for each first participant, dividing the number of parameter uploads corresponding to the current participant by the local sample size in the local sample information to obtain a first division result, and using the first division result as the second confidence result corresponding to the current participant.

[0058] The local sample size can refer to the number of local samples. In this embodiment, the local sample size can also refer to the number of enterprise users. Furthermore, considering that the enterprise user size of some first participants is private information and is subjectively provided by the first participants, to avoid reducing the accuracy of participant screening due to inaccurate data collection, the total number of parameter uploads by all first participants in the current communication round can be used instead of the enterprise user size or local sample information.

[0059] S140: Perform data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine a third confidence result corresponding to each first participant.

[0060] The number of communication rounds may refer to the round number corresponding to the current communication round. For example, if the current communication round is the 9th communication round, the number of communication rounds is 9. The third confidence result may refer to a preliminary determination of the trustworthiness of the first participant based on the data processing results related to the effective participation of the participant. For example, a larger third confidence result indicates a higher degree of trustworthiness of the first participant.

[0061] Specifically, the number of communication rounds and historical usage information corresponding to the current communication round are quantified to determine a third confidence result corresponding to each first participant.

[0062] On the basis of the above technical solution, "performing data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine the third confidence result corresponding to each first participant" may include: for each first participant, dividing the number of historical usage times in the historical usage information corresponding to the current participant by the number of communication rounds corresponding to the current communication round to obtain a second division result, and using the second division result as the third confidence result corresponding to the current participant.

[0063] The historical usage count can be understood as the number of times a first participant was identified as a second participant through validity screening in historical communication rounds. The historical usage count can also be understood as the number of times a first participant was identified as a normal participant in historical communication rounds. The second participant can be a normal participant. The remaining participant, obtained by removing the second participant from all first participants, is the third participant, also known as the abnormal participant.

[0064] S150: Perform data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant.

[0065] The target confidence result may refer to a comprehensive confidence result obtained by integrating the preliminary judgment results of the three aspects. The target confidence result may be used to ultimately determine whether the first party is trustworthy.

[0066] Specifically, based on the preset weight corresponding to the first confidence result, the preset weight corresponding to the second confidence result, the weight corresponding to the third confidence result, the first confidence result, the second confidence result and the third confidence result, a weighted calculation is performed to determine the target confidence result corresponding to each first participant.

[0067] S160: Screen each first participant based on a preset screening method and a target confidence result to determine a second participant, and train the global model using local model parameters corresponding to the second participant.

[0068] The preset screening method may refer to a pre-set method for screening participants based on a threshold. For example, the target confidence results may be sorted in descending order, and the first participants corresponding to the first preset number of target confidence results may be determined as the second participants. Alternatively, the target confidence results may be screened by taking the average of all target confidence results, and the first participants corresponding to target confidence results exceeding the average may be determined as the second participants. After global model training is completed in the current communication round, the global model parameters are distributed to all first participants. The global model may be, but is not limited to, a risk control model or an anti-fraud risk control model.

[0069] The technical solution of the embodiment of the present invention obtains the local model parameters, parameter upload times, local sample information and historical usage information corresponding to each first participant in the current communication round; performs clustering processing based on a preset clustering method and local model parameters to determine the first confidence result corresponding to each first participant; performs data processing based on the parameter upload times and local sample information to determine the second confidence result corresponding to each first participant; performs data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine the third confidence result corresponding to each first participant; performs data fusion based on the first confidence result, the second confidence result and the third confidence result to determine the target confidence result corresponding to each first participant; screens each first participant based on a preset screening method and the target confidence result to determine the second participant, and uses the local model parameters corresponding to the second participant to train the global model, so that the participants can be screened, and then the second participants that can participate in the global model training and their corresponding valid local model parameters are accurately and conveniently determined, so as to achieve effective training of the global model and improve the efficiency and accuracy of model training.

[0070] Example 2

[0071] Figure 2 This is a flowchart of a model training method provided in the second embodiment of the present invention. Based on the above embodiment, this embodiment describes in detail the process of determining the second participant. The explanations of the terms that are the same or corresponding to the above embodiments are not repeated here. Figure 2 As shown, the method includes:

[0072] S210: Obtain local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round.

[0073] S220: Perform clustering processing based on a preset clustering method and local model parameters to determine a first confidence result corresponding to each first participant.

[0074] S230: Perform data processing based on the number of parameter uploads and local sample information to determine a second confidence result corresponding to each first participant.

[0075] S240: Perform data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine a third confidence result corresponding to each first participant.

[0076] S250: Perform data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant.

[0077] Based on the above technical solution, "performing data fusion based on the first confidence result, the second confidence result and the third confidence result to determine the target confidence result corresponding to each first participant" may include: performing data fusion based on the conflict factor between the first confidence result and the second confidence result, the first confidence result and the second confidence result to determine the fourth confidence result corresponding to each first participant; performing data fusion based on the conflict factor between the third confidence result and the fourth confidence result, the third confidence result and the fourth confidence result to determine the target confidence result corresponding to each first participant.

[0078] The conflict factor may be an indicator of the degree of inconsistency between two confidence results. The conflict factor may be used to reflect the contradiction or opposition between the two confidence results. A larger conflict factor indicates a more severe conflict between the two confidence results. A zero conflict factor indicates that the two confidence results are completely consistent. The fourth confidence result may be a confidence result obtained by combining the preliminary judgment results of the two aspects.

[0079] Specifically, since the first confidence result, the second confidence result, and the third confidence result determined are independent and incompatible with each other, the target confidence result can be obtained by the following formula:

[0080]

[0081] Among them, B i (reliable) represents the target confidence result of the i-th first participant. Represents the first confidence result of the i-th first participant. Represents the second confidence result of the i-th first participant. Represents the third confidence result of the i-th first participant. For the first confidence result, the second confidence result, and the third confidence result, the three confidence results can be fused first with the two confidence results, and then fused with the third confidence result after the results are obtained. The fusion rules of the two confidence results are as follows:

[0082]

[0083] Among them, F is the conflict factor, which indicates the degree of conflict between two confidence results. The calculation formula is as follows:

[0084]

[0085] According to the above formula, after obtaining the fourth confidence result obtained by fusing the two confidence results, the fourth confidence result is fused with the third confidence result to obtain the final target confidence result.

[0086] It should be noted that the second confidence result and the third confidence result may be fused first, and then the fused confidence result may be fused with the first confidence result to obtain the target confidence result.

[0087] S260: Perform data processing on the target confidence results corresponding to each first participant to determine the average confidence result and the median confidence result corresponding to each first participant.

[0088] The average confidence result may refer to the average value of the target confidence results corresponding to all first participants, and the median confidence result may refer to the median of the target confidence results corresponding to all first participants.

[0089] Specifically, after determining the target confidence result, it is necessary to judge whether the local model (i.e., local model) of the i-th first participant can be aggregated into a robust global model according to pre-defined rules. Therefore, for the target confidence result set {B i (reliable)}, 1≤i≤N, sort the results in the set in descending order, and calculate the corresponding average value (i.e., average confidence result) and median (i.e., median confidence result) of the set.

[0090] S270: Screen each first participant based on the average confidence result and the median confidence result to determine a second participant that can participate in the global model training in the current communication round.

[0091] Specifically, two pointers t1 and t2 can be set. By comparing the average and median values, the smaller value is assigned to t1 and the larger value is assigned to t2. In this way, the target confidence result of the first participant in the set will be distributed in three intervals, namely less than t1, between t1 and t2, and greater than t2. When aggregating, only local models with confidence functions greater than t2 are aggregated. The first participant less than t1 will be considered unreliable, and the number of abnormalities will be recorded once. The first participant between t1 and t2 will not be recorded. The first participant greater than t2 is determined to be the second participant, and the number of times as a normal participant is recorded once to update the historical usage information. Finally, in each round of communication, the calculation and screening of the target confidence result are repeated, and the model aggregation process is performed until the global model gradually converges or reaches the target training number.

[0092] For example, during the model aggregation process, a verify-before-aggregate approach can be used. This approach verifies the local model using a validation dataset before the global model is aggregated. If the local model fails to obtain good results after verification, the model is removed from the aggregation of the global model. A GAN network can also be trained using a validation dataset to obtain an anomaly detector and set an anomaly threshold. If the outlier value of a participant is higher than the threshold, the local model in the participant will be excluded from the aggregation.

[0093] S280: Train the global model using the local model parameters corresponding to the second participant.

[0094] The technical solution of the embodiment of the present invention determines the average confidence result and median confidence result corresponding to each first participant by performing data processing on the target confidence result corresponding to each first participant; based on the average confidence result and the median confidence result, each first participant is screened to determine the second participant that can participate in the global model training in the current communication round, thereby effectively screening the first participant using multiple indicators such as the average confidence result and the median confidence result to obtain the second participant, further improving the reliability of model aggregation.

[0095] For example, to avoid invalid or even harmful parameters in uploaded local models, this embodiment of the present invention also proposes a participant screening method based on DS evidence theory. The DS evidence theory framework consists of four parts, thus defining a quadruple (A, E, P, B). A represents an identification framework, i.e., the hypothesis space of all propositions. In this embodiment, two propositions are defined: reliable and unreliable, so A can be expressed as A = {reliable, unreliable}. E represents a set of evidence used to assess the trustworthiness of the first participant, and each element in this set, i.e., evidence, must be independent of each other. For different scenarios, customized evidence can be used that does not involve specific numerical values ​​of local samples. Based on the characteristics of federated learning and privacy protection, this embodiment defines three pieces of evidence to demonstrate the trustworthiness of the first participant. These are based on clustering, the ratio of upload times to the number of enterprise users, and historical trust values. P represents a set of basic probability distribution functions. This describes the extent to which the evidence supports the two propositions in A, i.e., the process of quantifying the evidence. B represents the trust function or confidence function, which represents the sum of the trust of each proposition, that is, the trust function of the total reliable or unreliable proposition calculated after integrating the three evidences of different first participants.

[0096] Evidence 1: Clustering-based evidence. Considering that each first-party participant needs to encrypt their local model parameters using the same encryption method and upload them to the central server, machine learning models must be identical to ensure final model convergence. Excessive parameter outliers among participants will slow model convergence. Therefore, K-means clustering can be used to cluster local model parameters. Smaller clusters indicate higher parameter outliers and lower credibility. Ultimately, the size of the cluster in which the participant's model resides is quantified as the first piece of evidence. Evidence 2: The ratio of upload times to the number of enterprise users. Given that federated learning involves participants continuously uploading local model parameters, model uploads indicate an update to the model parameters. However, if a participant intentionally uploads harmful parameters to impact the global model, they will only be able to continuously upload parameters, resulting in a high upload frequency. For a normal participant with insufficient data, the frequency of local model updates will be low. Therefore, quantifying the ratio of upload times to the number of enterprise users provides the second piece of evidence. Evidence 3: The participant's historical trustworthiness. During each round of communication, the participants will be judged based on the confidence results of the DS evidence theory. We can define the historical trust value as the ratio of the number of times the participants are judged to be normal to the total number of communications.

[0097] The following is an embodiment of the model training device provided in an embodiment of the present invention. The device and the model training methods of the above embodiments belong to the same inventive concept. For details not fully described in the embodiment of the model training device, please refer to the embodiment of the above model training method.

[0098] Example 3

[0099] Figure 3 This is a structural diagram of a model training device provided by the third embodiment of the present invention. Figure 3 As shown, the apparatus includes: an information acquisition module 310 , a first confidence result determination module 320 , a second confidence result determination module 330 , a third confidence result determination module 340 , a target confidence result determination module 350 and a second participant determination module 360 ​​.

[0100] Among them, the information acquisition module 310 is used to obtain the local model parameters, parameter upload times, local sample information and historical usage information corresponding to each first participant in the current communication round; the first confidence result determination module 320 is used to perform clustering processing based on the preset clustering method and local model parameters to determine the first confidence result corresponding to each first participant; the second confidence result determination module 330 is used to perform data processing based on the parameter upload times and local sample information to determine the second confidence result corresponding to each first participant; the third confidence result determination module 340 is used to perform data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine the third confidence result corresponding to each first participant; the target confidence result determination module 350 is used to perform data fusion based on the first confidence result, the second confidence result and the third confidence result to determine the target confidence result corresponding to each first participant; the second participant determination module 360 ​​is used to screen each first participant based on the preset screening method and the target confidence result to determine the second participant, and use the local model parameters corresponding to the second participant to train the global model.

[0101] The technical solution of the embodiment of the present invention obtains the local model parameters, parameter upload times, local sample information and historical usage information corresponding to each first participant in the current communication round; performs clustering processing based on a preset clustering method and local model parameters to determine the first confidence result corresponding to each first participant; performs data processing based on the parameter upload times and local sample information to determine the second confidence result corresponding to each first participant; performs data processing based on the number of communication rounds corresponding to the current communication round and historical usage information to determine the third confidence result corresponding to each first participant; performs data fusion based on the first confidence result, the second confidence result and the third confidence result to determine the target confidence result corresponding to each first participant; screens each first participant based on a preset screening method and the target confidence result to determine the second participant, and uses the local model parameters corresponding to the second participant to train the global model, so that the participants can be screened, and then the second participants that can participate in the global model training and their corresponding valid local model parameters are accurately and conveniently determined, so as to achieve effective training of the global model and improve the efficiency and accuracy of model training.

[0102] Based on the above technical solution, the first confidence result determination module 320 is specifically used to: perform clustering based on a preset clustering method and the local model parameters corresponding to each first participant, determine multiple model parameter sets and the number of participants corresponding to each model parameter set; perform data processing based on the number of participants and the total number of participants corresponding to the first participant, and determine the first confidence result corresponding to each first participant.

[0103] Based on the above technical solution, the second confidence result determination module 330 is specifically used to: for each first participant, divide the number of parameter uploads corresponding to the current participant by the local sample size in the local sample information to obtain a first division result, and use the first division result as the second confidence result corresponding to the current participant.

[0104] Based on the above technical solution, the third confidence result determination module 340 is specifically used to: for each first participant, divide the number of historical usages in the historical usage information corresponding to the current participant by the number of communication rounds corresponding to the current communication round to obtain a second division result, and use the second division result as the third confidence result corresponding to the current participant.

[0105] Based on the above technical solution, the target confidence result determination module 350 is specifically used to: perform data fusion based on the conflict factor between the first confidence result and the second confidence result, the first confidence result and the second confidence result to determine the fourth confidence result corresponding to each first participant; perform data fusion based on the conflict factor between the third confidence result and the fourth confidence result, the third confidence result and the fourth confidence result to determine the target confidence result corresponding to each first participant.

[0106] Based on the above technical solution, the second participant determination module 360 ​​is specifically used to: perform data processing on the target confidence results corresponding to each first participant, and determine the average confidence results and median confidence results corresponding to each first participant; screen each first participant based on the average confidence results and median confidence results, and determine the second participant that can participate in the global model training in the current communication round.

[0107] The model training device provided in the embodiment of the present invention can execute the model training method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the model training method.

[0108] It is worth noting that in the above-mentioned model training embodiment, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0109] Example 4

[0110] Figure 4A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0111] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0112] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0113] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors for running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the model training method.

[0114] In some embodiments, the model training method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model training method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the model training method in any other appropriate manner (e.g., by means of firmware).

[0115] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0116] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0119] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0120] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0121] An embodiment of the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the model training method provided in any embodiment of the present application.

[0122] In the process of implementation, the computer program product can be written in one or more programming languages ​​or a combination thereof to write a computer program code for performing the operation of the present invention, and the programming language includes an object-oriented programming language, such as Java, Smalltalk, C++, and also includes a conventional procedural programming language, such as "C" language or similar programming language. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). The program product and the model training method disclosed in each embodiment of the present application belong to the same inventive concept, so they are not described here.

[0123] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0124] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A model training method, characterized in that: include: Obtain the local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round; Performing clustering processing based on a preset clustering method and the local model parameters to determine a first confidence result corresponding to each first participant; Performing data processing based on the number of parameter uploads and the local sample information to determine a second confidence result corresponding to each first participant; performing data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant; Performing data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant; Based on a preset screening method and the target confidence result, each first participant is screened to determine the second participant, and the global model is trained using the local model parameters corresponding to the second participant.

2. The method according to claim 1, characterized in that The performing clustering processing based on the preset clustering method and the local model parameters to determine the first confidence result corresponding to each first participant includes: Clustering is performed based on a preset clustering method and local model parameters corresponding to each first participant, and determining multiple model parameter sets and the number of participants corresponding to each model parameter set; Data processing is performed based on the number of participants and the total number of participants corresponding to the first participant to determine a first confidence result corresponding to each first participant.

3. The method according to claim 1, characterized in that The performing data processing based on the parameter upload times and the local sample information to determine the second confidence result corresponding to each first participant includes: For each first participant, the number of parameter uploads corresponding to the current participant is divided by the local sample size in the local sample information to obtain a first division result, and the first division result is used as the second confidence result corresponding to the current participant.

4. The method according to claim 1, wherein The performing data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant includes: For each first participant, the number of historical usages in the historical usage information corresponding to the current participant is divided by the number of communication rounds corresponding to the current communication round to obtain a second division result, and the second division result is used as the third confidence result corresponding to the current participant.

5. The method according to claim 1, characterized in that The performing data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant includes: performing data fusion based on a conflict factor between the first confidence result and the second confidence result, the first confidence result, and the second confidence result to determine a fourth confidence result corresponding to each first participant; A target confidence result corresponding to each first participant is determined by performing data fusion based on the conflict factor between the third confidence result and the fourth confidence result, the third confidence result, and the fourth confidence result.

6. The method according to claim 1, characterized in that The step of screening each first participant based on the preset screening method and the target confidence result to determine the second participant includes: Performing data processing on the target confidence results corresponding to each first participant to determine an average confidence result and a median confidence result corresponding to each first participant; The first participants are screened based on the average confidence result and the median confidence result to determine a second participant that can participate in the global model training in the current communication round.

7. A model training device, characterized in that: The device comprises: An information acquisition module is used to obtain the local model parameters, parameter upload times, local sample information, and historical usage information corresponding to each first participant in the current communication round; a first confidence result determination module, configured to perform clustering processing based on a preset clustering method and the local model parameters to determine a first confidence result corresponding to each first participant; a second confidence result determination module, configured to perform data processing based on the number of parameter uploads and the local sample information to determine a second confidence result corresponding to each first participant; a third confidence result determination module, configured to perform data processing based on the number of communication rounds corresponding to the current communication round and the historical usage information to determine a third confidence result corresponding to each first participant; a target confidence result determination module, configured to perform data fusion based on the first confidence result, the second confidence result, and the third confidence result to determine a target confidence result corresponding to each first participant; The second participant determination module is used to screen each first participant based on a preset screening method and the target confidence result to determine the second participant, and use the local model parameters corresponding to the second participant to train the global model.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the model training method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the model training method as described in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements the model training method as described in any one of claims 1 to 6.