Participant selection method and equipment
By evaluating the local data features, label correlation and distribution differences of the participating nodes, appropriate participating nodes are selected, which solves the problem of low client selection efficiency in distributed machine learning and improves the model training efficiency and performance.
Patent Information
- Application Number
- CN202510756923.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
In distributed machine learning, as the number of clients increases, existing technologies find it difficult to efficiently select better clients, resulting in decreased model training efficiency and performance.
By determining the correlation and distribution differences between the local data features and data labels of each participant's nodes, a predetermined number of target participant nodes are selected, and indicators such as mutual information and Jensen-Shannon divergence are used to evaluate the node contribution and diversity to avoid insufficient and redundant data information.
The efficiency and performance of model training are improved, and appropriate participant nodes are selected while ensuring data privacy, thereby improving the training effect of the model.
Smart Images

Figure CN120671777A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of machine learning technology, and in particular to a participant selection method and electronic device. Background Art
[0002] With the development of machine learning technology, the use of sample data to train machine learning models is becoming increasingly common. How to efficiently use data from different clients for machine learning has become a focus of attention.
[0003] In related technical solutions, such as distributed machine learning (Federated Learning), multiple clients (such as terminal devices and edge nodes) jointly train models without sharing raw data. However, as the number of clients increases, adding more clients can reduce model training efficiency and affect model performance, as not all clients are equally important. Therefore, how to select the best client from among the multiple clients in a distributed learning platform has become a technical challenge that needs to be solved urgently.
[0004] The content of the background technology section is merely information known to the inventor personally, and does not mean that the above information has entered the public domain before the application date of this disclosure, nor does it mean that it can become the prior art of the present disclosure. Summary of the Invention
[0005] This specification provides a participant selection method and electronic device that can accurately and efficiently select a better client from multiple clients of a distributed learning platform, thereby improving the efficiency of model training and model performance.
[0006] In a first aspect, embodiments of this specification provide a participant selection method, applied to a distributed learning platform, the distributed learning platform including multiple participant nodes, each of the participant nodes including local data features of each of the data instances in a plurality of data instances, one of the multiple participant nodes including a data label of the data instance, the method comprising:
[0007] Determining the correlation between the local data features of each of the participant nodes and the data labels;
[0008] determining distribution differences in data distribution of the local data features between the participant nodes; and
[0009] Based on the correlation corresponding to the participant nodes and the distribution difference between the participant nodes, a predetermined number of target participant nodes are selected from the multiple participant nodes.
[0010] In some example embodiments, based on the above solution, selecting a predetermined number of target participant nodes from the plurality of participant nodes based on the correlations corresponding to the participant nodes and the distribution differences between the participant nodes includes:
[0011] Determine a node score for each of the participant nodes based on the correlation corresponding to each of the participant nodes;
[0012] Based on the distribution difference between the participant nodes, updating the node score of each of the participant nodes;
[0013] A target participant node is selected from the plurality of participant nodes based on the updated node score.
[0014] In some example embodiments, based on the above solution, updating the node score of each of the participant nodes based on the distribution difference between the participant nodes includes:
[0015] Selecting the participant node with the highest node score among the multiple participant nodes as the current participant node;
[0016] The node score of another participant node is updated based on the distribution difference between the current participant node and the another participant node.
[0017] In some example embodiments, based on the above solution, selecting a target participant node from the plurality of participant nodes based on the updated node score includes:
[0018] For each round of iterative update: the current participant node is added to the target participant node set until the number of elements in the target participant node set reaches the predetermined number.
[0019] In some example embodiments, based on the above solution, determining the correlation between the local data features of each participant node and the data tag includes:
[0020] Determine the mutual information between the local data features of each of the participant nodes and the data label, where the mutual information is used to represent the correlation between the local data features of the participant nodes and the data label.
[0021] In some example embodiments, based on the above solution, determining the mutual information between the local data features of each of the participant nodes and the data label includes:
[0022] Obtaining a characteristic distance between each of the data instances of the participant node and a query instance in the query instance set;
[0023] Based on the characteristic distance between the data instance of the participant node and the query instance, the mutual information of each participant node is determined by a K-nearest neighbor method.
[0024] In some example embodiments, based on the above solution, determining the mutual information of each of the participant nodes using a K-nearest neighbor approach based on the characteristic distance between the data instance of the participant node and the query instance includes:
[0025] Determine, from the multiple data instances, a set of instances with the same label as the query instance and the corresponding number of instances with the same label;
[0026] Determine a set of K nearest neighbor instances corresponding to the query instance from the set of instances with the same label based on the feature distance;
[0027] Determining a number of global neighbors corresponding to the query instance from the plurality of data instances based on the feature distance and the set of K nearest neighbor instances; and
[0028] The mutual information of each of the participant nodes is determined based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K nearest neighbors and the number of global neighbors through the K nearest neighbor mutual information function.
[0029] In some example embodiments, based on the above solution, determining the number of global neighbors corresponding to the query instance from the multiple data instances based on the feature distance and the K-nearest neighbor instance set includes:
[0030] Determining a distance threshold corresponding to the query instance based on the feature distance and the K nearest neighbor instance set, the distance threshold being the maximum feature distance for the query instance in the K nearest neighbor instance set; and
[0031] The number of global neighbors is determined based on the number of instances in the plurality of data instances whose feature distance to the query instance is less than the distance threshold.
[0032] In some example embodiments, based on the above solution, determining the mutual information of each of the participating nodes using a K-nearest neighbor mutual information function based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K-nearest neighbors, and the number of global neighbors includes:
[0033] Determining, based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K nearest neighbors, and the number of global nearest neighbors, a first mutual information function of each of the participant nodes for the query instance;
[0034] The second mutual information of the participant node is determined based on an average value of the first mutual information corresponding to each of the query instances in the query instance set.
[0035] In some example embodiments, based on the above solution, determining the mutual information between the local data features of each of the participant nodes and the data label includes:
[0036] generating a predetermined number of test groups based on the plurality of participant nodes, each of the test groups including a number of the participant nodes;
[0037] Determining group mutual information corresponding to each of the test groups, where the group mutual information represents mutual information between the local data features of the plurality of participant nodes included in the test group and the data labels; and
[0038] Based on the average value of the group mutual information of each of the test groups in which the participant node participates, the mutual information corresponding to each of the participant nodes is determined.
[0039] In some example embodiments, based on the above solution, the distributed learning platform further includes an aggregation node, and before determining the group mutual information corresponding to each of the test groups, the method further includes:
[0040] Obtaining, through the aggregation node, a local feature distance between the data instance sent by each participant node in the current test group and the query instance, where the local feature distance represents a distance between the local data feature of the data instance of the participant node and the query local feature of the query instance;
[0041] The local feature distances of the participant nodes in the current test group for the query instance are aggregated by the aggregation node to obtain the global feature distances of the participant nodes in the current test group for the query instance.
[0042] In some example embodiments, based on the above solution, aggregating the local feature distances of the participant nodes in the current test group for the query instance to obtain the global feature distances of the participant nodes in the current test group for the query instance includes:
[0043] An addition operation is performed on each local feature distance of each participant node in the current test group with respect to the query instance to obtain a global feature distance of each participant node in the current test group with respect to the query instance.
[0044] In some example embodiments, based on the above solution, before determining the group mutual information corresponding to each of the test groups, the method further includes:
[0045] Obtaining, through the aggregation node, a feature distance ranking for the query instance sent by each of the participant nodes in the current test group, wherein the feature distance ranking is a ranking of the size of the local feature distances between each of the data instances of the participant nodes and the query instance;
[0046] Based on the feature distance ranking of each participant node for the query instance, a candidate data instance set corresponding to the current test group is obtained,
[0047] The determining, based on the feature distance, from the set of instances with the same label, a set of K nearest neighbor instances corresponding to the query instance includes:
[0048] Determining a same-label candidate instance set of the current test group for the query instance based on an intersection of the candidate data instance set and the same-label instance set;
[0049] A set of K nearest neighbor instances of the current test group for the query instance is determined from the set of candidate instances with the same label based on the feature distance.
[0050] In some example embodiments, based on the above solution, obtaining the set of candidate data instances corresponding to the current test group based on the feature distance sorting of each of the participant nodes for the query instance includes:
[0051] Scanning the feature distance sorting of each participant node in the current test group to obtain a common data instance of each scanned participant node;
[0052] Determine the sum of the characteristic distances of the common data instances of the nodes of each participant to obtain a global characteristic distance corresponding to the common data instance;
[0053] A preset number of candidate data instances are determined based on the order of the global feature distances from largest to smallest.
[0054] In some example embodiments, based on the above solution, before obtaining the candidate data instance set, the method further includes:
[0055] Randomly sort the data instances of the participant nodes using the same random seed to generate pseudo-identities corresponding to the data instances.
[0056] After obtaining the candidate data instance set, the method further includes:
[0057] The pseudo identifier corresponding to each candidate data instance in the candidate data instance set is re-indexed into an original identifier.
[0058] In some example embodiments, based on the above solution, determining the distribution difference of the data distribution of the local data feature between the participant nodes includes:
[0059] Determining a preset divergence corresponding to the data distribution of the local data feature between each pair of the participant nodes;
[0060] A distribution difference of the data distribution between each pair of the participant nodes is determined based on the preset divergence.
[0061] In some example embodiments, based on the above solution, determining a preset divergence corresponding to the data distribution of the local data feature between each pair of participant nodes includes:
[0062] Determine the density of feature points of each participant node for each query instance in the query instance set;
[0063] Performing density aggregation on the feature point density of each query instance of the participant node to obtain a global density corresponding to the participant node;
[0064] Based on the global density corresponding to each pair of the participant nodes, the JS divergence between each pair of the participant nodes is determined.
[0065] In some example embodiments, based on the above solution, determining the density of feature points of each participant node for each query instance in the query instance set includes:
[0066] Determining a feature distance between the local data features of each of the data instances of the participant nodes and the query local features of the query instance;
[0067] Based on the feature distance and the hypersphere volume formula, the feature point density of the participant node for the query instance is determined.
[0068] In some example embodiments, based on the above solution, determining the feature point density of the participant node for the query instance based on the feature distance and the hypersphere volume formula includes:
[0069] Determine a set of K nearest neighbor instances of the participant node for the query instance based on the feature distance;
[0070] Determining a distance threshold corresponding to the query instance based on the feature distances corresponding to the data instances in the K nearest neighbor instance set, where the distance threshold is the maximum feature distance for the query instance in the K nearest neighbor instance set; and
[0071] Based on the distance threshold and the hypersphere volume formula, the feature point density of the participant node for the query instance is determined.
[0072] In a second aspect, this specification also provides an electronic device comprising: at least one storage medium storing at least one instruction set for performing participant selection processing; and at least one processor communicatively connected to the at least one storage medium, wherein, when the electronic device is running, the at least one processor reads the at least one instruction set and executes the participant selection method described in the first aspect of this specification according to the instructions of the at least one instruction set.
[0073] As can be seen from the above technical solutions, the participant selection method and device provided in the embodiments of this specification are applied to a distributed learning platform, which includes multiple participant nodes, each participant node includes local data features of each data instance in multiple data instances, and one participant node in the multiple participant nodes includes a data label of the data instance. On the one hand, the correlation between the local data features and data labels of the data instances of each participant node is determined, and the contribution of the participant data to the model training can be effectively evaluated through the correlation, guiding the feature selection or weight allocation in distributed learning; on the other hand, the distribution difference of the data distribution of the local data features of the data instances between the participant nodes is determined, which can help identify the diversity of the data distribution of the participant nodes and reduce the selection of participant nodes with similar data distribution; on the other hand, based on the correlation corresponding to the participant nodes and the distribution difference between the participant nodes, a predetermined number of target participant nodes are selected from the multiple participant nodes, which can accurately and efficiently select participant nodes with high correlation with the task label and diversified data distribution from the multiple participant nodes, avoiding the problem of insufficient data information of the participant nodes and excessive data redundancy, thereby obtaining a better participant selection result, and thus improving the efficiency of model training and model performance under the premise of protecting data privacy.
[0074] The following description partially outlines the features of the participant selection method, participant selection method, and device provided in this specification. The following figures and examples will provide insights that will be readily apparent to those skilled in the art. The inventive aspects of the participant selection method, participant selection method, and device provided in this specification can be fully demonstrated through practice or use of the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0076] Figure 1 A schematic diagram showing an implementation environment of a participant selection method provided in an embodiment of this specification;
[0077] Figure 2 FIG2 shows a hardware structure diagram of an electronic device 200 provided according to an embodiment of this specification;
[0078] Figure 3 A schematic diagram illustrating a process of a participant selection method provided according to some embodiments of this specification is shown;
[0079] Figure 4 A schematic diagram of a process for performing group testing according to some embodiments of this specification is shown;
[0080] Figure 5 A schematic diagram showing a distributed learning platform provided according to an embodiment of this specification; and
[0081] Figure 6 A flowchart of a participant selection method provided according to other embodiments of this specification is shown. DETAILED DESCRIPTION
[0082] The following description provides specific application scenarios and requirements for this specification, with the goal of enabling those skilled in the art to make and use the contents of this specification. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of this specification. Therefore, this specification is not limited to the embodiments shown, but is intended to be accorded the broadest scope consistent with the claims.
[0083] The terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. For example, as used herein, the singular forms "a," "an," and "the" may also include the plural forms unless the context clearly indicates otherwise. When used in this specification, the terms "comprise," "include," and / or "contain" are intended to refer to the presence of the associated integers, steps, operations, elements, and / or components, but do not preclude the presence of one or more other features, integers, steps, operations, elements, components, and / or groups or the addition of other features, integers, steps, operations, elements, components, and / or groups in the system / method.
[0084] These and other features of this specification, as well as the operation and function of the associated elements of the structure, and the economical assembly and manufacture of the components, can be significantly improved with consideration of the following description. Reference is made to the accompanying drawings, all of which form a part of this specification. However, it should be expressly understood that the drawings are for illustration and description purposes only and are not intended to limit the scope of this specification. It should also be understood that the drawings are not drawn to scale.
[0085] The flowcharts used in this specification illustrate operations implemented by systems according to some embodiments of the present specification. It should be clearly understood that the operations of the flowcharts may not be implemented in sequence. Rather, the operations may be implemented in reverse order or simultaneously. Furthermore, one or more additional operations may be added to the flowcharts. One or more operations may be removed from the flowcharts.
[0086] It should be noted that the user data obtained in this manual is authorized by the user and does not involve user privacy.
[0087] First, the terms involved in one or more embodiments of this specification are explained.
[0088] Privacy-preserving computing: Privacy-preserving computing is a broad category of technologies designed to enable data sharing, computation, and modeling among multiple data owners without exposing the original data, while ensuring data privacy and security. These technologies typically utilize encryption, differential privacy, and homomorphic encryption to safeguard data during computation.
[0089] Homomorphic encryption: Homomorphic encryption is an encryption technique that allows mathematical operations to be performed directly on encrypted data, with the decrypted result being the same as the result of the operation on the original data. For example, additive homomorphic encryption allows addition operations to be performed on encrypted data, with the decrypted result being the same as the result of the addition operation on the original data.
[0090] Mutual Information (MI): A statistical measure of the dependency between two random variables, quantifying the amount of information shared between them. If two variables are completely independent, their MI is zero; if two variables are completely correlated, their MI is large. Mutual Information is often used for feature selection and assessing correlation between variables.
[0091] Jensen-Shannon Divergence (JS Divergence): JS Divergence is a statistical measure of the difference between two probability distributions. It is based on the Kullback-Leibler (KL) Divergence, but is symmetric to make it fairer than the KL Divergence. JS Divergence is used to measure the similarity between probability distributions and is commonly used in information theory and machine learning.
[0092] Fagin's algorithm: This efficient sorting algorithm is used to merge multiple sorted lists and is particularly well-suited for top-k queries. It quickly finds the top k elements that meet the criteria by partially scanning multiple sorted lists. By optimizing the sorting process, the Fagin algorithm improves query efficiency and is suitable for solving top-k queries in large datasets.
[0093] Submodularity: Submodularity is an important property of set functions that describes their "diminishing marginal returns." That is, as the size of the set increases, the gain in the function value from adding new elements decreases. The submodularity property states that the gain from adding an element is greater in small sets and decreases as the set grows. If the objective function of a problem is submodular, a greedy algorithm can usually obtain an optimal solution.
[0094] Federated Learning: A machine learning technology that allows multiple participants to jointly train a machine learning model without sharing the original data. In federated learning, the original data remains local to each participant's device. Through encryption technology and communication protocols, participants only exchange model parameters or intermediate results. This enables joint data utilization and collaborative model optimization while protecting data privacy and security.
[0095] In related technical solutions, on a distributed machine learning platform such as a federated learning platform, multiple clients train machine learning models using local data. Therefore, the selection of clients is crucial to the efficiency and performance of model training. However, the client selection in the above technical solution has the following problems: (1) Insufficient information problem: The data of some clients may have low relevance to the task label. Selecting these clients for training will reduce model performance and result in the inability to obtain a better client selection result; (2) Data redundancy problem: The data distribution of the selected clients may be highly similar, resulting in information redundancy during the training process and reducing the efficiency of model training.
[0096] Based on the above content, the embodiments of this specification provide a participant selection method and electronic device, which are applied to a distributed learning platform, wherein the distributed learning platform includes multiple participant nodes, each participant node includes local data features of each data instance in multiple data instances, and one participant node in the multiple participant nodes includes a data label of the data instance. On the one hand, the correlation between the local data features and data labels of the data instances of each participant node is determined, and the contribution of the participant data to model training can be effectively evaluated through the correlation, guiding feature selection or weight allocation in distributed learning; on the other hand, the distribution difference of the data distribution of the local data features of the data instances between the participant nodes is determined, which can help identify the diversity of the data distribution of the participant nodes and reduce the selection of participant nodes with similar data distribution; on the other hand, based on the correlation corresponding to the participant nodes and the distribution difference between the participant nodes, a predetermined number of target participant nodes are selected from the multiple participant nodes, which can accurately and efficiently select participant nodes with high correlation with the task label and diversified data distribution from the multiple participant nodes, avoiding the problem of insufficient data information and excessive data redundancy of the participant nodes, thereby obtaining a better participant selection result, and thus improving the efficiency and performance of model training under the premise of protecting data privacy.
[0097] The technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings.
[0098] Figure 1 A schematic diagram showing an implementation environment of a participant selection method provided in an embodiment of this specification is shown.
[0099] See also Figure 1 As shown, the implementation environment 100 may include multiple participant nodes 110 and an aggregation node 120. Each participant node 110 contains partial data features, i.e., local data features, of a data instance from multiple data instances. The leading participant node among the multiple participant nodes 110 contains the data tag of the data instance. The aggregation node 120 is used to aggregate the encrypted data items sent by the multiple participant nodes 100.
[0100] The participant node 110 is connected to the aggregation node 120 via a wireless network or a wired network. The participant node 110 can be a computing node of a distributed system, or a mobile phone, tablet computer, laptop computer, desktop computer, etc., but is not limited thereto.
[0101] The participant node 110 may store data or instructions for executing the participant selection method described in this specification. The participant node 110 may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device to work.
[0102] Aggregator node 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, distribution networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms. Aggregator node 120 provides background services for applications running on participant nodes 110.
[0103] Aggregator node 120 is equipped with an integrated development platform (IDE). An IDE, also known as an integrated development environment (IDE), is an application program that provides a program development environment. It typically includes tools such as a code editor, compiler, debugger, and human-computer interface. Developers can write program code (i.e., develop programs) on the IDE. The IDE server can be a computing device specifically used by the IDE to implement the participant selection method. Aggregator node 120 can communicate data with participant nodes 110 and a database.
[0104] Furthermore, the aggregation node 120 may store data or instructions for executing the participant selection method described herein. The aggregation node 120 may include a hardware device with data processing capabilities and the necessary programs to drive the hardware device. Of course, the aggregation node 120 may also be solely a hardware device with data processing capabilities, or simply a program running on the hardware device.
[0105] The database may store data and / or instructions. In some embodiments, the database may store local data features of a data instance. In some embodiments, the database may store data and / or instructions executed by or used by the aggregator node 120 to execute the participant selection method described in this specification. The participant node 110 and the aggregator node 120 have access to the database of the aggregator node 120, and the participant node 110 and the aggregator node 120 may access the data or instructions stored in the database via a network. In some embodiments, the database may be directly connected to the participant node 110 and the aggregator node 120. In some embodiments, the database may be part of the participant node 110 or the aggregator node 120. In some embodiments, the database may include mass storage, removable storage, volatile read-write memory, read-only memory (ROM), or the like, or any combination thereof. Exemplary mass storage may include non-transitory storage media such as magnetic disks, optical disks, solid-state drives, etc. Example removable storage may include flash drives, floppy disks, optical disks, memory cards, zip disks, magnetic tape, etc. Typical volatile read-write memory may include random access memory (RAM). Example RAMs may include dynamic RAM (DRAM), double date rate synchronous dynamic RAM (DDRSDRAM), static RAM (SRAM), thyristor RAM (T-RAM), zero capacitance RAM (Z-RAM), etc. Example ROMs may include mask ROM (MROM), programmable ROM (PROM), virtually programmable ROM (PEROM), electronically programmable ROM (EEPROM), compact disc ROM (CD ROM), digital versatile disk ROM, etc.
[0106] Those skilled in the art will appreciate that the number of participant nodes 110 may be greater or lesser. For example, the number of participant nodes may be three, or the number of participant nodes may be dozens, hundreds, or even greater. In this case, the implementation environment may also include other participant nodes. The embodiments of this specification do not limit the number or device type of the participant nodes.
[0107] After introducing the implementation environment of the embodiments of this specification, the application scenarios of the embodiments of this specification will be introduced in conjunction with the above implementation environment. In the following description, the client node is also the participant node 110 in the above implementation environment, and the aggregation server is also the aggregation node 120 in the above implementation environment. The technical solutions provided by the embodiments of this specification can be applied to distributed machine learning scenarios of various machine learning models, such as large medical models, large financial models, and large recommendation models.
[0108] Taking the distributed learning platform of the medical big model in which the technical solution provided in the embodiment of this specification is applied as an example, the distributed learning platform includes multiple participant nodes, such as multiple medical institution nodes. Each participant node contains local data features of each data instance in multiple data instances. The leading participant node among the multiple participant nodes includes data labels of the data instances, such as a label indicating whether the patient has a certain disease. The leading participant node 110 determines the correlation between the local data features and the data labels of each medical institution node; determines the distribution differences of the data distribution of the local data features between the medical institution nodes; and selects a predetermined number of target medical institution nodes from the multiple medical institution nodes based on the corresponding correlations of the medical institution nodes and the distribution differences between the medical institution nodes.
[0109] It should be noted that the above is explained using the example of the technical solution provided in the embodiments of this specification applied in the distributed learning platform of the medical big model. The technical solution provided in the embodiments of this specification can also be applied to other appropriate distributed learning scenarios of machine learning models, such as distributed learning scenarios of financial big models or recommendation big models, etc. The implementation process belongs to the same inventive concept as the above description and will not be repeated here.
[0110] It should be noted that the steps in the participant selection method in the example embodiment of this specification can be partially executed by the client, partially executed by the server, or all executed by the server or all executed by the client, and this specification does not specifically limit this.
[0111] based on Figure 1 The implementation environment shown below will be combined with Figure 2-Figure 6 , a detailed introduction to the participant selection method and electronic device provided in the embodiments of this specification is provided. It should be noted that the aforementioned implementation environment is provided only to facilitate understanding of the spirit and principles of this specification, and the embodiments of this specification are not limited in this respect. Rather, the embodiments of this specification can be applied to any applicable scenario.
[0112] Figure 2 This is a schematic diagram of the structure of an electronic device 200 provided according to some embodiments of this specification. The electronic device 200 can execute the participant selection method described in this specification. The participant selection method is introduced in other parts of this specification. The electronic device 200 can be a general-purpose computer or a special-purpose computer. For example, the electronic device 200 can be a server, a personal computer, a portable computer (such as a notebook computer, a tablet computer, etc.), or other electronic devices with computing capabilities. Of course, the electronic device can be Figure 1 The intermediate participant node 110 or the aggregation node 120 may also be a participant node device used by multiple developers to develop programs on an integrated development platform.
[0113] The electronic device in this specification may include one or more of the following components: a processor 210 , a memory 220 , an input device 230 , an output device 240 , and a bus 250 . The processor 210 , the memory 220 , the input device 230 , and the output device 240 may be connected via the bus 250 .
[0114] The processor 210 may include one or more processing cores. The processor 210 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 220, as well as accesses data stored in the memory 220, to perform the participant selection method or participant selection method described in this specification. Optionally, the processor 210 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 210 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and applications; the GPU is responsible for rendering and drawing display content or training machine learning models; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 210 and may be implemented separately via a communications chip.
[0115] The memory 220 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 220 includes a non-transitory computer-readable storage medium. The memory 220 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 220 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system, including a system deeply developed based on the IOS system, or other systems.
[0116] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0117] The input device 230 is used to receive input commands or data and includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch-sensitive device. The output device 240 is used to output commands or data and includes, but is not limited to, a display device and a speaker. In one example, the input device 230 and the output device 240 may be combined, and the input device 230 and the output device 240 may be a touch-sensitive display.
[0118] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which will not be described in detail here.
[0119] Figure 3A flowchart of a participant selection method provided according to an embodiment of this specification is shown. As previously described, electronic device 200 can execute the participant selection method of the embodiment of this specification. Specifically, processor 210 can read an instruction set stored in its local storage medium and then execute the participant selection method of the embodiment of this specification according to the instructions set. Steps S310 to S330 of the participant selection method will be described in detail below with reference to the accompanying drawings.
[0120] like Figure 3 As shown, in step S310, the correlation between the local data features and data labels of each participant node is determined.
[0121] In this example embodiment, each participant node is a computing node of a distributed learning platform, which may be a federated learning platform. Each participant node contains partial features, i.e., local data features, of each of the multiple data instances. The leading participant node among the multiple participant nodes contains the data labels of the data instances. Let P be the participant node set P = {P1, P2, …, PM} of the multiple participant nodes, where P1 is the leading participant. Each participant Pi holds partial features of multiple data instances, which may be shared by multiple participant nodes.
[0122] For example, let's consider a participant node, P = {P1, P2, ..., PM}, representing multiple medical institutions, such as hospitals. Each medical institution holds different patient characteristics. For example, hospital P1 contains patient characteristic X1, representing blood sugar levels; hospital P2 contains patient characteristic X2, representing blood pressure; and hospital P3 contains patient characteristics, such as cholesterol. P1 is the lead medical institution, which contains data labels for patients' diseases. For example, Y = 1 indicates diabetes, and Y = 0 indicates health. Multiple data instances represent patients shared by multiple medical institutions.
[0123] It should be noted that although the participating node is a medical institution as an example, ordinary technicians in this field should understand that the participating node can also be other appropriate nodes such as financial platform nodes or shopping platform nodes, which is also within the scope of the embodiments of this specification.
[0124] Mutual information is a statistical metric that measures the dependency between two random variables and quantifies the amount of information shared between the two variables. In an exemplary embodiment, electronic device 200 determines the mutual information between the local data features and data labels of each participant node, and uses this mutual information to determine the correlation between the local data features and data labels of each participant node. The mutual information between the local data feature X and the data label Y can be expressed as follows:
[0125] I(X;Y)=H(X)+H(Y)-H(X,Y) (1)
[0126] Here, I(X;Y) represents mutual information, X represents the local data features of the data instance of each participating node, Y represents the data label of the data instance, H(X) and H(Y) represent entropy, and H(X, Y) represents the joint entropy. A larger mutual information value indicates a stronger correlation between X and Y.
[0127] According to the technical solution in the above embodiment, by calculating mutual information, each participant can quantify the correlation between local features and labels, guide feature selection or weight distribution in distributed learning, and thus improve model performance.
[0128] In some example embodiments, the electronic device 200 calculates the mutual information between the local data features and data labels of each participant node using the above formula (1). For example, the electronic device 200 counts the probability distribution of each local data feature and data label of the participant node, such as the joint distribution and the marginal distribution, calculates H(X), H(Y) and H(X, Y) based on the above probability distribution, and substitutes them into the above formula (1) to obtain the mutual information between the local data features and data labels of the participant node.
[0129] It should be noted that although the correlation between local data features and data labels is described using mutual information as an example, ordinary technicians in this field should understand that other appropriate indicators can also be used to quantify the correlation between data features and data labels, such as the maximum information coefficient or Kendall rank correlation coefficient, which is also within the scope of the embodiments of this specification.
[0130] In other example embodiments, the mutual information between the local data features and data labels of the participant nodes is determined based on a K-nearest neighbor approach. The electronic device 200 obtains the feature distances between each data instance of the participant node and the query instance in the query instance set; based on the feature distances between the data instance of the participant node and the query instance, the mutual information of each participant node is determined using a K-nearest neighbor approach. In the case of a participant node, the feature distance represents the distance between the local data features of each data instance of the participant node and the query local features of the query instance; in the case of several participant nodes in a test group, the feature distance may represent the global feature distance obtained by aggregating the local feature distances of each participant node in the test group for the target data instance and the target query instance.
[0131] Assume that N data instances constitute a data instance set D = {(x i ,y i)|i∈{1,2,…,N}}, the data instances in the data instance set D are data instances shared by multiple participating nodes, and the data elements in the data instance set D contain all the data features of the data instances, where xi is the data feature and yi is the data label. The query instance set can be a subset of the data instances collected from the data instance set D, namely the query instance set Q. The electronic device 200 determines a set of instances with the same label as the query instance and the corresponding number of instances with the same label from N data instances; determines a set of K nearest neighbor instances corresponding to the query instance from the set of instances with the same label based on the above-mentioned feature distance; determines a number m of global nearest neighbors corresponding to the query instance from multiple data instances based on the above-mentioned feature distance and the set of K nearest neighbor instances. q ; and based on the total number of instances N, the number of instances with the same label N q , K nearest neighbors and global nearest neighbors m q The mutual information of each participating node is determined by a K-nearest neighbor mutual information function, where the K-nearest neighbor mutual information function may be a bi-gamma function.
[0132] For example, let the total number of instances of the data set D be N and the number of K nearest neighbors be K. The electronic device 200 determines the label y of the query instance. q The data instances in the data instance set D with the same label are obtained as the instance set with the same label and the number of instances with the same label N q , that is: N q ={(x,y)∈D:y=y q}, based on the above feature distance, determine the K nearest neighbor instance set corresponding to the query instance from the instance set with the same label The K nearest neighbor instance set contains K nearest neighbor data instances with the same label as the query instance. Further, the electronic device 200 determines the distance threshold corresponding to the query instance based on the above feature distance and the K nearest neighbor instance set, where the distance threshold is the maximum feature distance of the query instance in the K nearest neighbor instance set; and determines the global nearest neighbor number m based on the number of instances in the N data instances whose feature distance to the query instance is less than the distance threshold. q .set up is the set of K nearest neighbor instances K nearest neighbor instances and query instance x q The distance threshold between them is the maximum feature distance; calculate the number of global neighbors m q , that is, with the query instance x q The feature distance between them is less than the distance threshold The number of data instances in the dataset D; based on the total number of instances N, the number of instances with the same label N q , K nearest neighbors and global nearest neighbors m q The mutual information MI of each participating node is determined by the following formula (2).
[0133] MI(X;Y)=ψ(N)-ψ(N q )+ψ(K)-ψ(m q ) (2)
[0134] Among them, X is the local data feature of the participating node, Y is the data label, N is the total number of instances in the data instance set D, Nq is the number of instances with the same label in the instance set, K is the number of K nearest neighbors in the instance set; m q is the number of global neighbors, and ψ() is the double gamma function. The first term ψ(N) represents the entropy correction of the global data scale, which is used for benchmark adjustment and reflects the overall amount of information; the second term -ψ(N q ) represents the entropy correction of the local category, which is used to balance the influence of the sample size of different categories; ψ(K) represents the entropy correction term of the K nearest neighbors, which is used to control the influence of the neighbor selection on the resolution; -ψ(m q ) represents the entropy correction term for the number of global neighbors, which is used to quantify the contrast between local and global distributions.
[0135] According to the technical solution in the above-mentioned example embodiment, on the one hand, by combining the entropy correction of global data and local data, the statistical dependence between local data features and data labels can be accurately captured; on the other hand, the double gamma function ψ(·) is used to estimate the mutual information, thereby improving the accuracy and efficiency of the mutual information calculation.
[0136] Assume that the participant node is a medical institution, there are 1000 patient instances, 300 of which are diabetic. The local data feature of the participant node is blood sugar level. The data label Y is 1 for diabetic patients and Y is 0 for healthy patients. Analyze the correlation between blood sugar level and diabetes. Assume that the query instance is diabetic patient L. The electronic device 200 determines the label y related to diabetic patient L. q The data instances in the data instance set D with the same label are obtained, and a set of 300 diabetes instances and 300 instances with the same label are obtained; from the 300 diabetes instance sets, a set of 5 nearest neighbor instances is found. For example, the 300 diabetes instances are sorted in order of feature distance from small to large, and the first 5 nearest neighbor instances are taken to obtain a set of 5 nearest neighbor instances. The global feature distances in the 5 nearest neighbor instance sets are set to 0.2, 0.3, 0.5, 0.6 and 0.7, and the maximum feature distance between the blood glucose level features of the data instances in the 5 nearest neighbor instance sets and the blood glucose level features of patient L is recorded. For example, 0.7 (the characteristic distance of the 5th nearest neighbor); among 1000 patients, the characteristic distance between the blood glucose level feature of patient L is less than The number of neighbors m qFor example, 50; Based on the above formula (2), the correlation between type 1 diabetes and blood glucose level of patient L is calculated, that is, mutual information MI(blood glucose level; 1) = ψ(1000)-ψ(300)+ψ(5)-ψ(50). The larger the mutual information MI value, the stronger the correlation between the feature and the label.
[0137] According to the technical solution in the above example embodiment, the distance threshold is determined by the K-nearest neighbor method, which avoids the tedious feature distance calculation that needs to be calculated for each instance, thereby improving data processing efficiency.
[0138] In some example embodiments, the electronic device 200 determines the number of instances based on the total number of instances, the number of instances with the same label Nq, the number of K nearest neighbors, and m q , the first mutual information of each participant node for each query instance is determined by the K-nearest neighbor mutual information function; the mutual information of the participant nodes is determined based on the average of the first mutual information of each query instance in the query instance set. For example, the first mutual information of each participant node for each query instance is determined by the following formula (3); the mutual information of the participant nodes is determined based on the average of the first mutual information of multiple query instances.
[0139]
[0140] Where Q is the query instance set, and |Q| represents the number of instances in the query instance set. For example, suppose that after calculating the mutual information (MI) for 50 patients, the average mutual information (MI) is 0.45, indicating that there is a moderate correlation between blood glucose level and diabetes type.
[0141] According to the technical solution in the above example embodiment, the final mutual information of the participant nodes is determined by averaging the mutual information of the participant nodes for multiple query instances, which can improve the accuracy of the mutual information of the participant nodes.
[0142] In step S320 , the distribution difference of the data distribution of the local data features between the participant nodes is determined.
[0143] In an example embodiment, in a distributed learning scenario, each participant node has its own local data set, and the local data set contains local data features of each data instance in multiple data instances. The distribution of data features of the local data sets of the participant nodes is different. This distribution difference will affect the training effect of the model, and the distribution difference of local data features of different participant nodes can be measured by a preset divergence. The electronic device 200 determines the preset divergence corresponding to the data distribution of the local data features between each pair of participant nodes, for example, KL divergence or JS divergence; and determines the distribution difference of the data distribution between each pair of participant nodes based on the preset divergence. Among them, the KL divergence is used to measure the difference between two probability distributions, and the JS divergence is a symmetric KL divergence.
[0144] It should be noted that although the distribution difference is described as KL divergence or JS divergence, ordinary technicians in this field should understand that the distribution difference can also be other appropriate indicators such as Hellinger characteristic distance or total variation, which is also within the scope of the embodiments of this specification.
[0145] Furthermore, the electronic device 200 determines the density of feature points of each participant node for each query instance in the query instance set; performs density aggregation on the density of feature points of each participant node for each query instance to obtain a global density corresponding to the participant node, for example, performs feature splicing on the density of feature points of each query instance to obtain a global density corresponding to the participant node; and determines the JS divergence between each pair of participant nodes based on the global density corresponding to each pair of participant nodes. For example, the JS divergence between each pair of participant nodes is determined by the following formula (4):
[0146]
[0147] Among them, ρ Si represents the global density of participant node i, ρ Sj represents the global density of participant node j, and M represents the global density ρ Si and the global density ρ Sj The mean value, D KL Denotes KL divergence. D KL The calculation of divergence is shown in the following formula (5):
[0148]
[0149] Where k represents the kth density interval, ρ Si (k) represents the global density of the i-th participant node in the k-th density interval, and M(k) represents the average density of participant node i and participant node j in the k-th density interval.
[0150] Assuming that the participant node is a medical institution, the distribution differences of the data distribution of the local data features of multiple medical institutions are analyzed, and the electronic device 200 determines the feature point density of each query instance in the query instance set for the local data instances of medical institutions i and medical institutions j; density aggregation is performed on the feature point density corresponding to each query instance to obtain the global density corresponding to medical institutions i and medical institutions j; based on the global density corresponding to medical institutions i and medical institutions j, the JS divergence between medical institutions i and medical institutions j is determined by the above formula (4).
[0151] According to the technical solutions in the above example embodiments, by selecting an appropriate distribution difference metric such as JS divergence, the distribution differences of local data features of each participant in distributed learning can be effectively evaluated.
[0152] Furthermore, the electronic device 200 calculates the JS divergence between each pair of participant nodes to obtain the JS divergence matrix JS corresponding to the multiple participant nodes. The node scores of the participant nodes can be updated in the following steps using the JS divergence matrix corresponding to the multiple participant nodes. The JS divergence matrix is shown in the following formula (6):
[0153]
[0154] According to the technical solution in the above-mentioned example embodiment, by introducing the JS divergence matrix, in the participant selection process, not only the importance score is considered, but also the distribution difference of the data distribution between the participants is considered, thereby ensuring the diversity of the data of the selected participant nodes.
[0155] In step S330, a predetermined number of target participant nodes are selected from the plurality of participant nodes based on the correlations corresponding to the participant nodes and the distribution differences between the participant nodes.
[0156] In an exemplary embodiment, the correlation corresponding to a participant node represents the degree of association between the data features of the participant node and the data label, which can effectively assess the contribution of the participant data to model training. For example, the correlation between the local data features and data labels of each participant node is determined through mutual information. The distribution differences of the data distributions between the participant nodes can help identify the diversity of the data distributions of the participant nodes. For example, the distribution differences of the local data features of different participant nodes can be measured using a preset divergence, such as the KL divergence or the JS divergence.
[0157] In some example embodiments, the electronic device 200 determines the comprehensive score of the participant node based on the correlation corresponding to the participant node and the distribution difference between the participant nodes, and selects a predetermined number of target participant nodes from the multiple participant nodes in descending order of the comprehensive score of the participant nodes. For example, the electronic device 200 performs a weighted operation on the mutual information and preset divergence of the participant nodes based on the preset correlation weight and distribution difference weight to determine the comprehensive score of the participant node. For example, if the task requires a high-precision model, the correlation weight is increased, focusing on the correlation between data features and data labels; if the task requires improved generalization ability, the distribution difference weight is increased to increase the difference between data distributions.
[0158] In other example embodiments, the electronic device 200 determines a node score for each participant node based on the corresponding correlation of each participant node; updates the node score for each participant node based on the distribution difference between the participant nodes; and selects a target participant node from the plurality of participant nodes based on the updated node score. For example, the participant node with the highest node score among the plurality of participant nodes may be selected as the current participant node; and the node score of the other participant node may be updated based on the distribution difference between the current participant node and the other participant node.
[0159] For example, for each round of iterative update: the electronic device 200 selects the participant node with the highest node score among multiple participant nodes as the current participant node; and adds the current participant node to the target participant node set until the number of elements in the target participant node set reaches a predetermined number.
[0160] Assume the initialized set of participant nodes U (0) =U is the initial scoring vector of the participant nodes, each element of the initial scoring vector is the mutual information value of the participant nodes, and K is the preset number of participant nodes to be selected. Each round of iterative update performs the following iterative operation until the number of elements in the target participant node set I reaches the predetermined number K:
[0161] (1) Select the participant node i with the highest current node score through the following formula (7):
[0162]
[0163] (2) By using the following formula (8), the participant node i * Join the target participant node set I:
[0164] I←I∪i * (8)
[0165] (3) Update the node scores of the remaining participating nodes through the following formula (9):
[0166]
[0167] in, represents the updated node score of participant node j, a is the experience weight, represents the JS divergence matrix, such as the JS weight matrix in equation (6), Represents the node score of participant node j before the update.
[0168] According to the technical solution in the above-mentioned example embodiment, on the one hand, an optimization objective with submodularity is constructed so that the participant selection problem can be solved by a greedy algorithm, thereby reducing the computational complexity while ensuring the quality of selection and improving the scalability of the algorithm; on the other hand, through a greedy strategy, a set of optimized participant nodes is gradually constructed, and a better participant can be selected from multiple participant nodes.
[0169] according to Figure 3 The technical solution in the example embodiment, on the one hand, determines the correlation between the local data features of the data instances of each participant node and the data label, can effectively evaluate the contribution of the participant data to the model training through the correlation, and guide the feature selection or weight allocation in distributed learning; on the other hand, determines the distribution difference of the data distribution of the local data features of the data instances between the participant nodes, can help identify the diversity of the data distribution of the participant nodes, and reduce the selection of participant nodes with similar data distribution; on the other hand, based on the corresponding correlation of the participant nodes and the distribution difference between the participant nodes, a predetermined number of target participant nodes are selected from multiple participant nodes, which can accurately and efficiently give priority to the participant nodes with high correlation with the task label and diversified data distribution from multiple participant nodes, avoiding the problems of insufficient data information and excessive data redundancy of the participant nodes, thereby obtaining a better participant selection result, and further improving the efficiency of model training and model performance under the premise of protecting data privacy.
[0170] In addition, in some example embodiments, the electronic device 200 determines the feature distance between the local data features of each data instance of the participant node and the query local features of the query instance; based on the feature distance and the hypersphere volume formula, determines the feature point density of the participant node for the query instance.
[0171] For example, the electronic device 200 determines the K nearest neighbor instance set of the participating node for the query instance based on the feature distance; determines the distance threshold corresponding to the query instance based on the feature distance corresponding to each data instance in the K nearest neighbor instance set, and the distance threshold is the maximum feature distance for the query instance in the K nearest neighbor instance set; and determines the feature point density of the participating node for the query instance based on the distance threshold and the hypersphere volume formula.
[0172] For example, the electronic device 200 calculates the Euclidean feature distance between the query instance q in the query instance set and each data instance x in the data instance set, and finds the nearest K neighbors, namely the K nearest neighbor instance set. The distance threshold is the maximum feature distance d for the query instance in the K nearest neighbor instance set. q,K , the distance threshold is determined by the following formula (10):
[0173] d q,K=max(||x i -x q ||),x i ∈N K (x q ) (10)
[0174] Among them, d q,K represents the Euclidean feature distance of the Kth nearest neighbor, N K (x j ) are the first K nearest neighbors in the dataset.
[0175] Furthermore, the electronic device 200 determines the data feature point x of the participant node for the query instance q by the following formula (11) based on the distance threshold and the hypersphere volume formula: q The feature point density ρ q :
[0176]
[0177] Among them, V is the volume of the hypersphere, K is the number of K nearest neighbors, and D is the characteristic dimension of the data feature (i.e. x i dimension); Γ(·) is the Gamma function: Γ(n) = (n-1)!. The volume V of the hypersphere is determined by the following equation (12):
[0178]
[0179] The electronic device 200 calculates the corresponding feature point density value for each query instance q in the query instance set Q in turn, and obtains a density value list ρ i ={ρ q |q∈Q}, uploaded to the aggregation server for subsequent calculations.
[0180] According to the technical solution in the above exemplary embodiment, the density of feature points is estimated by the K-nearest neighbor method, which avoids the complexity of directly calculating high-dimensional data and improves data processing efficiency.
[0181] Figure 4 A schematic diagram of a process for performing group testing according to some embodiments of this specification is shown.
[0182] Reference Figure 4 As shown, in step S410, a predetermined number of test groups are generated based on multiple participant nodes, and each test group includes several participant nodes.
[0183] In an exemplary embodiment, the target participant nodes are gradually screened out by grouping the test participant nodes. The electronic device 200 generates a predetermined number of test groups based on multiple participant nodes, each test group including several participant nodes. For example, the electronic device 200 generates T test groups G = {S1, S2, ..., S N}, where each test group S t Contains several participating nodes,
[0184] Furthermore, before determining the group mutual information corresponding to each test group, the electronic device 200 obtains the feature distance ranking for the query instance sent by each participant node in the current test group through the aggregation node. The feature distance ranking is the size ranking of the local feature distances between each data instance of the participant node and the query instance; based on the feature distance ranking of each participant node for the query instance, the candidate data instance set corresponding to the current test group is obtained.
[0185] For example, the electronic device 200 scans the feature distance ranking of each participant node in the current test group to obtain the common data instance of each participant node that has been scanned; determines the sum of the feature distances of the common data instance of each participant node to obtain the global feature distance corresponding to the common data instance; and determines a preset number of candidate data instances based on the order of the global feature distance from large to small. For example, the electronic device 200 runs the Fagin algorithm on the sorted list of feature distance rankings of each participant node in the current test group, and obtains the sorted list of candidate data instances I through the following formula (13): k ′:
[0186] I′ k,t =Fagin({I′ i,t |P i,t ∈F}) (13)
[0187] Among them, I′ k Represents the participant node P of the current test group i,t The corresponding ranked list of candidate data instances, P i,t represents the participant node Pi of the current test group, F represents the node set of multiple participant nodes, I′ i,t Indicates the feature distance ranking corresponding to the participant node Pi.
[0188] According to the technical solution in the above exemplary embodiment, the K nearest neighbor candidates corresponding to the current test group are efficiently found through the Fagin algorithm, thereby reducing tedious distance calculation and communication overhead in mutual information calculation.
[0189] In step S420, the group mutual information corresponding to each test group is determined, where the group mutual information represents the mutual information between the local data features and data labels of several participant nodes included in the test group.
[0190] In an exemplary embodiment, the electronic device 200 determines the group mutual information corresponding to the current test group based on the feature distance between the local data features and the data labels of each participant node included in the current test group. For example, the electronic device 200 determines the same-label candidate instance set for the query instance of the current test group based on the intersection of the above-mentioned candidate data instance set and the above-mentioned same-label instance set, where the candidate data instance set is a set of candidate data instances obtained based on the Fagin algorithm, and the same-label instance set is a set of instances with the same data label as the query instance among N data instances; determines the K nearest neighbor instance set for the query instance of the current test group from the same-label candidate instance set based on the above-mentioned feature distance; determines the global nearest neighbor number m corresponding to the query instance from multiple data instances based on the above-mentioned feature distance and the K nearest neighbor instance set. q ; and based on the total number of instances N, the number of instances with the same label Nq, the number of K nearest neighbors, and the number of global neighbors m q The group mutual information of the current test group is determined by the following formula (14).
[0191] For example, the group mutual information corresponding to the test group St is tested by the following formula (10):
[0192]
[0193] Where ψ(·) is the digamma function, which is expressed as: St represents the current test group, [X t,1 ,X t,2 ,…] represents the participating nodes The corresponding local data features, Q represents the query instance set.
[0194] In step S430, the mutual information corresponding to each participant node is determined based on the average value of the group mutual information of each test group in which the participant node participates.
[0195] In an exemplary embodiment, the electronic device 200 determines each target test group in which the participant node participates; obtains group mutual information of each target test group; averages the group mutual information of each target test group, and determines the average as the mutual information corresponding to the participant node.
[0196] For each test group S t , estimate the mutual information of the test group corresponding to the group through the above formula (14), and use it as the score U(S t ) until the traversal of T test groups is completed. For each participant node P i, calculate the mutual information corresponding to the participant node, i.e., the node score u, through the following formula (15): i , which is the average score of all test groups in which the participant participated:
[0197]
[0198] Among them, u i represents the node score of the participant node Pi, Ti represents the test group in which the participant node Pi participates, and U(St) represents the group mutual information corresponding to the test group St. The leading participant obtains the node score U of multiple participant nodes = {u i |P i ∈P}.
[0199] According to the technical solution in the above example embodiment, by calculating the average value through group testing, there is no need to calculate the mutual information of each participating node separately and completely, thereby improving data calculation efficiency and reducing communication costs.
[0200] Furthermore, the distributed learning platform also includes an aggregation node. Before determining the group mutual information corresponding to each test group, the electronic device 200 obtains the local feature distance between the data instance sent by each participant node in the current test group and the query instance through the aggregation node. The local feature distance represents the distance between the local data feature of the data instance of the participant node and the query local feature of the query instance; the local feature distance of each participant node in the current test group for the query instance is aggregated through the aggregation node to obtain the global feature distance of the participant node of the current test group for the query instance.
[0201] For example, the electronic device 200 obtains a current query instance from a query instance set, wherein the query features of the current query instance include multiple query local features; determines the local data features of the data instances of each participant node in the current test group and the local feature distances between the query local features of the current query instance; and performs an addition operation on the local feature distances of each participant node in the current test group for the query instance through an aggregation node to obtain a global feature distance of each participant node for the query instance. For example, the local feature distances of each participant node for the query instance are aggregated using the following formula (16):
[0202]
[0203] in, represents the global feature distance of the query instance after aggregation, HE.Sum(·) represents the addition operation of the encrypted local distance, d i,q,tIt represents the local feature distance corresponding to the participant node i of the current test group, Pi represents the participant node i, F represents the set of participant nodes of the federated learning platform, and pk represents the key.
[0204] According to the technical solution in the above example embodiment, by aggregating the local feature distances of each participant node, the global feature distance between the data instance and the query instance can be obtained, so that each participant node of the test group can be screened from the global feature perspective.
[0205] Figure 5 A schematic diagram of a distributed learning platform provided according to an embodiment of this specification is shown.
[0206] Reference Figure 5 As shown, the distributed learning platform includes: multiple participant nodes 110, an aggregation node 120, and a key server 130. Each participant node 110 contains partial data features of a data instance in N data instances, namely local data features, and the leading participant node P1 includes the data label of the data instance. The key server 130 generates a homomorphic encryption public key and private key (pk, sk), assigns the private key sk to the leading participant node P1, and assigns the public key pk to each participant node and the aggregation server. The aggregation server 120 provides aggregation calculations for the data items of the participant nodes, such as the addition operation HE.Sum, for securely aggregating multiple encrypted data items sent by the participant nodes: [d] = HE.Sum({[di]}; pk). The leading participant node P1 determines the correlation between the local data features and data labels of each participant node 110; determines the distribution differences of the data distribution of the local data features between the participant nodes 110; and based on the corresponding correlations of the participant nodes 110 and the distribution differences between the participant nodes 110, selects a predetermined number of target participant nodes from multiple participant nodes 110.
[0207] Figure 6 A flowchart of a participant selection method according to some other example embodiments of the present specification is shown.
[0208] Reference Figure 6 As shown, in step S610, initialization is performed.
[0209] In an example embodiment, key server 130 generates homomorphic encryption key pairs and distributes them to participant nodes and aggregator nodes. Each participant node shuffles its local data and generates pseudo-IDs for each data instance. The lead participant node generates a test group and query instance set for evaluating candidate participant nodes, i.e., clients.
[0210] For example, suppose P is the set of all participants and P1 is the leading participant: P = {P1, P2, ..., PM}, the local data features held by each participant node correspond to X1, X2, ...X M N data instances are data instance sets D consisting of N data pairs: D = {(x i ,y i )|i∈{1,2,…,N}}, where xi is the global data feature corresponding to data instance i, and yi is the data label corresponding to data instance i; the leading participant node collects M data instances from the data instance set D and generates a query instance set Q of size M:
[0211]
[0212] in, represents the global data feature corresponding to the query instance i, Represents the data label corresponding to the query instance i.
[0213] The key server 130 generates a pair of homomorphically encrypted public and private keys using the following formula (17):
[0214] (pk,sk)=HE.KeyGen() (17)
[0215] The key server 130 sends the public key pk to the aggregation node 120; sends the private key sk to each participant node 110;
[0216] Each participant node 110 shuffles its local data using the same random seed to generate a pseudo ID for the data instance, denoted as I′;
[0217] The leading participant node generates T test groups G = {S1, S2, ..., S N}, where each test group element S t Involving several parties,
[0218] In step S620 , mutual information calculation is performed.
[0219] In this example embodiment, each participant node 110 calculates and encrypts local feature distances, then sends the encrypted local feature distances along with the sorted list to the aggregator node 120. Aggregator node 120 aggregates the encrypted local feature distances to obtain a global feature distance, which it then sends to the leading participant P1, generating a list of candidate instances using the Fagin algorithm. Each participant node 110 restores the pseudo-ID. Leading participant node P1 decrypts the global feature distances and calculates mutual information based on the decrypted global feature distances. The implementation process for calculating mutual information is described in detail below, in conjunction with steps S6210 to S6240.
[0220] In step S6210, feature distance calculation of local features is performed.
[0221] In the example embodiment, the current test group S t Each participant node in calculates the feature distance between the local local data feature and the query local feature of the query instance q in the query instance set Q. Assume that the participant P i The participating node calculates the feature distance d between the local feature corresponding to the query instance q and the local feature xi of the data instance j through the following formula (18): i :
[0222] d i,q,t =[(x q,i -x j,i ) 2 for(x j ,y j )∈D] (18)
[0223] Among them, the data feature x j,i is the local data feature of the participant node Pi corresponding to the data instance j in the data instance set D, x q,i Represents the local features corresponding to the query instance q and the local data features xi.
[0224] The participant node 110 sorts the local feature distances of multiple data instances for the query instance and generates a feature distance sorting I′ for the query instance. i,q :
[0225] I′ i,q,t =argsort(d i,q )
[0226] The participant node 110 uses homomorphic encryption to encrypt the local feature distance and obtains the encrypted feature distance [d i ]:
[0227] [d i,t ]=HE.Enc(d i ;pk)
[0228] The participant node 110 performs the above process on each query instance q in the query instance set in turn, and obtains the local feature distance set and ranking list of the participant nodes in the current test group for each query instance:
[0229] d i,t ={[d i,q,t ]|q∈Q}
[0230] I′ i,t ={I′ i,q,t |q∈Q}
[0231] The participant node 110 will use the encrypted local feature distance set d i,t and sorted list I′ i,t Sent to the aggregation server.
[0232] According to the technical solution in the above-mentioned example embodiment, on the one hand, through homomorphic encryption technology, the participating nodes do not need to expose the original data when calculating and transmitting feature distance data, thereby ensuring the privacy of the data; on the other hand, by sorting the local feature distances between each data instance and the query instance, each participating node can independently calculate and sort the local feature distances, thereby reducing communication overhead.
[0233] In step S6220, global feature distance aggregation and candidate generation are performed.
[0234] In the example embodiment, the aggregation node 120 performs an addition operation on the encrypted local feature distances of each participant node for the query instance through the above formula (16), and obtains the encrypted global feature distance for the query instance q: The aggregation node 120 performs the above operations on each query instance q in turn, and obtains the set The aggregation node 120 runs the Fagin algorithm on the sorted list and obtains the candidate instance sorted list I′ of the candidate data instance through the above formula (13): k ; Sort the candidate instances into a list I′ k,t Send to each participant node 110. The participant node 110 sorts the candidate instances into a list I′ k,t The pseudo IDs corresponding to the data instances in are re-indexed to the original IDs.
[0235] According to the technical solution in the above example embodiment, on the one hand, the Fagin algorithm is used to efficiently find the k-nearest neighbor candidates corresponding to the query instance, reducing the computational and communication overhead; on the other hand, under the premise of protecting privacy, the pseudo ID of the data instance is remapped to the original ID, ensuring the correctness of the data.
[0236] In step S6230, group mutual information calculation is performed.
[0237] In an exemplary embodiment, the leading participant node P1 decrypts the global feature distance returned by the aggregation node 120 using a private key:
[0238]
[0239] The candidate instance sorting list I′ returned by the leading participant node P1 to the aggregation node 120 k,t , mapping it back to the original IDI k,t, calculate the group mutual information of the participating nodes of the test group for the data labels through the above formula (14):
[0240] The leading participant node P1 for each test group S t , estimate the mutual information of the test group and use it as the score U(S t ) until T test groups are traversed.
[0241] According to the technical solution in the above exemplary embodiment, the dual gamma function ψ(·) is used to estimate the mutual information, thereby improving the accuracy and efficiency of mutual information calculation.
[0242] In step S6240, participant scores are calculated.
[0243] The leading participant node P1 is responsible for each participant node P i , calculate its importance score u through the above formula (15) i , that is, the average score of all test groups in which the participant node participates, and obtain the score U of each participant node in the participant node set P = {u i |P i ∈P}.
[0244] In step S630 , divergence calculation is performed.
[0245] In this example embodiment, each participant node 110 calculates the density of feature points for each query instance in the query instance set and sends it to the aggregator node 120. Aggregator node 120 aggregates the density and sends it to the lead participant. The lead participant node calculates the JS divergence between different participant nodes. The implementation process of divergence calculation is described in detail below, in conjunction with steps S6310 to S6320.
[0246] In step S6310, local density calculation is performed.
[0247] The participant node 110 calculates the Euclidean feature distance between the query instance q in the query instance set and each data instance in the data instance set using the above formula (10), and finds the nearest K neighbors, namely the K nearest neighbor instance set. The distance threshold is the maximum feature distance d for the query instance in the K nearest neighbor instance set. q,K .
[0248] Furthermore, the participant node 110 determines the data feature point x of the participant node for the query instance q based on the distance threshold and the hypersphere volume formula through the following formula (11): q The feature point density ρ q :
[0249] The participant node 110 calculates the density value for each query instance q in turn, obtains the following density value list, and uploads the density value list to the aggregation node 120 for subsequent calculations.
[0250] In step S6320, density aggregation and divergence calculation are performed.
[0251] The aggregation node 120 performs density aggregation on the feature point density of each query instance of the participant node, for example, performs feature splicing on the feature point density of each query instance, obtains the global density ρ corresponding to the participant node, and sends the global density to the leading participant P1:
[0252] ρ F ={ρ i ;P i ∈F}
[0253] The leading participant node P1 is based on the density value corresponding to each pair of participant nodes i and j. and The Jensen-Shannon divergence between participant node i and participant node j is calculated using equations (4) and (5):
[0254] Finally, the leading participant node P1 calculates the divergence of each pair of participants and obtains the JS divergence matrix JS corresponding to multiple participant nodes.
[0255] According to the technical solution in the above example embodiment, the distribution differences between participants are quantified by JS divergence, which provides diversity guarantee for subsequent participant selection.
[0256] In step S640, participant node selection is performed.
[0257] In the example embodiment, the leading participant node P1 initializes the scoring of each participant node 110 based on the mutual information, obtains the first score of each participant node, and selects the participant node with the highest score to join the target participant node set; the first score of each participant node is updated based on the JS divergence matrix until a predetermined number of participant nodes are selected, as shown in the above formulas (7) to (9).
[0258] According to the technical solution in the above-mentioned example embodiment, on the one hand, by introducing the JS divergence matrix, not only the importance score but also the distribution differences between participants are considered in the participant selection process, thereby ensuring the diversity of selection; on the other hand, the participant nodes are selected through a greedy strategy, and a better set of participant nodes is gradually constructed.
[0259] Based on the above content, according to the technical solution of the embodiment of this specification, on the one hand, mutual information is used to measure the correlation between the data and labels of the participating nodes, ensuring that the selected data has a higher contribution to the model training; on the other hand, JS divergence is used to measure the differences in the data distribution of the participating nodes, ensuring that the selected data has sufficient diversity and improving the generalization ability of the model; on the other hand, the mutual information calculation is accelerated by the Fagin algorithm, the computational overhead is reduced, and the efficiency of the participating node selection is improved; on the other hand, an optimization objective with submodularity is constructed, so that the participating node selection problem can be solved by a greedy algorithm, while ensuring the selection quality, reducing the computational complexity and improving the scalability of the algorithm.
[0260] Another aspect of this specification provides a non-transitory storage medium storing at least one set of executable instructions for performing participant selection. When the executable instructions are executed by a processor, the executable instructions direct the processor to implement the steps of the participant selection method described herein. In some possible implementations, various aspects of this specification may also be implemented in the form of a program product comprising program code. When the program product is executed on an electronic device 200, the program code is used to cause the electronic device 200 to perform the steps of the participant selection method described herein. The program product for implementing the above method may utilize a portable compact disc read-only memory (CD-ROM) comprising the program code and may be executed on the electronic device 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system. The program product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of computer-readable storage media include: an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the electronic device 200, partially on the electronic device 200, as a stand-alone software package, partially on the electronic device 200 and partially on a remote computing device, or entirely on the remote computing device.
[0261] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0262] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0263] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0264] It should be understood that in the foregoing descriptions of the embodiments of this specification, to facilitate understanding of a feature and to simplify this specification, various features are combined in a single embodiment, figure, or description thereof. However, this does not necessarily mean that these features are combined. When reading this specification, a person skilled in the art may label some of the devices as separate embodiments. In other words, the embodiments of this specification can also be understood as the integration of multiple sub-embodiments. This also applies when each sub-embodiment contains fewer than all the features of a single previously disclosed embodiment.
[0265] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, articles, etc., cited herein is hereby incorporated by reference, except for any content of the same that appears in the relevant documents that may be inconsistent or conflicting with this document, or that may have a limiting effect on the broadest scope of the claims. For example, if there is any inconsistency or conflict between the description, definition, and / or use of terms associated with any incorporated material and the terminology, description, definition, and / or use associated with this document, the terminology in this document will control.
[0266] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A participant selection method, applied to a distributed learning platform, comprising a plurality of participant nodes, each of the participant nodes comprising local data features of a plurality of data instances, one of the plurality of participant nodes comprising a data label of the data instance, the method comprising: Determining the correlation between the local data features of each of the participant nodes and the data labels; determining distribution differences in data distribution of the local data features between the participant nodes; and Based on the correlation corresponding to the participant nodes and the distribution difference between the participant nodes, a predetermined number of target participant nodes are selected from the multiple participant nodes.
2. The method according to claim 1, wherein The selecting a predetermined number of target participant nodes from the plurality of participant nodes based on the correlation corresponding to the participant nodes and the distribution difference between the participant nodes includes: Determine a node score for each of the participant nodes based on the correlation corresponding to each of the participant nodes; Based on the distribution difference between the participant nodes, updating the node score of each of the participant nodes; A target participant node is selected from the plurality of participant nodes based on the updated node score.
3. The method according to claim 2, wherein: The updating of the node score of each of the participant nodes based on the distribution difference between the participant nodes includes: Selecting the participant node with the highest node score among the multiple participant nodes as the current participant node; The node score of another participant node is updated based on the distribution difference between the current participant node and the another participant node.
4. The method according to claim 3, wherein: The selecting a target participant node from the plurality of participant nodes based on the updated node score includes: For each round of iterative update: the current participant node is added to the target participant node set until the number of elements in the target participant node set reaches the predetermined number.
5. The method according to claim 1, wherein Determining the correlation between the local data features of each participant node and the data label includes: Determine the mutual information between the local data features of each of the participant nodes and the data label, where the mutual information is used to represent the correlation between the local data features of the participant nodes and the data label.
6. The method according to claim 5, wherein: The determining of the mutual information between the local data feature of each of the participant nodes and the data label includes: Obtaining a characteristic distance between each of the data instances of the participant node and a query instance in the query instance set; Based on the characteristic distance between the data instance of the participant node and the query instance, the mutual information of each participant node is determined by a K-nearest neighbor method.
7. The method according to claim 6, wherein: The determining, based on the characteristic distance between the data instance of the participant node and the query instance, the mutual information of each participant node by a K-nearest neighbor method includes: Determine, from the multiple data instances, a set of instances with the same label as the query instance and the corresponding number of instances with the same label; Determine a set of K nearest neighbor instances corresponding to the query instance from the set of instances with the same label based on the feature distance; Determining a number of global neighbors corresponding to the query instance from the plurality of data instances based on the feature distance and the set of K nearest neighbor instances; and The mutual information of each of the participant nodes is determined based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K nearest neighbors and the number of global neighbors through the K nearest neighbor mutual information function.
8. The method according to claim 7, wherein: The determining, from the plurality of data instances, the number of global neighbors corresponding to the query instance based on the feature distance and the K-nearest neighbor instance set includes: Determining a distance threshold corresponding to the query instance based on the feature distance and the K nearest neighbor instance set, the distance threshold being the maximum feature distance for the query instance in the K nearest neighbor instance set; and The number of global neighbors is determined based on the number of instances in the plurality of data instances whose feature distance to the query instance is less than the distance threshold.
9. The method according to claim 7, wherein: The determining the mutual information of each of the participant nodes by a K-nearest neighbor mutual information function based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K-nearest neighbors, and the number of global neighbors includes: Determining, based on the total number of instances of the multiple data instances, the number of instances with the same label, the number of K nearest neighbors, and the number of global nearest neighbors, a first mutual information function of each of the participant nodes for the query instance; The second mutual information of the participant node is determined based on an average value of the first mutual information corresponding to each of the query instances in the query instance set.
10. The method according to any one of claims 5 to 9, wherein The determining of the mutual information between the local data feature of each of the participant nodes and the data label includes: generating a predetermined number of test groups based on the plurality of participant nodes, each of the test groups including a number of the participant nodes; Determining group mutual information corresponding to each of the test groups, where the group mutual information represents mutual information between the local data features of the plurality of participant nodes included in the test group and the data labels; and Based on the average value of the group mutual information of each of the test groups in which the participant node participates, the mutual information corresponding to each of the participant nodes is determined.
11. The method according to claim 10, wherein: The distributed learning platform further includes an aggregation node. Before determining the group mutual information corresponding to each of the test groups, the method further includes: Obtaining, through the aggregation node, a local feature distance between the data instance sent by each participant node in the current test group and the query instance, where the local feature distance represents a distance between the local data feature of the data instance of the participant node and the query local feature of the query instance; The local feature distances of the participant nodes in the current test group for the query instance are aggregated by the aggregation node to obtain the global feature distances of the participant nodes in the current test group for the query instance.
12. The method according to claim 11, wherein The aggregating the local feature distances of the participant nodes in the current test group for the query instance to obtain the global feature distances of the participant nodes in the current test group for the query instance includes: An addition operation is performed on each local feature distance of each participant node in the current test group with respect to the query instance to obtain a global feature distance of each participant node in the current test group with respect to the query instance.
13. The method according to claim 11, wherein Before determining the group mutual information corresponding to each of the test groups, the method further includes: Obtaining, through the aggregation node, a feature distance ranking for the query instance sent by each of the participant nodes in the current test group, wherein the feature distance ranking is a ranking of the size of the local feature distances between each of the data instances of the participant nodes and the query instance; Based on the feature distance ranking of each participant node for the query instance, a candidate data instance set corresponding to the current test group is obtained, The determining, based on the feature distance, from the set of instances with the same label, a set of K nearest neighbor instances corresponding to the query instance includes: Determining a same-label candidate instance set of the current test group for the query instance based on an intersection of the candidate data instance set and the same-label instance set; A set of K nearest neighbor instances of the current test group for the query instance is determined from the set of candidate instances with the same label based on the feature distance.
14. The method according to claim 13, wherein The obtaining of a set of candidate data instances corresponding to the current test group based on the feature distance sorting of each of the participant nodes for the query instance includes: Scanning the feature distance sorting of each participant node in the current test group to obtain a common data instance of each scanned participant node; Determine the sum of the characteristic distances of the common data instances of the nodes of each participant to obtain a global characteristic distance corresponding to the common data instance; A preset number of candidate data instances are determined based on the order of the global feature distances from largest to smallest.
15. The method according to claim 13, wherein Before obtaining the candidate data instance set, the method further includes: Randomly sort the data instances of the participant nodes using the same random seed to generate pseudo-identities corresponding to the data instances. After obtaining the candidate data instance set, the method further includes: The pseudo identifier corresponding to each candidate data instance in the candidate data instance set is re-indexed into an original identifier.
16. The method according to claim 1, wherein Determining the distribution difference of the data distribution of the local data feature between the participant nodes includes: Determining a preset divergence corresponding to the data distribution of the local data feature between each pair of the participant nodes; A distribution difference of the data distribution between each pair of the participant nodes is determined based on the preset divergence.
17. The method according to claim 16, wherein The determining of a preset divergence corresponding to the data distribution of the local data feature between each pair of the participant nodes includes: Determine the density of feature points of each participant node for each query instance in the query instance set; Performing density aggregation on the feature point density of each query instance of the participant node to obtain a global density corresponding to the participant node; Based on the global density corresponding to each pair of the participant nodes, the JS divergence between each pair of the participant nodes is determined.
18. The method according to claim 16, wherein The determining of the density of feature points of each participant node for each query instance in the query instance set includes: Determining a feature distance between the local data features of each of the data instances of the participant nodes and the query local features of the query instance; Based on the feature distance and the hypersphere volume formula, the feature point density of the participant node for the query instance is determined.
19. The method according to claim 18, wherein The determining, based on the feature distance and the hypersphere volume formula, the feature point density of the participant node for the query instance includes: Determine a set of K nearest neighbor instances of the participant node for the query instance based on the feature distance; Determining a distance threshold corresponding to the query instance based on the feature distances corresponding to the data instances in the K nearest neighbor instance set, where the distance threshold is the maximum feature distance for the query instance in the K nearest neighbor instance set; and Based on the distance threshold and the hypersphere volume formula, the feature point density of the participant node for the query instance is determined.
20. An electronic device comprising: at least one storage medium storing at least one instruction set for performing a participant selection process; as well as at least one processor, in communication with the at least one storage medium; Wherein, when the electronic device is running, the at least one processor reads the at least one instruction set and executes the participant selection method as described in any one of claims 1-19 according to the instructions of the at least one instruction set.