System and method for performing longitudinal federated learning

By selecting the appropriate encoder and data source in longitudinal federated learning, customizing the encoder, combined with the optimized VFL configuration, the problem of inefficient processing of heterogeneous data features and overlapping features in the prior art is solved, and a more efficient training process is achieved.

CN120092428APending Publication Date: 2025-06-03HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101282.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing vertical federated learning (VFL) architectures have problems of inefficiency and high training costs when dealing with heterogeneous data features and overlapping features.

Method used

Selecting the appropriate encoder and data source through the encoder selector, the encoder customizer customizes the encoder structure by silencing part of the neuron or link, and the VFL configurator configures the evaluator and joint classifier to optimize data exchange and backpropagation routes.

Benefits of technology

The efficiency and training speed of vertical federated learning are improved, and the problems of relative increase in weights and slow convergence speed due to overlapping data features are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092428A_ABST
    Figure CN120092428A_ABST
Patent Text Reader

Abstract

The present invention relates to a method and a server for performing vertical federated learning (VFL) on a communication network, and more particularly, to a method and a server for performing VFL on a communication network. Some methods include encoder and data selection, encoder customization, and VFL configuration. The encoder and data selection includes: receiving a VFL request and analyzing the VFL request to determine a data requirement and an encoder requirement; selecting a first encoder and a second encoder based on the encoder requirements; a data source is selected for the first encoder and the second encoder based on the data requirement. The encoder customization comprises the following steps: selecting a first encoder and a second encoder; the first encoder is customized by muting a portion of the first encoder. The VFL configuration comprises: receiving indications of annotated data information and a loss function; configuring an evaluator using the annotation data information and the indication of the loss function; an encoder output configured to the joint classifier; and configuring the private router.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of federated learning, and in particular, to systems and methods for performing vertical federated learning. Background Art

[0002] In the telecommunications field, "6G" is the sixth-generation standard currently under development. The 6G standard is a wireless communication technology for supporting cellular data networks. Compared with previous standards such as 5G networks, 6G networks exhibit more heterogeneity and support applications beyond current mobile usage scenarios, such as virtual and augmented reality, ubiquitous instant communication, all-pervasive intelligence, and the Internet of Things (IoT). Mobile network operators can adopt a flexible and decentralized business model for 6G networks, supported by mobile edge computing, artificial intelligence (AI), short-packet communication, and blockchain technology, to achieve local spectrum authorization, spectrum sharing, infrastructure sharing, and intelligent automation management.

[0003] Federated learning (FL) is one of various machine learning (ML) techniques that can be used to realize 6G network capabilities. Broadly speaking, FL is an ML technique that trains one or more AI models using its local datasets among multiple decentralized participants (e.g., edge devices or servers) without raw data exchange. This approach contrasts with traditional "centralized" ML techniques, where all local datasets are uploaded to an executor or it is assumed that raw local data samples are shared among decentralized participants.

[0004] In other words, FL enables multiple participants to build a joint AI model without sharing data, thereby solving key issues such as data privacy, data security, data access rights, and heterogeneous data access. There are mainly two methods for performing FL: horizontal federated learning (HFL) and vertical federated learning (VFL).

[0005] Broadly speaking, HFL is applicable to datasets of multiple participants that share the same feature space but different ID spaces. The feature space refers to the set of attributes / dimensions of each data entity / sample in a dataset, while the ID space refers to the identifiers of each corresponding data entity. For example, in a bank dataset, the feature space can be (account balance, mortgage amount, credit score, etc.), and the ID space is the list of users of all accounts.

[0006] In contrast to HFL, VFL is applicable to datasets of multiple participants who share the same ID space but have different feature spaces. For example, Bell TM and Amazon TM have a large number of overlapping entities in the ID space (e.g., the same users in the real world), but different data features (e.g., Bell TM hosts mobile usage features, and Amazon TM hosts online shopping features). If one or both of these participants want to train a joint ML model to evaluate the credit score of a given user by considering the impacts of both mobile usage features and online shopping features simultaneously, VFL can be used to train the joint ML model without any leakage of raw data and local model information between the two participants.

[0007] Since the participants of VFL may belong to different business or application categories, their local ML models usually differ in structure and hyperparameters. Therefore, a better VFL architecture is needed to enable deployment and use for a large number of participants. Summary of the Invention

[0008] The objective of the technology of the present invention is to at least improve some inconveniences existing in the traditional VFL architecture.

[0009] Typical VFL Architecture

[0010] Referring to FIG. 1, a typical VFL architecture 100 is depicted, where a first participant has a first ML model 102 and a second participant has a second ML model 104. The first ML model 102 and the second ML model 104 together form a joint model 110.

[0011] In contrast to HFL, training and using the VFL architecture 100 may require two specific functions, namely, the exchange of intermediate results 106 between the ML models and the coordinator 108.

[0012] According to the VFL paradigm, in each training step and each usage step of the joint model 110, the intermediate outputs or calculation results of the first local ML model 102 and the second local ML model 104 may need to be encrypted first through homomorphic encryption, and then can be transmitted to the other local ML model among the first local ML model 102 and the second local ML model 104 for them to perform local calculations. Each local ML model among the first local ML model 102 and the second local ML model 104 participating in the intermediate result exchange process 106 does not decrypt the intermediate results from the other local ML model among the first local ML model 102 and the second local ML model 104, that is, they perform local calculations on homomorphically encrypted data.

[0013] To ensure the confidentiality of data during the training process, the coordinator 108 is required to perform VFL. For example, the coordinator 108 can be a third-party entity trusted by the first participant and the second participant. As a non-limiting example, the coordinator 108 can be embodied as a secure computing node. The coordinator 108 creates encryption pairs and sends the public keys to the first participant and the second participant to enable the encrypted intermediate result exchange process 106. Once the intermediate result exchange process 106 is completed for each training and use step of the federated model 110, the first participant and the second participant send their encrypted outputs and locally generated masks to the coordinator 108. The coordinator 208 decrypts their masked outputs with the private key, calculates the decrypted gradients and losses, and sends this information back to the corresponding participants. Each participant un-masks the gradients and losses from the coordinator 208 and thus updates its corresponding local ML model parameters.

[0014] Developers have recognized that VFL may require the support of the NET4AI architecture to provide native support for the training and use phases of AI-based computing services. Broadly speaking, NET4AI is a service-oriented architecture for 6G wireless systems that provides end-to-end support for AI applications from the deployment phase to the operation phase. The NET4AI architecture is generally described in the article titled "Nine Challenges in Artificial Intelligence and Wireless Communications for 6G" written by Wen Tong and Geoffry Ye Li, published in 2021, DOI 10.1109 / MWC.006.2100543, the entire content of which is incorporated herein by reference. Supporting VFL in the NET4AI architecture is a challenging task.

[0015] Developers have recognized that in existing VFL paradigms, VFL clients can access information indicating multiple other VFL participants, data characteristics of other VFL participants, datasets of other VFL participants, labeled data locations / ownership, and encoder structures of other VFL participants. However, in implementing an automated VFL system, VFL clients may not be able to access information about other VFL participants. Therefore, the developers of the present invention's technology have recognized the need for systems and methods that can enable the NET4AI architecture to select datasets and encoders with necessary characteristics for VFL clients.

[0016] Developers have also recognized that in the existing VFL paradigm, the labeled data is provided by a single VFL participant, and the loss function can be split without violating homomorphic encryption. Therefore, a privacy-preserving intermediate result exchange process can be performed to calculate the joint loss. However, in some cases, the labeled data may be at least partially owned by multiple participants. In addition, some loss functions may be different from the linear polynomial functions that ensure homomorphic encryption after calculation. Therefore, developers have recognized that it may be necessary to configure a joint classifier and an evaluator for different encoders. This can simplify the interaction between the local classifiers of different VFL participants, and / or can support loss functions that violate homomorphic encryption, and / or can allow the labeled data to be provided by multiple data sources.

[0017] Developers have also recognized that in the existing VFL paradigm, the encoder structure maintained by the corresponding AI enablers remains unchanged. In these solutions, there is a lack of procedures for customizing the encoder structure for specific scenarios, such as collaborative training and using multiple encoders to form a joint model. However, when applying VFL in the NET4AI architecture or other architectures, the training data of the selected participants may have overlapping data features.

[0018] For example, in a VFL model with a "bank" participant and a "mobile operator" participant, the data of these two participants may have a common feature called "mobile bill payment records". In this example, if the VFL joint model is formed by the original encoders of the two participants, then due to the repetition in the two model inputs, the weight of the feature "mobile bill payment records" in the joint output may increase significantly. In the same example, the weights of other features may decrease in the joint output without repetition.

[0019] Developers have recognized that although the problem of the relatively increased weight of overlapping features can be alleviated by setting strict convergence requirements, the overall training cost may increase significantly. Therefore, developers have recognized that customizing the encoder structure may be desirable to reduce drawbacks such as the relatively increased weight and slow convergence speed due to overlapping data features.

[0020] Typical Encoder-Classifier Model

[0021] Referring to FIG. 2, a simplified representation of a typical local ML model 200 is shown. For example, the local ML model 200 can be a deep neural network (DNN) executed by a corresponding computing service (e.g., a participant).

[0022] The local ML model 200 includes an encoder 202 and a classifier 204, and each of the encoder 202 and the classifier 204 includes multiple layers. As shown in Figure 2, layers 2 to 4 are sometimes referred to as "hidden" layers. The encoder 202 includes an input layer 206 and multiple hidden layers, while the classifier 204 includes the remaining hidden layers and an output layer.

[0023] The encoder 202 captures low-order features of the local ML model 200, and these features can be shared with other local ML models implemented in a similar manner to the local ML model 200 (e.g., local ML models with the goal of solving similar problems). The classifier 204 captures high-order features of the local ML model 200, and these features are typically target-specific. For the input layer 206 of the encoder 202, each neuron corresponds to a feature of the input data. The dimension of the input layer 206, i.e., the number of neurons in the input layer 206, is equal to the number of features associated with the input data of the local ML model 200.

[0024] The developers of the present invention's technology have recognized that when performing automated VFL, the encoder can be an encoder provided by the network, i.e., provided by the network and implemented by a network entity called an "AI enabler". It is conceivable that during automated VFL execution, the classifier can be at least one of the following: (i) a classifier provided by the customer, i.e., a local classifier provided by the participant corresponding to the ML enabler, and (ii) a classifier provided by the network, i.e., a joint classifier provided by the orchestrator in the NET4AI architecture or other architectures. It is conceivable that both the classifier provided by the customer and the classifier provided by the network can be implemented through the service functions of computing services without departing from the scope of the present invention's technology.

[0025] The developers of the present invention's technology have recognized that neural network (NN) pruning techniques can be used to reduce the size of the NN by "muting" specific neurons or links of the NN while maintaining the same or similar performance as the corresponding full-size NN. For any input, muting a neuron can be achieved by forcing its output to be equal to zero. Muting a link can be achieved by setting the weight value of the link to zero.

[0026] Generally, during NN pruning, the neurons of a given input layer are not muted. However, in some embodiments of the present invention's technology, at least some neurons of at least some input layers of the encoder of the corresponding VFL participant can be muted to reduce the adverse effects caused by overlapping data features. The process of changing the encoder structure by muting specific neurons or links of a given layer (possibly the input layer) can be called an encoder customization program, or encoder pruning.

[0027] In some embodiments of the present invention's technology, the VFL participants include an encoder, a local classifier (if available), and corresponding data sources. The data sources of the VFL participants can include multiple data sets implemented at different network locations. All data entities provided by the data sources of the VFL participants can have the same format and the same features. Since each encoder is associated with an application category, the data sources of the VFL participants can be associated with the same application category as the application category of the encoder of the VFL participants.

[0028] The present invention's technology

[0029] In a first broad aspect of the present invention's technology, a method is provided, which includes: an encoder selector receives a VFL request from a vertical federated learning (VFL) client; the encoder selector determines, according to the VFL request, data requirements and encoder requirements for solving a VFL task, the VFL task can be split into multiple subtasks, and the multiple subtasks at least include a first subtask and a second subtask; the encoder selector selects a first encoder and a second encoder from an encoder pool in a communication network based on the encoder requirements, the first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask; the encoder selector selects a first data source for the first encoder and a second data source for the second encoder based on the data requirements, the selected data sources are from a data source pool, the first data source includes the first feature set, and the second data source includes the second feature set.

[0030] In some embodiments of the method, the data requirements indicate the requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

[0031] In some embodiments of the method, the encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

[0032] In some embodiments of the method, the method further includes: the encoder selector splits the VFL task into the multiple subtasks.

[0033] In some embodiments of the method, the selecting the first encoder includes: the encoder selector determines the application category of the first subtask based on the encoder requirements, and the application category is stored in a memory associated with the first encoder.

[0034] In some embodiments of the method, the selecting of the first encoder includes: the encoder selector identifies the encoder pool having the first encoder structure; the encoder selector selects the first encoder from the encoder pool having the first encoder structure.

[0035] In some embodiments of the method, the selecting of the first data source and the second data source includes: the encoder selector sends a request for data source information to the data source manager, the request including the data requirements; the encoder selector receives the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data sources in the data source pool; the encoder selector uses the data source information to select the first data source for the first encoder and the second data source for the second encoder.

[0036] In some embodiments of the method, the first data source and the second data source are associated with corresponding IDs, and the method further includes: the encoder selector sends the corresponding IDs of the first data source and the second data source to the data source manager; the encoder selector receives the network locations of the first data source and the second data source from the data source manager.

[0037] In a second broad aspect of the present invention's technology, a method is provided, the method including: an encoder customizer customizes the first encoder by silencing a part of the first encoder, thereby generating a customized first encoder. The first encoder and the second encoder have been selected from an encoder pool in a communication network based on encoder requirements, the encoder requirements having been determined according to a vertical federated learning (VFL) request for solving a VFL task. The VFL task can be split into multiple subtasks, the multiple subtasks including at least a first subtask and a second subtask. The first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask.

[0038] In some embodiments of the method, the customization includes: determining overlapping features between the first feature set and the second feature set; silencing the part of the first encoder that processes the overlapping features.

[0039] In some embodiments of the method, the customization includes: the encoder customizer determines that the first encoder is to be customized.

[0040] In some embodiments of the method, the encoder customizer determining that the first encoder is to be customized includes: the encoder customizer determining that the first encoder is to be customized based on data feature weights that have been received from a referenced AI enabler.

[0041] In some embodiments of the method, the customizing of the first encoder includes: the encoder customizer silencing the portion of the first encoder that processes the overlapping features. The silencing includes silencing at least one of the neurons of the first encoder and the links of the first encoder.

[0042] In a third broad aspect of the present invention's technology, a method is provided, the method including: a vertical federated learning (VFL) configurator configuring at least one of an evaluator and a joint classifier; the VFL configurator configuring a privacy router to determine at least one of the following: a data exchange route between the joint classifier and the evaluator, an annotation data merging route between the evaluator and one or more data sources, a backpropagation route between the joint classifier and customized first and second encoders for backpropagating gradient values, wherein the joint classifier, the evaluator, and the privacy router are used to perform VFL.

[0043] In some embodiments of the method, the configuring of the evaluator includes: the VFL configurator receiving annotation data information from an encoder selector; the VFL configurator determining a loss function for performing VFL; the VFL configurator configuring the evaluator using the annotation data information and the loss function.

[0044] In some embodiments of the method, the configuring of the evaluator includes: the VFL configurator configuring a combination rule for annotation data and the loss function, the annotation data being sourced from data sources selected for the first and second encoders.

[0045] In some embodiments of the method, the configuring of the joint classifier includes: the VFL configurator receiving information indicating the output layer structures of the first and second encoders from an encoder customizer; the VFL configurator configuring the encoder outputs to the joint classifier based on the output layer structures of the first and second encoders.

[0046] In some embodiments of the method, the second encoder is a customized second encoder.

[0047] In a fourth broad aspect of the technology of the present invention, an encoder selector is provided, which is configured to: receive a VFL request from a vertical federated learning (VFL) client; according to the VFL request, determine the data requirements and encoder requirements for solving the VFL task, where the VFL task can be split into multiple subtasks, and the multiple subtasks at least include a first subtask and a second subtask; based on the encoder requirements, select a first encoder and a second encoder from an encoder pool in a communication network, where the first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask; based on the data requirements, select a first data source for the first encoder and a second data source for the second encoder, where the selected data sources are from a data source pool, the first data source includes the first feature set, and the second data source includes the second feature set.

[0048] In some embodiments of the encoder selector, the data requirements indicate the requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

[0049] In some embodiments of the encoder selector, the encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

[0050] In some embodiments of the encoder selector, the encoder selector is further configured to split the VFL task into the multiple subtasks.

[0051] In some embodiments of the encoder selector, the selection of the first encoder includes that the encoder selector is configured to: determine the application category of the first subtask based on the encoder requirements, where the application category is stored in a memory associated with the first encoder.

[0052] In some embodiments of the encoder selector, the selection of the first encoder includes that the encoder selector is configured to: identify the encoder pool having the first encoder structure; select the first encoder from the encoder pool having the first encoder structure.

[0053] In some embodiments of the encoder selector, selecting the first data source and the second data source includes the encoder selector being configured to: send a request for data source information to a data source manager, the request including the data requirements; receive the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data sources in the data source pool; use the data source information to select the first data source for the first encoder and the second data source for the second encoder.

[0054] In some embodiments of the encoder selector, the first data source and the second data source are associated with corresponding IDs, and the encoder selector is further configured to: send the corresponding IDs of the first data source and the second data source to the data source manager; receive the network locations of the first data source and the second data source from the data source manager.

[0055] In a fifth broad aspect of the present invention's technology, an encoder customizer is provided, the encoder customizer being configured to: customize the first encoder by silencing a part of the first encoder, thereby generating a customized first encoder. The first encoder and the second encoder have been selected from an encoder pool in a communication network based on encoder requirements that have been determined according to a vertical federated learning (VFL) request for solving a VFL task. The VFL task can be split into multiple subtasks, the multiple subtasks including at least a first subtask and a second subtask. The first encoder is configured to process a first feature set for solving the first subtask, and the second encoder is configured to process a second feature set for solving the second subtask.

[0056] In some embodiments of the encoder customizer, the customization includes the encoder customizer being configured to: determine overlapping features between the first feature set and the second feature set; silence the part of the first encoder that processes the overlapping features.

[0057] In some embodiments of the encoder customizer, the customization includes the encoder customizer being configured to: determine that the first encoder is to be customized.

[0058] In some embodiments of the encoder customizer, determining that the first encoder is to be customized includes the encoder customizer being configured to: determine that the first encoder is to be customized based on data feature weights received from a referenced AI enabler.

[0059] In some embodiments of the encoder customizer, customizing the first encoder includes the encoder customizer being configured to: silence the part of the first encoder that processes the overlapping features, where silencing includes silencing at least one of the neurons of the first encoder and the links of the first encoder.

[0060] In a sixth broad aspect of the present invention's technology, a vertical federated learning (VFL) configurator is provided, where the VFL configurator is configured to: configure at least one of an evaluator and a joint classifier; configure a privacy router to determine at least one of the following: a data exchange route between the joint classifier and the evaluator, an annotation data merging route between the evaluator and one or more data sources, a backpropagation route between the joint classifier and a customized first encoder and a second encoder for backpropagating gradient values, where the joint classifier, the evaluator, and the privacy router are configured to perform VFL.

[0061] In some embodiments of the VFL configurator, configuring the evaluator includes the VFL configurator being configured to: receive annotation data information from an encoder selector; determine a loss function for performing VFL; configure the evaluator using the annotation data information and the loss function.

[0062] In some embodiments of the VFL configurator, configuring the evaluator includes the VFL configurator being configured to: configure a combination rule for annotation data and the loss function, where the annotation data is sourced from data sources selected for the first encoder and the second encoder.

[0063] In some embodiments of the VFL configurator, configuring the joint classifier includes the VFL configurator being configured to: receive information indicating the output layer structures of the first encoder and the second encoder from an encoder customizer; configure the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

[0064] In some embodiments of the VFL configurator, the second encoder is a customized second encoder.

[0065] In a seventh broad aspect of the technology of the present invention, a method is provided, the method comprising: an encoder selector receiving a VFL request from a vertical federated learning (VFL) client; the encoder selector determining, according to the VFL request, data requirements and encoder requirements for solving a VFL task, the VFL task being splittable into a plurality of subtasks, the plurality of subtasks including at least a first subtask and a second subtask; the encoder selector selecting a first encoder and a second encoder from an encoder pool in a communication network based on the encoder requirements, the first encoder being used to process a first feature set for solving the first subtask, and the second encoder being used to process a second feature set for solving the second subtask; the encoder selector selecting a first data source for the first encoder and a second data source for the second encoder based on the data requirements, the selected data sources being from a data source pool, the first data source including the first feature set, and the second data source including the second feature set; an encoder customizer customizing the first encoder by silencing a part of the first encoder, thereby generating a customized encoder; a VFL configurator configuring at least one of an evaluator and a joint classifier; the VFL configurator configuring a privacy router to determine at least one of the following: a data exchange route between the joint classifier and the evaluator, an annotation data merging route between the evaluator and one or more of the selected data sources, a backpropagation route between the joint classifier and each of the customized first encoder and the second encoder for backpropagating gradient values; the joint classifier, the evaluator, and the privacy router being used to perform VFL.

[0066] In some embodiments of the method, the data requirements indicate requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

[0067] In some embodiments of the method, the encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

[0068] In some embodiments of the method, the method further comprises: the encoder selector splitting the VFL task into the plurality of subtasks.

[0069] In some embodiments of the method, the selecting of the first encoder comprises: the encoder selector determining an application category of the first subtask based on the encoder requirements, the application category being stored in a memory associated with the first encoder.

[0070] In some embodiments of the method, the selecting the first encoder includes: the encoder selector identifying the encoder pool having the first encoder structure; the encoder selector selecting the first encoder from the encoder pool having the first encoder structure.

[0071] In some embodiments of the method, the selecting the first data source and the second data source includes: the encoder selector sending a request for data source information to the data source manager, the request including the data requirements; the encoder selector receiving the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data sources in the data source pool; the encoder selector using the data source information to select the first data source for the first encoder and the second data source for the second encoder.

[0072] In some embodiments of the method, the first data source and the second data source are associated with corresponding IDs. The method further includes: the encoder selector sending the corresponding IDs of the first data source and the second data source to the data source manager; the encoder selector receiving the network locations of the first data source and the second data source from the data source manager.

[0073] In some embodiments of the method, the customization includes: determining the overlapping features between the first feature set and the second feature set; silencing the part of the first encoder that processes the overlapping features.

[0074] In some embodiments of the method, the customization includes: the encoder customizer determining that the first encoder is to be customized.

[0075] In some embodiments of the method, the encoder customizer determining that the first encoder is to be customized includes: the encoder customizer determining that the first encoder is to be customized based on data feature weights received from a referenced AI enabler.

[0076] In some embodiments of the method, the customizing the first encoder includes: the encoder customizer silencing the part of the first encoder that processes the overlapping features, the silencing including silencing at least one of the neurons of the first encoder and the links of the first encoder.

[0077] In some embodiments of the method, the configuring the evaluator includes: the VFL configurator receiving labeled data information from the encoder selector; the VFL configurator determining a loss function for performing VFL; the VFL configurator configuring the evaluator using the labeled data information and the loss function.

[0078] In some embodiments of the method, configuring the evaluator includes: the VFL configurator configures a combination rule of labeled data and the loss function, and the labeled data is derived from data sources selected for a first encoder and a second encoder.

[0079] In some embodiments of the method, configuring the joint classifier includes: the VFL configurator receives information indicating output layer structures of the first encoder and the second encoder from an encoder customizer; the VFL configurator configures encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

[0080] In some embodiments of the method, the second encoder is a customized second encoder.

[0081] In an eighth broad aspect of the present invention's technology, a system is provided. The system is configured to: an encoder selector receives a VFL request from a vertical federated learning (VFL) client; the encoder selector determines data requirements and encoder requirements for solving a VFL task according to the VFL request, the VFL task can be split into multiple subtasks, and the multiple subtasks at least include a first subtask and a second subtask; the encoder selector selects a first encoder and a second encoder from an encoder pool in a communication network based on the encoder requirements, the first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask; the encoder selector selects a first data source for the first encoder and a second data source for the second encoder based on the data requirements, the selected data sources are from a data source pool, the first data source includes the first feature set, and the second data source includes the second feature set; an encoder customizer customizes the first encoder by silencing a part of the first encoder to generate a customized encoder; a VFL configurator configures at least one of an evaluator and a joint classifier; the VFL configurator configures a privacy router to determine at least one of the following: a data exchange route between the joint classifier and the evaluator, a labeled data merging route between the evaluator and one or more of the selected data sources, a backpropagation route between the joint classifier and each of the customized first encoder and the second encoder for backpropagating gradient values; the joint classifier, the evaluator, and the privacy router are used to perform VFL.

[0082] In some embodiments of the system, the data requirement indicates a requirement for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

[0083] In some embodiments of the system, the encoder requirement indicates a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

[0084] In some embodiments of the system, the encoder selector is further configured to split the VFL task into the multiple subtasks.

[0085] In some embodiments of the system, selecting the first encoder includes the encoder selector being configured to: determine an application category of the first subtask based on the encoder requirement, where the application category is stored in a memory associated with the first encoder.

[0086] In some embodiments of the system, selecting the first encoder includes the encoder selector being configured to: identify an encoder pool having the first encoder structure; select the first encoder from the encoder pool having the first encoder structure.

[0087] In some embodiments of the system, selecting the first data source and the second data source includes the encoder selector being configured to: send a request for data source information to a data source manager, where the request includes the data requirement; receive the data source information from the data source manager, where the data source information indicates the quality and characteristics of data included in corresponding data sources in the data source pool; use the data source information to select the first data source for the first encoder and the second data source for the second encoder.

[0088] In some embodiments of the system, the first data source and the second data source are associated with corresponding IDs, and the encoder selector is further configured to: send the corresponding IDs of the first data source and the second data source to the data source manager; receive the network locations of the first data source and the second data source from the data source manager.

[0089] In some embodiments of the system, the customization includes the encoder customizer being configured to: determine overlapping features between the first feature set and the second feature set; silence the part of the first encoder that processes the overlapping features.

[0090] In some embodiments of the system, the customization includes the encoder customizer for: obtaining (i) the encoder structure parameters of the first encoder and the second encoder, (ii) the position of the target AI enabler of the customized encoder, (iii) the position of the referenced AI enabler of the trained encoder; sending a request for the data feature weights of the trained encoder to the referenced AI enabler; receiving the data feature weights from the referenced AI enabler; determining the overlapping features between the first feature set of the first encoder and the second feature set of the second encoder; determining that the first encoder is to be customized based on the data feature weights; silencing the part of the first encoder that processes the overlapping features, the silencing including silencing at least one of the neurons of the first encoder and the links of the first encoder.

[0091] In some embodiments of the system, configuring the evaluator includes the VFL configurator for: receiving labeled data information from the encoder selector; determining a loss function for performing VFL; configuring the evaluator using the labeled data information and the loss function.

[0092] In some embodiments of the system, configuring the evaluator includes the VFL configurator for: configuring the combination rule of the labeled data and the loss function, the labeled data being sourced from the data sources selected for the first encoder and the second encoder.

[0093] In some embodiments of the system, configuring the joint classifier includes the VFL configurator for: receiving information indicating the output layer structures of the first encoder and the second encoder from the encoder customizer; configuring the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

[0094] In some embodiments of the system, the second encoder is a customized second encoder.

[0095] Implementations of the techniques of the present invention all include at least one of the above objects and / or aspects, but not necessarily all of these objects and / or aspects. It should be understood that certain aspects of the techniques of the present invention are attempts to achieve the above objects, but may not meet that object, and / or may meet other objects not specifically set forth herein.

[0096] Other and / or alternative features, aspects, and advantages of the techniques of the present invention will be apparent from the following description, drawings, and appended claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Embodiments of the present invention are described below by way of example only and with reference to the drawings, in which:

[0098] Figure 1 shows a schematic representation of a typical VFL architecture with two participants.

[0099] Figure 2 shows a schematic representation of a typical local ML model.

[0100] Figure 3 A schematic representation of an automated VFL architecture as envisioned in some embodiments of the technology of the present invention is shown.

[0101] Figure 4 Shows Figure 3 A schematic representation of the privacy excavator of the automated VFL architecture of

[0102] Figure 5 Shows the Figure 3 Schematic representation of the encoder and data selection procedure performed by the encoder selector of the VFL architecture of

[0103] Figure 6 Shows the Figure 3 Schematic representation of the encoder customization procedure performed by the encoder customizer of the VFL architecture of

[0104] Figure 7 Shows Figure 6 Two customization modes of the encoder customization procedure of

[0105] Figure 8 Shows the Figure 3 Schematic representation of the VFL configuration procedure performed by the VFL configurator of the VFL architecture of

[0106] Figure 9 Shows three VFL frameworks, including traditional VFL, automated VFL, and spliced single NN.

[0107] Figure 10 Shows the performance comparison between traditional VFL applied to two test datasets and Figure 9 Automated VFL of

[0108] Figure 11 Shows the Figure 9 Performance comparison between automated VFL of Detailed Description

[0109] Broadly speaking, the sixth generation (6G) communication system is an end-to-end (E2E) system that supports AI-based services and applications in various contexts and scenarios. For example, 6G network components can natively integrate communication, computing, and sensing capabilities to facilitate the transition from centralized intelligence in the cloud to pervasive intelligence at the "deep" edge. This integration can be achieved through an architecture referred to herein as "network for AI" (NET4AI). The NET4AI architecture can support the provision of AI as a service (AIaaS). However, as will become clear from the description below, at least some aspects of the techniques of the present invention can be embodied in a variety of AIaaS architectures.

[0110] In the context of the techniques of the present invention, developers have developed methods and systems for automated vertical federated learning (VFL). Broadly speaking, VFL is a machine learning (ML) technique that uses heterogeneous AI models and datasets belonging to different participants to train a joint model. The developers of the techniques of the present invention have recognized that it may be beneficial to support automated VFL in the NET4AI architecture or other architectures.

[0111] According to some non-limiting embodiments of the techniques of the present invention, methods and systems for automated VFL are provided to enable a customer to use the data and computing resources of other customers, even if those other customers belong to different organizations or even industry sectors. As will be described in more detail below, VFL automation can be achieved using at least some of the following functions: encoder and data selection, encoder customization, and VFL configuration.

[0112] In other non-limiting embodiments of the techniques of the present invention, an automated VFL solution that supports AI model customization is provided. Broadly speaking, the customization of an AI model is a function of an automated VFL solution for improving the scalability of the AI model across heterogeneous devices and scenarios. As will be described in more detail below, AI model customization can be enabled through the encoder customization function.

[0113] Automated VFL architecture

[0114] Reference Figure 3 , a non-limiting embodiment of an automated VFL architecture 300 is described. In this embodiment, the automated VFL architecture 300 includes an application layer 302, a network layer 304, and a data layer 306.

[0115] Broadly speaking, the various layers of the automated VFL architecture 300 are part of the framework used to describe the functions of the network system implementing VFL. Different layers can be used to characterize the computing functions performed by one or more network nodes (e.g., servers) to at least partially support the interoperability between the hardware components and software components of the automated VFL architecture 300.

[0116] In this embodiment, the application layer 302 includes a VFL client 310. However, without departing from the scope of the technology of the present invention, the application layer 302 may include multiple VFL clients. In this embodiment, the VFL client 310 maintains a joint classifier 314 and is used to provide VFL requests, such as VFL request 312, to the components of the subsequent network 304.

[0117] In this embodiment, the data layer 306 includes multiple data sources 342 that are potentially used during automated VFL. The ID alignment function 344 can be enabled on the data layer 306. Broadly speaking, one or more data sources can be used as a logical network entity memory for storing and / or maintaining the necessary raw data (as VFL model inputs) and annotation data (as a benchmark for calculating training losses) for training the VFL model of the technology of the present invention. It is envisioned that the entire batch of raw data or annotation data required for training the VFL model can be provided by one or more data sources. In some embodiments, the data sources can be centrally deployed in a network entity or deployed in multiple distributed entities without departing from the scope of the technology of the present invention.

[0118] In the illustrated non-limiting embodiment, the network layer 304 maintains an AI hyperparameter optimizer 320, multiple encoders 328, an evaluator 330, and a privacy router 332 (forwarding plane). In the context of the technology of the present invention, the AI hyperparameter optimizer 320 is used to execute one or more computer-implemented programs for enabling the following functions: encoder and data selection, encoder customization, and VFL configuration. In the illustrated non-limiting embodiment, the AI hyperparameter optimizer 320 includes an encoder selector 322, an encoder customizer 324, and a VFL configurator 326 for enabling the above functions respectively. The components of the AI hyperparameter optimizer 320 will now be described in turn.

[0119] Encoder Selector

[0120] In the illustrated non - limiting embodiment, the encoder selector 322 is used to receive a VFL request 312 from the VFL client 310 through an interface between the application controller and the service manager (e.g., as defined in the NET4AI architecture). The VFL client 310 can be said to be a user of automated VFL, which provides the VFL request 312 that contains the VFL problem to be solved by the network layer 304 and maintains its own federated classifier 314 to compute the federated ML model output from multiple encoders 328. In some embodiments, the VFL client 310 can be one of the VFL participants with its associated local encoder and dataset. In other embodiments, the VFL client 310 may not have an associated local encoder or dataset and requests a federated model for solving one or more VFL problems, which can be regarded as VFL tasks.

[0121] In the illustrated non - limiting embodiment, the encoder selector 322 maps the VFL request 312 to the data requirements and encoder structure requirements that may be needed to solve the VFL problem described in the VFL request 312. The data requirements include "must - contain" data features and an indication of the labeled data requirements (e.g., the minimum amount of labeled data, the labeled data format, etc.), which may need to be provided by the data sources of the VFL participants. The data requirements indication is used to select data sources to solve the VFL problem.

[0122] The encoder requirements indication is the encoder structure that may be needed to provide a specific function for solving the VFL problem. For example, assuming the problem description is "identifying traffic violations from surveillance videos", an indication of the encoder structure "CNN" can be provided to indicate that the federated model must include a CNN encoder to process the video data. The encoder selector 322 can interact with the data source manager through multiple rounds of communication to complete the encoder and data source selection procedure. How to implement the encoder and data selection process in some embodiments of the present invention will be described in more detail below.

[0123] Encoder customizer

[0124] In the illustrated non - limiting embodiment, when selecting a set of encoders (with associated AI enablers) for the VFL problem, the encoder customizer 324 is used to "filter out" overlapping data features of the selected encoders and can determine the customization parameters for each encoder. To calculate the customization parameters for the corresponding encoder, the encoder customizer 324 can determine the weights of the corresponding data features in the input layer of the corresponding encoder. The calculation process requires interaction between the encoder customizer 324 and another AI enabler that maintains a trained encoder having the same structure as the corresponding encoder. In some cases, after encoder customization is implemented for the corresponding selected encoder, the corresponding AI enabler can feedback the customized encoder features to the encoder customizer 324, thus completing the encoder customization procedure.

[0125] VFL configuration

[0126] In the illustrated non - limiting embodiment, when the encoder customizer 324 provides the customized encoder features, the VFL configurator 326 configures the parameters for the joint classifier 314 and the evaluator 330, as well as the routing rules for data exchange between the joint classifier 314, the evaluator 330, and the multiple encoders 328. This configuration of the parameters and one or more router rules is referred to herein as the VFL configuration procedure.

[0127] Reference Figure 4 , the VFL configuration procedure involves the joint classifier 314, the evaluator 330, and the privacy router 332. In the illustrated non - limiting embodiment, the joint classifier 314 is implemented as a service function run by a third - party VFL customer 310 on the application server and / or by the network. The joint classifier 314 can perform the functions of intermediate result exchange and coordinator defined in the conventional VFL framework, such as calculating the output of the joint ML model from the outputs of multiple selected VFL encoders, interacting with the evaluator 330 to obtain the joint loss, and backpropagating the gradients of each encoder.

[0128] In some non - limiting embodiments of the present invention's technology, the joint classifier 314 can be pre - configured with a fixed structure. In these non - limiting embodiments, the VFL configurator 326 can configure the encoder outputs into the joint classifier 314 according to the customized encoder output layer features.

[0129] In the illustrated non - limiting embodiment, the evaluator 330 is implemented as a service function running on the network layer 304. The evaluator 330 receives the labeled data forwarded by the privacy router 332 from one or more data sources 308 and calculates the joint loss value by applying a specific loss function to compare the joint ML model output from the joint classifier 314 and the corresponding labeled data.

[0130] In some non - limiting embodiments, the VFL configurator 326 can configure the selected loss function and the labeled data format (including combination rules if the labeled data consists of data from multiple data sources) according to the labeled data information of the selected encoder and data source 308.

[0131] In the illustrated non - limiting embodiment, the privacy router 332 is a predefined NET4AI function. However, the privacy router 332 can be implemented as a predefined function with a different architecture without departing from the scope of the technology of the present invention. In automated VFL, the privacy router 332 enables secure data exchange between the joint classifier 314, the evaluator 330, the selected encoder, and the labeled data source 308.

[0132] In the illustrated non - limiting embodiment, the data connection function 412 can be performed by the privacy router 332 to cascade unbalanced data entities (e.g., in terms of the amount of data entities) associated with the same data ID during the training phase. The data union function 412 can provide an interface between the encoder output 410 and the joint classifier 314. Broadly speaking, the data connection function 412 combines multiple outputs from different encoders (each output can be in the form of a data batch) into a single joint output batch according to predefined combination rules (e.g., cascading all output batches into one), and then sends it to the joint classifier 314 for further training / inference. The privacy router 332 can also perform joint loss forwarding 416 between the joint classifier 314 and the evaluator 330, and labeled data forwarding 414 from the labeled data source 408 to the evaluator 330.

[0133] Now, various embodiments of the technology of the present invention that may contribute to at least some of the encoder and data selection function, encoder customization function, and VFL configuration function for implementing an automated VFL architecture will be described.

[0134] Encoder and Data Selection Embodiments

[0135] Reference Figure 5 shows a schematic diagram of an encoder and data selection procedure 500. In the illustrated non - limiting embodiment, the procedure starts at step 506, where the encoder selector 504 receives a VFL request 506 from the VFL client 502. It is contemplated that the VFL request 506 can be received by the encoder selector 504 in a manner similar to how the encoder selector 322 receives the VFL request 312 without departing from the scope of the technology of the present invention. As described above, the VFL request 506 contains a description of the VFL problem to be solved / the VFL problem for which the VFL is trained to solve.

[0136] Broadly speaking, the VFL problem is a general problem that can be "understood" by the encoder selector 504 to determine specific data requirements and / or encoder requirements. In those embodiments where the VFL client 502 partially or fully understands the detailed data requirements and encoder requirements, the VFL client 502 can directly indicate them in the VFL request 502 for the encoder selector 504 to process.

[0137] In other embodiments where the VFL client 502 only provides the general VFL problem in the VFL request 506, the program can proceed to step 508 to determine the data requirements and encoder requirements. To this end, the encoder selector 504 decouples or splits the VFL task into subtasks. For example, the subtasks include a first subtask and a second subtask. It can be understood that the number of subtasks can be more than two.

[0138] Performing a given subtask requires a list of necessary data characteristics and requirements for labeled data (e.g., content or quantity, data, etc.), which will be referred to as the necessary data requirements of the subtask hereinafter. Each subtask also corresponds to an application category associated with a specific encoder structure, which is called the necessary encoder requirement of the subtask.

[0139] For example, the VFL task "traffic violation recognition" can be decomposed into a combination of the following subtasks: (1) a subtask of recognizing traffic violations by camera recording, which requires video stream data characteristics and labeled traffic violation video segments (necessary data requirements) and a CNN encoder for video processing (necessary encoder requirement); (2) a subtask of recognizing traffic violations by RSU data, which requires RSU perception data characteristics and labeled data and an encoder structure for RSU data processing. It should be noted that the application category can be a predefined NET4AI concept. For example, it describes a set of applications / problems that can be solved by a specific encoder structure. Therefore, it can be said that the encoder structure can be stored in association with the corresponding application category.

[0140] In some embodiments, the VFL task can be implemented as a list of attributes representing the subtask ID and its importance (weight). The composition and attribute values of each VFL task can be maintained by the encoder selector 504 and learned from historical data and / or pre-existing knowledge. In the previous example of the "traffic violation recognition" task, the VFL task can be defined as [(1, 0.4); (2, 0.6)], where (1, 0.4) corresponds to the "camera recording" subtask with an ID equal to "1" and an importance equal to "0.4", and (2, 0.6) corresponds to the RSU subtask with an ID of "2" and an importance equal to "0.6".

[0141] In this embodiment, the program proceeds to step 508, where the encoder selector 504 selects one or more encoders for solving the VFL task. The one or more encoders can be selected from an encoder pool in the communication network. Since a given encoder structure corresponds to an application category, the encoder selector 504 can select the encoders for participating in the VFL according to the necessary encoder requirements of multiple subtasks. The number of encoders can be consistent with the number of subtasks. When the subtasks include a first subtask and a second subtask, the encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

[0142] In some embodiments, the encoder selector 504 can select encoders with a relatively simple structure to accelerate the convergence rate of the joint model training to ensure a predetermined effective threshold (e.g., classification / prediction accuracy).

[0143] In other embodiments, the encoder selector 504 may need to ensure that the set of selected encoders includes the necessary encoder structures required to solve all subtasks of the VFL task. For example, in the previously proposed "traffic violation recognition" task, the set of selected encoders may need to include at least a CNN for solving the "camera recording" subtask and a DNN for solving the "RSU data" subtask. The CNN can be regarded as the first encoder, and the DNN can be regarded as the second encoder. The first encoder is used to process the first feature set for solving the first subtask, and the second encoder is used to process the second feature set for solving the second subtask.

[0144] In the illustrated non-limiting embodiment, the program proceeds to step 510 to select the data for solving the VFL task. Given the selected encoders in step 508, the encoder selector 504 is used to determine the data sources for each selected encoder. The encoder selector 504 is used to select a first data source for solving the first subtask from the data source pool, where the first data source includes a first feature set, and a second data source for solving the second subtask from the data source pool, where the second data source includes a second feature set.

[0145] In those embodiments where all data sources are owned by the network, the encoder selector 504 can directly execute the data selection process based on the complete information of all available data sources. However, in other embodiments, when the available data sources are not owned by the network, the encoder selector 504 can interact with the data source manager to access the information of the available data sources (including data source location, data quality, data features that can be provided for each data source, etc.) for data selection.

[0146] For example, the data source manager can be a data analytics and management (DAM) function or entity that is designed to monitor data sources and provide data analysis capabilities for different network applications, including VFL applications.

[0147] In this example, the data selection process can be completed through the following set of interactions between the data source manager and the encoder selector 504. The first interaction can include the encoder selector 504 sending all necessary data requirements to the data source manager to request information about available data sources, and the data requirements are included in the request sent by the encoder selector 504 to the data source manager. The second interaction can include the data source manager providing feedback on the data quality and data characteristic information of all available data sources in the data source pool.

[0148] In response to the received information about available data sources, the encoder selector 504 can select detailed data sources for each selected encoder, as well as the data sources for the labeled data. During data source selection, the encoder selector 504 may need to ensure that (i) the set of selected data sources provides all the necessary data characteristics required for the subtasks, (ii) the selected data sources have a minimum number of overlapping characteristics, and (iii) the labeled data can be the original labeled data from one or more selected sources, or combined labeled data composed of labeled data with partial characteristics from multiple data sources.

[0149] In this example, the third interaction can include the encoder selector 504 sending all the selected data source IDs to the DAM tool. The fourth interaction can include the DAM tool providing feedback on the locations of all data sets including the selected data sources. It should be noted that each data source can be partially hosted by multiple data sets located at different network locations (e.g., servers, edge devices, etc.). In response to the received detailed data set locations, the encoder selector 504 can determine the locations of the AI enablers (AI enabler locations) for implementing the corresponding encoders, thereby minimizing the communication cost between the data sets and the AI enablers.

[0150] Encoder customization embodiments

[0151] Reference Figure 6 shows a schematic diagram of an encoder customization program 600. The customization program 600 can involve communication and data transfer between an encoder selector 602, an encoder customizer 604, a referenced AI enabler 606, and a target AI enabler 608. The encoder selector 602 can be implemented in a manner similar to Figure 3 the encoder selector 322 and / or Figure 5 the encoder selector 504. The encoder customizer 604 can be implemented in a manner similar to Figure 3implemented in the manner of the encoder customizer 324. The cited AI enabler 606 is a well-trained "full" encoder available on the network. The target AI enabler 608 is an encoder that will be "customized" according to at least some embodiments of the technology of the present invention.

[0152] In step 610, the encoder selector 602 sends the encoder selection result to the encoder customizer 604. For example, step 610 may occur after the above encoder and data selection procedures are completed. In some embodiments, the customization program 600 may be triggered by the encoder customizer 602 in response to receiving the encoder selection result.

[0153] The content of the encoder selection result is not restricted and may vary particularly according to various implementations of the technology of the present invention. However, in some embodiments, the content of the encoder selection result includes: the encoder structure parameters of the corresponding selected encoder (e.g., number of layers, size, etc.), the location of the target AI enabler (TE) of the selected encoder that will implement the corresponding customization, and the location of the cited AI enabler (RE) of the selected encoder that implements the corresponding well-trained encoder.

[0154] In Figure 6 In the non-limiting embodiment shown, the encoder selection result particularly includes the encoder structure parameters of the first encoder of the first subtask, the location of the corresponding TE 608, and the location of the corresponding RE 606.

[0155] Without wishing to be bound by any particular theory, developers have recognized that the purpose of involving RE 606 in the encoder customization program 600 is to calculate different weights of data features in a given selected encoder, and these weights may be the required parameters for making customization decisions. It is conceivable that the data features for which RE 606 is to calculate weights are data features derived from the first data source. In at least one case, since all AI enablers for different applications can be maintained by a network layer (e.g., the NET4AI network), the encoder selector can directly select the best-trained RE 606. However, to ensure security and privacy requirements, the encoder selector only "knows" the encoder structure in RE 606 (the same as the selector encoder structure), and the degree to which the encoder is trained (e.g., the number of times of training / use). The detailed values (biases / weights) of the neurons and links in RE 606 are not known to the encoder selector itself.

[0156] Continue Figure 6Description: At step 620, the encoder customizer 604 sends a request for data feature weights to the corresponding RE 606. At step 630, the encoder customizer 604 receives the calculated data feature weights. It should be noted that for a given encoder, the weight of each data feature in its input layer indicates the influence / importance of the corresponding data feature on the output of the given encoder. For example, the higher the weight of a given data feature, the more obvious the change in the encoder output can be observed when the value of the given data feature is changed.

[0157] Reference Figure 7 , an example of the structures of the first encoder 700 and the second encoder 710 selected by the encoder and the data selection program is shown. For example, the first encoder 700 has input data features (A, B, C, D) with weights (0.3, 0.4, 0.1, 0.2), and the second encoder 710 has input data features (A, B, E, F) with weights (0.1, 0.2, 0.4, 0.3). In this example, the input data features (A, B, C, D) are from the first data source, while the input data features (A, B, E, F) are from the second data source.

[0158] Developers of the technology of the present invention have recognized that there are various ways to calculate the weights of the corresponding data features of a customized encoder based on a well-trained encoder. However, due to security and privacy requirements, the encoder customizer 604 may not have access to the well-trained values of the neurons and links in the corresponding RE 606. Therefore, the calculation of the weights of the corresponding data features of a given selected encoder can be completed through the interaction between the encoder customizer 604 and the corresponding RE 606 during steps 620 and 630. Thus, the encoder customizer 604 sends a request for the data feature weights to the corresponding RE 606 (step 620), and the corresponding RE 606 locally calculates the data feature weights according to known techniques, for example, and then feeds back the calculated data feature weights to the encoder customizer 604 (step 630). For example, the calculated data feature weights include the data feature weights corresponding to each feature in the first feature set and the data feature weights corresponding to each feature in the second feature set, the first feature set being included in the first data source of the first encoder and the second feature set being included in the second data source of the second encoder. At least one such technique is disclosed in the article titled "Problems with Shapley-value-based explanations as feature importance measures" written by Kumar, I. Elizabeth, et al., published on June 30, 2020, the entire content of which is incorporated herein by reference. Other techniques may also be considered without departing from the scope of the technology of the present invention.

[0159] The customization program 600 proceeds to step 640, in which, given the calculated data feature weights of all the selected encoders calculated by the corresponding RE 606, the encoder customizer 604 calculates the encoder customization parameters, and the encoder customizer 604 is used to select "primary features" for the joint model being constructed.

[0160] In some embodiments, non-overlapping data features can be automatically selected as primary features. Broadly speaking, non-overlapping data features include data features that exist in their corresponding selected encoders but do not exist in other selected encoders. In Figure 7 the example shown, the non-overlapping features include features C, D, E, and F.

[0161] In other embodiments, for a given overlapping data feature, the encoder customizer 604 compares its data feature weights in the corresponding selected encoders and selects the feature with the highest relative weight as the primary feature. Broadly speaking, overlapping data features include data features that are repeated in at least two selected encoders. In Figure 7In the example shown, the overlapping features include features A and B.

[0162] For example, it should be noted that in the first encoder 700, features A and B are the main features. The relative weights are different from the data feature weights calculated by the RE. The calculation of the relative weights takes into account both the data feature weights of each RE and other factors, such as but not limited to: the size of the encoder structure (a large size means high complexity, which affects the output convergence performance), the dependence on other features, and the overall importance of the features. It should be noted that the relative weights are used to determine which data features to retain and which to delete in the first encoder 700 and the second encoder 710. In this example, it can be said that the relative weights are used to determine whether to delete features A and B from the first encoder 700 or the second encoder 710. Given the selected main features, the encoder customizer will delete all non-main data features from the corresponding selected encoder. By deleting the non-main features, the side effects caused by data feature overlap on the joint model can be effectively alleviated.

[0163] In some embodiments of the technology of the present invention, the encoder selector 604 can selectively execute three modes of encoder customization especially according to different application scenarios. For example, the first mode 720 can be called the minimum customization mode, in which only the input layer neurons corresponding to the non-main features in the selected encoder are deleted. This is an encoder customization mode with the least impact on the joint model structure.

[0164] In another example, the second mode can be called the weight-based customization mode 730, in which, for a given selected encoder from whose input layer the non-primitive features have been deleted, the encoder customizer 604 sends the IDs of the neurons to be deleted to the corresponding RE of the encoder with the corresponding well-trained encoder. The corresponding RE will delete all the neurons and links in the encoder that are more affected by the deleted features. The influence of each data feature on a specific neuron or link in the encoder can be calculated according to various techniques. At least one such technique is disclosed in the article titled "Problems with Shapley-value-based explanations as feature importance measures" mentioned above. Then, the corresponding RE feeds back the IDs / locations of the deleted neurons and links to the encoder customizer 604 to complete the customization of the encoder.

[0165] In another example, the third mode may be referred to as a learning-based customization mode, in which, for a given selected encoder from which non-primary features have been removed from its input layer, the encoder customizer 604 applies a learning-based NN pruning scheme to train the customized encoder. At least one such scheme is disclosed in the article titled "Learning both Weights and Connections for Efficient Neural Network" written by Han, Song, et al. and published in 2015, the entire content of which is incorporated herein by reference.

[0166] Developers have recognized that such a scheme requires a large amount of computing resources during training. It should be noted that the training phase should be completed by the corresponding TE that implements the customized encoder to reduce the computational cost of executing the encoder customizer 604.

[0167] In the illustrated non-limiting embodiment, at step 650, the encoder customizer 604 sends encoder customization parameters to the corresponding TE 608 for performing the encoder customization decision process. The content of the encoder customization parameters may particularly depend on different encoder customization modes. For the minimal customization mode 720, information indicating the input layer neurons that have been removed (non-primary features) may be included in the encoder customization parameters. For the weight-based customization mode 730, information indicating the neurons and links that have been removed from the full encoder may be included in the encoder customization parameters. For the learning-based customization mode, information indicating the neurons removed from the input layer and hyperparameters for the customization training may be included in the encoder customization parameters. It is contemplated that in some embodiments of the present invention technology, hyperparameters may include a loss threshold for determining convergence, a weight threshold for neurons and links to be removed, and the like.

[0168] In the illustrated non-limiting embodiment, at step 660, the corresponding TE 608 implements an encoder customization decision based on the received encoder customization parameters of step 650. It is contemplated that the decision may be made based on relative weights. The relative weights may be calculated based on the encoder customization parameters received at step 650. For the minimal customization mode 720 and the weight-based customization mode 730, the corresponding TE 608 may directly remove the corresponding neurons and / or links. For the learning-based customization mode, the corresponding TE 608 may run the customization training until convergence without departing from the scope of the present invention technology. The customization training may be triggered in response to receiving the encoder customization parameters.

[0169] For illustration purposes only, assume that the corresponding TE 608 implements an encoder customization decision according to a learning-based customization mode. In the non-limiting embodiment shown, since the final customization result of the learning-based customization is obtained at the corresponding TE 608, after convergence, the corresponding TE 608 can feedback the customized encoder features (e.g., the customized structure of the encoder) to the encoder customizer 604 at step 660. For example, the customized encoder features can be used for further VFL configuration.

[0170] VFL Configuration Embodiment

[0171] Reference Figure 8 , a schematic representation of a VFL configuration program 800 is shown. The VFL configuration program 800 involves an encoder selector 802, an encoder customizer 804, a VFL configurator 806, a joint classifier 808, an evaluator 810, and a privacy router 332. It is contemplated that at least some of the components involved in the VFL configuration program 800 can be implemented in a manner similar to the components of the VFL architecture 300 shown without departing from the scope of the technology of the present invention. Figure 3 In the non-limiting embodiment shown, at step 820, the encoder selector 802 sends annotation data information to the VFL configurator 806, and at step 830, the encoder customizer 804 sends customized encoder information to the VFL configurator 806. In some embodiments, the VFL configurator 806 receiving the annotation data information and the customized encoder information can trigger the following steps of the VFL configuration program 800.

[0172] Broadly speaking, the annotation data information includes the location of the data set constituting one or more annotation data sources and the provided data features. In other words, the annotation data information indicates the location (which part of the annotation data comes from which of the first data source and the second data source) and how to combine the annotation data. It is contemplated that the loss function indicated by the VFL client can also be transmitted as part of the annotation data information. The customized encoder information includes information indicating the output layer structure of the corresponding customized encoder. For example, when the encoder customizer 804 determines that a first encoder needs to be customized according to the data feature weights, the customized encoder information includes information indicating the output layer structure of the customized first encoder. When the encoder customizer 804 determines that a first encoder and a second encoder need to be customized according to the data feature weights, the customized encoder information includes information indicating the output layer structures of the customized first encoder and the customized second encoder.

[0173]

[0174] ​In the illustrated non - limiting embodiment, at step 840, the VFL configurator 806 configures the encoder output to the joint classifier based on customized encoder information. The configuration of the encoder output to the joint classifier can be performed in a variety of ways. In some embodiments, the output layers of multiple encoders can be connected to the input layer of the joint classifier through the concatenation of the output layers of the multiple encoders. At least one other connection pattern is disclosed in the article titled "SplitNN - driven vertical partitioning" written by Iker Ceballos et al. and published on August 7, 2020, the entire content of which is incorporated herein by reference.

[0175] In the illustrated non - limiting embodiment, at step 850, the VFL configurator 806 sends joint classifier configuration parameters to the joint classifier 808. The content of the joint classifier configuration parameters is used by the joint classifier to customize or modify its local configuration to match the encoder output (which is the input of the joint classifier). For example, in some embodiments, as described above, the output layer neurons of the encoder customized in a weight - based or learning - based manner can be partially silenced and / or deleted. The VFL configurator can put the information of the silenced output layer neurons (e.g., of all encoders) into the joint classifier configuration parameters. By receiving the joint classifier configuration parameters, the joint classifier can customize its input layer architecture (e.g., by silencing neurons and / or deleting neurons) to match the encoder output.

[0176] In the illustrated non - limiting embodiment, at step 860, the VFL configurator 806 configures the evaluator 810. As part of step 860, when the partially labeled data from different data sets are merged in the evaluator 810, the VFL configurator 806 uses the labeled data information to configure the combination rule for the partially labeled data from different data sets. Also as part of step 860, the VFL configurator 806 configures the loss function to be used during VFL.

[0177] It is conceivable that partially labeled data can be said to be labeled data whose associated label information does not represent complete label information. All partially labeled data associated with the same ID / sample should be combined together to form complete labeled data. For example, the completed labeled data can be represented as "[A, blue cat]", where "A" is the associated ID / sample (i.e., the represented entity), and where "blue cat" is the label. In some embodiments, the labeled data can be provided from two different data sources, one data source providing "[A, blue]" and the other data source providing "[A, cat]". In this example, both "[A, blue]" and "[A, cat]" are named partially labeled data.

[0178] In the illustrated non - limiting embodiment, at step 880, the VFL configurator 806 configures the inter - encoder or "privacy" router 812. It is contemplated that the privacy router 812 may be implemented similar to the privacy router 332 in Figure 3 . As part of step 880, the VFL configurator 806 configures the routing lines for data transfer for the following: (i) the data exchange line between the joint classifier 808 and the evaluator 810, (ii) the labeled data merging line between the evaluator 810 and the corresponding data sources, and (iii) the backpropagation line for backpropagating gradient values between the joint classifier 808 and the corresponding encoders. The "corresponding encoders" in (iii) can be understood as: all the selected encoders selected by the encoder selector. When some of the selected encoders are customized by the encoder customizer, the "corresponding encoders" in (iii) include the customized encoders and the remaining non - customized encoders among all the selected encoders; when the selected encoders selected by the encoder selector include a first encoder and a second encoder, and the first encoder is customized, the "corresponding encoders" in (iii) include the customized first encoder and the second encoder.

[0179] In the illustrated non - limiting embodiment, at step 890, the VFL configurator 806 sends the privacy router configuration parameters to the privacy router 812. It is contemplated that in at least some embodiments of the present technology, the lines supported by the privacy router 812 are discussed above with reference to Figure 4 .

[0180] Performance Evaluation

[0181] Referring to Figure 9 , three simplified representations of the VFL framework are depicted. To verify the effectiveness and efficiency of the automated VFL architecture envisioned in at least some embodiments of the present technology, simulations have been performed to compare the performance of three scenarios in terms of a binary classification task, namely traditional VFL 910, automated VFL 920, and spliced single NN 930.

[0182] Broadly speaking, the traditional VFL 910 includes two participants (P1 and P2) with different encoders (E1 and E2). The detailed parameters of E1 and E2 are shown in Table 1 below. Both participants have independent classifiers with the same structure (C0). The inputs of the two participants (I1 and I2) completely overlap in the ID space and partially overlap in the feature space. The automated VFL 920 includes two participants (PA1 and PA2), with the encoders and input parts of P1 and P2. The outputs of the two participants are cascaded to a joint classifier with the same structure as C0. It should be noted that all classifiers used in these three cases are identical in structure. The spliced single NN 930 cascades I1 and I2 as an integrated input layer and uses full connections to splice E1 and E2 layer by layer to form a spliced encoder. It should be noted that the remaining hidden layers of the encoder (e.g., the remaining layers of E2) are fully connected to the spliced structure. The independent classifier C0 is fully connected to the output layer of the spliced encoder.

[0183] Simulation Parameters Value E1 Structure (Neurons per Layer) [32,64] E2 Structure (Neurons per Layer) [8,16,16] C0 Structure (Neurons per Layer) [8, 2 (using softmax)] Activation Function SELU Number of Epochs 100 Batch Size 32 Learning Rate 1e-3

[0184] Table 1: Simulation Parameters

[0185] In the simulation, the experiments were conducted on two datasets, namely the anonymized bank dataset (Germanbank) containing German bank user information and the network traffic dataset (TLS22) for deep packet inspection of encrypted Internet traffic.

[0186] Reference Figure 10 , shows a performance comparison between the traditional VFL 910 and the automated VFL 920 applied to the above two test datasets. To verify the effectiveness of encoder customization as envisioned in at least some embodiments of the present invention's technology, an automated VFL "autoVFL_noCut" without applying encoder customization and an automated VFL "autoVFL" with minimal encoder customization have been simulated. As shown in Charts 1010 and 1030, it can be seen that in some implementations of the present invention's technology, compared with the traditional VFL 910, the automated VFL 920 (with or without encoder customization) can converge to better optimal results with higher accuracy and lower variance. As shown in Charts 1020 and 1040, the classification performance of the traditional VFL 910 and the automated VFL 920 is compared. In at least some implementations of the present invention's technology, it can be said that compared with the traditional VFL 910, the automated VFL 920 can improve general performance metrics, especially in terms of accuracy (increased by 5% to 8%) and precision (increased by 10%).

[0187] Developers have recognized that in at least some embodiments of the technology of the present invention, better performance achieved through automated VFL can be realized through two aspects of automated VFL. First, compared with the intermediate result exchange of traditional VFL, the joint classifier can make the interaction between participating encoders more intensive, thereby improving the efficiency of training the joint model. Second, encoder customization can further reduce the negative impact brought by the overlapping data features of the inputs of multiple participants. Although both the joint classifier and encoder customization can improve the overall performance, according to Figure 10 the results shown, the impact of the joint classifier is higher than that of encoder customization. In addition, it should be noted that Figure 10 the simulation only applied the minimum encoder customization mode. When more complex encoder customization algorithms are fully implemented (other modes described above), encoder customization can improve the performance of automated VFL without departing from the scope of the technology of the present invention.

[0188] Referring to Figure 11 , a performance comparison between the automated VFL 920 applied to the above two test datasets and the concatenated single NN930 is shown. For the concatenated single NN 930, the single NN930 "singleNN_noCut" without applying encoder customization to its cascaded input layer and the single NN "singleNN_cut" applying minimum encoder customization to its cascaded input layer have both been simulated. It can be envisioned that in at least some embodiments of the technology of the present invention, the automated VFL can achieve comparable convergence and classification performance without significant loss compared with the concatenated single NN. In addition, as seen on charts 1110 and 1120, it should be noted that due to the simple concatenated NN structure may not be suitable for specific datasets or problems, the concatenated single NN 930 will result in significant performance defects. Although appropriate new NNs can reduce these drawbacks, developers have recognized that the re-design process of new NNs may be very costly because there are no clear design principles and no existing models can be reused. Therefore, it can be said that in at least some implementations of the technology of the present invention, the automated VFL 920 can achieve general performance metrics comparable to or better than those of the concatenated single NN 930.

[0189] In some embodiments of the technology of the present invention, an automated VFL system is provided, and the automated VFL system is used to perform at least some of the following functions: encoder selector, encoder customizer, VFL configurator, and joint classifier. In some implementations of the technology of the present invention, an automated VFL system can be implemented for the NET4AI architecture.

[0190] In some embodiments, compared with traditional VFL and concatenated single NNs, the automated VFL system can achieve better AI model performance because (i) the joint classifier can achieve deep mutual influence among participating encoders (compared with traditional VFL), and (ii) encoder customization can reduce the negative impact caused by overlapping data features. In other embodiments, the automated VFL system can achieve automatic encoder selection and reuse. Thus, this supports reusing existing encoders in the network to solve VFL tasks and may not require designing a specific NN for this purpose. There is no need to design a specific NN. The selection and customization of encoders (and datasets) are both automatically done by NET4AI. Additionally, this can avoid VFL configuration through the application layer, so there is no need for the VFL management overhead of the application layer. In other embodiments, the automated VFL system can solve complex tasks by supporting loss functions that violate homomorphic encryption and support labeled data being provided separately by multiple data sources.

[0191] Those of ordinary skill in the art will recognize that the descriptions of the various embodiments are merely exemplary and are not intended to limit in any way. Other embodiments will be readily proposed to those of ordinary skill in the art who benefit from the present invention. Additionally, at least some of the disclosed embodiments can be customized to provide valuable solutions to address existing needs and problems related to VFL. For clarity, not all conventional features of the implementation of at least some of the disclosed embodiments are shown and described.

[0192] Specifically, the combinations of features are not limited to those presented in the above description, as the combinations of elements listed in the appended claims form part of the present disclosure. Of course, it should be understood that in developing any such actual implementation of at least some of the disclosed embodiments, many specific implementation decisions may be required to achieve the developer's specific goals, such as complying with application, system, and business-related constraints, and these specific goals will vary depending on the implementation and the developer. Additionally, it should be understood that the development work may be complex and time-consuming, but it is a routine engineering task for those of ordinary skill in the art who benefit from the present invention in feedback equalization at high data rates.

[0193] In accordance with the present invention, the components, process operations, and / or data structures described herein can be implemented using various types of operating systems, computing platforms, network devices, computer programs, and / or general-purpose machines. Additionally, those of ordinary skill in the art will recognize that less general-purpose devices, such as hardwired devices, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc., can also be used. When a method comprising a series of operations is implemented by a computer, a processor operatively connected to a memory, or a machine, these operations can be stored as a series of instructions readable by the machine, processor, or computer and can be stored on a non-transitory tangible medium.

[0194] The systems and modules described herein can include software, firmware, hardware, or any one or more combinations of software, firmware, or hardware suitable for the purposes described herein. The software and other modules can be executed by a processor and reside in the memory of a server, workstation, personal computer, computerized tablet, personal digital assistant (PDA), and other devices suitable for the purposes described herein. The software and other modules can be accessed via local memory, via a network, via a browser or other application, or by other means suitable for the purposes described herein. The data structures described herein can include computer files, variables, programming arrays, programming structures, or any electronic information storage scheme or method suitable for the purposes described herein, or any combination thereof.

[0195] The present invention has been described in the foregoing specification by way of non-limiting illustrative embodiments provided as examples. These illustrative embodiments can be modified arbitrarily. The scope of the claims should not be limited by the embodiments set forth in the examples, but should be given the broadest interpretation consistent with the entire description.

Claims

1. A method, characterized in that, comprising: An encoder selector receives a VFL request from a vertical federated learning VFL client; The encoder selector determines, according to the VFL request, data requirements and encoder requirements for solving the VFL task, the VFL task being splittable into multiple subtasks, the multiple subtasks including at least a first subtask and a second subtask; Based on the encoder requirements, the encoder selector selects a first encoder and a second encoder from an encoder pool in the communication network, the first encoder being used to process a first feature set for solving the first subtask, and the second encoder being used to process a second feature set for solving the second subtask; Based on the data requirements, the encoder selector selects a first data source for the first encoder and a second data source for the second encoder, the selected data sources being from a data source pool, the first data source including the first feature set, and the second data source including the second feature set.

2. The method according to claim 1, characterized in that, The data requirements indicate requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

3. The method according to claim 1 or 2, characterized in that, The encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises: the encoder selector splits the VFL task into the multiple subtasks.

5. The method according to any one of claims 1 to 4, characterized in that, The selection of the first encoder comprises: The encoder selector determines an application category of the first subtask based on the encoder requirements, the application category being stored in a memory associated with the first encoder.

6. The method according to any one of claims 1 to 4, characterized in that, The selection of the first encoder comprises: The encoder selector identifies the encoder pool having the first encoder structure; The encoder selector selects the first encoder from the encoder pool having the first encoder structure.

7. The method according to claim 1, characterized in that, The selection of the first data source and the second data source comprises: The encoder selector sends a request for data source information to a data source manager, the request including the data requirements; The encoder selector receives the data source information from the data source manager, the data source information indicating the quality and features of the data included in the corresponding data sources in the data source pool; The encoder selector uses the data source information to select the first data source for the first encoder and the second data source for the second encoder.

8. The method according to claim 7, characterized in that, The first data source and the second data source are associated with corresponding IDs, and the method further comprises: The encoder selector sends the corresponding IDs of the first data source and the second data source to the data source manager; The encoder selector receives the network locations of the first data source and the second data source from the data source manager.

9. A method characterized in that it includes: The encoder customizer customizes the first encoder by silencing a part of the first encoder, thereby generating a customized first encoder. The first encoder and the second encoder have been selected from an encoder pool in a communication network based on encoder requirements, and the encoder requirements have been determined according to a vertical federated learning VFL request for solving a VFL task; the VFL task can be split into multiple subtasks, and the multiple subtasks at least include a first subtask and a second subtask. The first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask.

10. The method according to claim 9, characterized in that the customization includes: determining overlapping features between the first feature set and the second feature set; silencing the part of the first encoder that processes the overlapping features.

11. The method according to claim 9 or 10, characterized in that the customization includes: the encoder customizer determines that the first encoder is to be customized.

12. The method according to claim 11, characterized in that the encoder customizer determines that the first encoder is to be customized, including: the encoder customizer determines that the first encoder is to be customized based on data feature weights received from a referenced AI enabler.

13. The method according to any one of claims 9 to 12, characterized in that the customizing the first encoder includes: the encoder customizer silences the part of the first encoder that processes the overlapping features, and the silencing includes silencing at least one of neurons of the first encoder and links of the first encoder.

14. A method characterized in that it includes: A vertical federated learning VFL configurator configures at least one of an evaluator and a joint classifier; The VFL configurator configures a privacy router to determine at least one of the following: a data exchange route between the joint classifier and the evaluator, a labeled data merging route between the evaluator and one or more data sources, a backpropagation route between the joint classifier and the customized first encoder and the second encoder for backpropagating gradient values, where the joint classifier, the evaluator, and the privacy router are used to perform VFL.

15. The method according to claim 11, characterized in that the configuring the evaluator includes: the VFL configurator receives labeled data information from an encoder selector; the VFL configurator determines a loss function for performing VFL; the VFL configurator configures the evaluator using the labeled data information and the loss function.

16. The method according to claim 12, characterized in that The configuration of the evaluator includes: The VFL configurator configures the combination rule of the labeled data and the loss function, and the labeled data is sourced from the data sources selected for the first encoder and the second encoder.

17. The method according to any one of claims 11 to 13, wherein, The configuration of the joint classifier includes: The VFL configurator receives information indicating the output layer structures of the first encoder and the second encoder from the encoder customizer; The VFL configurator configures the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

18. The method according to any one of claims 14 to 17, wherein, The second encoder is a customized second encoder.

19. An encoder selector, wherein, for: Receiving a VFL request from a vertical federated learning (VFL) client; According to the VFL request, determining the data requirements and encoder requirements for solving the VFL task, the VFL task being splittable into multiple subtasks, the multiple subtasks including at least a first subtask and a second subtask; Based on the encoder requirements, selecting a first encoder and a second encoder from the encoder pool in the communication network, the first encoder being used to process a first feature set for solving the first subtask, and the second encoder being used to process a second feature set for solving the second subtask; Based on the data requirements, selecting a first data source for the first encoder and a second data source for the second encoder, the selected data sources being from a data source pool, the first data source including the first feature set, and the second data source including the second feature set.

20. The encoder selector according to claim 19, wherein, The data requirements indicate the requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

21. The encoder selector according to claim 19 or 20, wherein, The encoder requirements indicate a first encoder structure for solving the first subtask and a second encoder structure for solving the second subtask.

22. The encoder selector according to any one of claims 19 to 21, wherein, The encoder selector is further configured to split the VFL task into the multiple subtasks.

23. The encoder selector according to any one of claims 19 to 22, wherein, The selection of the first encoder includes that the encoder selector is configured to: Determine the application category of the first subtask based on the encoder requirements, The application category is stored in a memory associated with the first encoder.

24. The encoder selector according to any one of claims 19 to 22, wherein, The selection of the first encoder includes that the encoder selector is configured to: Identify the encoder pool having the first encoder structure; Select the first encoder from the encoder pool having the first encoder structure.

25. The encoder selector according to claim 19, wherein, the selection of the first data source and the second data source includes the encoder selector being configured to: send a request for data source information to the data source manager, the request including the data requirements; receive the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data source in the data source pool; use the data source information to select the first data source for the first encoder and the second data source for the second encoder.

26. The encoder selector according to claim 19, wherein, the first data source and the second data source are associated with corresponding IDs, and the encoder selector is further configured to: send the corresponding IDs of the first data source and the second data source to the data source manager; receive the network locations of the first data source and the second data source from the data source manager.

27. An encoder customizer, wherein, for: customizing the first encoder by silencing a part of the first encoder, thereby generating a customized first encoder, the first encoder and the second encoder have been selected from an encoder pool in a communication network based on encoder requirements, the encoder requirements having been determined according to a vertical federated learning VFL request for solving a VFL task; the VFL task can be split into multiple subtasks, the multiple subtasks at least including a first subtask and a second subtask, the first encoder being used to process a first feature set for solving the first subtask, and the second encoder being used to process a second feature set for solving the second subtask.

28. The encoder customizer according to claim 27, wherein, the customization includes the encoder customizer being configured to: determine the overlapping features between the first feature set and the second feature set; silence the part of the first encoder that processes the overlapping features.

29. The encoder customizer according to claim 27 or 28, wherein, the customization includes the encoder customizer being configured to: determine that the first encoder is to be customized.

30. The encoder customizer according to claim 29, wherein, the determination that the first encoder is to be customized includes the encoder customizer being configured to: determine that the first encoder is to be customized based on data feature weights received from a referenced AI enabler.

31. The encoder customizer according to any one of claims 27 to 30, wherein, the customization of the first encoder includes the encoder customizer being configured to: silence the part of the first encoder that processes the overlapping features, the silencing includes silencing at least one of the neurons of the first encoder and the links of the first encoder.

32. A vertical federated learning VFL configurator, wherein, for: configuring at least one of an evaluator and a joint classifier; configuring a privacy router to determine at least one of the following: The data exchange route between the joint classifier and the evaluator The labeled data merging route between the evaluator and one or more data sources The backpropagation route between the joint classifier and the customized first encoder and second encoder for backpropagating gradient values, wherein the joint classifier, the evaluator, and the privacy router are used to perform VFL 33. The VFL configurator according to claim 32 wherein Configuring the evaluator includes the VFL configurator for Receiving labeled data information from an encoder selector Determining a loss function for performing VFL Configuring the evaluator using the labeled data information and the loss function 34. The VFL configurator according to claim 33 wherein Configuring the evaluator includes the VFL configurator for Configuring the combination rule of the labeled data and the loss function, and the labeled data is sourced from the data sources selected for the first encoder and the second encoder 35. The VFL configurator according to any one of claims 32 to 34 wherein Configuring the joint classifier includes the VFL configurator for Receiving information indicating the output layer structures of the first encoder and the second encoder from an encoder customizer Configuring the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder 36. The VFL configurator according to any one of claims 32 to 35 wherein The second encoder is a customized second encoder 37. A method wherein includes An encoder selector receives a VFL request from a vertical federated learning VFL client The encoder selector determines, according to the VFL request, the data requirements and encoder requirements for solving the VFL task, and the VFL task can be split into multiple subtasks, and the multiple subtasks at least include a first subtask and a second subtask The encoder selector selects a first encoder and a second encoder from an encoder pool in the communication network based on the encoder requirements, where the first encoder is used to process a first feature set for solving the first subtask, and the second encoder is used to process a second feature set for solving the second subtask The encoder selector selects a first data source for the first encoder and a second data source for the second encoder based on the data requirements, and the selected data sources are from a data source pool, the first data source includes the first feature set, and the second data source includes the second feature set An encoder customizer customizes the first encoder by silencing a part of the first encoder, thereby generating a customized encoder A VFL configurator configures at least one of an evaluator and a joint classifier The VFL configurator configures a privacy router to determine at least one of the following The data exchange route between the joint classifier and the evaluator The labeled data merging route between the evaluator and one or more of the selected data sources The backpropagation routes between the joint classifier and each of the customized first encoder and second encoder are used to backpropagate gradient values; The joint classifier, the evaluator, and the privacy router are used to perform VFL.

38. The method according to claim 37, wherein, The data requirement indicates the requirements for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

39. The method according to claim 37 or 38, wherein, The encoder requirement indicates the first encoder structure for solving the first subtask and the second encoder structure for solving the second subtask.

40. The method according to any one of claims 37 to 39, wherein, The method further includes: the encoder selector splits the VFL task into the multiple subtasks.

41. The method according to any one of claims 37 to 40, wherein, The selecting the first encoder includes: The encoder selector determines the application category of the first subtask based on the encoder requirement, The application category is stored in the memory associated with the first encoder.

42. The method according to any one of claims 37 to 41, wherein, The selecting the first encoder includes: The encoder selector identifies the encoder pool having the first encoder structure; The encoder selector selects the first encoder from the encoder pool having the first encoder structure.

43. The method according to claim 37, wherein, The selecting the first data source and the second data source includes: The encoder selector sends a request for data source information to the data source manager, the request including the data requirement; The encoder selector receives the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data sources in the data source pool; The encoder selector uses the data source information to select the first data source for the first encoder and the second data source for the second encoder.

44. The method according to claim 43, wherein, The first data source and the second data source are associated with corresponding IDs, and the method further includes: The encoder selector sends the corresponding IDs of the first data source and the second data source to the data source manager; The encoder selector receives the network locations of the first data source and the second data source from the data source manager.

45. The method according to any one of claims 37 to 44, wherein, The customization includes: Determining the overlapping features between the first feature set and the second feature set; Silencing the part of the first encoder that processes the overlapping features.

46. The method according to any one of claims 37 to 45, wherein, The customization includes: The encoder customizer determines that the first encoder is to be customized.

47. The method according to claim 46, wherein, the encoder customizer determining that the first encoder is to be customized includes: the encoder customizer determining that the first encoder is to be customized based on data feature weights received from the referenced AI enabler.

48. The method according to any one of claims 37 to 47, wherein, customizing the first encoder includes: the encoder customizer silencing the part of the first encoder that processes the overlapping features, the silencing including silencing at least one of the neurons of the first encoder and the links of the first encoder.

49. The method according to claim 37, wherein, configuring the evaluator includes: the VFL configurator receiving labeled data information from the encoder selector; the VFL configurator determining a loss function for performing VFL; the VFL configurator configuring the evaluator using the labeled data information and the loss function.

50. The method according to claim 49, wherein, configuring the evaluator includes: the VFL configurator configuring a combination rule for the labeled data and the loss function, the labeled data being sourced from data sources selected for the first encoder and the second encoder.

51. The method according to any one of claims 47 to 50, wherein, configuring the joint classifier includes: the VFL configurator receiving information indicating the output layer structures of the first encoder and the second encoder from the encoder customizer; the VFL configurator configuring the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

52. The method according to any one of claims 37 to 51, wherein, the second encoder is a customized second encoder.

53. A system, wherein, for: an encoder selector receiving a VFL request from a vertical federated learning (VFL) client; the encoder selector determining, according to the VFL request, data requirements and encoder requirements for solving a VFL task, the VFL task being splittable into multiple subtasks, the multiple subtasks including at least a first subtask and a second subtask; the encoder selector selecting a first encoder and a second encoder from an encoder pool in the communication network based on the encoder requirements, the first encoder being used to process a first feature set for solving the first subtask, and the second encoder being used to process a second feature set for solving the second subtask; the encoder selector selecting a first data source for the first encoder and a second data source for the second encoder based on the data requirements, the selected data sources being from a data source pool, the first data source including the first feature set, and the second data source including the second feature set; an encoder customizer customizing the first encoder by silencing a part of the first encoder, thereby generating a customized encoder; a VFL configurator configuring at least one of an evaluator and a joint classifier; The VFL configurator configures the privacy router to determine at least one of the following: The data exchange route between the federated classifier and the evaluator; The labeled data merging route between the evaluator and one or more of the selected data sources among the selected data sources; The backpropagation route between the federated classifier and each of the customized first encoder and second encoder for backpropagating gradient values; The federated classifier, the evaluator, and the privacy router are used to perform VFL.

54. The system according to claim 53, wherein, The data requirement indicates the requirement for selecting the first data source for solving the first subtask and the second data source for solving the second subtask.

55. The system according to claim 53 or 54, wherein, The encoder requirement indicates the first encoder structure for solving the first subtask and the second encoder structure for solving the second subtask.

56. The system according to any one of claims 53 to 55, wherein, The encoder selector is further configured to split the VFL task into the multiple subtasks.

57. The system according to any one of claims 53 to 56, wherein, The selecting the first encoder includes that the encoder selector is configured to: Determine the application category of the first subtask based on the encoder requirement, The application category is stored in the memory associated with the first encoder.

58. The system according to any one of claims 53 to 57, wherein, The selecting the first encoder includes that the encoder selector is configured to: Identify the encoder pool having the first encoder structure; Select the first encoder from the encoder pool having the first encoder structure.

59. The system according to claim 53, wherein, The selecting the first data source and the second data source includes that the encoder selector is configured to: Send a request for data source information to the data source manager, the request including the data requirement; Receive the data source information from the data source manager, the data source information indicating the quality and characteristics of the data included in the corresponding data sources in the data source pool; Use the data source information to select the first data source for the first encoder and the second data source for the second encoder.

60. The system according to claim 59, wherein, The first data source and the second data source are associated with corresponding IDs, and the encoder selector is further configured to: Send the corresponding IDs of the first data source and the second data source to the data source manager; Receive the network locations of the first data source and the second data source from the data source manager.

61. The system according to any one of claims 53 to 60, wherein, The customization includes that the encoder customizer is configured to: Determine the overlapping features between the first feature set and the second feature set; Silence the part in the first encoder that processes the overlapping features.

62. The system according to claim 61, wherein, the customization includes that the encoder customizer is used for: obtaining (i) the encoder structure parameters of the first encoder and the second encoder, (ii) the position of the target AI enabler of the customized encoder, and (iii) the position of the referenced AI enabler of the trained encoder; sending a request for the data feature weights of the trained encoder to the referenced AI enabler; receiving the data feature weights from the referenced AI enabler; determining the overlapping features between the first feature set of the first encoder and the second feature set of the second encoder; determining that the first encoder is to be customized based on the data feature weights; silencing the part of the first encoder that processes the overlapping features, wherein the silencing includes silencing at least one of the neurons of the first encoder and the links of the first encoder.

63. The system according to claim 53, wherein, configuring the evaluator includes that the VFL configurator is used for: receiving labeled data information from the encoder selector; determining a loss function for performing VFL; configuring the evaluator using the labeled data information and the loss function.

64. The system according to claim 63, wherein, configuring the evaluator includes that the VFL configurator is used for: configuring a combination rule for the labeled data and the loss function, where the labeled data is sourced from data sources selected for the first encoder and the second encoder.

65. The system according to claim 63 or 64, wherein, configuring the joint classifier includes that the VFL configurator is used for: receiving information indicating the output layer structures of the first encoder and the second encoder from the encoder customizer; configuring the encoder outputs to the joint classifier based on the output layer structures of the first encoder and the second encoder.

66. The system according to any one of claims 53 to 65, wherein, the second encoder is a customized second encoder.