A model training method, apparatus, device, and readable storage medium
By deploying memory components on the first node to store the training sample features of the second node, using federated learning methods for feature association and model training, the problems of data privacy security and model accuracy are solved, and a privacy, secure and efficient model training effect is achieved.
Patent Information
- Application Number
- CN202311008652.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-08-10
AI Technical Summary
In federated learning, how to effectively use the data information of multiple nodes to train models while ensuring data privacy and security, and improve the accuracy and generalization capabilities of the model.
Using the idea of federated learning, the memory component is deployed at the first node to store the training sample features from the second node, and feature association and model training are performed through the first encoder and classifier to avoid original data transmission and ensure privacy and security.
It improves the privacy, security and accuracy of the model, expands the breadth and dimension of data characteristics, and improves the training accuracy and generalization capabilities of the model.
Smart Images

Figure CN117151250B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to a model training method, apparatus, device, and readable storage medium. Background Art
[0002] With the increasing attention to privacy data and multi-party secure computing, using deep learning to automatically learn effective feature representations from data and improve the accuracy of prediction models has been widely applied in fields such as speech recognition, image recognition, object detection, and risk recognition.
[0003] Currently, the federated learning (FL) method can be used to train a model. Multiple nodes are used to jointly train a model with the training samples they each hold, and there is no need to transmit the training samples to a central server, thus avoiding the leakage of training samples.
[0004] Based on this, this specification provides a model training method. Summary of the Invention
[0005] This specification provides a model training method, apparatus, device, and readable storage medium to partially solve the above problems existing in the prior art.
[0006] This specification adopts the following technical solutions:
[0007] This specification provides a model training method. The method is applied to a first node that deploys a target model to be trained. The target model includes a first encoder and a classifier;
[0008] The method includes:
[0009] Determine a first training sample and the annotation of the first training sample;
[0010] Receive the features of a second training sample sent by a second node, and write the features of the second training sample into a memory component deployed in the first node; wherein, the features of the second training sample are determined by a second encoder deployed in the second node;
[0011] Input the first training sample into the first encoder to obtain the features of the first training sample output by the first encoder;
[0012] According to the features of the first training sample, read target features associated with the features of the first training sample from the features of the second training sample stored in the memory component;
[0013] Input the features of the first training sample and the target features into the classifier to obtain the prediction result of the first training sample output by the classifier;
[0014] Train the target model according to the prediction result and the annotation of the first training sample.
[0015] This specification provides a model training device, which is applied to a first node. The first node deploys a target model to be trained, and the target model includes a first encoder and a classifier;
[0016] The device includes:
[0017] A first training sample determination module, configured to determine a first training sample and the annotation of the first training sample;
[0018] A writing module, configured to receive the features of a second training sample sent by a second node, and write the features of the second training sample into a memory component deployed in the first node; wherein, the features of the second training sample are determined by a second encoder deployed in the second node;
[0019] A first feature determination module, configured to input the first training sample into the first encoder to obtain the features of the first training sample output by the first encoder;
[0020] A target feature determination module, configured to read, according to the features of the first training sample, target features associated with the features of the first training sample from the features of the second training sample stored in the memory component;
[0021] A prediction result determination module, configured to input the features of the first training sample and the target features into the classifier to obtain the prediction result of the first training sample output by the classifier;
[0022] A training module, configured to train the target model according to the prediction result and the annotation of the first training sample.
[0023] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above model training method is implemented.
[0024] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above model training method is implemented.
[0025] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0026] In the model training method provided in this specification, the first node determines the first training sample and its annotation, and when receiving the features of the second training sample sent by the second node, writes the features of the second training sample into the memory component. Thus, when determining the features of the first training sample through the first encoder, the target features associated with the features of the first training sample are read from the memory component according to the features of the first training sample, and then the target features and the features of the first training sample are input into the classifier to obtain the prediction result of the first training sample. Therefore, based on the prediction result and annotation of the first training sample, the target model is trained. It can be seen that the above solution is based on the idea of federated learning. By deploying the memory component on the first node to store the features from the second node, while effectively utilizing the information of the second training sample during the training process of the target model, it ensures the privacy and security of the second training sample and the first training sample, and improves the privacy security and accuracy of the model. Description of the Drawings
[0027] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The schematic embodiments and descriptions thereof in this specification are used to explain this specification and do not constitute an improper limitation to this specification. In the attached
[0028] In the figures:
[0029] Figure 1 is a schematic flowchart of a model training method in this specification;
[0030] Figure 2 is a schematic structural diagram of a model training system in this specification;
[0031] Figure 3 is a schematic flowchart of a model training method in this specification;
[0032] Figure 4 is a schematic flowchart of a model training method in this specification;
[0033] Figure 5 is a schematic diagram of a model training device provided in this specification;
[0034] Figure 6 corresponds to the Figure 1 schematic diagram of the electronic device. Detailed Embodiments
[0035] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0036] In addition, it should be noted that all actions of obtaining signals, information, or data in this specification are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.
[0037] Deep learning is an important branch field of computer science and artificial intelligence and is a further extension of neural networks. By automatically learning effective feature representations from data to improve the accuracy of prediction models, it has been widely applied in fields such as speech recognition, image recognition, transaction risk recognition, and object detection.
[0038] With the rapid expansion of the data scale of deep learning and the increasing attention to data security, how different institutions can share data to achieve win-win cooperation at the model level while ensuring security is an important issue in digital transformation. Federated learning has been the main paradigm in joint modeling in recent years. Especially in the scenario of transaction risk recognition in the financial industry, the federated learning method of using multiple nodes to execute model training tasks in parallel has been widely applied. While ensuring the privacy and security of data, it uses the data of multiple nodes to improve the training accuracy of the model, thereby accelerating the accuracy of the model's actual application. Among them, federated learning can meet the premise that the local data separately saved by each party does not leave the domain, and on this premise, conduct multi-party joint modeling, and realize the joint training of the model by exchanging intermediate model results (such as the representation of the original data) between different parties.
[0039] Based on this, this specification provides a model training method. Based on the idea of federated learning, the memory component deployed in the first node stores the features of the second training samples from the second node. While effectively using the information of the second training samples during the training process of the target model, it ensures the privacy and security of the second training samples and the first training samples, and improves the privacy security and accuracy of the model.
[0040] The following will detail the technical solutions provided by each embodiment of this specification in conjunction with the drawings.
[0041] Figure 1 It is a schematic flowchart of a model training method provided by this specification.
[0042] S100: Determine the first training sample and the annotation of the first training sample.
[0043] A model training method provided in an embodiment of this specification is applied to a first node. Therefore, the execution process of this model training method can be executed by the first node for model training, where the first node can be an electronic device such as a server. In addition, after the target model is trained, to ensure the privacy and security of the model structure and model parameters of the target model, the node that performs prediction on the data to be predicted based on the trained target model can be the first node, or it can also be other nodes trusted by the first node. This specification does not limit this.
[0044] For the convenience of description, taking the example where both the node that trains the target model by executing this model training method and the node that performs prediction based on the trained target model are the first node, the specific technical solution will be described.
[0045] In various actual different scenarios, there are privacy computing and data security requirements. Generally speaking, for the data stored by a single data holder, the characteristics of its data are restricted by the business type executed between the data holder and the user. Therefore, even if the data holder has a relatively wide user group, the data it stores will have the problem of single characteristics. Therefore, the method of federated learning can be adopted to comprehensively combine the data separately stored by multiple data holders, expand the breadth and dimension of data characteristics, so as to jointly train a model with better performance while ensuring data security.
[0046] In the model training method provided in this specification, both the first node and the second node can be data holders that store data. According to different application scenarios, the models to be jointly trained also have different functions. For example, in a speech recognition scenario, the target model can be a speech recognition model. Then, the data stored by the first node can be the speech data of the voice chat between the chatbot and the user, and the data stored by the second node can be the speech data of the voice conversation between the in-vehicle voice system and the user. Or in an image classification scenario, the target model can be an image classification model. Then, the data stored by the first node can be the pedestrian images captured by the urban road surveillance cameras, and the data stored by the second node can be the images of consumers captured by the indoor surveillance cameras in the mall. Another example is in a transaction risk control scenario, the target model can be a transaction risk identification model. Then, the data stored by the first node can be the fund flow data of the user in the bank, and the data stored by the second node can be the transaction data of the user's purchase of goods in the merchant.
[0047] It can be seen that the data separately saved by different nodes (data holders) may intersect in terms of users, but there are differences in terms of data characteristics. The model trained only based on the data mastered by one node may have problems such as low data accuracy and inability to be generalized and applied to actual scenarios.
[0048] Therefore, in this specification, a method is adopted in which a target model including a first encoder and a classifier is deployed on a first node, and a second encoder is deployed on a second node. The features of the data mastered by the second node are obtained through the second encoder, and the second node sends the features of the data to the first node. That is, the data held by the first node and the data held by the second node are combined, and the original data saved by the second node will not be leaked due to the need to be transmitted to the first node, improving the privacy and security of the data.
[0049] Generally, based on the data saved by the first node and the scenario applicable to the target model, a first training sample is determined. For example, in the field of financial risk control, when the first node is a bank, the fund flow of users saved by the bank may include the transaction record data and loan data of users. If the target model is a transaction risk identification model, the first training sample needs to be determined based on the transaction record data of users saved by the bank.
[0050] Furthermore, the annotation of the first training sample can be obtained by manual annotation or automatically annotated based on other machine learning models. This specification does not make any limitations in this regard.
[0051] Taking the target model as a transaction risk identification model as an example, the first training sample can be the transaction record data of users saved by the bank. The transaction record data may include transaction time, transaction amount, transaction object (such as local account for transaction, commodity name or merchant name), transaction location, transaction remarks (extra transaction description or remarks information, such as purpose of transfer, transaction postscript, etc.), the account balance of the user, and transaction type (deposit, transfer, withdrawal, payment, receipt, etc.). The annotation of the first training sample is whether there is a transaction risk or what type of transaction risk (black production transactions such as fraud, money laundering, gambling, etc.) exists in this transaction record of the user.
[0052] S102: Receive the features of the second training sample sent by the second node, and write the features of the second training sample into the memory component deployed in the first node; wherein, the features of the second training sample are determined by the second encoder deployed in the second node.
[0053] In practical applications, to improve the accuracy and generalization of the target model, the target model can be jointly trained by combining data held by multiple nodes (data holders). Since the target model with a complete structure is deployed on the first node, the combination of data held by multiple nodes can be the combination of the features of the training samples determined by each node respectively, thus avoiding the transmission of the training samples themselves.
[0054] It can be seen that after the first node determines the first training sample based on the data it stores and the scenario applicable to the target model, the second node also needs to determine the second training sample based on the data it stores and the scenario applicable to the target model. Still taking the above financial risk control field as an example, if the target model is a transaction risk identification model, the purpose of the target model is actually to identify whether there is a transaction risk in the transactions executed by users from the transaction data of users. Therefore, the second node can be a merchant, and the second training sample can be the transaction data of the transactions between the merchant and the user, such as product information (product name, description, model, specification, quantity, price, etc.), transaction time and location, transaction amount, payment method, user information (user name, contact information (mobile phone number, email address), delivery address, etc.) and merchant information (merchant name, contact information, business address, etc.).
[0055] To ensure that the second training sample of the second node does not leak, the features of the second training sample can be extracted based on the second encoder deployed on the second node, and the features of the second training sample can be used as the object to be transmitted and sent to the first node. Even if the features of the second training sample leak, it is difficult for an attacker to reverse the second training sample itself from the features of the second training sample, improving the privacy security of the second training sample.
[0056] For this reason, the second node can input the second training sample into the second encoder to obtain the features of the second training sample output by the second encoder. It should be noted here that the second encoder deployed on the second node can be a pre-trained encoder or an encoder to be trained. This specification does not make any limitations in this regard. When the second encoder is a pre-trained encoder, the training process of this second encoder can be completed by the second node based on the training samples, or by other nodes based on the training samples. This specification does not make any limitations in this regard. When the second encoder is an encoder to be trained, during the process of training the target model, when the first node performs backpropagation of gradients, it can send the gradient information to the second node, and the second node can optimize the model parameters of the second encoder based on the gradient descent algorithm, that is, the second encoder is trained simultaneously during the training process of the target model.
[0057] Such as Figure 2The following is a schematic structural diagram of a model training system. In this model training system, there are a first node and a second node. Among them, the complete target model is deployed on the first node, which includes a first encoder and a classifier. Only the second encoder is deployed on the second node. Moreover, in order to save the features of the second training samples sent by the second node, a memory component is also deployed on the first node. When the first node receives the features of the second training samples sent by the second node, it can directly write the features of the second training samples into this memory component.
[0058] Among them, the memory component is a component containing a memory network. The memory network in the memory component can store data and can also forget data. The first node can perform write operations and read operations on this memory component. Among them, the write operation means that the first node stores the features of the second training samples sent by the second node, and the read operation means that the first node obtains the required features from the features of the second training samples stored in the memory component. Generally, the memory component includes an input module, an output module, a storage module, and a query module. When the first node receives the features of the second training samples sent by the second node, it can use the input module in the memory component to write the features of the second training samples into the storage module. When the first node needs to obtain features from the memory component, it can screen out the features that the first node needs to obtain from the features stored in the storage module through the query module, and then read out the screened features through the output module.
[0059] In addition, since in the process of iterative training of the target model, in each round of iteration, the second node will send the features of the second training samples to the first node. In this specification, the storage module in the configurable memory component can store the features of all the second training samples sent by the second node during the complete training process. That is to say, if the first node is not receiving the features of the second training samples for the first time currently, when the first node writes the features of the second training samples into the memory component, it will not update and overwrite the features written into the memory component before. And after the target model training is completed, the memory component will store the features of all the second training samples during the model training process. Of course, in order to save the storage resources of the storage module, the features of the second training samples that have not contributed to the training of the target model can be forgotten. That is, if the first node is not receiving the features of the second training samples for the first time currently, before the first node writes the features of the second training samples into the memory component, it can also judge whether there are other features stored in the memory component deployed on the first node. If so, update the features stored in the memory component according to the features of the second training samples. That is, use the features of the second training samples sent by the current second node to overwrite at least part of the features already stored in the memory component.
[0060] S104: Input the first training sample into the first encoder to obtain the features of the first training sample output by the first encoder.
[0061] Specifically, the first encoder can map the first training sample to the representation space to extract the features of the first training sample.
[0062] The model structures and model parameters of the first encoder and the second encoder can be the same or different, and this specification does not make any limitations in this regard.
[0063] S106: According to the features of the first training sample, read the target features associated with the features of the first training sample from the features of the second training sample stored in the memory component.
[0064] As described above, a complete target model is deployed on the first node, and the first training sample and its annotation are determined based on the scenario applicable to the target model. That is, the first training sample and its annotation are closely related to the purpose and task that the target model can achieve. Although the second training sample is also determined based on the scenario applicable to the target model, the second training sample and the first training sample may have a large overlap in users but a small overlap in features. If the features of the first training sample and the second training sample are simply concatenated and fused to obtain fused features and used for downstream tasks (classifiers), although it will greatly expand the dimension of the features included in the fused features, it may be difficult to train the target model or impossible to train due to the sparse features. Therefore, based on the attention mechanism method, the features of the first training sample can be used as the query vector to find the target features associated with the features of the first training sample from the features of the second training sample stored in the memory component. Since the target features are associated with the features of the first training sample, in subsequent steps, when the target features and the features of the first training sample are fused and used for downstream tasks, not only the dimension of the fused features is reduced, but also the correlation degree between the features in each dimension of the fused features is improved, which is more conducive to the training of the downstream classifier.
[0065] For the method based on the attention mechanism, the process of using the features of the first training sample as the query vector to find the target features associated with the features of the first training sample from the features of the second training samples stored in the memory component can be as follows: For the method based on the attention mechanism, determine the similarity between the features of the first training sample and the features of the second training samples stored in the memory component. The similarity here can be obtained based on any existing similarity determination method, such as Euclidean distance, dot product, or non-linear transformation using a neural network, etc. This specification does not limit this. Furthermore, perform normalization processing on the determined similarity to determine the attention weights. The attention weights can represent the importance and correlation degree of the features of the second training samples stored in the memory component to the features of the first training sample. The higher the attention weight, the more relevant the features of the second training sample are to the features of the first training sample and the more relevant to the training objective of the current target model. Therefore, based on the attention weights between the features of the first training sample and the features of each second training sample, several second training samples with higher attention weights can be selected as the target features. As for the number of the selected target features, it can be dynamically determined based on the actual situation or a preset fixed number. This specification does not limit this.
[0066] S108: Input the features of the first training sample and the target features into the classifier to obtain the prediction result of the first training sample output by the classifier.
[0067] Specifically, the features of the first training sample and the target features can both be used as the input of the classifier, or the features of the first training sample and the target features can be fused to obtain the fused features, and the fused features are used as the input of the classifier. The fusion method can be splicing, average fusion, or weighted fusion. This specification does not limit this.
[0068] According to the specific application scenario, determine the specific type of the classifier in the target model, such as a linear classifier, decision tree, naive Bayes, artificial neural network, K-nearest neighbor, support vector machine, etc. This specification does not limit this.
[0069] Based on the input features of the first training sample and the target features, the classifier can identify that the first training sample belongs to one of the preset multiple categories, and use the identified category as the prediction result of the first training sample.
[0070] For example, if the target model is a transaction risk identification model, and the multiple preset categories corresponding to the first training sample include no transaction risk, fraud risk, money laundering risk, and gambling risk, then the prediction result obtained by the classifier based on the features of the first training sample and the target features can be one of the above categories, such as no transaction risk.
[0071] S110: Train the target model according to the prediction result and the annotation of the first training sample.
[0072] Specifically, the loss can be determined based on the difference between the prediction result of the first training sample and the annotation of the first training sample, and the minimization of the loss can be used as the training objective to iteratively train the target model. The training termination condition of the target model can be that the number of iterations is greater than a preset number threshold, or the difference between the prediction result of the first training sample and the annotation of the first training sample is less than a preset difference threshold. This specification does not limit this.
[0073] After the target model is trained, since the target model itself is deployed on the first node, the first node can directly predict the data to be predicted based on the trained target model, thus eliminating the need to transfer the trained target model and avoiding the leakage of the model structure and model parameters of the target model.
[0074] The method of predicting the data to be predicted based on the target model can be to input the data to be predicted into the target model, obtain the features of the data to be predicted through the first encoder of the target model, and the classifier obtains the prediction result of the data to be predicted only based on the features of the data to be predicted. In addition, since a memory component is also deployed in the first node, and the features of the second training sample sent by the second node to the first node during the training process of the target model are stored in the memory component, that is, the features stored in the memory component can represent the information of the second training sample held by the second node. Therefore, when using the trained target model to predict the data to be predicted, the information of the second training sample can also be combined with the features of the data to be predicted. That is, the classifier obtains the prediction result of the data to be predicted based on the features of the data to be predicted output by the first encoder and the features stored in the memory component. In this way, not only the information of the data to be predicted is utilized, but also the information of the second training sample held by the second node is fully utilized, thereby expanding the feature dimension of the data to be predicted and improving the prediction accuracy of the data to be predicted.
[0075] In the model training method provided in this specification, when the first node receives the features of the second training sample sent by the second node, it writes the features of the second training sample into the memory component. When determining the features of the first training sample through the first encoder, it reads the target features associated with the features of the first training sample from the memory component according to the features of the first training sample, and then inputs the target features and the features of the first training sample into the classifier to obtain the prediction result of the first training sample, and trains the target model based on the prediction result and annotation of the first training sample.
[0076] It can be seen that the above solution is based on the idea of federated learning. By using the memory component to store the features from the second node, it effectively utilizes the information of the second training samples during the training process of the target model, while ensuring the privacy and security of the second training samples and the first training samples, and improving the privacy security and accuracy of the model.
[0077] In an optional embodiment of this specification, since a federated learning architecture with a central server and multiple nodes is not adopted in this specification, but instead the complete target model is deployed on the first node and the features of the second training samples sent by the second node are used to expand the dimension of the features of the first training samples. Therefore, the embodiments of this specification can be based on the idea of vertical federated learning (VFL). It is a solution for joint training through cross-samples in the case where the data features of the training samples held by each node (data holding method) overlap less and the user overlap is more, expanding the dimension and diversity of the features, thereby improving the performance of the target model. For this purpose, it is necessary to determine the first training samples and the second training samples respectively based on the intersection between the original data held by the first node and the original data held by the second node (more user overlap but less feature overlap), so as to perform joint training based on the original data held by the first node and the original data held by the second node. The following solution can be specifically adopted for implementation:
[0078] First, encrypt the identifier of the original data stored in the first node to obtain a first encrypted identifier.
[0079] As mentioned above, there is a situation where there is more user overlap and less data feature overlap between the original data stored in the first node and the original data stored in the second node. Among them, the information of the user in the original data can be determined by the identifier of the original data. For example, when the target model is a transaction risk identification model, the original data of the first node can be the transaction flow data of users stored by a bank, and the original data of the second node can be the transaction data of users stored by a merchant. Thus, in order to distinguish different transaction flow data, the bank can configure an identifier that can represent the user identity for each transaction flow data. Similarly, the merchant can configure an identifier that can represent the user identity of the user who executes the transaction for each transaction data. Thus, in order to use the cross-data with the same user but different data features as the training samples of the target model, the cross-data can be determined based on the overlap of the identifiers of the original data. The identifier of the original data can be information representing the user identity, such as the user's phone number, ID number, bank card number, etc., and this specification does not limit this.
[0080] Furthermore, based on the intersection between the identifiers of the original data stored in the first node and the identifiers of the original data stored in the second node, the intersection of the original data stored in the first node and the original data stored in the second node can be determined, thereby determining the first training sample and the second training sample for training the target model.
[0081] In order to compare the identifiers of the original data stored in the first node and the identifiers of the original data stored in the second node to determine the intersection, it is necessary to place the identifiers of the original data stored in the first node and the identifiers of the original data stored in the second node on the same node for comparison. That is, the process of determining the intersection of the identifiers needs to be carried out on a certain node, which can be the first node or the second node, and this specification does not make any restrictions on this. However, for the convenience of description, this specification takes the first node executing the identifier comparison process as an example to elaborate on the specific technical solution.
[0082] Since there may be a risk of data leakage during the process of the second node sending the identifiers of the original data to the first node, the identifiers sent by the second node are encrypted. In order to improve the efficiency of identifier comparison, the identifiers of the original data of the first node can also be encrypted. Therefore, in this step, the identifiers of the original data stored in the first node are encrypted to obtain the first encrypted identifier. The encryption process used to obtain the first encrypted identifier can be any existing encryption method, such as symmetric key, asymmetric key, hash processing, etc., and this specification does not make any restrictions on this.
[0083] Secondly, receive the second encrypted identifier sent by the second node, where the second encrypted identifier is obtained by the second node encrypting the identifiers of the original data stored in the second node.
[0084] As mentioned above, in order to prevent the identifiers of the original data stored in the second node from being leaked during transmission, the second node can encrypt the identifiers of the original data stored in the second node. This encryption process can be any existing encryption method, and the encryption method for the first node to obtain the first encrypted identifier can be the same or different from the encryption method for the second node to obtain the second encrypted identifier, and this specification does not make any restrictions on this.
[0085] Then, determine the intersection of the first encrypted identifier and the second encrypted identifier, and use the identifiers included in the intersection as the target identifiers.
[0086] Specifically, the second node should not only avoid the leakage of the original data, but also avoid the leakage of the identifiers of the original data. Similarly, the objects to be protected from leakage include not only attackers but also the first node. Therefore, the intersection of the first encrypted identifier and the second encrypted identifier can be obtained based on the method of Private Set Intersection (PSI).
[0087] An optional PSI method is as follows: The first node performs first encryption processing on the identifier of the original data stored by the first node to obtain a first encrypted identifier. The second node performs second encryption processing on the identifier of the original data stored by the second node to obtain a second encrypted identifier. The first node sends the first encrypted identifier to the second node. The second node encrypts the first encrypted identifier using the second encryption process to obtain a third encrypted identifier, and returns the second encrypted identifier and the third encrypted identifier to the first node. The first node processes the third encrypted identifier through the reverse process of the first encryption process to obtain a fourth encrypted identifier. At this time, the identifiers of the original data of the first node and the second node are both encrypted based on the second encryption process used by the second node. Therefore, by determining the intersection between the second encrypted identifier and the fourth encrypted identifier, the intersection of the first encrypted identifier and the second encrypted identifier can be determined.
[0088] Finally, according to the target identifier and the identifier of the original data stored by the first node, the first training sample is determined from the original data stored by the first node.
[0089] The target identifier represents the overlap of the original data held by the first node and the original data held by the second node in terms of users. The first training sample determined based on the target identifier is actually the overlapping partial transaction data of the users.
[0090] Similarly, the first node can send the target identifier to the second node, so that the second node determines the second training sample based on the target identifier. Thus, the first training sample and the second training sample are samples from different nodes that have a large overlap in terms of users and a small overlap in terms of data features. By jointly training the target model using the first training sample and the second training sample, the multi-dimensional data from different nodes can be effectively utilized, thereby improving the accuracy and generalization ability of the trained target model.
[0091] Furthermore, after the first training sample is determined based on the above solution, the second training sample of the second node also needs to be determined by the target identifier. Therefore, before the first node receives the features of the second training sample sent by the second node in the foregoing step S102, the first node may send the target identifier to the second node, so that the second node determines the second training sample from the original data stored in the second node according to the target identifier and the identifier of the original data stored in the second node, inputs the second training sample into the second encoder, obtains the features of the second training sample output by the second encoder, and sends them to the first node.
[0092] In addition, in one or more embodiments of the present specification, in order to further improve the performance of the trained target model, the target model can be jointly trained by the first node and multiple second nodes. Each second node determines the second training sample based on the data stored in itself, and determines the features of the second training sample through each second encoder deployed on each second node itself, and sends them to the first node. Thus, the first node needs to deploy multiple memory components, and each memory component corresponds to each second node one by one. Then, step S102 can be specifically as follows:
[0093] The first step: Obtain the corresponding relationship between each memory component and each second node.
[0094] Specifically, multiple memory components are deployed in the first node, and different memory components are used to store the features of the second training samples from different second nodes. For this reason, before training the target model, the first node can determine the number and identifier of the second nodes that need to be combined for training the target model, deploy the same number of memory components according to the number of second nodes, and establish the corresponding relationship between each memory component and each second node based on the identifier of the second node.
[0095] The second step: For each second node, receive the features of the second training sample sent by the second node.
[0096] This step is similar to the foregoing step S102 and will not be elaborated here.
[0097] The third step: According to the corresponding relationship, determine the memory component corresponding to the second node, and write the features of the second training sample into the memory component corresponding to the second node.
[0098] When receiving the features of the second training samples sent by the second node, the identifier of the second node as the sender can be obtained. Based on the corresponding relationship obtained in the first step and the identifier of the second node, the memory component of the second node is found from each memory component, so as to write the features of the second training samples sent by the second node into the memory component corresponding to the second node. The solution of writing the features of the second training samples into the memory component is similar to the foregoing step S102 and will not be elaborated here.
[0099] Furthermore, the foregoing step S104 can be specifically as follows:
[0100] First step: For each second node, according to the features of the first training sample, candidate features associated with the features of the first training sample are read from the features of the second training samples stored in the memory component corresponding to the second node.
[0101] This step is similar to the foregoing step S106 and will not be elaborated here.
[0102] Second step: According to the candidate features read from the memory components corresponding to the respective second nodes, target features associated with the features of the first training sample are determined.
[0103] Specifically, since the candidate features are respectively read from each memory component based on the features of the first training sample, the candidate features can be directly used as the target features associated with the features of the first training sample. Alternatively, when the number of candidate features is too large, the target features can be further screened based on the attention weights (degree of association) between the features of the first training sample and each candidate feature. For example, multiple candidate features with higher attention weights can be used as the target features. This specification does not make any limitations in this regard.
[0104] In an optional embodiment of this specification, since in the process of training the target model based on the solution as Figure 1 shown, it is possible that the first training sample is actually data with the same identifier as the original data stored by the second node. Therefore, not all the data stored by the first node are used as the first training sample. That is, compared with the full amount of data stored by the first node, the scale of the first training sample is smaller. Therefore, there may be a problem that the generalization performance of the target model is poor. For this reason, the following solution can be executed on the basis of the Figure 1 solution shown, so as to improve the performance of the target model, as Figure 3 shown:
[0105] S200: Determine the third training sample and the annotation of the third training sample according to the original data stored by the first node.
[0106] As described above, in practical applications, the first training sample and the second training sample can be cross samples with the same identifier. That is, the first training sample and the second training sample have a large overlap in users and a small overlap in data features. When there is a high degree of heterogeneity in the data, the number of cross samples is small, that is, the sample sizes of both the first training sample and the second training sample are small, making it uncertain whether the trained target model can exhibit high accuracy and generalization ability. Therefore, after the target model undergoes the Figure 1 training process shown, it can be tuned again based on the full amount of data held by the first node to compensate for the poor accuracy and generalization ability caused by the small size of the cross samples.
[0107] Therefore, in this step, the third training sample and its representation are determined based on the original data stored in the first node. The sample size of the third training sample is larger than that of the first training sample. The method for determining the annotation of the third training sample is similar to the method for determining the label of the first training sample in S100 described above, and will not be elaborated here.
[0108] For example, when the first node trains a target model as a transaction risk identification model, the first training sample and the second training sample used are cross samples. Therefore, the size of the first training sample is less than the size of the original data held by the first node. For example, if the first node is a bank, the bank may hold 100,000 transaction flow data, but only 10,000 transaction flow data can be used in the training of the transaction risk identification model. After the transaction risk identification model is trained, the first node can tune the trained transaction risk identification model again based on the 100,000 transaction flow data to improve the accuracy and generalization ability of the transaction risk identification model.
[0109] S202: Input the third training sample into the trained target model, and obtain the features of the third training sample through the first encoder of the target model.
[0110] Similar to the previous step S104, it will not be elaborated here.
[0111] However, it should be noted that since Figure 3 the scheme shown is to retune the trained target model, the target model used here is the target model trained based on the Figure 1 scheme shown.
[0112] S204: According to the features of the third training sample, read the target features associated with the features of the third training sample from the memory component.
[0113] Similar to the previous step S106, it will not be elaborated here.
[0114] It can be understood that since the target model has been trained and the complete, trained target model is deployed on the first node, in the Figure 3 scheme shown, there is no need for the second node to communicate with the first node, that is, the features stored in the memory component do not need to be updated.
[0115] S206: Input the features of the third training sample and the target features into the classifier of the target model to obtain the prediction result of the third training sample.
[0116] Similar to the foregoing step S108, details are not described herein again.
[0117] S208: Optimize the model parameters of the target model with the minimization of the difference between the prediction result of the third training sample and the annotation of the third training sample as the optimization objective.
[0118] Similar to the foregoing step S110, details are not described herein again.
[0119] In one or more embodiments of this specification, after the target model is obtained by combined training based on Figure 1 or Figure 1 and Figure 3 , the first node can, based on the trained target model, execute a prediction task in response to a prediction request. Specifically, it can be implemented according to the following steps, as Figure 4 shown:
[0120] S300: In response to a prediction request, determine the data to be predicted corresponding to the prediction request.
[0121] In the currently adopted federated learning method, there may be a situation where the target model is split into multiple sub-models and each sub-model is deployed on different nodes (data holders) for joint training. After model training, in order to avoid leakage of the model structure and model parameters, generally, the sub-models deployed on each node are not sent to the central server. Instead, the trained sub-models are still deployed on different nodes. When executing a prediction task, the intermediate results of model predictions are sequentially transmitted between the nodes according to the connection order of the sub-models, so as to obtain the prediction result of the data to be predicted.
[0122] It can be seen that in the current federated learning method, there are problems of complex deployment and large communication resource consumption during the prediction process. In this specification, since the target model is completely deployed on the first node during the training process, when the first node is trained and put into actual use, as long as a prediction request carrying the data to be predicted is sent to the first node, the first node can respond to the prediction request and directly obtain the prediction result of the data to be measured through the target model.
[0123] In this step, the data to be predicted is generally data related to the scenario applicable to the target model. For example, in the field of financial risk control, if the target model is a transaction risk identification model, the data to be predicted is the transaction data to be predicted, which includes transaction time, transaction amount, transaction object, etc. The purpose of the prediction request is to determine whether there is a transaction risk in the transaction data to be predicted or what kind of transaction risk exists through the transaction risk identification model.
[0124] S302: Input the data to be predicted into the first encoder of the trained target model to obtain the features of the data to be predicted output by the first encoder.
[0125] This step is similar to the previous step S104 and will not be elaborated here.
[0126] S304: According to the features of the data to be predicted, read the target features associated with the features of the data to be predicted from the features stored in the memory component; where the features stored in the memory component include the features of the second training samples written during the iterative training of the target model.
[0127] The memory component in the first node stores the features of the second training samples sent by the second node during the iterative training of the target model. Therefore, during the process of predicting the data to be predicted through the target model, the target features associated with the features of the data to be predicted can also be queried from the memory component, so as to obtain the prediction result of the data to be predicted by combining different data features. Therefore, in this step, the method based on the attention mechanism can also be used to use the features of the data to be predicted as the query vector to find the target features associated with the features of the data to be predicted from the features of the second training samples stored in the memory component.
[0128] In addition, if there are multiple memory components deployed in the first node, the candidate features associated with the features of the data to be predicted can be respectively queried from each memory component based on the features of the data to be predicted, and the target features associated with the features of the data to be predicted can be determined based on the candidate features.
[0129] The specific method of reading the target features associated with the features of the data to be predicted is similar to the previous step S106 and will not be elaborated here.
[0130] S306: Input the features of the data to be predicted and the target features into the classifier to obtain the prediction result of the data to be predicted output by the classifier.
[0131] This step is similar to the previous step S108 and will not be elaborated here.
[0132] Through the above solution, a complete target model is deployed on the first node. Thus, when executing a prediction request, there is no need to transmit the intermediate results of the data to be predicted between multiple nodes. Therefore, the target model will not be restricted by the system architecture (such as cross-network transmission) during the deployment phase, which not only saves communication resources, but also reduces the system risk of the model and improves the efficiency of the prediction task based on the confidence of the target model.
[0133] Figure 5 A schematic diagram of a model training device provided in this specification. The device is applied to a first node, and the first node deploys a target model to be trained. The target model includes a first encoder and a classifier. Specifically, it includes:
[0134] A first training sample determination module 400, configured to determine a first training sample and the annotation of the first training sample;
[0135] A writing module 402, configured to receive the features of a second training sample sent by a second node and write the features of the second training sample into a memory component deployed in the first node; wherein, the features of the second training sample are determined by a second encoder deployed in the second node;
[0136] A first feature determination module 404, configured to input the first training sample into the first encoder to obtain the features of the first training sample output by the first encoder;
[0137] A target feature determination module 406, configured to read, according to the features of the first training sample, target features associated with the features of the first training sample from the features of the second training sample stored in the memory component;
[0138] A prediction result determination module 408, configured to input the features of the first training sample and the target features into the classifier to obtain the prediction result of the first training sample output by the classifier;
[0139] A training module 410, configured to train the target model according to the prediction result and the annotation of the first training sample.
[0140] Optionally, the first training sample determination module 400 is specifically configured to encrypt the identifier of the original data stored in the first node to obtain a first encrypted identifier; receive a second encrypted identifier sent by a second node, where the second encrypted identifier is obtained by the second node encrypting the identifier of the original data stored in the second node; determine the intersection of the first encrypted identifier and the second encrypted identifier, and use the identifiers included in the intersection as target identifiers; and determine first training samples from the original data stored in the first node according to the target identifiers and the identifiers of the original data stored in the first node.
[0141] Optionally, the apparatus further includes:
[0142] A sending module 412, specifically configured to send the target identifier to the second node, so that the second node determines second training samples from the original data stored in the second node according to the target identifier and the identifier of the original data stored in the second node, input the second training samples into a second encoder to obtain features of the second training samples output by the second encoder, and send them to the first node.
[0143] Optionally, the writing module 402 is specifically configured to determine whether other features are stored in a memory component deployed in the first node; if so, update the features already stored in the memory component according to the features of the second training samples; if not, write the features of the second training samples into the memory component.
[0144] Optionally, the apparatus further includes:
[0145] An optimization module 414, specifically configured to determine third training samples and labels of the third training samples according to the original data stored in the first node; input the third training samples into a trained target model to obtain features of the third training samples through a first encoder of the target model; read target features associated with the features of the third training samples from the memory component according to the features of the third training samples; input the features of the third training samples and the target features into a classifier of the target model to obtain a prediction result of the third training samples; and optimize model parameters of the target model with the minimization of the difference between the prediction result of the third training samples and the labels of the third training samples as an optimization objective.
[0146] Optionally, multiple memory components are deployed in the first node, and each memory component corresponds to one of the second nodes respectively;
[0147] Optionally, the writing module 402 is specifically configured to obtain the correspondence between each memory component and each second node; for each second node, receive the features of the second training sample sent by the second node; determine the memory component corresponding to the second node according to the correspondence, and write the features of the second training sample into the memory component corresponding to the second node;
[0148] Optionally, the target feature determination module 406 is specifically configured to, for each second node, according to the features of the first training sample, read candidate features associated with the features of the first training sample from the features of the second training sample stored in the memory component corresponding to the second node; determine the target features associated with the features of the first training sample according to the candidate features read from the memory components corresponding to the respective second nodes.
[0149] Optionally, the apparatus further includes:
[0150] A prediction module 416, specifically configured to, in response to a prediction request, determine the data to be predicted corresponding to the prediction request; input the data to be predicted into the first encoder of the trained target model to obtain the features of the data to be predicted output by the first encoder; according to the features of the data to be predicted, read target features associated with the features of the data to be predicted from the features stored in the memory component; wherein the features stored in the memory component include the features of the second training sample written during the iterative training of the target model; input the features of the data to be predicted and the target features into a classifier to obtain the prediction result of the data to be predicted output by the classifier.
[0151] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 model training method shown.
[0152] This specification also provides Figure 6 a schematic structural diagram of the electronic device shown. As Figure 6 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 model training method shown. Of course, in addition to the software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0153] In the 1990s, it was obvious to distinguish whether an improvement in a technology was a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method processes). However, with the development of technology, many method process improvements today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method process into the hardware circuit. Therefore, it cannot be said that an improvement in a method process cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. The designer can program by himself to "integrate" a digital system on a PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, today, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). And there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that as long as the method process is slightly logically programmed with the above-mentioned several hardware description languages and programmed into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method process.
[0154] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0155] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0156] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0157] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0158] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0161] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0162] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0163] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0164] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0165] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0166] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0167] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0168] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A model training method, which is applied to a first node. The first node deploys a target model to be trained, and the target model includes a first encoder and a classifier; The method includes: Determine a first training sample and the annotation of the first training sample; Receive the features of a second training sample sent by a second node, and write the features of the second training sample into a memory component deployed in the first node; wherein, the features of the second training sample are determined by a second encoder deployed in the second node; Input the first training sample into the first encoder to obtain the features of the first training sample output by the first encoder; According to the features of the first training sample, read target features associated with the features of the first training sample from the features of the second training sample stored in the memory component; Input the features of the first training sample and the target features into the classifier to obtain the prediction result of the first training sample output by the classifier; Train the target model according to the prediction result and the annotation of the first training sample.
2. The method according to claim 1, wherein determining the first training sample specifically includes: Encrypt the identifier of the original data stored in the first node to obtain a first encrypted identifier; Receive a second encrypted identifier sent by the second node, wherein the second encrypted identifier is obtained by the second node encrypting the identifier of the original data stored in the second node; Determine the intersection of the first encrypted identifier and the second encrypted identifier, and use the identifiers included in the intersection as target identifiers; Determine the first training sample from the original data stored in the first node according to the target identifier and the identifier of the original data stored in the first node.
3. The method according to claim 2, before receiving the features of the second training sample sent by the second node, the method further includes: Send the target identifier to the second node, so that the second node determines a second training sample from the original data stored in the second node according to the target identifier and the identifier of the original data stored in the second node, input the second training sample into the second encoder to obtain the features of the second training sample output by the second encoder, and send them to the first node.
4. The method according to claim 1, wherein writing the features of the second training sample into a memory component deployed in the first node specifically includes: Judge whether other features are stored in the memory component deployed in the first node; If so, update the features already stored in the memory component according to the features of the second training sample; If not, write the features of the second training sample into the memory component.
5. The method according to claim 1, the method further includes: Determine a third training sample and the annotation of the third training sample according to the original data stored in the first node; Input the third training sample into the trained target model, and obtain the features of the third training sample through the first encoder of the target model; Read, according to the features of the third training sample, a target feature associated with the features of the third training sample from the memory component; Input the features of the third training sample and the target feature into the classifier of the target model to obtain the prediction result of the third training sample; Minimize the difference between the prediction result of the third training sample and the annotation of the third training sample as the optimization objective to optimize the model parameters of the target model.
6. The method according to claim 1, wherein a plurality of memory components are deployed on the first node, and each memory component corresponds to each second node one by one; Receiving the features of the second training sample sent by the second node and writing the features of the second training sample into the memory component deployed on the first node, specifically including: Obtain the corresponding relationship between each memory component and each second node; For each second node, receive the features of the second training sample sent by the second node; According to the corresponding relationship, determine the memory component corresponding to the second node, and write the features of the second training sample into the memory component corresponding to the second node; According to the features of the first training sample, read, from the features of the second training sample stored in the memory component, a target feature associated with the features of the first training sample, specifically including: For each second node, according to the features of the first training sample, read, from the features of the second training sample stored in the memory component corresponding to the second node, a candidate feature associated with the features of the first training sample; Determine a target feature associated with the features of the first training sample according to the candidate features read from the memory components corresponding to the respective second nodes.
7. The method according to claim 1, wherein the method further comprises: In response to a prediction request, determine the data to be predicted corresponding to the prediction request; Input the data to be predicted into the first encoder of the trained target model to obtain the features of the data to be predicted output by the first encoder; According to the features of the data to be predicted, read, from the features stored in the memory component, a target feature associated with the features of the data to be predicted; wherein the features stored in the memory component include the features of the second training sample written during the iterative training of the target model; Input the features of the data to be predicted and the target feature into the classifier to obtain the prediction result of the data to be predicted output by the classifier.
8. A model training device, the device is applied to a first node, the first node deploys a target model to be trained, and the target model includes a first encoder and a classifier; The device comprises: A first training sample determination module, configured to determine a first training sample and the annotation of the first training sample; A writing module, configured to receive the features of the second training sample sent by the second node and write the features of the second training sample into the memory component deployed on the first node; wherein the features of the second training sample are determined by a second encoder deployed on the second node; The first feature determination module is configured to input the first training sample into the first encoder to obtain the feature of the first training sample output by the first encoder; The target feature determination module is configured to, according to the feature of the first training sample, read, from the features of the second training samples stored in the memory component, the target features that have an association relationship with the feature of the first training sample; The prediction result determination module is configured to input the feature of the first training sample and the target feature into the classifier to obtain the prediction result of the first training sample output by the classifier; The training module is configured to train the target model according to the prediction result and the annotation of the first training sample.
9. A computer-readable storage medium storing a computer program, which when executed by a processor, implements the method according to any one of claims 1 to 7 above.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the method according to any one of claims 1 to 7 above.
Citation Information
Patent Citations
Model training method and device and keyword classification method and device
CN113887221A
Image processing method, device and equipment and readable storage medium
CN116524295A