Text model training method, recognition method, device, equipment and storage medium

Through the text model training method based on federated learning, the problems of data leakage and inaccurate model prediction are solved, and high-accuracy violation text prediction and reduction of model training time are achieved.

CN112734050BActive Publication Date: 2025-05-23PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011446681.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-11
Publication Date
2025-05-23
Estimated Expiration
2040-12-11

AI Technical Summary

Technical Problem

When the prior art uploads data sets to the cloud for model training, it is easy to cause data leakage, damage user security, and the training model predicts violations inaccurately.

Method used

The text model training method based on federated learning is adopted. By obtaining the data to be trained, the preset language model is trained, the model parameter information is encrypted and uploaded to the aggregated federated model for federated learning, and the preset language model is updated to generate a text model.

Benefits of technology

On the basis of protecting data privacy, implement joint training of multiple models, improve the accuracy of violation text prediction, and reduce model training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112734050B_ABST
    Figure CN112734050B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a training method for a text model based on federated learning, a text model recognition method, an apparatus, a computer device and a computer-readable storage medium. The method comprises: acquiring a set of data to be trained, training a preset language model based on the set of data to be trained, and obtaining model parameter information of the preset language model; encrypting the model parameter information and uploading it to a preset aggregated federated model to obtain aggregated model parameter information returned by the preset aggregated federated model after performing federated learning on the model parameter information; updating the preset language model based on the aggregated model parameter information to obtain a corresponding text model, thereby realizing joint training of multiple models on the basis of protecting data privacy, improving the accuracy of predicting illegal texts and reducing the training time of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a training method for a text model based on federated learning, a text model recognition method, apparatus, computer equipment, and computer-readable storage medium. Background Art

[0002] Identification of illegal content is widely used in the Internet world. The widespread dissemination of illegal content on the Internet will cause potential or obvious negative impacts and harm to the country and society. Therefore, how to quickly analyze and identify illegal content on the Internet has become a challenge faced by the industry. There are many carriers of illegal content, such as text, pictures, videos, audio, etc.

[0003] The traditional approach to detecting illegal content is to hire professionals to screen, label, and filter. Although AI filtering has been introduced and semantic recognition and classification technologies have been used, different enterprise platforms receive different illegal content. However, these illegal content data are difficult to achieve joint modeling due to privacy, insecurity, and the inability to be disseminated and shared. Summary of the invention

[0004] The main purpose of this application is to provide a text model training method based on federated learning, a text model recognition method, apparatus, computer equipment and computer-readable storage medium, aiming to solve the technical problems that in the existing process of uploading data sets to the cloud as model training data, data set leakage is prone to occur, user security is damaged, and the obtained training model predicts illegal content inaccurately.

[0005] In a first aspect, the present application provides a method for training a text model based on federated learning, the method for training a text model based on federated learning comprising the following steps:

[0006] Acquire the data of the to-be-trained set, train a preset language model based on the data of the to-be-trained set, and obtain model parameter information of the preset language model;

[0007] Encrypting the model parameter information and uploading it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information;

[0008] The preset language model is updated based on the aggregated model parameter information to obtain a corresponding text model.

[0009] In a second aspect, the present application provides a method for recognizing a text model based on federated learning, the method for recognizing a text model based on federated learning comprising the following steps:

[0010] Get the text to be predicted;

[0011] Based on the text encoding model and the text to be predicted, obtaining second text semantic vector information of the text to be predicted output by the text encoding model;

[0012] Based on the text recognition model and the second text semantic vector information, obtaining label information of the second text semantic vector information output by the text recognition model;

[0013] According to the label information, it is determined whether the text to be predicted violates the rules, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0014] In a third aspect, the present application further provides a training device for a text model based on federated learning, wherein the training device for a text model based on federated learning comprises:

[0015] A first acquisition module is used to acquire the to-be-trained set data, train a preset language model based on the to-be-trained set data, and obtain model parameter information of the preset language model;

[0016] A second acquisition module is used to encrypt the model parameter information and upload it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information;

[0017] A generation module is used to update the preset language model based on the aggregated model parameter information to obtain a corresponding text model.

[0018] In a fourth aspect, the present application further provides a training device for a text model based on federated learning, wherein the training device for a text model based on federated learning comprises:

[0019] A first acquisition module, used for acquiring the text to be predicted;

[0020] A second acquisition module, configured to acquire, based on the text encoding model and the text to be predicted, second text semantic vector information of the text to be predicted output by the text encoding model;

[0021] A third acquisition module, configured to acquire label information of the second text semantic vector information output by the text recognition model based on the text recognition model and the second text semantic vector information;

[0022] A determination module is used to determine whether the text to be predicted violates the rules based on the label information, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0023] In a fifth aspect, the present application also provides a computer device, comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the training method of the text model based on federated learning and the steps of the text recognition method based on federated learning are implemented.

[0024] In a sixth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the training method of the text model based on federated learning and the steps of the text recognition method based on federated learning are implemented.

[0025] The present application provides a training method for a text model based on federated learning, a text model recognition method, an apparatus, a computer device, and a computer-readable storage medium. The method comprises the following steps: obtaining a set of data to be trained, training a preset language model based on the set of data to be trained, and obtaining model parameter information of the preset language model; encrypting the model parameter information and uploading it to a preset aggregated federated model to obtain aggregated model parameter information returned by the preset aggregated federated model after performing federated learning on the model parameter information; updating the preset language model based on the aggregated model parameter information to obtain a corresponding text model, thereby realizing joint training of multiple models on the basis of protecting data privacy, improving the accuracy of predicting illegal texts, and reducing the training time of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0027] Figure 1 A flowchart of a method for training a text model based on federated learning provided in an embodiment of the present application;

[0028] Figure 2 for Figure 1 A flowchart of the sub-steps of the training method of the text model based on federated learning in ;

[0029] Figure 3 It is a schematic diagram of encrypting a plurality of first model parameter information and a plurality of second model parameter information and uploading them to a preset aggregate federated model provided by an embodiment of the present application;

[0030] Figure 4 for Figure 1A flowchart of the sub-steps of the training method of the text model based on federated learning in ;

[0031] Figure 5 A flowchart of a text model recognition method based on federated learning provided in an embodiment of the present application;

[0032] Figure 6 A schematic block diagram of a training device for a text model based on federated learning provided in an embodiment of the present application;

[0033] Figure 7 A schematic block diagram of a text model recognition device based on federated learning provided in an embodiment of the present application;

[0034] Figure 8 The present invention is a block diagram showing the structure of a computer device according to an embodiment of the present application.

[0035] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0037] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0038] The embodiments of the present application provide a method for training a text model based on federated learning, a method for identifying a text model, an apparatus, a computer device, and a computer-readable storage medium. The method for training a text model based on federated learning and the method for identifying a text model based on federated learning can be applied to a computer device, which can be an electronic device such as a laptop computer, a desktop computer, or a server.

[0039] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0040] Please refer to Figure 1 , Figure 1 A flowchart of a method for training a text model based on federated learning is provided in an embodiment of the present application.

[0041] like Figure 1 As shown, the training method of the text model based on federated learning includes steps S101 to S103.

[0042] Step S101: Acquire a set of data to be trained, train a preset language model based on the set of data to be trained, and obtain model parameter information of the preset language model.

[0043] Exemplarily, a set of data to be trained is obtained, the set of data to be trained includes a plurality of texts to be trained, wherein the texts to be trained include illegal content, and the set of data to be trained is stored as a preset storage path or a preset blockchain. When the set of data to be trained is obtained, a preset language model is trained by the set of data to be trained to obtain model parameter information of the preset language model, wherein the preset language model includes a preset neural network model, wherein the preset language model is multiple, the specific number is not limited, and the preset language model is located at the user end.

[0044] In one embodiment, specifically, referring to Figure 2 , step S101 includes: sub-step S1011 to sub-step S1022.

[0045] Sub-step S1011, training the preset pre-trained language model based on the text to be trained, obtaining the first semantic vector information corresponding to the text to be trained output by the preset pre-trained language model, and obtaining the first model parameter information of the preset pre-trained language model after training.

[0046] Exemplarily, the preset language model includes a preset pre-trained language model and a preset dual propagation model, and the model parameters of the preset language model include a first model parameter and a second model parameter. Figure 3The pre - set pre - trained language model is shown. Obtain the first semantic vector information output by the pre - set pre - trained language model for the text to be trained. Among them, the text to be trained includes illegal words and the label values annotated for the illegal words. The label values can be 1 - 10, or can also be numerical values between 1 - 100. For example, the label values annotated for the illegal words are 1, 5, 10, etc.; or the label values are 1, 20, 50, 100, etc. The pre - set pre - trained language model is a pre - set BERT model. The full name of the BERT model is Bidirectional Encoder Representations from Transformer. The role of the BERT model is to obtain the text containing rich semantic information, that is, to represent the semantics of the text. Among them, the pre - set pre - trained language model is at the user side, and at least one pre - trained language model can be set on one user side. Extract the words in the text to be trained through the hidden layer of the pre - set pre - trained language model, obtain the semantic vector information of the word through the weight matrix of the hidden layer, use the semantic vector information of the illegal word as the first semantic vector information, and output it through the output layer.

[0047] After obtaining the first semantic vector information of the text to be trained output by the pre - set pre - trained model, obtain the first model parameter information of the current pre - set pre - trained model. Extract the features of the text to be trained through the network layer in the pre - set pre - trained model to obtain the gradient value of the text to be trained. For example, obtain the vector feature information of the word through the full weight matrix of the hidden layer in the pre - set pre - trained language model, and obtain the vector feature information of the word through the full weight matrix of the hidden layer in the pre - set pre - trained language model, and obtain the corresponding gradient value according to the vector feature information. Update the model parameters of the pre - set pre - trained language model through the gradient value of the text to be trained to obtain the updated first model parameter information of the pre - set pre - trained language model. Among them, when there are multiple pre - set pre - trained language models, respectively obtain the first model parameter information and the first semantic vector information of each pre - set pre - trained language model.

[0048] Sub - step S1021: Train the pre - set double - propagation model based on the first semantic vector information to obtain the second model parameter information of the pre - set double - propagation model after training.

[0049] Exemplarily, the preset dual propagation model is a BiLSTM model (Bi-directional Long Short-Term Memory), which is composed of a forward LSTM and a backward LSTM. After obtaining the first semantic vector information corresponding to the text to be trained output by the preset pre-trained language model, the preset dual propagation model is trained by the first semantic vector information to obtain the second model parameter information of the preset dual propagation model after training. For example, the first semantic vector information is feature extracted by the network layer in the preset dual propagation model to obtain the gradient value of the label value corresponding to the first semantic vector information. For example, the vector feature information of the label value is obtained by the full weight matrix of the hidden layer in the preset dual propagation model, and the vector feature information of the label value is obtained by the full weight matrix of the hidden layer in the preset dual propagation model, and the corresponding gradient value is obtained according to the vector feature information. The model parameters of the preset dual propagation model are updated by the gradient value of the text to be trained to obtain the updated second model parameter information of the preset dual propagation model. Wherein, when there are multiple preset dual propagation models, the second model parameter information of each preset dual propagation model is obtained respectively.

[0050] Step S102: encrypt the model parameter information and upload it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information.

[0051] Exemplarily, the preset aggregated federated model is located in the server, an upload request is sent to the server, an encrypted public key is received from the server, the model parameters of each preset language model are encrypted by the encrypted public key, and the encrypted model parameters are sent to the server. When the server receives the encrypted model parameters, it decrypts each encrypted model parameter respectively to obtain the model parameters of each preset language model after decryption. Each model parameter is learned by the preset aggregated federated model in the server to obtain the corresponding aggregated model parameters, and the obtained aggregated model parameters are returned to each preset language model. Among them, the aggregated federated model includes types such as aggregated horizontal federated model, aggregated vertical federated model and aggregated federated migration model.

[0052] It should be noted that federated learning refers to a method of machine learning modeling by uniting different clients or participants. In federated learning, the client does not need to expose its own data to other clients and coordinators (also called servers), so federated learning can well protect user privacy and data security, and can solve the problem of data silos. Federated learning has the following advantages: data isolation, data will not be leaked to the outside, meeting the needs of user privacy protection and data security; it can ensure that the quality of the federated learning model is lossless, there will be no negative transfer, and the federated learning model is better than the fragmented independent model; it can ensure that each client can perform encrypted exchange of information and model parameters while maintaining independence, and grow at the same time.

[0053] In a real-time example, the model parameter information includes first model parameter information and second model parameter information; the encrypting the model parameter information and uploading it to a preset aggregate federated model to obtain the aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information, includes: encrypting the first model parameter information and uploading it to a preset aggregate federated model, obtaining the first aggregate model parameter information returned by the preset aggregate federated model after performing horizontal federated learning on the first model parameter information; encrypting the second model parameter information and uploading it to a preset aggregate federated model, obtaining the second aggregate model parameter information returned by the preset aggregate federated model after performing horizontal federated learning on the second model parameter information.

[0054] Exemplarily, a public key sent by a server is received, wherein the number of the public keys is multiple. For example, when the number of the public keys is two, namely, a first public key and a second public key. The first model parameter information of each preset pre-trained language model and the second model parameter information of each preset dual propagation model are encrypted respectively by the received public keys. For example, when the first public key and the second public key are received, the first public key and the second public key are used to encrypt the first model parameter information of each preset pre-trained language model and the second model parameter information of each preset dual propagation model.

[0055] After encrypting the first model parameter information of each preset pre-trained language model and the second model parameter information of each preset dual propagation model by a public key, as shown in FIG. Figure 3As shown, each preset pre-trained language model and the preset dual propagation model adopt a construction method of inadvertent transmission to establish a secret communication channel, and send the encrypted first model parameter information of each preset pre-trained language model and the second model parameter of each preset dual propagation model to the server through the secret communication channel. When the first public key and the second public key are used to encrypt the first model parameter information of each preset pre-trained language model, and the first public key and the second public key are used to encrypt the second model parameter information of each preset dual propagation model, the first model parameter information of each preset pre-trained language model encrypted by the first public key and the second model parameter of each preset pre-trained language model encrypted by the second public key, as well as the second model parameter of each preset dual propagation model encrypted by the first public key and the second model parameter information of each preset dual propagation model encrypted by the second public key are sent to the server through the secret communication channel.

[0056] The server decrypts the first model parameter information of each preset pre-trained language model and the second model parameter information of each preset dual propagation model after receiving encryption. For example, when receiving the first model parameter information of each preset pre-trained language model encrypted by the first public key and the first model parameter information of each preset pre-trained language model encrypted by the second public key, as well as the first model parameter information of each preset dual propagation model encrypted by the first public key and the first model parameter information of each preset dual propagation model encrypted by the second public key, the server randomly decrypts the first model parameter information of each preset pre-trained language model encrypted by the first public key and the first model parameter information of each preset pre-trained language model encrypted by the second public key, as well as the first model parameter information of each preset dual propagation model encrypted by the first public key and the first model parameter information of each preset dual propagation model encrypted by the second public key through the private key. Among them, the private key corresponds to the first public key or the second public key, that is, the private key decrypts the first public key or decrypts the second public key. The first model parameter information of the preset pretrained language model encrypted by the first public key and the first model parameter information of each preset pretrained language model encrypted by the second public key are decrypted by the private key to obtain the first model parameter information of each preset pretrained language model, and the first model parameter information of each preset dual propagation model is decrypted by the private key and the first model parameter information of each preset dual propagation model encrypted by the second public key to obtain the first model parameter information of each preset dual propagation model.

[0057] The parameters corresponding to the intersection features of the first model parameter information of each preset pre-trained language model are learned through the horizontal federated learning mechanism in the server, the corresponding first aggregate model parameters are obtained by averaging the parameters corresponding to the intersection features, and the first aggregate model parameters are returned to each preset pre-trained language model. The parameters corresponding to the intersection features of the second model parameter information of each preset dual propagation model are learned through the horizontal federated learning mechanism in the server, the corresponding second aggregate model parameters are obtained by averaging the parameters corresponding to the intersection features, and the second aggregate model parameters are returned to each preset dual propagation model.

[0058] Step S103: Update the preset language model based on the aggregated model parameter information to obtain a corresponding text model.

[0059] In an exemplary embodiment, after receiving the aggregate model parameter information returned by the aggregate federated model, the model parameter information of each preset language model is updated by using the aggregate model parameter information, and the corresponding text model is generated for each preset language model after the aggregate model parameter information is updated.

[0060] In one embodiment, specifically, referring to Figure 4 , step S103 includes: sub-step S1031 to sub-step S1032.

[0061] Sub-step S1031: updating the first model parameter information of the preset pre-trained language model based on the first aggregated model parameter information to generate a corresponding text encoding model.

[0062] In an exemplary embodiment, the first model parameter information of a preset pre-trained language model is updated by the first aggregate model parameter information returned by the aggregate federated model, and the updated preset pre-trained language model generates a corresponding text encoding model. Wherein, when there are multiple preset pre-trained language models, when they are respectively in different user terminals, the second model parameter information of each preset pre-trained language model is updated by the first aggregate model parameter information returned by the aggregate federated model, and the updated preset pre-trained language models respectively generate corresponding text encoding models.

[0063] Sub-step S1032: updating the second model parameters of the preset dual propagation model based on the second aggregation model parameters to generate a corresponding text recognition model.

[0064] In an exemplary embodiment, the second model parameter information of a preset dual propagation model is updated by using the second aggregate model parameter information returned by the aggregate federated model, and a corresponding text recognition model is generated from the updated preset dual propagation model. When there are multiple preset dual propagation models and they are respectively located at different user terminals, the second model parameter information of each preset dual propagation model is updated by using the second aggregate model parameter information returned by the aggregate federated model, and corresponding text recognition models are generated from each updated preset dual propagation model.

[0065] In a real-time example, before generating a corresponding text encoding model and / or generating a corresponding text recognition model, it includes: determining whether the preset pre-trained language model and / or the preset dual propagation model is in a convergence state; if it is determined that the preset pre-trained language model and / or the preset dual propagation model is in a convergence state, using the preset pre-trained language model as a text encoding model and / or using the preset dual propagation model as a text recognition model; if the preset pre-trained language model and / or the preset dual propagation model is not in a convergence state, training the preset pre-trained language model and / or the preset dual propagation model according to preset sample data to be trained, and obtaining third model parameter information of the preset pre-trained language model and / or fourth model parameter information of the preset dual propagation model after training.

[0066] Exemplarily, determine whether the preset pre-trained language model and / or the preset dual propagation model is in a convergence state. For example, compare the first aggregate model parameter information with the previously recorded first aggregate model parameter information. If the first aggregate model parameter information is the same as the previously recorded first aggregate model parameter information, or the difference between the first aggregate model parameter information and the previously recorded first aggregate model parameter information is less than the preset difference, then determine that the preset pre-trained language model is in a convergence state; and / or, compare the second aggregate model parameter information with the previously recorded second aggregate model parameter information. If the second aggregate model parameter information is the same as the previously recorded second aggregate model parameter information, or the difference between the second aggregate model parameter information and the previously recorded second aggregate model parameter information is less than the preset difference, then determine that the preset dual propagation model is in a convergence state.

[0067] For example, the first aggregation model parameter information is compared with the previously recorded first aggregation model parameter information. If the first aggregation model parameter information is different from the previously recorded first aggregation model parameter information, or the difference between the first aggregation model parameter information and the previously recorded first aggregation model parameter information is greater than or equal to the preset difference, it is determined that the preset pre-trained language model is not in a converged state; and / or, the second aggregation model parameter information is compared with the previously recorded second aggregation model parameter information. If the second aggregation model parameter information is different from the previously recorded second aggregation model parameter information, or the difference between the second aggregation model parameter information and the previously recorded second aggregation model parameter information is greater than or equal to the preset difference, it is determined that the preset dual propagation model is not in a converged state.

[0068] If it is determined that the preset pre-trained language model is in a convergence state, the preset pre-trained language model is used as a text encoding model; and / or, if it is determined that the preset dual propagation model is in a convergence state, the preset dual propagation model is used as a text recognition model.

[0069] If it is determined that the preset pre-trained language model is not in a convergence state, the preset pre-trained language model is continued to be trained according to the preset sample data to be trained, and the third model parameter information and second semantic vector information of the preset pre-trained language model after training are obtained; and / or, if it is determined that the preset dual propagation model is not in a convergence state, the preset dual propagation model is continued to be trained according to the second semantic vector information to obtain the fourth model parameter information after training, and the third model parameter information and / or the fourth model parameter information are uploaded to the aggregated federated model for federated learning.

[0070] In an embodiment of the present invention, a preset language model is trained using the data of a training set to obtain model parameter information of the preset language model, the model parameter information is federated learned using an aggregated federated model to obtain aggregated model parameter information, and the model parameter information of the preset language model is updated using the aggregated model parameter information to generate a corresponding text model, thereby achieving joint training of multiple models on the basis of protecting data privacy, improving the accuracy of predicting illegal texts, and reducing the training time of the model.

[0071] Please refer to Figure 5 , Figure 5 A flowchart of a text model recognition method based on federated learning is provided in an embodiment of the present application.

[0072] like Figure 5 As shown, the text model recognition method based on federated learning includes steps S201 to S204.

[0073] Step S201: Obtain text to be predicted.

[0074] Exemplarily, a text to be predicted is obtained, where the text to be predicted contains illegal words or non-illegal words and is a sentence or short sentence sent by a user detected through a network.

[0075] Step S202: Based on the text encoding model and the text to be predicted, obtain second text semantic vector information of the text to be predicted output by the text encoding model.

[0076] Exemplarily, semantic prediction is performed on the text to be predicted through the text encoding model to obtain the second text semantic vector information of the text to be predicted. For example, the semantic vectors of each word in the text to be predicted are extracted through the hidden layer of the text encoding model, and the obtained semantic vectors are combined to obtain the second text semantic vector information of the text to be predicted.

[0077] Step S203: Based on the text recognition model and the second text semantic vector information, obtain label information of the second text semantic vector information output by the text recognition model.

[0078] Exemplarily, the second text semantic vector information is predicted by a text recognition model to obtain label information of the second text semantic vector information. For example, the semantic vectors of each word in the second text semantic vector information are extracted by a hidden layer of the text recognition model, and the semantic vectors of each word are mapped to obtain label information of the second text semantic vector information.

[0079] Step S204: determining whether the text to be predicted violates the rules based on the label information, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0080] Exemplarily, when the label information is obtained, it is determined whether the text to be predicted is illegal based on the label information. For example, when the label information is a label value, the label value is compared with a preset label value. If the label value is greater than or equal to the preset label value, it is determined that the text to be predicted is illegal content; if it is determined that the label value is less than the preset label value, it is determined that the text to be predicted is not illegal content, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0081] In an embodiment of the present invention, second text semantic vector information of a text to be predicted is obtained through a text encoding model, label information of the second text semantic vector information is obtained through a text recognition model, and whether the text to be predicted is illegal content is determined through the label information, wherein the text encoding model and the text recognition model are both obtained through federated learning, thereby improving the accuracy of the text encoding model and the text recognition model.

[0082] Please refer to Figure 6 , Figure 6 A schematic block diagram of a training device for a text model based on federated learning provided in an embodiment of the present application.

[0083] like Figure 6 As shown, the training device 400 of the text model based on federated learning includes: a first acquisition module 401, a second acquisition module 402, and a generation module 403.

[0084] The first acquisition module 401 is used to acquire the data of the to-be-trained set, train the preset language model based on the data of the to-be-trained set, and obtain model parameter information of the preset language model;

[0085] A second acquisition module 402 is used to encrypt the model parameter information and upload it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information;

[0086] The generation module 403 is used to update the preset language model based on the aggregated model parameter information to obtain a corresponding text model.

[0087] The first acquisition module 401 is further configured to:

[0088] Training the preset pre-trained language model based on the text to be trained, obtaining first semantic vector information corresponding to the text to be trained output by the preset pre-trained language model, and obtaining first model parameter information of the preset pre-trained language model after training;

[0089] The preset dual propagation model is trained based on the first semantic vector information, and second model parameter information of the preset dual propagation model after training is obtained.

[0090] The second acquisition module 402 is further configured to:

[0091] Encrypting the first model parameter information and uploading it to a preset aggregate federated model, and obtaining first aggregate model parameter information returned by the preset aggregate federated model after performing horizontal federated learning on the first model parameter information;

[0092] The second model parameter information is encrypted and uploaded to a preset aggregate federated model, and the second aggregate model parameter information returned by the preset aggregate federated model after horizontal federated learning of the second model parameter information is obtained.

[0093] The generating module 403 is further used for:

[0094] Based on the first aggregated model parameter information, first model parameter information of the preset pre-trained language model is updated to generate a corresponding text encoding model;

[0095] The second model parameters of the preset dual propagation model are updated based on the second aggregation model parameters to generate a corresponding text recognition model.

[0096] The generating module 403 is further used for:

[0097] Determining whether the preset pre-trained language model and / or the preset dual propagation model is in a convergence state;

[0098] If it is determined that the preset pre-trained language model and / or the preset dual propagation model is in a convergence state, the preset pre-trained language model is used as a text encoding model and / or the preset dual propagation model is used as a text recognition model;

[0099] If the preset pre-trained language model and / or the preset dual propagation model is not in a convergence state, the preset pre-trained language model and / or the preset dual propagation model is trained according to the preset sample data to be trained to obtain the third model parameter information of the preset pre-trained language model and / or the fourth model parameter information of the preset dual propagation model after training.

[0100] It should be noted that those skilled in the art can clearly understand that, for the sake of convenience and conciseness of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned training method embodiment of the text model based on federated learning, and will not be repeated here.

[0101] Please refer to Figure 7 , Figure 7 A schematic block diagram of a text model recognition device based on federated learning provided in an embodiment of the present application.

[0102] like Figure 7 As shown, the text model recognition device 500 based on federated learning includes: a first acquisition module 501, a second acquisition module 502, a third acquisition module 503, and a determination module 504.

[0103] A first acquisition module 501 is used to acquire the text to be predicted;

[0104] A second acquisition module 502 is used to acquire second text semantic vector information of the text to be predicted output by the text encoding model based on the text encoding model and the text to be predicted;

[0105] A third acquisition module 503 is used to acquire label information of the second text semantic vector information output by the text recognition model based on the text recognition model and the second text semantic vector information;

[0106] The determination module 504 is used to determine whether the text to be predicted violates the rules according to the label information, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0107] It should be noted that technical personnel in the relevant field can clearly understand that, for the convenience and conciseness of description, the specific working process of the above-described device and each module and unit can refer to the corresponding process in the aforementioned embodiment of the recognition method of the text model based on federated learning, and will not be repeated here.

[0108] The apparatus provided in the above embodiment may be implemented in the form of a computer program. The computer program may be Figure 8 Runs on the computer device shown.

[0109] See also Figure 8 , Figure 8 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. The computer device may be a terminal.

[0110] like Figure 8 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0111] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the training methods of the text model based on federated learning and the recognition method of the text model based on federated learning.

[0112] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0113] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the training methods of the text model based on federated learning and the recognition method of the text model based on federated learning.

[0114] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 8The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0115] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0116] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0117] Acquire the data of the to-be-trained set, train a preset language model based on the data of the to-be-trained set, and obtain model parameter information of the preset language model;

[0118] Encrypting the model parameter information and uploading it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information;

[0119] The preset language model is updated based on the aggregated model parameter information to obtain a corresponding text model.

[0120] In one embodiment, the processor said to-be-trained set data includes to-be-trained text, said preset language model includes a preset pre-trained language model and a preset dual propagation model, and said model parameter information includes first model parameter information and second model parameter information;

[0121] The method of training a preset language model based on the to-be-trained set data to obtain model parameter information of the preset language model is used to implement:

[0122] Training the preset pre-trained language model based on the text to be trained, obtaining first semantic vector information corresponding to the text to be trained output by the preset pre-trained language model, and obtaining first model parameter information of the preset pre-trained language model after training;

[0123] The preset dual propagation model is trained based on the first semantic vector information, and second model parameter information of the preset dual propagation model after training is obtained.

[0124] In one embodiment, the processor said model parameter information includes first model parameter information and second model parameter information;

[0125] The encrypting of the model parameter information and uploading it to the preset aggregate federated model to obtain the aggregate model parameter information returned by the preset aggregate federated model after the model parameter information is federated learned is used to achieve:

[0126] Encrypting the first model parameter information and uploading it to a preset aggregate federated model, and obtaining first aggregate model parameter information returned by the preset aggregate federated model after performing horizontal federated learning on the first model parameter information;

[0127] The second model parameter information is encrypted and uploaded to a preset aggregate federated model, and the second aggregate model parameter information returned by the preset aggregate federated model after horizontal federated learning of the second model parameter information is obtained.

[0128] In one embodiment, the processor said preset language model includes a preset pre-trained language model and a preset dual propagation model, and the text model includes a text encoding model and a text recognition model;

[0129] When the preset language model to be trained is updated based on the aggregated model parameter information to obtain the corresponding text model implementation, it is used to implement:

[0130] Based on the first aggregated model parameter information, first model parameter information of the preset pre-trained language model is updated to generate a corresponding text encoding model;

[0131] The second model parameters of the preset dual propagation model are updated based on the second aggregation model parameters to generate a corresponding text recognition model.

[0132] In one embodiment, the processor is used to implement the following before generating the corresponding text encoding model and / or generating the corresponding text recognition model:

[0133] Determining whether the preset pre-trained language model and / or the preset dual propagation model is in a convergence state;

[0134] If it is determined that the preset pre-trained language model and / or the preset dual propagation model is in a convergence state, the preset pre-trained language model is used as a text encoding model and / or the preset dual propagation model is used as a text recognition model;

[0135] If the preset pre-trained language model and / or the preset dual propagation model is not in a convergence state, the preset pre-trained language model and / or the preset dual propagation model is trained according to the preset sample data to be trained to obtain the third model parameter information of the preset pre-trained language model and / or the fourth model parameter information of the preset dual propagation model after training.

[0136] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0137] Get the text to be predicted;

[0138] Based on the text encoding model and the text to be predicted, obtaining second text semantic vector information of the text to be predicted output by the text encoding model;

[0139] Based on the text recognition model and the second text semantic vector information, obtaining label information of the second text semantic vector information output by the text recognition model;

[0140] According to the label information, it is determined whether the text to be predicted violates the rules, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning.

[0141] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. The computer program includes program instructions. The method implemented when the program instructions are executed can refer to the various embodiments of the present application's training method for a text model based on federated learning and the recognition method for a text model based on federated learning.

[0142] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.

[0143] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0144] The blockchain referred to in the present invention is a new application mode of computer technologies such as storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. of preset pre-trained language models, preset dual propagation models, text encoding models and text recognition models. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, platform product service layer, and application service layer.

[0145] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0146] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A training method for a text model based on federated learning, characterized in that, it includes: Obtain the data of the training set to be trained, and train a pre-set language model based on the data of the training set to be trained to obtain the model parameter information of the pre-set language model; Encrypt and upload the model parameter information to a pre-set aggregated federated model to obtain the aggregated model parameter information returned by the pre-set aggregated federated model after performing federated learning on the model parameter information; Update the pre-set language model based on the aggregated model parameter information to obtain a corresponding text model; Wherein, the pre-set language model includes a pre-set pre-trained language model and a pre-set double-propagation model, and the aggregated model parameter information includes first aggregated model parameter information and second aggregated model parameter information; updating the pre-set language model based on the aggregated model parameter information to obtain a corresponding text model includes: Updating the first model parameter information of the pre-set pre-trained language model based on the first aggregated model parameter information, and updating the second model parameter information of the pre-set double-propagation model based on the second aggregated model parameter information; Determine whether the updated pre-set pre-trained language model and / or the updated pre-set double-propagation model is in a converged state; If it is determined that the pre-set pre-trained language model and / or the pre-set double-propagation model is in a converged state, then use the pre-set pre-trained language model as a text encoding model and / or use the pre-set double-propagation model as a text recognition model, wherein the text encoding model is used to determine the second text semantic vector information of the text to be predicted, and the text recognition model is used to determine the label information of the text to be predicted according to the second text semantic vector information; If the pre-set pre-trained language model and / or the pre-set double-propagation model is not in a converged state, then train the pre-set pre-trained language model and / or the pre-set double-propagation model according to pre-set training sample data to obtain the third model parameter information of the pre-trained pre-set language model and / or the fourth model parameter information of the pre-set double-propagation model after training, and upload the third model parameter information and / or the fourth model parameter information to the pre-set aggregated federated model for federated learning.

2. The training method for a text model based on federated learning according to claim 1, characterized in that, the data of the training set to be trained includes texts to be trained; training the pre-set language model based on the data of the training set to be trained to obtain the model parameter information of the pre-set language model includes: Training the pre-set pre-trained language model based on the text to be trained, obtaining the first semantic vector information corresponding to the text to be trained output by the pre-set pre-trained language model, and obtaining the first model parameter information of the pre-set pre-trained language model after training; Training the pre-set double-propagation model based on the first semantic vector information to obtain the second model parameter information of the pre-set double-propagation model after training.

3. The training method for a text model based on federated learning according to claim 1, characterized in that, The model parameter information of the preset language model includes first model parameter information and second model parameter information; The encrypting the model parameter information of the preset language model and uploading it to the preset aggregated federated model to obtain the aggregated model parameter information returned by the preset aggregated federated model after the preset aggregated federated model performs federated learning on the model parameter information of the preset language model, includes: Encrypting the first model parameter information and uploading it to a preset aggregate federated model, and obtaining first aggregate model parameter information returned by the preset aggregate federated model after performing horizontal federated learning on the first model parameter information; The second model parameter information is encrypted and uploaded to a preset aggregate federated model, and the second aggregate model parameter information returned by the preset aggregate federated model after horizontal federated learning of the second model parameter information is obtained.

4. A text model recognition method based on federated learning, It is characterized in that include: Get the text to be predicted; Based on the text encoding model and the text to be predicted, obtaining second text semantic vector information of the text to be predicted output by the text encoding model; Based on the text recognition model and the second text semantic vector information, obtaining label information of the second text semantic vector information output by the text recognition model; According to the label information, it is determined whether the text to be predicted violates the rules, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning as described in any one of claims 1-3.

5. A training device for a text model based on federated learning, It is characterized in that include: A first acquisition module is used to acquire the to-be-trained set data, train a preset language model based on the to-be-trained set data, and obtain model parameter information of the preset language model; A second acquisition module is used to encrypt the model parameter information and upload it to a preset aggregate federated model to obtain aggregate model parameter information returned by the preset aggregate federated model after performing federated learning on the model parameter information; A generation module, configured to update the preset language model based on the aggregate model parameter information to obtain a corresponding text model; The preset language model includes a preset pre-trained language model and a preset dual propagation model, and the aggregation model parameter information includes first aggregation model parameter information and second aggregation model parameter information; and the generation module is further used for: Updating first model parameter information of the preset pre-trained language model based on the first aggregate model parameter information, and updating second model parameter information of the preset dual propagation model based on the second aggregate model parameter information; Determining whether the updated preset pre-trained language model and / or the updated preset dual propagation model is in a convergence state; If it is determined that the preset pre-trained language model and / or the preset dual propagation model is in a convergence state, the preset pre-trained language model is used as a text encoding model and / or the preset dual propagation model is used as a text recognition model, wherein the text encoding model is used to determine the second text semantic vector information of the text to be predicted, and the text recognition model is used to determine the label information of the text to be predicted according to the second text semantic vector information; If the preset pre-trained language model and / or the preset dual propagation model is not in a convergence state, the preset pre-trained language model and / or the preset dual propagation model is trained according to the preset sample data to be trained, and the third model parameter information of the preset pre-trained language model and / or the fourth model parameter information of the preset dual propagation model after training is obtained, and the third model parameter information and / or the fourth model parameter information is uploaded to the preset aggregation federated model for federated learning.

6. A text model recognition device based on federated learning, It is characterized in that include: A first acquisition module, used for acquiring the text to be predicted; A second acquisition module, configured to acquire, based on the text encoding model and the text to be predicted, second text semantic vector information of the text to be predicted output by the text encoding model; A third acquisition module, configured to acquire label information of the second text semantic vector information output by the text recognition model based on the text recognition model and the second text semantic vector information; A determination module is used to determine whether the text to be predicted violates the rules based on the label information, wherein the text encoding model and the text recognition model are obtained by the training method of the text model based on federated learning as described in any one of claims 1-3.

7. A computer device, It is characterized in that The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the training method of the text model based on federated learning as described in any one of claims 1 to 3, and implements the steps of the recognition method of the text model based on federated learning as described in claim 4.

8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the training method of a text model based on federated learning as described in any one of claims 1 to 3 and the steps of the recognition method of a text model based on federated learning as described in claim 4 are implemented.

Citation Information

Patent Citations

  • Negative text pushing method, device and system and computer device

    CN110457585A

  • SQL statement generation method and device, computer equipment and storage medium

    CN111581229A

  • Sensitive information identification method and device

    CN111966875A