Method and device for multi-party joint fine tuning of language model based on data protection

By obfuscating the parameter matrix and embedded tables of the large language model, and using column permutation and other methods, the problem of model parameters and data privacy leakage is solved, and multi-party joint fine-tuning is realized while protecting model parameters and data privacy.

CN120471139AActive Publication Date: 2025-08-12ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Patent Information

Application Number
CN202510394732.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-12
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the process of transmission, storage and application of large language models, model parameters and data privacy leakage is difficult to effectively protect model parameters and data privacy, especially when the model party deploys the model to the computing power party and the data party sends the data to the computing power party.

Method used

By determining the multi-layer obfuscation layer and obfuscation embedding table by the first party, obfuscation methods such as column permutation are used to obfuscate the parameter matrix and embedding table of the target language model to ensure that the model parameters are not leaked at the second party, and adjust the multi-layer obfuscation layer through the differences in obfuscation word embedding sequence and labels to achieve joint fine-tuning of the model.

Benefits of technology

On the premise of protecting model parameters and data privacy, a multi-party joint fine-tuning language model is implemented to avoid the leakage of model parameters and data on the computing power side, and ensure the security and privacy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471139A_ABST
    Figure CN120471139A_ABST
Patent Text Reader

Abstract

A multi-party joint fine-tuning language model method and device based on data protection, multiple parties comprise a first party holding a target language model and a second party providing computing power, the target language model comprises an embedding table and a plurality of serially arranged network layers, and the first party determines a plurality of confusion layers and a confusion embedding table; at least sending the multi-layer confusion layer to a second party; the confusion embedding table is obtained by performing second confusion on the embedding table, the confusion layer is obtained by performing first confusion on a parameter matrix in the corresponding network layer, and the first confusion mode corresponds to the second confusion mode; the second party obtains a training sample comprising a target confusion word embedding sequence and a label thereof, wherein the training sample is determined based on the text and the confusion embedding table; based on the target confusion word embedding sequence, obtaining a first output result through multiple confusion layers; and based on the difference between the first output result and the tag, adjusting a multi-layer confusion layer to realize a multi-party joint fine tuning model on the premise of protecting model parameters and data privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the field of artificial intelligence technology, and in particular, relate to a method and device for multi-party joint fine-tuning of a language model based on data protection. Background Art

[0002] Large language models have achieved tremendous success in recent years and are widely used in various fields. During the application of large language models, the model owner (hereinafter referred to as the model party), the user of the large language model, or the data owner (hereinafter referred to as the data party), and the computing power provider (hereinafter referred to as the computing power party) are often different. Considering that the model parameters of large language models are often obtained by the model party through a considerable investment of time, data, and computing power in training, the model party accordingly hopes that the model parameters of its large language model will be protected from theft during the transmission, storage, and application of the large language model. On the other hand, the data party often has data in specific fields and hopes that their own data will not be exposed when fine-tuning the large language model and using it.

[0003] In some example scenarios, assume there is a three-party scenario. One party (the model party) owns a large language model and can provide large language model inference services; one party (the computing power party) has strong computing power and can provide large language model fine-tuning and deployment services; one party (the data party) owns the data and hopes to fine-tune the large language model to meet its own inference needs.

[0004] For the above scenario, the three parties can currently collaborate as follows: the model party hands over its model to the computing power party, the data party sends its data to the computing power party, and the computing power party uses the data to fine-tune the model and deploy the model.

[0005] The above scenario poses a serious risk of leaking model parameters and data privacy. First, the data provider's data is private, and they don't want their data to be directly exposed to others (such as the computing power provider) outside the domain. The model provider, on the other hand, wants to protect the model parameters of their large language model to prevent them from being leaked. However, in the above scenario, deploying the large language model directly on the computing power provider will directly leak the model parameters, and directly sending the data provider's data to the computing power provider will also leak their data privacy. Therefore, a solution that can protect model parameters and data privacy in the above scenario is urgently needed to promote the further implementation of large language models. Summary of the Invention

[0006] The purpose of the present invention is to provide a method and device for multi-party joint fine-tuning of a language model based on data protection, so as to realize joint fine-tuning of the language model while protecting the model parameters and data privacy of the model.

[0007] In a first aspect, the present disclosure provides a method for multi-party joint fine-tuning of a language model based on data protection, wherein the multi-party includes a first party and a second party, the first party holds a target language model, the target language model includes an embedding table and a serially arranged multi-layer network layer, and the second party is used to provide computing power, including:

[0008] The first party determines multiple obfuscation layers and obfuscation embedding tables; and sends at least the multiple obfuscation layers to the second party; wherein a single obfuscation layer is obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer, and the obfuscation embedding table is obtained by performing a second obfuscation on the embedding table, the first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method, and the second obfuscation includes column permutation;

[0009] The second party obtains a training sample, wherein the training sample includes a target confused word embedding sequence and a corresponding label; based on the target confused word embedding sequence, a first output result is obtained through the multi-layer confusion layer; based on the difference between the first output result and the label, the multi-layer confusion layer is adjusted, wherein the target confused word embedding sequence and the label are determined based on the text and the confusion embedding table.

[0010] A second aspect of this specification provides a method for multi-party joint fine-tuning of a language model based on data protection, wherein the multiple parties include a first party and a second party, the first party holds a target language model, the target language model includes an embedding table and a serially arranged multi-layer network layer, the second party is used to provide computing power, and the method is implemented by the second party, including:

[0011] Obtaining multiple obfuscation layers from the first party, wherein the multiple obfuscation layers are obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer by the first party, and the first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method performed by the embedding table;

[0012] Obtaining a training sample, the training sample including a target obfuscated word embedding sequence and a corresponding label, the target obfuscated word embedding sequence and the label being determined based on text and an obfuscated embedding table, the obfuscated embedding table being obtained by performing the second obfuscation on the embedding table, the second obfuscation including column permutation;

[0013] Based on the target confused word embedding sequence, obtain a first output result through the multi-layer confusion layer;

[0014] The multiple confusion layers are adjusted based on a difference between the first output result and the label.

[0015] A third aspect of this specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the method described in the second aspect.

[0016] A fourth aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the second aspect is implemented.

[0017] A fifth aspect of this specification provides a computer program product, comprising a computer program / instruction, which implements the steps of the method described in the second aspect when executed by a processor.

[0018] According to the multi-party joint fine-tuning language model method and device based on data protection provided in the embodiments of this specification, the first party determines multiple confusion layers and confusion embedding tables to achieve parameter confusion of the target language model to avoid its network layer parameters and embedding table being exposed to other parties. The first confusion mode of the confusion matrix in each confusion layer corresponds to the second confusion mode of the confusion embedding table, ensuring that the two can still perform correct calculations after being confused; then, the first party sends at least multiple confusion layers to the second party, so that the model parameters can be avoided from being exposed at the second party under the premise of using the computing power service provided by the second party; the second party obtains a training sample including a target confusion word embedding sequence and its corresponding label, and the target confusion word embedding sequence and label are based on the text. This and the confusion embedding table can avoid the exposure of the plaintext of the text to the second party; the second party obtains a first output result through multiple confusion layers based on the target confusion word embedding sequence; since the first confusion method of the confusion matrix in each confusion layer corresponds to the second confusion method including column permutation of the confusion embedding table, the mutual coordination of the confusion methods of the parameter matrices of each network layer can ensure that the confusion situation of the output of the confused model matches the confusion situation of the label determined based on the confusion embedding table, and then based on the difference between the first output result and the label of the confusion situation match, the multi-layer confusion layers of the target language model can be adjusted to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 It is a schematic diagram of an implementation framework of an embodiment disclosed in an embodiment;

[0021] Figure 2A This is a schematic diagram of an implementation framework of another embodiment disclosed in an embodiment;

[0022] Figure 2B This is a schematic diagram of an implementation framework of another embodiment disclosed in an embodiment;

[0023] Figure 3 This is a flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in one embodiment of this specification;

[0024] Figure 4 This is a schematic diagram of the structure of a network layer of the target language model in the embodiment of this specification;

[0025] Figure 5 This is another flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in an embodiment of this specification;

[0026] Figure 6 This is another flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in an embodiment of this specification;

[0027] Figure 7 This is another flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in an embodiment of this specification;

[0028] Figure 8 This is a schematic block diagram of an apparatus for multi-party joint fine-tuning of a language model based on data protection in one embodiment of this specification. DETAILED DESCRIPTION

[0029] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. All other embodiments derived by those skilled in the art based on the embodiments in this specification without creative effort shall fall within the scope of protection of this specification.

[0030] The embodiments of this specification disclose a method and device for multi-party joint fine-tuning of a language model based on data protection. The following first introduces the application scenarios and technical concepts of the method as follows:

[0031] As mentioned above, in a scenario where one party (the model party) owns a large language model and can provide large language model inference services, one party (the computing power party) has strong computing power and can provide large language model fine-tuning and deployment services, and one party (the data party) owns data and hopes to fine-tune the model party's large language model to meet its own inference needs, if the model party deploys the large language model to the computing power party, and the data party sends the data it holds in plain text to the computing power party so that the large language model deployed by the computing power party can be fine-tuned through the data, there will be a problem that the model party's large language model and the data party's data are both leaked on the computing power party.

[0032] Alternatively, in a scenario where one party (called the model data party) has a large language model and data for fine-tuning the large language model, but its own computing power is insufficient and it needs to rely on the powerful computing power provided by the other party (the computing power party) to fine-tune and deploy the model, the model data party sends its large language model and data to the computing power party to use the computing power provided by the computing power party to fine-tune the large language model deployed by the computing power party through the data. However, there will also be the problem of the parameters and data of the large language model being leaked on the computing power party.

[0033] In view of this, there is an urgent need for a solution that can protect model parameters and data privacy in the above scenarios to promote the further implementation and application of large language models.

[0034] In the embodiments of this specification, the inventors propose a method for multi-party joint fine-tuning of language models based on data protection. Figure 1 A schematic diagram illustrating an implementation scenario according to an embodiment disclosed herein is provided. In this implementation scenario, the multiple parties include a first party A and a second party B. First party A holds a target language model LM, which includes an embedding table and a serially arranged multi-layer network layer. Second party B provides computing power to fine-tune the target language model LM and provide inference services based on the fine-tuned target language model LM. The second party can be referred to as the computing power party.

[0035] The embedding table includes at least a word embedding subtable, which includes the word embedding corresponding to each word (or token), namely word_embedding, which can exist in the form of a vector. In some cases, the embedding table may also include a position embedding subtable, which includes the position embedding corresponding to each position, namely position_embedding, which can also exist in the form of a vector and is used to represent the position of each word in its text; the embedding table may also include a token type embedding subtable, which may include the token type embedding corresponding to each token type, namely token_type_embedding, which can also exist in the form of a vector, and each token type is used to distinguish different types or parts in the input (for example, tokens belonging to different sentences, such as marking the beginning and end of a sentence, etc.).

[0036] The process of multi-party joint fine-tuning of language models can include two stages: model preprocessing stage and actual model fine-tuning stage;

[0037] In the model preprocessing phase: Party A determines multiple layers of confusion layers and confusion embedding tables based on its target language model LM. The i-th confusion layer is obtained by performing a first confusion on the parameter matrix in the i-th network layer, and the confusion embedding table is obtained by performing a second confusion on the embedding table. The second confusion includes column permutation. The first and second confusion modes of the parameter matrices in each network layer correspond to each other. The confusion modes of the parameter matrices of each network layer cooperate with each other so that the confusion of the output of the obfuscated model matches the confusion of the labels determined based on the confusion embedding table. Where i is a positive integer, and its maximum value is the total number of layers in the multi-layer network.

[0038] It can be understood that in the embodiments of this specification, the “first” in the first obfuscation and the “second” in the second obfuscation are only used to distinguish the obfuscation operations performed on different objects and do not have any other limiting meanings.

[0039] Afterwards, the first party A sends at least multiple obfuscation layers to the second party B. In response, the second party B obtains the multiple obfuscation layers and deploys them. Then, the actual model fine-tuning phase begins.

[0040] Specifically, during the model fine-tuning stage, the second party B obtains a training sample, which includes a target confused word embedding sequence and its corresponding label; the target confused word embedding sequence and label are determined based on the text and the confused embedding table. Accordingly, the second party B can only obtain the obfuscated content of the text used to fine-tune the target language model LM, and cannot obtain the plaintext of the text, thereby achieving privacy protection of the text, i.e., data, used to fine-tune the target language model LM.

[0041] Then, the second party B obtains the first output result based on the target confused word embedding sequence through multiple confusion layers; since the first confusion method of the parameter matrix in each network layer corresponds to the second confusion method, it can ensure that the output result of the multi-layer confusion layer matches the confusion situation of the label determined based on the confusion embedding table; accordingly, the second party B adjusts the multi-layer confusion layer based on the difference between the first output result and the label to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy.

[0042] Understandably, Figure 1As shown in FIG, in some possible examples, the text used to fine-tune the target language model LM may be held by the first party. In this case, the first party A can directly determine the training sample based on its confusion embedding table and the text, and then send it to the second party B so that the second party B obtains the training sample. The aforementioned second party B adjusts the multi-layer confusion layer based on the difference between the first output result and the label, as shown in FIG. Figure 1 As shown, the second party B sends the first output result to the first party A, and the first party A determines the confused word embedding corresponding to the generated word corresponding to each corresponding position in the first output result based on a series of probabilities corresponding to each corresponding position in the first output result and the confusion embedding table, thereby obtaining a predicted confused word embedding sequence; then the predicted confused word embedding sequence is sent to the second party B. The second party B adjusts the multiple confusion layers based on the difference between the predicted confused word embedding sequence and the label determined based on the confusion embedding table.

[0043] In the above process, the first party determines multiple layers of confusion layers and confusion embedding tables to achieve parameter confusion of the target language model to avoid exposing its network layer parameters and embedding tables to other parties. The first confusion mode of the confusion matrix in each confusion layer corresponds to the second confusion mode of the confusion embedding table, ensuring that the two can still perform correct calculations after being confused. Then, the first party sends at least multiple layers of confusion layers to the second party, so that the model parameters can be avoided from being exposed at the second party under the premise of using the computing power service provided by the second party. The second party obtains training samples including the target confused word embedding sequence and its corresponding label, and the target confused word embedding sequence and label are determined based on the text and the confusion embedding table, which can be used. Avoid exposing the plaintext of the text to the second party; the second party obtains a first output result through multiple confusion layers based on the target obfuscated word embedding sequence; since the first obfuscation method of the confusion matrix in each confusion layer corresponds to the second obfuscation method including column permutation of the obfuscation embedding table, the mutual coordination of the obfuscation methods of the parameter matrices of each network layer can ensure that the obfuscation situation of the output of the obfuscated model matches the obfuscation situation of the label determined based on the obfuscation embedding table, and then based on the difference between the first output result and the label of the confusion situation match, the multi-layer obfuscation layers of the target language model can be adjusted to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy.

[0044] like Figure 2AAs shown, it is a schematic diagram of another implementation scenario according to an embodiment disclosed in this specification. In this implementation scenario, the multiple parties that jointly fine-tune the language model include a first party A, a second party B, and a third party C, wherein the first party A holds the target language model, which includes an embedding table and a multi-layer network layer arranged in series, and the embedding table includes at least a word embedding sub-table, the second party B is used to provide computing power, and the third party C holds the text required for fine-tuning the target language model, as well as the text required for subsequent reasoning. In this implementation scenario, the first party A can also be called the model party, the second party B can be called the computing power party, and the third party C can be called the data party.

[0045] The process of multi-party joint fine-tuning of language models can include two stages: model preprocessing stage and actual model fine-tuning stage;

[0046] In the model preprocessing stage: determine the multiple layers of confusion layers and the confusion embedding table; further, determine the first mapping relationship and the second mapping relationship corresponding to the confusion word embedding sub-table in the confusion embedding table, wherein the confusion word embedding sub-table is obtained by performing row permutation on the word embedding sub-table after column permutation; the first mapping relationship is the mapping relationship between each word in the confusion word embedding sub-table and the position identifier of its position in the confusion word embedding sub-table, which can be expressed as word->id; the second mapping relationship is the mapping relationship between the confusion word embedding of each word in the confusion word embedding sub-table and the position identifier of each word, which can be expressed as id->embedding.

[0047] It can be understood that by performing column permutation on the word embeddings corresponding to each word in the word embedding subtable in the embedding table to obtain the confused word embeddings corresponding to each word, the word embeddings of each word can be confused. Then, at least the word embedding subtable after the column permutation is subjected to row permutation, that is, the words in the word embedding subtable and their confused word embeddings are permuted again to achieve position confusion of each word and its confused word embedding in the word embedding subtable, so as to better avoid the exposure of the embedded table text and better achieve the confidentiality of the text of the third party C.

[0048] It can be understood that row permutation does not affect the correspondence between a word and its confused word. A word and its confused word are always in the same row in the table. For example, before row permutation, word C and its confused word embedding are in row 4 of the word embedding subtable (i.e., in the embedding table). Correspondingly, both of them correspond to position ID 4. After row permutation, word C and its confused word embedding are permuted to row 13 of the word embedding subtable (i.e., in the embedding table). Correspondingly, both of them correspond to position ID 13.

[0049] For the word embedding subtable before and after row permutation, its content, i.e., the word and its confused word embedding, remains unchanged. Only the position of the word and its confused word embedding in the table changes. This change in position does not affect the accuracy of the determined confused word embedding corresponding to the word.

[0050] Next, party A sends the multi-layered obfuscation layer and the aforementioned second mapping relationship id->embedding to party B. In response, party B obtains the multi-layered obfuscation layer, deploys it, obtains the second mapping relationship id->embedding, and stores it. Party A sends the first mapping relationship word->id to party C, which obtains and stores it. This is the actual model fine-tuning phase.

[0051] Specifically, in the model fine-tuning stage: the third party C determines the position identifier id corresponding to each word in the text based on the first mapping relationship word->id it obtained to obtain the target position identifier sequence; then, the third party C sends the target position identifier sequence to the second party B.

[0052] After the second party B obtains the target position identification sequence sent by the third party C, it can determine the training sample, that is, the target confusion word embedding sequence and its corresponding label based on its second mapping relationship id->embedding and the target position identification sequence.

[0053] Then, the second party B obtains a first output result based on the target confused word embedding sequence through multiple confusion layers; since the first confusion method of the parameter matrix in each network layer corresponds to the second confusion method, it can be ensured that the output result of the multi-layer confusion layer matches the confusion situation of the label determined based on the confusion embedding table; accordingly, based on the difference between the first output result and the label, the multi-layer confusion layer is adjusted to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy.

[0054] Illustratively, the aforementioned process of adjusting the multiple obfuscation layers based on the difference between the first output result and the label may include: determining, based on a series of probabilities corresponding to each corresponding position in the first output result and a second mapping relationship, a position identifier corresponding to a generated word corresponding to each corresponding position in the first output result, thereby obtaining a predicted position identifier sequence; and then, adjusting the multiple obfuscation layers based on the difference between the predicted position identifier sequence and the label (which is a series of position identifiers) determined based on at least the first mapping relationship.

[0055] In the above process, the first party determines the first mapping relationship and the second mapping relationship corresponding to the multi-layer confusion layer and the confusion embedding table and the confusion word embedding sub-table; then sends the first mapping relationship word->id to the third party, which can not only hide the original position of each word in the embedding table, but also prevent the third party from knowing the word embedding corresponding to each word, thereby achieving the confidentiality of the parameters of the target language model at the third party; the first party sends the multi-layer confusion layer and the second mapping relationship id->embedding to the second party, so that the second party cannot know the true embedding of each word and the original position of each word in the embedding table, and cannot know the correspondence between each embedding and the word, thereby achieving the confidentiality of the parameters of the target language model at the second party, and the second party only obtains the second mapping relationship id->embedding, which cannot deduce the plaintext of the third party's text, thereby avoiding the leakage of the third party's private data, i.e., text, at the second party. Moreover, the first obfuscation mode of the parameter matrix in each network layer corresponds to the second obfuscation mode of the obfuscation embedding table. The mutual coordination of the obfuscation modes of the parameter matrices of each network layer can ensure that the obfuscation status of the output of the obfuscated model matches the obfuscation status of the label determined based on the obfuscation embedding table. Based on the difference between the first output result and the label of the confusion status match, the multi-layer obfuscation layer of the target language model can be adjusted to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy.

[0056] like Figure 2B The figure shows another implementation scenario diagram according to an embodiment disclosed in this specification. In the case where the multiple parties jointly fine-tuning the language model include the aforementioned first party A, second party B, and third party C, after first party A determines the multi-layer confusion layer and the confusion embedding table, it can also send the multi-layer confusion layer to second party B and send the confusion embedding table to third party C; then third party C determines to obtain a training sample based on the confusion embedding table and the text it holds, that is, the target confusion word embedding sequence and its corresponding label; and sends the training sample to second party B; then, second party B obtains a first output result based on the target confusion word embedding sequence through the multi-layer confusion layer; second party B adjusts the multi-layer confusion layer based on the difference between the first output result and the label, so as to achieve multi-party joint fine-tuning of the language model while protecting model parameters and data privacy.

[0057] The aforementioned second party B adjusts the process of multiple confusion layers based on the difference between the first output result and the label, such as Figure 2BAs shown, the second party B sends the first output result to the third party C, and the third party C determines the confused word embedding corresponding to the generated word corresponding to each corresponding position in the first output result based on a series of probabilities and the confusion embedding table corresponding to each corresponding position in the first output result, thereby obtaining a predicted confused word embedding sequence; then the predicted confused word embedding sequence is sent to the second party B. The second party B adjusts the multiple confusion layers based on the difference between the predicted confused word embedding sequence and the aforementioned label.

[0058] The method for multi-party joint fine-tuning of the language model based on data protection provided in this specification is described in detail below with reference to specific embodiments.

[0059] Figure 3 A flowchart of a method for jointly fine-tuning a language model based on data protection by multiple parties in one embodiment of this specification is shown. The aforementioned multiple parties may include a first party A and a second party B. The first party A holds a target language model, which includes an embedding table and a serially arranged multi-layer network layer. The second party B is used to provide computing power. In some possible examples, when the first party A and the second party B jointly fine-tune the target language model, the data for fine-tuning the target language model, namely the text mentioned later, may be held by the first party A.

[0060] In some possible examples, the target language model is a network model based on a Transformer structure. Multiple network layers can be referred to as Transformer layers. The network structures in the multiple network layers are similar. For example, the target language model can be a BERT (Bidirectional Encoder Representations from Transformers) model or an improved model thereof, or a GPT (Generative Pre-trained Transformer) model or an improved model thereof.

[0061] like Figure 3 As shown, the method for multi-party joint fine-tuning of the language model based on data protection includes the following steps S310-S360:

[0062] Considering that during the process of jointly fine-tuning the target language model, first party A and second party B need to protect the model parameters of the target language model (including the parameters in the embedding table and each network layer) to prevent the model parameters from being leaked at second party B, they also need to protect the data used to fine-tune the target language model, that is, the text held by them for fine-tuning the target language model, to prevent its leakage at second party B. Accordingly, before jointly fine-tuning the target language model, first party A and second party B need to perform model preprocessing on the target language model to prevent the model parameters of the target language model from being leaked at second party B, and to prevent the textual version of the fine-tuned target language model from being leaked at second party B.

[0063] Specifically, in step S310, the first party A determines multiple obfuscation layers and obfuscation embedding tables. A single obfuscation layer is obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer. The obfuscation embedding table is obtained by performing a second obfuscation on the embedding table. The second obfuscation includes column permutation. The first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method. That is, the first obfuscation method of the confusion matrix in each obfuscation layer corresponds to the second obfuscation method.

[0064] Among them, the embedding table is a table used to vectorize the input data of the model (such as text used to fine-tune the target language model, and other texts that need to be inferred using the fine-tuned target language model). A row in the embedding table can represent a word embedding corresponding to a word (i.e., a segmentation word), or can represent a position embedding corresponding to a position, or can represent a tag type embedding corresponding to a tag type, wherein each type of embedding can exist in the form of a vector. Accordingly, the embedding table can be said to include at least a word embedding subtable, and the word embedding subtable includes words (i.e., segmentations) and their corresponding word embeddings. The embedding table can also include a position embedding subtable and a tag type embedding subtable, and the position embedding subtable includes each position and its corresponding position embedding, and the tag type embedding subtable includes each tag type and its corresponding tag type embedding.

[0065] The obfuscated embedding table is obtained by performing a second obfuscation on the embedding table. The second obfuscation includes at least column permutation, which can be understood as: first party A performs at least column permutation on the embedding table to obtain the obfuscated embedding table. The column permutation on the embedding table may include: performing column permutation on each embedding in the embedding table.

[0066] The parameter matrix in each network layer may include one or more, and a single parameter matrix may include at least a weight matrix and may also include a bias matrix corresponding to the weight matrix.

[0067] It can be understood that after the first party A confuses the model parameters of the target language model (i.e., the embedding table and the parameter matrix of the multi-layer network layer) to obtain the confusion model, it is also necessary to ensure that the confusion of the output result obtained for the input using the confusion model matches the confusion of the label corresponding to the input determined based on the confusion embedding table. Accordingly, when confusing the target language model, the first confusion method of the parameter matrix in each network layer corresponds to the second confusion method of the embedding table.

[0068] In some possible examples, the first party A may obtain the obfuscation embedding table and the multiple obfuscation layers from its designated storage location. The obfuscation embedding table and the multiple obfuscation layers are pre-generated by the first party A for the target language model.

[0069] In some further possible examples, step S310 may include steps 11-13: in step 11, the first party A generates a first permutation matrix corresponding to the embedding table and second permutation matrices corresponding to the parameter matrices in the network layers.

[0070] In this step, the first party A can randomly generate a corresponding permutation matrix based on the size (or shape) of the embedding table, which can be called the first permutation matrix π e ; and based on the size of each parameter matrix in each network layer, randomly generate the permutation matrix corresponding to each parameter matrix, which can be called the second permutation matrix π corresponding to each parameter matrix name , where "name" is associated with the parameter matrix corresponding to the second permutation matrix.

[0071] It is understandable that a permutation matrix (including the first and second permutation matrices, as well as the third permutation matrix mentioned later) can also be called a permutation matrix or a shuffle matrix. It is a special type of square matrix in which each row and column contains exactly one element with a value of 1, and the rest of the elements have a value of 0. A permutation matrix can reorder the rows or columns of other matrices. A permutation matrix is an orthogonal matrix, and its inverse matrix is its transposed matrix.

[0072] In some examples, the first permutation matrix and each second permutation matrix may be used multiple times to confuse the target language model (including its embedding table and parameter matrix), or a new first permutation matrix and each second permutation matrix may be generated each time there is a need for confusion.

[0073] The second permutation matrices corresponding to the parameter matrices of different network layers may be the same or different; the second permutation matrices corresponding to different parameter matrices may be the same or different.

[0074] Wherein, each parameter matrix may include a weight matrix and may also include a bias matrix corresponding to the weight matrix. The aforementioned randomly generating a permutation matrix corresponding to each parameter matrix based on the size of each parameter matrix in each network layer may refer to: randomly generating a permutation matrix corresponding to each parameter matrix based on the size of the weight matrix included in each parameter matrix in each network layer. It can be understood that the size of the bias matrix corresponding to the weight matrix is the same as the size of the weight matrix.

[0075] The randomly generating a corresponding permutation matrix based on the size (or shape) of the embedding table may refer to randomly generating a corresponding permutation matrix based on the sizes (or shapes) of multiple embeddings in the embedding table.

[0076] Next, in step 12, the first party A right-multiplies the embedding table by at least the first permutation matrix to obtain the obfuscated embedding table.

[0077] In some possible examples, the aforementioned second obfuscation may include at least column permutation. Accordingly, in step 12, the first party may right-multiply the embedding table by the first permutation matrix to perform column permutation on the embedding table, that is, perform column permutation on each embedding in the embedding table to obtain a confused embedding table. In the case where the embedding table includes word embedding e1, position embedding e2, and token type embedding e3, right-multiplying the embedding table by the first permutation matrix may include: right-multiplying the word embedding e1, position embedding e2, and token type embedding e3 by the first permutation matrix π e .

[0078] In some other possible examples, the aforementioned second confusion may include column permutation and row permutation. Accordingly, in step 12, it may specifically include: the first party A multiplies the embedding table by the first permutation matrix to obtain the embedding table after column permutation; then, the word embedding subtable after column permutation in the embedding table after column permutation is multiplied by the third permutation matrix to obtain the confused word embedding subtable, and then the confused embedding table is obtained. Continuing with the above example, in this step, the first permutation matrix π is multiplied to the right e The word embedding e1, continue to left multiply the third permutation matrix π w , to obtain the confusing word embedding subtable, and realize row permutation of each word and its confusing word embedding.

[0079] In step 13, the first party A processes each parameter matrix in each network layer using at least a second permutation matrix corresponding to the parameter matrix to obtain a confusion matrix corresponding to the parameter matrix, thereby obtaining a multi-layer confusion layer.

[0080] In which, a single parameter matrix may include a weight matrix and may also include a bias matrix corresponding to the weight matrix. When the parameter matrix includes the weight matrix, in step 13, the first party A processes the weight matrix in the parameter matrix for each parameter matrix in each network layer using at least the second permutation matrix corresponding to the parameter matrix to obtain a confusion matrix corresponding to the parameter matrix, wherein the confusion matrix corresponding to the parameter matrix includes: the weight matrix after confusion, which can also be called the weight matrix after column permutation) to obtain a multi-layer confusion layer.

[0081] In some possible examples, in order to ensure that the target language model after obfuscation (i.e., multiple obfuscation layers) can correctly calculate its input, and to ensure that the output of the obfuscated target language model matches the obfuscation of the label determined based on the obfuscation embedding table, the first obfuscation mode of the parameter matrix in each network layer corresponds to the second obfuscation mode.

[0082] On the basis of the above, we also consider that the row permutation performed by the confusing word embedding subtable will only confuse the words and their positions in the table, and will not affect the subsequent vectorization of each word in the text. That is, the vectorization of the text using the confusing word embedding subtable is consistent with the vectorization of the text using the word embedding subtable without row permutation (only column permutation), and has no effect on the subsequent intermediate processing process of the target language model.

[0083] Accordingly, the embodiments of this specification also provide the following replacement idea to ensure that the confusion model (i.e., the obfuscated target language model) correctly processes the input and ensures that its output matches the confusion of the label determined based on the obfuscation embedding table.

[0084] Specifically, for the weight matrix in each parameter matrix, if the weight matrix needs to perform matrix multiplication on its input, then its input (which has also undergone column permutation) can be first restored by column permutation, and then the second permutation matrix corresponding to the weight matrix (i.e., the parameter matrix to which it belongs) can be used to perform column permutation on the output of the weight matrix (i.e., the result of matrix multiplication of the input after column permutation restoration and the weight matrix); if the weight matrix needs to perform element-wise multiplication on its input (which has also undergone column permutation), it is necessary to ensure that the column permutation of the weight matrix is consistent with the column permutation of its input. This ensures that during the inference process, the input to the j-th parameter matrix (i.e., the output of the j-1-th parameter matrix) is only permuted (i.e., obfuscated) based on the second permutation matrix corresponding to the j-1-th parameter matrix, relative to the input plaintext that has not undergone column permutation.

[0085] In view of the above-mentioned permutation idea, in some possible examples, a single parameter matrix includes at least a weight matrix. Exemplarily, the weight matrix can be a first weight matrix that needs to be element-wise multiplied with its input; accordingly, in step 13, the process of obtaining the confusion matrix corresponding to the parameter matrix can include: right-multiplying the first weight matrix by its corresponding second permutation matrix to obtain the confusion matrix corresponding to the first weight matrix, wherein the second permutation matrix corresponding to the first weight matrix makes the confusion of the confusion matrix corresponding to the first weight matrix consistent with the confusion of its input. Specifically, the second permutation matrix corresponding to the first weight matrix can be the same as the second permutation matrix corresponding to the previous weight matrix of the first weight matrix.

[0086] In some possible examples, a single parameter matrix includes at least a weight matrix. Exemplarily, the weight matrix can be a second weight matrix that needs to be matrix multiplied with its input; accordingly, in step 13, the process of obtaining the confusion matrix corresponding to the parameter matrix can include: multiplying the second weight matrix on the left by the transposed matrix of the second permutation matrix corresponding to the previous weight matrix, and multiplying it on the right by the corresponding second permutation matrix to obtain the confusion matrix corresponding to the second weight matrix.

[0087] In some possible examples, the target language model is a network model based on a Transformer structure, and each network layer may include a network sublayer based on a self-attention mechanism and a feedforward neural network sublayer, wherein the network sublayer based on the self-attention mechanism and the feedforward neural network sublayer both include layer normalization; wherein the aforementioned first weight matrix may include the weight matrix in layer normalization. The aforementioned second weight matrix may include other weight matrices in the network sublayer based on the self-attention mechanism except the weight matrix in layer normalization, and other weight matrices in the feedforward neural network sublayer except the weight matrix in layer normalization.

[0088] In some examples, the last network layer of the target language model may further include a dimension mapping sublayer, namely, a LM_head (Language Model Head) layer, relative to other network layers. This dimension mapping sublayer may be provided after the feedforward neural network sublayer of the current network layer. This dimension mapping sublayer is used to map its input dimensions to the word dimensions of the confusion word embedding subtable.

[0089] In some possible examples, a single parameter matrix may also include a bias matrix corresponding to the weight matrix. Accordingly, in step 13, obtaining the confusion matrix corresponding to the parameter matrix may also include: right-multiplying the bias matrix by its corresponding second permutation matrix to obtain a confused bias matrix, that is, a bias matrix after column permutation, and the second permutation matrix corresponding to the bias matrix is the same as the second permutation matrix corresponding to the weight matrix.

[0090] It can be understood that during the processing, the bias matrix needs to be added to the product of its corresponding weight matrix and the input. For the bias matrix, it is sufficient to ensure that the column permutation of the bias matrix after confusion is consistent with the column permutation of the product of its corresponding weight matrix and the input.

[0091] For example, for a second weight matrix W a1 (a weight matrix that needs to be matrix multiplied with its input) and its bias matrix, the second weight matrix W a1 The input is based on the second weight matrix W a1 The previous weight matrix W a2 The corresponding second permutation matrix The second weight matrix W is permuted a1 The input can be expressed as @ represents matrix multiplication, H represents the second weight matrix W a1 The input plaintext.

[0092] The second weight matrix W a1 Multiply the previous weight matrix W on the left a2 The corresponding second permutation matrix The transposed matrix of After that, right multiply the second weight matrix W a1 The corresponding second permutation matrix It can be achieved by first inputting the second weight matrix Perform column permutation recovery, then perform column permutation recovery on the input H and the second weight matrix W a1 The multiplication result 1 of the two is then permuted. Specifically, the process can be expressed as in is the identity matrix, which can be simplified as

[0093] At this time, the column permutation of the multiplication result 1 is based on the second permutation matrix corresponding to the second weight matrix Wa1 Accordingly, the second weight matrix W can be a1 The corresponding bias matrix is right-multiplied by its corresponding second permutation matrix This ensures that the aforementioned multiplication result is consistent with the column permutation of the bias matrix after column permutation.

[0094] Similarly, for a first weight matrix W a3 (a weight matrix that needs to be element-wise multiplied with its input) and its bias matrix, the first weight matrix W a3 The input is based on the first weight matrix W a3 The previous weight matrix Wa4 The corresponding second permutation matrix After column permutation, the first weight matrix W a3 The input can be expressed as F represents the first weight matrix W a3 The input plaintext.

[0095] In order to ensure that the first weight matrix W a3 The situation after column permutation is consistent with the column permutation of the input. The first weight matrix W a3 The corresponding second permutation matrix With the previous weight matrix W a4 The corresponding second permutation matrix The same, then the first weight matrix W a3 Right multiply by its second permutation matrix Afterwards, its column permutations are consistent with those of its input.

[0096] Subsequently, the first weight matrix W a3 Input With the first weight matrix W a3 Right multiply by its second permutation matrix The result of element multiplication is 2, which can be expressed as * represents element multiplication. The column permutation of the multiplication result 2 is the same as the first weight matrix W a3 Right multiply by its second permutation matrix (and the first weight matrix W a3 Input ) is consistent with the column permutation. Accordingly, the first weight matrix W a3 The corresponding bias matrix is right-multiplied by its corresponding second permutation matrix This ensures that the aforementioned multiplication result 2 is consistent with the column permutation of the bias matrix after column permutation.

[0097] In summary, the bias matrix can be directly multiplied by its corresponding second permutation matrix. This will not cause any error in the output of the obfuscated target language model compared to the output obtained using the target language model.

[0098] In some possible examples, the network sublayer based on the self-attention mechanism of each network layer can be a network sublayer based on the multi-head self-attention mechanism, and the network sublayer based on the self-attention mechanism of each network layer includes multiple groups of query parameter matrices Q W , key parameter matrix K W and the value parameter matrix V W .

[0099] At this point, it is necessary to calculate the query parameter matrix Q in each self-attention head W , key parameter matrix K W and the value parameter matrix V W Generate the corresponding second permutation matrix respectively. For each self-attention head, use the query parameter matrix Q W , key parameter matrix K W and the value parameter matrix V W The corresponding second permutation matrix is respectively the query parameter matrix Q W , key parameter matrix K W and the value parameter matrix V W Permutations are performed to ensure that each permutation occurs only within each self-attention head and not between two self-attention heads, to ensure that subsequent processing remains correct after permutation of multiple self-attention heads and does not introduce errors into the model output results (confusion).

[0100] Due to the calculation method inside each self-attention head, the query parameter matrix Q in each attention head W and key parameter matrix K W Can correspond to the same second permutation matrix, which can be expressed as π QK For example, for multiple heads, i.e. multiple groups (e.g. n groups), the query parameter matrix Q W or multiple sets of key parameter matrices K W , generate the second permutation matrix π QK , the second permutation matrix π QK It can be expressed as in, represents the query parameter matrix Q in the kth attention head W and key parameter matrix K W The corresponding second permutation matrix, each The dimension is equal to the dimension d corresponding to each attention head. diag means that n second permutation matrices are arranged diagonally from the upper left to the lower right, and the remaining positions are filled with 0 to form an nd-dimensional square matrix. For example, n is equal to 3, the second permutation matrix π QK , which can be expressed as follows:

[0101]

[0102] For multiple heads, that is, multiple groups (for example, n groups) of value parameter matrix V W , generate its corresponding second permutation matrix π V The method can be referred to generate the second permutation matrix π QK The method is not described here.

[0103] In some possible examples, each network layer may further include a residual connection operation, and the second permutation matrix corresponding to the previous parameter matrix corresponding to the residual connection operation may be the same as the first permutation matrix.

[0104] It is understandable that any network layer c in each network layer may include a network sublayer based on the self-attention mechanism and a feedforward neural network sublayer, and both the network sublayer based on the self-attention mechanism and the feedforward neural network sublayer may include a residual connection operation. For the residual connection operation in the network sublayer based on the self-attention mechanism, it is necessary to perform the previous operation (for example, Figure 4 The output of the first linear transformation operation (dense Linear) shown in is added to the input of the network sublayer based on the self-attention mechanism. If the output of the previous operation and the input of the network sublayer based on the self-attention mechanism are permuted, the column permutations of the two must be consistent to ensure the accuracy of the residual connection operation.

[0105] Considering that when network layer c is the first network layer, the input of its network sublayer based on the self-attention mechanism is the initial embedding sequence (that is, the target confused word embedding sequence obtained based on the confusion embedding table). When network layer c is not the first network layer, the input of its network sublayer based on the self-attention mechanism is the output of the previous network layer.

[0106] In view of this, if the target confused word embedding sequence (or the processing result of the target confused word embedding sequence after layer normalization) is permuted in columns, the column permutation is based on the first permutation matrix corresponding to the confused embedding table. In combination with the aforementioned permutation idea and the network structure of the network layer, if the output of each network layer is permuted in columns, the column permutation is also based on the first permutation matrix corresponding to the confused embedding table. Correspondingly, the column permutation of the output of the previous operation (for example, the first linear transformation operation) of the residual connection operation is also based on the first permutation matrix corresponding to the confused embedding table. Then, the second permutation matrix corresponding to the previous parameter matrix corresponding to the residual connection operation is the same as the first permutation matrix.

[0107] Similarly, for the residual connection operation in the feedforward neural network sublayer, it is necessary to add the input of the feedforward neural network sublayer and the output of its previous operation. If the input of the feedforward neural network sublayer and the output of its previous operation are both column-permuted, it is necessary to ensure that the column permutations of the two are consistent in order to ensure the accuracy of the result of the residual connection operation.

[0108] The input to the feedforward neural network sublayer can be the output of the residual connection operation in the self-attention mechanism-based network sublayer of network layer c. The column permutation of the input to the feedforward neural network sublayer is performed based on the aforementioned first permutation matrix. To ensure the accuracy of the results of the residual connection operation, the column permutation of the output of the previous residual connection operation in the feedforward neural network sublayer also needs to be performed based on the first permutation matrix. Accordingly, the second permutation matrix corresponding to the previous parameter matrix corresponding to the residual connection operation in the feedforward neural network sublayer is also the same as the first permutation matrix.

[0109] In some possible examples, each network layer includes layer normalization, which requires element-wise multiplication of its input. The permutation of its parameter matrix after permutation needs to be consistent with the permutation of its input. Accordingly, based on the network layer structure, the second permutation matrix corresponding to the layer normalization parameter matrix is the same as the first permutation matrix to ensure correct calculation.

[0110] It can be understood that the network structures of multiple layers in the target language model are similar. The following takes any one of the network layers as an example to illustrate the process of confusing the parameter matrices in each network layer. Figure 4 As shown, a network structure of a network layer in a target language model, for example, the i-th network layer, is shown.

[0111] like Figure 4 As shown, the i-th network layer may include a network sublayer based on the self-attention mechanism and a feedforward neural network sublayer, where Figure 4 As shown, the network sublayer based on the self-attention mechanism can include a first layer normalization (such as Figure 4 LayerNorm1 in QKV linear transformation operations (such as Figure 4 QKV Linear in ), attention value calculation operations (such as Figure 4 Attention shown), attention value and value multiplication operation (such as Figure 4 The A@V shown in the figure, @ represents matrix multiplication, A represents the output of the attention value calculation operation, V represents the value matrix, and V is the value parameter matrix W The input and value parameter matrix V W Multiplication), the first linear transformation operation (such as Figure 4 dense Linear in) and the first residual connection operation Skip Add (such as Figure 4 As shown in ).

[0112] The first layer normalization includes a parameter matrix, namely the weight matrix L w1 and the bias matrix L bias1The second permutation matrix corresponding to the first layer normalization can be expressed as The corresponding confusion matrix can be expressed as: confusion weight matrix Confusion bias matrix

[0113] The network sublayer based on the self-attention mechanism in the i-th network layer can be a network sublayer based on a multi-head (e.g., n-head) or single-head self-attention mechanism. When the network sublayer based on the self-attention mechanism in the i-th network layer is a network sublayer based on a multi-head self-attention mechanism, each self-attention head includes a QKV linear transformation operation (e.g., Figure 4 QKV Linear in ), attention value calculation operations (such as Figure 4 Attention shown), attention value and value multiplication operation (such as Figure 4 A@V shown in Figure 2), each self-attention head is combined with the first linear transformation operation (such as Figure 4 dense Linear) connection in .

[0114] Among them, the attention value calculation operation of the kth self-attention head is calculated as follows represents the key parameter matrix K of the kth self-attention head k The @ denotes matrix multiplication.

[0115] The QKV linear transformation operation of the kth self-attention head includes three parameter matrices, namely the query parameter matrix (including the query weight matrix Q w-k and its corresponding query bias matrix Q bias-k ), key parameter matrix (including the key weight matrix K w-k and its corresponding key bias matrix K qbias-k ), and the value parameter matrix (including the value weight matrix V w-k and its corresponding value bias matrix V bias-k ).

[0116] Among them, the query parameter matrix in the k-th self-attention head is and key parameter matrix The corresponding second permutation matrix is the same and can be expressed as The query parameter matrix in the kth self-attention head The corresponding confusion matrix can be expressed as:

[0117] Confusion weight matrix Confusion bias matrix π e Tis the transposed matrix of the first permutation matrix corresponding to the confusion embedding table.

[0118] Key parameter matrix in the kth self-attention head The corresponding confusion matrix can be expressed as:

[0119] Confusion weight matrix Confusion bias matrix

[0120] The value parameter matrix in the kth self-attention head The corresponding second permutation matrix can be expressed as Value parameter matrix The corresponding confusion matrix can be expressed as:

[0121] Confusion weight matrix Confusion bias matrix

[0122] The attention value calculation operation and the attention value multiplication operation do not include the parameter matrix that needs to be permuted.

[0123] The first linear transformation operation may include a parameter matrix, which may include a weight matrix D w and the bias matrix D bias The second permutation matrix corresponding to the first linear transformation operation is the same as the first permutation matrix and can be expressed as π e , the confusion matrix corresponding to the first linear transformation operation can be expressed as: confusion weight matrix π V T @D w @π e , confusion bias matrix π V T @D bias @π e , where π V T Represents the value parameter matrix V W The corresponding second permutation matrix π V The transposed matrix of .

[0124] The first residual connection operation itself does not include a parameter matrix that needs to be permuted, and its two inputs must ensure that the column permutations are consistent.

[0125] The feedforward neural network sublayer in the i-th network layer includes a second layer of normalization (such as Figure 4 LayerNorm2 shown in ), the second linear transformation operation (such as Figure 4 Linear1 shown in ), activation function (such as Figure 4 ReLU (Rectified Linear Unit) shown), the third linear transformation operation (such as Figure 4 Linear2 shown in ), the second residual connection operation Skip Add (as shown in Figure 4 As shown in ) and the third layer normalization (such as Figure 4 LayerNorm3 shown in ).

[0126] Among them, the second layer normalization includes a parameter matrix, which includes the weight matrix L w2 and the bias matrix L bias2 The second permutation matrix corresponding to the second layer normalization is the same as the first permutation matrix and can be expressed as π e ; The corresponding confusion matrix can be expressed as: confusion weight matrix L w2 @π e , confusion bias matrix L bias2 @π e .

[0127] The second linear transformation operation includes a parameter matrix, which includes the weight matrix L w3 and the bias matrix L bias3 The second permutation matrix corresponding to the second linear transformation operation can be expressed as The corresponding confusion matrix can be expressed as: confusion weight matrix Confusion bias matrix

[0128] The activation function and the second residual connection operation themselves do not include a parameter matrix that needs to be permuted. The two inputs of the second residual connection operation need to keep the column permutation consistent.

[0129] The third linear transformation operation includes a parameter matrix, which includes the weight matrix L w4 and the bias matrix L bias4 The second permutation matrix corresponding to the third linear transformation operation can be expressed as The corresponding confusion matrix can be expressed as: confusion weight matrix Confusion bias matrix

[0130] The third layer normalization includes a parameter matrix, which includes the weight matrix L w5 and the bias matrix L bias5 The second permutation matrix corresponding to the second layer normalization is the same as the first permutation matrix and can be expressed as π e ; The corresponding confusion matrix can be expressed as: confusion weight matrix L w5 @π e , confusion bias matrix L bias5 @π e .

[0131] Exemplarily, when the i-th network layer is the last network layer of the target language model, the i-th network layer may further include a dimension mapping sublayer, namely LM_head, which is used to map the dimension of its input to the dimension of the word embedding subtable, which is also equivalent to the dimension of the confused word embedding subtable; the embedding table includes the word embedding subtable.

[0132] As mentioned above, when processing the input based on each confusion matrix in the confusion layer, the permutation concept, or the confusion concept, is that, on the one hand, for the parameter matrix that needs to be matrix multiplied with its input, the column permutation of its input is first restored, and then the column permutation of its processing result is performed based on its own corresponding permutation matrix, such as the second weight matrix and its bias matrix; on the other hand, for the parameter matrix that needs to be element-wise multiplied with its input, its own column permutation needs to be kept consistent with the permutation of its input, such as the first weight matrix and its bias matrix. In this way, for the corresponding output obtained by each confusion matrix, its permutation is only related to the permutation of the confusion matrix. It can also be understood that the output corresponding to each confusion matrix is only permuted based on the second permutation matrix corresponding to the confusion matrix itself.

[0133] Accordingly, for the dimension mapping sublayer, its input is only column-permuted based on the second permutation matrix corresponding to the parameter matrix in the previous sublayer of the dimension mapping sublayer. Considering that the weight matrix in the parameter matrix of the dimension mapping sublayer is the second weight matrix that needs to be matrix multiplied with its input, the weight matrix in the dimension mapping sublayer needs to be first left-multiplied by the transposed matrix of the second permutation matrix corresponding to the parameter matrix in the previous sublayer in order to restore the column permutation of its input, that is, to restore its input plaintext.

[0134] Next, consider that the purpose of the dimension mapping sublayer, or LM head, is to map its input dimensions to the word dimensions of the word embedding subtable, so as to determine the corresponding word based on the output of the dimension mapping sublayer and the word embedding subtable. The output of the dimension mapping sublayer may include a series of probabilities corresponding to each position predicted by the model, which represents the probability of each word in the word embedding subtable at that position. The positions may refer to each position in the output sequence predicted by the model.

[0135] It can be understood that the size of the parameter matrix of the dimensional mapping sub-layer after confusion is still equal to the parameter matrix LM of the dimensional mapping sub-layer before confusion. wThe size of the dimensionality mapping sublayer is expressed as (hidden_size, vocab_size). Among them, the value of hidden_size is the number of hidden layer dimensions. Assume that the size of the input of the dimension mapping sublayer is expressed as (seq_length (maximum sequence length), hidden_dim (hidden layer dimension)), and the value of the aforementioned hidden_size is equal to the value of hidden_dim. vocab_size represents the word dimension of the confused word embedding subtable, that is, its value is equal to the total number of words in the confused word embedding subtable. For example, the confused word embedding subtable includes 30,000 confused word embeddings, and the vocab_size can be 30,000. The value of hidden_dim is equal to the number of dimensions of the confused word embedding corresponding to each word in the confused word embedding subtable (that is, the vector length of a single confused word embedding). The size of the output of the dimension mapping sublayer can be expressed as (seq_length, vocab_size). The seq_length represents the sequence length of the output sequence corresponding to the model output.

[0136] It will be appreciated that, before obfuscation, the series of probabilities corresponding to each position in the output of the dimension mapping sublayer, i.e., the order of the probabilities of the words in the word embedding subtable at each position, is related to, e.g., has the same order as, the order of the words in the original word embedding subtable. For example, the mth probability in the series of probabilities corresponding to position p in the output of the dimension mapping sublayer may correspond to the word in the mth row of the word embedding subtable.

[0137] Since the order of the words and their confused word embeddings in the confused word embedding subtable and their confused word embeddings has changed based on the aforementioned third permutation matrix relative to the words and their confused word embeddings in the word embedding subtable after column permutation, in order to ensure that the accurate words and their confused word embeddings can be found from the confused word embedding subtable based on the output of the confused dimension mapping sublayer, the second permutation matrix corresponding to the parameter matrix of the confused dimension mapping sublayer is required to ensure that the confusion of the output of the confused dimension mapping sublayer matches the confusion of the confused word embedding subtable, that is, to ensure that the order of the series probabilities corresponding to each position in the output of the confused dimension mapping sublayer is still related to the order of arrangement between the words (and their confused word embeddings) in the confused word embedding subtable, for example, the order of arrangement is the same.

[0138] Continuing with the previous example, assume that the word in the mth row of the word embedding subtable and its confused word embedding after column permutation are processed by the third permutation matrix and permuted to the oth row of the confused word embedding subtable; then the second permutation matrix corresponding to the parameter matrix in the dimension mapping sublayer after confusion is required to ensure that the oth probability in the series of probabilities corresponding to each position in the output of the dimension mapping sublayer after confusion is the mth probability in the series of probabilities corresponding to each position in the output of the dimension mapping sublayer before confusion.

[0139] The size of the parameter matrix of the aforementioned dimensionality mapping sublayer and the size of its output can be used to determine the order of the rows of the word embedding subtable (each row is the word embedding corresponding to a word), which corresponds to the order of the columns of the output of the dimensionality mapping sublayer (a column contains the probabilities corresponding to each word in the word embedding subtable). This can be proven as follows:

[0140] Parameter matrix LM of the dimension mapping sublayer w The size of can be expressed as (hidden_size, vocab_size), the size of the word embedding subtable can be expressed as: (vocab_size, hidden_size), that is, the word embedding subtable W after column permutation v The size of can also be expressed as (vocab_size, hidden_size). Without paying attention to the values of their respective elements and only paying attention to their respective sizes, there exists: LM w =Wv T ;

[0141] And there exists: Among them, π w @W v Represents the use of the third permutation matrix π w The word embedding subtable W after column permutation v Perform row permutation, Represents the transposed matrix using the third permutation matrix The parameter matrix LM of the dimension mapping sub-layer w Perform column permutation.

[0142] In view of the above situation, the second weight matrix in the dimension mapping sublayer is left-multiplied by the transposed matrix of the second permutation matrix corresponding to the parameter matrix in the previous sublayer to restore the column permutation of its input, and then it needs to be right-multiplied by the second permutation matrix corresponding to itself; and the second permutation matrix corresponding to the second weight matrix (i.e., the parameter matrix) in the dimension mapping sublayer is related to the third permutation matrix corresponding to the confused word embedding subtable, wherein the third permutation matrix is used to perform row permutation on the word embedding subtable after column permutation. Thereby, the confusion of the output of the confused dimension mapping sublayer matches the row permutation of the confused word embedding subtable. In some examples, when the second confusion includes column permutation and row permutation, that is, when the word embedding subtable after column permutation is further permuted based on the aforementioned third permutation matrix, the second permutation matrix corresponding to the dimension mapping sublayer is the transposed matrix of the third permutation matrix.

[0143] For example, the previous sublayer of the dimension mapping sublayer may be a feedforward neural network sublayer in the last layer of the network. The parameter matrix in the dimension mapping sublayer may include a second weight matrix LM w The second permutation matrix corresponding to the dimension mapping sublayer can be expressed as π w T , and its corresponding confusion matrix can be expressed as: confusion weight matrix π e T @LM w @π w T .

[0144] In some other examples, when the second confusion includes column permutation, that is, when only the embedding table is permuted, the second permutation matrix corresponding to the dimension mapping sublayer can be the identity matrix. Accordingly, the confusion matrix corresponding to the dimension mapping sublayer can be expressed as: confusion weight matrix π e T @LM w In some other examples, in order to better protect the privacy of the first-party model parameters and text, the confusion matrix corresponding to the dimension mapping sublayer can also be expressed as: confusion weight matrix π e T @LM w @π x , to permute the output of the dimension mapping sublayer, where π x It can be the output permutation matrix set by the first party A.

[0145] Afterwards, the second party B sends the row-permuted output of the dimension mapping sublayer back to the first party A. The first party can decrypt the row-permuted output of the dimension mapping sublayer based on the decryption matrix corresponding to the output permutation matrix to obtain the row-permuted output of the dimension mapping sublayer, so that the row order of the series of probabilities at each position in the row-permuted output of the dimension mapping sublayer is consistent with the row order of each word and its corresponding confused word embedding in the word embedding table in the embedding table after column permutation. Afterwards, the first party A can determine the generated word at each position based on each word and its corresponding confused word embedding in the word embedding table in the embedding table after column permutation and the row-permuted output of the dimension mapping sublayer, and determine the model loss based on the difference between the generated word at each position and the label word at the corresponding position in the label, and send the model loss to the second party B. The second party B adjusts the multi-layer confusion layer to minimize the model loss.

[0146] First Party A can obtain multiple obfuscation layers and an obfuscation embedding table using the aforementioned method. Subsequently, if First Party A possesses the data, i.e., text, used to fine-tune the target language model, in step S320, First Party A determines a training sample based on the text and the obfuscation embedding table. The training sample includes the target obfuscated word embedding sequence and its corresponding label.

[0147] The confusion embedding table can include a confusion word embedding subtable, a confusion position embedding subtable, and a confusion tag type embedding subtable. Based on these three and each word segment in the text, the corresponding target confusion word embedding sequence and its corresponding label can be obtained, and the training sample can be remembered.

[0148] Exemplarily, the specific form of the text is related to the service that the target language model needs to provide. In some possible examples, the target language model can be a generative model. Accordingly, the text can include a first text, and the target confused word embedding sequence can be determined directly based on the first text and the confusion embedding table, and the target confused word embedding sequence is used as the label corresponding to the target confused word embedding sequence to obtain a training sample. In this example, the first text can be any type of text, for example, but not limited to news articles, scientific articles, or medical articles. The type of the first text is related to the specific needs of the first party A.

[0149] In some other possible examples, the target language model can be a seq2seq model. In this case, the aforementioned text includes an input text and its corresponding target text. Accordingly, the target confusion word embedding sequence corresponding to the input text and the label corresponding to the target text can be determined based on the confusion embedding table. When the first party A needs to fine-tune the target language model to build a question-answering system, the input text can be a question, and its corresponding target text can be the answer corresponding to the question. When the first party A needs to fine-tune the target language model to build a translation system, the input text can be a text to be translated, and its corresponding target text can be the translation corresponding to the text to be translated.

[0150] After the first party A obtains the training sample, then, in step S330, the first party A sends the aforementioned training sample and multiple confusion layers to the second party B. Correspondingly, in step S340, the second party B obtains the training sample and multiple confusion layers.

[0151] The first party A may send the aforementioned training samples and the multiple confusion layers to the second party B via a network or other data transmission method. The second party B obtains the training samples and the multiple confusion layers, deploys the multiple confusion layers, and fine-tunes the multiple confusion layers based on the training samples.

[0152] Specifically, in step S350, the second party B obtains a first output result based on the target confused word embedding sequence through multiple confusion layers. In this step, the second party B inputs the target confused word embedding sequence into the multiple confusion layers to process the target confused word embedding sequence through the multiple confusion layers to obtain the first output result.

[0153] In the case where the target language model is a generative model, the first output result may at least include the probability of each word in the confusing word embedding subtable at each position corresponding to one or more words masked in the target confusing word embedding sequence. Thereafter, for each position, the word corresponding to the probability with the largest value at the position can be used as the generated word at the position p.

[0154] When the target language model is a seq2seq model, the first output result may include the probability of each word in the confusion word embedding subtable at each position in the output sequence predicted by the multi-layer confusion layer. Thereafter, for each position, the word corresponding to the probability with the largest value at the position may be used as the generated word at the position p.

[0155] Next, in step S360 , the second party B adjusts the multiple confusion layers based on the difference between the first output result and the label.

[0156] In this step, since the second party B has not obtained the confusion embedding table, it is unable to determine the corresponding generated word at the corresponding position based on the first output result. At this time, the second party B will send the first output result it obtained to the first party A. The first party A determines the confused word embedding corresponding to the generated word at the corresponding position based on the first output result and the confused word embedding sub-table; then sends the confused word embedding corresponding to the generated word at the corresponding position determined by it to the second party B.

[0157] Next, when the target language model is a generative model, the confused word embedding corresponding to the generated word at the corresponding position obtained from the first party A is the confused word embedding corresponding to the generated word at the position corresponding to the obscured one or more words. At this time, the second party B can determine the model loss based on the confused word embedding corresponding to the generated word at the position corresponding to the obscured one or more words, and the confused word embedding at the position corresponding to the obscured one or more words in the label, that is, the target confused word embedding subtable, and adjust the parameters of the multi-layer confusion layer with the goal of minimizing the model loss.

[0158] The following example demonstrates that the confusion matrices in a multi-layer confusion layer can still maintain the correspondence between the parameter matrices in the original network layers. This demonstrates that fine-tuning the multi-layer confusion layer based on the difference between the output results and labels of text processed by the multi-layer confusion layer and the confusion embedding table is equivalent to fine-tuning the multi-layer network layer based on the difference between the output results and labels of text processed by the multi-layer network layer and the embedding table.

[0159] Taking the derivation process of matrix multiplication as an example, when processing text based on the original embedding table and multi-layer network layer, assuming that the model input X is obtained based on the text and the original embedding table, the output result obtained based on the original multi-layer network layer and model input X is expressed as D, D = XW all , where W all The parameters of the multi-layer network layer are the parameter set of all its parameter matrices. Assuming that the model loss is expressed as L, the corresponding derivative formula or gradient formula can be expressed as follows:

[0160]

[0161] Based on the above permutation idea, the confusion (column permutation) between the intermediate parameter matrices in the parameter matrix of the multi-layer confusion layer has been mutually eliminated. Accordingly, the overall confusion mode of the confusion matrix of the multi-layer confusion layer is related to the second confusion mode corresponding to the embedding table, for example, the first permutation matrix and the third permutation matrix mentioned above. Accordingly, continuing the above example, the permutation result of the confusion matrix of the multi-layer confusion layer can be expressed as in, The parameters of the multi-layer confusion layer are the parameter set of the overall confusion matrix.

[0162] Based on the above substitution idea, it can also be known that the confusion method of the output result obtained by processing the text based on multiple confusion layers and confusion embedding tables is only related to the second substitution matrix corresponding to the last confusion matrix of the last layer in the multi-layer confusion layer. In other words, the output result obtained by processing the text based on multiple confusion layers and confusion embedding tables is only permuted based on the second substitution matrix corresponding to the last confusion matrix of the last layer in the multi-layer confusion layer, relative to the output result plaintext obtained by processing the text based on multiple network layers and embedding tables. Continuing with the above example, the output result D′ obtained based on multiple confusion layers and the obfuscated model input X′ can be expressed as D'=D@π w T .

[0163] It can be understood that permuting the word embedding subtable after column permutation does not affect the correspondence between the word and its corresponding obfuscated word embedding. Accordingly, the obfuscated model input X′ obtained based on the text and the obfuscated embedding table can be expressed as X′=X@π e Assuming that the model loss (confusion) is expressed as L′, when processing text based on the confusion embedding table and multi-layer confusion layers, its derivative formula, that is, its gradient formula, can be expressed as follows:

[0164]

[0165] From the derivation formulas in the above two cases, we can see that when processing text based on the obfuscation embedding table and multi-layer obfuscation layer, the gradient obtained is The confusion situation and the parameters of the multi-layer confusion layer The confusion is consistent with that of and the parameters W of the multi-layer confusion layer a '' ll The same permutation is performed. When updating the parameters of the multi-layer confusion layer, the update method is to add the gradient to the parameters of the multi-layer confusion layer at the corresponding position. Therefore, it can be understood that the updated parameters of the multi-layer confusion layer are correspondingly permuted, or confused, with respect to the updated parameters of the original multi-layer network layer. Therefore, it is only necessary to ensure that the model loss plaintext L and the model loss (confusion) L′ are consistent or identical.

[0166] For the target language model, its output is the probability of each word in the word embedding subtable (or confused word embedding subtable) appearing at each predicted position; then, for each position, from the corresponding word embedding subtable (or confused word embedding subtable), find the word corresponding to the maximum probability at that position as the generated word for that position.

[0167] Based on the above example, when the multi-layer network layer and the embedding table are obfuscated, the obfuscation of the output result of the multi-layer obfuscation layer is maintained, which is consistent with the obfuscation of the obfuscated word embedding subtable. It can be considered that when processing text based on the obfuscated embedding table and the multi-layer obfuscation layer, compared with the case of processing text based on the embedding table and the multi-layer network layer, the two processes only replace a word embedding subtable. The correspondence between each word in the word embedding subtable and its corresponding word embedding in the two processes remains unchanged, but the row where each word and its corresponding word embedding are located in the word embedding subtable in the two processes has changed. It does not affect the result of finding the generated word at the corresponding position based on the probability of each word in the word embedding subtable (or the obfuscated word embedding subtable) appearing at each position in the model output. The only difference is that a generated word YY in the word embedding subtable is in row A, while in the obfuscated word embedding subtable, the generated word YY is in row B. Therefore, the model loss (obfuscation) L′ obtained by processing the text based on the multi-layer obfuscation layer and the obfuscation embedding table and its label is the same as the model loss plaintext L obtained by the difference between the output result and its label obtained by processing the text based on the multi-layer network layer and the embedding table.

[0168] Accordingly, the effect of fine-tuning multiple obfuscation layers based on the difference between the output results obtained by processing the text with multiple obfuscation layers and obfuscation embedding tables and its labels is equivalent to the effect of fine-tuning multiple network layers based on the difference between the output results obtained by processing the text with multiple network layers and embedding tables and its labels.

[0169] The above process can achieve multi-party joint fine-tuning of the target language model while protecting model parameters and text privacy.

[0170] The above process only describes fine-tuning the target language model using a single text. In some examples, multiple texts can be used, and the target language model can be fine-tuned multiple times through multiple iterations based on the aforementioned multi-party joint fine-tuning process until the target language model reaches a preset convergence condition, thereby determining that the target language model fine-tuning is complete. The preset convergence condition may include, but is not limited to, exceeding a specified number of iterations, exceeding a specified fine-tuning duration, or the aforementioned model loss falling below a specified threshold.

[0171] After fine-tuning the target language model, Party B can deploy the fine-tuned target language model to provide data inference services to Party A. The specific data inference process is similar to that of fine-tuning the target language model and will not be detailed here.

[0172] Figure 5A flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in another embodiment of the present specification is shown. The multi-party method includes a first party A, a second party B, and a third party C. The first party A holds a target language model, which includes an embedding table and a serially arranged multi-layer network layer. The embedding table includes a word embedding sub-table. The second party B is used to provide computing power. The third party C holds text, which is used to fine-tune the target language model.

[0173] In some possible examples, the target language model is a network model based on a Transformer structure. Multiple network layers can be referred to as Transformer layers. The network structures in the multiple network layers are similar. For example, the target language model can be a BERT (Bidirectional Encoder Representations from Transformers) model or an improved model thereof, or a GPT (Generative Pre-trained Transformer) model or an improved model thereof.

[0174] like Figure 5 As shown, the method for multi-party joint fine-tuning of the language model based on data protection includes the following steps S510-S580:

[0175] In step S510, the first party A determines a first mapping relationship and a second mapping relationship corresponding to multiple confusion layers, a confusion embedding table, and a confusion word embedding sub-table. A single confusion layer is obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer; the confusion embedding table is obtained by performing a second obfuscation on the embedding table, wherein the second obfuscation includes column permutation, and the first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method; the confusion word embedding sub-table is obtained by performing a row permutation on the word embedding sub-table after column permutation; the first mapping relationship is a mapping relationship between each word in the confusion word embedding sub-table and its position identifier in the confusion word embedding sub-table; and the second mapping relationship is a mapping relationship between the confusion word embedding of each word in the confusion word embedding sub-table and the position identifier of each word.

[0176] The implementation process of the first party A determining the multiple obfuscation layers and the obfuscation embedding table can refer to the implementation process of determining the multiple obfuscation layers and the obfuscation embedding table in the above embodiment, which will not be repeated here.

[0177] After obtaining the confusion embedding table, a confusion word embedding sub-table is determined from the confusion embedding table, and based on each word in the confusion word embedding sub-table and the position of its confusion word embedding in the confusion word embedding sub-table, a corresponding first mapping relationship and a second mapping relationship are determined. The first mapping relationship is a mapping relationship between each word in the confusion word embedding sub-table and the position identifier of its position in the confusion word embedding sub-table, which can be expressed as word->id; the second mapping relationship is a mapping relationship between the confusion word embedding of each word in the confusion word embedding sub-table and the position identifier of each word, which can be expressed as id->embedding.

[0178] Next, in step S520 , the first party A sends the multiple obfuscation layers and the second mapping relationship to the second party B. In step S530 , the first party A sends the first mapping relationship to the third party C.

[0179] In some possible examples, the embedding table may further include a position embedding subtable and a type embedding subtable. First party A further determines an obfuscated position embedding subtable and an obfuscated type embedding subtable, and sends the obfuscated position embedding subtable and the obfuscated type embedding subtable to second party B. The obfuscated position embedding subtable and the obfuscated type embedding subtable are obtained by performing column permutation on the position embedding subtable and the type embedding subtable. Accordingly, second party B also obtains the obfuscated position embedding subtable and the obfuscated type embedding subtable.

[0180] In step S540 , the third party C determines the position identifier corresponding to each word segment in the text it holds based on the first mapping relationship to obtain a target position identifier sequence.

[0181] Exemplarily, the specific form of the text is related to the service that the target language model needs to provide. In some possible examples, the target language model can be a generative model, that is, the target language model is required to provide text generation services. Accordingly, the text may include a first text. In this case, the position identifiers corresponding to each word segment in the first text can be determined directly based on the first mapping relationship to obtain a target position identifier sequence. Specifically, the third party C processes the first text it holds, obtains each word segment therein, matches each word segment with the word in the first mapping relationship, and based on the matching results, determines the position identifier corresponding to each word segment from the first mapping relationship; based on the order of each word segment in the text, the position identifiers corresponding to each word segment are combined into a target position identifier sequence.

[0182] In this example, the first text may be any type of text, such as, but not limited to, news articles, technology articles, or medical articles, etc. The type of the first text is related to the specific needs of the first party A.

[0183] In some other possible examples, the target language model can be a seq2seq model. In this case, the aforementioned text includes the input text and its corresponding target text. Accordingly, based on the first mapping relationship and the input text, the position identifiers corresponding to each word segment of the input text can be determined to obtain a first position identifier sequence; and based on the first mapping relationship and the target text, the position identifiers corresponding to each word segment of the target text can be determined to obtain a second position identifier sequence, and both the first position identifier sequence and the second position identifier sequence are included in the target position identifier sequence.

[0184] In some examples, when first party A needs to fine-tune a target language model to build a question-answering system, the input text can be a question, and the corresponding target text can be the answer to the question. When first party A needs to fine-tune a target language model to build a translation system, the input text can be the text to be translated, and the corresponding target text can be the translation of the text to be translated.

[0185] Then, in step S550, the third party C sends the target location identification sequence to the second party B. Correspondingly, the second party B obtains the target location identification sequence sent by the third party C. Then, in step S560, the second party B determines a training sample based on the obtained target location identification sequence and the second mapping relationship.

[0186] In some possible implementations, when the target language model provides text generation services, the target position identifier sequence is the position identifier sequence corresponding to the aforementioned first text. Accordingly, the second party B determines the confused word embedding corresponding to each position identifier in the target position identifier sequence based on each position identifier in the target position identifier sequence and the mapping relationship between each position identifier and the confused word embedding in the second mapping relationship to obtain a target confused word embedding sequence, and uses the target confused word embedding sequence as its corresponding label to obtain a training sample.

[0187] In some other possible implementations, when the target language model is a seq2seq model, the target position identifiers include the aforementioned first position identifier sequence corresponding to the input text and the second position identifier sequence corresponding to the target text. Subsequently, the second party B can determine the confused word embedding corresponding to each position identifier in the first position identifier sequence based on each position identifier in the first position identifier sequence and the mapping relationship between each position identifier and the confused word embedding in the second mapping relationship, to obtain a target confused word embedding sequence.

[0188] In the latter case, the second position identifier sequence can be directly used as the target confused word embedding sequence as its corresponding label to obtain a training sample. In another case, based on each position identifier in the second position identifier sequence and the mapping relationship between each position identifier and the confused word embedding in the second mapping relationship, the confused word embedding corresponding to each position identifier in the second position identifier sequence can be determined to obtain the label corresponding to the target confused word embedding sequence. In this case, the label is the confused word embedding sequence corresponding to the aforementioned target text, thereby obtaining a training sample.

[0189] In some possible examples, the second party B also obtains the aforementioned confusion position embedding subtable and confusion type embedding subtable from the first party A, and the second party B can determine the training sample based on the confusion position embedding subtable and confusion type embedding subtable, the target position identification sequence and the second mapping relationship.

[0190] Next, in step S570, the second party B obtains a first output result based on the target confusion word embedding sequence through multiple confusion layers. The implementation principle of step S570 is similar to that of the aforementioned step S350, and its implementation process can refer to the implementation process of the aforementioned step S350.

[0191] Then, in step S580, the second party B adjusts the multiple confusion layers based on the difference between the first output result and the label.

[0192] In this step, the first output result may include a series of probabilities corresponding to each position, wherein the series of probabilities corresponding to each position represents the probability of each word in the confusing word embedding subtable at that position, wherein the order between the series of probabilities corresponding to a single position corresponds to the order between the words in the confusing word embedding subtable.

[0193] For each position, based on the maximum probability in a series of probabilities corresponding to it, the position identifier and / or confused word embedding of the generated word corresponding to the position is determined from the second mapping relationship; then, based on the difference between the position identifier and / or confused word embedding of the generated word corresponding to each position and the label, the model loss is determined; and with the goal of minimizing the model loss, the multi-layer confusion layer is adjusted.

[0194] In some specific examples, the label can be the target position identification sequence of the aforementioned first text. Accordingly, the second party B can obtain the first predicted position identification sequence predicted by the model based on the first output result and the second mapping relationship obtained therefrom, and then adjust the multiple confusion layers based on the difference between the target position identification sequence and the first predicted position identification sequence. Alternatively, the label can be the target confusion word embedding sequence of the aforementioned first text. Accordingly, the second party B can obtain the generated word embedding sequence predicted by the model based on the first output result and the second mapping relationship obtained therefrom, and then adjust the multiple confusion layers based on the difference between the target confusion word embedding sequence and the generated word embedding sequence.

[0195] In some other specific examples, the label can be the aforementioned second position identification sequence. Accordingly, the second party B can obtain the second predicted position identification sequence predicted by the model based on the first output result and the second mapping relationship obtained therefrom, and then adjust the multiple confusion layers based on the difference between the second position identification sequence and the second predicted position identification sequence.

[0196] In some other specific examples, the label can be a label confusion word embedding sequence determined based on the aforementioned second position identification sequence and the second mapping relationship. Accordingly, the second party B can obtain the predicted confusion word embedding sequence predicted by the model based on the first output result and the second mapping relationship obtained therefrom, and then adjust the multi-layer confusion layer based on the difference between the label confusion word embedding sequence and the predicted confusion word embedding sequence.

[0197] In the above process, the confused word embedding subtable after the second obfuscation is split into a first mapping relationship and a second mapping relationship, and the first mapping relationship is sent to the third party C, so that the third party C can only obtain each word and its position identifier in the confused word embedding subtable, and cannot obtain the confused word embedding corresponding to the word, which can better achieve the privacy protection of the model parameters of the first party A without affecting the model fine-tuning process; the second mapping relationship is sent to the second party B, so that the second party B can only obtain the confused word embedding corresponding to each word and its position identifier in the confused word embedding subtable, and cannot obtain the specific word plaintext. On the one hand, it can achieve the privacy protection of the model parameters of the first party A, and can also achieve the privacy protection of the text of the third party C without affecting the model fine-tuning process.

[0198] In the embodiments of this specification, the target language model is secure after obfuscation and is not easily cracked by malicious actors. Taking a target language model with a dimension of 768 as an example, if a brute force search is used to crack the parameter matrix of a non-multi-head target language model (i.e., a network sublayer based on the self-attention mechanism includes one self-attention head), 768 cracking attempts are required, while for a multi-head target language model, 64!^12 cracking attempts are required.

[0199] In addition, it should be noted that for the nonlinear layer in the target language model, since the nonlinear operation is often performed on a row of its input, the parameter matrix is confused by column permutation, and after the column permutation of the embedding table, the row permutation of the word embedding sub-table is performed. Since the setting of the second permutation matrix corresponding to the last parameter matrix of the last network layer will not affect the calculation results.

[0200] Model obfuscation is performed through permutation, which has low computational cost, no impact on model performance, and no impact on model size. It only increases the communication cost of transmitting the decryption matrix and the computational cost of generating the permutation matrix and permuting the parameters. It is very friendly to practical use and has good versatility.

[0201] After fine-tuning the target language model using the above method, the second party B can directly deploy the multi-layer confusion layer of the fine-tuned target language model to provide data inference services to the third party C. During the inference process, the third party C can obtain a position identifier sequence corresponding to the text to be inferred based on its first mapping relationship, and send the position identifier sequence corresponding to the text to be inferred to the second party B. The second party B determines the confusing word embedding sequence corresponding to the text to be inferred based on the position identifier sequence corresponding to the text to be inferred and the second mapping relationship it obtained. The confusing word embedding sequence corresponding to the text to be inferred is input into the multi-layer confusion layer, and the confusing word embedding sequence corresponding to the text to be inferred is processed by the multi-layer confusion layer to obtain an inference result. The second party B then determines the position identifier corresponding to the generated word corresponding to each position in the inference result based on a series of probabilities corresponding to each position and the second mapping relationship. The second party B then sends the position identifier corresponding to the generated word corresponding to each position in the inference result to the third party C. The third party C obtains the output sequence corresponding to the inference result based on the position identifier corresponding to the generated word corresponding to each position in the inference result and the first mapping relationship it holds, thereby obtaining the output sequence corresponding to the text to be inferred. It enables joint data processing by multiple parties while protecting model parameters and third-party data privacy.

[0202] Corresponding to the above method embodiment, Figure 6 A flowchart of a method for fine-tuning a language model based on multi-party data protection in another embodiment of the present specification is shown. The aforementioned multi-party includes a first party A, a second party B, and a third party C. The first party A holds a target language model, which includes an embedding table and a multi-layer network layer arranged in series. The embedding table includes a word embedding sub-table. The second party B is used to provide computing power. The third party C holds text, which is used to fine-tune the target language model. Figure 6 As shown, the method may include the following steps S610-S670:

[0203] In step S610, first party A determines multiple obfuscation layers and obfuscation embedding tables. A single obfuscation layer is obtained by performing a first obfuscation on the parameter matrix in the corresponding network layer, and the obfuscation embedding table is obtained by performing a second obfuscation on the embedding table, where the second obfuscation includes column permutation. The first obfuscation method for the parameter matrix in each network layer corresponds to the second obfuscation method.

[0204] In step S620, first party A sends the multi-layer confusion layers to second party B. In step S630, first party A sends the confusion embedding table to third party C. In step S640, third party C determines a training sample based on the confusion embedding table and the text it holds, where the training sample includes the target confused word embedding sequence and its corresponding label. In step S650, third party C sends the training sample to second party B. In step S660, second party B obtains a first output result based on the target confused word embedding sequence through the multi-layer confusion layers.

[0205] The implementation principle of steps S610-S660 is similar to the implementation principle of the aforementioned steps S310-S350. The implementation process thereof can refer to the implementation process of the aforementioned steps S310-S350, and will not be repeated here.

[0206] In step S670, the second party B adjusts the multi-layer confusion layer based on the difference between the first output result and the label. In step S670, the second party B can send the first output result to the third party C. The third party C determines the confusion word embedding of the generated word corresponding to each position in the first output result from the confusion embedding table based on a series of probabilities corresponding to each position in the first output result, obtains the inferred confusion word embedding sequence, and feeds it back to the second party B; the second party B adjusts the multi-layer confusion layer based on the difference between the inferred confusion word embedding sequence and the label. The process of the second party B adjusting the multi-layer confusion layer based on the difference between the inferred confusion word embedding sequence and the label can refer to the process of adjusting the multi-layer confusion layer in the aforementioned step S360, which will not be repeated here.

[0207] In the above process, the first party determines multiple layers of confusion layers and confusion embedding tables to achieve parameter confusion of the target language model to avoid exposure of its network layer parameters and embedding tables to other parties. The first confusion method of the parameter matrix in each network layer corresponds to the second confusion method of the confusion embedding table to ensure that the two can still perform correct calculations after confusion; the multiple layers of confusion layers are sent to the second party, so that the model parameters of the model can be avoided from being exposed at the second party under the premise of using the computing power service provided by the second party; the confusion embedding table is sent to a third party to avoid the true embedding of each word in the confusion embedding table from being exposed at the third party. Then, the third party determines a training sample including a target confused word embedding sequence and its corresponding label based on the obtained confusion embedding table and text, and sends it to the second party, which can protect the privacy and security of its text, and the plaintext of the text will not be exposed at the second party; after the second party obtains the training sample, based on the target confused word embedding sequence, a first output result is obtained through multiple confusion layers; since the first confusion method of the parameter matrix in each network layer corresponds to the second confusion method including column permutation of the confusion embedding table, the confusion methods of the parameter matrices of each network layer cooperate with each other, so that the confusion situation of the output of the multi-layer confusion layer matches the confusion situation of the label determined based on the confusion embedding table, and then based on the difference between the first output result of the confusion situation match and the label, the multi-layer confusion layer of the target language model can be adjusted to achieve multi-party joint fine-tuning of the language model under the premise of protecting model parameters and data privacy.

[0208] After fine-tuning the target language model using the above method, the second party B can directly deploy the multiple obfuscation layers of the fine-tuned target language model to provide inference services to the third party C. The process of providing inference services to the third party C can be referred to the process of jointly fine-tuning the target language model in the above example, and will not be repeated here.

[0209] Corresponding to the above method embodiment, Figure 7 A flowchart of a method for multi-party joint fine-tuning of a language model based on data protection in another embodiment of this specification is shown. The multiple parties include a first party and a second party. The first party holds a target language model, which includes an embedding table and a multi-layer network layer arranged in series. The second party is used to provide computing power. The method is implemented by the second party, such as Figure 7 As shown, the method for multi-party joint fine-tuning of the language model based on data protection may include steps S710-S740:

[0210] In step S710, a multi-layer obfuscation layer is obtained from the first party, wherein the multi-layer obfuscation layer is obtained by the first party performing a first obfuscation on a parameter matrix in a corresponding network layer, and the first obfuscation mode of the parameter matrix in each network layer corresponds to the second obfuscation mode performed by the embedding table.

[0211] In step S720, a training sample is obtained, where the training sample includes a target obfuscated word embedding sequence and a corresponding label, where the target obfuscated word embedding sequence and the label are determined based on the text and an obfuscated embedding table, where the obfuscated embedding table is obtained by performing the second obfuscation on the embedding table, where the second obfuscation includes column permutation.

[0212] In step S730, based on the target confused word embedding sequence, a first output result is obtained through multiple confusion layers;

[0213] In step S740 , the multiple confusion layers are adjusted based on the difference between the first output result and the label.

[0214] The foregoing description describes specific embodiments of the present disclosure, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily need to be performed in the specific order shown or in a sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0215] Corresponding to the above method embodiment, the embodiment of this specification provides a device 800 for fine-tuning a language model by multiple parties based on data protection. The multiple parties include a first party and a second party. The first party holds a target language model, and the target language model includes an embedding table and a multi-layer network layer arranged in series. The second party is used to provide computing power. The device is deployed on the second party. Its schematic block diagram is as follows: Figure 8 As shown, it includes: a first acquisition module 810, configured to obtain a multi-layer confusion layer from the first party, wherein the multi-layer confusion layer is a single confusion layer of the first party, which is obtained by the first party performing a first confusion on the parameter matrix in the corresponding network layer, and the first confusion method of the parameter matrix in each network layer corresponds to the second confusion method performed by the embedding table; a second acquisition module 820, configured to obtain a training sample, the training sample includes a target confusion word embedding sequence and a corresponding label, the target confusion word embedding sequence and the label are determined based on text and a confusion embedding table, the confusion embedding table is obtained by performing the second confusion on the embedding table, and the second confusion includes column permutation; an acquisition module 830, configured to obtain a first output result based on the target confusion word embedding sequence through the multi-layer confusion layer; an adjustment module 840, configured to adjust the multi-layer confusion layer based on the difference between the first output result and the label.

[0216] The above-mentioned device embodiments correspond to the method embodiments. For detailed descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For detailed descriptions, please refer to the corresponding method embodiments.

[0217] An embodiment of this specification also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the method for multi-party joint fine-tuning of a language model based on data protection provided in this specification.

[0218] An embodiment of this specification also provides a computing device including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of multi-party joint fine-tuning of the language model based on data protection provided in this specification is implemented.

[0219] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0220] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0221] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the future development of computer technology, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0222] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flow charts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not clearly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any particular order.

[0223] For the convenience of description, the above devices are described in terms of functions divided into various modules. Of course, when implementing one or more of the present specifications, the functions of each module can be implemented in the same or multiple software and / or hardware, or the module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0224] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0225] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0226] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0227] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0228] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0229] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0230] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0231] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0232] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referenced across them. Each embodiment focuses on the differences from the other embodiments. In particular, since the system embodiments are generally similar to the method embodiments, their description is relatively simple. For relevant parts, reference can be made to the description of the method embodiments. Throughout this specification, reference to the terms "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples, and features of different embodiments or examples, described in this specification, without conflict.

[0233] The foregoing description is merely an example of one or more embodiments of this specification and is not intended to limit the one or more embodiments of this specification. Those skilled in the art will appreciate that various modifications and variations of one or more embodiments of this specification are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this specification are intended to be included within the scope of the claims.

Claims

1. A method for multi-party joint fine-tuning of language models based on data protection, wherein: The multiple parties include a first party and a second party. The first party holds a target language model, the target language model including an embedding table and a serially arranged multi-layer network layer. The second party is used to provide computing power, including: The first party determines multiple obfuscation layers and an obfuscation embedding table; and sends at least the multiple obfuscation layers to the second party; wherein a single obfuscation layer is obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer, and the obfuscation embedding table is obtained by performing a second obfuscation on the embedding table, wherein the second obfuscation includes column permutation, and the first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method; The second party obtains a training sample, wherein the training sample includes a target confused word embedding sequence and a corresponding label; based on the target confused word embedding sequence, a first output result is obtained through the multi-layer confusion layer; based on the difference between the first output result and the label, the multi-layer confusion layer is adjusted, wherein the target confused word embedding sequence and the label are determined based on the text and the confusion embedding table.

2. The method of claim 1, wherein the plurality of parties further comprises a third party holding the text, and the embedding table comprises a word embedding subtable; The first party further determines a first mapping relationship and a second mapping relationship corresponding to the confusing word embedding sub-table; and sends the first mapping relationship to the third party; The second mapping relationship is sent to the second party; wherein the confused word embedding subtable is obtained by performing row permutation on the word embedding subtable after the column permutation; the first mapping relationship is a mapping relationship between each word in the confused word embedding subtable and the position identifier of its position in the confused word embedding subtable, and the second mapping relationship is a mapping relationship between the confused word embedding of each word in the confused word embedding subtable and the position identifier of each word; The third party determines the position identifier corresponding to each word in the text based on the first mapping relationship to obtain a target position identifier sequence, and sends the target position identifier sequence to the second party; The second party obtains the training sample, including: The second party determines the training sample based on the target location identifier sequence and the second mapping relationship.

3. The method according to claim 1, wherein The plurality of parties also includes a third party holding the text; The third party obtains the obfuscation embedding table from the first party; and determines the training sample based on the text and the obfuscation embedding table; and sending the training sample to the second party.

4. The method according to claim 1, wherein The target language model is a generative model, the text includes a first text, and the label is the target confused word embedding sequence itself; or the target language model is a seq2seq model, the text includes an input text and its corresponding target text, and the label is determined based on the target text and the confused embedding table; the target confused word embedding sequence is determined based on the input text and the confused embedding table.

5. The method according to any one of claims 1 to 4, wherein: The determining of multiple obfuscation layers and obfuscation embedding tables includes: The first party generates a first permutation matrix corresponding to the embedding table and second permutation matrices corresponding to the parameter matrices in the network layers; at least right-multiplying the embedding table by the first permutation matrix to obtain an obfuscated embedding table; For each parameter matrix in each network layer, at least a second permutation matrix corresponding to the parameter matrix is used to process the parameter matrix to obtain a confusion matrix corresponding to the parameter matrix, so as to obtain a multi-layer confusion layer.

6. The method according to claim 5, wherein: The parameter matrix includes a first weight matrix that needs to be element-wise multiplied with its input; Obtaining the confusion matrix corresponding to the parameter matrix includes: The first weight matrix is right-multiplied by its corresponding second permutation matrix to obtain a confusion matrix corresponding to the first weight matrix, wherein the second permutation matrix corresponding to the first weight matrix makes the confusion condition of the confusion matrix corresponding to the first weight matrix consistent with the confusion condition of its input.

7. The method according to claim 5, wherein: The parameter matrix includes a second weight matrix that needs to be matrix multiplied with its input; Obtaining the confusion matrix corresponding to the parameter matrix includes: The second weight matrix is multiplied on the left by the transposed matrix of the second permutation matrix corresponding to the previous weight matrix, and on the right by the corresponding second permutation matrix, to obtain a confusion matrix corresponding to the second weight matrix.

8. The method of claim 5, wherein: The last network layer also includes a dimension mapping sublayer, which is used to map the dimension of its input to the word dimension of the word embedding subtable; the embedding table includes the word embedding subtable; The second permutation matrix corresponding to the dimension mapping sublayer is the transpose of the third permutation matrix corresponding to the confused word embedding subtable, wherein the confused word embedding subtable is obtained by performing row permutation on the word embedding subtable after the column permutation, and the third permutation matrix is used to perform row permutation on the word embedding subtable after the column permutation.

9. A method for multi-party joint fine-tuning of language models based on data protection, wherein: The multiple parties include a first party and a second party, the first party holds a target language model, the target language model includes an embedding table and a serially arranged multi-layer network layer, the second party is used to provide computing power, and the method is implemented by the second party, including: Obtaining multiple obfuscation layers from the first party, wherein the multiple obfuscation layers are obtained by performing a first obfuscation on a parameter matrix in a corresponding network layer by the first party, and the first obfuscation method of the parameter matrix in each network layer corresponds to the second obfuscation method performed by the embedding table; Obtaining a training sample, the training sample including a target obfuscated word embedding sequence and a corresponding label, the target obfuscated word embedding sequence and the label being determined based on text and an obfuscated embedding table, the obfuscated embedding table being obtained by performing the second obfuscation on the embedding table, the second obfuscation including column permutation; Based on the target confused word embedding sequence, obtain a first output result through the multi-layer confusion layer; The multiple confusion layers are adjusted based on a difference between the first output result and the label.

10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to claim 9 is implemented.

Citation Information

Patent Citations

  • Error correction method and device for Chinese text in power field, storage medium and computing equipment

    CN114118065A

  • Complex control flow obfuscation method based on LLVM compiling optimization stage

    CN115438318A

  • Multi-party joint data processing method and device based on data protection

    CN118862124A

  • Model fine tuning method, text processing method, medium, equipment and program product

    CN119150862A

  • SYSTEM AND METHOD FOR TRAINING HETEROGENEOUS MODELS WITH STACKED ENSEMBLES ON DECENTRALIZED DATA

    DE102022108864A1

Cited By

  • Multi-party joint data processing method and device based on data protection

    CN118862124A

  • A method and apparatus for multi-party joint processing of data based on data protection

    CN118862124B