Model fine-tuning method, text processing method, medium, device and program product

By obtaining and obfuscating the parameters of the large language model on the server, only sending the embedded layer parameters to the server, and retaining the word list at the user, the risk of customer data leakage in the large language model service is solved, and the balance between data protection and model effect is achieved.

CN119150862BActive Publication Date: 2025-05-16BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411132183.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2025-05-16
Estimated Expiration
2044-08-16

AI Technical Summary

Technical Problem

The fine-tuning and inference services provided by the large language model on the cloud service platform poses a risk of customer data leakage.

Method used

By obtaining the vocabulary list and embedding layer parameters of the large language model on the server, obfuscation processing is performed, target vocabulary list and target embedding layer parameters are generated, and target embedding layer parameters are only sent to the server, and the target vocabulary list is retained on the server, using the target vocabulary list for text word segmentation and indexing operations, and sending word element index to the server for fine-tuning.

Benefits of technology

It effectively protects the data of the server users, avoids the server obtaining the correspondence between embedded vectors and word elements, guarantees the model effect, and reduces the resource occupation of the server users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119150862B_ABST
    Figure CN119150862B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a model fine-tuning method, a text processing method, a medium, a device and a program product. The method includes: obtaining the vocabulary and embedding layer parameters of a large language model from a server; performing obfuscation processing on the vocabulary and embedding layer parameters respectively to obtain a target vocabulary and target embedding layer parameters; sending the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; using the target vocabulary, performing word segmentation and index conversion operations on the text sample to obtain the first word element index corresponding to the text sample; sending the first word element index to the server, so that the server fine-tunes the new large language model based on the first word element index. In this way, the server can avoid obtaining the corresponding relationship between the embedding vector and the word element, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of privacy protection, and in particular, to a model fine-tuning method, a text processing method, a medium, a device, and a program product. Background Art

[0002] Large Language Models (LLM) are becoming more and more widely used, and many customers want to use the LLM services provided by cloud service platforms. Customers usually use two typical functions provided by LLM services: fine-tuning and inference (i.e., prediction). First, customers upload private datasets to the cloud service platform to fine-tune the pre-trained model deployed on the platform to improve the model's task performance in specific fields. After that, customers will call the inference function of the fine-tuned model through the platform and enter the inference text to obtain the prediction result.

[0003] However, while the above solution provides customers with efficient and customizable LLM services, it also brings the risk of customer data leakage. Summary of the invention

[0004] This summary is provided to introduce concepts in a brief form that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a large language model fine-tuning method, comprising: obtaining a vocabulary and embedding layer parameters of a large language model from a server, wherein the pre-trained large language model is deployed on the server; performing obfuscation processing on the vocabulary and the embedding layer parameters respectively to obtain a target vocabulary and target embedding layer parameters; sending the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; using the target vocabulary, performing word segmentation and index conversion operations on a text sample to obtain a first word element index corresponding to the text sample; sending the first word element index to the server, so that the server fine-tunes the new large language model based on the first word element index.

[0006] In a second aspect, the present disclosure provides a large language model fine-tuning method, which is applied to a server, comprising: in response to receiving a target embedding parameter sent by a server user, updating the embedding layer parameters of the large language model of the server to the target embedding layer parameters to obtain a new large language model, wherein the target embedding layer parameters are obtained by the server user performing obfuscation processing on the embedding layer parameters; in response to receiving a first word element index sent by the server user, fine-tuning the new large language model based on the first word element index, wherein the first word element index is generated by the server user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server user performing obfuscation processing on the vocabulary of the large language model.

[0007] In a third aspect, the present disclosure provides a text processing method, which is applied to a server user, comprising: obtaining a text to be processed; performing word segmentation and index conversion operations on the text to be processed using a target vocabulary to obtain a second word element index corresponding to the text to be processed; sending the second word element index to the server, so that the server generates a text processing result of the text to be processed based on the second word element index through a pre-fine-tuned large language model, and sending the text processing result to the server user, wherein the large language model is obtained according to the large language model fine-tuning method provided by the first and second aspects of the present disclosure; receiving the text processing result.

[0008] In a fourth aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the large language model fine-tuning method provided in the first and second aspects of the present disclosure or the steps of the text processing method provided in the third aspect of the present disclosure.

[0009] In a fifth aspect, the present disclosure provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of the large language model fine-tuning method provided in the first and second aspects of the present disclosure or the steps of the text processing method provided in the third aspect of the present disclosure.

[0010] In a sixth aspect, the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the large language model fine-tuning method provided in the first and second aspects of the present disclosure or the steps of the text processing method provided in the third aspect of the present disclosure.

[0011] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0012] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale. In the drawings:

[0014] Figure 1 It is a schematic diagram of the structure of a large language model according to an exemplary embodiment.

[0015] Figure 2It is a schematic diagram of the structure of a large language model according to another exemplary embodiment.

[0016] Figure 3 The figure is a schematic diagram of a process of fine-tuning a large language model and processing text according to an exemplary embodiment.

[0017] Figure 4 The present invention is a flowchart of a method for fine-tuning a large language model applied to a server user according to an exemplary embodiment.

[0018] Figure 5 The present invention is a flowchart of a method for fine-tuning a large language model applied to a server according to an exemplary embodiment.

[0019] Figure 6 The figure is a flowchart of a text processing method according to an exemplary embodiment.

[0020] Figure 7 The present invention is a block diagram of a large language model fine-tuning device applied to a server user according to an exemplary embodiment.

[0021] Figure 8 It is a block diagram of a large language model fine-tuning device applied to a server according to an exemplary embodiment.

[0022] Fig. 9 It is a block diagram of a text processing device according to an exemplary embodiment.

[0023] Fig.10 The diagram is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0026] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0027] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0028] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0030] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0031] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0032] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0033] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0034] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0035] Before introducing the specific implementation of the present disclosure, the structure of the large language model involved in the present disclosure is first explained.

[0036] A large language model is a model used to analyze, understand, and generate text data. A large language model can be used to perform a variety of tasks, including but not limited to: classification tasks, generation tasks, etc.

[0037] In one implementation, the large language model is an LLM, such as Figure 1 and Figure 2 As shown in the figure, the large language model includes a word segmenter containing a vocabulary, an input embedding layer, a backbone network composed of transformers connected in series, and an output layer related to the task.

[0038] Among them, the tokenizer with a vocabulary is used to divide a piece of input text into multiple tokens, and map these tokens into a series of indexes according to the order of these tokens in the vocabulary, and input the token indexes into the input embedding layer, such as Figure 1 and Figure 2 As shown, the vocabulary includes multiple word units and an index (Index) corresponding to each word unit.

[0039] The input embedding layer is used to obtain the embedding vectors corresponding to the above multiple words from the input embedding layer parameters according to the above word index, and obtain the embedding vector matrix corresponding to multiple tokens in the input text, wherein the input embedding layer parameters include the embedding vectors corresponding to each word in the vocabulary.

[0040] The backbone network is usually a neural network composed of multiple transformer structures connected in series. For example, the Bidirectional Encoder Representation from Transformers (BERT) model contains 12 layers of encoder-decoder type transformers, and the Large Language Model Meta AI (LLaMA) model contains 32 layers of decoder-only type transformers. The input and output data of each transformer structure have the same form: they are all embedding vector matrices of the same dimension.

[0041] When the above large language model performs classification tasks, it is specifically a text classification model, such as Figure 1 As shown, the output layer of the text classification model generally adopts a multilayer perceptron (MLP), where the MLP takes the first vector of the embedding vector matrix output by the backbone network (i.e., the first row of the embedding vector matrix) as input, and finally outputs a vector to represent the probability of each preset category.

[0042] When the above large language model performs a generation task, it is specifically a text generation model, such as Figure 2 As shown in the figure, the output layer of the text generation model usually includes an output embedding layer (called Language model head or Outputembedding layer). The output embedding layer is used to convert the embedding vector matrix output by the backbone network into the probability value of each Token in the vocabulary according to the output embedding layer parameters, and determine the next generated Token according to the size of the probability value, where the output embedding layer parameters include the embedding vector corresponding to each word in the vocabulary. The text generation model adopts different processing logic in the model fine-tuning stage and the inference stage:

[0043] During the fine-tuning phase, all embedding vectors (i.e., the complete embedding vector matrix) output by the last Transformer of the backbone network are usually input into the output embedding layer to calculate the probability value of each token corresponding to the next token, and update the model based on the difference with the input text.

[0044] In the inference phase, usually only the last embedding vector output by the last Transformer of the backbone network (that is, the last row of the embedding vector matrix output by the last Transformer of the backbone network) is taken as the input of the output embedding layer to calculate the next token predicted based on the input text. Figure 2 As shown in the figure, the output token will be concatenated into the input text and input into the model again for prediction until the model outputs a stop identifier or the length of the predicted text reaches the maximum length limit.

[0045] like Figure 3 As shown, the large language model fine-tuning method in the present disclosure may include a model conversion phase and a model fine-tuning phase.

[0046] Among them, in the model conversion stage: the server user confuses the vocabulary and embedding layer parameters of the large language model respectively, and then retains the vocabulary obtained after obfuscation (i.e., the target vocabulary), and only sends the obfuscated embedding layer parameters to the server, so that the server uses the obfuscated embedding layer parameters to update the embedding layer parameters of the large language model on the server. The server user is a client-side device, which can be a terminal device or a server, and the server is a cloud service platform.

[0047] In the model fine-tuning stage: the server user uses the target vocabulary to perform word segmentation and indexing operations on the text samples, obtains a word-unit index set (i.e., the first word-unit index), and sends it to the server, so that the server can fine-tune the large language model obtained after updating the embedding layer parameters based on the first word-unit index.

[0048] Specifically, if Figure 4 As shown, the large language model fine-tuning method provided by the present disclosure and applied to a server user may include the following S101 to S105.

[0049] In S101, the vocabulary and embedding layer parameters of the large language model are obtained from the server.

[0050] In the present disclosure, a pre-trained large language model is deployed on the server side, which can be pre-trained by the server side using a large amount of text in a general field. The large language model can be a text classification model or a text generation model.

[0051] Among them, when the large language model is a text classification model, such as Figure 1 and Figure 3 As shown, the text classification model includes an input embedding layer, and in this case, the embedding layer parameters of the large language model include input embedding layer parameters (e.g., input embedding layer parameter matrix).

[0052] When the large language model is a text generation model, such as Figure 2 and Figure 3 As shown, the text generation model includes an input embedding layer and an output embedding layer. At this time, the embedding layer parameters of the large language model include input embedding layer parameters (e.g., input embedding layer parameter matrix) and output embedding layer parameters (e.g., output embedding layer parameter matrix).

[0053] The server user can interact with the server to obtain the vocabulary and embedding layer parameters of the pre-trained large language model on the server.

[0054] In S102, the vocabulary and the embedding layer parameters are obfuscated respectively to obtain a target vocabulary and a target embedding layer parameters.

[0055] In S103, the target embedding layer parameters are sent to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model.

[0056] In this disclosure, Figure 3 As shown, after the server user confuses the vocabulary and embedding layer parameters respectively, the target vocabulary and target embedding layer parameters can be obtained. After that, the target vocabulary is retained locally, and only the target embedding layer parameters (i.e., the obfuscated embedding layer parameter matrix) are sent to the server, so that the server can update the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model.

[0057] In one embodiment, when the above-mentioned large language model is a text classification model, the embedding layer parameters only include the input embedding layer parameters. At this time, the server-side user performs obfuscation processing on the vocabulary and the input embedding layer parameters respectively to obtain the target vocabulary and the obfuscated input embedding layer parameters, that is, the target embedding layer parameters include the obfuscated input embedding layer parameters; then, the server-side user sends the obfuscated input embedding layer parameters to the server; next, the server updates the parameters of the input embedding layer of the local large language model to the obfuscated input embedding layer parameters to obtain a new large language model.

[0058] In another embodiment, when the above-mentioned large language model is a text generation model, the embedding layer parameters include input embedding layer parameters and output embedding layer parameters. At this time, the server-side user performs obfuscation processing on the vocabulary, input embedding layer parameters and output embedding layer parameters respectively to obtain the target vocabulary, the obfuscated input embedding layer parameters and the obfuscated output embedding layer parameters, that is, the target embedding layer parameters include the obfuscated input embedding layer parameters and the obfuscated output embedding layer parameters; then, the server-side user sends the obfuscated input embedding layer parameters and the obfuscated output embedding layer parameters to the server; next, the server updates the parameters of the input embedding layer of the local large language model to the obfuscated input embedding layer parameters, and at the same time, updates the parameters of the output embedding layer of the local large language model to the obfuscated output embedding layer parameters, thereby obtaining a new large language model.

[0059] In S104, the target vocabulary is used to perform word segmentation and index conversion operations on the text sample to obtain a first word element index corresponding to the text sample.

[0060] In S105 , the first word-unit index is sent to the server, so that the server fine-tunes the new large language model based on the first word-unit index.

[0061] In this disclosure, Figure 3As shown, after obtaining the target vocabulary, the server user can first use the target vocabulary to segment the text sample; then, according to the correspondence between the target vocabulary segmentation and the index (such as Figure 3 The target vocabulary shown in ), determines the index corresponding to each word unit obtained after word segmentation, that is, obtains the first word unit index corresponding to the text sample, and sends the first word unit index to the server; after receiving the first word unit index, the server fine-tunes the new large language model based on the first word unit index.

[0062] It should be noted that the above steps S104 and S105 can be performed after the above step S103 (e.g. Figure 3 As shown), it may also be executed before the above S103, or it may be executed synchronously with the above S103, and the present disclosure does not make any specific limitation on this.

[0063] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0064] The following is a detailed description of the specific implementation method of performing obfuscation processing on the vocabulary and embedding layer parameters in the above S102 to obtain the target vocabulary and target embedding layer parameters. Specifically, it can be achieved by the following steps (1) to (3):

[0065] Step (1): Generate random permutations.

[0066] Step (2): Perform a permutation operation on the vocabulary using random permutation to obtain a target vocabulary, and perform a permutation operation on the embedding layer parameters using random permutation to obtain permuted embedding layer parameters.

[0067] In the present disclosure, the random permutation may be an index sequence of length n, and then the vocabulary may be rearranged according to the order of the indexes in the index sequence to obtain the target vocabulary, where n is the number of tokens in the vocabulary and also the number of tokens in the target vocabulary. The same random permutation is used to perform permutation operations on the vocabulary and the embedding layer parameters.

[0068] For example, n=3, and the random permutation is {2, 3, 1}. Then the original second word in the vocabulary and its index can be placed in the first position of the vocabulary, the original third word in the vocabulary and its index can be placed in the second position of the vocabulary, and the original first word in the vocabulary and its index can be placed in the third position of the vocabulary, thereby obtaining the target vocabulary.

[0069] For example, Figure 3 In the model conversion stage shown in the figure, by random permutation, Figure 3 The vocabulary in is replaced by Figure 3 The target vocabulary shown in .

[0070] Step (3): According to the target vocabulary, the permutation embedding layer parameters are subjected to noise perturbation processing to obtain the target embedding layer parameters.

[0071] In this disclosure, Figure 3 As shown, for the embedding layer parameters, it is necessary to perform a permutation operation on them first, and then, according to the target vocabulary, the permuted embedding layer parameters (i.e., permuted embedding layer parameters) are subjected to noise processing (i.e., noise perturbation processing) to obtain the target embedding layer parameters, specifically, the confused embedding parameter matrix.

[0072] Specifically, the noise perturbation process can be performed on the permutation embedding layer parameters through the following steps (31) and (32):

[0073] Step (31): Cluster the words in the target vocabulary according to the permuted embedding layer parameters.

[0074] In the present disclosure, the word units in the target vocabulary can be clustered using clustering methods such as Neighbor Distance Preserving Obfuscation (NDPO) algorithm and K-means clustering (K-means) according to the permutation embedding layer parameters.

[0075] Preferably, the NDPO algorithm can be used to cluster the word elements in the target vocabulary to reduce the overhead of the server-side user.

[0076] Step (32): According to the clustering results, the permuted embedding layer parameters are subjected to noise perturbation processing to obtain the target embedding layer parameters.

[0077] The following is a detailed description of the specific implementation method of clustering the word elements in the target vocabulary according to the replacement embedding layer parameters in the above step (31). Specifically, it can be implemented by using the NDPO algorithm according to the replacement embedding layer parameters through the following steps (311) to (316):

[0078] Step (311): Take any word in the target vocabulary as the current word.

[0079] For example, the first word-gram in the target vocabulary may be used as the current word-gram.

[0080] Step (312): Calculate the second similarity between the current word and each other word in the target vocabulary that has not been clustered according to the permuted embedding layer parameters.

[0081] Step (313): The current word-gram and K-1 other words-grams that have the second highest similarity to the current word-gram and have not been clustered are grouped as a cluster.

[0082] In the present disclosure, the current word and the K-1 other words that are not clustered and have the second highest similarity to the current word can be used as a cluster cluster, or the index of the current word and the index of the K-1 other words that are not clustered and have the second highest similarity to the current word can be used as a cluster cluster. K is used to represent the number of words in the cluster cluster, that is, the cluster size, which can be a default value or a value set by the user.

[0083] Step (314): Determine whether the number of clusters reaches N.

[0084] In this disclosure, n is the number of word-grams in the target vocabulary, that is, the number of word-grams in the vocabulary of the large language model.

[0085] If the number of clusters does not reach N, the following step (315) is executed, and then the process returns to the above step (312); if the number of clusters reaches N, it indicates that the clustering is completed, and at this time, the following step (316) can be executed.

[0086] Step (315): Take any word-gram among the unclustered word-grams in the target vocabulary as the current word-gram.

[0087] Step (316): Take N clusters as clustering results.

[0088] The following is a detailed description of the specific implementation method of calculating the second similarity between the current word and each other word in the target vocabulary that has not been clustered according to the replacement embedding layer parameters in the above step (312). Specifically, it can be achieved by the following steps (a1) and (a2).

[0089] Step (a1): For each other word-unit that is not clustered in the target vocabulary, the semantic similarity and edit distance between the other word-unit and the current word-unit are calculated according to the permutation embedding layer parameters.

[0090] In the present disclosure, edit distance, also known as Levenshtein distance, is a method of measuring the difference between two strings, that is, the minimum number of single-character editing (insertion, deletion or substitution) operations required to convert from one string to another.

[0091] Specifically, the semantic similarity between the other word-unit and the current word-unit can be calculated in the following way: obtain the first embedding vector corresponding to the other word-unit and the second embedding vector corresponding to the current word-unit from the permutation embedding layer parameters; then, measure the semantic similarity between the first embedding vector and the second embedding vector by cosine similarity, Euclidean distance, Manhattan distance, Pearson Correlation Coefficient, etc., as the semantic similarity between the other word-unit and the current word-unit.

[0092] Since the specific implementation of calculating the edit distance between two character strings is well known to those skilled in the art, the present disclosure will not further elaborate on the specific implementation of calculating the edit distance between the other word-gram and the current word-gram.

[0093] Step (a2): Determine a second similarity between the other word-gram and the current word-gram based on the semantic similarity and the edit distance between the other word-gram and the current word-gram.

[0094] For example, the second similarity between the other word-gram and the current word-gram may be determined according to the semantic similarity and the edit distance by the following equation (1):

[0095]

[0096] Among them, s ij is the second similarity between other words and the current word; e′ i is the second embedding vector corresponding to the current word; e′ j is the first embedding vector corresponding to the other word; sim(e′ i ,e′ j ) is the semantic similarity between the other word and the current word; edit_dist(v′ i ,v′ j ) is the other word v′ j With the current word v′ i The edit distance between |v′ i | is the current word v′ i The character length of |v′ j | is the other word v′ j The character length of

[0097] The following describes in detail a specific implementation method of obtaining the first embedding vector corresponding to the other word unit and the second embedding vector corresponding to the current word unit from the replacement embedding layer parameters.

[0098] Specifically, the index of the other word and the index of the current word can be obtained according to the target vocabulary; then, according to the index of the other word, the embedding vector corresponding to the index of the other word is obtained from the replacement embedding layer parameters as the first embedding vector, and at the same time, according to the index of the current word, the embedding vector corresponding to the index of the current word is obtained from the replacement embedding layer parameters as the second embedding vector.

[0099] The following is a detailed description of the specific implementation method of performing noise perturbation processing on the replacement embedding layer parameters according to the clustering results in the above step (32) to obtain the target embedding layer parameters. Specifically, it can be achieved by the following steps (321) to (323):

[0100] Step (321): for each cluster in the clustering result, determine the embedding vector set corresponding to the cluster from the permutation embedding layer parameters.

[0101] In the present disclosure, the embedding vector corresponding to each word unit or word unit index in the cluster can be determined from the permutation embedding layer parameters, and these embedding vectors constitute an embedding vector set.

[0102] Step (322): Use the differential privacy mechanism to perform noise perturbation processing on the embedding vector set corresponding to the cluster cluster to obtain the perturbation vector set corresponding to the cluster cluster.

[0103] In the present disclosure, differential privacy is a mathematical technique that adds controlled random noise to an embedded vector set to prevent the server from obtaining private data about the embedded vector set. Among them, differential privacy mechanisms such as Laplace and Gaussian can be used to perform noise perturbation processing on the embedded vector set.

[0104] Preferably, the Laplace differential privacy mechanism can be used to perform noise perturbation processing on the embedded vector set to make the data of the server user more secure.

[0105] Step (323): Integrate the perturbation vector set corresponding to each cluster to obtain the target embedding layer parameters.

[0106] In the present disclosure, the perturbation vectors corresponding to each word in the N clusters can be integrated according to the order of each word in the target vocabulary to obtain the target embedding layer parameters. The perturbation vector set corresponding to the cluster includes the perturbation vector corresponding to each word in the cluster.

[0107] The following is a detailed description of the specific implementation method of using the differential privacy mechanism in the above step (322) to perform noise perturbation processing on the embedded vector set corresponding to the cluster cluster to obtain the perturbation vector set corresponding to the cluster cluster. Specifically, the Laplace differential privacy mechanism can be used to perform noise perturbation processing on the embedded vector set through the following steps (b1) to (b6):

[0108] Step (b1): for each embedding vector in the embedding vector set corresponding to the cluster, calculate a first similarity between the embedding vector and each embedding vector in the embedding vector set.

[0109] In the present disclosure, the first similarity between the embedding vector and each embedding vector in the embedding vector set can be measured by cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, etc.

[0110] Step (b2): According to a preset privacy budget, all first similarities corresponding to the embedded vector are mapped into a first probability distribution.

[0111] In the present disclosure, a preset privacy budget is used to control the smoothness of the first probability distribution.

[0112] For example, according to a preset privacy budget, all first similarities corresponding to the embedding vector can be mapped to a first probability distribution through the following equation (2):

[0113]

[0114] Wherein, the first probability distribution W = {w j|1≤j≤|T|}; T is the embedding vector set corresponding to the cluster; |T| is the number of words or word indexes in the embedding vector set corresponding to the cluster; σ j is the jth embedding vector u in the embedding vector set corresponding to the embedding vector and the cluster j The first similarity between k is the first similarity between the embedding vector and the kth embedding vector in the embedding vector set corresponding to the cluster, 1≤k≤|T|; ∈ is the preset privacy budget; w j is the jth probability in the first probability distribution.

[0115] Step (b3): ​​According to the preset privacy budget, add Laplace noise to the first probability distribution to obtain a second probability distribution.

[0116] Specifically, the difference between the maximum probability and the minimum probability in the first probability distribution can be determined as the sensitivity s; then, according to the sensitivity s and the preset privacy budget, Laplace noise is added to the first probability distribution to obtain the second probability distribution W′=W+ Among them, laplace() is the Laplace noise function.

[0117] Step (b4): normalize the second probability distribution to obtain a third probability distribution.

[0118] For example, the second probability distribution W′ may be normalized by the following equation (3) to obtain a third probability distribution W″:

[0119]

[0120] Among them, w″ j is the jth probability in the third probability distribution W″; w′ j is the jth probability in the second probability distribution W′; w′ k is the kth probability in the second probability distribution W′.

[0121] Step (b5): Determine a perturbation vector corresponding to the embedding vector according to the third probability distribution and the embedding vector set.

[0122] For example, the perturbation vector corresponding to the embedding vector can be determined according to the third probability distribution and the embedding vector set by the following equation (4):

[0123]

[0124] in, is the perturbation vector corresponding to the embedding vector, t i is the index of the embedding vector in the target embedding layer parameters.

[0125] Step (b6): Integrate the perturbation vector corresponding to each embedded vector to obtain the perturbation vector set corresponding to the cluster.

[0126] In the present disclosure, after obtaining the perturbation vectors corresponding to the respective embedding vectors in the embedding vector set corresponding to the cluster, they are aggregated to obtain the perturbation vector set corresponding to the cluster.

[0127] It should be noted that when the above-mentioned large language model is a text classification model, the obfuscation processing of the input embedding layer parameters can be implemented through steps (1) to (3) for the input embedding layer parameters.

[0128] When the above-mentioned large language model is a text generation model, the input embedding layer parameters can be obfuscated through steps (1) to (3) for the input embedding layer parameters. At the same time, the output embedding layer parameters can be obfuscated through steps (1) to (3) for the output embedding layer parameters.

[0129] Among them, the cluster size of the clustering cluster used in the obfuscation processing of the input embedding layer parameters can be the same as or different from the cluster size of the clustering cluster used in the obfuscation processing of the output embedding layer parameters. The privacy budget used in the obfuscation processing of the input embedding layer parameters can be the same as or different from the privacy budget used in the obfuscation processing of the output embedding layer parameters, and the present disclosure does not make specific limitations.

[0130] In the above implementation, the larger the privacy budget and the smaller the cluster size, the better the model effect, but the weaker the privacy protection strength; conversely, the smaller the privacy budget and the larger the cluster size, the worse the model effect, but the stronger the privacy protection strength. Users can balance the relationship between model effect and privacy by adjusting the cluster size and privacy budget during obfuscation, thereby providing users with adjustable data protection functions while ensuring model effect.

[0131] The following describes in detail a specific implementation method for fine-tuning the new large language model based on the first word unit index on the server.

[0132] In one implementation, the large language model is a text classification model; in this case, the large language model fine-tuning method applied to the server user may further include the following steps:

[0133] Get the classification label corresponding to the text sample;

[0134] At this time, the above S105 may include:

[0135] The first word unit index and the classification label are sent to the server, so that the server can fine-tune the new large language model based on the first word unit index and the classification label.

[0136] In the present disclosure, the server user first uses the target vocabulary to segment the text sample; then, according to the correspondence between the target vocabulary segmentation and the index (such as Figure 3 The target vocabulary shown in ), determines the index corresponding to each word unit obtained after word segmentation, that is, obtains the first word unit index corresponding to the text sample, and sends the first word unit index and the classification label to the server; after the server receives the first word unit index and the classification label, it fine-tunes the new large language model based on the first word unit index and the classification label.

[0137] Specifically, the server can fine-tune the new large language model based on the first word index and classification label in the following ways:

[0138] The server can use the first word element index as the input of the input embedding layer of the new large language model, use the output of the input embedding layer as the input of the backbone network, use the first vector of the embedding vector matrix output by the backbone network as the input of the MLP, and use the classification label corresponding to the above text sample as the target output of the MLP to fine-tune the new large language model.

[0139] The server can calculate the model loss based on the difference between the predicted category output by the MLP and the classification label corresponding to the above text sample, and adjust the model parameters of the new large language model according to the model loss.

[0140] In another embodiment, the large language model is a text generation model, such as Figure 3 As shown above, the server user first uses the target vocabulary to segment the text sample, and then, according to the correspondence between the target vocabulary segmentation and the index (such as Figure 3 The target vocabulary shown in ), determines the index corresponding to each word unit obtained after word segmentation, that is, obtains the first word unit index corresponding to the text sample, and sends the first word unit index to the server; after receiving the first word unit index, the server fine-tunes the new large language model based on the first word unit index.

[0141] Specifically, the server can fine-tune the new large language model based on the first word index in the following ways:

[0142] like Figure 3 As shown, the server can use the first word element index as the input of the input embedding layer of the new large language model, use the output of the input embedding layer as the input of the backbone network, use the embedding vector matrix output by the backbone network as the input of the output embedding layer, and use the above text sample as the target output of the output embedding layer to fine-tune the new large language model.

[0143] The server may calculate the model loss based on the difference between the predicted word unit index output by the output embedding layer and the first word unit index, and adjust the model parameters of the new large language model according to the model loss.

[0144] Figure 5 FIG. 1 is a flow chart showing a method for fine-tuning a large language model applied to a server according to an exemplary embodiment. Figure 5 As shown, the large language model fine-tuning method applied to the server may include the following S201 and S202.

[0145] In S201, in response to receiving the target embedding parameters sent by the server user, the embedding layer parameters of the large language model on the server are updated to the target embedding layer parameters to obtain a new large language model.

[0146] The target embedded layer parameters are obtained by the server user through obfuscating the embedded layer parameters.

[0147] In S202, in response to receiving the first word-unit index sent by the server user, the new large language model is fine-tuned based on the first word-unit index.

[0148] The first word-meta index is generated by the server-side user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server-side user through obfuscation processing on the vocabulary of a large language model.

[0149] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0150] In a possible implementation, the large language model is a text classification model; in response to receiving the first word-meta index sent by the server-side user, fine-tuning the new large language model based on the first word-meta index includes: in response to receiving the first word-meta index sent by the server-side user and a classification label corresponding to the text sample, fine-tuning the new large language model based on the first word-meta index and the classification label.

[0151] In a possible implementation, the large language model is a text generation model, wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

[0152] The specific implementation methods of each step in the large language model fine-tuning method applied to the server according to the embodiment of the present disclosure have been described in detail in the large language model fine-tuning method applied to the server user according to the embodiment of the present disclosure, and will not be repeated here.

[0153] Figure 6 is a flowchart of a text processing method according to an exemplary embodiment. The text processing method can be applied to a server user. Figure 6 As shown, the text processing method may include the following S301 to S304.

[0154] In S301, the text to be processed is obtained.

[0155] In S302, the target vocabulary is used to perform word segmentation and index conversion operations on the text to be processed to obtain a second word element index corresponding to the text to be processed.

[0156] In the present disclosure, a similar method as using the target vocabulary to perform word segmentation and indexing operations on text samples in S104 above can be adopted to use the target vocabulary to perform word segmentation and indexing operations on the text to be processed, which will not be repeated in the present disclosure.

[0157] In S303, the second word-unit index is sent to the server, so that the server generates a text processing result of the text to be processed based on the second word-unit index by using a pre-fine-tuned large language model, and sends the text processing result to the server user.

[0158] In this disclosure, Figure 3 As shown, after obtaining the second word-unit index, the server user sends it to the server; after receiving the second word-unit index, the server generates a text processing result of the text to be processed based on the second word-unit index using the pre-adjusted large language model. When the large language model is a text classification model, the text processing result is the category of the text to be processed. When the large language model is a text generation model, the text processing result is the predicted word-unit index corresponding to the text to be processed.

[0159] Among them, the large language model is obtained according to the above-mentioned large language model fine-tuning method provided in the present disclosure.

[0160] In S304, the text processing result is received.

[0161] In the present disclosure, when the large language model is a text classification model, the above-mentioned text processing method can be used to analyze the sentiment category of the text to be processed, for example, dividing the text to be processed into positive, negative or neutral sentiment categories, and can also be used for news classification. At this time, the text to be processed is news text. The text processing method can be used to divide the news text into categories such as sports, entertainment, and technology. It can also be used for spam detection, for example, dividing emails into normal emails and spam.

[0162] When the large language model is a text generation model, the above text processing method can be used in scenarios such as text translation and summary generation.

[0163] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0164] When the large language model is a text generation model, the text processing result is a predicted word element index corresponding to the text to be processed; in this case, the text processing method may further include the following steps:

[0165] According to the target word list, the text corresponding to the predicted word element index is determined to obtain the predicted text corresponding to the text to be processed.

[0166] like Figure 3 As shown, after obtaining the predicted word element index corresponding to the text to be processed, the server sends it to the server user; after receiving the predicted word element index, the server uses the target word list to restore it to the predicted text.

[0167] Figure 7 is a block diagram of a large language model fine-tuning device applied to a server user according to an exemplary embodiment. Figure 7 As shown, the large language model fine-tuning device 400 applied to the server user includes: a first acquisition module 401, which is used to obtain the vocabulary and embedding layer parameters of the large language model from the server, wherein the pre-trained large language model is deployed on the server; an obfuscation module 402, which is used to perform obfuscation processing on the vocabulary and the embedding layer parameters respectively to obtain a target vocabulary and a target embedding layer parameter; a first sending module 403, which sends the target embedding layer parameter to the server, so that the server updates the embedding layer parameter of the large language model to the target embedding layer parameter to obtain a new large language model; a first word segmentation module 404, which is used to use the target vocabulary to perform word segmentation and index conversion operations on a text sample to obtain a first word element index corresponding to the text sample; the first sending module 403 is also used to send the first word element index to the server, so that the server fine-tunes the new large language model based on the first word element index.

[0168] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0169] Optionally, the obfuscation module 402 includes: a generation submodule for generating random permutations; a substitution submodule for performing a substitution operation on the vocabulary using the random permutations to obtain a target vocabulary, and performing a substitution operation on the embedding layer parameters using the random permutations to obtain substitution embedding layer parameters; and a first noise perturbation submodule for performing noise perturbation processing on the substitution embedding layer parameters according to the target vocabulary to obtain target embedding layer parameters.

[0170] Optionally, the first noise perturbation submodule includes: a clustering submodule, used to cluster the words in the target vocabulary according to the replacement embedding layer parameters; and a second noise perturbation submodule, used to perform noise perturbation processing on the replacement embedding layer parameters according to the clustering results to obtain target embedding layer parameters.

[0171] Optionally, the second noise perturbation submodule includes: a first determination submodule, used to determine, for each cluster in the clustering result, an embedding vector set corresponding to the cluster cluster from the permutation embedding layer parameters; a third noise perturbation submodule, used to perform noise perturbation processing on the embedding vector set using a differential privacy mechanism to obtain a perturbation vector set corresponding to the cluster cluster; and a first integration submodule, used to integrate the perturbation vector set corresponding to each cluster cluster to obtain the target embedding layer parameters.

[0172] Optionally, the third noise perturbation submodule includes: a first calculation submodule, used to calculate, for each embedded vector in the embedded vector set, a first similarity between the embedded vector and each embedded vector in the embedded vector set; a mapping submodule, used to map all the first similarities corresponding to the embedded vector to a first probability distribution according to a preset privacy budget, wherein the preset privacy budget is used to control the smoothness of the first probability distribution; a noise addition submodule, used to add Laplace noise to the first probability distribution according to the preset privacy budget to obtain a second probability distribution; a normalization submodule, used to normalize the second probability distribution to obtain a third probability distribution; a second determination submodule, used to determine the perturbation vector corresponding to the embedded vector according to the third probability distribution and the embedded vector set; a second integration submodule, used to integrate the perturbation vector corresponding to each of the embedded vectors to obtain a perturbation vector set corresponding to the cluster cluster.

[0173] Optionally, the clustering submodule includes: a third determination submodule, used to take any word in the target vocabulary as the current word; a second calculation submodule, used to calculate the second similarity between the current word and each other non-clustered word in the target vocabulary according to the replacement embedding layer parameters; a fourth determination submodule, used to take the current word and K-1 other non-clustered words with the highest second similarity to the current word as a cluster cluster; a fifth determination submodule, used to take any word in the non-clustered word of the target vocabulary as the current word if the number of cluster clusters does not reach N, wherein, n is the number of words in the target vocabulary; a triggering submodule is used to trigger the second calculation submodule to calculate the second similarity between the current word and each other non-clustered word in the target vocabulary according to the replacement embedding layer parameters, until the number of clusters reaches N; a sixth determination submodule is used to take N clusters as the clustering results if the number of clusters reaches N.

[0174] Optionally, the second calculation submodule includes: a third calculation submodule, used to calculate the semantic similarity and edit distance between each other word-unit that is not clustered in the target vocabulary and the current word-unit according to the replacement embedding layer parameters; and a seventh determination submodule, used to determine the second similarity between the other word-unit and the current word-unit based on the semantic similarity and the edit distance.

[0175] Optionally, the large language model is a text classification model; the large language model fine-tuning device 400 also includes: a third acquisition module, used to obtain a classification label corresponding to the text sample; the first sending module 403 is used to send the first word element index and the classification label to the server, so that the server can fine-tune the new large language model based on the first word element index and the classification label.

[0176] Optionally, the large language model is a text generation model, wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

[0177] Figure 8 is a block diagram of a large language model fine-tuning device applied to a server according to an exemplary embodiment. Figure 8 As shown, the large language model fine-tuning device 500 applied to the server includes: an updating module 501, which is used to update the embedding layer parameters of the large language model of the server to the target embedding layer parameters in response to receiving the target embedding parameters sent by the server user, so as to obtain a new large language model, wherein the target embedding layer parameters are obtained by the server user performing obfuscation processing on the embedding layer parameters; a fine-tuning module 502, which is used to fine-tune the new large language model based on the first word element index in response to receiving the first word element index sent by the server user, wherein the first word element index is generated by the server user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server user performing obfuscation processing on the vocabulary of the large language model.

[0178] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0179] Optionally, the large language model is a text classification model; the fine-tuning module 502 is used to fine-tune the new large language model based on the first word index and the classification label in response to receiving the first word index sent by the server user and the classification label corresponding to the text sample.

[0180] Optionally, the large language model is a text generation model, wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

[0181] Fig. 9 is a block diagram of a text processing device according to an exemplary embodiment, wherein the text processing device 600 is applied to a server user. Fig. 9As shown, the text processing device 600 includes: a second acquisition module 601, used to acquire a text to be processed; a second word segmentation module 602, used to perform word segmentation and index conversion operations on the text to be processed using a target vocabulary, and obtain a second word element index corresponding to the text to be processed; a second sending module 603, used to send the second word element index to the server, so that the server generates a text processing result of the text to be processed based on the second word element index through a pre-fine-tuned large language model, and sends the text processing result to the server user, wherein the large language model is obtained according to the above-mentioned large language model fine-tuning method provided in the present disclosure; a receiving module 604, used to receive the text processing result.

[0182] In the above technical solution, before using the text sample of the server user to fine-tune the large language model pre-trained by the server, the server user first obtains the vocabulary and embedding layer parameters of the large language model by interacting with the server; then, the server user performs obfuscation processing on the vocabulary and embedding layer parameters of the large language model respectively to obtain the target vocabulary and target embedding layer parameters, and sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; the server user uses the target vocabulary to perform word segmentation and index conversion operations on the text sample to obtain the first word unit index corresponding to the text sample; next, the server user sends the first word unit index to the server, so that the server fine-tunes the new large language model based on the first word unit index. Among them, the vocabulary and embedding layer parameters in the pre-trained large language model are obfuscated separately, and only the obfuscated embedding layer parameters (i.e., the target embedding layer parameters) are sent to the server, while the obfuscated vocabulary (i.e., the target vocabulary) is retained by the server user. In this way, the server can avoid obtaining the correspondence between the embedding vector and the word unit, so that the model effect can be better guaranteed through fine-tuning operations under the premise of effectively protecting the data of the server user. In addition, the target vocabulary is retained by the server user. During the model fine-tuning stage and the inference stage, the server user performs text segmentation based on the target vocabulary, so that the server does not directly contact the private data of the server user, and can only obtain the index corresponding to the text word unit, which can better protect the data of the server user. In addition, the server-side user only needs to bear the one-time embedding perturbation overhead (i.e., obfuscating the vocabulary and embedding layer parameters of the large language model) and the text segmentation overhead in the fine-tuning and inference stages. In this way, the server-side user does not need to occupy a large amount of CPU, memory, or battery power and other resources when performing text processing tasks, making text processing tasks easier to run on different types of server-side users, allowing more users to enjoy the convenience of this technology.

[0183] Optionally, the large language model is a text classification model, and the text processing result is the category of the text to be processed.

[0184] Optionally, the large language model is a text generation model, and the text processing result is a predicted word element index corresponding to the text to be processed; the text processing device 600 also includes: a determination module, which is used to determine the text corresponding to the predicted word element index according to the target vocabulary, and obtain the predicted text corresponding to the text to be processed.

[0185] The present disclosure also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the large language model fine-tuning method provided by the present disclosure or the steps of the text processing method provided by the present disclosure.

[0186] The present disclosure also provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of the above-mentioned large language model fine-tuning method provided by the present disclosure or the steps of the above-mentioned text processing method provided by the present disclosure.

[0187] The present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the large language model fine-tuning method provided by the present disclosure or the steps of the text processing method provided by the present disclosure.

[0188] Reference below Fig.10 , which shows a schematic diagram of the structure of an electronic device (such as a terminal device or a server) 700 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0189] like Fig.10As shown, the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0190] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although Fig.10 The electronic device 700 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0191] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 709, or installed from a storage device 708, or installed from a ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0192] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0193] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0194] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0195] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the vocabulary and embedding layer parameters of the large language model from the server, wherein the pre-trained large language model is deployed on the server; performs obfuscation processing on the vocabulary and the embedding layer parameters respectively to obtain a target vocabulary and target embedding layer parameters; sends the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; uses the target vocabulary to perform word segmentation and index conversion operations on a text sample to obtain a first word element index corresponding to the text sample; sends the first word element index to the server, so that the server fine-tunes the new large language model based on the first word element index.

[0196] Alternatively, the computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: in response to receiving a target embedding parameter sent by a server-side user, updates the embedding layer parameters of the server-side large language model to the target embedding layer parameters to obtain a new large language model, wherein the target embedding layer parameters are obtained by the server-side user performing obfuscation processing on the embedding layer parameters; in response to receiving a first word element index sent by the server-side user, fine-tunes the new large language model based on the first word element index, wherein the first word element index is generated by the server-side user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server-side user performing obfuscation processing on the vocabulary of the large language model.

[0197] Alternatively, the computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the text to be processed; performs word segmentation and index conversion operations on the text to be processed using the target vocabulary to obtain a second word element index corresponding to the text to be processed; sends the second word element index to the server, so that the server generates a text processing result of the text to be processed based on the second word element index through a pre-fine-tuned large language model, and sends the text processing result to the server user, wherein the large language model is obtained by the large language model fine-tuning method provided in the present disclosure; and receives the text processing result.

[0198] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0199] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0200] The modules described in the embodiments of the present disclosure may be implemented by software or hardware. The name of a module does not limit the module itself in some cases. For example, the first acquisition module may also be described as a "module for acquiring the vocabulary and embedding layer parameters of a large language model from a server."

[0201] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0202] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0203] According to one or more embodiments of the present disclosure, Example 1 provides a large language model fine-tuning method, including: obtaining a vocabulary and embedding layer parameters of a large language model from a server, wherein the pre-trained large language model is deployed on the server; performing obfuscation processing on the vocabulary and the embedding layer parameters respectively to obtain a target vocabulary and a target embedding layer parameters; sending the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; using the target vocabulary, performing word segmentation and indexing operations on a text sample to obtain a first word element index corresponding to the text sample; sending the first word element index to the server, so that the server fine-tunes the new large language model based on the first word element index.

[0204] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the vocabulary and the embedding layer parameters are respectively obfuscated to obtain a target vocabulary and a target embedding layer parameters, including: generating a random permutation; using the random permutation to perform a permutation operation on the vocabulary to obtain a target vocabulary, and using the random permutation to perform a permutation operation on the embedding layer parameters to obtain permuted embedding layer parameters; according to the target vocabulary, performing noise perturbation processing on the permuted embedding layer parameters to obtain target embedding layer parameters.

[0205] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein the replacement embedding layer parameters are subjected to noise perturbation processing according to the target vocabulary to obtain target embedding layer parameters, including: clustering the word elements in the target vocabulary according to the replacement embedding layer parameters; and based on the clustering results, performing noise perturbation processing on the replacement embedding layer parameters to obtain target embedding layer parameters.

[0206] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, wherein, based on the clustering results, the replacement embedding layer parameters are subjected to noise perturbation processing to obtain target embedding layer parameters, including: for each cluster cluster in the clustering results, determining an embedding vector set corresponding to the cluster cluster from the replacement embedding layer parameters; using a differential privacy mechanism to perform noise perturbation processing on the embedding vector set to obtain a perturbation vector set corresponding to the cluster cluster; and integrating the perturbation vector sets corresponding to each of the cluster clusters to obtain the target embedding layer parameters.

[0207] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 4, which uses a differential privacy mechanism to perform noise perturbation processing on the embedding vector set to obtain a perturbation vector set corresponding to the cluster cluster, including: for each embedding vector in the embedding vector set, calculating the first similarity between the embedding vector and each embedding vector in the embedding vector set; according to a preset privacy budget, mapping all the first similarities corresponding to the embedding vector to a first probability distribution, wherein the preset privacy budget is used to control the smoothness of the first probability distribution; according to the preset privacy budget, adding Laplace noise to the first probability distribution to obtain a second probability distribution; normalizing the second probability distribution to obtain a third probability distribution; according to the third probability distribution and the embedding vector set, determining the perturbation vector corresponding to the embedding vector; integrating the perturbation vectors corresponding to each of the embedding vectors to obtain the perturbation vector set corresponding to the cluster cluster.

[0208] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 3, wherein the word-grams in the target vocabulary are clustered according to the replacement embedding layer parameters, including: taking any word-gram in the target vocabulary as the current word-gram; calculating the second similarity between the current word-gram and each other non-clustered word-gram in the target vocabulary according to the replacement embedding layer parameters; taking the current word-gram and the K-1 other non-clustered word-grams with the highest second similarity to the current word-gram as a cluster; if the number of cluster clusters does not reach N, taking any word-gram in the non-clustered word-grams of the target vocabulary as the current word-gram, wherein, n is the number of words in the target vocabulary; returning to the step of respectively calculating the second similarity between the current word and each other word that has not been clustered in the target vocabulary according to the replacement embedding layer parameters, until the number of clusters reaches N; if the number of clusters reaches N, then taking the N clusters as the clustering results.

[0209] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 6, which calculates the second similarity between the current word and each other non-clustered word in the target vocabulary according to the replacement embedding layer parameters, including: for each other non-clustered word in the target vocabulary, calculating the semantic similarity and edit distance between the other word and the current word according to the replacement embedding layer parameters; determining the second similarity between the other word and the current word according to the semantic similarity and the edit distance.

[0210] According to one or more embodiments of the present disclosure, Example 8 provides a method of any one of Examples 1-7, wherein the large language model is a text classification model; the method further includes: obtaining a classification label corresponding to the text sample; sending the first word element index to the server so that the server fine-tunes the new large language model based on the first word element index, including: sending the first word element index and the classification label to the server so that the server fine-tunes the new large language model based on the first word element index and the classification label.

[0211] According to one or more embodiments of the present disclosure, Example 9 provides the method of any one of Examples 1-7, wherein the large language model is a text generation model, and wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

[0212] According to one or more embodiments of the present disclosure, Example 10 provides a large language model fine-tuning method, which is applied to a server, comprising: in response to receiving a target embedding parameter sent by a server user, updating the embedding layer parameters of the large language model of the server to the target embedding layer parameters to obtain a new large language model, wherein the target embedding layer parameters are obtained by the server user performing obfuscation processing on the embedding layer parameters; in response to receiving a first word element index sent by the server user, fine-tuning the new large language model based on the first word element index, wherein the first word element index is generated by the server user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server user performing obfuscation processing on the vocabulary of the large language model.

[0213] According to one or more embodiments of the present disclosure, Example 11 provides the method of Example 10, wherein the large language model is a text classification model; in response to receiving the first word-meta index sent by the server-side user, fine-tuning the new large language model based on the first word-meta index, including: in response to receiving the first word-meta index sent by the server-side user and the classification label corresponding to the text sample, fine-tuning the new large language model based on the first word-meta index and the classification label.

[0214] According to one or more embodiments of the present disclosure, Example 12 provides the method of Example 10, wherein the large language model is a text generation model, and wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

[0215] According to one or more embodiments of the present disclosure, Example 13 provides a text processing method, which is applied to a server-side user, including: obtaining a text to be processed; performing word segmentation and indexing operations on the text to be processed using a target vocabulary to obtain a second word element index corresponding to the text to be processed; sending the second word element index to the server, so that the server generates a text processing result of the text to be processed based on the second word element index using a pre-fine-tuned large language model, and sending the text processing result to the server-side user, wherein the large language model is obtained according to the large language model fine-tuning method described in any one of Examples 1-12; receiving the text processing result.

[0216] According to one or more embodiments of the present disclosure, Example 14 provides the method of 13, wherein the large language model is a text classification model, and the text processing result is the category of the text to be processed.

[0217] According to one or more embodiments of the present disclosure, Example 15 provides the method of Example 13, wherein the large language model is a text generation model, and the text processing result is a predicted word element index corresponding to the text to be processed; the method also includes: determining the text corresponding to the predicted word element index according to the target vocabulary, and obtaining the predicted text corresponding to the text to be processed.

[0218] According to one or more embodiments of the present disclosure, Example 16 provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-15 when executed by a processing device.

[0219] According to one or more embodiments of the present disclosure, Example 17 provides an electronic device, comprising: a storage device on which a computer program is stored; and a processing device for executing the computer program in the storage device to implement the steps of any one of the methods described in Examples 1-15.

[0220] According to one or more embodiments of the present disclosure, Example 18 provides a computer program product, including a computer program, which implements the steps of any one of the methods of Examples 1-15 when executed by a processor.

[0221] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0222] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0223] Although the subject matter has been described in language specific to structural features and / or method logic actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims. Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment related to the method, and will not be elaborated here.

Claims

1. A large language model fine-tuning method, characterized in that: include: Obtaining a vocabulary and embedding layer parameters of a large language model from a server, wherein the pre-trained large language model is deployed on the server; Obfuscating the vocabulary and the embedding layer parameters respectively to obtain a target vocabulary and a target embedding layer parameter; Sending the target embedding layer parameters to the server, so that the server updates the embedding layer parameters of the large language model to the target embedding layer parameters to obtain a new large language model; Using the target vocabulary, performing word segmentation and index conversion operations on the text sample to obtain a first word element index corresponding to the text sample; The first word-unit index is sent to the server, so that the server fine-tunes the new large language model based on the first word-unit index.

2. The method according to claim 1, characterized in that The obfuscating the vocabulary and the embedding layer parameters to obtain a target vocabulary and a target embedding layer parameter comprises: Generate random permutations; Performing a permutation operation on the vocabulary using the random permutation to obtain a target vocabulary, and performing a permutation operation on the embedding layer parameters using the random permutation to obtain permuted embedding layer parameters; According to the target vocabulary, the replacement embedding layer parameters are subjected to noise perturbation processing to obtain target embedding layer parameters.

3. The method according to claim 2, characterized in that The step of performing noise perturbation processing on the replacement embedding layer parameters according to the target vocabulary to obtain target embedding layer parameters includes: Clustering the word-units in the target vocabulary according to the permutation embedding layer parameters; According to the clustering result, the replacement embedding layer parameters are subjected to noise perturbation processing to obtain target embedding layer parameters.

4. The method according to claim 3, characterized in that The step of performing noise perturbation processing on the replacement embedding layer parameters according to the clustering result to obtain target embedding layer parameters includes: For each cluster in the clustering result, determine an embedding vector set corresponding to the cluster from the permutation embedding layer parameters; perform noise perturbation processing on the embedding vector set using a differential privacy mechanism to obtain a perturbation vector set corresponding to the cluster; The perturbation vector set corresponding to each of the clusters is integrated to obtain the target embedding layer parameters.

5. The method according to claim 4, characterized in that The method of performing noise perturbation processing on the embedding vector set by using the differential privacy mechanism to obtain a perturbation vector set corresponding to the cluster cluster includes: For each embedding vector in the embedding vector set, calculate a first similarity between the embedding vector and each embedding vector in the embedding vector set; according to a preset privacy budget, map all the first similarities corresponding to the embedding vector to a first probability distribution, wherein the preset privacy budget is used to control the smoothness of the first probability distribution; according to the preset privacy budget, add Laplace noise to the first probability distribution to obtain a second probability distribution; normalize the second probability distribution to obtain a third probability distribution; determine a perturbation vector corresponding to the embedding vector according to the third probability distribution and the embedding vector set; The perturbation vector corresponding to each of the embedding vectors is integrated to obtain a perturbation vector set corresponding to the cluster.

6. The method according to claim 3, characterized in that The clustering of word elements in the target vocabulary according to the replacement embedding layer parameters comprises: Taking any word in the target vocabulary as the current word; Calculating, according to the permuted embedding layer parameters, a second similarity between the current word and each other word in the target vocabulary that is not clustered; The current word-element and K-1 other words-element that have not been clustered and have the second highest similarity to the current word-element are grouped as a cluster; If the number of clusters does not reach N, any word in the target vocabulary that has not been clustered is used as the current word, where N= , n is the number of word-grams in the target vocabulary; Returning to the step of respectively calculating the second similarity between the current word and each other word in the target vocabulary that is not clustered according to the replacement embedding layer parameters, until the number of the clusters reaches N; If the number of the clusters reaches N, the N clusters are used as the clustering results.

7. The method according to claim 6, characterized in that The step of calculating the second similarity between the current word and each other word in the target vocabulary that is not clustered according to the replacement embedding layer parameters comprises: For each other word-gram that is not clustered in the target vocabulary, calculating the semantic similarity and edit distance between the other word-gram and the current word-gram according to the permutation embedding layer parameters; A second similarity between the other word-gram and the current word-gram is determined according to the semantic similarity and the edit distance.

8. The method according to any one of claims 1 to 7, characterized in that The large language model is a text classification model; The method further comprises: Obtaining a classification label corresponding to the text sample; The sending the first word-unit index to the server so that the server fine-tunes the new large language model based on the first word-unit index includes: The first word-unit index and the classification label are sent to the server, so that the server fine-tunes the new large language model based on the first word-unit index and the classification label.

9. The method according to any one of claims 1 to 7, characterized in that: The large language model is a text generation model, wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

10. A large language model fine-tuning method, characterized in that: Applied to the server, including: In response to receiving the target embedding layer parameters sent by the server user, updating the embedding layer parameters of the large language model of the server to the target embedding layer parameters to obtain a new large language model, wherein the target embedding layer parameters are obtained by the server user performing obfuscation processing on the embedding layer parameters; In response to receiving a first word-meta index sent by the server-side user, fine-tuning the new large language model based on the first word-meta index, wherein the first word-meta index is generated by the server-side user based on a target vocabulary and a text sample, and the target vocabulary is obtained by the server-side user performing obfuscation processing on the vocabulary of the large language model.

11. The method according to claim 10, characterized in that The large language model is a text classification model; In response to receiving the first word-unit index sent by the server user, fine-tuning the new large language model based on the first word-unit index includes: In response to receiving the first word-unit index sent by the server user and the classification label corresponding to the text sample, the new large language model is fine-tuned based on the first word-unit index and the classification label.

12. The method according to claim 10, characterized in that The large language model is a text generation model, wherein the embedding layer parameters include input embedding layer parameters and output embedding layer parameters.

13. A text processing method, characterized in that: Applicable to the server user, including: Get the text to be processed; Using the target vocabulary to perform word segmentation and index conversion operations on the text to be processed, to obtain a second word element index corresponding to the text to be processed; The second word-unit index is sent to the server, so that the server generates a text processing result of the to-be-processed text based on the second word-unit index by using a pre-fine-tuned large language model, and sends the text processing result to the server user, wherein the large language model is obtained according to the large language model fine-tuning method according to any one of claims 1 to 12; The text processing result is received.

14. The method according to claim 13, characterized in that The large language model is a text classification model, and the text processing result is the category of the text to be processed.

15. The method according to claim 13, characterized in that The large language model is a text generation model, and the text processing result is a predicted word element index corresponding to the text to be processed; The method further comprises: According to the target word list, the text corresponding to the predicted word element index is determined to obtain the predicted text corresponding to the text to be processed.

16. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 15 are implemented.

17. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 15.

18. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Model federation fine tuning method and device, text classification method and device, medium and equipment

    CN117744145A

  • Large model encryption method and device, equipment and storage medium

    CN117768219A