Word vector-based large language model input disturbance method, medium and system
Through the large language model input perturbation method based on word vectors, the input data of the large language model is fine-tuned and perturbed, which solves the problem that privacy protection affects the accuracy of model output, and realizes effective privacy protection and efficient model output.
Patent Information
- Application Number
- CN202510686093.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
When privacy protection of input data in large language models, adding a lot of noise will affect the accuracy of the output results and reduce work efficiency.
The large language model input perturbation method based on word vectors is adopted to determine the word vector perturbation range by obtaining the text data to be processed, preprocessing to obtain a set of sensitive words, word vector representation based on word vector model, fine-tuning the word vector matrix, comparing the reference output and fine-tuning the output to determine the word vector perturbation range, and perturbing the input data according to this range.
Effectively protect user privacy, while reducing the impact of privacy protection on the output results of large language models, improving the working efficiency of the model and the accuracy of the output results.
Smart Images

Figure CN120197715A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of large language model applications, and in particular to a large language model input perturbation method, medium and system based on word vectors. Background Art
[0002] With the rapid development of big data and artificial intelligence technology, the importance of data privacy protection has become increasingly prominent. Data encryption, differential privacy and other methods, as key technologies to protect data privacy, have been widely used in various fields.
[0003] In the related art, in the process of protecting the privacy of large language model inputs, a large amount of noise is often added to meet strict privacy protection requirements; and large language models usually need to parse and understand the input data to mine the features and patterns in the data for learning and reasoning. The addition of a large amount of noise makes it difficult for these inputs to be directly processed by the large language model, which greatly affects the accuracy of the large language model's output results and reduces the working efficiency of the large language model. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems in the related art to a certain extent. To this end, one purpose of the present invention is to propose a large language model input perturbation method based on word vectors, which can effectively protect the privacy of users and reduce the impact of privacy protection on the output results of the large language model.
[0005] In a first aspect, an embodiment of the present invention proposes a large language model input perturbation method based on word vectors, comprising the following steps: obtaining text data to be processed, and preprocessing the text data to be processed to obtain a set of sensitive words corresponding to the text data to be processed; representing the text data to be processed by word vectors based on a word vector model to obtain a benchmark word vector matrix corresponding to the text data to be processed, and inputting the benchmark word vector matrix into a large language model to obtain a corresponding benchmark output; fine-tuning the text data to be processed based on the sensitive word set to generate a fine-tuning word vector matrix, and inputting the fine-tuning word vector matrix into a large language model to obtain a corresponding fine-tuning output; comparing the benchmark output and the fine-tuning output to determine a word vector perturbation range; perturbing the text data to be processed according to the word vector perturbation range to obtain a final large language model input.
[0006] According to the method for perturbing the input of a large language model based on word vectors according to an embodiment of the present invention, first, obtain the text data to be processed, and preprocess the text data to be processed to obtain a set of sensitive words corresponding to the text data to be processed; then, perform word vector representation on the text data to be processed based on a word vector model to obtain a benchmark word vector matrix corresponding to the text data to be processed, and input the benchmark word vector matrix into the large language model to obtain a corresponding benchmark output; then, fine-tune the text data to be processed based on the set of sensitive words to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into the large language model to obtain a corresponding fine-tuned output; then, compare the benchmark output and the fine-tuned output to determine the word vector perturbation range; then, perturb the text data to be processed according to the word vector perturbation range to obtain the final input of the large language model, thereby effectively protecting the privacy of the user. At the same time, the impact of privacy protection on the output result of the large language model is reduced.
[0007] In some embodiments, comparing the benchmark output and the fine-tuned output to determine the word vector perturbation range includes: calculating the difference value between the benchmark output and the fine-tuned output; determining whether the difference value is greater than a preset difference threshold; if so, return to the step of fine-tuning the text data to be processed based on the set of sensitive words; if not, use the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determine whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; if the number of compliance adjustment values is greater than the preset number threshold, end the fine-tuning operation for the current fine-tuning vector, and return to reselect the current fine-tuning vector based on the set of sensitive words.
[0008] In some embodiments, after comparing the benchmark output and the fine-tuned output to determine the word vector perturbation range, it further includes: calculating the cosine similarity between each word vector in the word vector perturbation range and the corresponding benchmark word vector in the benchmark word vector matrix; determining whether the cosine similarity is greater than a preset similarity threshold; if so, re-fine-tune the word vector.
[0009] In some embodiments, preprocessing the text data to be processed includes: constructing a conditional random field model to identify the named entities corresponding to the text data to be processed through the conditional random field model, and obtaining a set of sensitive words corresponding to the text data to be processed based on the named entities; wherein, the conditional random field model determines the label corresponding to each word segment in the text data to be processed by maximizing the objective function; The objective function is expressed by the following formula: ; wherein, represents the objective function, represents the conditional probability, represents the th word segment, represents the th label of the word segment, represents the model parameters.
[0010] In some embodiments, perturbing the to-be-processed text data according to the word vector perturbation range includes: Selecting a Gaussian function as the perturbation function to perturb the to-be-processed text data according to the word vector perturbation range; wherein, the Gaussian function is expressed by the following formula: ; where represents the Gaussian function, represents the input vector, represents the mean vector, represents the standard deviation, represents the vector dimension.
[0011] In some embodiments, perturbing the to-be-processed text data according to the word vector perturbation range includes: For each word vector within the word vector perturbation range, calculating the sensitivity value and the text importance value corresponding to the word vector; calculating the corresponding perturbation intensity value according to the sensitivity value and the text importance value to determine the perturbation intensity corresponding to each word vector.
[0012] In some embodiments, the method further includes: Inputting the to-be-processed text data into a large language model so that the large language model outputs the original output corresponding to the to-be-processed text data, and obtaining the corresponding first model performance; Inputting the final large language model input into the large language model so that the large language model outputs the perturbed output corresponding to the final large language model input, and obtaining the corresponding second model performance; Comparing the original output and the perturbed output to calculate the effective protection value corresponding to the perturbation operation; Comparing the first model performance and the second model performance to calculate the model performance impact value corresponding to the perturbation operation; Evaluating the perturbation operation according to the effective protection value and the model performance impact value.
[0013] In a second aspect, an embodiment of the present invention proposes a computer-readable storage medium, on which a large language model input perturbation program based on word vectors is stored. When the large language model input perturbation program based on word vectors is executed by a processor, the above-mentioned large language model input perturbation method based on word vectors is implemented.
[0014] In a third aspect, an embodiment of the present invention provides a large language model input perturbation system based on word vectors, including: a preprocessing module, which is configured to obtain text data to be processed and preprocess the text data to be processed to obtain a set of sensitive words corresponding to the text data to be processed; a word vector representation module, which is configured to perform word vector representation on the text data to be processed based on a word vector model to obtain a reference word vector matrix corresponding to the text data to be processed, and input the reference word vector matrix into a large language model to obtain a corresponding reference output; a fine-tuning module, which is configured to fine-tune the text data to be processed based on the set of sensitive words to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into the large language model to obtain a corresponding fine-tuned output; a comparison module, which is configured to compare the reference output and the fine-tuned output to determine a word vector perturbation range; and a perturbation module, which is configured to perturb the text data to be processed according to the word vector perturbation range to obtain a final large language model input.
[0015] In some embodiments, comparing the reference word vector matrix and the fine-tuned word vector matrix to determine the word vector perturbation range includes: calculating a difference value between the reference word vector matrix and the fine-tuned word vector matrix; determining whether the difference value is greater than a preset difference threshold; if so, returning to the step of fine-tuning the text data to be processed based on the set of sensitive words; if not, taking the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determining whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; if the number of compliance adjustment values is greater than the preset number threshold, ending the fine-tuning operation for the current fine-tuning vector, and returning to reselect the current fine-tuning vector based on the set of sensitive words.
[0016] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of a method for perturbing the input of a large language model based on word vectors according to an embodiment of the present invention; Figure 2 is a block diagram of a large language model input perturbation system based on word vectors according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0019] The following describes a method for perturbing the input of a large language model based on word vectors according to an embodiment of the present invention with reference to the accompanying drawings.
[0020] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for perturbing the input of a large language model based on word vectors according to an embodiment of the present invention. As shown in Figure 1 , the method for perturbing the input of a large language model based on word vectors includes the following steps: S101: Obtain the text data to be processed, and preprocess the text data to be processed to obtain a set of sensitive words corresponding to the text data to be processed.
[0021] In some embodiments, preprocessing the text data to be processed includes: Construct a conditional random field model to identify the named entities corresponding to the text data to be processed through the conditional random field model, and obtain a set of sensitive words corresponding to the text data to be processed based on the named entities; wherein, the conditional random field model determines the label corresponding to each word segment in the text data to be processed by maximizing the objective function; The objective function is expressed by the following formula: ; wherein, represents the objective function, represents the conditional probability, represents the label of the th word segment, represents the th word segment, represents the model parameters.
[0022] As an example, first, parse the text data to be processed, and use a named entity recognition algorithm to identify the named entities in the text data to be processed by constructing a conditional random field model. Assume that the text data to be processed consists of a series of word segments . The conditional random field model determines the label corresponding to each word segment in the text data to be processed by maximizing the following objective function: ; wherein, represents the objective function, represents the conditional probability (that is, given the entire sequence of text data to be processed and model parameters Under the condition that the th token is the label indicating the label of the th token, which is used to indicate whether the token is a sensitive word; specifically, it can take values according to the model setting. For example, 1 can represent a sensitive word and 0 can represent a non - sensitive word), indicating the th token indicating the model parameters. At the same time, combining part - of - speech tagging technology, according to the pre - trained part - of - speech tagging model, each token is tagged with a part of speech. By integrating the results of named entity recognition algorithms and POS, a set of words that may contain sensitive information in the text data to be processed is identified, that is, the sensitive word set.
[0023] S102, Represent the text data to be processed based on the word vector model to obtain the corresponding benchmark word vector matrix of the text data to be processed, and input the benchmark word vector matrix into the large - language model to obtain the corresponding benchmark output.
[0024] As an example, first, according to the specific application scenario and the characteristics of the large - language model, select a suitable word vector model. For example, Word2Vec, whose Skip - Gram model learns word vectors by maximizing the following objective function: ; where is the center word, is the context word within the window, is 's context word set, is the model parameter. Set the window size to (when processing each center word , the words centered on with words on each side as context words will be considered. When is odd, one side has one more word than the other side), the vector dimension is , and the text is trained to obtain the word vector corresponding to each center word .
[0025] S103, Fine - tune the text data to be processed based on the sensitive word set to generate a fine - tuned word vector matrix, and input the fine - tuned word vector matrix into the large - language model to obtain the corresponding fine - tuned output.
[0026] S104, Compare the benchmark output and the fine - tuned output to determine the word vector perturbation range.
[0027] In some embodiments, comparing the baseline output and the fine-tuned output to determine the word vector perturbation range includes: calculating the difference value between the baseline output and the fine-tuned output; determining whether the difference value is greater than a preset difference threshold; if so, returning to the step of fine-tuning the text data to be processed based on the sensitive word set; if not, taking the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determining whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; if the number of compliance adjustment values is greater than the preset number threshold, ending the fine-tuning operation for the current fine-tuning vector, and returning to reselect the current fine-tuning vector based on the sensitive word set.
[0028] In some embodiments, after comparing the baseline output and the fine-tuned output to determine the word vector perturbation range, it further includes: calculating the cosine similarity between each word vector in the word vector perturbation range and the corresponding baseline word vector in the baseline word vector matrix; determining whether the cosine similarity is greater than a preset similarity threshold; if so, re-fine-tuning the word vector.
[0029] As an example, first, input the baseline word vector matrix corresponding to the text data to be processed into a large language model to obtain the corresponding baseline output; then, select a word vector as the current fine-tuning vector based on the sensitive word set, and fine-tune the current fine-tuning vector to generate a fine-tuned word vector matrix; then, input the fine-tuned word vector matrix into the large language model to obtain the corresponding fine-tuned output; then, calculate the difference value between the baseline output and the fine-tuned output; then, determine whether the difference value is greater than a preset difference threshold, if so, return to the step of fine-tuning the current fine-tuning vector; if not, take the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determine whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; if the number of compliance adjustment values is greater than the preset number threshold, it means that the perturbation range corresponding to the current fine-tuning vector has been determined, end the fine-tuning operation for the current fine-tuning vector, and return to the step of selecting a word vector as the current fine-tuning vector based on the sensitive word set until each word vector in the sensitive word set has been traversed.
[0030] As another example, in order to further improve the accuracy of determining the word vector perturbation range, after comparing the baseline output and the fine-tuned output to determine the word vector perturbation range, the final word vector perturbation range is also determined by calculating the cosine similarity between word vectors; first, for any two word vectors, the cosine similarity is defined as: ; where represents the dot product of vector and vector , and Represents vectors and vector The norm of (here is the bi-norm, i.e. the length of the vector). Set a cosine similarity threshold , for sensitive word vectors , filter out Similarity is higher than A set of vectors to determine the final word vector perturbation range.
[0031] S105, perturbing the text data to be processed according to the word vector perturbation range to obtain a final large language model input.
[0032] In some embodiments, perturbing the text data to be processed according to the word vector perturbation range includes: According to the word vector perturbation range, a Gaussian function is selected as the perturbation function to perturb the text data to be processed; The Gaussian function is expressed by the following formula: ; in, represents the Gaussian function, represents the input vector, represents the mean vector, represents the standard deviation, Represents the vector dimension.
[0033] As an example, first, according to the word vector perturbation range, the Gaussian function is selected as the perturbation function. The Gaussian function is expressed by the following formula: ; in, represents the Gaussian function, represents the input vector, represents the mean vector, represents the standard deviation, Represents the vector dimension.
[0034] In this way, the Gaussian function parameters can be adjusted according to the word vector perturbation range and the privacy protection strength requirements. If stronger privacy protection is required, the Gaussian function parameters can be appropriately increased. If you want to minimize the impact on the data, you can reduce . Let the privacy protection strength parameter be , established through experiments and relationship, such as ,in is the function obtained through experimental fitting.
[0035] In some embodiments, perturbing the text data to be processed according to the word vector perturbation range includes: for each word vector within the word vector perturbation range, calculating the sensitivity value and text importance value corresponding to the word vector; calculating the corresponding perturbation intensity value according to the sensitivity value and text importance value to determine the perturbation intensity corresponding to each word vector.
[0036] That is to say, in order to improve the accuracy of perturbation, not all word vectors in the sensitive word set are perturbed to the same degree. Instead, different degrees of perturbation are performed according to sensitivity and text importance.
[0037] Specifically, for calculating the sensitivity value: construct a sensitivity evaluation system and assign corresponding weights according to the types of sensitive words. For example, sensitive words related to personal identity information such as ID numbers and names are assigned high weights; relatively less important sensitive words such as general industry terms containing business sensitive information are assigned low weights. The weight system is determined with reference to industry standards, laws and regulations, and actual application scenarios.
[0038] For calculating the text importance value: starting from the text semantic structure and logical relationship, with the help of natural language processing techniques (such as dependency syntax analysis), analyze the dependency relationship between sensitive words and other words to judge their importance in the text. If a sensitive word is in the core position of a sentence and plays a key role in semantic expression, its importance is high; otherwise, it is low. For example, in a text describing an event, the core participants (sensitive words) of the event are more important than the modifying sensitive words.
[0039] For calculating the perturbation intensity value: integrate the sensitivity value and text importance value, and calculate a comprehensive score for each sensitive word through weighted summation. For example: perturbation intensity value = sensitivity weight × sensitivity value + text importance weight × text importance value. Accordingly, set a threshold. Sensitive word vectors with perturbation intensity values greater than the threshold are used as key perturbation objects and a larger perturbation intensity is adopted; for those with perturbation intensity values less than or equal to the threshold, the perturbation intensity is reduced or even not perturbed, so as to balance privacy protection and the processing effect of large language models. Or, preset the perturbation intensity corresponding to a benchmark perturbation intensity value, and further determine the final perturbation intensity according to the ratio between the perturbation intensity value and the benchmark perturbation intensity value.
[0040] Next, select the Gaussian perturbation function to process the selected sensitive word vectors The Gaussian perturbation function can flexibly control the degree and range of perturbation by adjusting parameters to balance the requirements of privacy protection and semantic preservation. Substitute the selected sensitive word vectors into the Gaussian perturbation function to obtain the perturbed word vectors That is: ; Among them, is a random vector sampled from a Gaussian distribution Here, represents the zero vector, indicating that the center of the perturbation is around the original word vector; is the variance of the Gaussian distribution, which determines the magnitude of the perturbation. The larger the variance, the greater the magnitude of the perturbation; is the identity matrix with the same dimension as the word vector, ensuring that the randomly sampled vector has the same dimension as the word vector so that operations can be performed. For example, if the word vector is a -dimensional vector, then the identity matrix is also a -dimensional matrix, and the randomly sampled vector from the Gaussian distribution is also a -dimensional vector, thus ensuring the consistency and feasibility of mathematical operations.
[0041] Meanwhile, considering the semantic coherence of the text, if a sensitive word plays a key semantic connection role in the text, its perturbation degree should be carefully evaluated to avoid significant damage to the text semantics caused by excessive perturbation. At this time, the variance of the Gaussian distribution can be adjusted to perturb the sensitive word vectors at different positions in the text to different degrees, obtaining the desired perturbation results.
[0042] Finally, replace the word vectors corresponding to the sensitive words in the original text with the perturbed word vectors to form the perturbed text data. During the replacement process, strictly follow the structural rules of the text to ensure that the word order and part - of - speech collocations meet the grammatical requirements, and ensure that the structure and basic semantics of the text remain relatively intact after replacement for the subsequent normal processing of the large - language model.
[0043] As an example, before inputting the final large - language model into the large - language model, the perturbed text data to be processed is also formatted to meet the input requirements of the large - language model. For example, the text is tokenized to obtain the corresponding sub - sequences, and they are converted into the tensor form acceptable to the model, such as adding position encoding, segment encoding, etc. operations to obtain the input tensor.
[0044] Finally, input the prepared input tensor into the large - language model , the model performs calculations through a series of neural network layers. For example, a large language model composed of multiple Transformer blocks, each Transformer block contains a self-attention mechanism and a feed-forward neural network. Taking the self-attention mechanism as an example, its calculation process is as follows: ; Among them, , , are the query matrix, key matrix, and value matrix respectively, which together determine the focus of the model's attention to the input information and the direction of feature extraction. is the key matrix The dimension of is used for The calculation result of is scaled so that The input of the function is within an appropriate range, thus avoiding The problem of gradient disappearance or explosion during the calculation of the function. After the above calculation process, the model can generate a weighted output according to the input query, key, and value matrices. This output integrates feature information from different positions to help the model better capture long-range dependencies in the sequence. After layers of calculations by multiple Transformer blocks, the final inference result is output .
[0045] In some embodiments, the method further includes: inputting the text data to be processed into the large language model so that the large language model outputs the original output corresponding to the text data to be processed and obtaining the corresponding first model performance; inputting the final large language model input into the large language model so that the large language model outputs the perturbed output corresponding to the final large language model input and obtaining the corresponding second model performance; comparing the original output and the perturbed output to calculate the effective protection value corresponding to the perturbation operation; comparing the first model performance and the second model performance to calculate the model performance impact value corresponding to the perturbation operation; evaluating the perturbation operation according to the effective protection value and the model performance impact value.
[0046] That is to say, further analyze the impact of the perturbed data on the model performance and whether it effectively protects privacy, and adjust the perturbation process according to the analysis results.
[0047] Specifically, first, the text data to be processed and the perturbed text data to be processed are input into a privacy detection model . The privacy detection model evaluates the privacy risk by calculating the sensitive information leakage index , and measures it by calculating the entropy difference of sensitive information in the original text and the perturbed text: ; Among them, is the information entropy, which is used to measure the uncertainty or chaos degree of text information. For the text , the calculation formula of its information entropy is , where represents the word in the text , and the probability of each word in the text is obtained by summing and taking the negative value, and the information entropy of the text is obtained, so as to reflect the information chaos degree of the text , and then the sensitive information leakage index is obtained by taking the difference with .
[0048] Alternatively, it can be evaluated manually: that is, invite professionals to review the perturbed text to judge whether the sensitive information is effectively protected and whether there is a potential risk of privacy leakage.
[0049] Next, perform model performance evaluation: First, perform index selection: According to the application scenario of the large language model, select appropriate performance indicators. In the text generation task, the perplexity (PP) index can be used to evaluate the quality of the generated text, and the perplexity is defined as: ; Among them, represents the total number of words in the generated text, is the probability of generating the -th word. The lower the perplexity, the relatively better the quality of the text generated by the model. In the text classification task, indicators such as accuracy, recall, and F1 value are selected to evaluate the performance.
[0050] As a specific embodiment of the present invention, the large language model input perturbation method based on word vectors includes the following steps: S201, obtain the text data to be processed, and preprocess the text data to be processed to obtain the sensitive word set corresponding to the text data to be processed.
[0051] S202, perform word vector representation on the text data to be processed based on the word vector model to obtain the benchmark word vector matrix corresponding to the text data to be processed, and input the benchmark word vector matrix into the large language model to obtain the corresponding benchmark output.
[0052] S203, randomly select an unfinely tuned word vector from the sensitive word set as the current fine-tuning vector.
[0053] S204, randomly fine-tune the current fine-tuning vector to generate a fine-tuning word vector matrix.
[0054] S205, input the fine-tuning word vector matrix into the large language model to obtain the corresponding fine-tuning output.
[0055] S206, calculate the difference value between the baseline output and the fine-tuning output.
[0056] S207, determine whether the difference value is greater than the preset difference threshold; if yes, execute step S204; if no, execute step S208.
[0057] S208, take the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector.
[0058] S209, determine whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than the preset number threshold; if yes, execute step S210; if no, execute step S204.
[0059] S210, determine whether all the word vectors in the sensitive word set have been fine-tuned; if yes, execute step S211; if no, execute step S203.
[0060] S211, output the preselected word vector perturbation range.
[0061] S212, randomly select any word vector in the preselected word vector perturbation range as the current matching vector.
[0062] S213, calculate the cosine similarity between the current matching vector and the baseline word vector corresponding to this word vector in the baseline word vector matrix.
[0063] S214; determine whether the cosine similarity is greater than the preset similarity threshold; if yes, take the current matching vector as the current fine-tuning vector and return to step S204; if no, execute step S215.
[0064] S215, determine whether all the word vectors in the preselected word vector perturbation range have been traversed; if yes, execute step S216; if no, execute step S212.
[0065] S216, output the final word vector perturbation range.
[0066] S217, for each word vector in the final word vector perturbation range, calculate the sensitivity value and text importance value corresponding to the word vector.
[0067] S218. Calculate the corresponding perturbation intensity value according to the sensitivity value and the text importance value to determine the perturbation intensity corresponding to each word vector.
[0068] S219. Perturb the reference word vector based on the determined perturbation intensity to obtain the final input for the large language model.
[0069] In summary, according to the method for perturbing the input of a large language model based on word vectors in the embodiments of the present invention, first, obtain the text data to be processed, and preprocess the text data to be processed to obtain the set of sensitive words corresponding to the text data to be processed; then, perform word vector representation on the text data to be processed based on the word vector model to obtain the reference word vector matrix corresponding to the text data to be processed, and input the reference word vector matrix into the large language model to obtain the corresponding reference output; then, fine-tune the text data to be processed based on the set of sensitive words to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into the large language model to obtain the corresponding fine-tuned output; then, compare the reference output and the fine-tuned output to determine the word vector perturbation range; then, perturb the text data to be processed according to the word vector perturbation range to obtain the final input for the large language model, thereby effectively protecting the privacy of users. At the same time, the impact of privacy protection on the output results of the large language model is reduced.
[0070] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a program for perturbing the input of a large language model based on word vectors is stored. When the program for perturbing the input of a large language model based on word vectors is executed by a processor, the method for perturbing the input of a large language model based on word vectors as described above is implemented.
[0071] In a third aspect, an embodiment of the present invention provides a system for perturbing the input of a large language model based on word vectors, as Figure 2 shown. The system for perturbing the input of a large language model based on word vectors includes a preprocessing module 10, a word vector representation module 20, a fine-tuning module 30, a comparison module 40, and a perturbation module 50.
[0072] Among them, the preprocessing module 10 is used to obtain the text data to be processed, and preprocess the text data to be processed to obtain the set of sensitive words corresponding to the text data to be processed; The word vector representation module 20 is used to perform word vector representation on the text data to be processed based on the word vector model to obtain the reference word vector matrix corresponding to the text data to be processed, and input the reference word vector matrix into the large language model to obtain the corresponding reference output; The fine-tuning module 30 is used to fine-tune the to-be-processed text data based on the set of sensitive words to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into the large language model to obtain a corresponding fine-tuned output; The comparison module 40 is used to compare the baseline output and the fine-tuned output to determine the word vector perturbation range; The perturbation module 50 is used to perturb the to-be-processed text data according to the word vector perturbation range to obtain the final large language model input.
[0073] In some embodiments, comparing the baseline output and the fine-tuned output to determine the word vector perturbation range includes: calculating the difference value between the baseline output and the fine-tuned output; determining whether the difference value is greater than a preset difference threshold; if so, returning to the step of fine-tuning the to-be-processed text data based on the set of sensitive words; if not, taking the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determining whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; if the number of compliance adjustment values is greater than the preset number threshold, ending the fine-tuning operation for the current fine-tuning vector and returning to reselect the current fine-tuning vector based on the set of sensitive words.
[0074] It should be noted that the above description of the method for perturbing the input of the large language model based on word vectors also applies to this system for perturbing the input of the large language model based on word vectors, which will not be elaborated here.
[0075] Note that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0076] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0077] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0078] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation on the present invention.
[0079] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0080] In the present invention, unless otherwise clearly specified and limited, the terms "mounted", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0081] In the present invention, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.
[0082] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as a limitation on the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A large language model input perturbation method based on word vectors, characterized in that, It includes the following steps: Obtain the text data to be processed, and preprocess the text data to be processed to obtain a set of sensitive words corresponding to the text data to be processed; Perform word vector representation on the text data to be processed based on a word vector model to obtain a benchmark word vector matrix corresponding to the text data to be processed, and input the benchmark word vector matrix into a large language model to obtain a corresponding benchmark output; Fine-tune the text data to be processed based on the set of sensitive words to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into a large language model to obtain a corresponding fine-tuned output; Compare the benchmark output and the fine-tuned output to determine the word vector perturbation range; Perturb the text data to be processed according to the word vector perturbation range to obtain the final input for the large language model.
2. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein Comparing the benchmark output and the fine-tuned output to determine the word vector perturbation range includes: Calculate the difference value between the benchmark output and the fine-tuned output; Determine whether the difference value is greater than a preset difference threshold; If so, return to the step of fine-tuning the text data to be processed based on the set of sensitive words; If not, take the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and determine whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; If the number of compliance adjustment values is greater than the preset number threshold, end the fine-tuning operation for the current fine-tuning vector, and return to reselect the current fine-tuning vector based on the set of sensitive words.
3. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein After comparing the benchmark output and the fine-tuned output to determine the word vector perturbation range, it further includes: Calculate the cosine similarity between each word vector in the word vector perturbation range and the corresponding benchmark word vector in the benchmark word vector matrix; Determine whether the cosine similarity is greater than a preset similarity threshold; If so, re-fine-tune the word vector.
4. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein Preprocessing the text data to be processed includes: Construct a conditional random field model to identify the named entities corresponding to the text data to be processed through the conditional random field model, and obtain a set of sensitive words corresponding to the text data to be processed based on the named entities; Wherein, the conditional random field model determines the label corresponding to each word segment in the text data to be processed by maximizing the objective function; The objective function is expressed by the following formula: ; Among them, represents the objective function, represents the conditional probability, represents the label of the th word segmentation, represents the th word segmentation, represents the model parameter.
5. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein Perturbing the text data to be processed according to the word vector perturbation range includes: According to the word vector perturbation range, select a Gaussian function as the perturbation function to perturb the text data to be processed; Wherein, the Gaussian function is expressed by the following formula: ; Among them, represents the Gaussian function, represents the input vector, represents the mean vector, represents the standard deviation, represents the vector dimension.
6. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein Perturbing the text data to be processed according to the word vector perturbation range includes: For each word vector within the word vector perturbation range, calculate the sensitivity value and text importance value corresponding to the word vector; Calculate the corresponding perturbation intensity value according to the sensitivity value and the text importance value to determine the perturbation intensity corresponding to each word vector.
7. The method for perturbing the input of a large language model based on word vectors according to claim 1, wherein It also includes: Input the text data to be processed into a large language model so that the large language model outputs the original output corresponding to the text data to be processed, and obtain the corresponding first model performance; Input the input of the final large language model into the large language model so that the large language model outputs the perturbed output corresponding to the input of the final large language model, and obtain the corresponding second model performance; Compare the original output and the perturbed output to calculate the effective protection value corresponding to the perturbation operation; Compare the first model performance and the second model performance to calculate the model performance impact value corresponding to the perturbation operation; Evaluate the perturbation operation according to the effective protection value and the model performance impact value.
8. A computer-readable storage medium, characterized in that, Stored thereon is a large language model input perturbation program based on word vectors. When the large language model input perturbation program based on word vectors is executed by a processor, it implements the large language model input perturbation method based on word vectors as described in any one of claims 1-7.
9. A large language model input perturbation system based on word vectors, characterized in that, Comprising: A preprocessing module, which is used to obtain the text data to be processed and preprocess the text data to be processed to obtain a sensitive word set corresponding to the text data to be processed; A word vector representation module, which is used to perform word vector representation on the text data to be processed based on a word vector model to obtain a benchmark word vector matrix corresponding to the text data to be processed, and input the benchmark word vector matrix into a large language model to obtain a corresponding benchmark output; A fine-tuning module, which is used to fine-tune the text data to be processed based on the sensitive word set to generate a fine-tuned word vector matrix, and input the fine-tuned word vector matrix into a large language model to obtain a corresponding fine-tuned output; A comparison module, which is used to compare the benchmark output and the fine-tuned output to determine the word vector perturbation range; A perturbation module, which is used to perturb the text data to be processed according to the word vector perturbation range to obtain the input of the final large language model.
10. The large language model input perturbation system based on word vectors according to claim 9, characterized in that, Comparing the benchmark output and the fine-tuned output to determine the word vector perturbation range includes: Calculating the difference value between the benchmark output and the fine-tuned output; Judging whether the difference value is greater than a preset difference threshold; If so, return to the step of fine-tuning the text data to be processed based on the sensitive word set; If not, use the fine-tuning operation corresponding to the current fine-tuning vector as the compliance adjustment value of the current fine-tuning vector, and judge whether the number of compliance adjustment values corresponding to the current fine-tuning vector is greater than a preset number threshold; If the number of compliance adjustment values is greater than the preset number threshold, end the fine-tuning operation for the current fine-tuning vector, and return to reselect the current fine-tuning vector based on the sensitive word set.
Citation Information
Patent Citations
Large NLP language model privacy protection method based on differential privacy
CN116502263A
Personal data privacy protection and recovery method based on large language model service
CN118278040A
Method for generating desensitized text by utilizing diffusion model guided by bidirectional gradient
CN118468332A
Text classification model robustness detection method
CN118520105A
Methods of providing data privacy for neural network based inference
US20210390188A1