A privacy protection method based on a large model and related devices

By dividing a large language model into sub-models and utilizing privacy-preserving modules and compensation coefficients, the problem of performance loss in privacy protection of large models is solved, achieving a balance between privacy protection and performance.

CN119416254BActive Publication Date: 2026-05-29HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
Filing Date
2024-10-21
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing large models lack sufficient research on privacy protection, especially affecting performance on complex tasks, resulting in poor practicality and failing to effectively protect user privacy without impacting model performance.

Method used

The large language model is divided into a first large language sub-model and a second large language sub-model. Data processing is performed using a privacy protection module, and the privacy protection sub-model and compensation coefficients ensure that the server cannot obtain user data while maintaining model performance.

Benefits of technology

This approach achieves the goal of protecting user privacy while reducing performance loss in large models, thus ensuring both model availability and the effectiveness of privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416254B_ABST
    Figure CN119416254B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a privacy protection method based on a large model and related equipment, which are used to protect privacy while not affecting the usability of the large model. The method of the embodiments of the present application comprises: cutting a large language model into a first large language sub-model and a second large language sub-model with a first privacy protection sub-module in a privacy protection module as a cutting point; obtaining to-be-analyzed data input by a user in the large language model, inputting the to-be-analyzed data into the first large language sub-model to obtain intermediate output data; transmitting the intermediate output data to the second large language sub-model to obtain model output data corresponding to the second large language sub-model; compensating the model output data according to a target compensation coefficient to obtain hidden state data corresponding to the to-be-analyzed data, and sending the hidden state data to a server provided with the large language model to obtain analysis result data returned by the server after analyzing the hidden state data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model privacy protection, and more particularly to a privacy protection method and related equipment based on large models. Background Technology

[0002] Currently, large-scale models offer tremendous convenience to our daily work. For example, when accessing certain cloud services, we can upload articles for large-scale models to quickly summarize or upload text for translation. These large-scale model cloud service scenarios all require uploading data. However, this data may involve privacy concerns. Currently, existing research on large-scale models focuses very little on privacy protection, and these studies have only verified feasibility on very simple tasks. On more complex tasks (such as mathematical calculations and coding), these methods severely impact the original performance of the large-scale models, thus lacking practical applicability.

[0003] Therefore, how to minimize the performance loss of large models while ensuring privacy protection is a pressing technical challenge that needs to be addressed. Summary of the Invention

[0004] This application provides a privacy protection method and related device based on a large model, which can protect privacy without affecting the usability of the large model.

[0005] The first aspect of this application provides a privacy protection method based on a large language model, applied to a user terminal. The large language model runs on the user terminal, and the large language model includes a privacy protection module. The method includes:

[0006] Using the first privacy protection sub-module in the privacy protection module as the dividing point, the large language model is divided into a first large language sub-model and a second large language sub-model; wherein, the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking.

[0007] The system acquires the data to be analyzed input by the user into the large language model, inputs the data to be analyzed into the first large language sub-model to obtain intermediate output data, and transmits the intermediate output data to the second large language sub-model to obtain model output data corresponding to the second large language sub-model.

[0008] The model output data is compensated according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed, and the hidden state data is sent to the server equipped with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data.

[0009] Optionally, the step of dividing the large language model into a first large language sub-model and a second large language sub-model, using the first privacy protection sub-module in the privacy protection module as the dividing point, includes:

[0010] Obtain the ranking order of all privacy-preserving sub-modules in the large language model, and set the privacy-preserving sub-module with the target ranking order as the first privacy-preserving sub-module;

[0011] Set a preset number of privacy protection sub-modules in the first large language sub-model, and set the first privacy protection sub-module as the last privacy protection sub-module in the first large language sub-model;

[0012] The privacy protection submodules whose ranking order is after the first privacy protection submodule are set in the second major language submodel.

[0013] Optionally, the step of inputting the data to be analyzed into the first large language sub-model to obtain intermediate output data includes:

[0014] The data to be analyzed is input into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed; wherein, the second privacy protection submodule is the first privacy protection submodule in the first large language submodel;

[0015] The first vector data to be analyzed is subjected to privacy protection processing based on privacy parameters to obtain the first output data corresponding to the second privacy protection submodule;

[0016] The first output data is used as the input data of the next privacy protection submodule of the second privacy protection submodule, and according to the ranking order, until the first privacy protection submodule of the first large language submodel receives the input data of the third privacy protection submodule and outputs the intermediate output data; wherein, the third privacy protection submodule is the previous privacy protection submodule of the first privacy protection submodule.

[0017] Optionally, the step of inputting the data to be analyzed into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed includes:

[0018] The data to be analyzed is input into the second privacy protection submodule, which performs word segmentation on the data to be analyzed to obtain a dimension matrix corresponding to the vector dimension of each word segment in the data to be analyzed; wherein, the dimension matrix is ​​used to represent the first vector data to be analyzed;

[0019] Wherein, the dimension matrix is ​​h (i) The h (i) = l×d, where i represents the ranking order of any privacy protection submodule in the first large language submodel, l represents the total length of all words after the data to be analyzed is processed, and d is a preset vector dimension; when any privacy protection submodule is the second privacy protection submodule, i = 1.

[0020] Optionally, the step of performing privacy protection processing on the first vector data to be analyzed according to privacy parameters to obtain the first output data corresponding to the second privacy protection submodule includes:

[0021] The first vector data to be analyzed is input into the privacy protection function formula to obtain the first output data; wherein, the privacy protection function formula is: The The p is used to characterize the output data of any privacy-preserving submodule in the first large language submodel, where p is an l-dimensional vector, and any element p in p... i The parameter is obtained by sampling from the interval set U[1,1+δ], where δ is the privacy parameter. A vector of all 1s with dimension d d The transpose of , where ⊙ is the Hadama product.

[0022] Optionally, the step of compensating the model output data according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed includes:

[0023] The model output data is input into the compensation function calculation formula to obtain the hidden state data; wherein, the compensation function calculation formula is: The The m is used to characterize the hidden state data obtained after compensation processing in the second major language sub-model, and the h is used to represent the ranking order of the last privacy-preserving sub-module in the second major language sub-model. (m) The output data used to characterize the last privacy-preserving submodule in the second major language submodel, where C = c·1 l , where c is the target compensation coefficient, and C is an l-dimensional vector in which all elements are c.

[0024] Optionally, before compensating the model output data according to the target compensation coefficient, the method further includes:

[0025] Set multiple initial compensation coefficients and initial training data;

[0026] The initial training data is input into all privacy-preserving sub-modules in the first large language sub-model, and based on different initial compensation coefficients, the training output data corresponding to the initial training data is obtained from the output of the first large language sub-model.

[0027] The output data to be trained is compensated to obtain the hidden state data to be trained corresponding to the initial training data, and the hidden state data to be trained is sent to the large language model to obtain the training result data returned by the large language model.

[0028] Identify the target result data with the highest accuracy among all training result data, and use the initial compensation coefficient corresponding to the target result data as the target compensation coefficient.

[0029] A second aspect of this application provides a privacy protection system based on a large language model, applied to a user terminal. The large language model runs on the user terminal, and the large language model includes a privacy protection module. The system includes:

[0030] The segmentation unit is used to segment the large language model into a first large language sub-model and a second large language sub-model, using the first privacy protection sub-module in the privacy protection module as the segmentation point; wherein, the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking.

[0031] The acquisition unit is used to acquire the data to be analyzed input by the user into the large language model, input the data to be analyzed into the first large language sub-model to obtain intermediate output data, and transmit the intermediate output data to the second large language sub-model to obtain model output data corresponding to the second large language sub-model.

[0032] The sending unit is used to compensate the model output data according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed, and send the hidden state data to the server equipped with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data.

[0033] The large-model-based privacy protection system provided in the second aspect of this application is used to execute the large-model-based privacy protection method described in the first aspect.

[0034] A third aspect of this application provides a privacy protection device based on a large model, comprising:

[0035] Central processing unit, memory, input / output interfaces, wired or wireless network interfaces, and power supply;

[0036] The memory is either a short-term storage memory or a persistent storage memory;

[0037] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the privacy protection method based on a large model as described in the first aspect.

[0038] A fourth aspect of this application provides a computer-readable storage medium including instructions that, when executed on a computer, cause the computer to perform the privacy protection method based on a large model as described in the first aspect.

[0039] A fifth aspect of this application provides a computer program product including instructions that, when executed on a computer, cause the computer to perform the privacy protection method based on a large model as described in the first aspect.

[0040] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: The privacy protection method based on a large model disclosed in this application can ensure that the server with a large language model cannot directly obtain user data, while also compensating for the performance loss caused by the large language model in terms of privacy protection. Through the combination of these two aspects, user privacy is effectively protected without affecting model performance. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0042] Figure 1 This is a diagram of a distributed inference architecture disclosed in an embodiment of this application;

[0043] Figure 2 This is a flowchart illustrating a privacy protection method based on a large model disclosed in an embodiment of this application;

[0044] Figure 3 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application;

[0045] Figure 4 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application;

[0046] Figure 5 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application;

[0047] Figure 6 This is a schematic diagram of a privacy protection module disclosed in an embodiment of this application;

[0048] Figure 7 This is a schematic diagram of the structure of a privacy protection system based on a large model disclosed in an embodiment of this application;

[0049] Figure 8 This is a schematic diagram of the structure of a privacy protection device based on a large model disclosed in an embodiment of this application. Detailed Implementation

[0050] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0051] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of those features. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0053] To resolve the technical issues raised above, please refer to [link / reference]. Figure 1 , Figure 1This is a diagram illustrating a distributed inference architecture disclosed in an embodiment of this application. It should be noted that this embodiment primarily employs a distributed inference architecture, where the provider of the large language model (or described as a large model) reaches an agreement with the user to deploy a small portion of the large model on the user's end. Figure 1 As shown in (a), a small number of modules (including the embedding layer of the large language model and a small number of attention layers. These modules are then transformed into "privacy-preserving modules" through subsequent operations) are deployed on the user's end to protect user privacy. Figure 1 Example (a) illustrates a scenario where a user uploads a "context" and provides a series of "instructions" for the large model to return routes to city B. In this case, the large model should infer from the context that the user is currently in city A and provide routes from city A to city B. Furthermore, in Figure 1 (a) The user transmits a hidden state to the large model after it has been processed by the privacy protection module. We assume that if a malicious service provider obtains the hidden state, even if it uses advanced reconstruction attack methods, it will be unable to reconstruct the user's original data (e.g., ...). Figure 1 (as shown in the green box in (a)). The red box shows a failed privacy protection, where the service provider completely reconstructed the user's original data. Figure 1 (b) demonstrates the objective of the present invention, namely, that after passing through the privacy protection module we designed, the large model service provider is unable to reconstruct the user's original data, while the large model is able to return the correct response, namely, to give the correct travel route from city A to city B.

[0054] Please see Figure 2 , Figure 2 This is a flowchart illustrating a privacy protection method based on a large model disclosed in an embodiment of this application. It includes steps 201-203.

[0055] 201. Using the first privacy protection sub-module in the privacy protection module as the dividing point, the large language model is divided into the first large language sub-model and the second large language sub-model.

[0056] It should be noted in advance that this embodiment applies to the user end, where a large language model runs. The user end can refer to a business server or a terminal cluster. The aforementioned business server can be an independent physical server, a server cluster composed of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud databases, cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Terminal devices can be smartphones, tablets, laptops, desktop computers, PDAs, mobile internet devices (MIDs), wearable devices (such as smartwatches and smart bracelets), smart computers, smart in-vehicle systems, and other smart terminals.

[0057] In this embodiment, to design a reliable user-side privacy protection module, the privacy protection module in the large language model needs to be segmented. Specifically, taking the first privacy protection sub-module in the privacy protection module as the segmentation point, the large language model is segmented into a first large language sub-model and a second large language sub-model. It should be noted that the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking.

[0058] In one specific embodiment, the first privacy protection submodule can be the second-to-last privacy protection submodule in the large language model. That is, the last privacy protection submodule is located in the second large language submodel, and the remaining privacy protection submodules are located in the first large language submodel. Furthermore, the first and second large language submodels described above can also be understood as dividing the large language model into submodels that perform different module functions, rather than splitting the large language model into two independent submodels. Moreover, the privacy protection submodule refers to the understanding of each layer of privacy protection modules in the large language model. See details for further information. Figure 6 , Figure 6 This is a schematic diagram of a privacy protection module disclosed in an embodiment of this application. Figure 6 It can be seen that, let Figure 1 The privacy protection module in the middle consists of (m-1) layers. Figure 6 Submodules in (a) and Level 1 Figure 6The sub-modules in (b) are chained together (therefore, the privacy protection module on the user side has a total of m layers). In the first (m-1) layers, a privacy protection PrivScale operation is performed on the output of each layer (this operation will be explained in detail later). That is, the PrivScale operation is performed on every i-th layer (i < m, and i and m are both positive integers) to achieve privacy protection. In the m-th layer, a CompScale operation is performed on the output of the m-th layer (CompScale combines word compensation and scale transformation, which can be understood as "scaling transformation for compensation". The specific operation is explained later, which is roughly to scale up the output of the large language model proportionally) to compensate for performance. Finally, the output of the m-th layer is sent as the final hidden state to the large model service provider to obtain the response returned by the large model. In one feasible technical solution, m can be set to 10, 20, etc., which is not limited here.

[0059] 202. Obtain the data to be analyzed input by the user in the large language model, input the data to be analyzed into the first large language sub-model to obtain intermediate output data; transfer the intermediate output data to the second large language sub-model to obtain the model output data corresponding to the second large language sub-model.

[0060] Then, based on step 201, the data to be analyzed input by the user in the large language model can be obtained. Thus, the large language model inputs the data to be analyzed into the first large language sub-model to obtain intermediate output data. Then, the intermediate output data is transmitted to the second large language sub-model to obtain the model output data corresponding to the second large language sub-model.

[0061] In one specific embodiment, combined with Figure 1 As shown, users can input various statements, prompts, or commands (using English as an example) into the input interface of a user terminal equipped with a large language model, such as a travel route to a certain place. The large language model can then recognize these inputs and generate data to be analyzed. This data is then input into the first large language sub-model (first into the privacy-preserving sub-module corresponding to the first layer of the first large language sub-model; then the output of the first layer is used as the input of the second layer, and the data is analyzed sequentially according to its order). This yields the data output by the privacy-preserving sub-module of the first large language sub-model. The PrivScale operation is then performed on this data to obtain intermediate output data (the execution steps are the same in the first large language sub-model). It should be noted that the intermediate output data can be understood as the data output by the layer corresponding to the first privacy-preserving sub-module.

[0062] Then, the intermediate output data is used as the input data of the second major language sub-model. The intermediate output data is then analyzed by the privacy protection sub-module of the second major language sub-model to obtain the output data of the second major language sub-model, i.e., the model output data.

[0063] 203. Compensate the model output data according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed, and send the hidden state data to the server with the large language model to obtain the analysis result data replied by the server after analyzing the hidden state data.

[0064] Then, the model output data is compensated according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed. The hidden state data is then sent to the server with a large language model to obtain the analysis result data returned by the server after analyzing the hidden state data.

[0065] In one specific embodiment, the model output data is subjected to a CompScale operation by filtering the target compensation coefficients, thereby obtaining the latent state data output by the second large language sub-model. It should be noted that this latent state data corresponds to the user-input data to be analyzed. Then, the user-side large language model can send this latent state data to the server (which also has a similar large language model as the user-side model). The server then uses the large language model to analyze the latent state data and obtain the analysis results. The user-side can then receive the analysis results returned by the server, which represent responses to various statements, prompts, or commands input by the user.

[0066] The privacy protection method based on a large model disclosed in this embodiment ensures that the server with a large language model cannot directly obtain user data, while also compensating for the performance loss caused by the large language model in terms of privacy protection. By combining these two approaches, user privacy is effectively protected without affecting model performance.

[0067] To further explain Figure 2 For step 201 shown, please refer to [link / reference]. Figure 3 , Figure 3 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application. It includes steps 301-303.

[0068] 301. Obtain the ranking order of all privacy protection sub-modules in the large language model, and set the privacy protection sub-module whose ranking order is the target ranking as the first privacy protection sub-module.

[0069] Corresponding to Figure 2As shown in step 201, in this embodiment, the ranking order of all privacy protection sub-modules in the large language model can be obtained first, and then the privacy protection sub-module with the target ranking order can be set as the first privacy protection sub-module.

[0070] In one specific embodiment, it is necessary to first obtain the privacy protection module set in the large language model, then determine all privacy protection sub-modules contained in the privacy protection module and their ranking order, and then set the privacy protection sub-module with the target ranking order as the first privacy protection sub-module. Furthermore, this first privacy protection sub-module is used as the splitting point between the first large language sub-module and the second large language sub-model.

[0071] Furthermore, in combination Figure 6 As is known, the first privacy-preserving submodule can be understood as the privacy-preserving submodule corresponding to the (m-1)th layer in the large language model. Here, m is the number of all privacy-preserving submodules contained in the large language model.

[0072] 302. Set a preset number of privacy protection sub-modules in the first major language sub-model, and set the first privacy protection sub-module as the last privacy protection sub-module in the first major language sub-model.

[0073] Therefore, based on step 301, a preset number of privacy protection sub-modules can be set in the first major language sub-model, and the first privacy protection sub-module can be set as the last privacy protection sub-module in the first major language sub-model.

[0074] In one specific embodiment, combined with Figure 6 As described in the preceding steps, the (m-1) privacy protection sub-modules can be set in the first large language sub-model. Correspondingly, the first privacy protection sub-module is the last privacy protection sub-module in the first large language sub-model.

[0075] 303. Set the privacy protection sub-module that ranks after the first privacy protection sub-module to the second major language sub-model.

[0076] By combining step 302, the privacy protection submodule, which ranks after the first privacy protection submodule, can be set in the second major language submodel.

[0077] In one specific embodiment, the privacy protection submodule ranked m-th is placed in the second largest language submodel. Furthermore, m privacy protection submodules are deployed in the first and second largest language submodels, respectively. Specifically, the first (m-1) privacy protection submodules are located in the first largest language submodel, and the m-th privacy protection submodule is located in the second largest language submodel.

[0078] In other feasible technical solutions, privacy protection sub-modules can be set for the first and second major language sub-models in other ways, or other numbers of privacy protection sub-modules can be set for the first and second major language sub-models. For example, at least two privacy protection sub-modules can be set for the second major language sub-model. No specific restrictions are imposed here.

[0079] Furthermore, for ease of subsequent understanding, the privacy protection sub-module is divided as follows:

[0080] The privacy protection method based on a large model disclosed in this embodiment can design a reliable user-side privacy protection module by splitting the privacy protection module into large language sub-models that implement different privacy protection operations. After the hidden state after passing through this module is transmitted to the service provider of the large language model, they will not be able to reconstruct the user's original data based on these hidden states.

[0081] To further explain Figure 2 For steps 202-203 shown, please refer to [link / reference]. Figure 4 , Figure 4 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application. It includes steps 401-405.

[0082] 401. Input the data to be analyzed into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed.

[0083] It should be noted that steps 401-405 in this embodiment mainly involve... Figure 2 The following is a detailed description of steps 202-203. In this embodiment, the acquired data to be analyzed is first input into the second privacy protection submodule of the first large language submodel, thereby obtaining the first vector data to be analyzed corresponding to the data to be analyzed. It should be noted that the second privacy protection submodule is the first privacy protection submodule in the first large language submodel, that is, the privacy protection submodule ranked first (i=1) in the first large language submodel.

[0084] In one specific embodiment, the data to be analyzed can be input into a second privacy protection submodule, which performs word segmentation on the data to obtain a dimension matrix corresponding to the vector dimension of each segmented word in the data. Specifically, the dimension matrix is ​​used to represent the first vector data to be analyzed. The dimension matrix is ​​h. (i) h (i)= l×d formula (1), where i is used to characterize the ranking order of any privacy-preserving submodule in the first large language submodel (i.e., the layer number or ranking order of the privacy-preserving submodule at this time), l is used to characterize the total length of all words after the data to be analyzed has been processed by word segmentation, and d is the preset vector dimension. Furthermore, when any privacy-preserving submodule is the second privacy-preserving submodule, i = 1.

[0085] Therefore, as can be understood from the above description, for the first major language sub-model, let the original output of the i-th layer be h. (i) , where h (i) It is an l×d tensor (which can be understood as a matrix), where l is the length of the user input prompt words (which can be understood as the total length of the user output after word segmentation; for example, the word "playing" will be segmented into "play" and "ing," while the word "play" will still be segmented into "play"), and d is the dimension of the intermediate state of each word token (which can be a pre-set vector dimension, such as 4096 or 3028, etc.). Since there are l tokens in total, each token corresponds to an output d-dimensional vector, and concatenating them together gives the l×d h. (i) It should be noted that large language models typically convert each segmented word into a vector before performing calculations on this vector. Here, we assume the vector has a dimension of d, and since there are l segmented words in total, the final result is a 1*d matrix.

[0086] 402. Perform privacy protection processing on the first vector data to be analyzed according to the privacy parameters to obtain the first output data corresponding to the second privacy protection submodule.

[0087] Then, based on step 401, the first vector data to be analyzed can be subjected to privacy protection processing according to the privacy parameters, thereby obtaining the first output data corresponding to the second privacy protection submodule.

[0088] In one specific embodiment, the first vector data to be analyzed can be input into the privacy protection function formula to obtain the first output data. It should be noted that the privacy protection function formula is... Formula (2), Used to represent the output data of any privacy-preserving submodule in the first major language submodel, p is an l-dimensional vector, and any element p in p... i It is obtained by sampling from the interval set U[1,1+δ], where δ is a privacy parameter. A vector of all 1s with dimension d d The transpose of ⊙ is the Hadama product.

[0089] Therefore, as described above, it can be understood that after obtaining the first vector data to be analyzed from the second privacy-preserving submodule, a PrivScale operation can be performed on this first vector data to obtain the first output data from the first-layer privacy-preserving submodule. Furthermore, for the privacy-preserving submodules of other layers of the first large language submodel, combined with step 402, assuming that after the PrivScale operation, the new i-th layer output is obtained as follows: In other words, the output h of the i-th layer (i) Then, after the PrivScale operation, the result is obtained. Where p is an l-dimensional vector (p multiplied by...) It's a 1*d matrix with exactly dimension h. (i) Like the same, they can be multiplied together), each p in p i The elements are sampled from a uniform distribution U[1, 1+δ], where δ is a user-defined privacy parameter; a larger δ provides stronger privacy protection, and a smaller δ provides weaker privacy protection. Furthermore, U[1, 1+δ] is a random sample, meaning a number is randomly selected from 1 to 1+δ. It is a d-dimensional vector of all 1s. d transpose, It is an h (i) The same l×d matrix as before, ⊙ is the Hadamard product. Note that this method preserves the direction of the size contraction because for h... (i) Each d-dimensional vector in After the above operations, the direction of the vector will not change; only the scale will change.

[0090] Therefore, it can be understood that if it is the 3rd layer, then the output is...

[0091] 403. Use the first output data as the input data of the next privacy protection submodule of the second privacy protection submodule, and follow the ranking order until the first privacy protection submodule of the first large language submodel receives the input data of the third privacy protection submodule and outputs intermediate output data.

[0092] Therefore, based on step 402, the first output data can be used as the input data for the next privacy protection submodule of the second privacy protection submodule, and so on, until the first privacy protection submodule of the first large language submodel receives the input data of the third privacy protection submodule and outputs intermediate output data. It should be noted that the third privacy protection submodule is the previous privacy protection submodule of the first privacy protection submodule.

[0093] In one specific embodiment, the first output data of the second privacy protection submodule located in the first layer of the first large language submodel is used as the input data of the privacy protection submodule located in the next layer below the second privacy protection submodule. Then, following the order of steps 401-402, the operations of formula (1) and formula (2), i.e., the PrivScale operation, are performed sequentially. Then, according to the ranking order, until the first privacy protection submodule of the first large language submodel receives the input data of the privacy protection submodule located in the layer above the first privacy protection submodule and outputs intermediate output data.

[0094] Furthermore, assuming the first output data of the first layer is Therefore, it can be understood that at this time, it is to The input data for the second-layer privacy protection submodule is then used to obtain h. (2) and Until the first privacy protection submodule receives input data from the privacy protection submodule located at the layer above it. Then, the first privacy protection submodule can output intermediate output data, which is...

[0095] 404. Transfer the intermediate output data to the second major language sub-model to obtain the model output data corresponding to the second major language sub-model.

[0096] In this embodiment, step 404 is the same as described above. Figure 2 Step 202 is similar, and will not be elaborated here. However, it should be noted that in this embodiment, the intermediate output data at this time is... Will The input is fed into the privacy-preserving submodule (m-th block) of the second largest language submodel, thereby enabling the privacy-preserving submodule of the second largest language submodel to achieve privacy protection using formula (1). The analysis yields h. (m) This refers to the model output data described above.

[0097] 405. Input the model output data into the compensation function calculation formula to obtain the hidden state data, and send the hidden state data to the server with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data.

[0098] Based on step 405, the model output data is input into the compensation function calculation formula to obtain the hidden state data. This hidden state data is then sent to a server configured with a large language model to obtain the analysis results returned by the server after analyzing the hidden state data. It should be noted that the compensation function calculation formula is... The hidden state data obtained after compensation processing in the second major language sub-model is used to represent the hidden state data. 'm' represents the ranking order of the last privacy-preserving sub-module in the second major language sub-model. 'h' represents the position of the last privacy-preserving sub-module in the second major language sub-model. (m) The output data used to characterize the output of the last privacy-preserving submodule in the second largest language submodel, C = c·1 l c is the target compensation coefficient, and C is an l-dimensional vector in which all elements are c.

[0099] Therefore, as can be understood from the above description, it is necessary to output data h for the model at the m-th layer. (m) Perform the CompScale operation. Specifically, after (m-1) consecutive layers of PrivScale operations, the original output h of the m-th layer is calculated. (m) The following formula (3) is used for compensation, where C = c·1 l It is an l-dimensional vector with all elements equal to c, where c is the compensation coefficient. It's important to note that the compensation coefficient is a constant, such as 1, 2, 3, etc. Essentially, it's equivalent to multiplying the output of the m-th layer by c.

[0100] This embodiment discloses a privacy protection method based on a large model. First, it employs orientation-preserving stochastic scaling of the hidden state, a technique that can protect user privacy data while minimizing the impact on model usability. Then, it proposes an orientation-preserving adaptive hidden state compensation technique to further mitigate the performance degradation caused by the previous operation. By combining these two methods, privacy is protected without compromising the usability of the large model.

[0101] For other feasible technical solutions, please refer to Figure 5 , Figure 5 This is a flowchart illustrating another privacy protection method based on a large model disclosed in an embodiment of this application. It includes steps 501-504.

[0102] 501. Set initial training data and multiple initial compensation coefficients.

[0103] It should be noted that this embodiment mainly describes the strategy for obtaining compensation coefficients. Specifically, it is necessary to first set initial training data and multiple initial compensation coefficients.

[0104] In one specific implementation, several mathematical problems are first selected from the GSM8K mathematical computation dataset. GSM8K is a dataset with many mathematical problems, such as "Xiaoming has 2 apples, eats one, how many are left?" (In reality, the problems are usually more difficult than this example). This problem is then input into the model, and the model's ability to answer it is evaluated. The compensation coefficient *c* is continuously adjusted until the model answers the most problems correctly. Then, these selected mathematical problems are used as initial training data, and multiple different compensation coefficients are set as initial compensation coefficients.

[0105] 502. Input the initial training data into all privacy-preserving sub-modules in the first large language sub-model, and based on different initial compensation coefficients, obtain the training output data corresponding to the initial training data output by the first large language sub-model.

[0106] Then, the initial training data is input into all privacy-preserving sub-modules in the first large language sub-model, and based on different initial compensation coefficients, the output data to be trained corresponding to the initial training data is obtained from the output of the first large language sub-model.

[0107] In one specific embodiment, see the above. Figure 2 , Figure 3 and Figure 4 In the embodiment shown, the initial training data is used as the input data of the large language model (understood as the data to be analyzed). These mathematical problems are input into the privacy protection submodule of the first (m-1) layer (using formula (1) and formula (2)) to obtain the training output data of the initial training data output by the first large language submodel.

[0108] 503. Perform compensation processing on the output data to be trained to obtain the hidden state data to be trained corresponding to the initial training data, and send the hidden state data to be trained to the large language model to obtain the training result data returned by the large language model.

[0109] Then, the output data to be trained is compensated to obtain the hidden state data to be trained corresponding to the initial training data. The hidden state data to be trained is then sent to the large language model to obtain the training result data returned by the large language model.

[0110] In one specific embodiment, the training output data output by the first privacy protection submodule is again processed using formula (2). Then, the data obtained according to formula (2) is compensated using different values ​​of c. It should be noted that a value of c is used for compensation at the m-th layer, for example, c is increased by 0.5 each time from 1 to 5. Then, the training hidden state data corresponding to different values ​​of c can be obtained respectively, i.e. In short, the following simplified steps describe how the mathematical problem is input into the model, and the results are obtained sequentially. h (m) Then, use c against h (m) To obtain compensation Then The data is transmitted to the server, and the results returned by the server's large language model are obtained, which is the training result data.

[0111] 504. Determine the target result data with the highest accuracy among all training result data, and use the initial compensation coefficient corresponding to the target result data as the target compensation coefficient.

[0112] Then, the target result data with the highest accuracy among all training result data can be determined, and the initial compensation coefficient corresponding to the target result data can be used as the target compensation coefficient.

[0113] In one specific embodiment, when the result returned by the large language model has the highest accuracy among these mathematical problems, the c value at this time is selected as the final compensation coefficient, i.e., the target compensation coefficient.

[0114] The privacy protection method based on a large model disclosed in this embodiment can effectively adjust the compensation value in the large language model, thereby achieving the highest accuracy of the results returned by the large language model and improving the feasibility of the solution.

[0115] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0116] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a privacy protection system based on a large model disclosed in an embodiment of this application.

[0117] The segmentation unit 701 is used to segment the large language model into a first large language sub-model and a second large language sub-model, with the first privacy protection sub-module in the privacy protection module as the segmentation point; wherein, the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking.

[0118] The acquisition unit 702 is used to acquire the data to be analyzed input by the user in the large language model, input the data to be analyzed into the first large language sub-model to obtain intermediate output data, and transmit the intermediate output data to the second large language sub-model to obtain the model output data corresponding to the second large language sub-model.

[0119] The sending unit 703 is used to compensate the model output data according to the target compensation coefficient, obtain the hidden state data corresponding to the data to be analyzed, and send the hidden state data to the server with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data.

[0120] For example, the system further includes: a setting unit 704;

[0121] The acquisition unit 702 is specifically used to acquire the ranking order of all privacy-preserving sub-modules in the large language model, and set the privacy-preserving sub-module with the target ranking order as the first privacy-preserving sub-module;

[0122] Setting unit 704 is used to set a preset number of privacy protection sub-modules in the first large language sub-model, and set the first privacy protection sub-module as the last privacy protection sub-module in the first large language sub-model;

[0123] Setting unit 704 is also used to set the privacy protection sub-modules that are ranked after the first privacy protection sub-modules in the second major language sub-model.

[0124] For example, the system includes:

[0125] The acquisition unit 702 is specifically used to input the data to be analyzed into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed; wherein, the second privacy protection submodule is the first privacy protection submodule in the first large language submodel;

[0126] The acquisition unit 702 is also used to perform privacy protection processing on the first vector data to be analyzed according to privacy parameters, so as to obtain the first output data corresponding to the second privacy protection submodule;

[0127] Setting unit 704 is specifically used to take the first output data as the input data of the next privacy protection submodule of the second privacy protection submodule, and according to the ranking order, until the first privacy protection submodule of the first large language submodel receives the input data of the third privacy protection submodule and outputs intermediate output data; wherein, the third privacy protection submodule is the previous privacy protection submodule of the first privacy protection submodule.

[0128] For example, the system further includes: an input unit 705;

[0129] The input unit 705 is used to input the data to be analyzed into the second privacy protection submodule, perform word segmentation on the data to be analyzed, and obtain a dimension matrix corresponding to the vector dimension of each word length in the data to be analyzed; wherein, the dimension matrix is ​​used to represent the first vector data to be analyzed;

[0130] Where the dimension matrix is ​​h (i) h (i) = l×d, where i represents the ranking of any privacy protection submodule in the first large language submodel, l represents the total length of all words after the data to be analyzed is processed, and d is the preset vector dimension; when any privacy protection submodule is the second privacy protection submodule, i = 1.

[0131] For example, the system includes:

[0132] Input unit 705 is specifically used to input the first vector data to be analyzed into the privacy protection function calculation formula to obtain the first output data; wherein, the privacy protection function calculation formula is: Used to represent the output data of any privacy-preserving submodule in the first major language submodel, p is an l-dimensional vector, and any element p in p... i It is obtained by sampling from the interval set U[1,1+δ], where δ is a privacy parameter. A vector of all 1s with dimension d d The transpose of ⊙ is the Hadama product.

[0133] For example, the system includes:

[0134] Input unit 705 is specifically used to input model output data into the compensation function calculation formula to obtain hidden state data; wherein, the compensation function calculation formula is: The hidden state data obtained after compensation processing in the second major language sub-model is used to represent the hidden state data. 'm' represents the ranking order of the last privacy-preserving sub-module in the second major language sub-model. 'h' represents the position of the last privacy-preserving sub-module in the second major language sub-model. (m) The output data used to characterize the output of the last privacy-preserving submodule in the second largest language submodel, C = c·1 l c is the target compensation coefficient, and C is an l-dimensional vector in which all elements are c.

[0135] For example, the system further includes: a setting unit 706 and a determining unit 707;

[0136] The setting unit 706 is used to set multiple initial compensation coefficients and initial training data;

[0137] The input unit 705 is also used to input the initial training data into all privacy-preserving sub-modules in the first large language sub-model, and to obtain the training output data corresponding to the initial training data output by the first large language sub-model based on different initial compensation coefficients.

[0138] The acquisition unit 702 is also used to perform compensation processing on the output data to be trained to obtain the hidden state data to be trained corresponding to the initial training data, so as to send the hidden state data to be trained to the large language model to obtain the training result data returned by the large language model.

[0139] The determination unit 707 is used to determine the target result data with the highest accuracy among all training result data, and to use the initial compensation coefficient corresponding to the target result data as the target compensation coefficient.

[0140] Please refer to the following: Figure 8 The schematic diagram of a privacy protection device based on a large model disclosed in this application includes:

[0141] Central processing unit 801, memory 805, input / output interface 804, wired or wireless network interface 803, and power supply 802;

[0142] Memory 805 is either a short-term storage memory or a persistent storage memory;

[0143] The central processing unit 801 is configured to communicate with the memory 805 and execute instructions stored in the memory 805 to perform the aforementioned operations. Figures 2 to 5 Privacy protection methods based on large models in any of the illustrated embodiments.

[0144] This application also provides a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the aforementioned... Figures 2 to 5 Privacy protection methods based on large models in any of the illustrated embodiments.

[0145] This application also provides a computer-readable storage medium, which includes instructions that, when executed on a computer, cause the computer to perform the aforementioned actions. Figures 2 to 5 Privacy protection methods based on large models in any of the illustrated embodiments.

[0146] This application also provides a computer program product containing instructions, which, when run on a computer, causes the computer to perform the aforementioned... Figures 2 to 5 Privacy protection methods based on large models in any of the illustrated embodiments.

[0147] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0148] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0149] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0150] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0151] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A privacy protection method based on a large model, characterized in that, Applied to the user terminal, the large language model runs on the user terminal, the large language model includes a privacy protection module, and the method includes: Using the first privacy protection sub-module in the privacy protection module as the dividing point, the large language model is divided into a first large language sub-model and a second large language sub-model; wherein, the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking. The system acquires the data to be analyzed input by the user into the large language model, inputs the data to be analyzed into the first large language sub-model to obtain intermediate output data, and transmits the intermediate output data to the second large language sub-model to obtain model output data corresponding to the second large language sub-model. The model output data is compensated according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed, and the hidden state data is sent to the server with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data. The step of inputting the data to be analyzed into the first large language sub-model to obtain intermediate output data includes: The data to be analyzed is input into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed; wherein, the second privacy protection submodule is the first privacy protection submodule in the first large language submodel; The first vector data to be analyzed is subjected to privacy protection processing based on privacy parameters to obtain the first output data corresponding to the second privacy protection submodule; The first output data is used as the input data of the next privacy protection submodule of the second privacy protection submodule, and according to the ranking order, until the first privacy protection submodule of the first large language submodel receives the input data of the third privacy protection submodule and outputs the intermediate output data; wherein, the third privacy protection submodule is the previous privacy protection submodule of the first privacy protection submodule; The step of inputting the data to be analyzed into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed includes: The data to be analyzed is input into the second privacy protection submodule, which performs word segmentation on the data to be analyzed to obtain a dimension matrix corresponding to the vector dimension of each word segment in the data to be analyzed; wherein, the dimension matrix is ​​used to represent the first vector data to be analyzed; Wherein, the dimension matrix is The The The ranking order of any privacy-preserving sub-module within the first large language sub-model is used to characterize the position of that sub-module. This is used to characterize the total length of all words in the data to be analyzed after word segmentation processing. The preset vector dimension; when any of the privacy protection submodules is the second privacy protection submodule, the ; The step of performing privacy protection processing on the first vector data to be analyzed according to privacy parameters to obtain the first output data corresponding to the second privacy protection submodule includes: The first vector data to be analyzed is input into the privacy protection function formula to obtain the first output data; wherein, the privacy protection function formula is: The The output data used to characterize the output of any privacy-preserving submodule in the first large language submodel, for A vector of dimension, and the any element in For the set of intervals The sample obtained from the middle, the For the privacy parameter, the for A vector of all 1s of dimension The transpose, the For Hadama accumulation.

2. The privacy protection method based on a large model according to claim 1, characterized in that, The step of dividing the large language model into a first large language sub-model and a second large language sub-model, using the first privacy protection sub-module in the privacy protection module as the dividing point, includes: Obtain the ranking order of all privacy-preserving sub-modules in the large language model, and set the privacy-preserving sub-module with the target ranking order as the first privacy-preserving sub-module; Set a preset number of privacy protection sub-modules in the first large language sub-model, and set the first privacy protection sub-module as the last privacy protection sub-module in the first large language sub-model; The privacy protection submodules whose ranking order is after the first privacy protection submodule are set in the second major language submodel.

3. The privacy protection method based on a large model according to claim 1, characterized in that, The step of compensating the model output data according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed includes: The model output data is input into the compensation function calculation formula to obtain the hidden state data; wherein, the compensation function calculation formula is: The The hidden state data obtained after compensation processing in the second major language sub-model is used to characterize the hidden state data. The ranking order used for the last privacy-preserving submodule in the second major language submodel, the The output data used to characterize the output of the last privacy-preserving submodule in the second major language submodel, The The target compensation coefficient is the... For one All elements of the dimension are The vector.

4. The privacy protection method based on a large model according to claim 1, characterized in that, Before compensating the model output data according to the target compensation coefficient, the method further includes: Set initial training data and multiple initial compensation coefficients; The initial training data is input into all privacy-preserving sub-modules in the first large language sub-model, and based on different initial compensation coefficients, the training output data corresponding to the initial training data is obtained from the output of the first large language sub-model. The output data to be trained is compensated to obtain the hidden state data to be trained corresponding to the initial training data, and the hidden state data to be trained is sent to the large language model to obtain the training result data returned by the large language model. Identify the target result data with the highest accuracy among all training result data, and use the initial compensation coefficient corresponding to the target result data as the target compensation coefficient.

5. A privacy protection system based on a large model, characterized in that, Applied to the user terminal, the large language model runs on the user terminal, the large language model includes a privacy protection module, and the system includes: The segmentation unit is used to segment the large language model into a first large language sub-model and a second large language sub-model, using the first privacy protection sub-module in the privacy protection module as the segmentation point; wherein, the privacy protection module includes multiple privacy protection sub-modules, and the first privacy protection sub-module is the privacy protection sub-module in the large language model whose ranking order is the target ranking. The acquisition unit is used to acquire the data to be analyzed input by the user into the large language model, input the data to be analyzed into the first large language sub-model to obtain intermediate output data, and transmit the intermediate output data to the second large language sub-model to obtain model output data corresponding to the second large language sub-model. The sending unit is used to compensate the model output data according to the target compensation coefficient to obtain the hidden state data corresponding to the data to be analyzed, and send the hidden state data to the server with the large language model to obtain the analysis result data returned by the server after analyzing the hidden state data. The system includes: The acquisition unit is specifically used to input the data to be analyzed into the second privacy protection submodule of the first large language submodel to obtain the first vector data to be analyzed corresponding to the data to be analyzed; wherein, the second privacy protection submodule is the first privacy protection submodule in the first large language submodel; The acquisition unit is also used to perform privacy protection processing on the first vector data to be analyzed according to privacy parameters to obtain the first output data corresponding to the second privacy protection submodule; The setting unit is specifically used to take the first output data as the input data of the next privacy protection submodule of the second privacy protection submodule, and according to the ranking order, until the first privacy protection submodule of the first large language sub-model receives the input data of the third privacy protection submodule and outputs the intermediate output data; wherein, the third privacy protection submodule is the previous privacy protection submodule of the first privacy protection submodule; The system also includes: an input unit; The input unit is used to input the data to be analyzed into the second privacy protection submodule, perform word segmentation processing on the data to be analyzed, and obtain a dimension matrix corresponding to the vector dimension of each word length in the data to be analyzed; wherein, the dimension matrix is ​​used to represent the first vector data to be analyzed; Wherein, the dimension matrix is The The The ranking order of any privacy-preserving sub-module within the first large language sub-model is used to characterize the position of that sub-module. This is used to characterize the total length of all words in the data to be analyzed after word segmentation processing. The preset vector dimension; when any of the privacy protection submodules is the second privacy protection submodule, the ; The system includes: The input unit is specifically used to input the first vector data to be analyzed into the privacy protection function formula to obtain the first output data; wherein, the privacy protection function formula is: The The output data used to characterize the output of any privacy-preserving submodule in the first large language submodel, for A vector of dimension, and the any element in For the set of intervals The sample obtained from the middle, the For the privacy parameter, the for A vector of all 1s of dimension The transpose, the For Hadama accumulation.

6. A privacy protection device based on a large model, characterized in that, The device includes: Central processing unit, memory, input / output interfaces, wired or wireless network interfaces, and power supply; The memory is either a short-term storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the privacy protection method based on a large model as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on a computer, cause the computer to perform the privacy protection method based on a large model as described in any one of claims 1 to 4.