Parameter space conversion-based split learning large language model privacy protection method

By splitting the large language model, performing parameter space conversion and freezing the parameters of the head segment model, combined with multi-objective loss function training, the problem of data reconstruction attacks in split learning is solved, and the balance between security and downstream task performance is achieved.

CN120296781APending Publication Date: 2025-07-11HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359076.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the split learning scenario, there is a risk of privacy leakage of data reconstruction attacks during fine-tuning of large language models. The existing privacy protection methods cannot be effectively prevented, and it is difficult to balance privacy protection and downstream task performance.

Method used

By splitting the large language model into the head segment model, the middle segment model and the tail segment model, the projection layer is connected and preheated training is used to convert the model parameters to the specified parameter space, freeze the parameters of the head segment model, and collaborately train the middle segment model and the tail segment model on the device side to build a multi-objective loss function for privacy protection.

Benefits of technology

It effectively prevents data reconstruction attacks, improves the security of large language models, and realizes privacy protection without affecting the performance of downstream tasks, and can defend against two-way attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296781A_ABST
    Figure CN120296781A_ABST
Patent Text Reader

Abstract

The invention relates to a parameter space conversion-based split learning large language model privacy protection method, which comprises the following steps of: splitting a pre-trained large language model to obtain a head section model, a middle section model and a tail section model of the large language model, and deploying the head section model and the tail section model at a first equipment end, deploying the middle section model at a second equipment end; the head section model and the tail section model are connected based on the projection layer, and preheating training is performed on the connected head section model and tail section model, so that parameters of the head section model and the tail section model are converted to a specified parameter space; and freezing parameters in the head section model, and cooperatively training the middle section model and the tail section model based on the first equipment end and the second equipment end to obtain a fine-tuned large language model. By adopting the method, the safety in the large language model splitting learning process can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and particularly to a privacy protection method for split learning large language models based on parameter space conversion. Background Art

[0002] With the wide application of large language model technology, more and more small and medium-sized enterprises hope to train and deploy their own private large models through a first data set. However, in the split learning scenario, although the fine-tuning of large models can be achieved through the collaborative training of local clients and remote servers without transmitting raw data, there is still a risk of privacy leakage from data reconstruction attacks by malicious attackers through the transmitted intermediate activations and gradient information, and the current privacy protection means cannot fully and effectively protect the data privacy in this scenario. During the fine-tuning process of large language models based on split learning, it is difficult to effectively prevent data reconstruction attacks, and there is a problem of low security in the fine-tuning process of split learning large language models.

[0003] Regarding the problem of low security in the fine-tuning process of split learning large language models in the related art, no effective solution has been proposed yet. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a privacy protection method for split learning large language models based on parameter space conversion that can solve the problem of low security in the fine-tuning process of split learning large language models.

[0005] In a first aspect, in the present embodiment, a privacy protection method for split learning large language models based on parameter space conversion is provided, and the method includes:

[0006] Split a pre-trained large language model to obtain a head segment model, an intermediate segment model, and a tail segment model of the large language model, deploy the head segment model and the tail segment model on a first device end, and deploy the intermediate segment model on a second device end;

[0007] Connect the head segment model and the tail segment model based on a projection layer, and perform warm-up training on the connected head segment model and tail segment model so that the parameters of the head segment model and the tail segment model are converted to a specified parameter space;

[0008] Freeze the parameters in the head segment model, and co-train the intermediate segment model and the tail segment model based on the first device end and the second device end to obtain the adjusted large language model.

[0009] In some of these embodiments, the tail segment model includes an identical first tail segment model and a second tail segment model, wherein the parameters of the second tail segment model are frozen; the preheating training of the connected head segment model and the tail segment model includes:

[0010] Connect the head segment model and the first tail segment model based on a projection layer, and construct a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model;

[0011] Construct an adversarial loss function, which includes a first adversarial loss function and / or a second adversarial loss function; wherein, the first adversarial loss function is used to maximize the difference between the output data of the inversion model corresponding to the head segment model and the input data of the head segment model; the second adversarial loss function is used to maximize the difference between the prediction data of the second tail segment model and the corresponding real data, and the second tail segment model is connected to the head segment model through the projection layer;

[0012] On the first device side, preheat and train the head segment model and the first tail segment model based on the first task loss function and the adversarial loss function.

[0013] In some of these embodiments, the constructing a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model includes:

[0014] Obtain the activation amount obtained by the forward propagation of the head segment model on the first data set;

[0015] Project the activation amount by the projection layer to obtain projection data;

[0016] Input the projection data into the first tail segment model to obtain first prediction data;

[0017] Construct the first task loss function for minimizing the difference between the first prediction data and the corresponding first real data in the first data set.

[0018] In some of these embodiments, the constructing the adversarial loss function includes:

[0019] Obtain the activation amount obtained by the forward propagation of the head segment model on the first data set;

[0020] Construct the inversion model according to the second data set and the head segment model, and input the activation amount into the inversion model to obtain the output data;

[0021] Obtain the activation amount obtained by the forward propagation of the head segment model on the first data set;

[0022] Construct the first adversarial loss function for maximizing the difference between the output data and the corresponding input data in the first data set.

[0023] In some embodiments, the constructing of the adversarial loss function includes:

[0024] Obtain the activation amount obtained by the forward propagation of the head segment model on the first data set;

[0025] Project the activation amount by the projection layer to obtain projection data;

[0026] Input the projection data into the second tail segment model to obtain second prediction data;

[0027] Construct a second adversarial loss function for maximizing the difference between the second prediction data and the corresponding second true data in the projection data.

[0028] In some embodiments, the co-training of the head segment model, the middle segment model, and the tail segment model based on the first device and the second device to obtain the large language model with globally parameter fine-tuned includes:

[0029] According to the first data set, construct a second task loss function corresponding to the head segment model, the middle segment model, and the first tail segment model; wherein, the first tail segment model is the tail segment model with parameters converted to the specified parameter space; the second task loss function is associated with the downstream task of the large language model and is used to minimize the prediction data output by the first tail segment model and the corresponding true data in the first data set;

[0030] According to the first data set, construct a third adversarial loss function corresponding to the head segment model, the middle segment model, and the second tail segment model; wherein, the second tail segment model is the tail segment model with parameters not converted to the specified parameter space; the third adversarial loss function is used to maximize the prediction data output by the second tail segment model and the corresponding true data in the first data set;

[0031] On the first device and the second device, co-train the middle segment model and the first tail segment model based on the second task loss function and the third adversarial loss function.

[0032] In a second aspect, in this embodiment, a privacy protection device for a split learning large language model based on parameter space conversion is provided, and the device includes:

[0033] A splitting module, configured to split a pre-trained large language model to obtain a head segment model, a middle segment model, and a tail segment model of the large language model, and deploy the head segment model and the tail segment model on a first device end, and deploy the middle segment model on a second device end;

[0034] A training module, configured to connect the head segment model and the tail segment model based on a projection layer, and perform warm-up training on the connected head segment model and tail segment model, so that the parameters of the head segment model and the tail segment model are transformed into a specified parameter space;

[0035] A fine-tuning module, configured to freeze the parameters in the head segment model, and co-train the middle segment model and the tail segment model based on the first device end and the second device end to obtain the adjusted large language model.

[0036] In a third aspect, in this embodiment, a split learning large language model privacy protection system based on parameter space conversion is provided, and the system includes:

[0037] A processing device, configured to implement the split learning large language model privacy protection method based on parameter space conversion described in the first aspect above;

[0038] A first device end, configured to deploy the head segment model and the tail segment model;

[0039] A second device end, configured to deploy the middle segment model.

[0040] In a fourth aspect, in this embodiment, a computer device is provided, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the split learning large language model privacy protection method based on parameter space conversion described in the first aspect above.

[0041] In a fifth aspect, in this embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the split learning large language model privacy protection method based on parameter space conversion described in the first aspect above.

[0042] The above privacy protection method for split learning large language models based on parameter space transformation connects the head segment model and the tail segment model through a projection layer, so that the parameters of the head segment model and the tail segment model after warm-up training are transformed into a specified parameter space; by freezing the parameters of the head segment model, it is avoided that when the first device and the second device cooperate in training, the parameters of the head segment model and the tail segment model return to the original pre-trained parameter space; it is ensured that attackers cannot use prior knowledge such as pre-trained models to train inversion models or perform gradient matching attacks, which will not have a negative impact on the downstream tasks of large language models, and improves the security during the split learning process of large language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is an application environment diagram of the privacy protection method for split learning large language models based on parameter space transformation in an embodiment;

[0044] Figure 2 It is a schematic flowchart of the privacy protection method for split learning large language models based on parameter space transformation in an embodiment;

[0045] Figure 3 It is a schematic diagram of the privacy protection method for split learning large language models in an embodiment;

[0046] Figure 4 It is a schematic flowchart of the privacy protection method for split learning large language models based on parameter space transformation in another embodiment;

[0047] Figure 5 It is a structural block diagram of the privacy protection device for split learning large language models based on parameter space transformation in an embodiment;

[0048] Figure 6 It is a structural block diagram of the privacy protection system for split learning large language models based on parameter space transformation in an embodiment;

[0049] Figure 7 It is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] The split learning strategy reduces the computational and storage burden on local devices by splitting the model layers, with a part of the layers computed on the local machine and the intermediate computation results and the computation tasks of another part of the layers sent to a remote server for delegation. Through the collaboration between the local and the remote, the fine-tuning training of large models can be completed. Therefore, split learning can provide an effective solution for small and medium-sized enterprises or research institutions with limited resources.

[0052] However, split learning still has deficiencies at present. Although split learning does not directly share the original data and can protect data privacy to a certain extent, since it still needs to transmit intermediate computation results such as forward activations and backward gradients between the client and the server, attackers can still carry out data reconstruction attacks through these intermediate computation results. For example, an attacker can perform a model inversion attack based on the received activation values of the client-side head model output to recover the original input; the attacker can also execute a gradient matching attack through the gradients transmitted from the client-side tail model to recover the label data. Therefore, additional data privacy protection methods are needed to protect the private data in the split fine-tuning process.

[0053] In related technologies, the privacy of private data is mainly protected by methods such as differential privacy, DP-forward (Differential Privacy-forward, differential privacy defense based on forward propagation), and DP-SGD (Differential Privacy Stochastic Gradient Descent, differential privacy defense based on gradient descent). The differential privacy method adds noise during the training process. For example, noise is added to the activation values and gradients, so that attackers cannot infer the original data through the reconstruction process. The DP-forward method protects the transmitted activation values by introducing noise in the forward propagation to prevent attackers from reconstructing private data from them. The DP-SGD method adds noise to the gradients in the backward propagation to prevent attackers from inferring private labels through gradient matching attacks.

[0054] Although these defense techniques are effective in preventing certain types of attacks, they still have some problems. For example, it is difficult for related technologies to achieve a balance between privacy protection and the performance of downstream tasks: in some scenarios that require strong privacy protection, a higher level of noise needs to be added to the data. However, as the noise level increases, the task performance of the model is often affected. If the noise level is too small, although the model's performance on the task may be well maintained, it may not be able to fully resist data reconstruction attacks. There is a significant trade-off between privacy protection and downstream tasks, and it is impossible to protect privacy without negatively affecting the performance of the model's downstream tasks. Moreover, related technologies do not effectively consider the problem of the "not-too-far" property of large model fine-tuning. This is because, during the model fine-tuning process, since large language models usually only update a small number of parameters, the parameter space between the fine-tuned model and the pre-trained model is very similar. In LLM-FT (Large Language Model Fine-Tuning), the parameters of the fine-tuned model are very similar to those of the pre-trained model, which provides a large amount of prior knowledge for attackers, resulting in the ability of attackers to still use this prior knowledge to conduct attacks even under defense measures.

[0055] Based on this, the present application proposes a privacy protection method for split learning large language models based on parameter space transformation.

[0056] In one embodiment, the privacy protection method for split learning large language models based on parameter space transformation provided by the embodiments of the present application can be applied to, for example Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. Through the collaborative training of the terminal 102 and the server 104, global parameter fine-tuning is performed on the split large language model. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0057] In one embodiment, as Figure 2 shown, a privacy protection method for split learning large language models based on parameter space transformation is provided. Taking the application of this method to Figure 1 the application environment shown as an example, the method includes the following steps:

[0058] Step S202: Split the pre-trained large language model to obtain the head segment model, the middle segment model, and the tail segment model of the large language model. Deploy the head segment model and the tail segment model on the first device end, and deploy the middle segment model on the second device end.

[0059] Among them, the head segment model is used to effectively extract data features, and the tail segment model is used to adapt to downstream tasks; the head segment model and the tail segment model contain a small number of model layers. The middle segment model is used to implement the core training and calculation of the model, and the middle model segment contains a large number of model layers. There are limitations in the computing and storage space of the first device end, and the computing and storage space of the second device end is larger than that of the first device end.

[0060] Optionally, the first device end is the user's client, and the second device end is the server end. By deploying the head segment model and the tail segment model on the first device end and deploying the middle segment model on the second device end, the computing and storage costs of the first device end during the model training process can be reduced.

[0061] Optionally, an existing open-source large model can be selected as the pre-trained large language model. The model is split according to the Transformer blocks hierarchical structure, and the model is divided into a head segment model, a middle segment model, and a tail segment model. The head segment model and the tail segment model contain a small number of blocks (basic building units) and are deployed on the first device end, while the middle segment model contains a large number of blocks and is deployed on the second device end model.

[0062] Step S204: Connect the head segment model and the tail segment model based on the projection layer, and perform warm-up training on the connected head segment model and tail segment model to transfer the parameters of the head segment model and the tail segment model to the specified parameter space.

[0063] Among them, a three-layer MLP (Multilayer Perceptron) projection network can be selected as the projection layer. It can be understood that other projection networks can also be selected as the projection layer, and the projection layer is not restricted here. Warm-up training refers to a method of parameter adjustment for specific model layers (the head segment model and the tail segment model) at the initial stage of training the split learning large language model. Optionally, introduce local warm-up training on the first device end: After using the projection layer to connect the head model segment and the tail model segment, the connected head segment model and tail segment model are trained through multi-task learning. The multi-task learning objectives in local warm-up training can be set according to requirements.

[0064] Optionally, when pre-training the concatenated head segment model and tail segment model, in the case of forward propagation, the output of the head segment model will first pass through the projection of the projection layer and then be input into the tail segment model; similarly, in the case of backward propagation, the output of the tail segment model will also first pass through the projection of the projection layer and then be input into the head segment model. Since the projection layer can project the input, such as through linear transformation or non-linear mapping, etc., into another parameter space, the parameters of the pre-trained head segment model and tail segment model after pre-training can be converted from the original pre-training parameter space of the large language model to a specified parameter space that is quite different from the pre-training parameter space through the local pre-training on the first device.

[0065] Step S206: Freeze the parameters in the head segment model, and co-train the middle segment model and the tail segment model based on the first device and the second device to obtain an adjusted large language model.

[0066] Among them, the parameters in the head segment model converted to the specified parameter space are set to an unmodifiable state to prevent the model parameters in the specified parameter space from transferring back to the original pre-training parameter space of the large language model during the fine-tuning training of the first device and the second device.

[0067] The first device and the second device work together to complete the fine-tuning of the large language model by using the transmitted intermediate activation values and gradient information. Optionally, during the forward propagation process, the first device and the second device communicate to transmit the intermediate activation values; during the backward propagation process, the first device and the second device communicate to transmit the gradient information. By co-training the middle segment model and the tail segment model, the adaptability of the large language model to downstream tasks can be improved.

[0068] In the above split learning large language model privacy protection method based on parameter space conversion, by training the head segment model and the tail segment model connected by the projection layer in the first device, the safe conversion of the parameter spaces of the head segment model and the tail segment model is realized. The parameter conversion of the head segment model and the tail segment model to the specified parameter space can ensure that attackers cannot use prior knowledge such as pre-trained models to train inversion models or perform gradient matching attacks, so as to avoid attackers reconstructing the original data or original labels of the large language model based on the intermediate data exchanged between the first device and the first device during the split learning process. By freezing the parameters in the head segment model, it is avoided that the fine-tuning of the large language model parameters causes the parameters to return to the original pre-training parameter space, further reducing the data privacy risk brought by the "not-too-far" problem (that is, the fine-tuned model parameters are too similar to the pre-trained model, resulting in the problem of leaking prior knowledge), and solving the problem of low security in the split learning process of the large language model.

[0069] At the same time, compared with the method of adding noise to the data, the split learning large language model privacy protection method based on parameter space transformation in this embodiment does not affect the training of downstream tasks, achieving a balance between privacy protection and downstream task performance.

[0070] Furthermore, the related technologies also have the problem of only being able to defend against one-way attacks. For example, DP-forward mainly defends against forward attacks, while DP-SGD focuses on defending against backward attacks. However, due to the autoregressive nature of large language models, there is a strong overlap between the input and the label. Therefore, attackers can improve the accuracy of reconstruction by combining forward and backward attacks, which makes the related technologies perform poorly in defending against two-way attacks, that is, considering both forward attacks and backward attacks at the same time. For example, the defense method against forward attacks cannot defend against the gradient matching attack from backpropagation, and the defense method against backward attacks cannot defend against the model inversion attack from forward propagation.

[0071] Based on this, in one embodiment, the tail segment model includes the same first tail segment model and second tail segment model, wherein the parameters of the second tail segment model are frozen; preheating training is performed on the connected head segment model and tail segment model, including: connecting the head segment model and the first tail segment model based on the projection layer, and constructing a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model; constructing an adversarial loss function, the adversarial loss function including a first adversarial loss function and / or a second adversarial loss function; wherein, the first adversarial loss function is used to maximize the difference between the output data of the inversion model corresponding to the head segment model and the input data of the head segment model; the second adversarial loss function is used to maximize the difference between the predicted data of the second tail segment model and the corresponding real data, and the second tail segment model is connected to the head segment model through the projection layer; at the first device, preheating training is performed on the head segment model and the first tail segment model based on the first task loss function and the adversarial loss function.

[0072] Among them, the first task loss function is used to quantify the deviation between the predicted data of the tail segment model and the real data. By constructing the first task loss function, while achieving privacy protection, the performance and generalization ability of the model in downstream tasks are improved. The first task loss function can be set according to specific downstream tasks.

[0073] The inversion model can be constructed based on the head segment model. The inversion model performs reverse operations relative to the head segment model and can form a complementary reverse mapping relationship with the head segment model. The output data of the inversion model corresponds to the input data in the initial head segment model. By setting the first adversarial loss function to train the head segment model, the difference between the output data of the inversion model and the input data of the head segment model is increased, reducing the possibility of an attacker attacking through the data in the forward propagation process during the subsequent collaborative training process between the first device and the second device.

[0074] The parameters of the second tail segment model cannot be modified. Optionally, two identical tail segment models are deployed on the first device to obtain the first tail segment model and the second tail segment model, and the parameters of the second tail segment model are frozen. By setting the second adversarial loss function and training the head segment model, the projection layer, and the first tail segment model based on the second adversarial loss function, the difference between the trained first tail segment model and the frozen second tail segment model can be increased, thereby reducing the possibility of an attacker attacking through the data in the backpropagation process during the subsequent collaborative training process between the first device and the second device.

[0075] Optionally, if defense against forward attacks is required, a first task loss function and a first adversarial loss function can be constructed; if defense against gradient matching attacks in backpropagation is required, a first task loss function and a second adversarial loss function can be constructed. If defense against two-way attacks is required, a first task loss function, a first adversarial loss function, and a second adversarial loss function can be constructed, solving the limitation that the defense methods in related technologies usually only target attacks in a single direction.

[0076] In this embodiment, while implementing privacy protection based on the projection layer, a loss function is constructed based on multi-objective learning for training. Among them, constructing the first task loss function can maintain the performance of the downstream task and ensure that the performance of the downstream task is not significantly affected while implementing privacy protection. Constructing the first adversarial loss function can transfer the weight parameters of the head segment model of the first device to the specified parameter space during training, increasing the difficulty for the attacker to reconstruct the original data through the inversion model. By constructing the second adversarial loss function, the applicability of the second tail segment model can be disrupted, making it impossible for the attacker to effectively infer the output of the downstream task through the original pre-trained weights.

[0077] Next, how to construct the first task loss function, the first adversarial loss function, and the second adversarial loss function will be described.

[0078] In one embodiment, constructing a first task loss function for the association between the model and the downstream task of the large language model based on the connected head segment model and the first tail segment model includes: obtaining the activation amount obtained by the forward propagation of the head segment model on the first data set; projecting the activation amount by a projection layer to obtain projection data; inputting the projection data into the first tail segment model to obtain first prediction data; constructing a first task loss function for minimizing the difference between the first prediction data and the corresponding first real data in the first data set.

[0079] Among them, the first data set refers to a set of private data that is non-public and restricted in access. The first data set is associated with the downstream task and can be selected according to requirements. Optionally, taking the private data set D as the first data set, for the k-th batch of data D in the private data set D k Input it into the head segment model and perform forward propagation to obtain the intermediate activation amount a h of this batch, that is, the activation amount. Pass the intermediate activation amount a h through the projection layer Proj for projection to obtain the projection result a p ; input a p to the tail segment model to obtain the first prediction data. According to the difference between the first prediction data and the corresponding first real data in the first data set, calculate the first task loss function L local_ft of the downstream task.

[0080] In this embodiment, constructing the first task loss function for the head segment model and the first tail segment model connected by the projection layer, during warm-up training, adjusting the parameters of the head segment model and the first tail segment model connected by the projection layer based on the first task loss function to reduce the gap between the prediction data of the tail segment model and the real data, while protecting the model privacy, will not significantly affect the performance of the large language model in the downstream task.

[0081] In one embodiment, constructing an adversarial loss function includes: obtaining the activation amount obtained by the forward propagation of the head segment model on the first data set; constructing an inversion model according to the second data set and the head segment model, and inputting the activation amount into the inversion model to obtain output data; obtaining the activation amount obtained by the forward propagation of the head segment model on the first data set; constructing a first adversarial loss function for maximizing the difference between the output data and the corresponding input data in the first data set.

[0082] Among them, the second data set refers to a set of public data that is public and unrestricted in access. The second data set is associated with the downstream task and can be selected according to requirements. Optionally, constructing an inversion model according to the second data set and the head segment model, so that the inversion model can obtain the input data of the head segment model through reverse operation according to the output data of the initially deployed head segment model.

[0083] Optionally, use a publicly available text generation dataset as the second dataset, and train an inversion model W based on the second dataset and the head segment model. sip_inv ; Use the private dataset D as the first dataset, and for the k-th batch of data D in the private dataset D k Perform forward propagation on the head model segment to obtain the intermediate activation a h , that is, the activation. Feed a h Into the inversion model W sip_inv , to obtain the inversion result, that is, the output data of the inversion model. Calculate the MSE loss L between the inversion result and the true input (the corresponding input data in the first dataset) inv , and generate the first adversarial loss function:

[0084]

[0085] It can be understood that L inv Can also be other existing types of loss functions. In addition to using the reciprocal of L inv As the first adversarial loss function, other mathematical transformation or mapping methods can also be used to generate a first adversarial loss function that aims to increase the difference between the inversion result and the true input.

[0086] In this embodiment, by combining the head segment model and the inversion model of the head segment model to construct the first adversarial loss function, and adjusting the parameters of the head segment model based on the first adversarial loss function during warm-up training, the parameters of the head segment model can be transformed into a more secure specified parameter space, and the difference between the output data of the inversion model and the corresponding input data in the first dataset can be increased, reducing the success rate of an attacker reconstructing the original input data based on the data transmitted during the collaborative training process between the first device and the second device.

[0087] In one embodiment, constructing an adversarial loss function includes: obtaining the activation obtained by the head segment model performing forward propagation on the first dataset; projecting the activation by a projection layer to obtain projection data; inputting the projection data into a second tail segment model to obtain second prediction data; constructing a second adversarial loss function that maximizes the difference between the second prediction data and the corresponding second true data in the projection data.

[0088] Among them, the second adversarial loss function is used to train the head segment module connected by the projection layer and the first tail segment model. The parameters of the second tail segment model are frozen and cannot be modified. Optionally, use a publicly available text generation dataset as the second dataset, and train an inversion model W based on the second dataset and the head segment model. sip_inv ; Use the private dataset D as the first dataset, and for the k-th batch of data D in the private dataset D kPerform forward propagation on the head model segment to obtain the intermediate activation a h , which is the activation. Input a h to the second tail segment model to obtain the second predicted data output by the second tail segment model. According to the second predicted data and the corresponding second ground truth data in the projection data, calculate the loss L lical_pt of the second tail segment model, and generate the loss for adversarial training of the second tail segment model, that is, the second adversarial loss function:

[0089]

[0090] It can be understood that the loss L local_pt of the second tail segment model can be obtained based on existing loss construction methods. In addition to taking the reciprocal of L local_pt as the second adversarial loss function, other mathematical transformation or mapping methods can also be used to generate the second adversarial loss function for increasing the difference between the second predicted data and the second ground truth data.

[0091] In this embodiment, by constructing the second adversarial loss function, there is a certain difference between the trained first tail segment model and the second tail segment model, making it difficult for attackers to infer the output of the large language model for downstream tasks based on the data related to the first tail end model.

[0092] Furthermore, in one embodiment, on the first device side, when pre-training the head segment model and the first tail segment model based on the first task loss function, the first adversarial loss function, and / or the second adversarial loss function, corresponding weights λ1 and λ2 can be set for the first adversarial loss function and / or the second adversarial loss function. Optionally, taking the case of simultaneously selecting the first adversarial loss function L anti_inv and the second adversarial loss function L local_pt , and the first task loss function is L local_ft as an example, integrate the first task loss function, the first adversarial loss function, and the second adversarial loss function to obtain the total loss L warm_up in the pre-training stage:

[0093] L warm_up = L local_ft + λ1 * L anti_inv + λ2 * L anti_local_pt

[0094] Based on the total loss, perform backpropagation on the head segment model, the projection layer, and the first tail segment model, and update the model through iterative training until the head segment model and the first tail segment model after being connected by the projection layer converge, or the number of iterations reaches the predetermined number of training epochs.

[0095] In one embodiment, according to the first data set, a second task loss function corresponding to the head segment model, the middle segment model, and the first tail segment model is constructed; wherein, the first tail segment model is a tail segment model whose parameters are transformed into a specified parameter space; the second task loss function is associated with the downstream task of the large language model and is used to minimize the predicted data output by the first tail segment model and the corresponding real data in the first data set; according to the first data set, a third adversarial loss function corresponding to the head segment model, the middle segment model, and the second tail segment model is constructed; wherein, the second tail segment model is a tail segment model whose parameters are not transformed into the specified parameter space; the third adversarial loss function is used to maximize the predicted data output by the second tail segment model and the corresponding real data in the first data set; on the first device and the second device, the middle segment model and the first tail segment model are co-trained based on the second task loss function and the third adversarial loss function.

[0096] Wherein, the head segment model and the first tail segment model are connected through a projection layer. Similarly, the head segment model and the second tail segment model are connected through a projection layer. The parameters in the head segment model and the first tail segment model are adjusted through warm-up training. Optionally, according to the first data set and the downstream task, a third task loss function corresponding to the head segment model, the middle segment model, and the second tail segment model is constructed, and the third task loss function is used to minimize the predicted data output by the second tail segment model and the corresponding real data in the first data set; the third task loss function is adjusted to obtain a third adversarial loss function for maximizing the predicted data output by the second tail segment model and the corresponding real data in the first data set.

[0097] When co-training based on the second task loss function and the third adversarial loss function, the second task loss function and the third adversarial loss function can be integrated, and a corresponding weight can be set for the third adversarial loss function.

[0098] Optionally, the first device freezes the head segment model, prevents the parameters of the head segment model from being transferred back to the insecure parameter space, performs forward propagation on the first data set using the frozen head segment model to generate SmashedData (intermediate output data) of the cutting layer; sends the Smashed Data to the second device, and performs forward propagation on the middle segment model of the second device to obtain the activation amount a output by the middle segment model. s 。Send a s back to the first device.

[0099] Perform forward propagation on a s on the first tail segment model of the first device, and according to the forward propagation, obtain the prediction result and the corresponding real data in the first data set, and calculate the second task loss function L corresponding to the downstream task. global_ft 。

[0100] On the first device side, input a s to the second tail segment model, and calculate the third task loss function L for the downstream task according to the prediction result obtained by performing forward propagation on the second tail segment model and the corresponding real data in the first dataset global_pt , and generate an adversarial pre-training tail model loss based on the third task loss function, that is, the third adversarial loss function;

[0101]

[0102] On the first device side and the second device side, co-training the middle segment model and the first tail segment model based on the second task loss function and the third adversarial loss function includes: obtaining the total loss L in the global fine-tuning stage according to the second task loss function and the third adversarial loss function global :

[0103] L global = L global_ft + λ3 * L anti_global_pt

[0104] According to L global Perform backpropagation on the middle segment model on the second device side and the first tail segment model on the first device side, and update the model through iterative training until the middle segment model and the first tail segment model converge, or the number of iterations reaches a predetermined number of training epochs.

[0105] Considering that during the fine-tuning process, since the middle segment model on the second device side may affect the parameters of the client tail model segment, causing it to return to the original pre-trained model parameter space, thus exposing privacy information. Therefore, in this embodiment, a model fine-tuning training based on the second task loss function and the third task loss function is realized by combining the frozen head segment model, the frozen second tail segment model, the middle segment model whose parameters need to be adjusted, and the first tail segment model. Among them, the frozen head segment model is used to avoid its transfer back to the insecure parameter space, and at the same time, the second task loss function and the third task loss function further ensure the difference in the parameter spaces of the second tail segment model and the first tail segment model during the training process, so that the first tail model segment maintains its converted specified parameter space during the entire fine-tuning process and is not affected by the middle segment model, thereby effectively preventing attacks based on the prior knowledge of the pre-trained model.

[0106] In one embodiment, Figure 3 A schematic diagram of a privacy protection method for a split learning large language model is provided. As Figure 3As shown in the figure, it includes the splitting of the pre-trained large model, the conversion of the local parameter space, and the global fine-tuning parameter space preservation strategy. When splitting the pre-trained large language model, the head segment model, the middle segment model, and the tail segment model of the large language model are obtained. The first device end is the client, which deploys the head segment model, the first tail segment model, and the second tail segment model. The second device end is the server end, which deploys the middle segment model.

[0107] The conversion of the local parameter space is achieved through local warm-up training, and the warm-up training can be carried out according to the total loss L in the warm-up training stage. warm_up Specifically, during local warm-up training, the head segment model and the first tail segment model are connected through a projection layer (Protection Layer). The private dataset, that is, the first dataset, is input into the head segment model, and the first prediction data output by the first tail segment model is obtained. And a first task loss function is constructed as L local_ft , so that the output of the first tail segment model approaches the correct prediction. A first adversarial loss function L anti_inv is constructed to make the output data of the inversion model of the head segment model far from the input data of the head segment model. A second adversarial loss function L local_pt is constructed according to the second tail segment model to make the output prediction data of the first tail segment model far from the prediction data output by the second tail segment model.

[0108] The global fine-tuning parameter space preservation strategy freezes the head segment model and trains the first tail segment model and the middle segment model. Among them, the global fine-tuning can be carried out according to the total loss L in the global fine-tuning stage. global Among them, is the gradient input from the first tail segment model to the middle segment model, and A s2t is the data of the gradient input from the middle segment model to the first tail segment model; is the gradient input from the first tail segment model to the middle segment model, and A s2h is the data of the gradient input from the middle segment model to the first tail segment model.

[0109] In this embodiment, protection is carried out through non-linear transformation and multi-task learning framework, so that personal privacy data cannot be reconstructed through activation or gradient in forward propagation. Combining the loss L warm_up and the loss L global can prevent forward data reverse attack and backward gradient matching attack, and effectively cope with two-way enhanced reconstruction attack.

[0110] In one embodiment, Figure 4 provides another schematic flow diagram of the privacy protection method for the split learning large language model based on parameter space conversion. Among them, the client is the first device end, and the server end is the second device end. As Figure 3As shown, it includes a local warm-up stage and a global fine-tuning stage:

[0111] Step S401, split the pre-trained large language model. Optionally, the pre-trained large model is Llama3-70b. Split the pre-trained large model with 80 layers of blocks. The head model segment contains 4 layers of blocks, and the tail model segment contains 4 layers of blocks, which are deployed on the client side; the middle model segment contains 72 layers of blocks, which are deployed on the remote server side.

[0112] Step S402, connect the head segment model and the tail segment model on the client through a projection layer. Optionally, on the client side, connect the head model (4 layers of blocks) directly to the tail model (4 layers of blocks) after passing through a projection layer (3 layers of MLP).

[0113] Step S403, perform local warm-up training using the training data until the model converges or reaches the specified number of training epochs. The training data includes the first data set and the second data set. Optionally, introduce a warm-up training on the client side. By means of multi-task learning, transform the parameters of the head model segment and the tail model segment after being linked by the projection layer from the original pre-trained space to a specified parameter space that is safe and has a large difference from the pre-trained model parameter space, minimizing the risk of privacy leakage. The loss function constructed based on multi-task learning includes the first task loss function, the first adversarial loss function, and the second adversarial loss function in the above embodiments.

[0114] Step S404, perform global fine-tuning of the large model based on split learning with parameter space preservation, and iteratively train until the model converges or reaches the predetermined number of training epochs.

[0115] Optionally, perform global split learning for large model fine-tuning, and adopt a parameter space preservation strategy during global fine-tuning, including: freezing the parameters in the head segment model, constructing the second task loss function and the third adversarial loss function in the above embodiments, and co-training the middle segment model and the first tail segment model based on the second task loss function and the third adversarial loss function. Thus, prevent the client model parameter space from returning to the parameter space consistent with the pre-trained model.

[0116] It should be understood that the pre-trained large model can be other types of models other than Llama3-70b, such as Llama (Large Language Model Meta AI), GLM (General Language Model), and other series of models launched by Baichuan. The method in this embodiment can be applied to scenarios where enterprises, organizations, or individuals hope to fine-tune the pre-trained large model using private data sets containing sensitive privacy to obtain a privatized large model.

[0117] In this embodiment, the strategy of implementing parameter space conversion through local warm-up training enables the local client to convert model parameters into a secure space that is significantly different from the pre-trained model through multi-task learning. This effectively prevents attackers from launching attacks using prior knowledge and greatly enhances the privacy protection ability. Through the parameter space conversion strategy and the global fine-tuning parameter space maintenance strategy, privacy protection can be effectively improved without sacrificing task performance. By performing local warm-up training on the client side to convert model parameters into a secure parameter space, both the effectiveness of task performance can be ensured and the leakage of prior knowledge can be prevented, thus achieving a better balance between privacy protection and task performance. Moreover, two-way defense is achieved, that is, both forward data reverse attack and backward gradient matching attack are defended against, effectively enhancing the security of the model and ensuring that attackers cannot successfully reconstruct private data through a single-direction attack.

[0118] Based on the same inventive concept, an embodiment of the present application also provides a privacy protection device for a split learning large language model based on parameter space conversion for implementing the above-mentioned privacy protection method for a split learning large language model based on parameter space conversion. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the privacy protection device for a split learning large language model based on parameter space conversion provided below can refer to the limitations on the privacy protection method for a split learning large language model based on parameter space conversion in the above text, and will not be elaborated here.

[0119] In one embodiment, as Figure 5 shown, a privacy protection device for a split learning large language model based on parameter space conversion is provided, including: a splitting module, a training module, and a fine-tuning module, where:

[0120] The splitting module is used to split the pre-trained large language model to obtain the head segment model, the middle segment model, and the tail segment model of the large language model, and deploy the head segment model and the tail segment model on the first device side, and deploy the middle segment model on the second device side;

[0121] The training module is used to connect the head segment model and the tail segment model based on the projection layer, and perform warm-up training on the connected head segment model and tail segment model so that the parameters of the head segment model and the tail segment model are converted into a specified parameter space;

[0122] The fine-tuning module is used to freeze the parameters in the head segment model, and co-train the middle segment model and the tail segment model based on the first device side and the second device side to obtain an adjusted large language model.

[0123] In one embodiment, the tail segment model includes the same first tail segment model and second tail segment model, wherein the parameters of the second tail segment model are frozen; the training module pre-trains the connected head segment model and tail segment model, including: connecting the head segment model and the first tail segment model based on a projection layer, and constructing a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model; constructing an adversarial loss function, the adversarial loss function including a first adversarial loss function and / or a second adversarial loss function; wherein, the first adversarial loss function is used to maximize the difference between the output data of the inversion model corresponding to the head segment model and the input data of the head segment model; the second adversarial loss function is used to maximize the difference between the predicted data of the second tail segment model and the corresponding real data, and the second tail segment model is connected to the head segment model through a projection layer; on the first device side, pre-train the head segment model and the first tail segment model based on the first task loss function and the adversarial loss function.

[0124] Optionally, the training module constructs a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model, including: obtaining the activation amount obtained by the head segment model for forward propagation of the first data set; projecting the activation amount by the projection layer to obtain projection data; inputting the projection data into the first tail segment model to obtain first prediction data; constructing a first task loss function for minimizing the difference between the first prediction data and the corresponding first real data in the first data set.

[0125] Optionally, the training module constructs an adversarial loss function, including: obtaining the activation amount obtained by the head segment model for forward propagation of the first data set; constructing an inversion model according to the second data set and the head segment model, and inputting the activation amount into the inversion model to obtain output data; obtaining the activation amount obtained by the head segment model for forward propagation of the first data set; constructing a first adversarial loss function for maximizing the difference between the output data and the corresponding input data in the first data set.

[0126] Optionally, the training module constructs an adversarial loss function, including: obtaining the activation amount obtained by the head segment model for forward propagation of the first data set; projecting the activation amount by the projection layer to obtain projection data; inputting the projection data into the second tail segment model to obtain second prediction data; constructing a second adversarial loss function for maximizing the difference between the second prediction data and the corresponding second real data in the projection data.

[0127] In one embodiment, the training module collaboratively trains the head segment model, the middle segment model, and the tail segment model based on the first device and the second device to obtain a large language model with fine-tuned global parameters, including: constructing a second task loss function corresponding to the head segment model, the middle segment model, and the first tail segment model according to the first data set; wherein, the first tail segment model is a tail segment model whose parameters are converted to a specified parameter space; the second task loss function is associated with the downstream task of the large language model and is used to minimize the predicted data output by the first tail segment model and the corresponding real data in the first data set; constructing a third adversarial loss function corresponding to the head segment model, the middle segment model, and the second tail segment model according to the first data set; wherein, the second tail segment model is a tail segment model whose parameters are not converted to the specified parameter space; the third adversarial loss function is used to maximize the predicted data output by the second tail segment model and the corresponding real data in the first data set; on the first device and the second device, collaboratively train the middle segment model and the first tail segment model based on the second task loss function and the third adversarial loss function.

[0128] Each module in the above split learning large language model privacy protection device based on parameter space conversion can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0129] In one embodiment, Figure 6 There is also provided a split learning large language model privacy protection system based on parameter space conversion, the system includes: a processing device for implementing the steps in the above method embodiments; a first device for deploying the head segment model and the tail segment model; a second device for deploying the middle segment model.

[0130] Among them, the first device and the second device can be communicatively connected to realize the transmission of intermediate data and activation quantities. Optionally, the first device is a client and the second device is a server. The processing device can be deployed on the first device or the second device, and the processing device can also be independently deployed.

[0131] In one embodiment, there is provided a computer device, which can be a server, and its internal structure diagram can be as Figure 7As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data of the large language model. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a privacy protection method for a split learning large language model based on parameter space transformation.

[0132] Those skilled in the art can understand that Figure 7 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0133] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0134] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0135] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0136] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., and are not limited thereto.

[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0138] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A privacy protection method for split learning large language models based on parameter space transformation, characterized in that The method includes: Splitting a pre-trained large language model to obtain a head segment model, a middle segment model, and a tail segment model of the large language model, deploying the head segment model and the tail segment model on a first device, and deploying the middle segment model on a second device; Connecting the head segment model and the tail segment model based on a projection layer, and performing warm-up training on the connected head segment model and tail segment model to transfer the parameters of the head segment model and the tail segment model to a specified parameter space; Freezing the parameters in the head segment model, and co-training the middle segment model and the tail segment model based on the first device and the second device to obtain the adjusted large language model.

2. The method according to claim 1, wherein The tail segment model includes an identical first tail segment model and a second tail segment model, wherein the parameters of the second tail segment model are frozen; the performing warm-up training on the connected head segment model and tail segment model includes: Connecting the head segment model and the first tail segment model based on a projection layer, and constructing a first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model; Constructing an adversarial loss function, the adversarial loss function including a first adversarial loss function and / or a second adversarial loss function; wherein, the first adversarial loss function is used to maximize the difference between the output data of the inversion model corresponding to the head segment model and the input data of the head segment model; the second adversarial loss function is used to maximize the difference between the prediction data of the second tail segment model and the corresponding real data, and the second tail segment model is connected to the head segment model through the projection layer; On the first device, performing warm-up training on the head segment model and the first tail segment model based on the first task loss function and the adversarial loss function.

3. The method according to claim 2, characterized in that, The constructing the first task loss function associated with the downstream task of the large language model according to the connected head segment model and the first tail segment model includes: Obtaining the activation obtained by the head segment model through forward propagation on a first data set; Projecting the activation by the projection layer to obtain projection data; Inputting the projection data into the first tail segment model to obtain first prediction data; Constructing the first task loss function for minimizing the difference between the first prediction data and the corresponding first real data in the first data set.

4. The method according to claim 2, wherein The constructing the adversarial loss function includes: Obtaining the activation obtained by the head segment model through forward propagation on a first data set; Constructing the inversion model according to a second data set and the head segment model, and inputting the activation into the inversion model to obtain the output data; Obtaining the activation obtained by the head segment model through forward propagation on the first data set; Constructing the first adversarial loss function for maximizing the difference between the output data and the corresponding input data in the first data set.

5. The method according to claim 2, wherein The constructing the adversarial loss function includes: Obtain the activation amount obtained by the forward propagation of the head segment model on the first data set; Project the activation amount by the projection layer to obtain projection data; Input the projection data into the second tail segment model to obtain second prediction data; Construct a second adversarial loss function for maximizing the difference between the second prediction data and the corresponding second true data in the projection data.

6. The method according to claim 1, characterized in that, The collaborative training of the head segment model, the middle segment model, and the tail segment model based on the first device and the second device to obtain the large language model with fine-tuned global parameters includes: According to the first data set, construct a second task loss function corresponding to the head segment model, the middle segment model, and the first tail segment model; wherein, the first tail segment model is the tail segment model whose parameters are converted to the specified parameter space; the second task loss function is associated with the downstream task of the large language model and is used to minimize the prediction data output by the first tail segment model and the corresponding true data in the first data set; According to the first data set, construct a third adversarial loss function corresponding to the head segment model, the middle segment model, and the second tail segment model; wherein, the second tail segment model is the tail segment model whose parameters are not converted to the specified parameter space; the third adversarial loss function is used to maximize the prediction data output by the second tail segment model and the corresponding true data in the first data set; On the first device and the second device, collaboratively train the middle segment model and the first tail segment model based on the second task loss function and the third adversarial loss function.

7. A privacy protection device for split learning large language models based on parameter space transformation, characterized in that, The device includes: A splitting module for splitting the pre-trained large language model to obtain the head segment model, the middle segment model, and the tail segment model of the large language model, deploying the head segment model and the tail segment model on the first device, and deploying the middle segment model on the second device; A training module for connecting the head segment model and the tail segment model based on a projection layer and performing warm-up training on the connected head segment model and tail segment model so that the parameters of the head segment model and the tail segment model are converted to the specified parameter space; A fine-tuning module for freezing the parameters in the head segment model and collaboratively training the middle segment model and the tail segment model based on the first device and the second device to obtain the fine-tuned large language model.

8. A split learning large language model privacy protection system based on parameter space transformation, characterized in that, The system includes: A processing device for implementing the steps of the method according to any one of claims 1 to 6; A first device for deploying the head segment model and the tail segment model; A second device for deploying the middle segment model.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.