End-cloud cooperative training method, system and device

By stripping the representation layer to the terminal and encrypting it in edge-cloud collaborative training, the problem of data leakage during transmission is solved, achieving data privacy security and improving training efficiency. It is suitable for large language models in edge-cloud collaborative training systems.

CN120880679APending Publication Date: 2025-10-31HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410548751.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In the process of edge-cloud collaborative training, how can we improve data privacy and security, prevent client data from being leaked and parsed into plaintext during transmission, especially in the privacy protection of large language models where the model is invisible, and how can we protect user data used for fine-tuning from being leaked as much as possible?

Method used

The client's representation layer is stripped to the terminal for training. Representations and representation vectors are embedded through encryption to avoid plaintext transmission. A low-rank matrix (Lora) is used to encrypt the representation vocabulary, and a feedforward neural network (FNN) is used for dimensionality reduction to ensure that the data is not parsed into plaintext during transmission.

Benefits of technology

It improves the privacy and security of client data, reduces the amount of data transmission, improves training efficiency, and even if the representation vocabulary is leaked in the previous round of collaborative training, it can prevent the embedded representation from being parsed as plaintext when it is leaked through encryption, thus enhancing privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880679A_ABST
    Figure CN120880679A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an end-cloud cooperative training method, system and device in the field of artificial intelligence, and can be used for stripping an embedding layer to terminal training in an end-cloud cooperative training process so as to improve the privacy security of end-side data. The method comprises the steps that a client uses training data as input of a representation layer to obtain a representation vector, the representation layer is used for obtaining a vector corresponding to the input data from a representation word list stored in the client, and the representation word list is stored in the client; the client sends the representation vector to the cloud platform, so that the cloud platform takes the representation vector as the input of a language model deployed at the cloud platform side to obtain an output feature; the client receives an output feature sent by the cloud platform, wherein the output feature is obtained by inputting a representation vector into a language model by the cloud platform; and the client calculates a loss value by using the output feature, updates the representation layer according to the loss value, and sends the loss value to the cloud platform, and the loss value is used for the cloud platform to update the language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to an edge-cloud collaborative training method, system, and device. Background Technology

[0002] The fields of artificial intelligence (AI) and deep learning have undergone tremendous changes, with natural language processing being one of the hottest areas. Edge-cloud collaborative training is becoming increasingly common, especially for large language models, maximizing the computing power of different devices and platforms to achieve flexible deployment and collaborative computing of data and models in the cloud and on the edge.

[0003] In edge-cloud collaborative environments, privacy protection is particularly critical. Data invisibility protection strategies ensure the security and privacy of user data during storage, processing, and transmission, using technologies such as encryption to prevent unauthorized theft or misuse during transmission. Model invisibility privacy protection, on the other hand, is primarily based on commercial protection needs. For example, large models like the GPT series typically only provide API service calls, while their internal model structure, parameters, training and fine-tuning processes are strictly confidential.

[0004] Therefore, how to improve data privacy during edge-cloud collaborative training is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a method, system, and apparatus for edge-cloud collaborative training, which can be used to strip the embedding layer to the terminal training during edge-cloud collaborative training to improve the privacy and security of terminal data.

[0006] In view of this, firstly, this application provides an edge-cloud collaborative training method applied to an edge-cloud collaborative training system, which includes a client and a cloud platform. The method includes: firstly, the client uses training data as input to a representation layer to obtain a representation vector. This representation layer is used to obtain vectors corresponding to the input data from a representation vocabulary stored on the client. The representation vocabulary is stored on the client. Subsequently, the client sends the representation vector to the cloud platform, so that the cloud platform uses the representation vector as input to a language model deployed on the cloud platform to obtain output features. The client receives the output features sent by the cloud platform, which are obtained by the cloud platform inputting the representation vector into the language model. The client calculates a loss value using the output features, updates the representation layer according to the loss value, and sends the loss value to the cloud platform. The loss value is used by the cloud platform to update the language model to obtain an updated language model.

[0007] In the method provided in this application, the client side updates the representation layer locally, so that the representation word table (embedding table) on the client side does not leave the client. When sending data to the cloud platform, the extracted features, i.e. the hidden state, are sent. This can prevent the client data from being leaked and parsed into plaintext during transmission, thereby improving the privacy and security of the client data.

[0008] In one possible implementation, the aforementioned client uses training data as input to the representation layer to obtain a representation vector, including: the client inputting training data into the representation layer and outputting an embedded representation; the client encrypting the embedded representation to obtain a representation vector.

[0009] In this embodiment, when generating the representation vector, the client can encrypt the embedded representation, so that the data transmitted by the client to the cloud platform is encrypted data. Even if the data transmitted by the client to the cloud platform is leaked, it can be further prevented from being parsed as plaintext, thereby improving the privacy and security of the client's data.

[0010] In one possible implementation, the aforementioned client inputs training data into the representation layer and outputs embedded representations, including: the client encrypts the representation vocabulary to obtain an encrypted representation vocabulary; the client inputs training data into the representation layer and determines the embedded representations from the encrypted representation vocabulary through the representation layer.

[0011] In this embodiment, the client can encrypt the representation vocabulary and use the encrypted representation vocabulary to output the embedded representation. The output embedded representation is thus an encrypted representation, further preventing it from being parsed as plaintext when leaked, improving client privacy and security. Furthermore, even if the representation vocabulary used in the previous round of collaborative training on the client side is leaked, in the current collaborative training process, after encrypting the representation vocabulary used in the previous round, the output embedded representation can only be parsed based on the encrypted representation vocabulary. Therefore, it further prevents the embedded representation from being parsed as plaintext when leaked, improving client privacy and security.

[0012] In one possible implementation, the aforementioned client encrypts the representation vocabulary to obtain an encrypted representation vocabulary, including: the client adding a low-rank matrix to the representation vocabulary to obtain the encrypted representation vocabulary.

[0013] In this embodiment, a low-rank matrix (Lora) can be used as the specific encryption method for the representation vocabulary, so that the representation vocabulary can be fine-tuned based on the increased Lora during the training process, thereby achieving model fine-tuning.

[0014] In one possible implementation, the aforementioned client performs encryption processing on the embedded representation to obtain a representation vector, including: the client performs dimensionality reduction processing on the embedded representation to obtain a representation vector.

[0015] In this embodiment, the embedded representation can be dimensionality reduced to decrease the amount of data transmitted between the client and the cloud platform, thereby improving data transmission efficiency and training efficiency. Furthermore, dimensionality reduction of the embedded representation also effectively encrypts it, further preventing it from being parsed as plaintext when leaked, thus enhancing client privacy and security.

[0016] In one possible implementation, the aforementioned client performs dimensionality reduction processing on the embedded representation to obtain a representation vector, including: the client uses the embedded representation as input to a feedforward neural network (FNN) to obtain a representation vector, and the FNN is used to perform linear transformation processing on the input embedded representation.

[0017] Therefore, in this embodiment, an FNN can be used to output the representation vector, and the FNN can be used to achieve dimensionality reduction, thereby reducing the amount of data transmitted between the client and the cloud platform. This further prevents the embedded representation from being parsed as plaintext when it is leaked, improving client privacy and security.

[0018] In one possible implementation, the aforementioned method further includes: the client performing a nonlinear transformation on the representation vector and using the transformed vector as a new representation vector.

[0019] In this embodiment, the representation vector can be further subjected to nonlinear transformation and further encrypted to further prevent the embedded representation from being parsed as plaintext when it is leaked, thereby improving the privacy and security of the client.

[0020] In one possible implementation, the aforementioned client calculates the loss value using the output features, including: the client obtaining the predicted score of each vector in the representation vocabulary as the next vector based on the output features; and the client calculating the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0021] In this embodiment, the client can use the output features sent by the cloud platform to perform a prediction task, and then calculate the corresponding loss value based on the prediction branch of the prediction task to realize the model training process of end-to-cloud collaboration.

[0022] Secondly, this application provides an edge-cloud collaborative training method applied to an edge-cloud collaborative training system, which includes a client and a cloud platform. The method includes: the cloud platform receiving a representation vector sent by the client, the representation vector being obtained by the client inputting training data into the representation layer, and the representation layer being used to obtain the vector corresponding to the input data from the representation vocabulary stored by the client; the cloud platform inputting the representation vector into a language model to obtain output features, and sending the output features to the client; the cloud platform receiving a loss value sent by the client, the loss value being calculated by the client based on the output features; and the cloud platform using the loss value to update the language model to obtain an updated language model.

[0023] In the method provided in this application, the client side updates the representation layer locally, so that the representation word table (embedding table) on the client side does not leave the client. When sending data to the cloud platform, the extracted features, i.e. the hidden state, are sent. This can prevent the client data from being leaked and parsed into plaintext during transmission, thereby improving the privacy and security of the client data.

[0024] In one possible implementation, the aforementioned representation vector is obtained by the client encrypting the embedded representation output by the representation layer. In this embodiment, when generating the representation vector, the client can encrypt the embedded representation, thereby ensuring that the data transmitted by the client to the cloud platform is encrypted. Even if the data transmitted by the client to the cloud platform is leaked, it can be further prevented from being parsed as plaintext, thus improving the client's data privacy and security.

[0025] Thirdly, this application provides a client application, including:

[0026] The input module is used to use training data as input to the representation layer to obtain representation vectors. The representation layer is used to obtain vectors corresponding to the input data from the representation vocabulary stored on the client.

[0027] The transceiver module is used to send representation vectors to the cloud platform; the client receives the output features sent by the cloud platform, and the output features are obtained by the cloud platform inputting the representation vectors into the language model.

[0028] The processing module is used to calculate the loss value using the output features and update the representation layer based on the loss value;

[0029] The transceiver module is also used to send loss values ​​to the cloud platform, which are then used by the cloud platform to update the language model and obtain the updated language model.

[0030] The effects achieved by the third aspect and any optional implementation of the third aspect can be found in the description of the first aspect or any optional implementation of the first aspect, and will not be repeated here.

[0031] In one possible implementation, the aforementioned input module is specifically used to: input training data into the representation layer and output embedded representations; and encrypt the embedded representations to obtain representation vectors.

[0032] In one possible implementation, the aforementioned input module is specifically used to: encrypt the representation vocabulary to obtain an encrypted representation vocabulary; input training data into the representation layer, and determine the embedded representation from the encrypted representation vocabulary through the representation layer.

[0033] In one possible implementation, the aforementioned input module is specifically used to add a low-rank matrix to the representation vocabulary to obtain an encrypted representation vocabulary.

[0034] In one possible implementation, the aforementioned input module is specifically used to perform dimensionality reduction processing on the embedded representation to obtain a representation vector.

[0035] In one possible implementation, the aforementioned input module is specifically used to take the embedded representation as input to the feedforward neural network (FNN) to obtain a representation vector, and the FNN is used to perform linear transformation processing on the input embedded representation.

[0036] In one possible implementation, the aforementioned input module is further configured to perform a nonlinear transformation on the representation vector, and use the transformed vector as a new representation vector.

[0037] In one possible implementation, the aforementioned processing module is specifically used to: obtain the predicted score of each vector in the representation vocabulary as the next vector based on the output features; and calculate the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0038] Fourthly, this application provides a cloud platform, including:

[0039] The transceiver module is used to receive representation vectors sent by the client. The representation vectors are obtained by the client inputting training data into the representation layer. The representation layer is used to obtain the vectors corresponding to the input data from the representation vocabulary stored by the client.

[0040] The input module is used to input representation vectors into the language model, obtain output features, and send the output features to the client.

[0041] The transceiver module is also used to receive the loss value sent by the client, which is calculated by the client based on the output features;

[0042] The update module is used to update the language model using the loss value, resulting in an updated language model.

[0043] In one possible implementation, the aforementioned representation vector is obtained by the client encrypting the embedded representation output by the representation layer.

[0044] Fifthly, this application provides an edge-cloud collaborative training method applied to an edge-cloud collaborative training system, which includes a client and a cloud platform. The method includes: the client using training data as input to a representation layer to obtain representation vectors, and sending the representation vectors to the cloud platform; the representation layer is used to obtain vectors corresponding to the input data from a representation vocabulary stored on the client; the cloud platform using the representation vectors as input to a language model to obtain output features, and sending the output features to the client; subsequently, the client uses the output features to calculate a loss value, updates its local representation layer based on the loss value, and sends the loss value to the cloud platform; the cloud platform uses the loss value to update the language model, obtaining the updated language model.

[0045] In the method provided in this application, the client side updates the representation layer locally, so that the representation word table (embedding table) on the client side does not leave the client. When sending data to the cloud platform, the extracted features, i.e. the hidden state, are sent. This can prevent the client data from being leaked and parsed into plaintext during transmission, thereby improving the privacy and security of the client data.

[0046] In one possible implementation, the aforementioned client uses training data as input to the representation layer to obtain a representation vector, including: the client inputs training data to the representation layer, outputs an embedded representation, and then encrypts the embedded representation to obtain a representation vector.

[0047] In this embodiment, when generating the representation vector, the client can encrypt the embedded representation, so that the data transmitted by the client to the cloud platform is encrypted data. Even if the data transmitted by the client to the cloud platform is leaked, it can be further prevented from being parsed as plaintext, thereby improving the privacy and security of the client's data.

[0048] In one possible implementation, the aforementioned client inputs training data into the representation layer and outputs embedded representations, including: the client encrypts the representation vocabulary to obtain an encrypted representation vocabulary, inputs the training data into the representation layer, and determines the embedded representations from the encrypted representation vocabulary through the representation layer.

[0049] In this embodiment, the client can encrypt the representation vocabulary and use the encrypted representation vocabulary to output the embedded representation. The output embedded representation is thus an encrypted representation, further preventing it from being parsed as plaintext when leaked, improving client privacy and security. Furthermore, even if the representation vocabulary used in the previous round of collaborative training on the client side is leaked, in the current collaborative training process, after encrypting the representation vocabulary used in the previous round, the output embedded representation can only be parsed based on the encrypted representation vocabulary. Therefore, it further prevents the embedded representation from being parsed as plaintext when leaked, improving client privacy and security.

[0050] In one possible implementation, the aforementioned client encrypts the representation vocabulary to obtain an encrypted representation vocabulary, including: the client adding a low-rank matrix to the representation vocabulary to obtain the encrypted representation vocabulary.

[0051] In this embodiment, a low-rank matrix (Lora) can be used as the specific encryption method for the representation vocabulary, so that the representation vocabulary can be fine-tuned based on the increased Lora during the training process, thereby achieving model fine-tuning.

[0052] In one possible implementation, the aforementioned client performs encryption processing on the embedded representation to obtain a representation vector, including: the client performs dimensionality reduction processing on the embedded representation to obtain a representation vector.

[0053] In this embodiment, the embedded representation can be dimensionality reduced to decrease the amount of data transmitted between the client and the cloud platform, thereby improving data transmission efficiency and training efficiency. Furthermore, dimensionality reduction of the embedded representation also effectively encrypts it, further preventing it from being parsed as plaintext when leaked, thus enhancing client privacy and security.

[0054] In one possible implementation, the aforementioned client performs dimensionality reduction processing on the embedded representation to obtain a representation vector, including: the client uses the embedded representation as input to a feedforward neural network (FNN) to obtain a representation vector, and the FNN is used to perform linear transformation processing on the input embedded representation.

[0055] Therefore, in this embodiment, an FNN can be used to output the representation vector, and the FNN can be used to achieve dimensionality reduction, thereby reducing the amount of data transmitted between the client and the cloud platform. This further prevents the embedded representation from being parsed as plaintext when it is leaked, improving client privacy and security.

[0056] In one possible implementation, the aforementioned method further includes: the client performing a nonlinear transformation on the representation vector and using the transformed vector as a new representation vector.

[0057] In this embodiment, the representation vector can be further subjected to nonlinear transformation and further encrypted to further prevent the embedded representation from being parsed as plaintext when it is leaked, thereby improving the privacy and security of the client.

[0058] In one possible implementation, the aforementioned client calculates the loss value using the output features, including: the client obtaining the predicted score of each vector in the representation vocabulary as the next vector based on the output features, and calculating the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0059] In this embodiment, the client can use the output features sent by the cloud platform to perform a prediction task, and then calculate the corresponding loss value based on the prediction branch of the prediction task to realize the model training process of end-to-cloud collaboration.

[0060] Sixthly, this application provides an edge-cloud collaborative training system, which is applied to an edge-cloud collaborative training system, comprising a client and a cloud platform;

[0061] The client is used to use training data as input to the representation layer to obtain representation vectors. The representation layer is used to obtain vectors corresponding to the input data from the representation vocabulary stored in the client.

[0062] The client is also used to send representation vectors to the cloud platform;

[0063] The cloud platform is used to take the representation vector as input to the language model and obtain the output features.

[0064] The cloud platform is also used to send output characteristics to the client;

[0065] The client is also used to calculate the loss value using the output features, update the representation layer based on the loss value, and send the loss value to the cloud platform;

[0066] The cloud platform is also used to update the language model using the loss value, resulting in an updated language model.

[0067] The effects achieved by the sixth aspect and any optional implementation of the sixth aspect can be found in the description of the fifth aspect or any optional implementation of the fifth aspect, and will not be repeated here.

[0068] In one possible implementation, the client is specifically configured to: input training data into the representation layer and output embedded representations; encrypt the embedded representations to obtain representation vectors.

[0069] In one possible implementation, the client is specifically configured to: encrypt the representation vocabulary to obtain an encrypted representation vocabulary; input training data into the representation layer, and determine the embedded representation from the encrypted representation vocabulary through the representation layer.

[0070] In one possible implementation, the client is specifically used to: add a low-rank matrix to the characterization vocabulary to obtain an encrypted characterization vocabulary.

[0071] In one possible implementation, the client is specifically used to perform dimensionality reduction on the embedded representation to obtain a representation vector.

[0072] In one possible implementation, the client is specifically used to take the embedded representation as input to the feedforward neural network (FNN) to obtain a representation vector, and the FNN is used to perform a linear transformation on the input embedded representation.

[0073] In one possible implementation, the client is also used to perform a nonlinear transformation on the representation vector and use the transformed vector as a new representation vector.

[0074] In one possible implementation, the client is specifically configured to: obtain the predicted score of each vector in the representation vocabulary as the next vector based on the output features; and calculate the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0075] In a seventh aspect, this application provides a computing device including a processor and a memory; the processor of the computing device is configured to execute instructions stored in the memory of the computing device to cause the computing device to perform method steps as performed by the client in the first aspect and any implementation thereof.

[0076] Eighthly, embodiments of this application provide a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs method steps as performed by the cloud platform in the second aspect and any implementation thereof.

[0077] Ninthly, embodiments of this application provide a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform a method as described in either the first or second aspect.

[0078] In a tenth aspect, embodiments of this application provide a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, enable the cluster of computing devices to perform a method as described in either the first or second aspect.

[0079] Eleventhly, embodiments of this application provide a chip, including at least one processor and an interface; at least one processor obtains program instructions or data through the interface; at least one processor is used to execute program line instructions to implement the method in any implementation of the first aspect or the second aspect. Attached Figure Description

[0080] Figure 1 This application provides an architectural diagram of an edge-cloud collaborative training system.

[0081] Figure 2 A flowchart illustrating an edge-cloud collaborative training method provided in this application;

[0082] Figure 3 A flowchart illustrating another edge-cloud collaborative training method provided in this application;

[0083] Figure 4 A flowchart illustrating another edge-cloud collaborative training method provided in this application;

[0084] Figure 5 A flowchart illustrating another edge-cloud collaborative training method provided in this application;

[0085] Figure 6 A flowchart illustrating another edge-cloud collaborative training method provided in this application;

[0086] Figure 7 A flowchart illustrating another edge-cloud collaborative training method provided in this application;

[0087] Figure 8 A schematic diagram of another edge-cloud collaborative training system provided in this application;

[0088] Figure 9 A schematic diagram of a client-side structure provided in this application;

[0089] Figure 10 This application provides a schematic diagram of the structure of a cloud platform;

[0090] Figure 11 A schematic diagram of another client structure provided in this application;

[0091] Figure 12 A schematic diagram of another cloud platform provided for this application. Detailed Implementation

[0092] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0093] First, the embodiments of this application involve the application of neural networks and natural language processing (NLP). In order to better understand the solutions of the embodiments of this application, the relevant terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.

[0094] (1) Neural Network

[0095] Neural networks can be composed of neural units, which can refer to units represented by x. s The arithmetic unit that takes input data and an intercept of 1 as input can output the following:

[0096]

[0097] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x s The weight parameters are denoted by b, which represents the bias of the neural unit. f is the activation function of the neural unit, used to introduce non-linear characteristics into the neural network to convert the input signal into the output signal. The output signal of this activation function can be used as the input to the next convolutional layer; the activation function can be the sigmoid function. A neural network is a network formed by connecting multiple individual neural units, meaning the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from the local receptive field, which can be a region composed of several neural units.

[0098] The work of each layer in a neural network can be described by the mathematical expression y = a(Wx + b). From a physical perspective, the work of each layer in a neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are performed by Wx, operation 4 by +b, and operation 5 by a(). The term "space" is used here because the objects being classified are not individual things, but a class of things, and space refers to the set of all individuals of this class of things. Here, W is the weight vector, and each value in this vector represents the weight value of a neuron in that layer of the neural network. This vector W determines the spatial transformation from the input space to the output space, that is, the weight W of each layer controls how the space is transformed. The purpose of training a neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (a weight matrix formed by the vectors W of many layers). Therefore, the training process of a neural network is essentially about learning how to control the transformation space, and more specifically, learning the weight matrix.

[0099] (2) Loss Function

[0100] In training deep neural networks, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value and update the weight vector of each layer based on the difference. (Of course, there's usually a pre-configuration process before the first update, where parameters are pre-configured for each layer.) For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network can predict the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training the deep neural network becomes a process of minimizing this loss. Common loss functions include mean squared error, cross-entropy, logarithmic, and exponential loss functions. For example, mean squared error can be used as the loss function, defined as... The specific loss function can be selected based on the actual application scenario.

[0101] (3) Backpropagation algorithm

[0102] An algorithm for calculating the gradient of model parameters based on a loss function and updating the model parameters. Neural networks can use backpropagation (BP) to correct the initial parameter values ​​during training, thus reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates error loss; this error loss information is then propagated back to update the parameters of the initial neural network model, thereby converging the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the neural network model, such as the weight matrix.

[0103] In the embodiments of this application, the BP algorithm can be used to train the model in either the training or inference stage of the large language model to obtain the trained model.

[0104] (4) Language model (LM)

[0105] Used in Natural Language Processing (NLP), it plays a crucial role in NLP, its task being to predict the probability of a sentence appearing in a language. For example, a language model is typically constructed as a probability distribution p(s) of string s, where p(s) attempts to reflect the frequency of string s as a sentence. It can be applied to scenarios such as text recognition or machine translation.

[0106] This technology can be applied to scenarios such as text recognition or machine translation.

[0107] (5) Neural Machine Translation: Neural machine translation is a typical task in natural language processing. This task involves outputting the corresponding sentence in the target language, given a sentence in the source language. In commonly used neural machine translation models, words in both the source and target language sentences are encoded as vector representations. The relationships between words and sentences are then calculated in the vector space to perform the translation task.

[0108] (6) Self-attention model refers to the effective encoding of a sequence of data (such as the natural language corpus "Your phone is very good.") into several multi-dimensional vectors to facilitate numerical calculations. The multi-dimensional vectors integrate the similarity information between each element in the sequence, and this similarity is called self-attention.

[0109] (7) Pre-trained Language Model (PLM): A PLM is a natural language sequence encoder that encodes each word in a natural language sequence into a vector representation for downstream tasks, such as word vector transformation networks, action vector transformation networks, or action generation networks mentioned below in this application. PLM training consists of two phases: pre-training and fine-tuning. In the pre-training phase, the model is trained on large-scale unsupervised text to learn word representations. In the fine-tuning phase, the model is initialized using the parameters learned in the pre-training phase and then trained on downstream tasks such as text classification or sequence labeling with fewer steps, successfully transferring the semantic information obtained from pre-training to downstream tasks.

[0110] In addition, fine-tuning can also include full fine-tuning, which means that all parameters of the language model can be adjusted.

[0111] (8) transformer

[0112] A transformer structure is a feature extraction network that includes both an encoder and a decoder (classified as a convolutional neural network). Of course, in some cases, a transformer structure may not include an encoder but may include a decoder.

[0113] Encoder: Learns features, such as pixel features, in the global receptive field through self-attention.

[0114] Decoder: Learns the features of the desired modules, such as the features of the output box, through self-attention and cross-attention.

[0115] The following is a description of attention (also known as the attention mechanism):

[0116] Attention mechanisms can quickly extract important features from sparse data. Attention occurs between the encoder and decoder, or more specifically, between the input and generated sentences. In contrast, the self-attention mechanism in a self-attention model occurs within the encoding matrix or the output sequence, extracting connections between distant words within the same sentence, such as syntactic features (phrase structure). Self-attention provides an effective modeling method for capturing global contextual information through QKV (key-value pairs). Assuming the input is Q (query), and the context is stored as key-value pairs (K, V), then the attention mechanism is essentially a mapping function from the query to a series of key-value pairs (key, value). The essence of the attention function can be described as a mapping from a query to a series of (key-value) pairs. Attention essentially assigns a weight coefficient to each element in the sequence, which can also be understood as soft addressing. If each element in the sequence is stored in (K, V) form, then attention performs addressing by calculating the similarity between Q and K. The similarity calculated between Q and K reflects the importance of the extracted V values, i.e., the weights, and then a weighted sum is obtained to obtain the final feature value.

[0117] Attention calculation mainly consists of three steps. The first step is to calculate the similarity between the query and each key to obtain weights. Common similarity functions include dot product, concatenation, and perceptron. The second step typically uses a softmax function (which can normalize the weights, resulting in a probability distribution where the sum of all weight coefficients is 1, and also highlights the weights of important elements) to normalize these weights. Finally, the weights and their corresponding key values ​​are weighted and summed to obtain the final feature value. The specific calculation formula is as follows:

[0118]

[0119] Where d is the dimension of matrix QK.

[0120] Furthermore, attention includes self-attention and cross-attention. Self-attention can be understood as a special type of attention where the inputs to the QKV features are consistent. Cross-attention, on the other hand, involves inconsistent inputs to the QKV features. Attention integrates the queried features as updated values ​​for the current features using the similarity between features (e.g., inner product) as weights. Self-attention is attention extracted based on the attention drawn from the feature map itself.

[0121] For convolutional networks, the kernel size limits the receptive field, often requiring multiple layers to focus on the entire feature map. Self-attention, on the other hand, offers the advantage of global focus; it can acquire global spatial information of the feature map through simple queries and assignments. A unique aspect of self-attention in query-key-value (QKV) models is that the inputs for each QKV value are consistent.

[0122] (9) Large Language Model (LLM)

[0123] A large language model (LLM) refers to a language model containing hundreds of billions (or more) parameters trained on massive amounts of text data. It is a natural language processing model based on deep learning. These models can process large amounts of text data to learn the grammatical and semantic rules of natural language. LLMs can be applied to text generation, machine translation, question answering systems, text summarization, and sentiment analysis, offering advantages such as strong generative capabilities, high adaptability, accurate prediction, and strong scalability. For example, in movie recommendation scenarios, a large language model can generate descriptions of movie scenes, including genre, main actors, and plot, enabling the system to better recommend similar movies. Large language models can also generate recommendation reasons; for example, e-commerce websites can use large language models to generate reasons for recommending products, such as product quality, price, and features, allowing users to better understand the value of the products. Furthermore, by connecting to a candidate list provided by a traditional recommendation system, large language models can also sort and summarize recommendation results, generating results that are easy for users to read and browse.

[0124] (10)Embedding

[0125] Embedding refers to the feature representation of a sample or word embedding representation.

[0126] (11)Embedding table

[0127] An embedding table can be understood as a collection of word embedding representations.

[0128] The application scenarios and architecture provided in this application are described below.

[0129] Large language models, such as OpenAI's GPT series, have become a hot topic in the tech world due to their superior natural language processing capabilities. These models not only support applications like chatbots and automated writing but also demonstrate powerful performance in multiple fields such as text summarization and sentiment analysis. Fine-tuning strategies for pre-trained models are also gradually becoming a mainstream approach. This involves pre-training models on large amounts of text data and then fine-tuning them on domain-specific data to optimize the model for specific tasks. Against this backdrop, edge-cloud / cloud-cloud collaboration technologies have emerged. Edge-cloud collaboration maximizes the computing power of different devices and platforms, enabling flexible deployment and collaborative computing of data and models in the cloud and on the edge, thereby expanding the application scenarios of AI and greatly optimizing the user experience. For example, lightweight models running on IoT devices and mobile devices can respond to user requests in real time, while complex model training and inference tasks can be offloaded to the cloud, achieving efficient utilization of computing resources. However, this collaborative computing model also brings new challenges to data security and privacy protection. How to ensure data security while achieving efficient model training and inference has become a focus of industry attention.

[0130] In edge-cloud collaborative environments, privacy protection is particularly critical. Data invisibility protection strategies ensure the security and privacy of user data during storage, processing, and transmission, using technologies such as encryption to prevent unauthorized theft or misuse during transmission. Model invisibility privacy protection is primarily based on commercial protection needs. For example, large models like the GPT series typically only provide API service calls, while their internal model structure, parameters, training and fine-tuning processes are strictly confidential. This is not only to protect the intellectual property rights of the model but also to prevent potential malicious attacks and ensure the stability and security of the model service.

[0131] Regarding model privacy, when the model's backbone is hosted in the cloud for training, the model can be considered invisible to the user. The key question then becomes: how to protect user data used for fine-tuning from leakage and ensure it is not parsed in plaintext on the cloud side?

[0132] In some existing solutions, such as differential privacy, a privacy protection framework is established by introducing randomness to protect individual information, addressing client-side privacy concerns. The core idea is to add random noise during data publishing or querying to ensure that the addition or removal of a single data record does not significantly affect the output, thus protecting individual privacy. However, since noise is added only at the client-side data level and the embedding layer weights from the cloud side are reused, the influence of plaintext cannot be eliminated. Furthermore, in this scenario, the client-side does not participate in training, thus failing to leverage its computing power.

[0133] For example, homomorphic encryption (HE), as a unique encryption technique, has the core advantage of performing computations on encrypted data without prior decryption. This technology is not only theoretically attractive but also shows great potential in practical applications. For instance, in fine-tuning training scenarios, the edge device can synchronize data to the cloud via homomorphic encryption, and the cloud device can directly train on the encrypted data and transmit the computation results (still encrypted) to the edge device. After decryption, the edge device can obtain the final computation result. However, compared to traditional encryption methods, homomorphic encryption introduces higher computational complexity; the decryption process also incurs additional computational overhead; in deep learning, especially in Transformer architecture models, most operators cannot achieve homomorphic encryption; and multiple computations on encrypted data may lead to error accumulation. Furthermore, while applying homomorphic encryption to large language models, i.e., Transformer architecture models, is theoretically feasible, it faces challenges and limitations in practical applications, such as computational efficiency and model accuracy. This determines that this solution is difficult to meet the needs of most user scenarios.

[0134] For example, federated learning (FL) is essentially an encrypted distributed machine learning technique. It can span multiple devices, allowing participating parties to collaboratively build models without disclosing the underlying data or its encrypted (obfuscated) form. Encryption mechanisms enable data exchange between companies without leaving their local machines, achieving the construction of a shared model without violating data privacy. However, in language model fine-tuning training, if federated learning is used, an attacker who obtains the embedding parameters can parse the token into plaintext, thereby leaking user data. Furthermore, in large language models, due to model size limitations, it's difficult for the edge devices to fully load the model. Treating the edge devices as part of the computing power in federated learning scenarios, and sharing the model across all edge devices, contradicts the principle of model visibility to users. In other words, while federated learning is a popular privacy protection method, when dealing with large language models, model parameters are shared among all participating edge devices. This contradicts vendors' requirements for closed-source large language models, especially when the model should not be visible to users. Additionally, if an attacker obtains the model's embedding parameters, there is a risk of parsing the token into plaintext and leaking user data.

[0135] Fine-tuning language models, especially large language models, requires significant computing power. Furthermore, due to the closed-source nature of these models, many users fine-tune their collected data using models and computing power provided by cloud vendors. Users typically want their data to remain private and unused, and they do not wish for it to be parsed in plaintext on the cloud.

[0136] Therefore, to address the aforementioned shortcomings, this application provides an edge-cloud collaborative training method that achieves a strategy that protects the model architecture from visibility while ensuring the cloud side's invisibility of user data in an edge-cloud collaborative environment. The method provided in this application decouples key components of the model, such as the embedding layer, from the cloud side and migrates them to the edge side for training. Furthermore, it adds additional trainable parameters, aiming to ensure that the data output from the network on the edge side cannot be reverse-analyzed by the cloud side without affecting model accuracy, thus achieving a certain degree of dual protection for both the model and the data.

[0137] First, the architecture of the edge-cloud collaborative training system provided in this application can be as follows: Figure 1 As shown, the cloud-edge collaborative training system may include a cloud platform 11 and a client 12.

[0138] The cloud platform 11 and the client 12 can be connected via wired or wireless means.

[0139] The cloud platform 11 may specifically include a server cluster with storage and processing capabilities. In this embodiment, the cloud platform 11 can deploy a transformer model, such as a language model or a large language model, and can be used to receive data sent by the client. In this embodiment, the data sent by the client may include at least two types: one is the representation vector output by the client, which the cloud platform uses as input to the transformer and returns the output features to the client; the other is the loss value uploaded by the client, which the cloud platform uses to update the transformer to achieve edge-cloud collaborative training.

[0140] Client 12 can achieve collaborative training of the model by interacting with cloud platform 11. This client can specifically include, but is not limited to, personal computers, computer workstations, smartphones, tablets, laptops, and smart cars. In this embodiment, to avoid data leakage on the client side, an embedding table can be deployed on the client side. An embedding layer can be deployed in client 12 to extract corresponding representation vectors from the embedding table and transmit them to the cloud platform. After receiving the output features from the cloud platform, a prediction result can be obtained based on these features, and a loss value can be calculated. This loss value is then fed back to the cloud platform, allowing the cloud platform to train the model based on the loss value, thus achieving end-to-cloud collaborative training.

[0141] Based on the aforementioned edge-cloud collaborative training architecture, the method and process provided in this application will be described below.

[0142] See Figure 2 The flowchart of an edge-cloud collaborative training method provided in this application is as follows.

[0143] 201. The client uses the training data as input to the representation layer to obtain the representation vector.

[0144] The training data can be data from a training set, which includes data used for model training. This training set can be divided into training data and corresponding labels (or ground truth values). Specifically, the training data can include text, speech, images, or point clouds, etc.

[0145] The representation layer, or model structure deployed on the client side, can be used to query the representation corresponding to the training data from the embedding table and obtain the representation vector based on the query result. The representation vector can be obtained by either using the queried embedding representation as the source or by processing the queried embedding representation.

[0146] Specifically, to improve the data privacy and security of the client, the embedded representation obtained from the Embedding table can be encrypted, and the encrypted embedded representation can be used as the representation vector. Various encryption methods can be used, which will be described below.

[0147] Method 1: Encrypt in the Embedding table.

[0148] In one possible implementation, the client can encrypt the embedding table to obtain an encrypted embedding table; input the training data into the representation layer, and determine the embedding representation from the encrypted representation vocabulary through the representation layer.

[0149] Therefore, in this embodiment, encryption can be performed on the embedding table. When outputting the embedding representation, the embedding representation can be output based on the encrypted embedding table, thereby avoiding the situation where the embedding representation is lost and is parsed as plaintext corresponding to the embedding table, thus improving the data security on the client side.

[0150] Optionally, to enable fine-tuning of the model on the client side, a low-rank matrix can be added to the representation vocabulary to obtain an encrypted representation vocabulary. For example, Low-Rank Adaptation (LoRA) can be introduced, using a low-rank matrix in the embedding table for encryption, so that the embedding table after the low-rank matrix is ​​introduced can be fine-tuned later.

[0151] Method 2: Dimensionality Reduction

[0152] In one possible implementation, the client can perform dimensionality reduction on the embedded representation to obtain a representation vector.

[0153] Optionally, dimensionality reduction can be implemented using a feedforward neural network (FNN). For example, the client inputs the embedded representation to the FNN to obtain a representation vector, and the FNN performs a linear transformation on the input embedded representation.

[0154] In this embodiment, the embedding representation can be dimensionality reduced. This reduces communication overhead and changes the embedding representation, thereby preventing it from being parsed as plaintext corresponding to the embedding table when the embedding representation is lost, thus improving data security on the client side.

[0155] Method 3: Nonlinear Transformation

[0156] In one possible implementation, the client performs a nonlinear transformation on the representation vector and uses the transformed vector as the new representation vector.

[0157] In this embodiment, a nonlinear transformation can be further performed, which helps to capture complex patterns in the input data and avoids parsing the data as plaintext corresponding to the Embedding table when the embedded representation is lost, thereby improving the data security on the client side.

[0158] 202. The client sends the representation vector to the cloud platform.

[0159] Once the client obtains the representation vector, it can send the representation vector to the cloud platform so that the cloud platform can perform the corresponding natural language processing task based on the representation vector.

[0160] 203. The cloud platform uses the representation vector as input to the language model to obtain the output features.

[0161] After receiving the representation vector from the client, the cloud platform uses the representation vector as input to the language model deployed locally on the cloud platform to obtain the output features.

[0162] Specifically, this language model can be used for natural language processing, such as predicting the probability of a corpus appearing in a language, or the probability of each word being the next word. For example, if a user's input text includes "The weather is really nice today," where the keywords can be divided into "today," "weather," and "really nice," then the language model can be used to identify whether the semantic composition of the input text is "The weather is really nice today" or "The weather is really nice today," in order to identify the semantics in the input data.

[0163] In one possible implementation, the language model can specifically be a large language model, thus possessing very strong processing capabilities. Typically, large language models have very high performance, but they are also highly complex and require significant computing power from the device. Therefore, large language models can be deployed on a cloud platform to utilize the cloud platform for operation.

[0164] 204. The cloud platform sends output characteristics to the client.

[0165] After obtaining the output features through the language model on the cloud platform, the output features can be sent to the client.

[0166] In this embodiment, the cloud platform and the client transmit a hidden state, eliminating the need to transmit plaintext, which greatly improves the data security between the client and the cloud platform.

[0167] 205. The client calculates the loss value using the output features, updates the representation layer based on the loss value, and sends the loss value to the cloud platform.

[0168] After receiving the output features, the client can use the output features to calculate the loss value, then use the loss value to update the model structure deployed locally on the client, and send the loss value to the cloud platform so that the cloud platform can use the loss value to update the language model.

[0169] Specifically, the client can obtain the predicted score of each vector in the representation vocabulary as the next vector based on the output features; in other words, the probability of each word being the next word. The client then calculates the loss value based on the predicted score of each vector as the next vector and the corresponding label in the training data. In other words, the client can use the output features sent by the cloud platform to calculate the prediction result and the difference between the prediction result and the true value to obtain the loss value.

[0170] 206. The cloud platform uses the loss value to update the language model, resulting in the updated language model.

[0171] After receiving the loss value sent by the client, the cloud platform can use the loss value to update the language model in reverse, and obtain the updated language model.

[0172] In this embodiment, in the edge-cloud collaborative training scenario, the embedding table is deployed on the client side. Non-plaintext data is transmitted between the client and the cloud platform, preventing the client's embedding table from being decoded into plaintext if it is leaked. This improves the security of data transmission between the client and the cloud platform, and enhances the client's data privacy. Furthermore, it fully utilizes cloud computing resources while protecting privacy. By performing preliminary data processing and model training on the edge, sensitive data is ensured to be transformed into a hidden state that cannot be directly parsed before leaving the local device. This effectively avoids the direct exposure of sensitive data in the cloud.

[0173] The foregoing has provided an overview of the method provided in this application. The following section will further describe the method flow provided in this application in conjunction with specific application scenarios.

[0174] First, the method provided in this application can include multiple parts, which can be understood as deploying an embedding table on the client side, stripping the embedding table, and further encrypting the transmitted data during client-side training. For ease of understanding, the method provided in this application will be described in several parts below. Specifically, it can be divided into embedding layer stripping training, embedding table encryption, dimensionality reduction processing, etc., which will be described separately below.

[0175] I. Embedding Layer Stripping Training

[0176] See Figure 3 The flowchart of another edge-cloud collaborative training method provided in this application is shown below.

[0177] 301. Input the Token into the Embedding layer.

[0178] The token can be a sequence of words in the text corresponding to the input data.

[0179] The input data can specifically be text, images, voice, or other types of input data entered by the user. After obtaining the input data, it can be preprocessed. For example, if the input data includes images, image recognition can be performed on the images to obtain one or more keywords (or key words) corresponding to the images, thus obtaining a token; if the input data includes voice data, speech recognition can be performed on the voice data to obtain one or more keywords (or key words) corresponding to the speech, thus obtaining a token, and so on; if the input data includes text data, the text data can be divided into one or more keywords (or key words) to obtain a token, and so on.

[0180] For the Embedding layer: In natural language processing (NLP), the embedding layer is a fundamental and crucial component. Its main function is to map discrete words to a continuous vector space, enabling neural networks to process text data more efficiently. The embedding layer can be understood as a lookup table that maps each word or vocabulary (usually represented as a one-hot encoding) to a fixed-length dense vector. These vectors are trainable parameters that learn the semantic representation of each word in the vector space during model training.

[0181] 302. The Embedding layer uses the Embedding table to obtain the Embedding vector and uploads it to the cloud platform.

[0182] The Embedding layer uses the Embedding table stored locally on the client to obtain the Embedding vector, which is the representation vector.

[0183] For example, when the model begins training, the input token_ids represent unique identifiers for each word or character in the input text, typically indices from the vocabulary. These token IDs, as input to the model, are first converted into corresponding vector representations through an embedding layer.

[0184] The mathematical representation of the Embedding layer is:

[0185] v i =E token_ids[i]

[0186] Where E is the embedding table, token_ids[i] is the index of the i-th token, and v i These are the corresponding token embeddings. That is, the Embedding Table provides a multi-dimensional vector representation for each unique token ID and outputs the token embedding.

[0187] In this embodiment, the client-side embedding table is stored only locally on the client, while the client needs to upload the embedding vector to the cloud platform, which is equivalent to a hidden state, thereby preventing the embedding table from being transmitted outside the client and improving the client's data security.

[0188] 303. The cloud outputs feature vectors through the transformer layer and sends them to the client.

[0189] The embedded vectors are fed into the Transformer layer in the cloud. The Transformer layer processes sequential data based on a self-attention mechanism, and each layer can capture the complex relationships between different tokens within the sequence. In the Transformer layer, the embedded vectors are transformed into new vector representations, namely feature vectors, or Token States, through a series of self-attention and feedforward networks. Simultaneously, both the input and output of the Transformer layer can be considered as hidden states, which are deep semantic representations of the tokens, rather than plaintext features.

[0190] 304. The client uses the feature vector to calculate Logits.

[0191] After processing by the Transformer layer, the final Hidden States are transformed into Logits through the output layer (usually a linear layer) deployed on the client side. Logits can be understood as a score prediction for each input Token as the next Token.

[0192] For example, Logits can be calculated as follows:

[0193] logits = EH

[0194] Where E is the Embedding Table and H is the Hidden States output by the Transformer, i.e., the output features.

[0195] Finally, the logits can be further processed using the Softmax function to obtain the probability distribution of the next token for each token.

[0196] 305. The client calculates the loss value based on Logits and uses this loss value for reverse updates.

[0197] The client then compares the result with the actual next token label and calculates the loss value. Specifically, cross-entropy loss (Loss) can be calculated. This loss value is used to update the model parameters, including the Embedding Table and the weights of the Transformer layers, during training via backpropagation.

[0198] Specifically, when performing a reverse update of the model, it is possible to update the model deployed locally on the client side, as well as the transformer layer deployed on the cloud platform. The client can send the loss value to the cloud platform, or feed back the gradient calculated based on the loss value to the cloud platform, so that the cloud platform can update the transformer layer based on the data fed back by the client, thereby realizing model updates in the cloud.

[0199] In some scenarios, during the training of traditional large language models, the embedding layer is responsible for mapping discrete input symbols (such as words or characters) to a continuous vector space. These vectors are then used as input to the model and passed to subsequent Transformer layers for processing. However, this approach carries privacy risks because the raw data needs to be transmitted to a cloud server, where it may be intercepted or illegally accessed.

[0200] In the method provided in this application, the embedding table is deployed only locally on the client side. Hidden states are transmitted between the client and the cloud platform, rather than in plaintext, thus avoiding data leakage on the client side and improving data security. The embedding table can be understood as the first gateway for processing text data in NLP models. By converting sensitive text into abstract vector representations, it plays a crucial role in protecting user privacy and data security. This conversion is particularly important in cloud-based collaborative training because it allows for the protection of user data privacy without sacrificing computational power and model performance. Through the method provided in this application, the embedding layer is stripped to the client side for training. That is, in the end-to-cloud collaborative training architecture provided in this application embodiment, the embedding table resides locally on the client (i.e., the user's device) for generating and updating embedding vectors. This means that sensitive raw data does not need to be transmitted to the cloud and can be protected on the client side without transmitting plaintext externally, providing very strong client privacy protection.

[0201] Therefore, the method provided in this application separates the embedding layer and other components from the cloud side to the edge side for training, while adding additional trainable parameters. This ensures that the hidden state output by the network on the edge side cannot be reverse-engineered by the cloud side without affecting model accuracy, thus protecting client data privacy and security.

[0202] II. Encryption of Embedding Table

[0203] Based on the aforementioned training of the Embedding layer stripping, when transmitting data between the client and the cloud platform, the data to be transmitted can be further encrypted to improve the data security of the client.

[0204] For example, see Figure 4 The flowchart of another edge-cloud collaborative training method provided in this application is shown below.

[0205] As mentioned above Figure 3 The difference lies in the fact that, during the generation of the embedding vector, the embedding table can be encrypted, such as... Figure 4 Step 40 in the process generates an Embedding vector based on the encrypted Embedding table.

[0206] In most current scenarios, users choose to incrementally fine-tune large models pre-trained by cloud vendors. In the edge-cloud collaborative training scenario provided in the application embodiment, the models deployed on the client and cloud platform can also be pre-trained models, and the cloud side also holds an initial Embedding Table. During incremental fine-tuning training, the Embedding Table may not be updated with parameters, meaning the Embedding Table remains the same on both the cloud and client sides. Once the cloud-side Embedding Table is leaked, attackers can directly use it to reverse-parse Hidden States during training into plaintext, thereby obtaining the user's personal data used for fine-tuning training, resulting in a privacy breach.

[0207] Therefore, in this application embodiment, considering the problem that the Embedding Table is not updated during training, the present invention proposes an encryption module based on a method of efficient parameter fine-tuning to ensure that the Embedding Table is updated during edge training, thereby forming a difference from the original Embedding Table on the cloud side.

[0208] like Figure 4 As shown, during the training of the embedding layer on the edge, LoRA can be introduced to generate and adjust the embedding vectors, such as adding extra layers or expanding or reducing the vocabulary. LoRA is a parameter-efficient model fine-tuning method that allows fine-tuning without significantly increasing the model size. It is achieved by adding low-rank updates to the original model weights, which is highly efficient compared to the full model parameters.

[0209] The method of using LoRA to update the Embedding Table can be represented as:

[0210] E′=E+ΔE

[0211] ΔE=LoRA A ×LoRA B

[0212] LoRA A and LoRA BThese are low-rank matrices, and their product ΔE provides an update to the original Embedding Table E. This allows the model to be fine-tuned while maintaining pre-trained parameters, rather than training the entire Embedding Table from scratch. ΔE can be considered an encryption component, the effectiveness of which depends on the distribution of training data on the endpoint and the number of training epochs. This ensures that the encryption component is differentiated across different user data during training. Furthermore, this encryption component is quite sensitive; different hyperparameters, training data, and model structures will all affect it. This significantly ensures that user data is not leaked during fine-tuning training based on the pre-trained model.

[0213] When the client iterates and updates the model, ΔE can also be updated. This ensures that as the number of iterations increases, the E on the client side remains different from the Embedding Table on the cloud side. Therefore, even if the Embedding vectors transmitted between the client and the cloud platform are leaked, the Embedding Table on the client side can be prevented from being parsed into plaintext, improving data security on the client side. Therefore, in this embodiment, trainable parameters are added to the Embedding layer, specifically by adjusting the original Embedding table using the LoRA (Lower Rank Matrix Factorization) method. This method improves the efficiency of model training and reduces the computational requirements on the client side.

[0214] III. Dimensionality Reduction

[0215] Furthermore, based on the aforementioned training by stripping the Embedding layer, dimensionality reduction processing can be added.

[0216] For example, see Figure 5 The flowchart of another edge-cloud collaborative training method provided in this application is shown below.

[0217] As mentioned above Figure 3 The difference lies in the fact that, during the generation of the embedding vector, the embedding table can be subjected to dimensionality reduction processing, such as... Figure 5 Step 50 shown can reduce the amount of data transmitted and encrypt the transmitted data to improve the security of data transmission between the client and the cloud platform.

[0218] While ensuring user data privacy during edge-cloud collaborative training, training performance is often paramount. If using a separate training architecture results in additional training time—sometimes several times longer than training directly on the cloud—this directly impacts the overall end-to-end training time, thus degrading the user's training experience.

[0219] In some common scenarios, large models often have 10B-500B parameters, where the word vector size after embedding table calculation is [vocab_size, embedding_size]. Typically, vocab_size is 10. 4 ~10 6 The embedding size is 10. 3 ~10 5 This means that word vectors stored in FP32 will reach 10KB to 100MB. This implies that the communication overhead between the edge and the cloud will be difficult to ignore, especially when the communication bandwidth is limited. The communication time will far exceed the training time per step, making the communication time the bottleneck in the entire training process.

[0220] In the method provided in this application embodiment, the vector extracted from the Embedding Table can be dimensionality reduced, thereby reducing communication overhead and also enabling data encryption to improve the data privacy and security of the client.

[0221] like Figure 5 As shown, the vector then passes through an FFN layer. A typical FFN layer usually includes linear transformations; for example, "down-projection" usually refers to the first linear transformation, which projects the high-dimensional representation of each token into a lower-dimensional space. This process can be understood as a form of data compression, reducing the amount of information required to represent each token. In the method provided in this application embodiment, the dimensionality reduction effect of the FFN layer is utilized. The dimensionality reduction method is expressed as follows:

[0222] z = W down H+b down

[0223] Where H is the original word vector, and W down It is a dimension-reduced weight matrix, b down By biasing the vector, we obtain the dimensionality-reduced vector z. Different degrees of dimensionality reduction can be achieved by adjusting the size of the dimensionality reduction weight matrix.

[0224] This method of dimensionality reduction, achieved by adding an MLP (Multilayer Perceptron) layer to the original embedding table, effectively reduces the amount of data required for communication, thereby reducing communication overhead during edge-cloud collaborative training. This optimizes the overall training time and improves the user experience. The dimensionality reduction method provided in this application can improve training efficiency while protecting privacy.

[0225] Furthermore, in the aforementioned Figure 5 On this basis, nonlinear processing can be added, such as Figure 6 As shown, after dimensionality reduction, the resulting representation vector can be further subjected to nonlinear transformation, i.e., step 60. That is, after the "Down-project" step, an activation function such as ReLU can be added. This entire process allows the FFN layer to add nonlinear processing while reducing data dimensionality, which helps to capture complex patterns in the input data.

[0226] The aforementioned components can also be implemented in combination, such as... Figure 7 As shown, by combining encryption and dimensionality reduction, the encryption effect is guaranteed while training performance is maintained to a certain extent. This makes the entire training scheme effectively acceptable and usable by users. Specifically, the client-side includes an Embedding layer responsible for converting input data into word vector representations. When generating the Embedding Table, for data encryption, the word vectors need to pass through an encryption module, i.e., a parameter-efficient fine-tuning layer, which is equivalent to adding additional trainable parameters. Subsequently, the word vectors undergo dimensionality reduction through a Feed Forward down-project layer to reduce data transmission volume. Then, a Nonlinearity layer introduces nonlinearity to enhance the model's expressive power and ensure that the model's accuracy does not decrease. The Transformer layer on the cloud side processes the word vectors from the client-side, generates prediction scores (logits), and calculates the error (loss) for gradient sharing in the model. The overall architecture design ensures that sensitive data is transmitted to the cloud side after being converted to a hidden state. Dimensionality reduction and encryption parameter updates / encryption ensure data security and prevent the recovery of original information, achieving effective protection of user privacy while fully utilizing cloud computing resources.

[0227] Therefore, in this embodiment, the edge-cloud collaborative and separate training structure enables preliminary data processing and embedding vector generation on the local device, while the complex model training process is completed in the cloud, thus protecting user data privacy without sacrificing training effectiveness. Simultaneously, encryption ensures data security even when data needs to be transmitted between the local device and the cloud, reducing the risk of data leakage and improving the overall security of the model training process. Furthermore, dimensionality reduction significantly reduces the communication overhead during end-to-end training, effectively alleviating communication bottlenecks, especially under bandwidth constraints, and accelerating model training and parameter updates.

[0228] Furthermore, in order to facilitate the description of the effects achieved by the method provided in the embodiments of this application, the effects achieved by the method provided in this application will be compared and described below with the training effects of existing solutions in more specific application scenarios.

[0229] The accuracy results of the current experiment are shown in Table 1.

[0230]

[0231] Table 1

[0232] By adding parameters to the embedding layer, the two downstream tasks can achieve experimental accuracies of 80% and 92% respectively, which are very close to the effect of full parameter fine-tuning of the entire model (84% and 94%). At the same time, the encryption effect experiment shows that the accuracy of cloud testing back-inferring user data uploaded to the end side is only 1%, which means that 99% of user content data cannot be back-inferred by cloud testing, which greatly improves the encryption effect of large model end-cloud collaborative fine-tuning.

[0233] Therefore, the method provided in this application embodiment can, to a certain extent, ensure user data security in the context of current large-scale model edge-cloud collaborative fine-tuning, while also avoiding the risk of closed-source model leakage; due to the trainability of parameters added to the embedding layer, it ensures good fine-tuning on different types of data, that is, it ensures the stability of model accuracy; in addition, it can also reduce the dimensionality of the embedding vector, reduce the communication overhead between edge and cloud, and improve training efficiency.

[0234] The foregoing has described the method steps provided in this application. The following describes the system and apparatus for performing the method steps provided in this application.

[0235] See Figure 8 This application provides a schematic diagram of another edge-cloud collaborative training system, which includes a cloud platform 82 and a client 81.

[0236] Client 81 is used to use training data as input to the representation layer to obtain representation vectors. The representation layer is used to obtain vectors corresponding to the input data from the representation vocabulary stored in client 81.

[0237] Client 81 is also used to send representation vectors to cloud platform 82;

[0238] The cloud platform 82 is used to take the representation vector as the input of the language model and obtain the output features;

[0239] The cloud platform 82 is also used to send output features to the client 81;

[0240] Client 81 is also used to calculate the loss value using the output features, update the representation layer based on the loss value, and send the loss value to cloud platform 82;

[0241] The cloud platform 82 is also used to update the language model using the loss value, resulting in an updated language model.

[0242] In one possible implementation, client 81 is specifically used to: input training data into the representation layer and output embedded representations; encrypt the embedded representations to obtain representation vectors.

[0243] In one possible implementation, client 81 is specifically used to: encrypt the representation vocabulary to obtain an encrypted representation vocabulary; input training data into the representation layer, and determine the embedded representation from the encrypted representation vocabulary through the representation layer.

[0244] In one possible implementation, client 81 is specifically used to: add a low-rank matrix to the characterization vocabulary to obtain an encrypted characterization vocabulary.

[0245] In one possible implementation, client 81 is specifically used to perform dimensionality reduction on the embedded representation to obtain a representation vector.

[0246] In one possible implementation, client 81 is specifically used to take the embedded representation as input to a feedforward neural network (FNN) to obtain a representation vector, and the FNN is used to perform a linear transformation on the input embedded representation.

[0247] In one possible implementation, client 81 is also used to perform a nonlinear transformation on the representation vector and use the transformed vector as a new representation vector.

[0248] In one possible implementation, client 81 is specifically configured to: obtain the predicted score of each vector in the representation vocabulary as the next vector based on the output features; and calculate the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0249] The foregoing has described the method and system architecture provided in this application. The following is a further description of the structure of the client and cloud platform provided in this application for performing the aforementioned method steps.

[0250] See Figure 9 This application provides a schematic diagram of a client architecture, which includes:

[0251] The input module 901 is used to use training data as input to the representation layer to obtain a representation vector. The representation layer is used to obtain the vector corresponding to the input data from the representation vocabulary stored on the client.

[0252] The transceiver module 902 is used to send representation vectors to the cloud platform; the client receives the output features sent by the cloud platform, and the output features are obtained by the cloud platform inputting the representation vectors into the language model.

[0253] Processing module 903 is used to calculate the loss value using the output features and update the representation layer based on the loss value;

[0254] The transceiver module 902 is also used to send loss values ​​to the cloud platform. The loss values ​​are used by the cloud platform to update the language model and obtain the updated language model.

[0255] In one possible implementation, the aforementioned input module 901 is specifically used to: input training data into the representation layer and output embedded representations; and encrypt the embedded representations to obtain representation vectors.

[0256] In one possible implementation, the aforementioned input module 901 is specifically used to: encrypt the representation vocabulary to obtain an encrypted representation vocabulary; input training data to the representation layer, and determine the embedded representation from the encrypted representation vocabulary through the representation layer.

[0257] In one possible implementation, the aforementioned input module 901 is specifically used to add a low-rank matrix to the representation vocabulary to obtain an encrypted representation vocabulary.

[0258] In one possible implementation, the aforementioned input module 901 is specifically used to perform dimensionality reduction processing on the embedded representation to obtain a representation vector.

[0259] In one possible implementation, the aforementioned input module 901 is specifically used to take the embedded representation as the input of the feedforward neural network FNN to obtain a representation vector, and the FNN is used to perform linear transformation processing on the input embedded representation.

[0260] In one possible implementation, the aforementioned input module 901 is further configured to perform a nonlinear transformation on the representation vector and use the transformed vector as a new representation vector.

[0261] In one possible implementation, the aforementioned processing module 903 is specifically used to: obtain the predicted score of each vector in the representation vocabulary as the next vector based on the output features; and calculate the loss value with the label corresponding to the training data based on the predicted score of each vector as the next vector.

[0262] See Figure 10 This application provides a schematic diagram of the structure of a cloud platform, which includes:

[0263] The transceiver module 1001 is used to receive the representation vector sent by the client. The representation vector is obtained by the client inputting the training data into the representation layer. The representation layer is used to obtain the vector corresponding to the input data from the representation vocabulary stored by the client.

[0264] Input module 1002 is used to input representation vectors into the language model, obtain output features, and send the output features to the client;

[0265] The transceiver module 1001 is also used to receive the loss value sent by the client, which is calculated by the client based on the output characteristics;

[0266] The update module 1003 is used to update the language model using the loss value to obtain the updated language model.

[0267] In one possible implementation, the aforementioned representation vector is obtained by the client encrypting the embedded representation output by the representation layer.

[0268] like Figure 11 The diagram shown is a hardware structure schematic of a client 100 provided in an embodiment of this application. This client 100 can be used to implement the aforementioned... Figures 2 to 8 The steps performed by the client in the method.

[0269] Figure 11 The client 110 shown may include a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. The processor 1101, the memory 1102, and the communication interface 1103 can be connected to each other via the bus 1104.

[0270] The processor 1101 is the control center of the client 110. It can be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor, such as a GPU or NPU, and can be adapted to the actual application scenario.

[0271] As an example, processor 1101 may include one or more CPUs, and may also include other processors, such as... Figure 11 The CPU, NPU, or GPU shown are examples of such devices.

[0272] The memory 1102 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0273] In one possible implementation, the memory 1102 may exist independently of the processor 1101. The memory 1102 can be connected to the processor 1101 via a bus 1104 and is used to store data, instructions, or program code. When the processor 1101 calls and executes the instructions or program code stored in the memory 1102, it can implement the methods provided in the embodiments of this application, for example, Figures 2 to 8 The steps performed by the client in the method shown.

[0274] In another possible implementation, the memory 1102 can also be integrated with the processor 1101.

[0275] Communication interface 1103 is used for client 110 to connect with other devices via a communication network, which may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. Communication interface 1103 may include a receiving unit for receiving data and a transmitting unit for sending data.

[0276] Bus 1104 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0277] It should be pointed out that, Figure 11The structure shown does not constitute a limitation on client 110, except Figure 11 In addition to the components shown, client 110 may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0278] like Figure 12 The diagram shown is a hardware structure schematic of a cloud platform 120 provided in an embodiment of this application. This cloud platform 120 can be used to implement the aforementioned... Figures 2 to 8 The steps in the cloud platform method.

[0279] Figure 12 The cloud platform 120 shown may include a processor 1201, a memory 1202, a communication interface 1203, and a bus 1204. The processor 1201, the memory 1202, and the communication interface 1203 can be connected to each other via the bus 1204.

[0280] Processor 1201 is the control center of cloud platform 120. It can be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor, such as a GPU or NPU, and can be adapted to the actual application scenario.

[0281] As an example, processor 1201 may include one or more CPUs, and may also include other processors, such as... Figure 12 The CPU, NPU, or GPU shown are examples of such devices.

[0282] The memory 1202 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), disk storage medium or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0283] In one possible implementation, the memory 1202 may exist independently of the processor 1201. The memory 1202 can be connected to the processor 1201 via a bus 1204 and is used to store data, instructions, or program code. When the processor 1201 calls and executes the instructions or program code stored in the memory 1202, it can implement the methods provided in the embodiments of this application, for example, Figures 2 to 8 The method shown.

[0284] In another possible implementation, the memory 1202 can also be integrated with the processor 1201.

[0285] The communication interface 1203 is used for the cloud platform 120 to connect with other devices via a communication network, which may be Ethernet, radio access network (RAN), wireless local area network (WLAN), etc. The communication interface 1203 may include a receiving unit for receiving data and a transmitting unit for sending data.

[0286] Bus 1204 can be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 12 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0287] It should be pointed out that, Figure 12 The structure shown does not constitute a limitation on the cloud platform 120, except Figure 12 In addition to the components shown, the cloud platform 120 may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0288] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0289] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0290] This application also provides a computer-readable storage medium storing a program for training a model or performing inference tasks, which, when run on a computer, causes the computer to perform the aforementioned... Figures 2 to 8 All or part of the steps in the method described in the embodiments shown.

[0291] This application also provides a digital processing chip. This digital processing chip integrates circuitry for implementing the aforementioned processor or processor functions, and one or more interfaces. When the digital processing chip integrates a memory, it can perform the method steps of any one or more of the foregoing embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the method steps of any one or more of the foregoing embodiments based on the program code stored in the external memory.

[0292] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0293] The data comparison device provided in this application embodiment can be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip in the server to perform the above-mentioned operations. Figures 3-8 The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, random access memory (RAM), etc.

[0294] Specifically, the aforementioned processing unit or processor can be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0295] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0296] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0297] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0298] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0299] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. The term "and / or" in this application is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps in this application does not imply that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved. The division of modules in this application is a logical division. In actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the modules shown or discussed may be through some ports, and the indirect coupling or communication connection between modules may be electrical or other similar forms, which are not limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed in multiple circuit modules. Some or all of the modules can be selected to achieve the purpose of the solution in this application according to actual needs.

Claims

1. A method for edge-cloud collaborative training, characterized in that, Applied to an edge-cloud collaborative training system, wherein the edge-cloud collaborative training system includes a client and a cloud platform, the method includes: The client uses training data as input to the representation layer to obtain a representation vector. The representation layer is used to obtain a vector corresponding to the input data from the representation vocabulary stored by the client. The client sends the representation vector to the cloud platform; the client receives the output features sent by the cloud platform, the output features being obtained by the cloud platform by inputting the representation vector into the language model; The client calculates a loss value using the output features, updates the representation layer based on the loss value, and sends the loss value to the cloud platform. The loss value is used by the cloud platform to update the language model, resulting in an updated language model.

2. The method according to claim 1, characterized in that, The client uses training data as input to the representation layer to obtain a representation vector, including: The client inputs the training data into the representation layer and outputs the embedded representation. The client encrypts the embedded representation to obtain the representation vector.

3. The method according to claim 2, characterized in that, The client inputs the training data into the representation layer and outputs embedded representations, including: The client encrypts the vocabulary list to obtain an encrypted vocabulary list. The client inputs the training data into the representation layer, and the representation layer determines the embedded representation from the encrypted representation vocabulary.

4. The method according to claim 3, characterized in that, The client encrypts the representation vocabulary to obtain an encrypted representation vocabulary, including: The client adds a low-rank matrix to the characterization vocabulary to obtain the encrypted characterization vocabulary.

5. The method according to any one of claims 2-4, characterized in that, The client encrypts the embedded representation to obtain the representation vector, including: The client performs dimensionality reduction on the embedded representation to obtain the representation vector.

6. The method according to claim 5, characterized in that, The client performs dimensionality reduction on the embedded representation to obtain the representation vector, including: The client takes the embedded representation as input to the feedforward neural network (FNN) to obtain the representation vector. The FNN is used to perform linear transformation on the input embedded representation.

7. The method according to any one of claims 4-6, characterized in that, The method further includes: The client performs a nonlinear transformation on the representation vector and uses the transformed vector as a new representation vector.

8. The method according to any one of claims 1-7, characterized in that, The client calculates the loss value using the output features, including: The client obtains the predicted score of each vector in the representation vocabulary as the next vector based on the output features; The client calculates the loss value based on the predicted score of each vector as the next vector and the label corresponding to the training data.

9. A method for edge-cloud collaborative training, characterized in that, Applied to an edge-cloud collaborative training system, wherein the edge-cloud collaborative training system includes a client and a cloud platform, the method includes: The cloud platform receives a representation vector sent by the client. The representation vector is obtained by the client inputting training data into the representation layer. The representation layer is used to obtain a vector corresponding to the input data from the representation vocabulary stored by the client. The cloud platform inputs the representation vector into the language model to obtain output features, and sends the output features to the client. The cloud platform receives the loss value sent by the client, and the loss value is calculated by the client based on the output features; The cloud platform uses the loss value to update the language model, resulting in an updated language model.

10. The method according to claim 9, characterized in that, The representation vector is obtained by the client after encrypting the embedded representation output by the representation layer.

11. A client, characterized in that, include: The input module is used to use training data as input to the representation layer to obtain a representation vector. The representation layer is used to obtain the vector corresponding to the input data from the representation vocabulary stored by the client. The transceiver module is used to send the representation vector to the cloud platform; the client receives the output features sent by the cloud platform, the output features being obtained by the cloud platform by inputting the representation vector into the language model; The processing module is used to calculate the loss value using the output features and update the representation layer based on the loss value; The transceiver module is also used to send the loss value to the cloud platform, and the loss value is used by the cloud platform to update the language model to obtain the updated language model.

12. The client according to claim 11, characterized in that, The input module is specifically used for: The training data is input into the representation layer, and the embedded representation is output. The embedded representation is encrypted to obtain the representation vector.

13. The client according to claim 12, characterized in that, The input module is specifically used for: The character vocabulary is encrypted to obtain an encrypted character vocabulary. The training data is input into the representation layer, and the embedding representation is determined from the encrypted representation vocabulary by the representation layer.

14. The client according to claim 13, characterized in that, The input module is specifically used to add a low-rank matrix to the representation vocabulary to obtain the encrypted representation vocabulary.

15. The client according to any one of claims 12-14, characterized in that, The input module is specifically used to perform dimensionality reduction processing on the embedded representation to obtain the representation vector.

16. The client according to claim 15, characterized in that, The input module is specifically used to take the embedded representation as the input of the feedforward neural network (FNN) to obtain the representation vector. The FNN is used to perform linear transformation processing on the input embedded representation.

17. The client according to any one of claims 14-16, characterized in that, The input module is also used to perform a nonlinear transformation on the representation vector and use the transformed vector as a new representation vector.

18. The client according to any one of claims 11-17, characterized in that, The processing module is specifically used for: Based on the output features, the predicted score of each vector in the representation vocabulary is obtained as the next vector; The loss value is calculated based on the predicted score of each vector as the next vector and the label corresponding to the training data.

19. A cloud platform, characterized in that, include: The transceiver module is used to receive the representation vector sent by the client. The representation vector is obtained by the client inputting training data into the representation layer. The representation layer is used to obtain the vector corresponding to the input data from the representation vocabulary stored by the client. The input module is used to input the representation vector into the language model, obtain the output features, and send the output features to the client; The transceiver module is also used to receive the loss value sent by the client, the loss value being calculated by the client based on the output features; The update module is used to update the language model using the loss value to obtain the updated language model.

20. The cloud platform according to claim 19, characterized in that, The representation vector is obtained by the client after encrypting the embedded representation output by the representation layer.

21. A computing device, characterized in that, The computing device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the operational steps performed by the client in the method as described in any one of claims 1 to 8.

22. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the operation steps of the cloud platform as described in any one of claims 9 to 10.

23. An edge-cloud collaborative training system, characterized in that, include: Client and cloud platform; The client is used to perform the operation steps performed by the client in any one of the methods described in claims 1 to 8; The cloud platform is used to perform the operation steps of the cloud platform in any one of the methods described in claims 9 to 10.

24. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the operation steps of the method as described in any one of claims 1 to 8 or 9 to 10.

25. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the operational steps of the method as described in any one of claims 1 to 8 or 9 to 10.

Citation Information

Cited By

  • Large-model end-cloud collaborative reasoning method and system giving consideration to efficiency and privacy

    CN121919915A