Federal fine tuning method and device based on interval compression and related product

By partitioning, sparse and quantizing the original gradient data, the target transmission data is generated, and the problem of inefficiency of traditional compression algorithms under the scale of big data is solved, and more efficient data transmission and federated fine-tuning are achieved.

CN120124691AActive Publication Date: 2025-06-10SHENZHEN UNIV

Patent Information

Application Number
CN202510193929.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

When facing huge data scales, traditional data compression algorithms have poor compression effects, resulting in low federated fine-tuning efficiency of general pre-trained models.

Method used

By dividing the original gradient data according to the target partition parameters, sparse processing and quantization processing are performed to generate target transmission data to improve the efficiency of data transmission.

Benefits of technology

It reduces the position coding cost of traditional sparse algorithms, solves the problem that quantitative algorithms cannot provide sufficient compression rate, saves communication resources, and improves the efficiency of federated fine-tuning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124691A_ABST
    Figure CN120124691A_ABST
Patent Text Reader

Abstract

The invention discloses a federal fine tuning method and device based on interval compression and a related product, and the method comprises the steps: carrying out the dividing of original gradient data according to a target partition parameter, and obtaining the partition gradient data corresponding to at least two partition position codes; for each partition position code, performing sparse processing on the partition gradient data corresponding to the partition position code to obtain sparse position codes and sparse gradient data, and performing quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position codes; determining target transmission data according to the at least two partition position codes and the sparse position code and the target gradient data corresponding to each partition position code; and sending the target transmission data to the server, so that the server performs fine tuning processing on the general pre-training model according to the received target transmission data sent by the at least two clients, thereby improving the data compression effect in the federal fine tuning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a federated fine-tuning method, device and related products based on interval compression. Background Art

[0002] With the rapid development of artificial intelligence technology, general pre-trained models are huge in scale and have numerous parameters, with powerful learning and reasoning capabilities, and have demonstrated unprecedented capabilities in many fields.

[0003] The fine-tuning process of general pre-trained models requires a large amount of computing resources and data. Federated Learning (FL), as a distributed learning method, can use local data scattered on each client to fine-tune general pre-trained models. Specifically, after the client obtains the updated gradient of the general pre-trained model using local data, it will send the updated gradient to the server, and the server fine-tunes the general pre-trained model through the aggregated updated gradient.

[0004] The federated fine-tuning framework is accompanied by frequent communication between the client and the server. The huge scale of general pre-trained models makes the disadvantages of traditional data compression algorithms used in the data transmission process more obvious, and the compression effect is not ideal, resulting in low federated fine-tuning efficiency of general pre-trained models. Summary of the Invention

[0005] Embodiments of the present invention provide a federated fine-tuning method, device and related products based on interval compression to solve the problem of poor compression effect of traditional compression algorithms when facing a huge data scale, save the occupation of communication resources in the data transmission process, and improve the federated fine-tuning efficiency of general pre-trained models.

[0006] According to an embodiment of the present invention, a federated fine-tuning method based on interval compression is provided. The method includes:

[0007] Dividing the original gradient data according to the target partition parameters to obtain partition gradient data corresponding to at least two partition position encodings respectively;

[0008] For each partition position encoding, sparsifying the partition gradient data corresponding to the partition position encoding to obtain a sparse position encoding and sparse gradient data, and quantifying the sparse gradient data to obtain target gradient data corresponding to the partition position encoding;

[0009] Determining target transmission data according to the at least two partition position encodings and the sparse position encoding and target gradient data corresponding to each partition position encoding;

[0010] Send the target transmission data to the server so that the server can perform fine-tuning on the general pre-trained model according to the target transmission data sent by at least two clients respectively.

[0011] According to another embodiment of the present invention, there is provided a federated fine-tuning device based on interval compression, and the device includes:

[0012] An original gradient data partitioning module, configured to partition the original gradient data according to the target partitioning parameters to obtain partition gradient data corresponding to at least two partition position encodings;

[0013] A target gradient data determination module, configured to perform sparsification processing on the partition gradient data corresponding to each partition position encoding to obtain a sparse position encoding and sparse gradient data, and perform quantization processing on the sparse gradient data to obtain the target gradient data corresponding to the partition position encoding;

[0014] A target transmission data determination module, configured to determine target transmission data according to the at least two partition position encodings and the sparse position encoding and target gradient data corresponding to each partition position encoding;

[0015] A target transmission data sending module, configured to send the target transmission data to the server so that the server can perform fine-tuning on the general pre-trained model according to the target transmission data sent by at least two clients respectively.

[0016] According to another embodiment of the present invention, there is provided an electronic device, and the electronic device includes:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the federated fine-tuning method based on interval compression according to any embodiment of the present invention.

[0020] According to another embodiment of the present invention, there is provided a computer-readable storage medium, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the federated fine-tuning method based on interval compression according to any embodiment of the present invention when executed by a processor.

[0021] According to another embodiment of the present invention, there is provided a computer program product, including a computer program, and the computer program implements the federated fine-tuning method based on interval compression according to any embodiment of the present invention when executed by a processor.

[0022] In the technical solution of the embodiment of the present invention, by dividing the original gradient data into multiple partition gradient data according to the target partition parameter, and sequentially performing sparsification processing and quantization processing on each partition gradient data, it not only reduces the position encoding cost generated by the traditional sparsification algorithm as the data volume increases, but also solves the problem that the traditional quantization algorithm cannot provide sufficient compression rate as the data volume increases, saves the occupation of communication resources in the data transmission process, and improves the efficiency of the federated fine-tuning of the general pre-trained model.

[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0025] Figure 1 It is a flowchart of a federated fine-tuning method based on interval compression provided by an embodiment of the present invention;

[0026] Figure 2 It is a flowchart of another federated fine-tuning method based on interval compression provided by an embodiment of the present invention;

[0027] Figure 3 It is a schematic diagram of a specific example of a federated fine-tuning method based on interval compression provided by an embodiment of the present invention;

[0028] Figure 4 It is a flowchart of another federated fine-tuning method based on interval compression provided by an embodiment of the present invention;

[0029] Figure 5 It is a schematic structural diagram of a federated fine-tuning device based on interval compression provided by an embodiment of the present invention;

[0030] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0031] The quantization algorithm realizes the purpose of data compression by converting the high-precision representation of gradient values in gradient data into a low-precision representation. However, the quantization algorithm has limited effectiveness in achieving a high compression rate. Especially when faced with the gradient data of large-scale general pre-trained models, the quantization algorithm cannot provide a sufficiently high compression rate.

[0032] Taking the case where each gradient value in the gradient data is represented by 32 bits and the quantization algorithm reduces it to y-bit representation, the compression rate Λ is defined as the ratio of the number of bits representing the original gradient value to the number of bits representing the compressed gradient value. The compression rate Λ of the quantization algorithm Q = 32 / y. According to existing research, when the number of bits representing the gradient value is less than 6 bits, the performance of the general pre-trained model will decrease significantly. It can be seen that the upper limit of the compression rate of the quantization algorithm is approximately Λ Q = 32 / 6 ≈ 5.4. Therefore, when the quantization algorithm faces a large gradient scale, its compression rate is limited and it is difficult to meet both the compression requirements and the performance requirements of the general pre-trained model simultaneously.

[0033] The sparsification algorithm realizes the purpose of data compression by retaining the key gradient values in the gradient data and deleting the remaining other gradient values from the gradient data. For example, the TopK algorithm is a common sparsification algorithm. Specifically, the TopK algorithm aims to retain the top K gradient values with the largest absolute values in the gradient data. However, when the sparsification algorithm is applied, it will generate position encodings for marking the positions of the retained gradient values in the gradient data. As the number of gradients increases, the representation cost of the position encodings also increases, thus limiting the compression effect of the sparsification method.

[0034] Taking the gradient data with the number of gradients being d as an example, the TopK algorithm only retains the top K gradient values with the largest absolute values. Ignoring the cost of the position encoding, the compression rate Λ of the TopK algorithm S = d / K, and this compression rate is much higher than that of the quantization algorithm. Different from the quantization algorithm that directly transmits all gradient values, the sparsification algorithm needs to transmit both the retained gradient values and the position encoding of each gradient value, and the advantage of the sparsification algorithm will be weakened due to the additional position encoding. Specifically, for the gradient data with the number of gradients being d, the position encoding in the sparsification algorithm requires at least bits to represent. Therefore, considering the cost of the position encoding, the compression rate of the TopK algorithm where y represents the number of bits for representing the gradient value.

[0035] When facing the large gradient scale of the general pre-trained model, the number of gradients d is very large, making the number of bits for representing the position encoding at risk of exceeding the number of bits y for representing the gradient value. Exemplarily, assuming y = 16, when the number of gradients d = 106 When , the number of bits representing the position code is 20. At this time, no matter how K is adjusted, the communication efficiency of the TopK algorithm will be affected.

[0036] In summary, when faced with the huge gradient scale of the general pre-training model, the quantization algorithm is simple but its compression rate is limited, while the sparsification algorithm may increase the amount of transmitted data due to the additional position encoding, making the communication efficiency low.

[0037] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0038] It should be noted that the terms "first", "second", "target", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0039] Figure 1 This is a flowchart of a federated fine-tuning method based on interval compression provided by an embodiment of the present invention. This embodiment can be applied to the case where a general pre-trained model is fine-tuned using a federated fine-tuning framework. The method can be executed by a federated fine-tuning device based on interval compression. The federated fine-tuning device based on interval compression can be implemented in the form of hardware and / or software. The federated fine-tuning device based on interval compression can be configured in a terminal device. Figure 1 As shown, the method includes:

[0040] S110. According to the target partition parameter, the original gradient data is divided to obtain partition gradient data corresponding to at least two partition position codes.

[0041] Specifically, the target partition parameter is used to characterize the set value under the specified partition dimension. In an optional embodiment, the target partition parameter is the target partition number, the target gradient number, or at least two gradient start and end ranges.

[0042] Among them, the target partition number is the number of partition gradient data. For example, assuming that the number of gradients corresponding to the original gradient data is 100, when the target partition number is 20, the number of gradients corresponding to each partition gradient data is 5.

[0043] Among them, the target gradient number is the gradient number corresponding to the partition gradient data, which represents the number of gradient values ​​contained in the partition gradient data. For example, assuming that the gradient number corresponding to the original gradient data is 100, when the target gradient number is 4, the number of partition gradient data is 25.

[0044] The gradient start and end ranges include the start position code and the end position code of the gradient value in the original gradient data. For example, assuming that the number of gradients corresponding to the original gradient data is 100, the two gradient start and end ranges included in the target partition parameters may be [0, 59] and [60, 99], respectively.

[0045] In a specific embodiment, the target partition parameters can be customized according to actual needs, and the numbers of gradients corresponding to different partition gradient data can be the same or different.

[0046] Specifically, the partition position code is used to uniquely identify the partition gradient data. Exemplarily, the partition position code may be an index-based position code or a hash-based position code. The encoding form used for the partition position code is not limited here.

[0047] S120. For each partition position code, perform sparse processing on the partition gradient data corresponding to the partition position code to obtain sparse position code and sparse gradient data, and perform quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position code.

[0048] In an optional embodiment, the sparseness algorithms used by different partition position codes can be the same or different. Exemplarily, the sparseness algorithms include a TopK algorithm and a threshold algorithm, but are not limited to the example case.

[0049] Specifically, the sparse gradient data includes at least one gradient value retained after the partitioned gradient data is sparsely processed, and the sparse position code includes at least one gradient position code, and the gradient position code is used to mark the position information of the gradient value in the sparse gradient data in the partitioned gradient data. The gradient position code in the sparse position code corresponds one-to-one to the gradient value in the sparse gradient data.

[0050] For example, assuming that the partitioned gradient data is [10, 20, 1, 2, 3, 4, 5, 30], and the sparsity K corresponding to the TopK algorithm is 3, then the sparse gradient data is [10, 20, 30], and the sparse position encoding is [0, 1, 7].

[0051] In an optional embodiment, the quantization algorithms used for encoding different partition positions can be the same or different. Exemplarily, the quantization algorithms include scalar quantization algorithms, vector quantization algorithms, product quantization algorithms (ProductQuantization, PQ) and quantized stochastic gradient descent algorithms (Quantized Stochastic Gradient Descent, QSGD), etc., but are not limited to the example scenarios.

[0052] Specifically, the gradient position coding in the sparse position coding corresponds one-to-one to the gradient value in the target gradient data.

[0053] In the traditional sparsification algorithm, the retained gradient values ​​are screened in the global range of the original gradient data, which requires the number of bits of the position encoding to cover the entire range of the gradient vector. This embodiment divides the original gradient data into multiple partitioned gradient data, and independently applies the sparsification algorithm to each partitioned gradient data, thereby greatly reducing the position encoding representation requirements of each partitioned gradient data. Specifically, taking the original gradient data with d gradients as an example, this embodiment divides it into q partitioned gradient data, where the i-th partitioned gradient data contains b i gradient values, and satisfy Since the number of gradients of each partition gradient data is b i Usually much smaller than the number of gradients d of the original gradient data, so the number of bits representing the gradient position encoding is Smaller than the number of bits of position encoding in traditional sparsification algorithms This reduces the location encoding cost of gradient values ​​in traditional sparsification algorithms.

[0054] S130. Determine target transmission data according to at least two partition position codes and the sparse position code corresponding to each partition position code and target gradient data.

[0055] In an optional embodiment, target transmission data is determined based on at least two partition position codes and sparse position codes and target gradient data corresponding to each partition position code, including: for each partition position code, encapsulating the partition position code, the sparse position code corresponding to the partition position code, and the target gradient data in a data payload of a data packet; and generating target transmission data based on the data packet corresponding to each partition position code.

[0056] Specifically, in the data payload of the data packet in this embodiment, the partition position code, the gradient position code and the gradient value correspond one to one.

[0057] Exemplarily, assuming that the target gradient data contains 3 gradient values, and correspondingly, the sparse position code contains 3 gradient position codes, the data payload of the data packet contains 3 data groups, each data group contains a partition position code, a gradient position code and a gradient value.

[0058] In another optional embodiment, target transmission data is determined based on at least two partition position codes and sparse position codes and target gradient data corresponding to each partition position code, including: for each partition position code, encapsulating the partition position code in a data packet header, and encapsulating the sparse position code and target gradient data corresponding to the partition position code in a data payload of the data packet; generating target transmission data based on the data packet corresponding to each partition position code.

[0059] Specifically, in the data payload of the data packet in this embodiment, the gradient position code and the gradient value correspond one to one. Taking the above example, the data packet header of the data packet includes the partition position code, and the data payload includes 3 data groups, each of which includes the gradient position code and the gradient value.

[0060] The advantage of this setting is that it avoids adding a partition position code to each gradient value in the target gradient data, thereby greatly reducing the additional overhead of the data packet and further improving the data transmission efficiency.

[0061] In an optional embodiment, the partition position code is encapsulated in a data packet header of a data packet, and the sparse position code and target gradient data corresponding to the partition position code are encapsulated in a data payload of the data packet, including: obtaining the data packet capacity corresponding to the data packet; rounding up the ratio of the occupied capacity corresponding to the sparse position code and the target gradient data to the data packet capacity as the partition data packet flow; encapsulating the partition position code in a data packet header of the data packet of the partition data packet flow, and according to the data packet capacity, encapsulating the sparse position code and target gradient data in the data payload of the data packet of the partition data packet flow.

[0062] Specifically, the data packet capacity indicates the amount of data that can be encapsulated in the data payload of the data packet, and the partition data packet flow indicates the number of data packets allocated to the partition position code.

[0063] For example, assume that the number of gradients in the i-th partition gradient data is b i The number of gradients of the target gradient data is K i The number of bits representing the gradient value in the target gradient data is y′ bits, so the number of bits representing the gradient position encoding is bits, sparse position encoding and the occupied capacity corresponding to the target gradient data The packet capacity is represented by P, then the partition packet flow Correspondingly, (r i -1) The gradient allocation amount corresponding to each data packet is P, and the gradient allocation amount for 1 data packet is M i -P(r i -1). The gradient allocation amount represents the number of gradient values ​​encapsulated in the data packet.

[0064] S140: Send the target transmission data to the server, so that the server can fine-tune the universal pre-training model according to the target transmission data received from at least two clients.

[0065] Exemplarily, the general pre-trained model is a large language model (LLMs), a general model in the visual field, a general multimodal model, or a general model in the audio field, etc., but is not limited to the example scenarios.

[0066] In an optional embodiment, the server performs fine-tuning training on the general pre-trained model in combination with low-rank adaptation technology (LoRA) under the federated fine-tuning framework.

[0067] In the federated fine-tuning framework, multiple clients collaborate to train a shared universal pre-trained model while maintaining data privacy. Each client uses local data to fine-tune the low-rank matrix in the LoRA technology without updating the entire universal pre-trained model. Specifically, the LoRA technology introduces two low-rank matrices to represent the update gradients of the universal pre-trained model, thereby greatly reducing the number of gradients that need to be updated and the computational overhead.

[0068] The fine-tuning process of the universal pre-trained model based on LoRA technology includes four stages: parameter distribution, local update, parameter aggregation, and repeated iteration. In the parameter distribution stage, in each iteration of fine-tuning training, the server distributes the current global model data of the universal pre-trained model (including the low-rank matrix provided by LoRA technology) to all clients participating in the training, and all clients use the same universal pre-trained model as the initial model.

[0069] In the local update phase, each client uses local data to fine-tune the low-rank matrix. Specifically, the client only adjusts the two low-rank matrices introduced by the LoRA technique, without the need to update the complete model parameters of the general pre-trained model. The updated low-rank matrix represents the gradient data of the general pre-trained model determined by the client according to local data. After each client performs several rounds of local training, the calculated gradient data is uploaded to the server. Specifically, the transmission method of the gradient data adopts S110 - S130 in this embodiment.

[0070] In the model aggregation phase, after the server receives the gradient data from all clients, it performs an aggregation operation to obtain global gradient data. Exemplarily, the server averages the gradient data uploaded by all clients and uses the obtained average gradient data to fine-tune the general pre-trained model. The global model data of the fine-tuned general pre-trained model is redistributed to each client for the next iteration process.

[0071] In the repeated iteration phase, the above steps will be repeated in multiple training rounds until the performance of the general pre-trained model reaches the expected goal.

[0072] The LoRA technique enables the amount of gradient uploaded by the client to the server in each iteration process to be only 0.1% to 1% of all the model parameters of the general pre-trained model.

[0073] The technical solution of this embodiment divides the original gradient data into multiple partition gradient data according to the target partition parameters, and performs sparsification processing and quantization processing on each partition gradient data in turn, which not only reduces the position encoding cost of the traditional sparsification algorithm as the data volume increases, but also solves the problem that the traditional quantization algorithm cannot provide sufficient compression rate as the data volume increases, saves the occupancy of communication resources in the data transmission process, and improves the efficiency of the federated fine-tuning of the general pre-trained model.

[0074] Figure 2A flowchart of another federated fine-tuning method based on interval compression provided by an embodiment of the present invention, this embodiment further refines the "determining target transmission data based on at least two partition position codes and the sparse position codes and target gradient data corresponding to each partition position code" in the above embodiment. In this embodiment, the target transmission data is determined based on at least two partition position codes and the sparse position codes and target gradient data corresponding to each partition position code, including: obtaining the maximum data packet flow of the communication device corresponding to the client; taking the ratio of the maximum data packet flow to the number of partitions of the partition position code as the partition data packet flow corresponding to each partition position code; for each partition position code, determining the data packets of the partition data packet flow based on the partition position code and the sparse position code and target gradient data corresponding to the partition position code; generating the target transmission data based on the data packets of the maximum data packet flow. Figure 2 As shown, the method includes:

[0075] S210. According to the target partition parameter, the original gradient data is divided to obtain partition gradient data corresponding to at least two partition position codes.

[0076] S220. For each partition position code, perform sparse processing on the partition gradient data corresponding to the partition position code to obtain sparse position code and sparse gradient data, and perform quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position code.

[0077] S210-S220 in this embodiment are similar to those in the above embodiment. Figure 1 The S110 - S120 shown correspond to the same or similar ones, and are not described in detail in this embodiment.

[0078] S230: Obtain the maximum data packet flow of the communication device corresponding to the client.

[0079] This embodiment is more suitable for the case where the number of gradients corresponding to at least two target gradient data is similar. Specifically, the difference in the number of gradients corresponding to each two target gradient data is less than the difference threshold. For example, the difference threshold may be 20, but is not limited to the example case.

[0080] Specifically, the maximum data packet flow rate indicates the maximum number of data packets that a communication device can upload to a server at a single time.

[0081] S240: Taking the ratio between the maximum data packet flow and the number of partitions of the partition position code as the partition data packet flow corresponding to each partition position code.

[0082] Exemplarily, the maximum data packet flow is represented by R, the number of partitions corresponding to the partition position code is q, and the partition data packet flow r=R / q.

[0083] S250. For each partition position code, determine the data packet of the partition data packet flow according to the partition position code and the sparse position code corresponding to the partition position code and the target gradient data.

[0084] In an optional embodiment, a data packet of a partitioned data packet flow is determined based on a partition position code and sparse position code and target gradient data corresponding to the partition position code, including: encapsulating the partition position code and the sparse position code and target gradient data corresponding to the partition position code in a data payload of the data packet of the partitioned data packet flow.

[0085] Specifically, in the data payload of the data packet in this embodiment, the partition position code, the gradient position code and the gradient value correspond one to one.

[0086] In another optional embodiment, a data packet of a partitioned data packet flow is determined based on a partition position code and a sparse position code and target gradient data corresponding to the partition position code, including: encapsulating the partition position code in a data packet header of the data packet of the partitioned data packet flow; and randomly encapsulating the sparse position code and target gradient data in a data payload of the data packet of the partitioned data packet flow.

[0087] In another optional embodiment, a data packet of a partitioned data packet flow is determined based on a partition position code and a sparse position code and target gradient data corresponding to the partition position code, including: encapsulating the partition position code in a data packet header of the data packet of the partitioned data packet flow; taking the ratio between the number of gradients of the target gradient data and the partitioned data packet flow as the gradient allocation amount corresponding to the data packet; wherein the gradient allocation amount represents the number of gradient values ​​encapsulated in each data packet; and according to the gradient allocation amount, encapsulating the sparse position code and the target gradient data in a data payload of the data packet of the partitioned data packet flow.

[0088] Specifically, in the data payload of the data packet in this embodiment, the gradient position code and the gradient value correspond one to one, and the gradient quantity of the target gradient data represents the quantity of the gradient values ​​included in the target gradient data.

[0089] For example, the number of gradients of the i-th target gradient data is k i Indicates that the gradient allocation amount k for each data packet corresponding to the i-th partition position encoding is i,c =k i / r.

[0090] S260: Generate target transmission data according to the data packet with the maximum data packet flow.

[0091] Figure 3A schematic diagram of a specific example of a federated fine-tuning method based on interval compression provided by an embodiment of the present invention. Specifically, the number of bits of the position encoding corresponding to the original gradient data with a gradient number of d is The number of bits representing the gradient value is y. The original gradient data is divided into q partitioned gradient data with a gradient number of b. Correspondingly, the number of bits representing the partition position encoding is bits, the number of bits representing each gradient position code in sparse position coding is The number of gradient values ​​contained in the target gradient data is K.

[0092] In the gradient interval stage, the compression rate It can be expressed as the following formula:

[0093]

[0094] In the gradient compression stage, the compression rate It can be expressed as the following formula:

[0095]

[0096] Here, y′ represents the number of bits representing the gradient value in the target gradient data.

[0097] In the gradient packaging stage, since the partition position code is encapsulated in the data packet header, the compression rate It can be expressed as the following formula:

[0098]

[0099] S270: Send the target transmission data to the server, so that the server can fine-tune the universal pre-training model according to the target transmission data received from at least two clients.

[0100] S270 in this embodiment is similar to that in the above embodiment. Figure 1 The S140 shown corresponds to the same or similar steps, and will not be described in detail in this embodiment.

[0101] The technical solution of this embodiment, by uniformly distributing data packets for each partition position code and uniformly encapsulating the gradient values ​​in the target gradient data in the data packets, avoids the situation where the data packets are too large or too small under the constraint of fixed communication traffic, thereby ensuring the stability of the data transmission process and the data processing process on the server side, and avoiding the differences in the processing logic of the server side caused by inconsistent data volumes of data packets.

[0102] Figure 4The flowchart of another federated fine-tuning method based on interval compression provided by an embodiment of the present invention further refines the federated fine-tuning method based on interval compression in the above embodiment. As Figure 4 shown, the method includes:

[0103] S310. Taking the total compression error of the original gradient data as the optimization target, construct a target optimization function.

[0104] In this embodiment, the target partition parameter is the number of target partitions, and the general pre-trained model is a large language model. Specifically, whether it is the global gradient vector of the large language model or the intervalized gradient vector, both satisfy the Gaussian distribution characteristics. Therefore, the gradient vector corresponding to the t-th fine-tuning training can be modeled as a Gaussian distribution with zero mean and standard deviation of σ t .

[0105] Taking the TopK algorithm as an example, the l-th gradient value in the descending order of the absolute values of the original gradient data u t corresponding to the t-th fine-tuning training satisfies the following relationship:

[0106]

[0107] where erf -1 represents the inverse error function, d represents the number of gradients of the original gradient data u t , and l ∈ {1, 2,..., d}.

[0108] Correspondingly, after the original gradient data u t is divided into multiple partition gradient data, the l-th gradient value in the absolute value sequence of the i-th partition gradient data u t,i satisfies the following relationship:

[0109]

[0110] where b i represents the number of gradients of the partition gradient data u t,i , and l ∈ {1, 2,..., b i}}.

[0111] In this embodiment, the target optimization function represents the sum result of the partition compression errors corresponding to at least two partition position encodings. The partition compression error represents the compression error of the target gradient data relative to the partition gradient data, and the partition compression error includes sparsification error and quantization error.

[0112] Specifically, the sparsification process introduces a sparsification error by only retaining the K gradient values with the largest absolute values in the partition gradient data and discarding the remaining other gradients. The quantization process introduces a quantization error due to the precision loss of the gradient values.

[0113] In an alternative embodiment, the target optimization function satisfies the following formula:

[0114]

[0115] where Γ t represents the target optimization function corresponding to the t-th fine-tuning training of the client and the large language model, q represents the number of partitions, and γ t,i represents the partition compression error corresponding to the i-th partition position encoding under the t-th fine-tuning, G S represents the sparsification error, G Q represents the quantization error, u t,i represents the i-th partition gradient data under the t-th fine-tuning, u t,i {l} represents the l-th gradient value in the descending order of absolute values corresponding to the partition gradient data u t,i in the corresponding absolute value descending sequence, represents the sparse gradient data corresponding to the i-th partition position encoding, and ζ i represents the error parameter introduced by the quantization algorithm corresponding to the i-th partition position encoding.

[0116] Specifically, the total compression error satisfies the following relationship:

[0117]

[0118] where represents all the target gradient data corresponding to the t-th fine-tuning training, represents the i-th target gradient data under the t-th fine-tuning, and represents the desired function.

[0119] In an alternative embodiment, when the quantization process uses the PQ algorithm as the quantization algorithm, the error parameter ζ i,c corresponding to the c-th data packet under the i-th partition position encoding satisfies the following formula:

[0120]

[0121] where k i,c represents the gradient allocation amount of the c-th data packet corresponding to the i-th partition position encoding, b i represents the number of gradients corresponding to the i-th partition gradient data, y′ represents the number of bits required to represent the gradient value in the target gradient data, and P represents the data packet capacity.

[0122] Correspondingly, in this embodiment, the partition compression error γ t,i satisfies the following formula:

[0123]

[0124] Among them, r i represents the partition data packet traffic corresponding to the i-th partition position encoding.

[0125] S320. Determine the constraint conditions according to the maximum data packet traffic and data packet capacity of the communication device corresponding to the client.

[0126] In this embodiment, the constraint conditions include that the gradient allocation amount of each data packet under the partition position encoding is the same, the occupied capacity corresponding to the gradient allocation amount is less than or equal to the data packet capacity, the number of gradients corresponding to each partition gradient data is the same, and the partition data packet traffic corresponding to each partition position encoding is the same.

[0127] In an alternative embodiment, the constraint conditions include:

[0128] k i,c (log 2 b i +y′)≤P, b i =d / q, r i =R / q, k i,c =k i / r i

[0129] Among them, k i,c represents the gradient allocation amount of the c-th data packet corresponding to the i-th partition position encoding in the gradient allocation matrix k, b i represents the number of gradients corresponding to the i-th partition gradient data, y′ represents the number of bits required to represent the gradient value in the target gradient data, P represents the data packet capacity, d represents the number of gradients corresponding to the original gradient data, q represents the number of partitions, k i represents the partition sparsity corresponding to the i-th partition position encoding, 1≤k i ≤b i , r i represents the partition data packet traffic corresponding to the i-th partition position encoding, 1≤r i ≤R, R represents the maximum data packet traffic.

[0130] Specifically, k i,c satisfies the relationship Among them, 1≤k i,c ≤k i .

[0131] S330. Under the condition of satisfying the constraint conditions, perform minimization solution on the target optimization function to obtain the target number of partitions and the target gradient allocation matrix.

[0132] Specifically, based on the target optimization function and the constraint conditions, the minimization solution can be expressed as follows Form:

[0133]

[0134] In an alternative embodiment, minimizing the objective optimization function to obtain the target number of partitions and the target gradient allocation matrix includes: in each iteration process, substituting the gradient allocation matrix in the previous iteration process into the objective optimization function to obtain a first optimization function, and minimizing the first optimization function to obtain the number of partitions in the current iteration process; substituting the number of partitions in the current iteration process into the objective optimization function to obtain a second optimization function, and minimizing each partition optimization function corresponding to the second optimization function respectively to obtain the gradient allocation matrix in the current iteration process; wherein, the partition optimization function represents the partition compression error; until the objective optimization function converges, taking the number of partitions in the current iteration process as the target number of partitions, and taking the gradient allocation matrix in the current iteration process as the target gradient allocation matrix.

[0135] Proved by theoretical analysis, the form is a non-convex optimization problem. Therefore, the minimization solution process is divided into problem decomposition and alternating solution. Specifically, the form is decomposed into the form of a convex optimization problem with respect to the number of partitions q and the form of a convex optimization problem with respect to the gradient allocation matrix k.

[0136] Wherein, the form represents that in the case of fixing the gradient allocation matrix k, the objective optimization function in the form can be converted into a first optimization function, and by optimizing the number of partitions q, minimizing the first optimization function in the form. Exemplarily, the form can be expressed as the following form:

[0137]

[0138] Exemplarily, the solution method adopted for minimizing the first optimization function can be the gradient descent method, but it is not limited to the example situation.

[0139] Wherein, the form represents that in the case of fixing the number of partitions q, the objective optimization function in the form can be converted into a second optimization function, and by optimizing the gradient allocation matrix k, minimizing the second optimization function in the form. Since k i and k i,cIt only affects the compression error of the $i$-th partition and is independent of the compression errors of other partitions. Therefore, the optimization of the gradient allocation matrix $k$ can be carried out independently under the gradient data of each partition. Exemplarily, when $q = q$ * is fixed, it can be expressed in the following form:

[0140] where, represents the maximum gradient allocation amount of the data packet,

[0141] For the quantization error $G$ Q (k i ) in the partition optimization function, it can be expressed as:

[0142]

[0143] where, represents the set of gradient values assigned to the $c$-th data packet corresponding to the position encoding of the $i$-th partition in the target gradient data, represents the quantization error of the $c$-th data packet corresponding to the position encoding of the $i$-th partition.

[0144] Exemplarily, using the sequential minimal optimization algorithm, the gradient allocation amount of the data packet is iteratively adjusted to obtain the minimum quantization error $G$ Q (k i ). Specifically, in two consecutive $c$-th data packets and $(c - 1)$-th data packets, when $k$ i,c +k i,c-1 =X, the quantization error is minimized by adjusting $k$ i,c .

[0145] S340. Determine the partition sparsity amount corresponding to the sparsification process of each partition position encoding according to the target gradient allocation matrix.

[0146] Specifically, for each partition position encoding, the sum of the gradient allocation amounts of the $r$ data packets corresponding to the partition position encoding in the target gradient allocation matrix is used as the partition sparsity amount corresponding to the sparsification process of the partition position encoding.

[0147] In this embodiment, a gradient compression scheme is obtained by minimizing the target optimization function. Specifically, the gradient compression scheme includes the target number of partitions and the partition sparsity amounts corresponding to each partition position encoding respectively.

[0148] S350. Divide the original gradient data according to the target number of partitions to obtain partition gradient data corresponding to at least two partition position encodings respectively.

[0149] Exemplarily, for the original gradient data of the number of gradients d, when the number of target partitions is q, the number of gradients b corresponding to the partition gradient data is b = d / q.

[0150] S360. For each partition position encoding, according to the partition sparsity amount corresponding to the partition position encoding, sparsify the partition gradient data corresponding to the partition position encoding to obtain the sparse position encoding and the sparse gradient data, and perform quantization processing on the sparse gradient data to obtain the target gradient data corresponding to the partition position encoding.

[0151] Exemplarily, when the partition sparsity amount corresponding to the i-th partition position encoding is k i then the data amounts corresponding to the sparse position encoding, the sparse gradient data, and the target gradient data are all k i .

[0152] S370. Determine the target transmission data according to at least two partition position encodings and the sparse position encoding and the target gradient data corresponding to each partition position encoding.

[0153] S380. Send the target transmission data to the server so that the server can perform fine-tuning processing on the large language model according to the target transmission data sent by at least two clients received.

[0154] S370 - S380 in this embodiment is the same or similar to S130 - S140 shown in the above embodiment, or is the same or similar to S230 - S270 shown in the above embodiment. This embodiment will not be elaborated here. Figure 1 shown in the above embodiment Figure 2 This embodiment will not be elaborated here.

[0155] The number of gradients corresponding to the partition gradient data and the partition sparsity amount are crucial. If the number of gradients is too large, the position encoding cost is still high. If the gradient data is too small, it is difficult to sparsify the gradients sufficiently. If the partition sparsity amount is too large, the compression effect is not good. If the partition sparsity amount is too small, it is easy to affect the model performance of the large language model.

[0156] The technical solution of this embodiment is based on the characteristic that the gradient vectors of the large language model conform to the Gaussian distribution. Taking the total compression error of the original gradient data as the optimization target, constructing a target optimization function, determining the constraint conditions according to the maximum data packet traffic and the data packet capacity of the communication device corresponding to the client, and minimizing the target optimization function under the condition of satisfying the constraint conditions to obtain a gradient compression scheme, achieving a balance between the sparsification effect and the quantization effect, and at the same time achieving a balance between the compression effect and the model performance of the large language model.

[0157] The above embodiments have conducted extensive experiments on public datasets. The experimental results show that compared with the federated fine-tuning method using traditional data compression algorithms, the large language model obtained by using the federated fine-tuning method based on interval compression provided in this embodiment has an accuracy improvement of 6.42% to 18.87%, and a reduction of 17.07% to 44.44% in communication traffic, achieving the purpose of synchronously improving the fine-tuning effect of the large language model and the compression effect in the communication process.

[0158] The following are embodiments of the federated fine-tuning device based on interval compression provided by the embodiments of the present invention. This device and the federated fine-tuning method based on interval compression in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the federated fine-tuning device based on interval compression, reference can be made to the content of the federated fine-tuning method based on interval compression in the above embodiments.

[0159] Figure 5 The structural schematic diagram of a federated fine-tuning device based on interval compression provided by an embodiment of the present invention. As Figure 5 shown, the device includes: an original gradient data partitioning module 410, a target gradient data determination module 420, a target transmission data determination module 430, and a target transmission data sending module 440.

[0160] The original gradient data partitioning module 410 is configured to partition the original gradient data according to the target partitioning parameters to obtain partition gradient data corresponding to at least two partition position encodings;

[0161] The target gradient data determination module 420 is configured to perform sparsification processing on the partition gradient data corresponding to each partition position encoding to obtain a sparse position encoding and sparse gradient data, and perform quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position encoding;

[0162] The target transmission data determination module 430 is configured to determine target transmission data according to at least two partition position encodings and the sparse position encoding and target gradient data corresponding to each partition position encoding;

[0163] The target transmission data sending module 440 is configured to send the target transmission data to the server, so that the server fine-tunes the general pre-trained model according to the target transmission data received from at least two clients respectively.

[0164] The technical solution of this embodiment divides the original gradient data into multiple partition gradient data according to the target partition parameters, and sequentially performs sparsification processing and quantization processing on each partition gradient data, which not only reduces the position encoding cost of the traditional sparsification algorithm with the increase of data volume, but also solves the problem that the traditional quantization algorithm cannot provide enough compression rate with the increase of data volume, saves the occupation of communication resources in the data transmission process, and improves the efficiency of the federated fine-tuning of the general pre-trained model.

[0165] In an alternative embodiment, the target transmission data determination module 430 includes:

[0166] The maximum data packet traffic acquisition unit is used to acquire the maximum data packet traffic of the communication device corresponding to the client;

[0167] The partition data packet traffic determination unit is used to use the ratio between the maximum data packet traffic and the number of partitions of the partition position encoding as the partition data packet traffic corresponding to each partition position encoding;

[0168] The data packet determination unit is used to, for each partition position encoding, determine the data packet of the partition data packet traffic according to the partition position encoding, the sparse position encoding corresponding to the partition position encoding, and the target gradient data;

[0169] The target transmission data generation unit is used to generate target transmission data according to the data packet of the maximum data packet traffic.

[0170] In an alternative embodiment, the data packet determination unit is specifically used for:

[0171] Encapsulate the partition position encoding in the data header of the data packet of the partition data packet traffic;

[0172] Use the ratio between the number of gradients of the target gradient data and the partition data packet traffic as the gradient allocation amount corresponding to the data packet; wherein, the gradient allocation amount represents the number of gradient values encapsulated in each data packet;

[0173] According to the gradient allocation amount, encapsulate the sparse position encoding and the target gradient data in the data payload of the data packet of the partition data packet traffic.

[0174] In an alternative embodiment, the target partition parameter is the target number of partitions, and the general pre-trained model is a large language model. Correspondingly, the device further includes:

[0175] The target optimization function construction module is used to construct a target optimization function with the total compression error of the original gradient data as the optimization target;

[0176] The constraint condition determination module is used to determine the constraint conditions according to the maximum data packet traffic and the data packet capacity of the communication device corresponding to the client;

[0177] A target optimization function solving module, configured to minimize the target optimization function to obtain the target number of partitions and the target gradient allocation matrix under the condition of satisfying the constraint conditions; wherein, the matrix parameter k in the target gradient allocation matrix i,c is the gradient allocation amount, indicating the number of gradient values encapsulated in the c-th data packet corresponding to the i-th partition position encoding;

[0178] A partition sparsity amount determination module, configured to determine the partition sparsity amount corresponding to the sparsification process of each partition position encoding according to the target gradient allocation matrix;

[0179] Wherein, the target optimization function represents the summation result of the partition compression errors corresponding to at least two partition position encodings respectively, the partition compression error represents the compression error of the target gradient data relative to the partition gradient data, the partition compression error includes the sparsification error and the quantization error, and the constraint conditions include that the gradient allocation amounts of each data packet under the partition position encoding are the same, the occupied capacity corresponding to the gradient allocation amount is less than or equal to the data packet capacity, the number of gradients corresponding to each partition gradient data is the same, and the partition data packet traffic corresponding to each partition position encoding is the same.

[0180] In an alternative embodiment, the target optimization function satisfies the following formula:

[0181]

[0182] Wherein, Γ t represents the target optimization function corresponding to the t-th fine-tuning training of the client and the large language model, q represents the number of partitions, and γ t,i represents the partition compression error corresponding to the i-th partition position encoding under the t-th fine-tuning, G S represents the sparsification error, G Q represents the quantization error, u t,i represents the i-th partition gradient data under the t-th fine-tuning, u t,i {l} represents the l-th gradient value in the absolute value descending sequence corresponding to the partition gradient data u t,i , represents the sparse gradient data corresponding to the i-th partition position encoding, and ζ i represents the error parameter introduced by the quantization algorithm corresponding to the i-th partition position encoding.

[0183] In an alternative embodiment, the constraint conditions include:

[0184] k i,c (log 2 b i +y′)≤P, b i =d / q, r i= R / q, k i,c = k i / r i

[0185] where k i,c represents the gradient allocation amount of the c-th data packet corresponding to the i-th partition position encoding, b i represents the number of gradients corresponding to the i-th partition gradient data, y' represents the number of bits required to represent the gradient value in the target gradient data, P represents the data packet capacity, d represents the number of gradients corresponding to the original gradient data, q represents the number of partitions, and k i represents the partition sparsity amount corresponding to the i-th partition position encoding, r represents the partition data packet traffic corresponding to the i-th partition position encoding, and R represents the maximum data packet traffic.

[0186] In an alternative embodiment, the target optimization function solving module is specifically configured to:

[0187] In each iteration process, substitute the gradient allocation matrix in the previous iteration process into the target optimization function to obtain a first optimization function, and perform a minimization solution on the first optimization function to obtain the number of partitions in the current iteration process;

[0188] Substitute the number of partitions in the current iteration process into the target optimization function to obtain a second optimization function, and perform a minimization solution on each partition optimization function corresponding to the second optimization function respectively to obtain the gradient allocation matrix in the current iteration process; wherein, the partition optimization function characterizes the partition compression error;

[0189] Until the target optimization function converges, use the number of partitions in the current iteration process as the target number of partitions, and use the gradient allocation matrix in the current iteration process as the target gradient allocation matrix.

[0190] The federated fine-tuning device based on interval compression provided by the embodiments of the present invention can execute the federated fine-tuning method based on interval compression provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0191] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0192] As Figure 6 shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor 11. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0193] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information or data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0194] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the federated fine-tuning method based on interval compression provided in the above embodiments.

[0195] In some embodiments, the federated fine-tuning method based on interval compression provided by the above embodiments can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the federated fine-tuning method based on interval compression described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the federated fine-tuning method based on interval compression by any other suitable means (e.g., by means of firmware).

[0196] The various embodiments of the systems and techniques described above in this document can be implemented in the following systems or combinations thereof: digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0197] The computer program for implementing the federated fine-tuning method based on interval compression of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0198] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media would include electrical connections based on at least one wire, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0199] To provide for interaction with a user, the systems and techniques described herein can be implemented on a terminal device that has: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the terminal device. Other kinds of devices can also provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0200] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0201] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is created by computer programs running on corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and Virtual Private Server (VPS) services.

[0202] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0203] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A federated fine-tuning method based on interval compression, characterized in that: include: According to the target partition parameter, the original gradient data is divided to obtain partition gradient data corresponding to at least two partition position codes respectively; For each partition position code, performing sparse processing on the partition gradient data corresponding to the partition position code to obtain sparse position code and sparse gradient data, and performing quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position code; Determine target transmission data according to the at least two partition position codes and the sparse position code corresponding to each partition position code and target gradient data; The target transmission data is sent to the server, so that the server fine-tunes the universal pre-training model according to the target transmission data received from at least two clients respectively.

2. The method according to claim 1, characterized in that The determining target transmission data according to the at least two partition position codes and the sparse position code corresponding to each partition position code and the target gradient data comprises: Obtain the maximum data packet flow of the communication device corresponding to the client; The ratio between the maximum data packet flow and the number of partitions of the partition position code is used as the partition data packet flow corresponding to each partition position code; For each partition position code, determine a data packet of the partition data packet flow according to the partition position code and the sparse position code corresponding to the partition position code and target gradient data; Generate target transmission data based on the data packets with the maximum data packet flow.

3. The method according to claim 2, characterized in that The step of determining the data packet of the partition data packet flow according to the partition position code and the sparse position code and target gradient data corresponding to the partition position code comprises: Encapsulating the partition position code in a data packet header of a data packet of the partition data packet flow; The ratio between the number of gradients of the target gradient data and the flow rate of the partitioned data packet is used as the gradient allocation amount corresponding to the data packet; wherein the gradient allocation amount represents the number of gradient values ​​encapsulated in each data packet; The sparse position code and the target gradient data are encapsulated in a data payload of a data packet of the partitioned data packet flow according to the gradient allocation amount.

4. The method according to any one of claims 1 to 3, characterized in that: The target partition parameter is the target partition number, the general pre-trained model is a large language model, and accordingly, the method further includes: Taking the total compression error of the original gradient data as an optimization target, constructing a target optimization function; Determine the constraint condition according to the maximum data packet flow and data packet capacity of the communication device corresponding to the client; Under the condition that the constraint conditions are met, the target optimization function is minimized to obtain the target number of partitions and the target gradient allocation matrix; wherein the matrix parameter k in the target gradient allocation matrix is i,c is the gradient allocation amount, which indicates the number of gradient values ​​encapsulated in the cth data packet corresponding to the i-th partition position code; Determine the partition sparsity corresponding to the sparsification processing of each partition position encoding according to the target gradient allocation matrix; Among them, the objective optimization function represents the sum of the partition compression errors corresponding to at least two partition position codes, the partition compression error represents the compression error of the target gradient data relative to the partition gradient data, the partition compression error includes sparsification error and quantization error, and the constraints include that the gradient allocation amount of each data packet under the partition position coding is the same, the occupied capacity corresponding to the gradient allocation amount is less than or equal to the data packet capacity, the number of gradients corresponding to each partition gradient data is the same, and the partition data packet flow corresponding to each partition position coding is the same.

5. The method according to claim 4, characterized in that The objective optimization function satisfies the following formula: Among them, Γ t represents the target optimization function corresponding to the t-th fine-tuning training of the client and the large language model, q represents the number of partitions, and γ t,i represents the partition compression error corresponding to the i-th partition position encoding under the t-th fine-tuning, G S represents the sparsification error, G Q represents the quantization error, u t,i represents the i-th partition gradient data under the t-th fine-tuning, u t,i {l} represents the partition gradient data u t,i The corresponding l-th gradient value in the descending sequence of absolute values, represents the sparse gradient data corresponding to the position encoding of the i-th partition, ζ i Represents the error parameter introduced by the quantization algorithm corresponding to the i-th partition position encoding.

6. The method according to claim 4, characterized in that The constraints include: k i,c (log2b i +y′)≤P,b i =d / q,r i =R / q,k i,c =k i / r i ; Among them, k i,c represents the gradient allocation of the cth data packet corresponding to the i-th partition position encoding in the gradient allocation matrix k, b i represents the number of gradients corresponding to the i-th partition gradient data, y′ represents the number of bits required to represent the gradient value in the target gradient data, P represents the data packet capacity, d represents the number of gradients corresponding to the original gradient data, q represents the number of partitions, and k i represents the partition sparsity corresponding to the i-th partition position encoding, r i represents the partition data packet flow corresponding to the i-th partition position code, and R represents the maximum data packet flow.

7. The method according to claim 4, characterized in that The minimization of the target optimization function to obtain the target partition quantity and the target gradient allocation matrix includes: In each iteration process, the gradient allocation matrix in the previous iteration process is substituted into the target optimization function to obtain a first optimization function, and the first optimization function is minimized to obtain the number of partitions in the current iteration process; Substituting the number of partitions of the current iteration process into the target optimization function to obtain a second optimization function, and minimizing each partition optimization function corresponding to the second optimization function to obtain a gradient allocation matrix of the current iteration process; wherein the partition optimization function represents the partition compression error; Until the target optimization function converges, the number of partitions in the current iteration process is used as the target number of partitions, and the gradient allocation matrix in the current iteration process is used as the target gradient allocation matrix.

8. A federated fine-tuning device based on interval compression, characterized in that: include: The original gradient data partitioning module is used to partition the original gradient data according to the target partition parameters to obtain partition gradient data corresponding to at least two partition position codes; a target gradient data determination module, for each partition position code, performing a sparse processing on the partition gradient data corresponding to the partition position code to obtain a sparse position code and sparse gradient data, and performing a quantization processing on the sparse gradient data to obtain target gradient data corresponding to the partition position code; A target transmission data determination module, configured to determine the target transmission data according to the at least two partition position codes and the sparse position code corresponding to each partition position code and the target gradient data; The target transmission data sending module is used to send the target transmission data to the server, so that the server can fine-tune the universal pre-training model according to the target transmission data respectively sent by at least two clients.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so as to enable the at least one processor to perform the federated fine-tuning method based on interval compression according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the federated fine-tuning method based on interval compression according to any one of claims 1 to 7 when executed.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the federated fine-tuning method based on interval compression according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Federal learning-based model training method and device, equipment and medium

    CN116776155A

  • Efficient federated learning method based on gradient compression

    CN118333182A

  • Federal learning client communication compression method, client device and federal learning system

    CN118410840A

  • Federal learning method based on dynamic local training and gradient compression strategy

    CN118839752A

  • Communication compression method based on model weight distribution in federated learning

    US11468370B1

Cited By

  • Federal training method and device based on segmented sparse coding, equipment and medium

    CN120874972A

  • Federal training method and device based on segment sparse coding, equipment and medium

    CN120874972B