Information Transmission Method, Apparatus, Device, and Computer-Readable Storage Medium
By using preset transformation rules in federated learning to transform information, the problem of large amount of calculation under encryption is solved, and efficient privacy protection and model training are achieved.
Patent Information
- Application Number
- CN202010409477.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-05-14
AI Technical Summary
现有联邦学习中通过加密方式传递梯度信息或模型参数信息导致计算量大,影响效率,并存在隐私信息泄露的风险。
The information to be transferred is transformed using preset transformation rules, transformed information is generated, and transformed information and transformed rules are sent to the receiver for the receiver to restore the processing to obtain the information to be transferred and continue to perform federated learning tasks.
It improves the privacy protection effect of federated learning, reduces the amount of calculation, improves the efficiency of information transmission, and ensures that the model training effect is not affected.
Smart Images

Figure CN111582503B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an information transmission method, apparatus, device, and computer-readable storage medium. Background Art
[0002] With the development of computer technology, more and more technologies (big data, distributed, blockchain, artificial intelligence, etc.) are applied in the financial field, and the traditional financial industry is gradually transforming into financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for technologies.
[0003] Currently, the federated learning modeling method of jointly training models by multiple parties without data leaving the local has been applied to more and more fields to solve the problem of "data islands". In the current federated learning process, each party will send gradient information or model parameter information to each other, and through the transmission of this information, cooperate to complete the federated learning task. However, in the case of directly sending gradient information or model parameter information, if the gradient information or model parameter information is leaked to a third-party attacker, the third-party attacker may use the chain rule to restore the input data of the model based on the gradient information, and these data are often private data, and the model parameter information is also important information of the model and cannot be leaked casually. Therefore, the method of directly sending gradient information or model parameter information will still lead to the leakage of private information.
[0004] Existing solutions use encryption algorithms to encrypt the gradient information or model parameter information to be transmitted and then send it, so that a third-party attacker can only obtain the ciphertext and cannot decrypt it to obtain the original information. However, in the encryption process, very large numbers are often required for encryption calculations, and coupled with the fact that the gradient information and model parameter information themselves are very large amounts of data, it leads to a large amount of calculation in the information transmission process, and thus the information transmission efficiency is low, which in turn affects the federated learning modeling efficiency. Summary of the Invention
[0005] The main object of the present invention is to provide an information transmission method, apparatus, device, and computer-readable storage medium, aiming to solve the problem of large amount of calculation caused by the existing method of protecting the information to be transmitted in the federated learning modeling process by encryption.
[0006] To achieve the above object, the present invention provides an information transmission method, which is applied to a device participating in federated learning. The information transmission method includes the following steps:
[0007] Obtain the information to be transmitted in the federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information;
[0008] Send the transformation information and the transformation rule to the recipient, so that the recipient can perform restoration processing on the transformation information based on the transformation rule to obtain the information to be transmitted, and continue to execute the federated learning task based on the information to be transmitted.
[0009] Optionally, the step of obtaining the information to be transmitted in the federated learning task includes:
[0010] Obtain the vector to be transmitted in the federated learning task, where the vector to be transmitted is a gradient vector or a parameter vector;
[0011] Perform quantization processing on the vector to be transmitted to obtain an approximate integer vector corresponding to the vector to be transmitted, and use the approximate integer vector as the information to be transmitted.
[0012] Optionally, the step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information includes:
[0013] Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformation vector, and use the transformation vector as the transformation information, where some elements in the transformation vector are different from the elements at the corresponding positions in the approximate integer vector.
[0014] Optionally, the step of performing transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformation vector includes:
[0015] Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain an orthogonal vector of the approximate integer vector, and use the orthogonal vector as the transformation vector.
[0016] Optionally, the step of performing quantization processing on the vector to be transmitted to obtain an approximate integer vector corresponding to the vector to be transmitted includes:
[0017] Perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted.
[0018] Optionally, the step of performing approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted includes:
[0019] Obtain the descending vector corresponding to the vector to be transmitted;
[0020] Substitute the descending vector and the candidate quantity value into the maximization formula for maximizing the target value for calculation, and select the candidate quantity value that maximizes the target value as the target quantity value. Among them, the maximization formula is constructed based on the cosine similarity formula. In the maximization formula, the target value is equal to the result obtained by summing the target elements in the descending vector and then dividing by the square root of the target quantity. The target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value;
[0021] Perform an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the vector to be transmitted.
[0022] Optionally, the step of obtaining the descending vector corresponding to the vector to be transmitted includes:
[0023] Perform unit normalization processing on the vector to be transmitted to obtain a unit normalized vector;
[0024] Take the absolute value of each element in the unit normalized vector and then perform a descending order arrangement to obtain a descending vector;
[0025] The step of performing an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the vector to be transmitted includes:
[0026] Determine the elements in the descending vector whose serial numbers are less than or equal to the target quantity value as descending elements, and use the elements in the unit normalized vector corresponding to the descending elements as retained elements;
[0027] Allocate quantization values corresponding to each element in the unit normalized vector to obtain an approximate integer vector corresponding to the vector to be transmitted. Among them, allocate +1 to the elements greater than zero in the retained elements, allocate -1 to the elements less than zero in the retained elements, and allocate 0 to the elements other than the retained elements in the unit normalized vector.
[0028] Optionally, the step of approximately quantizing the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector corresponding to the vector to be transmitted includes:
[0029] Substitute a preset integer vector and the vector to be transmitted into the cosine similarity formula, where the element values of the preset integer vector are to be determined;
[0030] Solve the cosine similarity formula to obtain the element values of the preset integer vector, and use the preset integer vector with determined element values as the approximate integer vector of the vector to be transmitted, where the preset integer vector with determined element values maximizes the cosine value of the cosine similarity formula. Optionally, after the step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information, the method further includes:
[0031] Project the vector to be transmitted onto the approximate integer vector to obtain a scaling factor of the approximate integer vector relative to the vector to be transmitted;
[0032] Send the scaling factor to the receiving party, so that after the receiving party restores the transformation information to obtain the approximate integer vector based on the transformation rule, the receiving party scales the approximate integer vector based on the scaling factor to continue executing the federated learning task based on the scaling result.
[0033] Optionally, the step of sending the transformation information and the transformation rule to the receiving party includes:
[0034] Send the transformation information and the transformation rule to the receiving party based on different communication channels; or,
[0035] Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party; or,
[0036] Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party based on different communication channels.
[0037] To achieve the above object, the present invention further provides an information transmission device, which is deployed on a device participating in federated learning. The information transmission device includes:
[0038] A transformation module, configured to obtain information to be transmitted in a federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information;
[0039] A sending module, configured to send the transformation information and the transformation rule to a receiving party, so that the receiving party restores the transformation information to obtain the information to be transmitted based on the transformation rule, and continue to execute the federated learning task based on the information to be transmitted.
[0040] To achieve the above object, the present invention further provides an information transmission device, which includes: a memory, a processor, and an information transmission program stored on the memory and executable on the processor. When the information transmission program is executed by the processor, the steps of the above information transmission method are implemented.
[0041] In addition, to achieve the above object, the present invention further provides a computer-readable storage medium, on which an information transmission program is stored. When the information transmission program is executed by a processor, the steps of the information transmission method as described above are implemented.
[0042] In the present invention, by obtaining the information to be transmitted in the federated learning task, performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information, and sending the transformed information and the transformation rule to the receiving party, so that the receiving party can perform restoration processing on the transformed information according to the transformation rule to obtain the information to be transmitted, and continue to execute the federated learning task based on the information to be transmitted. In the present invention, since the transmitted information is the information to be transmitted after being transformed according to the transformation rule, rather than the information to be transmitted itself, the third attacker can only obtain the transformed information to be transmitted. Without the transformation rule, it is impossible to restore the information to be transmitted itself, so that the private information in the federated learning cannot be stolen. And although the transformation rule is also sent to the receiving party, it is independent of the transformed information. Even if a third-party attacker intercepts the transformation rule, it is difficult to know that the transformation rule can be used to restore the information to be transmitted, so that the private information in the federated learning cannot be stolen, improving the privacy protection effect of the federated learning. Since the receiving party can use the transformation rule to restore the transformed information to obtain the information to be transmitted, during the information transmission process, the information finally obtained by the receiving party is consistent with the transmitted information, without information loss, thus not affecting the training effect of the model. Compared with the mapping protection method of adding noise to the information, the solution of the present invention provides more accurate information transmission. Compared with the method of protecting by encryption, in the present invention, since a preset transformation rule is used to perform transformation processing on the information to be transmitted, the transformation rule can be set relatively simply, without using complex encryption algorithms and encryption keys for calculation, thus reducing the calculation amount in the information transmission process, improving the information transmission efficiency in the federated learning process, and further improving the federated learning modeling efficiency. It should be noted that in the present invention, what the sender transmits to the receiver is the transformation rule itself, rather than the key of the encryption algorithm. Therefore, the transformation rule in the solution of the present invention is different from the encryption algorithm, and thus the solution of using the transformation rule for transformation processing in the present invention is different from the solution of using the encryption algorithm for privacy protection. Moreover, the transformation rule can be more random than the encryption algorithm, that is, different customized transformation rules can be used for transformation processing in each transmission process, making it impossible for third-party attackers to find a pattern and impossible to crack using existing encryption algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic structural diagram of the hardware operating environment related to the solution of the embodiment of the present invention;
[0044] Figure 2 It is a schematic flowchart of the first embodiment of the information transmission method of the present invention;
[0045] Figure 3 It is a formula (1) involved in the solution of the embodiment of the present invention;
[0046] Figure 4 It is a formula (2) involved in the solution of the embodiment of the present invention;
[0047] Figure 5 It is a formula (3) involved in the solution of the embodiment of the present invention;
[0048] Figure 6 It is a functional schematic diagram module diagram of the preferred embodiment of the information transmission device of the present invention.
[0049] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0050] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] As Figure 1 shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the present invention.
[0052] It should be noted that the information transmission device in the embodiment of the present invention can be a smart phone, a personal computer, a server, etc., and no specific limitation is made here.
[0053] As Figure 1 shown, the information transmission device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and optionally the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0054] Those skilled in the art can understand, Figure 1The device structure shown does not constitute a limitation on the information transfer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0055] As Figure 1 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an information transfer program. Among them, the operating system is a program that manages and controls the hardware and software resources of the device, and supports the operation of the information transfer program and other software or programs.
[0056] In Figure 1 the device shown, the user interface 1003 is mainly used for data communication with the client; the network interface 1004 is mainly used for establishing a communication connection with the server; and the processor 1001 can be used to call the information transfer program stored in the memory 1005 and perform the following operations:
[0057] Obtain the information to be transferred in the federated learning task, and perform transformation processing on the information to be transferred according to a preset transformation rule to obtain transformed information;
[0058] Send the transformed information and the transformation rule to the receiving party, so that the receiving party can perform restoration processing on the transformed information based on the transformation rule to obtain the information to be transferred, and continue to perform the federated learning task based on the information to be transferred.
[0059] Further, the step of obtaining the information to be transferred in the federated learning task includes:
[0060] Obtain the vector to be transferred in the federated learning task, where the vector to be transferred is a gradient vector or a parameter vector;
[0061] Perform quantization processing on the vector to be transferred to obtain an approximate integer vector corresponding to the vector to be transferred, and use the approximate integer vector as the information to be transferred.
[0062] Further, the step of performing transformation processing on the information to be transferred according to a preset transformation rule to obtain transformed information includes:
[0063] Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformed vector, and use the transformed vector as the transformed information, where some elements in the transformed vector are different from the elements at the corresponding positions in the approximate integer vector.
[0064] Further, the step of performing transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformed vector includes:
[0065] Perform transformation processing on the approximate integer vector according to preset transformation rules to obtain an orthogonal vector of the approximate integer vector, and use the orthogonal vector as the transformation vector.
[0066] Further, the step of quantizing the vector to be transmitted to obtain the approximate integer vector corresponding to the vector to be transmitted includes:
[0067] Perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain the approximate integer vector of the vector to be transmitted.
[0068] Further, the step of performing approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain the approximate integer vector of the vector to be transmitted includes:
[0069] Obtain the descending vector corresponding to the vector to be transmitted;
[0070] Substitute the descending vector and the candidate quantity value into the maximization formula for maximizing the target value for calculation, and select the candidate quantity value that maximizes the target value as the target quantity value. Among them, the maximization formula is constructed based on the cosine similarity formula. In the maximization formula, the target value is equal to the result obtained by summing the target elements in the descending vector and then dividing by the square root of the target quantity. The target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value;
[0071] Perform an operation of allocating quantization values based on the target quantity value to obtain the approximate integer vector corresponding to the vector to be transmitted.
[0072] Further, the step of obtaining the descending vector corresponding to the vector to be transmitted includes:
[0073] Perform unit normalization processing on the vector to be transmitted to obtain a unit normalized vector;
[0074] Take the absolute value of each element in the unit normalized vector and then perform a descending order arrangement to obtain a descending vector;
[0075] The step of performing an operation of allocating quantization values based on the target quantity value to obtain the approximate integer vector corresponding to the vector to be transmitted includes:
[0076] Determine the elements in the descending vector whose serial numbers are less than or equal to the target quantity value as descending elements, and use the elements in the unit normalized vector corresponding to the descending elements as reserved elements;
[0077] Assign quantization values to each element in the unit vector to obtain an approximate integer vector corresponding to the vector to be transmitted. Among them, assign +1 to the elements greater than zero in the reserved elements, assign -1 to the elements less than zero in the reserved elements, and assign 0 to the elements other than the reserved elements in the unit vector.
[0078] Further, the step of approximately quantizing the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted includes:
[0079] Substitute a preset integer vector and the vector to be transmitted into the cosine similarity formula, where the element values of the preset integer vector are to be determined;
[0080] Solve the cosine similarity formula to obtain the element values of the preset integer vector, and use the preset integer vector with determined element values as the approximate integer vector of the vector to be transmitted, where the preset integer vector with determined element values maximizes the cosine value of the cosine similarity formula.
[0081] Further, after the step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information, the processor 1001 can be used to call the information transmission program stored in the memory 1005 and further perform the following operations:
[0082] Project the vector to be transmitted onto the approximate integer vector to obtain a scaling coefficient of the approximate integer vector relative to the vector to be transmitted;
[0083] Send the scaling coefficient to the receiving party, so that after the receiving party restores the transformation information to obtain the approximate integer vector based on the transformation rule, scale the approximate integer vector based on the scaling coefficient to continue to perform the federated learning task based on the scaling result.
[0084] Further, the step of sending the transformation information and the transformation rule to the receiving party includes:
[0085] Send the transformation information and the transformation rule to the receiving party based on different communication channels; or,
[0086] Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party; or,
[0087] Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party based on different communication channels.
[0088] Based on the above structure, various embodiments of the information transmission method are proposed.
[0089] Reference Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the information transmission method of the present invention.
[0090] Embodiments of the present invention provide embodiments of an information transmission method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here. The execution subject of each embodiment of the information transmission method of the present invention may be devices such as smart phones, personal computers, and servers. This device may be a device participating in federated learning. For ease of description, the sender is used as the execution subject in the following embodiments for elaboration. In this embodiment, the information transmission method includes:
[0091] Step S10, obtain the information to be transmitted in the federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information;
[0092] During the process of federated learning, participating devices and coordinating devices participating in federated learning may send gradient information or model parameter information to each other. Among them, the model parameter information is the parameter information of the model to be trained during the training process, such as the weight information in a neural network model, and the gradient information is the gradient information corresponding to each model parameter, which is used to update the model parameters. For example, during the process of horizontal federated learning, each participating device uses its own local training data to perform local training on the local model to obtain local gradient information; each participating device sends its own local gradient information to the coordinating device, and the coordinating device fuses the local gradient information of each device. The fusion operation can specifically be to perform weighted averaging on each gradient information to obtain global gradient information; the coordinating device sends the global gradient information to each participating device, and each participating device uses the global gradient information to update the local model, and based on the updated model, performs local training again, and so on in a loop until the trained model meets the performance requirements and stops training, completing the federated learning task. Similarly, each participating device can also send the local model parameter information obtained by performing local training on the local model to the coordinating device, and the coordinating device fuses the local model parameter information of each device to obtain global model parameter information, and sends the global model parameter information to each participating device for each participating device to continue local training based on the global model parameter information.
[0093] In the above-mentioned federated learning process, in order to complete the federated learning task, both the participating devices and the coordinating device may act as senders to send gradient information or model parameter information to each other. Correspondingly, both may also act as receivers to receive the gradient information or model parameter information sent by the other party. During the transmission of gradient information or model parameter information, it may be attacked by a third-party attacker, resulting in the leakage of the transmitted information, and thus the leakage of private information. Therefore, an information transmission method is proposed in the embodiments of the present invention to solve the problem of possible leakage of private information during the information transmission process.
[0094] Specifically, the sender can obtain the information to be transmitted in the federated learning task, where the information to be transmitted can be gradient information or model parameter information. For example, when the sender is a participating device and has calculated the local gradient information and needs to send it to the coordinating device (the receiver) for fusion by the coordinating device, the sender can obtain the local gradient information and use it as the information to be transmitted.
[0095] After obtaining the information to be transmitted, the sender can use a preset transformation rule to perform transformation processing on the information to be transmitted to obtain the transformed information to be transmitted (referred to as the transformed information). Among them, the preset transformation rule can be a transformation rule set in advance. After transforming the information to be transmitted according to this transformation rule, all or part of the information in the information to be transmitted has changed. And because it is transformed based on this transformation rule, the change occurs according to the rules of this transformation rule. For example, the transformation rule can be to change all positive numbers in the information to be transmitted to negative numbers and all negative numbers to positive numbers. It should be noted that multiple sets of transformation rules can be preset in advance, and the sender can select different transformation rules to perform transformation processing on the information to be transmitted each time the information is transmitted. For example, the participating device can use different transformation rules for transformation every time it sends the local gradient information in each round. Since the original information to be transmitted is transformed each time the information is sent, it is difficult for a third-party attacker to obtain the original information to be transmitted based on the transformed information without the transformation rule, thus achieving privacy protection; further, if the transformation rules used each time the information is transmitted are random, and even may be different each time, it makes it even more difficult for the attacker to obtain the original information to be transmitted based on the transformed information each time, thereby further improving the privacy protection effect.
[0096] Step S20: Send the transformed information and the transformation rule to the receiver, so that the receiver can perform reduction processing on the transformed information based on the transformation rule to obtain the information to be transmitted, and continue to execute the federated learning task based on the information to be transmitted.
[0097] After obtaining the transformation information, the sender can send the transformation information and the transformation rule to the receiver. After receiving the transformation information and the transformation rule, the receiver performs a restoration process on the transformation information to obtain the information to be transmitted. That is, based on the transformation information, an inverse transformation operation opposite to the transformation rule is performed to restore the transformation information to obtain the information to be transmitted. After obtaining the information to be transmitted, the receiver can continue to perform subsequent federated learning tasks. For example, when the receiver is a coordinating device and the restored information to be transmitted is the local model parameter information sent by the participating devices, the receiver can fuse the local model parameter information of each device to obtain the global model parameter information, and then use the global model parameter information as the information to be transmitted. It should be noted that the transformation rule sent by the sender to the receiver can be a parsable description file, and the receiver can parse the description file to obtain the transformation rule and execute the transformation rule.
[0098] It should be noted that to prevent a third-party attacker from stealing the transformation rule when stealing the transformation information, further, it can also be that the sender and the receiver have pre-set the same transformation rule locally and numbered the transformation rule. The sender sends the number of the selected transformation rule to the receiver, so that the third-party attacker cannot steal the transformation rule itself and thus cannot use the transformation rule to restore the transformation information, thereby further improving the privacy protection effect.
[0099] In this embodiment, by obtaining the information to be transmitted in the federated learning task, performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information, and sending the transformed information and the transformation rule to the receiving party, so that the receiving party can perform restoration processing on the transformed information according to the transformation rule to obtain the information to be transmitted, and continue to execute the federated learning task based on the information to be transmitted. In this embodiment, since the transmitted information is the information to be transmitted after being transformed according to the transformation rule, rather than the information to be transmitted itself, a third-party attacker can only obtain the transformed information to be transmitted. Without the transformation rule, it is impossible to restore the information to be transmitted itself, so that the private information in the federated learning cannot be stolen, and the privacy protection effect of the federated learning is improved. Since the receiving party can use the transformation rule to restore the transformed information to obtain the information to be transmitted, the information finally obtained by the receiving party is consistent with the transmitted information during the information transmission process, and there is no information loss, so that the training effect of the model will not be affected. Compared with the mapping protection method of adding noise to the information, the solution of this embodiment provides more accurate information transmission. Compared with the method of protecting by using an encryption method, in this embodiment, since a preset transformation rule is used to perform transformation processing on the information to be transmitted, the transformation rule can be set relatively simply, without using complex encryption algorithms and encryption keys for calculation, thereby reducing the amount of calculation during the information transmission process, improving the information transmission efficiency during the federated learning process, and further improving the federated learning modeling efficiency. It should be noted that in this embodiment, what the sender transmits to the receiver is the transformation rule itself, rather than the key of the encryption algorithm. Therefore, the transformation rule in the solution of this embodiment is different from the encryption algorithm, and thus the solution of using the transformation rule for transformation processing in this embodiment is different from the solution of using the encryption algorithm for privacy protection. Moreover, the transformation rule can be more random than the encryption algorithm, that is, different custom transformation rules can be used for transformation processing during each transmission process, so that there is no pattern for a third-party attacker to follow, and it is also impossible to use existing encryption algorithms for cracking.
[0100] Further, based on the above first embodiment, a second embodiment of the information transmission method of the present invention is proposed. In this embodiment, the step of obtaining the information to be transmitted in the federated learning task in step S10 includes:
[0101] Step S101, obtaining the vector to be transmitted in the federated learning task, where the vector to be transmitted is a gradient vector or a parameter vector;
[0102] Further, the gradient information or model parameter information to be transmitted in the federated learning task may be in the form of a vector. Therefore, the sender can obtain the gradient vector or parameter vector to be transmitted as the vector to be transmitted. If the form of the gradient information or model parameter information is a tensor, the sender can unfold the tensor in the channel direction to obtain multiple matrices. For each of these matrices, the rows of the matrix can be taken out for vector concatenation to obtain a high-dimensional vector. Then, unfolding multiple matrices results in multiple high-dimensional vectors, and these high-dimensional vectors are used as the vectors to be transmitted.
[0103] Step S102: Quantize the vector to be transmitted to obtain an approximate integer vector corresponding to the vector to be transmitted, and use the approximate integer vector as the information to be transmitted.
[0104] The sender quantizes the vector to be transmitted to obtain an approximate integer vector corresponding to the vector to be transmitted. Among them, the gradient vector and the parameter vector are generally floating-point numbers with a large number of digits. Therefore, the sender can quantize the vector to be transmitted, quantize the floating-point elements in the vector to be transmitted into integer elements approximate to the floating-point numbers, and obtain an approximate integer vector. The approximate integer vector requires much less storage space than the vector to be transmitted. Therefore, during the transmission process, the transmission rate is also much higher than that of the vector to be transmitted. There are many quantization processing methods that can be adopted. For example, the piecewise quantization method can be used. Specifically, the range of the floating-point numbers can be segmented, and each segment corresponds to an integer element. For the elements in the vector to be transmitted, determine the segment where the element is located and replace the element with the integer element corresponding to the segment where it is located. Other common quantization methods can also be used, which will not be elaborated in detail here.
[0105] After obtaining the approximate integer vector, the sender uses the approximate integer vector as the information to be transmitted, that is, the approximate integer vector is transformed to obtain transformed information, and then the transformed information and the transformation rule are sent to the receiver.
[0106] In this embodiment, by quantizing the vector to be transmitted to obtain an approximate integer vector of the vector to be transmitted, the amount of information to be transmitted is greatly reduced, thereby improving the efficiency of information transmission in federated learning. Moreover, since the vector to be transmitted is quantized and then transformed, it is more difficult for a third-party attacker to obtain the original vector to be transmitted. That is, even if the third-party attacker steals the transformation rule by some means and restores the transformed information using the transformation rule, what is obtained is the quantized approximate integer vector, rather than the vector to be transmitted itself, thereby increasing the difficulty for the attacker to crack the original data, and further improving the privacy protection effect during the information transmission process of federated learning.
[0107] Further, the step of performing transformation processing on the information to be transmitted according to a preset transformation rule in step S10 to obtain transformation information includes:
[0108] Step S103, performing transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformation vector, and using the transformation vector as the transformation information, wherein some elements in the transformation vector are different from the elements at the corresponding positions in the approximate integer vector.
[0109] The sender performs transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformation vector, and uses the transformation vector as the transformation information. Some elements in the transformation vector are different from the elements at the corresponding positions in the approximate integer vector. That is, after the transformation processing by the transformation rule, a transformation vector is obtained. The length of the transformation vector can be the same as that of the original approximate integer vector. Then, some or all of the elements in the transformation vector are different from the elements at the corresponding positions in the approximate integer vector. That is, the transformation rule can be a set transformation law. According to this law, some elements in the approximate integer vector can be changed. Then, the changed elements can be restored according to this transformation law, so that the receiver can restore the approximate integer vector without information loss.
[0110] In this embodiment, by performing transformation processing on the approximate integer vector corresponding to the vector to be transmitted according to a preset transformation rule, a transformation vector is obtained. The transformation vector and the transformation rule are sent to the receiver for the receiver to use the transformation rule to restore the transformation vector to obtain the approximate integer vector, so that even if a third-party attacker can restore the approximate integer vector to obtain the vector to be transmitted, and because the transformation vector is transmitted in this embodiment instead of the approximate integer vector, it adds another layer of difficulty for the third-party attacker to crack, thereby further improving the privacy protection effect in the information transmission process of federated learning.
[0111] Further, step S103 includes:
[0112] Step S1031, performing transformation processing on the approximate integer vector according to a preset transformation rule to obtain the orthogonal vector of the approximate integer vector, and using the orthogonal vector as the transformation vector.
[0113] To further increase the attack difficulty of the attacker and improve the privacy protection effect, in this embodiment, the sender can perform transformation processing on the approximate integer vector as the information to be transmitted according to a preset transformation rule to obtain the orthogonal vector of the approximate integer vector, and use the orthogonal vector as the transformation vector.
[0114] Among them, an orthogonal vector refers to a vector whose inner product with the approximate integer vector is zero. That is, through the transformation rule, the approximate integer vector is transformed so that the inner product of the transformed vector and the vector before transformation is zero. There are various transformation rules that can make the inner product of the transformed vector and the vector before transformation zero, and multiple different sets of transformation rules can be set in advance. For example, the transformation rule can be set as follows: when the elements in the approximate integer vector obtained by quantization are 0, -1, or +1, and the number of elements that are -1 is equal to the number of elements that are +1, all -1s in the approximate integer vector can be changed to +1, and the other elements remain unchanged, thereby obtaining the orthogonal vector of the approximate integer vector.
[0115] By using the transformation rule to transform the approximate integer vector to obtain the orthogonal vector of the approximate integer vector, and sending the orthogonal vector as the transformation vector to the receiving party, so that a third-party attacker can only obtain the orthogonal vector, and the orthogonal vector is perpendicular to the direction of the approximate integer vector. Therefore, it is almost impossible for a third-party attacker to deduce the input data of the model based on the orthogonal vector, and thus it is impossible to obtain the private data. Moreover, there are many transformation rules that can obtain the orthogonal vector, and it is impossible for a third-party attacker to know which one it is, making it even more difficult to crack the private information, thereby further improving the privacy protection effect in the information transmission process of federated learning.
[0116] Furthermore, based on the above first and second embodiments, a third embodiment of the information transmission method of the present invention is proposed. In this embodiment, the step of quantizing the vector to be transmitted in step S102 to obtain the approximate integer vector corresponding to the vector to be transmitted includes:
[0117] Step S1021, perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain the approximate integer vector of the vector to be transmitted.
[0118] After the sender obtains the vector to be transmitted, it can perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain the approximate integer vector of the vector to be transmitted, and send the approximate integer vector as approximate quantization information to the receiving party. The sender can calculate the approximate integer vector of the vector to be transmitted according to the cosine similarity formula. The principle is to find the approximate integer vector with the highest similarity to the vector to be transmitted. Specifically, the sender can calculate an approximate integer vector based on the cosine similarity formula so that the similarity between the approximate integer vector and the vector to be transmitted satisfies a preset similarity condition. The preset similarity condition can be a condition set in advance, such that when this condition is met, the similarity between the approximate integer vector and the vector to be transmitted is relatively high. For example, the preset similarity condition can be: after calculating multiple candidate approximate integer vectors, it is required that the finally determined approximate integer vector is the one with the highest similarity to the vector to be transmitted among these multiple candidate approximate integer vectors.
[0119] It should be noted that if the form of the gradient information or model parameter information is a tensor form, the sender can expand the tensor in the channel direction to obtain multiple matrices. For each of these matrices, the rows of the matrix can be taken out and vector concatenation can be performed to obtain a high-dimensional vector. Then, multiple matrices are expanded to obtain multiple high-dimensional vectors, and these high-dimensional vectors are used as the vectors to be transmitted. Correspondingly, after the receiver obtains the approximate integer vector corresponding to the vector to be transmitted, the approximate integer vector can be converted into a matrix form. For example, the 1*MN approximate integer vector is converted into an M*N-dimensional matrix; then multiple matrices are combined into a tensor, and subsequent federated learning tasks are performed based on the tensor.
[0120] In this embodiment, by performing approximate quantization processing on the vector to be transmitted according to the cosine similarity formula, an approximate integer vector of the vector to be transmitted is obtained, so that the approximate integer vector is as similar as possible to the vector to be transmitted in direction. Thus, the approximate integer vector obtained after quantizing the vector to be transmitted has as small a difference as possible from the original vector to be transmitted. As a result, after the receiver obtains the approximate integer vector, there is no need to restore the approximate integer vector, and subsequent calculations can be directly performed on the basis of the approximate integer vector. Therefore, while improving the communication efficiency of federated learning, the calculation amount of the receiver is also reduced, and further, the calculation amount in the federated learning process is reduced, and the modeling efficiency of federated learning is improved. Moreover, by making the approximate integer vector obtained after quantizing the vector to be transmitted have as small a difference as possible from the original vector to be transmitted, while improving the privacy protection effect and the information propagation efficiency in the federated learning process, the information transmitted can be not lost as much as possible, ensuring that the model in the federated learning can converge.
[0121] Further, in an embodiment, the step S1021 includes:
[0122] Step a, substituting a preset integer vector and the vector to be transmitted into the cosine similarity formula, where the element values of the preset integer vector are to be determined;
[0123] Step b, solving the cosine similarity formula to obtain the element values of the preset integer vector, and taking the preset integer vector with the determined element values as the approximate integer vector of the vector to be transmitted, where the preset integer vector with the determined element values maximizes the cosine value of the cosine similarity formula.
[0124] The process by which the sender calculates the approximate integer vector based on the cosine similarity formula can be specifically as follows: Substitute the preset integer vector and the vector to be transmitted into the cosine similarity formula, that is, calculate the cosine value of the preset integer vector and the vector to be transmitted. The element values of the preset integer vector are to be determined. Specifically, the element values can be randomly initialized; based on the cosine similarity formula, solve for the element values of the preset integer vector. The obtained element values maximize the cosine value of the cosine similarity formula, that is, maximize the cosine value of the preset integer vector and the vector to be transmitted; use the preset integer vector with the determined element values as the approximate integer vector of the vector to be transmitted. The solution process can be: Select various possible element values of the preset integer vector, calculate the corresponding cosine values according to the various possible element values, select the largest cosine value from the multiple cosine values, and use the element values corresponding to this cosine value as the finally determined element values of the preset integer vector. It can be understood that when the cosine value of two vectors is the largest, the two vectors are closest in direction.
[0125] Further, in one embodiment, a method for calculating an approximate integer vector based on the cosine similarity formula is also provided. This method makes the calculation efficiency faster, thereby saving computing resources and improving the overall communication efficiency. Specifically, step S1021 includes:
[0126] Step c, obtain the descending-order vector corresponding to the vector to be transmitted;
[0127] The descending-order vector corresponding to the vector to be transmitted can be obtained. Specifically, it can be to calculate the absolute values of the element values of the vector to be transmitted and then perform a descending-order arrangement to obtain the descending-order vector corresponding to the vector to be transmitted. The element values of the obtained descending-order vector are all greater than zero and are arranged in descending order of numerical value.
[0128] Further, obtaining the descending-order vector corresponding to the vector to be transmitted can also be: perform unit normalization processing on the vector to be transmitted to obtain a unit-normalized vector; take the absolute values of the elements in the unit-normalized vector and then perform a descending-order arrangement to obtain a descending-order vector. Among them, the unit normalization process is to keep the direction of the vector to be transmitted unchanged and change the length of the vector to be transmitted to 1. The specific processing process can refer to the existing process of performing unit normalization processing on vectors and will not be elaborated here in detail.
[0129] Step d, substitute the descending-order vector and the candidate quantity value into the maximization formula for maximizing the target value for calculation, and select the candidate quantity value that maximizes the target value as the target quantity value. Among them, the maximization formula is constructed based on the cosine similarity formula. In the maximization formula, the target value is equal to the result obtained by summing the target elements in the descending-order vector and then dividing by the square root of the target quantity. The target elements are the elements in the descending-order vector whose serial numbers are less than or equal to the candidate quantity value;
[0130] A maximization formula can be constructed according to the cosine similarity formula. This maximization formula is a formula for maximizing the target value. In this formula, the target value is equal to the result obtained by dividing the sum of the target elements in the descending vector by the square root of the said target quantity, where the target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value. The process of solving the maximization formula is to substitute each candidate quantity value and the target elements in the descending vector into the maximization formula to find the candidate quantity value that maximizes the target value.
[0131] Step e: Based on the said target quantity value, perform an operation of allocating quantization values to obtain the approximate integer vector corresponding to the vector to be transmitted.
[0132] After calculating the target quantity value, an operation of allocating quantization values can be performed based on the target quantity value, that is, determine the values of each element of the approximate integer vector, and then obtain the approximate integer vector corresponding to the vector to be transmitted. Specifically, the number of elements of the approximate integer vector is the same as the number of elements of the unit vector, and the serial numbers of the elements in the approximate integer vector correspond to the serial numbers of the elements in the unit vector. Then, a quantization value can be allocated to each element of the unit vector respectively, and this quantization value is used as the value of the corresponding element in the approximate integer vector, so as to obtain the approximate integer vector.
[0133] Furthermore, the elements in the descending vector whose serial numbers are less than or equal to the target quantity value can be determined as the descending elements, that is, take the largest several elements in the front as the descending elements. The descending vector is the vector obtained by taking the absolute value of the unit vector and then sorting it in descending order. Then, each element in the unit vector corresponds to an element in the descending vector. The elements in the unit vector corresponding to each descending element can be taken as the retained elements, that is, take the largest several elements with the largest absolute value in the unit vector as the retained elements. Furthermore, assign 0 to the elements in the unit vector other than the retained elements, assign +1 to the retained elements greater than zero, assign -1 to the retained elements less than zero, and if the retained element is 0, assign 0. Thus, the quantization values are allocated to each element of the unit vector, and the approximate integer vector is obtained. The values of each element in the approximate integer vector are 0, -1 or +1.
[0134] The principle therein is: Use W=(w1, w2, …, w n ) to represent the vector to be transmitted, and W 0 =(a1, a2, …, a n ) is the vector obtained by unitizing W. Let the approximate integer vector T=(t1, t2, …, t i , …t n ), where t i ∈{0, ±1} or ti ∈ {±1}, according to the cosine similarity formula, establish the relationship formula (1) between W, W Figure 3 as shown. According to this relationship formula (1), to maximize the cosine value, it is necessary to make the denominator of the term on the rightmost side of the equation in the relationship formula larger and the numerator smaller. The numerator is the square root of the sum of the squares of the values of each element in T, which can be transformed into the square root of the number of elements with an absolute value of 1 in T. The denominator is the sum of the products of the elements of W 0 corresponding to the elements of T. Then, to make the denominator larger, when the absolute value of the elements in T is 1 and the signs are the same as the corresponding elements in W 0 , the denominator can be made larger. And when the number of elements with an absolute value of 1 in T is more, the numerator will also become larger. Then, it is necessary to find the appropriate number of elements with an absolute value of 1 in T to make the result the largest. And when the elements in T with an absolute value of 1 correspond to larger elements in W 0 , the denominator is larger. Therefore, assign the quantization value with an absolute value of 1 to the elements in T corresponding to the largest several elements in W 0 . 0
[0135] Based on the above reasoning process, transform the relationship formula (1) into the formula (2) as shown Figure 4 . This formula is the maximization formula. Among them, take the absolute value of each element of W 0 to get {|a1|, |a2|,..., |a n |}, sort {|a1|, |a2|,..., |a n |} in descending order to get the vector B = {b1, b2,..., b n}, and the quantization vector T' = (t1', t2',..., t i ',... t n ') corresponding to the vector B, where t i ' represents the reordering of the original t i . The number of non-zero elements in T' is M (M = 1,..., N). Then when j ∈ (M + 1,..., N), t i ' are all zero. By solving the formula (2), an M can be obtained to make the result of (2) the largest, that is, to make the target value the largest. After obtaining M, the number of elements with an absolute value of 1 in T is obtained, and according to the corresponding relationship between the descending vector and the elements of the unit vector, it is determined which elements in T have an absolute value of 1, and then according to the positive and negative of the elements in the unit vector, the positive and negative of the elements with an absolute value of 1 in T are determined. That is, when the element in the unit vector is positive, the corresponding element in T is +1, and when the element in the unit vector is negative, the corresponding element in T is -1.
[0136] In this embodiment, in the process of calculating the approximate integer vector of the vector to be transmitted based on the cosine similarity formula, the non-linear integer problem is transformed into a linear optimization problem, that is, it is transformed into a problem of solving the maximization formula, and the computational complexity of the optimization solution is reduced to O(n*log n). Compared with the computational complexity of O(N 2 ), the solution of this embodiment greatly reduces the computational complexity in the process of federated learning information transmission, thereby improving the overall efficiency in the process of federated learning modeling while improving the communication efficiency.
[0137] Further, after the step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information in the step S10, the following steps are further included:
[0138] Step S30, projecting the vector to be transmitted onto the approximate integer vector to obtain a scaling coefficient of the approximate integer vector relative to the vector to be transmitted;
[0139] Step S40, sending the scaling coefficient to the receiving party, so that after the receiving party restores the transformation information based on the transformation rule to obtain the approximate integer vector, the receiving party scales the approximate integer vector based on the scaling coefficient, so as to continue to perform the federated learning task based on the scaling result.
[0140] Based on the quantization of the vector to be transmitted by the cosine similarity algorithm, the sender can calculate the scaling coefficient of the approximate integer vector relative to the vector to be transmitted, and send the scaling coefficient to the receiving party. After the receiving party restores the transformed vector to obtain the approximate integer vector by using the transformation rule, the receiving party scales the approximate integer vector by using the scaling coefficient, so that the difference between the scaled approximate integer vector and the vector to be transmitted is smaller, thereby minimizing the information loss caused by the quantization operation, and thus being able to further ensure that the model can converge while achieving the privacy protection effect and improving the communication efficiency, and ensuring the federated learning modeling effect.
[0141] Specifically, the sender can project the vector to be transmitted onto the approximate integer vector to obtain the scaling coefficient of the approximate integer vector relative to the vector to be transmitted. The vector to be transmitted can be projected onto the approximate integer vector in the way of orthogonal projection to obtain a scaling coefficient.
[0142] To obtain the scaling coefficient based on the projection, the scaling coefficient can be represented by Equation (3) as shown in Figure 5 , and the scaling coefficient is obtained through Equation (3), where M is the maximized M.
[0143] In order to further reduce the difference between the scaled approximate integer vector and the vector to be transmitted, the scaling coefficient can also include a positive scaling coefficient and a negative scaling coefficient.
[0144] Specifically, the sender can extract the positive element vector and the negative element vector from the vector to be transmitted, and extract the positive integer vector and the negative integer vector from the approximate integer vector. Extracting the positive element vector means keeping the positive elements in the vector to be transmitted unchanged and converting the remaining elements to 0. The extraction processes of the negative element vector, the positive integer vector, and the negative integer vector are the same. The sender projects the positive element vector onto the positive integer vector to obtain the positive scaling coefficient, and projects the negative element vector onto the negative integer vector to obtain the negative scaling coefficient. Based on the above step principle, the above formula (3) can be used to calculate the positive scaling coefficient and the negative scaling coefficient; w i is the positive element in W, t i is the positive element in T, and the positive scaling coefficient is calculated; w i is the negative element in W, t i is the negative element in T, and the negative scaling coefficient is calculated.
[0145] The sender sends the positive scaling coefficient and the negative scaling coefficient to the receiver. After the receiver uses the transformation rule to restore the transformed vector to obtain the approximate integer vector, it can scale the positive elements in the approximate integer vector based on the positive scaling coefficient, that is, multiply all the positive elements in the approximate integer vector by the positive scaling coefficient; then scale the negative elements in the approximate integer vector based on the negative scaling coefficient, that is, multiply all the negative elements in the approximate integer vector by the negative scaling coefficient; through the above scaling using the positive scaling coefficient and the negative scaling coefficient, the scaling of the approximate integer vector is completed, and the receiver continues to execute the subsequent federated learning task based on the scaling result, that is, based on the scaled approximate integer vector.
[0146] By calculating the positive scaling coefficient and the negative scaling coefficient of the approximate integer vector relative to the vector to be transmitted, and using the positive scaling coefficient and the negative scaling coefficient to scale the approximate integer vector, the difference in length between the scaled approximate integer vector and the vector to be transmitted is smaller.
[0147] Furthermore, based on the above first, second, and third embodiments, a third embodiment of the information transmission method of the present invention is proposed. In this embodiment, the step of sending the transformation information and the transformation rule to the receiver in step S20 includes:
[0148] Step S201, sending the transformation information and the transformation rule to the receiver based on different communication channels;
[0149] Further, to minimize the possibility that a third-party attacker can obtain both the transformation information and the transformation rule, the sender can send the transformation information and the transformation rule to the receiver via different communication channels. Specifically, two different communication channels can be established between the sender and the receiver, both of which can be constructed using existing communication methods and are not limited herein. By sending the transformation rule and the transformation information via different communication channels, it becomes difficult for a third-party attacker to intercept both types of information simultaneously, making it difficult to steal the private information in federated learning and further enhancing the privacy protection of federated learning.
[0150] Step S202: Encrypt the transformation rule and send the transformation information and the encrypted transformation rule to the receiver.
[0151] Alternatively, the sender can encrypt the transformation rule and send the transformation information and the encrypted transformation rule to the receiver, and in this case, the same communication channel can be used for sending. Encryption can be performed using common encryption algorithms, such as RSA (a type of asymmetric encryption algorithm), EIGmal (a two-key cryptosystem based on the discrete logarithm problem), and DSA (Digital Signature Algorithm), etc. The receiver can then use the decryption algorithm corresponding to the sender's encryption algorithm for decryption. After decrypting to obtain the transformation rule, the transformation information is restored using the transformation rule. Through encryption, even if a third-party attacker obtains the transformation rule, it is an encrypted transformation rule and cannot be used to restore the transformation information, thus enhancing the privacy protection of federated learning. Moreover, compared with the existing privacy protection method that directly encrypts the information to be transmitted using an encryption algorithm, in this embodiment, only the transformation rule needs to be encrypted, and the data volume of the transformation rule is much smaller than the information to be transmitted, so the computational amount during information transmission is not too large, thereby accelerating the information transmission efficiency in the federated learning process.
[0152] Step S203: Encrypt the transformation rule and send the transformation information and the encrypted transformation rule to the receiver via different communication channels.
[0153] Alternatively, the solutions in steps S201 and S202 can be combined. Specifically, the transformation rule can be encrypted and the transformation information and the encrypted transformation rule can be sent to the receiver via different communication channels to achieve a further privacy protection effect.
[0154] In addition, the sender can also encrypt the scaling factor and send it to the receiver, or send the scaling factor and the transformation information via different communication channels.
[0155] In addition, an information transmission device is further proposed in an embodiment of the present invention. The information transmission device is deployed on a device participating in federated learning. Referring to Figure 6 , the information transmission device includes:
[0156] A transformation module 10, configured to obtain information to be transmitted in a federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information;
[0157] A sending module 20, configured to send the transformed information and the transformation rule to a receiving party, so that the receiving party restores the transformed information based on the transformation rule to obtain the information to be transmitted, and continue to perform the federated learning task based on the information to be transmitted.
[0158] Further, the transformation module 10 includes:
[0159] An acquisition unit, configured to acquire a vector to be transmitted in a federated learning task, where the vector to be transmitted is a gradient vector or a parameter vector;
[0160] A quantization processing unit, configured to perform quantization processing on the vector to be transmitted to obtain an approximate integer vector corresponding to the vector to be transmitted, and use the approximate integer vector as the information to be transmitted.
[0161] Further, the transformation module 10 further includes:
[0162] A transformation unit, configured to perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformation vector, and use the transformation vector as the transformed information, where some elements in the transformation vector are different from the elements at the corresponding positions in the approximate integer vector.
[0163] Further, the transformation unit is further configured to:
[0164] Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain an orthogonal vector of the approximate integer vector, and use the orthogonal vector as the transformation vector.
[0165] Further, the quantization processing unit is further configured to:
[0166] Perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted.
[0167] Further, the quantization processing unit includes:
[0168] An acquisition subunit, configured to acquire a descending vector corresponding to the vector to be transmitted;
[0169] A calculation subunit, configured to substitute the descending vector and the candidate quantity value into a maximization formula for maximizing a target value for calculation, and select the candidate quantity value that maximizes the target value as the target quantity value. Wherein, the maximization formula is constructed based on the cosine similarity formula. In the maximization formula, the target value is equal to the result obtained by summing the target elements in the descending vector and then dividing by the square root of the target quantity. The target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value;
[0170] An allocation subunit, configured to perform an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the vector to be transmitted.
[0171] Further, the obtaining subunit is further configured to: perform unit normalization processing on the vector to be transmitted to obtain a unit-normalized vector; perform a descending order arrangement on the absolute values of the elements in the unit-normalized vector to obtain a descending vector;
[0172] The allocation subunit is further configured to: determine the elements in the descending vector whose serial numbers are less than or equal to the target quantity value as descending elements, and use the elements in the unit-normalized vector corresponding to the descending elements as reserved elements; allocate quantization values to the elements in the unit-normalized vector respectively to obtain an approximate integer vector corresponding to the vector to be transmitted. Among them, for the reserved elements greater than zero, positive 1 is allocated correspondingly, for the reserved elements less than zero, negative 1 is allocated correspondingly, and for the elements in the unit-normalized vector other than the reserved elements, 0 is allocated correspondingly.
[0173] Further, the quantization processing unit includes:
[0174] A substitution subunit, configured to substitute a preset integer vector and the vector to be transmitted into the cosine similarity formula, where the element values of the preset integer vector are to be determined;
[0175] A solution subunit, configured to solve the cosine similarity formula to obtain the element values of the preset integer vector, and use the preset integer vector with the determined element values as the approximate integer vector of the vector to be transmitted, where the preset integer vector with the determined element values maximizes the cosine value of the cosine similarity formula.
[0176] Further, the information transmission device further includes:
[0177] A projection module, configured to project the vector to be transmitted onto the approximate integer vector to obtain a scaling coefficient of the approximate integer vector relative to the vector to be transmitted;
[0178] The sending module 20 is further configured to send the scaling factor to the receiving party, so that after the receiving party restores the transformation information based on the transformation rule to obtain the approximate integer vector, the receiving party scales the approximate integer vector based on the scaling factor, and continues to execute the federated learning task based on the scaling result.
[0179] Further, the sending module 20 includes:
[0180] A first sending unit, configured to send the transformation information and the transformation rule to the receiving party based on different communication channels; or,
[0181] A second sending unit, configured to encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party; or,
[0182] A third sending unit, configured to encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving party based on different communication channels.
[0183] The extended content of the specific implementation manner of the information transmission device of the present invention is basically the same as that of the embodiments of the above information transmission method, and will not be elaborated here.
[0184] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which an information transmission program is stored. When the information transmission program is executed by a processor, the steps of the information transmission method described below are implemented.
[0185] For each embodiment of the information transmission device and the computer-readable storage medium of the present invention, reference may be made to each embodiment of the information transmission method of the present invention, which will not be elaborated here.
[0186] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0187] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0188] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0189] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An information transmission method, characterized in that, The information transmission method is applied to devices participating in federated learning, and the information transmission method includes the following steps: Obtain the information to be transmitted in the federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information, where the transformation rule and the transformed information are independent; Send the transformed information and the transformation rule to the receiving device, so that the receiving device can perform reduction processing on the transformed information based on the transformation rule to obtain the information to be transmitted, and continue to perform the federated learning task based on the information to be transmitted; The step of obtaining the information to be transmitted in the federated learning task includes: Obtain the vector to be transmitted in the federated learning task, where the vector to be transmitted is a gradient vector or a parameter vector; Perform approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted, and use the approximate integer vector as the information to be transmitted; The step of performing approximate quantization processing on the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted includes: Obtain the descending vector corresponding to the vector to be transmitted; where the descending vector is obtained by arranging the vector to be transmitted in descending order; Substitute the descending vector and the candidate quantity value into the maximization formula for maximizing the target value for calculation, and select the candidate quantity value that makes the target value the largest as the target quantity value, where the maximization formula is constructed based on the cosine similarity formula, and in the maximization formula, the target value is equal to the result of summing the target elements in the descending vector and then dividing by the square root of the target quantity, and the target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value; Perform an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the vector to be transmitted.
2. The information transmission method according to claim 1, wherein The step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformed information includes: Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformed vector, and use the transformed vector as the transformed information, where some elements in the transformed vector are different from the elements at the corresponding positions in the approximate integer vector.
3. The information transmission method according to claim 2, wherein The step of performing transformation processing on the approximate integer vector according to a preset transformation rule to obtain a transformed vector includes: Perform transformation processing on the approximate integer vector according to a preset transformation rule to obtain an orthogonal vector of the approximate integer vector, and use the orthogonal vector as the transformed vector.
4. The information transmission method according to claim 1, characterized in that The step of obtaining the descending vector corresponding to the vector to be transmitted includes: Perform unit normalization processing on the vector to be transmitted to obtain a unit normalized vector; Take the absolute value of each element in the unit normalized vector and then arrange them in descending order to obtain a descending vector; The step of performing an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the vector to be transmitted includes: Determine the elements in the descending vector whose sequence numbers are less than or equal to the target quantity value as the descending elements, and use the elements in the unitized elements corresponding to the descending elements as the retained elements; Correspondingly allocate quantization values to each element in the unitized vector to obtain an approximate integer vector corresponding to the vector to be transmitted. Among them, allocate positive 1 to the elements greater than zero in the retained elements, allocate negative 1 to the elements less than zero in the retained elements, and allocate 0 to the elements other than the retained elements in the unitized vector.
5. The information transmission method according to claim 1, characterized in that The step of approximately quantizing the vector to be transmitted according to the cosine similarity formula to obtain an approximate integer vector of the vector to be transmitted includes: Substitute a preset integer vector and the vector to be transmitted into the cosine similarity formula, where the element values of the preset integer vector are to be determined; Solve the cosine similarity formula to obtain the element values of the preset integer vector, and use the preset integer vector with the determined element values as the approximate integer vector of the vector to be transmitted, where the preset integer vector with the determined element values maximizes the cosine value of the cosine similarity formula.
6. The information transmission method according to claim 1, wherein, After the step of performing transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information, it further includes: Project the vector to be transmitted onto the approximate integer vector to obtain a scaling coefficient of the approximate integer vector relative to the vector to be transmitted; Send the scaling coefficient to the receiving device, so that after the receiving device restores the transformation information to obtain the approximate integer vector based on the transformation rule, scale the approximate integer vector based on the scaling coefficient to continue to perform the federated learning task based on the scaling result.
7. The information transmission method according to any one of claims 1 to 6, characterized in that The step of sending the transformation information and the transformation rule to the receiving device includes: Send the transformation information and the transformation rule to the receiving device based on different communication channels; or, Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving device; or, Encrypt the transformation rule, and send the transformation information and the encrypted transformation rule to the receiving device based on different communication channels.
8. An information transmission device, characterized in that, The information transmission device is deployed on the device participating in federated learning, and the information transmission device includes: A transformation module, configured to obtain the information to be transmitted in the federated learning task, and perform transformation processing on the information to be transmitted according to a preset transformation rule to obtain transformation information, where the transformation rule and the transformation information are independent; A sending module, configured to send the transformation information and the transformation rule to the receiving device, so that the receiving device restores the transformation information to obtain the information to be transmitted based on the transformation rule, and continue to perform the federated learning task based on the information to be transmitted; The transformation module includes: An obtaining unit, configured to obtain the vector to be transmitted in the federated learning task, where the vector to be transmitted is a gradient vector or a parameter vector; A quantization processing unit, configured to perform approximate quantization processing on the to-be-transmitted vector according to the cosine similarity formula, obtain an approximate integer vector of the to-be-transmitted vector, and use the approximate integer vector as the to-be-transmitted information; The quantization processing unit includes: An acquisition subunit, configured to acquire a descending vector corresponding to the to-be-transmitted vector; wherein, the descending vector is obtained by arranging the to-be-transmitted vector in descending order; A calculation subunit, configured to substitute the descending vector and the candidate quantity value into a maximization formula for maximizing the target value for calculation, and select the candidate quantity value that maximizes the target value as the target quantity value, wherein the maximization formula is constructed based on the cosine similarity formula, and in the maximization formula, the target value is equal to the result obtained by summing the target elements in the descending vector and then dividing by the square root of the target quantity, and the target elements are the elements in the descending vector whose serial numbers are less than or equal to the candidate quantity value; An allocation subunit, configured to perform an operation of allocating quantization values based on the target quantity value to obtain an approximate integer vector corresponding to the to-be-transmitted vector.
9. An information transfer device, characterized in that, The information transmission device includes: a memory, a processor, and an information transmission program stored on the memory and executable on the processor. When the information transmission program is executed by the processor, the steps of the information transmission method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, An information transmission program is stored on the computer-readable storage medium. When the information transmission program is executed by a processor, the steps of the information transmission method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Communication efficient federated learning
CN107871160A
Data interaction method, device and device
CN109214196A