Heterogeneous accelerated computing optimization method, device, equipment and readable storage medium
By copying plaintext data to heterogeneous chip memory in the heterogeneous federated learning framework and separating memory copy and cryptographic computing, the problem of low computing efficiency caused by multiple copies of ciphertext data is solved, and more efficient computing performance is achieved.
Patent Information
- Application Number
- CN202110427011.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-20
AI Technical Summary
In the heterogeneous federated learning framework, multiple copies of ciphertext data between CPU memory and heterogeneous chip memory lead to inefficient computing.
By obtaining plaintext data and copying it from CPU memory to heterogeneous chip memory, the memory copy and cryptographic calculation process are separated, and the number of copies of ciphertext data between memory is reduced.
Reduces the number of memory copies and time-consuming, and improves the computing efficiency of the heterogeneous federated learning framework.
Smart Images

Figure CN113204502B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a heterogeneous accelerated computing optimization method, device, equipment and readable storage medium. Background Art
[0002] With the continuous development of computer technology, the application of artificial intelligence is becoming more and more extensive. At present, when federated learning is used to combine multi-party data for modeling, the participants of federated learning usually need to perform massive secret state calculations. In the heterogeneous federated learning framework, in order to improve the computing effect of the data, in one round of federated learning iterations, it is usually necessary to copy the ciphertext data back and forth between the CPU memory and the heterogeneous chip memory multiple times. However, since the number of bits of ciphertext data is usually high, the memory copy of ciphertext data takes a long time, which will result in the memory copy time of the ciphertext data being much longer than the secret state calculation time of the ciphertext data in one round of federated learning iterations, making the computational efficiency of the heterogeneous federated learning framework low. Summary of the invention
[0003] The main purpose of this application is to provide a heterogeneous accelerated computing optimization method, device, equipment and readable storage medium, aiming to solve the technical problem of low computing efficiency of heterogeneous federated learning framework in the prior art.
[0004] To achieve the above-mentioned object, the present application provides a heterogeneous accelerated computing optimization method, which is applied to a heterogeneous accelerated computing optimization device, and the heterogeneous accelerated computing optimization method includes:
[0005] Obtaining plaintext data, and based on a first memory copy operator set, copying the plaintext data from a CPU memory to a heterogeneous chip memory;
[0006] Based on the set of secret computing operators, performing secret computing resident in the memory of the heterogeneous chip on the plaintext data to obtain a secret computing result;
[0007] Feedback the secret state calculation result to the CPU memory.
[0008] The present application also provides a heterogeneous accelerated computing optimization device, which is a virtual device and is applied to a heterogeneous accelerated computing optimization device. The heterogeneous accelerated computing optimization device includes:
[0009] A memory copy module, used to obtain plaintext data, and based on a first memory copy operator set, copy the plaintext data from the CPU memory to the heterogeneous chip memory;
[0010] A secret computing module, used to perform secret computing resident in the memory of the heterogeneous chip on the plaintext data based on a secret computing operator set to obtain a secret computing result;
[0011] A feedback module is used to feed back the secret state calculation result to the CPU memory.
[0012] The present application also provides a heterogeneous accelerated computing optimization device, which is a physical device. The heterogeneous accelerated computing optimization device includes: a memory, a processor, and a program of the heterogeneous accelerated computing optimization method stored in the memory and executable on the processor. When the program of the heterogeneous accelerated computing optimization method is executed by the processor, the steps of the heterogeneous accelerated computing optimization method as described above can be implemented.
[0013] The present application also provides a readable storage medium, on which is stored a program for implementing the heterogeneous accelerated computing optimization method. When the program of the heterogeneous accelerated computing optimization method is executed by a processor, the steps of the heterogeneous accelerated computing optimization method as described above are implemented.
[0014] The present application also provides a computer program product, including a computer program, which implements the steps of the heterogeneous accelerated computing optimization method as described above when executed by a processor.
[0015] The present application provides a heterogeneous accelerated computing optimization method, device, equipment and readable storage medium. Compared with the technical means used in the prior art to copy ciphertext data back and forth between the CPU memory and the heterogeneous chip memory multiple times during a round of iterations of federated learning, the present application first obtains plaintext data, and based on a first memory copy operator set, copies the plaintext data from the CPU memory to the heterogeneous chip memory, wherein the purpose of separating the memory copy process from the secret state computing process is achieved by encapsulating the memory copy as an operator separately, and the plaintext data, rather than the ciphertext data, is copied between the heterogeneous chip memory and the CPU memory. Since the number of bits of the plaintext data is much smaller than the ciphertext data, the time consumption of the memory copy is reduced. It should be noted that at present, the memory copy and the secret state computing are usually encapsulated in the same operator, and the operator execution process is usually the alternation of secret state computing and memory copy, and finally the final secret state computing result is obtained, wherein the intermediate ciphertext data generated by multiple secret state computing all need to be written to the CPU memory, thereby resulting in multiple memory copies of the ciphertext data between the heterogeneous chip memory and the CPU memory. Furthermore, based on In the secret state computing operator set, the plaintext data is subjected to secret state computing resident in the memory of the heterogeneous chip to obtain the secret state computing result. That is, after separating the memory copy from the secret state computing, the ciphertext data is subjected to secret state computing resident in the memory of the heterogeneous chip, and the intermediate ciphertext data generated in the secret state computing process are all written into the memory of the heterogeneous chip until the final secret state computing result is calculated, thereby avoiding multiple memory copy processes in the secret state computing process. In one round of iteration of federated learning, only one back-and-forth memory copy is required between the CPU memory and the heterogeneous chip memory, thereby reducing the number of memory copies, further reducing the time consumption of memory copy in the federated learning process, and finally feeding back the secret state computing result to the CPU memory to complete one round of iterative computing process of federated learning, thereby overcoming the problem that the number of bits of ciphertext data is usually high, the memory copy of ciphertext data is time-consuming, and thus the memory copy time of ciphertext data will be much longer than the secret state computing time of ciphertext data in one round of iteration of federated learning, resulting in the technical defect of low computing efficiency of the heterogeneous federated learning framework, thereby improving the computing efficiency of the heterogeneous federated learning framework. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 This is a flow chart of the first embodiment of the heterogeneous accelerated computing optimization method of the present application;
[0019] Figure 2 A schematic diagram of a five-stage pipeline in the heterogeneous accelerated computing optimization method of this application;
[0020] Figure 3 A schematic diagram of a heterogeneous federated learning framework in the heterogeneous accelerated computing optimization method of this application;
[0021] Figure 4 This is a flow chart of the second embodiment of the heterogeneous accelerated computing optimization method of the present application;
[0022] Figure 5 A schematic diagram of the design of multiple parallel pipelines in the heterogeneous accelerated computing optimization method of this application;
[0023] Figure 6 A schematic diagram of the device structure of the hardware operating environment involved in the heterogeneous accelerated computing optimization method in the embodiment of the present application;
[0024] Figure 7 This is a schematic diagram of the structure of the heterogeneous accelerated computing optimization device of this application.
[0025] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0026] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] The present application embodiment provides a heterogeneous accelerated computing optimization method. In the first embodiment of the heterogeneous accelerated computing optimization method of the present application, refer to Figure 1 , the heterogeneous accelerated computing optimization method includes:
[0028] Step S10, obtaining plaintext data, and based on a first memory copy operator set, copying the plaintext data from the CPU memory to the heterogeneous chip memory;
[0029] In this embodiment, it should be noted that the plaintext data is unencrypted data, and the number of bits of the plaintext data is much smaller than that of the ciphertext data. For example, assuming that the ciphertext data is a 10000*10000 matrix, the corresponding plaintext data is a 50*50 matrix, and the memory copy operator set includes at least one memory copy operator, wherein the memory copy operator is a mapping for memory copy, which is used to map data from the underlying CPU memory to the heterogeneous chip memory, wherein the heterogeneous chip memory includes GPU video memory and FPGA memory, etc.
[0030] In addition, it should be noted that the heterogeneous accelerated computing optimization method is applied to a heterogeneous federated learning framework, wherein the heterogeneous federated learning framework includes a CPU side and a heterogeneous chip side, the upper layer on the CPU side is an application layer, wherein the application layer includes a Python language layer, etc., and the bottom layer on the CPU side is a C language layer, etc., so the CPU memory includes an upper-layer CPU memory and a bottom-layer CPU memory, and there is an upper-layer CPU memory corresponding to the upper layer of the CPU side and a bottom-layer CPU memory corresponding to the bottom layer of the CPU side on the CPU side, and there is a heterogeneous chip memory on the heterogeneous chip side, and memory copying can be performed between the upper-layer CPU memory and the bottom-layer CPU memory, and memory copying can be performed between the heterogeneous chip memory and the bottom-layer CPU memory.
[0031] Obtain plaintext data, and based on the first memory copy operator set, copy the plaintext data from the CPU memory to the heterogeneous chip memory, specifically, determine the original data in the upper CPU memory, wherein the original data includes model parameters, model output data, and model input data in the federated learning process, and then perform data conversion on the original data to convert the original data into data adapted to the C language layer, obtain the plaintext data, and then call the memory copy operator encapsulated by the preset computing interface in the memory copy operator set through the upper computing module on the CPU side, copy the plaintext data from the bottom CPU memory to the heterogeneous chip memory, and convert the data format of the plaintext data into data adapted to the heterogeneous The data format of the chip memory is written into the heterogeneous chip memory, wherein the upper-layer computing module is deployed on the upper layer of the CPU side for executing computing tasks, the preset computing interface is deployed on the bottom layer of the CPU side for uniformly encapsulating the heterogeneous chip operators into an interface for the upper-layer computing module to call, the memory control module is a module for operating the underlying CPU memory and the heterogeneous chip memory through address information such as pointers or references, and is used to perform memory operations, serialization operations, data conversion and site communication on the underlying CPU memory and the heterogeneous chip memory, wherein the memory operations include transposition, cutting and memory copying, the serialization operations include converting data into bit streams, etc., and the heterogeneous chip operators include GPU operators and FPGA operators, etc.
[0032] Step S20, based on the secret state computing operator set, performing secret state computing resident in the memory of the heterogeneous chip on the plaintext data to obtain a secret state computing result;
[0033] In this embodiment, based on the secret state computing operator set, the plaintext data is subjected to secret state computing that resides in the memory of the heterogeneous chip to obtain a secret state computing result. Specifically, through the upper-level computing module, the encryption operator in the secret state computing operator set is called to encrypt the plaintext data to obtain ciphertext data, and the ciphertext data is written into the heterogeneous chip memory through the memory control module, and then through the upper-level computing module, the secret state computing operator in the secret state computing operator set is called to perform secret state computing on the ciphertext data that resides in the memory of the heterogeneous chip to obtain a secret state computing result, wherein the data generated in the secret state computing are all stored in the heterogeneous chip memory.
[0034] The heterogeneous accelerated computing optimization method is applied to a first device participating in federated learning, and the set of secret state computing operators includes an encryption operator, a first secret state computing operator, and a second secret state computing operator.
[0035] The step of performing a secret state calculation resident in the memory of the heterogeneous chip on the plaintext data based on the secret state calculation operator set to obtain a secret state calculation result comprises:
[0036] Step S21, based on the encryption operator, homomorphically encrypt the plaintext data to obtain ciphertext data, and write the ciphertext data into the heterogeneous chip memory;
[0037] In this embodiment, it should be noted that the heterogeneous accelerated computing optimization method is applied to a first device participating in federated learning, and the second device is another participant in the federated learning.
[0038] Based on the encryption operator, the plaintext data is homomorphically encrypted to obtain ciphertext data, and the ciphertext data is written into the heterogeneous chip memory. Specifically, the encryption operator is called through the upper-layer computing module to map the plaintext data into ciphertext data, and the ciphertext data is written into the heterogeneous chip memory.
[0039] Step S22: based on the memory copy operator set, copy the second ciphertext data sent by the second device to the underlying CPU memory to the heterogeneous chip memory;
[0040] In this embodiment, based on the memory copy operator set, the second ciphertext data sent by the second device to the underlying CPU memory is copied to the heterogeneous chip memory. Specifically, the second ciphertext data sent by the second device participating in federated learning is received and written into the underlying CPU memory, and then the memory copy operator in the memory copy operator set is called through the upper-layer computing module to map the second ciphertext data from the underlying CPU memory to the heterogeneous chip memory, and the data format of the second ciphertext data is converted into a data format suitable for the heterogeneous chip memory, and then the second ciphertext data is written into the heterogeneous chip memory.
[0041] Step S23, performing a first secret state calculation on the ciphertext data and the second ciphertext data by using the first secret state calculation operator to obtain an intermediate secret state calculation result, and writing the intermediate secret state calculation result into the heterogeneous chip memory;
[0042] Step S24, calling the intermediate secret state calculation result in the heterogeneous chip memory, and performing a second secret state calculation on the intermediate secret state calculation result through the second secret state calculation operator to obtain the secret state calculation result, and writing the secret state calculation result into the heterogeneous chip memory.
[0043] In this embodiment, the first secret state calculation operator is called by the upper-layer calculation module, and the first secret state calculation is performed on the ciphertext data and the second ciphertext data to obtain an intermediate secret state calculation result, and the intermediate secret state calculation result is written into the heterogeneous chip memory. Further, the intermediate secret state calculation result is called by the memory control module, and the second secret state calculation operator is called by the upper-layer calculation module to perform a second secret state calculation on the intermediate secret state calculation result to obtain the secret state calculation result, and the secret state calculation result is written into the heterogeneous chip memory. For example, assuming that the product of the secret state matrices A, B and C needs to be calculated, where A and C are the first ciphertext data and B is the second ciphertext data, the intermediate secret state calculation result A*B is obtained by the first secret state calculation, and A*B is written into the heterogeneous chip memory. Then, when A*B*C needs to be calculated, A*B and C are called in the heterogeneous chip memory to perform a second secret state calculation to obtain A*B*C.
[0044] Step S30, feeding back the secret state calculation result to the CPU memory.
[0045] In this embodiment, the secret state calculation result is fed back to the CPU memory. Specifically, through the upper-level computing module, the decryption operator in the secret state calculation operator set is called to decrypt the secret state calculation result to obtain the plain text calculation result, and then through the upper-level computing module, the second memory copy operator is called to copy the plain text calculation result from the heterogeneous chip memory to the underlying CPU memory.
[0046] The step of feeding back the secret state calculation result to the CPU memory includes:
[0047] Step A10, decrypting the secret calculation result in the memory of the heterogeneous chip to obtain a plaintext calculation result;
[0048] In this embodiment, the secret state calculation result is decrypted in the memory of the heterogeneous chip to obtain a plain text calculation result. Specifically, the upper layer calculation module calls a decryption operator to map the secret state calculation result into a plain text calculation result.
[0049] Step A20, based on the second memory copy operator set, copy the plaintext calculation result from the heterogeneous chip memory to the underlying CPU memory.
[0050] In this embodiment, based on the second memory copy operator set, the plaintext calculation result is copied from the heterogeneous chip memory to the underlying CPU memory. Specifically, through the upper-layer computing module, the second memory copy operator in the second memory copy operator set is called to map the plaintext calculation result from the heterogeneous chip memory to the underlying CPU memory. Furthermore, the data format of the plaintext calculation result in the underlying CPU memory is converted into a data format adapted to the application layer, and written into the upper-layer CPU memory to obtain the target heterogeneous accelerated calculation result. The entire calculation process from plaintext data to the target heterogeneous accelerated calculation result can be represented by a five-stage pipeline, as shown in FIG. Figure 2 The figure shows a schematic diagram of the five-stage pipeline, which is composed of a first data conversion module, a first data copy module, a secret computing module, a second data copy module and a second data conversion module in sequence, wherein the first data conversion module is used to convert data in the application layer into data in a data format adapted to the C language layer, the first data copy module is used to copy data from the CPU memory to the heterogeneous chip memory, the secret computing module is used to perform secret computing on the data, the second data copy module is used to copy data from the heterogeneous chip memory to the CPU memory, and the second data conversion module is used to convert data in the C language layer into data in a data format adapted to the application layer.
[0051] Furthermore, the step S30 further includes:
[0052] Step B10: Based on the second memory copy operator set, the secret state calculation result is copied from the heterogeneous chip memory to the CPU memory.
[0053] In this embodiment, based on the second memory copy operator set, the secret state calculation result is copied from the heterogeneous chip memory to the CPU memory. Specifically, through the upper-level computing module, the second memory copy operator in the second memory copy operator set is called to map the secret state calculation result from the heterogeneous chip memory to the underlying CPU memory.
[0054] In one embodiment, Figure 3The figure shows a schematic diagram of a heterogeneous federated learning framework in federated learning, wherein the heterogeneous federated learning framework includes a CPU side and a GPU side, the upper-layer model is the upper layer of the CPU side, the CPU side in the heterogeneous framework is the bottom layer of the CPU side, the computing module is the upper-layer computing module, the computing interface is the preset computing interface, the CPU memory in the upper-layer model is the upper-layer CPU memory, the CPU memory in the heterogeneous framework is the bottom-layer CPU memory, the GPU operator includes the secret computing operator, the encryption operator and the decryption operator, etc., the resource scheduling module is used to schedule system resources in the heterogeneous accelerated computing process, and the memory control module includes the memory copy operator.
[0055] The embodiment of the present application provides a heterogeneous accelerated computing optimization method. Compared with the technical means used in the prior art to copy ciphertext data back and forth between the CPU memory and the heterogeneous chip memory multiple times during a round of iterations of federated learning, the embodiment of the present application first obtains plaintext data, and based on a first memory copy operator set, copies the plaintext data from the CPU memory to the heterogeneous chip memory, wherein the purpose of separating the memory copy process from the secret state computing process is achieved by encapsulating the memory copy separately as an operator, and the plaintext data, rather than the ciphertext data, is copied between the heterogeneous chip memory and the CPU memory, and since the number of bits of the plaintext data is much smaller than the ciphertext data, the time consumption of the memory copy is reduced, wherein it should be noted that the memory copy and the secret state computing are usually encapsulated in the same operator, and the operator execution process is usually the alternation of secret state computing and memory copy, and finally the final secret state computing result is obtained, wherein the intermediate ciphertext data generated by multiple secret state computing all need to be written to the CPU memory, thereby resulting in multiple memory copies of the ciphertext data between the heterogeneous chip memory and the CPU memory, and further, based on the secret state computing The operator set is calculated, and the secret state calculation resident in the memory of the heterogeneous chip is performed on the plaintext data to obtain the secret state calculation result. That is, after the memory copy is separated from the secret state calculation, the ciphertext data is resident in the memory of the heterogeneous chip for secret state calculation. The intermediate ciphertext data generated in the secret state calculation process are all written into the memory of the heterogeneous chip until the final secret state calculation result is calculated, thereby avoiding multiple memory copy processes in the secret state calculation process. In one round of iteration of federated learning, only one back-and-forth memory copy is required between the CPU memory and the heterogeneous chip memory, so the number of memory copies is reduced, and the time consumption of memory copy in the federated learning process is further reduced. Finally, the secret state calculation result is fed back to the CPU memory to complete one round of iterative calculation process of federated learning, which overcomes the problem that the number of bits of ciphertext data is usually high, and the memory copy time of ciphertext data is long, which will result in that in one round of iteration of federated learning, the memory copy time of ciphertext data will be much longer than the secret state calculation time of ciphertext data, making the calculation efficiency of the heterogeneous federated learning framework low. Therefore, the calculation efficiency of the heterogeneous federated learning framework is improved.
[0056] Further, refer to Figure 4 Based on the first embodiment of the present application, in another embodiment of the present application, further, in step S10, the plaintext data includes at least a plaintext subset data,
[0057] The step of obtaining plaintext data comprises:
[0058] Step C10, obtaining original data in the upper CPU memory, and grouping the original data to obtain original subset data;
[0059] In this embodiment, it should be noted that the original data is an original matrix in the federated model, for example, a model input matrix and a model output matrix.
[0060] The original data is obtained in the upper CPU memory, and the original data is grouped to obtain each original subset data. Specifically, the original matrix data is obtained in the upper CPU memory, and the original matrix is split into a plurality of original sub-matrices with similar data amounts, no data dependency and no return address conflict, and then each of the original sub-matrices is used as the original subset data.
[0061] Step C20, allocating a parallel pipeline to each of the original subset data, and converting the original subset data on each of the parallel pipelines into a data format that complies with the underlying CPU memory, to obtain each plaintext subset data.
[0062] In this embodiment, it should be noted that in order to further improve the computing efficiency of the heterogeneous federated learning framework, multiple parallel pipelines are designed in parallel, wherein each of the parallel pipelines at least includes a memory copy from the underlying CPU memory to the heterogeneous chip memory, a secret state calculation, and a memory copy from the heterogeneous chip memory to the underlying CPU memory. In an practicable manner, Figure 5 The figure shows a design schematic diagram of the multiple parallel pipelines, wherein the GPU is the heterogeneous chip, and the scheduling module is used to allocate computing resources to each parallel pipeline to ensure that the time consumption of each parallel pipeline is the same, so that the pipeline will not be suspended due to insufficient computing resources, so that the GPU can complete computing tasks and data copying tasks faster, thereby reducing the time that the ciphertext data resides in the GPU, thereby preventing too much ciphertext data from residing in the GPU and affecting the computing performance of the GPU, thereby achieving the purpose of hiding the impact of data copying on the heterogeneous framework.
[0063] Allocate a parallel pipeline for each of the original subset data, and convert the original subset data on each of the parallel pipelines into a data format that conforms to the underlying CPU memory to obtain each plaintext subset data. Specifically, perform the following steps in parallel for each of the original subset data:
[0064] A parallel pipeline is allocated to the original subset data, and based on the first data conversion module on the parallel pipeline, data format conversion is performed on the original subset data to convert the original subset data into data in a data format suitable for the underlying CPU memory, so as to obtain plaintext subset data corresponding to the original subset data.
[0065] Further, in step S20, the memory copy operator set includes at least one memory copy operator on a parallel pipeline, the plaintext data includes at least one plaintext subset data,
[0066] The step of copying the plaintext data from the underlying CPU memory to the heterogeneous chip memory based on the memory copy operator set includes:
[0067] Step C30, by calling the memory copy operator on each of the parallel pipelines, each of the plaintext subset data is copied in parallel from the underlying CPU memory to the heterogeneous chip memory.
[0068] In this embodiment, by calling the memory copy operator on each of the parallel pipelines, each of the plaintext subset data is copied in parallel from the underlying CPU memory to the heterogeneous chip memory. Specifically, the memory copy operator on each of the parallel pipelines is called by the upper-layer computing module, and each of the plaintext subset data is mapped in parallel from the underlying CPU memory to the heterogeneous chip memory, and after converting the data format of each of the plaintext subset data into a data format suitable for the heterogeneous chip memory, each of the plaintext subset data is written into the heterogeneous chip memory.
[0069] Further, in step S30, the set of dense state computing operators includes at least one dense state computing operator on the parallel pipeline, and the dense state computing result includes at least one dense state subset computing result.
[0070] The step of performing a secret state calculation resident in the memory of the heterogeneous chip on the plaintext data based on the secret state calculation operator set to obtain a secret state calculation result comprises:
[0071] Step C40, by calling the secret state calculation operators on each of the parallel pipelines, the secret state calculations resident in the memory of the heterogeneous chip are performed in parallel on each of the plaintext subset data to obtain the calculation results of each secret state subset.
[0072] In this embodiment, by calling the secret state computing operators on each of the parallel pipelines, the secret state computing resident in the memory of the heterogeneous chip is performed in parallel on each of the plaintext subset data to obtain each secret state subset computing result. Specifically, the encryption operators on each of the parallel pipelines are called by the upper-level computing module to encrypt each of the plaintext subset data in parallel to obtain each ciphertext subset data, and then the secret state computing operators on each of the parallel pipelines are called by the upper-level computing module to perform secret state computing on each of the ciphertext subset data in parallel, so as to reduce the residence time of the ciphertext data resident on the heterogeneous chip through parallel computing, and then obtain each secret state subset computing result.
[0073] Further, after step S30, the secret state calculation result includes at least a secret state subset calculation result, the CPU memory includes an upper CPU memory and a bottom CPU memory,
[0074] After the step of feeding back the secret state calculation result to the CPU memory, the heterogeneous accelerated computing optimization method further includes:
[0075] Step C50, converting the calculation results of each of the dense state subsets into a data format that conforms to the upper-layer CPU memory, and obtaining the calculation results of each target dense state subset, wherein the calculation results of the dense state subset conform to the data format of the bottom-layer CPU memory;
[0076] In this embodiment, each of the dense state subset calculation results is converted into a data format that conforms to the upper-level CPU memory to obtain each target dense state subset calculation result, wherein the dense state subset calculation result conforms to the data format of the underlying CPU memory. Specifically, through the second data conversion module on each of the parallel pipelines, each of the dense state subset calculation results is converted in parallel from a data format that conforms to the underlying CPU memory to a data format that conforms to the upper-level CPU memory to obtain each target dense state subset calculation result.
[0077] Step C60, when the calculation results of each target secret state subset in the upper CPU memory meet the preset calculation end condition, the calculation results of each target secret state subset are integrated to obtain the target secret state calculation result.
[0078] In this embodiment, when the calculation results of each target dense state subset in the upper CPU memory meet the preset calculation end conditions, the calculation results of each target dense state subset are integrated to obtain the target dense state calculation result. Specifically, when the calculation results of each target dense state subset in the upper CPU memory meet the preset calculation end conditions, it is proved that the number of calculation results of each target dense state subset is consistent with the number of original sub-matrices originally split, and then the calculation results of each target dense state subset are integrated to obtain the target dense state calculation result corresponding to the original matrix.
[0079] The embodiment of the present application provides a heterogeneous accelerated computing optimization method based on multiple parallel pipelines, that is, firstly, the original data is obtained in the upper CPU memory, and the original data is grouped to obtain each original subset data, and then a parallel pipeline is allocated to each original subset data, and data conversion, memory copy and secret state calculation are performed in parallel on the corresponding original subset data in each parallel pipeline to obtain each secret state subset calculation result, and then each secret state subset calculation result is integrated in the CPU memory to obtain the final secret state calculation result, wherein, since secret state calculation and data copy are performed based on multiple parallel pipelines, compared with the method of performing secret state calculation and data copy based on one pipeline, in the embodiment of the present application, during one round of federated learning iteration, the time that the ciphertext data resides in the heterogeneous chip memory will be shortened, thereby improving the data carrying capacity of the heterogeneous chip memory during pipeline operation, avoiding the situation where the secret state calculation cannot be performed or the computing efficiency of the secret state calculation is reduced due to too much ciphertext data stored in the heterogeneous chip memory, thereby improving the computing efficiency of the heterogeneous federated learning framework.
[0080] Reference Figure 6 , Figure 6 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present application.
[0081] like Figure 6 As shown, the heterogeneous accelerated computing optimization device may include: a processor 1001, such as a CPU, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to realize the connection and communication between the processor 1001 and the memory 1005. The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0082] Optionally, the heterogeneous accelerated computing optimization device may also include a rectangular user interface, a network interface, a camera, an RF (Radio Frequency) circuit, a sensor, an audio circuit, a WiFi module, etc. The rectangular user interface may include a display screen (Display), an input submodule such as a keyboard (Keyboard), and the optional rectangular user interface may also include a standard wired interface and a wireless interface. The network interface may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0083] Those skilled in the art will understand that Figure 6The heterogeneous accelerated computing optimization device structure shown in the figure does not constitute a limitation on the heterogeneous accelerated computing optimization device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0084] like Figure 6 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, and a heterogeneous accelerated computing optimization program. The operating system is a program that manages and controls the hardware and software resources of the heterogeneous accelerated computing optimization device, and supports the operation of the heterogeneous accelerated computing optimization program and other software and / or programs. The network communication module is used to realize the communication between the components inside the memory 1005, and the communication with other hardware and software in the heterogeneous accelerated computing optimization system.
[0085] exist Figure 6 In the heterogeneous accelerated computing optimization device shown, the processor 1001 is used to execute the heterogeneous accelerated computing optimization program stored in the memory 1005 to implement the steps of any of the above-mentioned heterogeneous accelerated computing optimization methods.
[0086] The specific implementation methods of the heterogeneous accelerated computing optimization device of the present application are basically the same as the above-mentioned embodiments of the heterogeneous accelerated computing optimization method, and will not be repeated here.
[0087] Reference Figure 7 The embodiment of the present application further provides a heterogeneous accelerated computing optimization device, which is applied to a heterogeneous accelerated computing optimization device, and the heterogeneous accelerated computing optimization device includes:
[0088] A memory copy module, used to obtain plaintext data, and based on a first memory copy operator set, copy the plaintext data from the CPU memory to the heterogeneous chip memory;
[0089] A secret computing module, used to perform secret computing resident in the memory of the heterogeneous chip on the plaintext data based on a secret computing operator set to obtain a secret computing result;
[0090] A feedback module is used to feed back the secret state calculation result to the CPU memory.
[0091] Optionally, the feedback module is further used for:
[0092] Decrypting the secret calculation result in the heterogeneous chip memory to obtain a plaintext calculation result; based on a second memory copy operator set, copying the plaintext calculation result from the heterogeneous chip memory to the underlying CPU memory; and / or
[0093] Based on the second memory copy operator set, the secret state calculation result is copied from the heterogeneous chip memory to the CPU memory.
[0094] Optionally, the memory copy module is further used for:
[0095] Acquire original data in the upper CPU memory, and group the original data to obtain original subset data;
[0096] A parallel pipeline is allocated to each of the original subset data, and the original subset data on each of the parallel pipelines is converted into a data format that conforms to the underlying CPU memory to obtain each plaintext subset data.
[0097] Optionally, the memory copy module is further used for:
[0098] By calling the memory copy operator on each of the parallel pipelines, each of the plaintext subset data is copied in parallel from the underlying CPU memory to the heterogeneous chip memory.
[0099] Optionally, the secret state computing module is further used for:
[0100] By calling the secret state calculation operators on each of the parallel pipelines, secret state calculations resident in the memory of the heterogeneous chip are performed in parallel on each of the plaintext subset data to obtain calculation results of each secret state subset.
[0101] Optionally, the heterogeneous accelerated computing optimization device is further used for:
[0102] Respectively converting the calculation results of each of the dense state subsets into a data format that conforms to the data format of the upper-layer CPU memory to obtain calculation results of each target dense state subset, wherein the calculation results of the dense state subset conform to the data format of the bottom-layer CPU memory;
[0103] When the calculation results of each target secret state subset in the upper CPU memory meet the preset calculation end condition, the calculation results of each target secret state subset are integrated to obtain the target secret state calculation result.
[0104] Optionally, the secret state computing module is further used for:
[0105] Based on the encryption operator, homomorphically encrypt the plaintext data to obtain ciphertext data, and write the ciphertext data into the heterogeneous chip memory;
[0106] Based on the memory copy operator set, copy the second ciphertext data sent by the second device to the underlying CPU memory to the heterogeneous chip memory;
[0107] Performing a first secret state calculation on the ciphertext data and the second ciphertext data through the first secret state calculation operator to obtain an intermediate secret state calculation result, and writing the intermediate secret state calculation result into the heterogeneous chip memory;
[0108] The intermediate secret state calculation result is called in the heterogeneous chip memory, and the second secret state calculation operator is used to perform a second secret state calculation on the intermediate secret state calculation result to obtain the secret state calculation result, and the secret state calculation result is written into the heterogeneous chip memory.
[0109] The specific implementation of the heterogeneous accelerated computing optimization device of the present application is basically the same as the various embodiments of the above-mentioned heterogeneous accelerated computing optimization method, and will not be repeated here.
[0110] An embodiment of the present application provides a readable storage medium, and the readable storage medium stores one or more programs, and the one or more programs can also be executed by one or more processors to implement the steps of any of the heterogeneous accelerated computing optimization methods described above.
[0111] The specific implementation of the readable storage medium of the present application is basically the same as the above-mentioned embodiments of the heterogeneous accelerated computing optimization method, and will not be repeated here.
[0112] An embodiment of the present application provides a computer program product, and the computer program product includes one or more computer programs, and the one or more computer programs can also be executed by one or more processors to implement the steps of any of the heterogeneous accelerated computing optimization methods described above.
[0113] The specific implementation methods of the computer program product of the present application are basically the same as the above-mentioned embodiments of the heterogeneous accelerated computing optimization method, and will not be repeated here.
[0114] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.
Claims
1. A heterogeneous accelerated computing optimization method, characterized in that: The heterogeneous accelerated computing optimization method is applied to a heterogeneous federated learning framework, which includes a CPU side and a heterogeneous chip side. The upper layer of the CPU side is an application layer, and the bottom layer of the CPU side is a C language layer. The CPU memory includes an upper CPU memory and a bottom CPU memory. The CPU side has an upper CPU memory corresponding to the upper layer of the CPU side, and a bottom CPU memory corresponding to the bottom layer of the CPU side. The heterogeneous chip side has a heterogeneous chip memory; memory copying is performed between the upper CPU memory and the bottom CPU memory, and memory copying is performed between the heterogeneous chip memory and the bottom CPU memory; The heterogeneous accelerated computing optimization method comprises: Obtain plaintext data, and based on a first memory copy operator set, copy the plaintext data from the CPU memory to the heterogeneous chip memory, wherein the first memory copy operator set includes at least a memory copy operator, and the memory copy operator is used to map data from the underlying CPU memory to the heterogeneous chip memory; Based on the set of secret computing operators, performing secret computing resident in the memory of the heterogeneous chip on the plaintext data to obtain a secret computing result; Feeding back the secret state calculation result to the CPU memory; The step of obtaining plaintext data and copying the plaintext data from the CPU memory to the heterogeneous chip memory based on the first memory copy operator set includes: Determine original data in the upper CPU memory, and convert the original data into data adapted to the C language layer to obtain plaintext data, wherein the original data at least includes model parameters, model output data, and model input data in the federated learning process; The upper-layer computing module on the CPU side calls the memory copy operator encapsulated by the preset computing interface in the memory copy operator set, copies the plaintext data from the bottom-layer CPU memory to the heterogeneous chip memory, and converts the data format of the plaintext data into a data format suitable for the heterogeneous chip memory through the memory control module, and then writes it to the heterogeneous chip memory; wherein, The memory control module is used to perform memory operations, serialization operations, data conversion and site communication on the underlying CPU memory and the heterogeneous chip memory, wherein the memory operations include transposition, cutting and memory copying, and the serialization operations include converting data into a bit stream; The step of performing a secret state calculation resident in the memory of the heterogeneous chip on the plaintext data based on the secret state calculation operator set to obtain a secret state calculation result comprises: Through the upper-level computing module, the encryption operator in the secret state computing operator set is called to encrypt the plaintext data to obtain the ciphertext data, and after the ciphertext data is written into the heterogeneous chip memory through the memory control module, the secret state computing operator in the secret state computing operator set is called through the upper-level computing module to perform secret state computing on the ciphertext data that resides in the heterogeneous chip memory to obtain the secret state computing result, wherein the data generated in the secret state computing are all stored in the heterogeneous chip memory.
2. The heterogeneous accelerated computing optimization method according to claim 1, characterized in that: The step of feeding back the secret state calculation result to the CPU memory comprises: Decrypting the secret calculation result in the heterogeneous chip memory to obtain a plaintext calculation result; based on a second memory copy operator set, copying the plaintext calculation result from the heterogeneous chip memory to the underlying CPU memory; and / or Based on the second memory copy operator set, the secret state calculation result is copied from the heterogeneous chip memory to the CPU memory.
3. The heterogeneous accelerated computing optimization method according to claim 1, characterized in that: The plaintext data at least includes a plaintext subset data, The step of obtaining plaintext data comprises: Acquire original data in the upper CPU memory, and group the original data to obtain original subset data; A parallel pipeline is allocated to each of the original subset data, and the original subset data on each of the parallel pipelines is converted into a data format that conforms to the underlying CPU memory to obtain each plaintext subset data.
4. The heterogeneous accelerated computing optimization method according to claim 1, characterized in that: The memory copy operator set includes at least one memory copy operator on a parallel pipeline, the plaintext data includes at least one plaintext subset data, Based on the memory copy operator set, the steps of copying the plaintext data from the underlying CPU memory to the heterogeneous chip memory include: By calling the memory copy operator on each of the parallel pipelines, each of the plaintext subset data is copied in parallel from the underlying CPU memory to the heterogeneous chip memory.
5. The heterogeneous accelerated computing optimization method according to claim 4, characterized in that: The set of dense state computing operators includes at least one dense state computing operator on the parallel pipeline, and the dense state computing result includes at least one dense state subset computing result. The step of performing a secret state calculation resident in the memory of the heterogeneous chip on the plaintext data based on the secret state calculation operator set to obtain a secret state calculation result comprises: By calling the secret state calculation operators on each of the parallel pipelines, secret state calculations resident in the memory of the heterogeneous chip are performed in parallel on each of the plaintext subset data to obtain calculation results of each secret state subset.
6. The heterogeneous accelerated computing optimization method according to claim 1, characterized in that: The secret state calculation result includes at least a secret state subset calculation result, the CPU memory includes an upper CPU memory and a bottom CPU memory, After the step of feeding back the secret state calculation result to the CPU memory, the heterogeneous accelerated computing optimization method further includes: Respectively converting the calculation results of each of the dense state subsets into a data format that conforms to the data format of the upper-layer CPU memory to obtain calculation results of each target dense state subset, wherein the calculation results of the dense state subset conform to the data format of the bottom-layer CPU memory; When the calculation results of each target secret state subset in the upper CPU memory meet the preset calculation end condition, the calculation results of each target secret state subset are integrated to obtain the target secret state calculation result.
7. The heterogeneous accelerated computing optimization method according to claim 1, characterized in that: The heterogeneous accelerated computing optimization method is applied to a first device participating in federated learning, the set of secret state computing operators includes an encryption operator, a first secret state computing operator, and a second secret state computing operator, The step of performing a secret state calculation resident in the memory of the heterogeneous chip on the plaintext data based on the secret state calculation operator set to obtain a secret state calculation result comprises: Based on the encryption operator, homomorphically encrypt the plaintext data to obtain ciphertext data, and write the ciphertext data into the heterogeneous chip memory; Based on the memory copy operator set, copy the second ciphertext data sent by the second device to the underlying CPU memory to the heterogeneous chip memory; Performing a first secret state calculation on the ciphertext data and the second ciphertext data through the first secret state calculation operator to obtain an intermediate secret state calculation result, and writing the intermediate secret state calculation result into the heterogeneous chip memory; The intermediate secret state calculation result is called in the heterogeneous chip memory, and the second secret state calculation operator is used to perform a second secret state calculation on the intermediate secret state calculation result to obtain the secret state calculation result, and the secret state calculation result is written into the heterogeneous chip memory.
8. A heterogeneous accelerated computing optimization device, characterized in that: The heterogeneous accelerated computing optimization device is applied to a heterogeneous federated learning framework, which includes a CPU side and a heterogeneous chip side. The upper layer of the CPU side is an application layer, and the bottom layer of the CPU side is a C language layer. The CPU memory includes an upper CPU memory and a bottom CPU memory. The CPU side has an upper CPU memory corresponding to the upper layer of the CPU side, and a bottom CPU memory corresponding to the bottom layer of the CPU side. The heterogeneous chip side has a heterogeneous chip memory; memory copying is performed between the upper CPU memory and the bottom CPU memory, and memory copying is performed between the heterogeneous chip memory and the bottom CPU memory; The heterogeneous accelerated computing optimization device comprises: A memory copy module, used to obtain plaintext data, and based on a first memory copy operator set, copy the plaintext data from the CPU memory to the heterogeneous chip memory, wherein the first memory copy operator set includes at least a memory copy operator, and the memory copy operator is used to map data from the underlying CPU memory to the heterogeneous chip memory; A secret computing module, used to perform secret computing resident in the memory of the heterogeneous chip on the plaintext data based on a secret computing operator set to obtain a secret computing result; A feedback module, used for feeding back the secret state calculation result to the CPU memory; The memory copy module is further used to determine the original data in the upper CPU memory, and convert the original data into data adapted to the C language layer to obtain plaintext data, wherein the original data at least includes model parameters, model output data and model input data in the federated learning process; the upper computing module on the CPU side calls the memory copy operator encapsulated by the preset computing interface in the memory copy operator set, copies the plaintext data from the bottom CPU memory to the heterogeneous chip memory, and converts the data format of the plaintext data into a data format adapted to the heterogeneous chip memory through the memory control module, and then writes it to the heterogeneous chip memory; wherein the memory control module is used to perform memory operations, serialization operations, data conversion and site communication on the bottom CPU memory and the heterogeneous chip memory, wherein the memory operations include transposition, cutting and memory copying, and the serialization operations include converting data into a bit stream; The secret computing module is also used to call the encryption operator in the secret computing operator set through the upper-level computing module to encrypt the plaintext data to obtain the ciphertext data, and after writing the ciphertext data into the heterogeneous chip memory through the memory control module, call the secret computing operator in the secret computing operator set through the upper-level computing module to perform secret computing on the ciphertext data that resides in the heterogeneous chip memory to obtain the secret computing result, wherein the data generated in the secret computing are all stored in the heterogeneous chip memory.
9. A heterogeneous accelerated computing optimization device, characterized in that: The heterogeneous accelerated computing optimization device includes: a memory, a processor, and a program stored in the memory for implementing the heterogeneous accelerated computing optimization method. The memory is used to store a program for implementing a heterogeneous accelerated computing optimization method; The processor is used to execute a program for implementing the heterogeneous accelerated computing optimization method to implement the steps of the heterogeneous accelerated computing optimization method as described in any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a program for implementing a heterogeneous accelerated computing optimization method, and the program for implementing a heterogeneous accelerated computing optimization method is executed by a processor to implement the steps of the heterogeneous accelerated computing optimization method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and device of heterogeneous computing platform and readable storage medium
CN111143272A