Data synchronization method, system, electronic device, and computer-readable storage medium
By optimizing partition storage and data transmission parameters, parallel data synchronization between global memory and local memory is achieved, solving the problem of low data synchronization efficiency and improving data processing efficiency.
Patent Information
- Application Number
- CN202510238300.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-02-28
AI Technical Summary
In the prior art, data synchronization between global memory and local memory is inefficient and cannot meet the user's demand for efficient data processing.
By partitioning the raw data, each partition serves as a channel for external data reading and writing. The correspondence between the number of storage areas and local memory is determined based on the data dimensions and the total number of data blocks. Data transmission parameters are determined based on the communication parameters between the global memory and the data transmission interface, thereby achieving parallel data synchronization between the global memory and the local memory.
Improves the data synchronization efficiency between global memory and local memory, ensuring the effectiveness and efficiency of data synchronization.
Smart Images

Figure CN120029553B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, and in particular to a data synchronization method and system, electronic equipment and computer readable storage medium. BACKGROUND
[0002] With the increasing size of data that users need to process, users have higher and higher requirements for data processing speed. Currently, different types of computing units are used in the same computing system to complete tasks together, and data that occupies a large amount of computing resources is sent to an acceleration computing unit with parallel capabilities for processing.
[0003] In related technologies, when an acceleration computing unit is used to perform a task, a data transmission interface first obtains original data issued by a host from a global memory, and then transmits the original data to a local memory for processing according to address information carried by the original data, and finally sends the processed data to the global memory. This data synchronization method between the global memory and the local memory is inefficient and cannot meet the user's efficient data processing needs. SUMMARY
[0004] The present application provides a data synchronization method and system, electronic equipment and computer readable storage medium, which can realize parallel data synchronization between global memory and local memory, and effectively improve data synchronization efficiency.
[0005] To solve the above technical problems, the present application provides the following technical solutions:
[0006] In one aspect, the present application provides a data synchronization method applied to a first computing unit including a global memory and a plurality of local memories. The global memory of the first computing unit includes a plurality of storage areas. The method comprises: when a second computing unit sends a data block of to-be-processed original data to a corresponding storage area, determining a quantity correspondence relationship between each storage area and a local memory based on a data amount of the to-be-processed original data and a total number of data blocks; determining data transmission parameters between the global memory and each local memory according to communication parameters of the global memory and a data transmission interface; reading data matching the quantity correspondence relationship according to the data transmission parameters, and performing corresponding read-write processing according to address signals; and wherein the data transmission interface is a port between the global memory and each local memory.
[0007] Another aspect of the present application provides a data synchronization method applied to a second computing unit, comprising: obtaining raw data to be processed, and dividing the raw data to be processed into a plurality of data blocks; sending each data block to a corresponding storage area of a global memory of a first computing unit, so that the first computing unit determines a quantity correspondence between each storage area and a local memory based on a data volume of the raw data to be processed and a total number of data blocks; determining data transmission parameters between the global memory and each local memory according to communication parameters of the global memory and a data transmission interface; and reading data matching the quantity correspondence according to the data transmission parameters, and performing corresponding read / write processing according to address signals.
[0008] The present application also provides an electronic device comprising a memory and a processor, wherein the processor is configured to implement the steps of any of the above data synchronization methods when executing a computer program stored in the memory.
[0009] The present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is configured to implement the steps of any of the above data synchronization methods when executed by a processor.
[0010] The present application also provides a data synchronization system comprising at least a first computing unit and a second computing unit, wherein the first computing unit comprises an internal memory, a plurality of computing sub-units, and a data synchronization processor; the internal memory comprises a global memory and a plurality of local memories; each computing sub-unit corresponds to a local memory, and each computing sub-unit only interacts with the corresponding local memory; the first computing unit implements the steps of the data synchronization method applied to the first computing unit when executing a computer program; and the second computing unit implements the steps of the data synchronization method applied to the second computing unit when executing a computer program.
[0011] The technical solution provided by the present application has the following advantages: the raw data to be processed by the first computing unit is stored in partitions, each partition can be used as a channel for external data read / write, and the number of local memories corresponding to each storage area can be determined according to the data dimension of the raw data and the total number of divided data blocks, so that the raw data can be transmitted from each storage area of the global memory to the local memories through multiple channels at the same time, thereby effectively improving the data synchronization efficiency between the global memory and the local memories. In addition, during the data synchronization process between the global memory and the local memories, the communication capability between the global memory and the data transmission interface is also considered, which not only helps to maximize the efficiency of data synchronization, but also ensures the effectiveness of data synchronization. In addition, the present application also provides a corresponding implementation system, an electronic device, and a computer-readable storage medium for the data synchronization method, which further makes the method more practical, and the system, the electronic device, and the computer-readable storage medium have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 A schematic diagram of the system structure in an exemplary application scenario provided by the present invention;
[0014] Figure 2 A flowchart of a data synchronization method in related art;
[0015] Figure 3 A flowchart of a data synchronization method provided by the present invention;
[0016] Figure 4 A schematic diagram of segmenting the raw data to be processed provided by the present invention in an exemplary application scenario;
[0017] Figure 5 A schematic diagram of storing the raw data to be processed provided by the present invention in an exemplary application scenario;
[0018] Figure 6 A flowchart of another data synchronization method provided by the present invention;
[0019] Figure 7 A structural framework diagram of an exemplary embodiment of the data synchronization device provided by the present invention;
[0020] Figure 8 A structural framework diagram of an exemplary embodiment of the data synchronization device provided by the present invention;
[0021] Figure 9 A schematic structural diagram of an exemplary embodiment of a data synchronization system provided by the present invention;
[0022] Figure 10 A schematic diagram of a framework of a global memory provided by the present invention in an exemplary application scenario;
[0023] Figure 11 A schematic diagram of a framework of the data synchronization processor provided by the present invention in an exemplary application scenario;
[0024] Figure 12 A schematic diagram of data reading and writing in an exemplary application scenario of the data synchronization system provided by the present invention;
[0025] Figure 13A protocol conversion schematic diagram of the data synchronization system provided by the present application in an exemplary application scenario is shown in FIG. 1.
[0026] Figure 14 A framework schematic diagram of the data synchronization system provided by the present application in an exemplary application scenario is shown in FIG. 2. DETAILED DESCRIPTION
[0027] In order to make the personnel in the technical field better understand the technical solutions of the present application, the present application is further described in detail below in combination with the drawings and specific embodiments. In the specification and the above drawings, the terms "first", "second", "third", "fourth" and the like are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "as an example, embodiment or illustration". Any embodiment described as "exemplary" herein is not necessarily interpreted as superior or better than other embodiments.
[0028] In order to improve the efficiency of computing task execution and enhance the data computing performance, a heterogeneous computing mode using different types of processors or computing units in the same computing system to jointly complete computing tasks is born. A heterogeneous computing platform composed of multiple different types of computing units with different processing capabilities, energy efficiency and application scope is used to execute tasks, such as a heterogeneous computing platform composed of CPU (Central Processing Unit) and FPGA (Field Programmable Gate Array).
[0029] When the heterogeneous computing platform executes tasks such as high-performance computing tasks or tasks with high real-time requirements or tasks containing large-scale data requiring computation, it will send operation tasks requiring a large amount of computing resources or high complexity to an acceleration computing unit with parallel computing capability for processing. The host will transmit original data requiring processing by the acceleration computing unit to the global memory of the acceleration computing unit, and after the acceleration computing unit finishes processing the data, the host reads the data processing result in the global memory of the acceleration computing unit. In order to improve the computing speed of the acceleration computing unit, the acceleration computing unit will instantiate multiple computing sub-units, each of which is uniquely configured with a local memory, and each computing sub-unit can only access data from the local memory. In order to determine the accurate execution of data tasks, the acceleration computing unit needs to perform data synchronization on the global memory and each local memory when processing data.
[0030] The related art defines a data transmission interface in advance to sequentially take out original data from the global memory, and then the data transmission interface transmits the original data to the corresponding local memory according to the address of the original data in the global memory. Each computing core reads data from the corresponding local memory for calculation, and after the data calculation is completed, the data transmission interface sequentially takes back the data from each local memory and stores it in the global memory, thereby realizing data synchronization between the global memory and the local memory. As can be seen, there is only one data transmission interface between data synchronization, and all the data of the local memory is transmitted from one data transmission interface, which leads to low data transmission efficiency and affects the data processing efficiency of the entire acceleration computing unit, and cannot meet the user's efficient data processing demand.
[0031] Therefore, in order to solve the problems in the related art, the original data to be processed by the first computing unit is stored in partitions, each partition can be used as a channel for external data reading and writing, the number of local memories corresponding to a storage area is determined according to the data dimension of the original data and the total number of divided data blocks, and the communication capability between the global memory and the data transmission interface, the object relationship between the storage area and the number of local memories are comprehensively considered. The plurality of data blocks of the global memory can be sent to the local memory in parallel, and the data processing results of the plurality of local memories can also be sent to the global memory, which not only helps to maximize the efficiency of data synchronization, but also ensures the effectiveness of data synchronization. The specific application environment architecture or the specific hardware architecture on which the data synchronization method depends will be described below.
[0032] One of the application scenarios of the embodiments of the present application is that the server is used as a host to process a user's high-performance computing task involving three-dimensional Fourier transform. In this application scenario, the FPGA is inserted into the PCIE (peripheral component interconnect express, high-speed serial computer expansion bus) card slot of the server, and a heterogeneous computing platform is formed by the CPU and the FPGA, which can be used to process the user's high-performance computing task involving three-dimensional Fourier transform. Figure 1As shown, the high-performance computing task is executed by the heterogeneous computing platform, and the three-dimensional Fourier transform task is executed by the FPGA. A plurality of Fourier transform cores are instantiated in the FPGA to perform Fourier transform in parallel, each Fourier transform core serving as a computing subunit to execute a corresponding Fourier transform task, so as to improve the computing speed of the three-dimensional Fourier transform. Each Fourier transform core is uniquely configured with a local memory, and the Fourier transform core only accesses three-dimensional Fourier data from the corresponding local memory, so as to improve the memory access speed during the computation. Thus, the heterogeneous computing efficiency of the FPGA for the three-dimensional Fourier transform is improved as a whole. In order to ensure the data processing accuracy of the FPGA, data synchronization needs to be performed on the global memory and the local memory of the FPGA.
[0033] In this application scenario, in the related art, data synchronization between the global memory and the local memory is achieved by a data loading module, as shown in FIG. 1. Figure 2 The host 101 transmits original three-dimensional matrix data to the global memory of the FPGA 102, the data loading module sequentially takes out the three-dimensional matrix data from the global memory, and then transmits each local memory according to the address of the three-dimensional matrix data in the global memory. After the three-dimensional Fourier transform is completed, the data is sequentially taken back from each local memory and stored in the global memory, so as to realize data synchronization between the global memory and the local memory. As can be seen, in the related art, there is only one data transmission port between the global memory and the data loading module, and the data of all local memories is transmitted from one port, which leads to low data transmission efficiency. In view of this, the global memory with a plurality of storage areas in the present application divides the original three-dimensional matrix data into a plurality of data blocks, and stores each data block in a corresponding storage area. Based on the data dimension of the original three-dimensional matrix data and the total number of data blocks, the number of local memories corresponding to one storage area is determined. According to the communication parameters of the global memory and the data transmission interface, the data transmission parameters between the global memory and each local memory are determined; according to the data transmission parameters, the data matched with the number corresponding relationship is read, and the corresponding read / write processing is performed according to the address signal, so as to realize parallel data synchronization between the global memory and the local memory, thereby effectively improving the efficiency of data synchronization. It should be noted that the above application scenario is only shown for the purpose of facilitating understanding of the ideas and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario. After introducing the technical solutions of the present application, various non-limiting embodiments of the present application will be described in detail in combination with the drawings and specific embodiments.
[0034] First, please refer to Figure 3 , Figure 3A flowchart of a data synchronization method provided by the embodiment is shown in the figure. The present application is an implementation scheme for synchronizing the global memory and the local memory of a computing unit with parallel computing capability or accelerated computing capability in a heterogeneous computing process. For ease of description, the computing unit is defined as a first computing unit, which includes a global memory and a plurality of local memories. The global memory includes a plurality of storage areas. The data synchronization method is described below with the first computing unit as the subject of execution. The embodiment can include the following contents.
[0035] S301: When the second computing unit sends the data block of the raw data to be processed to the corresponding storage area, the quantity correspondence between each storage area and the local memory is determined based on the data volume of the raw data to be processed and the total number of data blocks.
[0036] In the embodiment, the second computing unit and the first computing unit constitute a heterogeneous computing platform. When the second computing unit receives a user task or a system task, it decomposes the task and sends data that needs to be processed by the first computing unit to the first computing unit. Such data in the embodiment is defined as raw data to be processed, that is, the raw data to be processed is processed by the first computing unit. When the second computing unit sends the raw data to be processed to the global memory of the first computing unit, the global memory stores it in each storage area. In order to realize the partitioned storage of raw data, the raw data to be processed can be divided into a plurality of data blocks, and each data block can be stored in a storage area. When the global memory and the local memory are synchronized, the data of each storage area of the global memory needs to be transmitted to the local memory, or the data processing result obtained from the local memory needs to be stored in the global memory. After the raw data to be processed is partitioned and stored by the present application, it is further necessary to determine how many local memories each storage area can interact with at the same time. In other words, how many local memories are needed to store the data block stored in a storage area. The embodiment determines this by the data volume of the raw data to be processed and the total number of data blocks. The data volume is the total amount of data contained in the raw data to be processed or the data dimension of the raw data to be processed, which does not affect the implementation of the present application. The total number of data blocks is the total number of data blocks obtained by dividing the raw data to be processed. The quantity correspondence is that one storage area corresponds to several local memories, for example, the data interaction between one storage area and 8 local memories. At this time, the quantity correspondence is that one storage area corresponds to 8 local memories.
[0037] S302: Determine the data transmission parameters between the global memory and each local memory according to the communication parameters of the global memory and the data transmission interface.
[0038] It can be understood that the global memory and the local memory of the first computing unit transmit data through the data transmission interface when performing data synchronization, and the data transmission interface of the embodiment is the port between the global memory and each local memory. Due to the different communication capabilities between the global memory and the data transmission interface, the communication parameters are parameters for translating the data transmission capability, such as bit width, clock frequency, and the data transmission capability of the global memory is subject to the physical parameters of the hardware device used. In order to efficiently read data from the global memory and transmit it to the local memory, the communication parameters of the global memory and the data transmission interface need to be converted. For example, the bit width of each channel of the global memory is 256 bits, and each data is composed of a 32-bit real part and a 32-bit imaginary part, so 4 data can be transmitted in each clock cycle. The clock frequency of the global memory is 450 MHz, in order to fully utilize the high-speed characteristics of the global memory and efficiently read data from the global memory, the bit width of the data transmission interface can be set to 512 bits, and the clock frequency used is 250 MHz. In order to realize data communication between the two, clock and bit width conversion is needed. The data transmission parameters refer to how to read data and how to transmit data under the condition that the communication capabilities of the global memory and the data transmission interface are matched, including but not limited to data reading frequency, amount of data read each time, and data transmission frequency.
[0039] S303: According to the data transmission parameters, read the data matched with the quantity corresponding relationship, and perform corresponding read-write processing according to the address signal.
[0040] After the previous step determines how to read data and how to transmit data, the embodiment determines the data reading position according to whether the data is sent from the global memory to the local memory or the data processing result is obtained from the local memory, that is, the local memory or the global memory, and then reads the corresponding data from the data reading position according to the corresponding data reading mode, such as reading frequency and the amount of data read each time, until the condition of matching the quantity corresponding relationship is met. For example, the quantity corresponding relationship is that 1 storage area corresponds to 8 local memories, and in one clock cycle, the amount of data read from the global memory is the amount of data that 8 local memories can store. When obtaining the data processing result, 8 local memories are read simultaneously in one clock cycle. It can be understood that the interface protocols of the global memory and the local memory are different. For example, the interface protocol of the global memory is the Advanced Extensible Interface Protocol, and the interface protocol of the local memory is the local interface protocol of the random access memory. Therefore, during the data synchronization process, interface protocol conversion is also required. The interface protocol conversion refers to determining the position to which the data is written or the position from which the data is read according to the address related signals carried by the data during the data reading and writing operation. The address signals include read address related signals and write address related signals.
[0041] In the technical scheme provided by the embodiment, the original data to be processed by the first computing unit is stored in partitions, each partition can be used as a channel for external data reading and writing, and the number of local memories corresponding to each storage area can be determined according to the data dimension of the original data and the total number of the divided data blocks, so that the original data can be transmitted from the storage areas of the global memory to the local memories through multiple channels at the same time, thereby effectively improving the data synchronization efficiency between the global memory and the local memory. In addition, during the data synchronization process between the global memory and the local memory, the communication capability between the global memory and the data transmission interface is also considered, which not only helps to maximize the data synchronization efficiency, but also ensures the effectiveness of the data synchronization.
[0042] In the above embodiment, how to read data from the global memory and write it into the local memory, and how to read the data processing result from the local memory and write it into the global memory are not limited. Based on the above embodiment, the present application also provides an exemplary implementation, which can include the following contents:
[0043] For the process of reading data from the global memory and writing to the local memory: when each data block of the raw data to be processed is transmitted to the corresponding local memory, a plurality of target data is read from the corresponding address of the global memory according to the first read address signal; the target local memory position is determined according to the first write address signal, and a write valid signal is generated; based on the target local memory position, each target data is written into the corresponding target local memory.
[0044] For the process of storing the data processing result of the local memory to the global memory: when the data processing result of the raw data to be processed is transmitted to the global memory, the source local memory position is determined according to the second read address signal; based on the source local memory position, the corresponding target data processing result is read from each source local memory, and each target data processing result is combined into a data processing result; the storage address of the global memory is determined according to the second write address signal, and based on the storage address of the global memory, the data processing result is written to the corresponding position of the global memory.
[0045] In the present embodiment, for the convenience of description, the currently written local memory is defined as the target memory, and the data written into each target memory is defined as the target data. The target local memory position includes the address information of each target local memory, and the total number of target local memories is the same as the total number of target data. Since the target data is read from the storage area and written into the local memory, and the data correspondence relationship describes the quantity relationship between the storage area and the local memory, the total number of target data matches the quantity correspondence relationship. The local memory that reads the data processing result is defined as the source memory, and the data processing result read from the local memory and written into the global memory is defined as the target data processing result. The source local memory position includes the address information of each source local memory, and the total number of source local memories matches the quantity correspondence relationship. In the process of reading data from the global memory and writing to the local memory, the global memory is told from which address to read how much data through the read address related signal, and then the global memory feeds back the corresponding data to the data transmission interface. When the data transmission interface writes the read data into the local memory, a write valid signal can be generated by pulling up the signal, the local memory is told the writing position of the data through the write address signal, and the target data is written through the write data signal. When the target data processing result is read from the local memory and written into the global memory, the source local memories are told the position of the data to be read through the read address signal, and then the corresponding data is taken out from these source local memories at the same time through the read data signal and combined, the global memory is told the address of the data to be written through the write address related signal, and then the data is written into the global memory through the write data signal.
[0046] From the above, the embodiment can accurately read data from the global memory and accurately write to the local memory through interface protocol conversion, can accurately obtain data processing results from the local memory, and combine the data processing results into data of the interface bit width of the global memory, and efficiently write into the global memory, which is beneficial to improve the data synchronization efficiency between the global memory and the local memory.
[0047] The above embodiment does not make any limitation on how to determine the quantity correspondence relationship between the storage areas and the local memories, and based on the above embodiment, the application also gives an exemplary determination method of the quantity correspondence relationship between the storage areas and the local memories, which can include the following contents:
[0048] According to the data dimension of the to-be-processed original data, the local storage parameters of each local memory are determined; based on the local storage parameters, the number of data stored in each local memory is determined; and according to the number of data stored in each storage area and the number of data stored in each local memory, the quantity correspondence relationship between each storage area and the local memory is determined.
[0049] In the embodiment, the to-be-processed original data is three-dimensional matrix data, and correspondingly, the to-be-processed original data can be represented by data dimension. The local storage parameter refers to the way of storing data in the local memory and the total number of local memories. For example, the three-dimensional matrix data includes three planes, and the local storage parameter is that the data of each XZ plane is stored in one local memory, or the data of several XZ planes is stored in one local memory. When the way of storing data is determined, the number of local memories used can be determined according to the data amount. When the local storage parameter is determined, the number of data stored in each local memory can be determined. Taking the to-be-processed original data of 128*128*128 as an example, the data of each XZ plane is stored in one local memory, and 128 XZ planes use 128 local memories, that is, 128 local memories are used to store data, so the number of data stored in each local memory is 128*128. If one local memory stores data of other number of planes, the number of local memories used can be an integer multiple of 128, such as 16, 32, and 64. Taking the total number of data blocks as 16 as an example, the data is stored in 16 storage areas of the global memory, that is, the number of data stored in each storage area is 8*128*128, and each storage area interacts with 8 local memories.
[0050] From the above, the embodiment determines the quantity correspondence relationship between the storage areas and the local memories according to the data amount of the to-be-processed original data and the storage capacity of the local memories, so that the global memory and the local memory can realize parallel data synchronization, which is beneficial to improve the data synchronization efficiency.
[0051] The above embodiments do not limit how to determine the data transmission parameters, and based on the above embodiments, the application also provides an exemplary determination method of the data transmission parameters between the global memory and the local memories, which can include the following contents:
[0052] Obtaining the bit width and clock frequency of the global memory and the data transmission interface; determining the reading parameters and sending parameters for reading the global memory according to the bit width and clock frequency of the global memory and the data transmission interface, and taking the reading parameters and the sending parameters as the data transmission parameters.
[0053] The bit width is the width of processing or transmitting data in the computer system, that is, the number of data bits that can be processed or transmitted at a time, which can reflect the data transmission capability, and the clock frequency determines the number of basic operations or instructions that can be executed per second, and the embodiment uses the bit width and the clock frequency to measure the data transmission capability. The reading parameters at least include the data reading times, the data amount and the reading frequency, and the sending parameters include the sending frequency. Exemplarily, the data reading parameters for reading the global memory can be determined according to the global bit width of the global memory and the interface bit width of the data transmission interface; the data reading parameters at least include the data reading times and the data amount; the data reading frequency for reading the global memory and the data sending frequency of the global memory are determined according to the global clock frequency of the global memory and the interface clock frequency of the data transmission interface; and finally, the data reading parameters and the data reading frequency are taken as the reading parameters, and the data sending frequency is taken as the sending parameters. Taking the bit width of the global memory as 256 bits, the bit width of the data transmission interface as 512 bits, the clock frequency of the global memory as 450 MHz and the clock frequency of the data transmission interface as 250 MHz as an example, the conversion of the clock and the bit width is realized by using the related IP provided by the FPGA supplier. In short, the bit width of the high-bandwidth memory is 256 bits, the data loading module is 512 bits, and the data transmission interface reads twice 256-bit data from the global memory to form a 512-bit data for the data transmission interface. The clock frequency of the global memory is close to twice the data transmission interface, and the data is read twice from the global memory at a frequency of 450 MHz, and then transmitted once to the data transmission interface at a frequency of 250 MHz. If the data is not enough, the corresponding operation can be performed when the data amount reaches the requirement.
[0054] As known from the above, the embodiment represents the data transmission capability by the bit width and the clock frequency, and through the conversion of the bit width and the clock frequency between the global memory and the data transmission interface, the data can be efficiently and accurately transmitted, which is beneficial to improving the data synchronization efficiency.
[0055] To ensure that the local memory can store continuous data, facilitate data processing by the first computing unit, and improve the data processing efficiency of the first computing unit, based on the above embodiment, the present invention further provides a storage method for the raw data to be processed, which may include the following:
[0056] The original data to be processed is three-dimensional matrix data, and the data values of each data block on the target coordinate axis match the corresponding quantity relationship; when storing each data block, the storage area of the global memory regards the data of the data block to be stored in the direction of the target coordinate axis as a group, and stores the data of the data block to be stored in sequence, so that each group of data is stored in the corresponding local memory group in each clock cycle.
[0057] In this embodiment, the local memory group includes multiple local memories, and the number of local memories is the same as the number of data in each group. The raw data to be processed is three-dimensional matrix data, that is, it has data in the directions of the three coordinate axes, where the target coordinate axis is, for example, the direction of the raw data to be processed. Also, taking the data dimension of the raw data to be processed as 128×128×128 as an example, Figure 4 As shown in the figure, each small cube represents 8×8×8 data. The global memory includes 16 storage areas. The target coordinate axis is the Y axis. In the Y axis direction, the original data to be processed is divided into 16 data blocks. Each storage area of the global memory stores 128×8×128 data (Z direction×Y direction×X direction), which corresponds to Figure 4 A data block in the Z direction, the data of each plane, that is, the XY plane, such as Figure 5 shown.
[0058] Among them, since the data bit width of the data transmission interface is 512 bits and the bit width of each data is 64 bits, 8 data can be transmitted per clock cycle. Therefore, when the original data is to be processed, the 8 vertical data are regarded as a group, such as Figure 5 As shown, 1-1 to 8-1 can be stored first, followed by 1-2 to 8-2, and the subsequent data can be stored in the same way. When the data transmission interface takes out data from each storage area, it takes out 512 bits of data per clock cycle, that is, 8 data, and then stores these 8 data in 8 local memories respectively. In the next cycle, 8 data are taken out and stored in 8 local memories respectively, and the subsequent data can be transferred in the same way. When 128 clock cycles are completed, the data of one plane can be transmitted. Each storage area stores 128 planes of data, so 128 rounds of reading are required to complete the synchronization of all data in each storage area. This data storage and reading method can ensure that the data stored in the local memory is continuous in the X direction, as shown in the figure. Figure 51-1, 1-2, 1-3, …, 1-128. Taking the execution of the three-dimensional Fourier transform task by the first computing unit as an example, each sub-computing unit can directly read data from the local memory continuously when performing the X-dimensional Fourier transform, that is, when performing the first dimension of the three-dimensional Fourier transform. After the three-dimensional Fourier transform is completed, the data processing result is written back to the global memory from the local memory. The data transmission interface reads one data from each of the eight local memories every clock cycle to form a 512-bit data written back to the global memory. When the next clock cycle arrives, the data is read again to form a 512-bit data written back to the global memory, and this process continues until all data is written back to the global memory.
[0059] As can be seen from the above, the embodiment stores the to-be-processed raw data as a group of data in the target time axis direction into a local memory, which facilitates the first computing unit to directly read data from the local memory for processing when performing a task, and is beneficial to improving the data processing efficiency.
[0060] It can be understood that the present application is for heterogeneous computing, and the above embodiment is a method flow performed by one computing unit of the heterogeneous computing platform. The following embodiment describes a method flow performed by another computing unit in the data synchronization process. For ease of description, the computing unit is defined as a second computing unit, such as Figure 6 As shown, the embodiment can include the following contents:
[0061] S601: Obtain to-be-processed raw data, and split the to-be-processed raw data into a plurality of data blocks.
[0062] S602: Send each data block to a corresponding storage area of the global memory of the first computing unit, so that the first computing unit determines the quantity correspondence relationship between each storage area and the local memory based on the data amount of the to-be-processed raw data and the total number of data blocks; determines the data transmission parameter between the global memory and each local memory according to the communication parameter of the global memory and the data transmission interface; reads the data matched with the quantity correspondence relationship according to the data transmission parameter, and performs corresponding read-write processing according to the address signal.
[0063] The same steps of the embodiment and the above embodiment are described in the above embodiment, which will not be described here.
[0064] As can be seen from the above, the embodiment splits the to-be-processed raw data into a plurality of data blocks by the second computing unit, and distributes the data blocks to the second computing unit, so that the second computing unit can realize raw data partition storage, and further realize parallel data synchronization between the global memory and the local memory, thereby effectively improving the efficiency of data synchronization.
[0065] Based on the above embodiments, the application also provides an exemplary data splitting mode, which can include the following contents:
[0066] Based on the total number of storage areas of the global memory of the first computing unit, the data block splitting is performed on the raw data to be processed to obtain a plurality of data blocks; and each data block is stored to the corresponding storage area through the read-write channel of the global memory.
[0067] In the embodiment, the number of data blocks of the raw data to be processed is equal to the number of storage areas of the global memory, and each storage area stores one data block. For example, if 16 storage areas are used, the data is divided into 16 data blocks along the Y direction. In the embodiment, the three-dimensional matrix data of 128x128x128 is taken as the raw data to be processed, and the global memory includes 16 storage areas. As shown in the figure, Figure 4 Each small cube represents 8x8x8 data, and each dimension has 16 small squares. The second computing unit divides the three-dimensional matrix data into 16 data blocks along the Y direction, and then stores the raw data to be processed to the 16 storage areas through one channel of the global memory by using the direct memory access engine, thereby completing the raw data issued by the host.
[0068] It should be noted that there is no strict execution order between the steps in the application, as long as the logical order is met, the steps can be executed simultaneously, or executed in a certain preset order, Figure 3 and Figure 6 It is only an illustrative way, and does not mean that only such an execution order can be used.
[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. The application also provides a corresponding device for the data synchronization method, which further makes the method more practical. Among them, the device can be explained from the perspective of functional modules and hardware.
[0070] The data synchronization device provided by the application is introduced below from the perspective of functional modules. Please refer to Figure 7 The device is applied to a first computing unit including a global memory and a plurality of local memories. The global memory of the first computing unit includes a plurality of storage areas, which are used to realize the data synchronization method provided by the application. The device can include:
[0071] A quantity relationship determination module 701 is configured to determine the quantity corresponding relationship between each storage area and the local memory based on the data volume of the raw data to be processed and the total number of data blocks when the second computing unit sends the data block of the raw data to be processed to the corresponding storage area.
[0072] The parameter conversion module 702 is configured to determine data transmission parameters between the global memory and each local memory according to the communication parameters of the global memory and the data transmission interface.
[0073] The data read-write module 703 is configured to read data corresponding to the quantity correspondence according to the data transmission parameters, and perform corresponding read-write processing according to the address signal.
[0074] In this embodiment, the data synchronization apparatus can include or be divided into a plurality of program modules, namely, the quantity relationship determination module 701, the parameter conversion module 702, and the data read-write module 703. These program modules are stored in a storage medium and executed by one or more processors, so as to complete the data synchronization method disclosed by the above method embodiments. The program module referred to in this embodiment refers to a series of computer program instruction segments capable of completing a specific function, and is more suitable than the program itself for describing the execution process of the data synchronization apparatus in the storage medium. The description of the features in the embodiments corresponding to the data synchronization apparatus can be referred to the related description of the embodiments corresponding to the data synchronization method, which will not be repeated here.
[0075] For example, in some embodiments of this embodiment, the data read-write module 703 can also be configured to: when transmitting each data block of the to-be-processed original data to the corresponding local memory, read a plurality of target data from the corresponding address of the global memory according to the first read address signal; determine the target local memory position according to the first write address signal, and generate a write enable signal; and write each target data into the corresponding target local memory based on the target local memory position, wherein the target local memory position includes address information of each target local memory, the total number of target local memories is the same as the total number of target data, and the total number of target data matches the quantity correspondence.
[0076] For example, in some other embodiments of this embodiment, the data read-write module 703 can also be configured to: when transmitting the data processing result of the to-be-processed original data to the global memory, determine the source local memory position according to the second read address signal; read the corresponding target data processing result from each source local memory based on the source local memory position, and combine each target data processing result into a data processing result; determine the storage address of the global memory according to the second write address signal, and write the data processing result into the corresponding position of the global memory based on the storage address of the global memory; wherein the source local memory position includes address information of each source local memory, and the total number of source local memories matches the quantity correspondence.
[0077] Illustratively, in some other implementations of this embodiment, the above-mentioned quantity relationship determination module 701 can also be used to: determine the local storage parameters of each local memory according to the data dimension of the original data to be processed; determine the number of data stored in each local memory based on the local storage parameters; determine the quantity correspondence between each storage area and the local memory according to the number of data stored in each storage area and the number of data stored in each local memory.
[0078] Exemplarily, in some other implementations of this embodiment, the above-mentioned parameter conversion module 702 can also be used to: obtain the bit width and clock frequency of the global memory and the data transmission interface; determine the read parameters and send parameters for reading the global memory based on the bit width and clock frequency of the global memory and the data transmission interface, and use the read parameters and send parameters as data transmission parameters; wherein the read parameters include at least the number of data reads, the data amount and the reading frequency, and the send parameters include the sending frequency.
[0079] As an exemplary implementation of the above embodiment, the parameter conversion module 702 may be further configured to: determine a data read parameter for reading the global memory based on a global bit width of the global memory and an interface bit width of the data transmission interface; the data read parameter includes at least a number of data reads and a data volume; determine a data read frequency for reading the global memory and a data sending frequency for sending data from the global memory based on a global clock frequency of the global memory and an interface clock frequency of the data transmission interface; and use the data read parameter and the data read frequency as read parameters, and the data sending frequency as a sending parameter.
[0080] Exemplarily, in some other implementations of this embodiment, the above-mentioned device may also include an original data storage module, which can be used for: the original data to be processed is three-dimensional matrix data, and the data values of each data block in the target coordinate axis match the corresponding quantity; when storing each data block, the storage area of the global memory regards each data of the data block to be stored in the direction of the target coordinate axis as a group, and stores each data of the data block to be stored in sequence, so as to store each group of data in the corresponding local memory group in each clock cycle; wherein, the local memory group includes multiple local memories, and the number of local memories is the same as the number of each group of data.
[0081] For example, the following describes the data synchronization device applied to the second computing unit based on the perspective of functional modules. Figure 8 , the device may include:
[0082] The data cutting module 801 is used to cut the acquired raw data to be processed into multiple data blocks.
[0083] The data issuing module 802 is configured to send each data block to a corresponding storage area of the global memory of the first computing unit, so that the first computing unit determines the quantity correspondence between each storage area and the local memory based on the data amount of the raw data to be processed and the total number of data blocks; determines the data transmission parameter between the global memory and each local memory according to the communication parameter of the global memory and the data transmission interface; reads the data matched with the quantity correspondence according to the data transmission parameter, and performs corresponding read-write processing according to the address signal.
[0084] For example, in some embodiments of the present embodiment, the data splitting module 801 described above can also be configured to: split the raw data to be processed into a plurality of data blocks based on the total number of storage areas of the global memory of the first computing unit; and store each data block to the corresponding storage area through the read-write channel of the global memory.
[0085] The data synchronization device mentioned above is described from the perspective of functional modules. Further, the present application also provides an electronic device described from the perspective of hardware, which comprises a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above data synchronization method embodiments.
[0086] The embodiments of the present application also provide a computer readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above data synchronization method embodiments when running.
[0087] In an exemplary embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0088] The embodiments of the present application also provide a computer program product, which comprises a computer program. The computer program is executed by a processor to implement the steps in any of the above data synchronization method embodiments.
[0089] The embodiments of the present application also provide another computer program product, which comprises a non-volatile computer readable storage medium storing a computer program. The computer program is executed by a processor to implement the steps in any of the above data synchronization method embodiments.
[0090] Finally, the present application also provides a data synchronization system, which is described with reference to Figure 9The data synchronization system comprises at least a first computing unit 901 and a second computing unit 902. The first computing unit comprises an internal memory, a plurality of computing sub-units and a data synchronization processor. The internal memory comprises a global internal memory and a plurality of local internal memories. Each computing sub-unit corresponds to a local internal memory, and each computing sub-unit only interacts with the corresponding local internal memory. The first computing unit executes a computer program to implement the steps of the data synchronization method described in any one of the method embodiments executed by the first computing unit. The second computing unit executes a computer program to implement the steps of the data synchronization method described in any one of the method embodiments executed by the second computing unit.
[0091] As can be seen from the above, the data synchronization system of the embodiment can realize parallel data synchronization of the global internal memory and the local internal memory, thereby effectively improving the efficiency of data synchronization.
[0092] The above embodiment does not make any limitation on the data structure of the global internal memory. In order to improve the data synchronization efficiency, the embodiment further provides a high-performance data storage structure, which can comprise the following contents:
[0093] The global internal memory comprises at least a first interconnection unit and a second interconnection unit. The first interconnection unit and the second interconnection unit are connected to each other. The first interconnection unit comprises at least a first interconnection module and a second interconnection module. The first interconnection module and the second interconnection module are connected through the first interconnection unit. The first interconnection module comprises at least a first storage area and a second storage area. The first storage area and the second storage area are connected through the first interconnection module. The second interconnection module comprises at least a third storage area and a fourth storage area. The third storage area and the fourth storage area are connected through the second interconnection module. The second interconnection unit comprises at least a third interconnection module and a fourth interconnection module. The fourth interconnection module and the third interconnection module are connected through the second interconnection unit. The third interconnection module comprises at least a fifth storage area and a sixth storage area. The fifth storage area and the sixth storage area are connected through the third interconnection module. The fourth interconnection module comprises at least a seventh storage area and an eighth storage area. The seventh storage area and the eighth storage area are connected through the fourth interconnection module.
[0094] In the embodiment, the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area are independent of each other and correspond to one read-write channel respectively. At least one read-write channel is used to perform data read-write operation on any number of storage areas in the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area. That is, the global internal memory comprises a plurality of independent storage areas, and data access can be independently performed through the corresponding channels. Figure 10As shown, the global memory includes 32 independent storage areas, each two storage areas are fully connected through an interconnection module, each two interconnection modules are fully connected through an interconnection unit, and the interconnection units are also connected. The global memory provides 32 channels externally, and each channel can perform data read-write operation on all storage areas. In order to improve the data read-write efficiency, the global memory simultaneously performs read-write operation on at least the first storage area and the second storage area through at least the first read-write channel and the second read-write channel at the same time, in other words, the global memory simultaneously enables 32 channels to perform parallel data read-write operation on 32 storage areas.
[0095] As can be seen from the above, the embodiment provides a global memory with large capacity storage and high memory bandwidth, realizes shorter data path and higher data transmission rate, reduces access delay, and is beneficial to improve data read-write efficiency.
[0096] In actual application process, the embodiment can also instantiate the related method recorded in the above embodiment as a plurality of function modules in order to quickly apply. Based on the above embodiment, the data synchronization method can be instantiated as a plurality of data adaptation modules and a plurality of data loading modules by using any instantiation manner. The data adaptation module group is configured to determine the data transmission parameters between the global memory and each local memory according to the communication parameters of the global memory and the data transmission interface. The data loading module group is configured to read the data matched with the quantity corresponding relationship according to the data transmission parameters, and perform corresponding read-write processing according to the address signal. In order to facilitate description, the plurality of data adaptation modules can be defined as a data adaptation module group, and the plurality of data loading modules can be defined as a data loading module group. Correspondingly, the data synchronization processor includes the data adaptation module group and the corresponding data loading module group. Figure 11 As shown, the number of data adaptation modules included in the data adaptation module group is the same as the number of data loading modules included in the data loading module group, and the number of data adaptation modules is the same as the total number of storage areas. Each data loading module includes a first data transmission interface and a second data transmission interface. Each data adaptation module is connected to the corresponding data loading module through the first data transmission interface. Each data loading module is connected to the corresponding local memory group through the second data transmission interface. The local memory group includes a plurality of local memories, and the number of local memories is determined based on the quantity corresponding relationship between the storage areas and the local memories.
[0097] In the embodiment, in order to further improve the data synchronization efficiency, the first data transmission interface of the data loading module adopts the advanced extensible interface, the second data transmission interface adopts the random access memory local interface, and accordingly, the interface between the data loading module and the data adaptation module adopts the advanced extensible interface for data interaction, and the interface between the data loading module and the local memory adopts the random access memory local interface for data interaction. The data bit width of the advanced extensible interface is 512 bits, and the bit width of each data is 64 bits, so that the data loading module reads one data containing 8 actual data each time, and accordingly, writes the data into 8 local memories. As shown in Figure 12 , in the process of writing the to-be-processed raw data from the global memory into the local memory, the data loading module can inform the global memory from which address to read how much data through the read address related signal, and the global memory feeds back the corresponding data to the data loading module. As shown in Figure 13 , when the data loading module writes the to-be-processed raw data, the write valid signal can be generated by pulling up the signal, the local memory is informed of the writing position of the to-be-processed raw data through the write address related signal, and the data is written through the write data signal. As shown in Figure 13 , when the data processing result is read from the local memory and written into the global memory, the read address signal is used to inform the 8 local memories of the position of the read data, 8 64-bit data are obtained from the 8 local memories at the same time through the read data signal, and the data are combined into one 512-bit data. The global memory is informed of the address of the writing of the data processing result through the write address related signal, and finally the data are written into the global memory through the write data signal.
[0098] As can be seen from the above, the data loading module is formed by the advanced extensible interface and the random access memory local interface, the protocol conversion between the global memory and the local memory is realized through the data loading module, the data can be accurately read from the global memory and accurately written into the local memory, the data processing result can be accurately obtained from the local memory, the data processing result is combined into the data of the interface bit width of the global memory, and the data processing result is efficiently written into the global memory, which is beneficial to improving the data synchronization efficiency between the global memory and the local memory.
[0099] In order to make the technical scheme of the present application more clear to those skilled in the art, the present application further provides a schematic data synchronization system as shown in Figure 14As shown, the first computing unit is an FPGA, the FPGA uses a high-bandwidth memory as a global memory, the high-bandwidth memory includes 16 memory areas, and all the memory areas can be accessed through one channel due to the interconnection between the 16 memory areas. The host sends data to the global memory of the first computing unit through a direct memory access engine, and since the direct memory access engine has only one channel, the data can be sent to the 16 memory areas through only one channel, and the data in different memory areas are distinguished according to the address. In order to improve the efficiency of the first computing unit reading and writing data from the global memory, the first computing unit reads and writes data from the global memory to the local memory using 16 channels to read and write 16 memory areas at the same time. The second computing unit is a CPU on the host, and the parallel data synchronization process between the global memory and the local memory in the three-dimensional Fourier transform based on the data synchronization system can include the following contents:
[0100] The host runs a target upper-layer application, the target upper-layer application is an application that can use the direct memory access engine, and the target upper-layer application uses the direct memory access engine driver to send the to-be-processed raw data in the host memory to the 16 memory areas through the direct memory access engine. After the to-be-processed raw data is sent, the host sends an instruction to the first computing unit to inform the first computing unit that the data sending is completed. After receiving the instruction, the first computing unit uses 16 data loading modules to read data from the 16 memory areas in parallel and stores the data in 128 local memories, respectively. After the data synchronization between the global memory and the local memory is completed, each computing subunit of the first computing unit takes out data from the local memory for three-dimensional Fourier transform, and returns the data to the local memory after the computation is completed. Each data loading module reads the calculation results from the 128 local memories and stores the calculation results in the 16 memory areas. After the data synchronization between the global memory and the local memory is completed, the first computing unit can send a task completion message to the host through interruption. The target upper-layer application uses the access engine driver again to read the data in the high-bandwidth memory back to the host memory through the direct memory access engine, and completes the parallel data synchronization between the global memory and the local memory in the three-dimensional Fourier transform.
[0101] The data synchronization method, apparatus, electronic device, computer readable storage medium and computer program product are described in detail above. Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. The units and algorithm steps of each example described by each disclosed embodiment are executed in the form of electronic hardware or computer software, which depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, and such implementation should not be considered beyond the scope of the present application. Without departing from the principles of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the scope of the present application.
Claims
1. A method of data synchronization, the method comprising: The application is applied to a first computing unit including a global memory and a plurality of local memories, the global memory of the first computing unit includes a plurality of storage areas, including: When the second computing unit sends data blocks of to-be-processed raw data to the corresponding storage areas, the quantity correspondence relationship between each storage area and the local memory is determined based on the data quantity of the to-be-processed raw data and the total number of data blocks; According to the communication parameters of the global memory and the data transmission interface, the data transmission parameters between the global memory and each local memory are determined; wherein the data transmission interface is a port between the global memory and each local memory; According to the data transmission parameters, the data matched with the quantity correspondence relationship is read, and the corresponding read-write processing is performed according to the address signal.
2. The data synchronization method of claim 1, wherein, The reading of the data matched with the quantity correspondence relationship and the corresponding read-write processing according to the address signal includes: When the data blocks of the to-be-processed raw data are transmitted to the corresponding local memories, a plurality of target data is read from the corresponding addresses of the global memory according to the first read address signal; The target local memory position is determined according to the first write address signal, and a write valid signal is generated; Based on the target local memory position, each target data is written into the corresponding target local memory; Wherein, the target local memory position includes the address information of each target local memory, the total number of target local memories is the same as the total number of target data, and the total number of target data matches the quantity correspondence relationship.
3. The data synchronization method of claim 1, wherein, The reading of the data matched with the quantity correspondence relationship and the corresponding read-write processing according to the address signal includes: When the data processing result of the to-be-processed raw data is transmitted to the global memory, the source local memory position is determined according to the second read address signal; Based on the source local memory position, the corresponding target data processing result is read from each source local memory, and each target data processing result is combined into a data processing result; According to the second write address signal, the storage address of the global memory is determined, and based on the storage address of the global memory, the data processing result is written into the corresponding position of the global memory; Wherein, the source local memory position includes the address information of each source local memory, and the total number of source local memories matches the quantity correspondence relationship.
4. The data synchronization method of claim 1, wherein, The determination of the quantity correspondence relationship between each storage area and the local memory based on the data quantity of the to-be-processed raw data and the total number of data blocks includes: According to the data dimension of the to-be-processed raw data, the local storage parameters of each local memory are determined; Based on the local storage parameters, the number of data stored in each local memory is determined; According to the number of data stored in each storage area and the number of data stored in each local memory, the quantity correspondence relationship between each storage area and the local memory is determined.
5. The data synchronization method according to any one of claims 1 to 4, characterized in that, The determination of the data transmission parameters between the global memory and each local memory according to the communication parameters of the global memory and the data transmission interface includes: The bit width and clock frequency of the global memory and the data transmission interface are obtained; determining read parameters and sending parameters for reading the global memory according to a bit width and a clock frequency of the global memory and the data transmission interface, taking the read parameters and the sending parameters as the data transmission parameters; wherein the read parameters at least include a data read times, a data amount and a read frequency, and the sending parameters include a sending frequency.
6. The data synchronization method of claim 5, wherein, The determining read parameters and sending parameters for reading the global memory according to a bit width and a clock frequency of the global memory and the data transmission interface includes: determining data read parameters for reading the global memory according to a global bit width of the global memory and an interface bit width of the data transmission interface; the data read parameters at least include a data read times and a data amount; determining a data read frequency for reading the global memory according to a global clock frequency of the global memory and an interface clock frequency of the data transmission interface, and a data sending frequency for sending data by the global memory; taking the data read parameters and the data read frequency as read parameters, and the data sending frequency as the sending parameters.
7. The data synchronization method according to any one of claims 1 to 4, characterized in that, The to-be-processed original data is three-dimensional matrix data, and a data quantity value of each data block on a target coordinate axis matches the quantity correspondence relationship; After the second computing unit sends the data block of the to-be-processed original data to the corresponding storage area, the method further includes: When storing each data block, the storage area of the global memory stores each data of the to-be-stored data block as a group in the direction of the target coordinate axis, and sequentially stores each data of the to-be-stored data block, so as to store each group of data into the corresponding local memory group in each clock cycle. The local memory group includes a plurality of local memories, and the number of local memories is the same as the number of data in each group.
8. A data synchronization method, characterized by, The application is applied to a second computing unit, and includes: obtaining to-be-processed original data, and dividing the to-be-processed original data into a plurality of data blocks; sending each data block to a corresponding storage area of a global memory of a first computing unit, so that the first computing unit determines a quantity correspondence relationship between each storage area and a local memory based on a data quantity of the to-be-processed original data and a total number of data blocks, determines data transmission parameters between the global memory and each local memory based on communication parameters of the global memory and a data transmission interface, reads data matching the quantity correspondence relationship according to the data transmission parameters, and performs corresponding read-write processing according to an address signal.
9. The data synchronization method of claim 8, wherein, The dividing the to-be-processed original data into a plurality of data blocks includes: dividing the to-be-processed original data into data blocks based on a total number of storage areas of the global memory of the first computing unit, to obtain a plurality of data blocks; storing each data block to a corresponding storage area through a read-write channel of the global memory.
10. An electronic device, comprising: It includes: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the data synchronization method according to any one of claims 1 to 9.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the data synchronization method according to any one of claims 1 to 9.
12. A data synchronization system comprising at least a first computing unit and a second computing unit, the first computing unit comprising an internal memory, a plurality of computing sub-units and a data synchronization processor; wherein the internal memory comprises a global internal memory and a plurality of local internal memories, each computing sub-unit corresponds to a local internal memory, and each computing sub-unit only interacts with the corresponding local internal memory; the first computing unit executes a computer program to implement the steps of the data synchronization method according to any one of claims 1 to 7; and the second computing unit executes a computer program to implement the steps of the data synchronization method according to claim 8 or 9.
13. The data synchronization system of claim 12, wherein, The data synchronization processor comprises a data adaptation module group and a corresponding data loading module group; the data adaptation module group contains the same number of data adaptation modules as the data loading module group, and the number of data adaptation modules is the same as the total number of storage areas; Each data loading module comprises a first data transmission interface and a second data transmission interface, each data adaptation module is connected to the corresponding data loading module through the first data transmission interface, and each data loading module is connected to the corresponding local internal memory group through the second data transmission interface; wherein the local internal memory group comprises a plurality of local internal memories, and the number of local internal memories is determined based on the number correspondence between the storage areas and the local internal memories; The data adaptation module group is configured to determine the data transmission parameters between the global internal memory and each local internal memory according to the communication parameters of the global internal memory and the data transmission interface; The data loading module group is configured to read the data matching the number correspondence according to the data transmission parameters, and perform corresponding read / write processing according to the address signal.
14. The data synchronization system of claim 12, wherein, The global internal memory comprises at least a first interconnection unit and a second interconnection unit, and the first interconnection unit and the second interconnection unit are connected to each other; The first interconnection unit comprises at least a first interconnection module and a second interconnection module, and the first interconnection module and the second interconnection module are connected through the first interconnection unit; the first interconnection module comprises at least a first storage area and a second storage area, and the first storage area and the second storage area are connected through the first interconnection module; the second interconnection module comprises at least a third storage area and a fourth storage area, and the third storage area and the fourth storage area are connected through the second interconnection module; The second interconnection unit comprises at least a third interconnection module and a fourth interconnection module, and the fourth interconnection module and the third interconnection module are connected through the second interconnection unit; the third interconnection module comprises at least a fifth storage area and a sixth storage area, and the fifth storage area and the sixth storage area are connected through the third interconnection module; the fourth interconnection module comprises at least a seventh storage area and an eighth storage area, and the seventh storage area and the eighth storage area are connected through the fourth interconnection module; The first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area are independent of each other and correspond to one read-write channel respectively; at least one read-write channel is used to perform data read-write operation on any number of storage areas among the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area.
15. The data synchronization system of claim 14, wherein, The global memory is used to simultaneously perform read-write operation on the first storage area and the second storage area through at least the first read-write channel and the second read-write channel at the same time.
Citation Information
Patent Citations
Method for rapid communication within a parallel computer system, and a parallel computer system operated by the method
US20020078322A1
Information processing apparatus, communication method and information processing system
US20160132272A1