Data synchronization method and system, electronic equipment and computer readable storage medium

By storing the data partition and determining the quantitative correspondence between the memory area and the local memory, parallel data synchronization between the global memory and the local memory is realized, solving the problem of inefficient data synchronization in the prior art and improving data processing efficiency.

CN120029553AActive Publication Date: 2025-05-23LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202510238300.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-23
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

When performing tasks using an accelerated computing unit, the data transmission interface is inefficient in obtaining data from the global memory and transmitting it to the local memory, which cannot meet the user's efficient data processing needs.

Method used

By storing the original data partitions to be processed, each partition serves as a channel, and the number correspondence between the memory area and the local memory is determined based on the data dimension and the total number of data blocks, and the data transmission parameters are determined based on the communication parameters of the global memory and the data transmission interface, so as to realize parallel data synchronization between the global memory and the local memory.

Benefits of technology

It effectively improves the data synchronization efficiency between global memory and local memory, meets users' efficient data processing needs, and ensures the effectiveness of data synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029553A_ABST
    Figure CN120029553A_ABST
Patent Text Reader

Abstract

The invention discloses a data synchronization method and system, electronic equipment and a computer readable storage medium, and relates to the technical field of data processing. The method comprises the steps that a second calculation unit sends data blocks obtained by dividing original data to be processed to all storage areas of a global memory of a first calculation unit, and the number of local memories corresponding to the storage areas is determined based on the data size of the original data to be processed and the total number of the data blocks. Determining data transmission parameters between the global memory and the local memories according to the communication parameters of the global memory and the data transmission interface; and reading data matched with the quantity corresponding relation according to the data transmission parameters, and performing corresponding read-write processing according to the address signal. The problem of low synchronization efficiency in related technologies can be solved, parallel data synchronization of the global memory and the local memory can be realized, and the data synchronization efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data synchronization method, system, electronic device and computer-readable storage medium. Background Art

[0002] As the scale of data that users need to process grows, their requirements for data processing speed are getting higher and higher. Currently, different types of computing units are used in the same computing system to complete tasks together, and data that occupies a large amount of computing resources is sent to accelerated computing units with parallel capabilities for processing.

[0003] When using the accelerated computing unit to perform tasks, the data transmission interface first obtains the original data sent by the host from the global memory, and transmits it to the local memory for processing according to the address information carried by the original data, and finally sends the processed data to the global memory. This data synchronization method between the global memory and the local memory is inefficient and cannot meet the user's demand for efficient data processing. Summary of the invention

[0004] The present invention provides a data synchronization method, system, electronic device and computer-readable storage medium, which can realize parallel data synchronization between a global memory and a local memory, and effectively improve data synchronization efficiency.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] On one hand, the present invention provides a data synchronization method, which is applied to a first computing unit including a global memory and multiple local memories, wherein the global memory of the first computing unit includes multiple storage areas, including: when a second computing unit sends a data block of raw data to be processed to a corresponding storage area, based on the data volume of the raw data to be processed and the total number of data blocks, determining the quantity correspondence between each storage area and the local memory; determining the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; reading data matching the quantity correspondence according to the data transmission parameters, and performing corresponding read and write processing according to the address signal; wherein the data transmission interface is a port between the global memory and each local memory.

[0007] On the other hand, the present invention provides a data synchronization method, which is applied to a second computing unit, including: obtaining original data to be processed and dividing the original data to be processed into multiple data blocks; sending each data block to a corresponding storage area of ​​a global memory of a first computing unit, so that the first computing unit determines the quantity correspondence between each storage area and a local memory based on the data volume of the original data to be processed and the total number of data blocks; determining the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; reading data matching the quantity correspondence according to the data transmission parameters, and performing corresponding read and write processing according to the address signal.

[0008] The present invention also provides an electronic device, comprising a memory and a processor, wherein the processor is used to implement the steps of any of the above-mentioned data synchronization methods when executing a computer program stored in the memory.

[0009] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above-mentioned data synchronization methods are implemented.

[0010] Finally, the present invention also provides a data synchronization system, which includes at least a first computing unit and a second computing unit, the first computing unit includes an internal memory, multiple computing sub-units and a data synchronization processor; wherein the internal memory includes a global memory and multiple local memories, each computing sub-unit corresponds to each local memory, and each computing sub-unit only interacts with data with the corresponding local memory; when the first computing unit executes a computer program, the steps of the data synchronization method applied to the first computing unit are implemented; when the second computing unit executes the computer program, the steps of the data synchronization method applied to the second computing unit are implemented.

[0011] The advantage of the technical solution provided by the present invention is that the original data that needs to be processed by the first computing unit is partitioned and stored, and each partition can be used as a channel for external data reading and writing. According to the data dimension of the original data and the total number of data blocks divided, the number of local memories that a storage area can correspond to can be clarified, so that at the same time, the original data can be transmitted from each storage area of ​​the global memory to the local memory through multiple channels, effectively improving the data synchronization efficiency between the global memory and the local memory. In addition, in the process of data synchronization between the global memory and the local memory, the communication capability between the global memory and the data transmission interface is taken into account at the same time, which is not only conducive to maximizing the efficiency of data synchronization, but also ensuring the effectiveness of data synchronization. In addition, the present invention also provides a corresponding implementation system, electronic device and computer-readable storage medium for the data synchronization method, which further makes the method more practical, and the system, electronic device and computer-readable storage medium have corresponding advantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0013] Figure 1 A schematic diagram of a system structure in an exemplary application scenario provided by the present invention;

[0014] Figure 2 A flowchart of a data synchronization method in the related art;

[0015] Figure 3 A flowchart of a data synchronization method provided by the present invention;

[0016] Figure 4 A schematic diagram of segmenting the raw data to be processed provided by the present invention in an exemplary application scenario;

[0017] Figure 5 A schematic diagram of storing the raw data to be processed provided by the present invention in an exemplary application scenario;

[0018] Figure 6 A flowchart of another data synchronization method provided by the present invention;

[0019] Figure 7 A structural framework diagram of an exemplary implementation of a data synchronization device provided by the present invention;

[0020] Figure 8 A structural framework diagram of an exemplary implementation of a data synchronization device provided by the present invention;

[0021] Fig. 9 A schematic diagram of the structure of an exemplary implementation of the data synchronization system provided by the present invention;

[0022] Fig.10 A schematic diagram of a framework of a global memory provided by the present invention in an exemplary application scenario;

[0023] Fig.11 A schematic diagram of a framework of a data synchronization processor provided by the present invention in an exemplary application scenario;

[0024] Fig.12 A schematic diagram of data reading and writing in an exemplary application scenario of the data synchronization system provided by the present invention;

[0025] Fig.13Schematic diagram of protocol conversion of the data synchronization system provided by the present invention in an exemplary application scenario;

[0026] Fig.14 Schematic diagram of the framework of the data synchronization system provided by the present invention in an exemplary application scenario. Detailed implementation manners

[0027] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Among them, the terms "first", "second", "third", "fourth", etc. in the description and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0028] In order to improve the execution efficiency of computing tasks and enhance data computing performance, the heterogeneous computing method of using different types of processors or computing units in the same computing system to jointly complete computing tasks has emerged. A heterogeneous computing platform is composed of multiple different types of computing units with different processing capabilities, energy efficiencies, and applicable scopes, and the heterogeneous computing platform is used to execute tasks, such as a heterogeneous computing platform composed of a CPU (Central Processing Unit) and an FPGA (Field Programmable Gate Array).

[0029] When a heterogeneous computing platform executes tasks such as high-performance computing tasks, tasks with high real-time requirements, or tasks that involve large-scale data that need to be calculated, it will send computing tasks that require a large amount of computing resources or are highly complex to the acceleration computing unit with parallel computing capabilities for processing. The host will transmit the original data that needs to be processed by the acceleration computing unit to the global memory of the acceleration computing unit. After the acceleration computing unit finishes processing the data, the host reads the data processing result in the global memory of the acceleration computing unit. In order to improve the computing speed of the acceleration computing unit, the acceleration computing unit will instantiate multiple computing sub-units, and each computing sub-unit is uniquely configured with a local memory. Each computing sub-unit can only access data from its local memory. In order to ensure the accurate execution of data tasks, when the acceleration computing unit processes data, it is necessary to synchronize the data between the global memory and each local memory.

[0030] When the related technology synchronizes the global memory and the local memory, the predefined data transmission interface will first take out the original data from the global memory in sequence, and then the data transmission interface will transfer the original data to the corresponding local memory according to the address of the original data in the global memory. Each computing core reads data from the corresponding local memory for calculation. After the data calculation is completed, the data transmission interface retrieves the data from each local memory in sequence and stores it in the global memory, thereby realizing data synchronization between the global memory and the local memory. It can be seen that there is only one data transmission interface for data synchronization, and the data of all local memories are transmitted from one data transmission interface, which leads to low data transmission efficiency, affects the data processing efficiency of the entire accelerated computing unit, and cannot meet the user's efficient data processing needs.

[0031] In view of this, in order to solve the problems existing in the related technology, the present invention partitions and stores the original data that needs to be processed by the first computing unit, and each partition can be used as a channel for external data reading and writing. The number of local memories that a storage area can correspond to is determined according to the data dimension of the original data and the total number of data blocks divided. The communication capability between the global memory and the data transmission interface, the number object relationship between the storage area and the local memory are comprehensively considered. Multiple data blocks of the global memory are sent to the local memory in parallel, and the data processing results of multiple local memories can also be sent to the global memory, which is not only conducive to maximizing the efficiency of data synchronization, but also ensures the effectiveness of data synchronization. The specific application environment architecture or specific hardware architecture on which the execution of the data synchronization method depends is described below.

[0032] One of the application scenarios of the embodiment of the present invention is: the server is used as a host to process the user's high-performance computing tasks involving three-dimensional Fourier transform. In this application scenario, the FPGA is inserted into the PCIE (peripheral component interconnect express, high-speed serial computer expansion bus) card slot of the server, and the CPU and FPGA form a heterogeneous computing platform. Figure 1As shown, the high-performance computing task is performed through a heterogeneous computing platform, and the FPGA is used to perform the three-dimensional Fourier transform task. Multiple Fourier transform cores are instantiated in the FPGA to perform Fourier transform in parallel. Each Fourier transform core performs the corresponding Fourier transform task as a computing subunit to improve the calculation speed of the three-dimensional Fourier transform. Each Fourier transform core is uniquely configured with a local memory, and the Fourier transform core only accesses the three-dimensional Fourier data from its corresponding local memory to improve the data access speed during the calculation process. Thereby improving the overall heterogeneous computing efficiency of the FPGA's three-dimensional Fourier transform. In order to ensure the data processing accuracy of the FPGA, it is necessary to synchronize the data between the global memory and the local memory of the FPGA.

[0033] In this application scenario, in the related technology, data synchronization between the global memory and the local memory is achieved through a data loading module, such as Figure 2 As shown. The host 101 transfers the original three-dimensional matrix data to the global memory of FPGA102, and the data loading module sequentially takes out the three-dimensional matrix data from the global memory, and then transmits each local memory respectively according to the address of the three-dimensional matrix data in the global memory. After the three-dimensional Fourier transform is completed, the data is retrieved from each local memory in turn and stored in the global memory to achieve data synchronization between the global memory and the local memory. It can be seen that there is only one data transmission port between the global memory and the data loading module of the related art, and the data of all local memories are transmitted from one port, resulting in low data transmission efficiency. In view of this, the present invention divides the original three-dimensional matrix data into multiple data blocks through a global memory with multiple storage areas, and stores each data block in a corresponding storage area, and determines the number of local memories corresponding to a storage area based on the data dimension of the original three-dimensional matrix data and the total number of data blocks. According to the communication parameters between the global memory and the data transmission interface, the data transmission parameters between the global memory and each local memory are determined; according to the data transmission parameters, the data matching the quantity correspondence is read, and the corresponding read and write processing is performed according to the address signal, so as to realize parallel data synchronization between the global memory and the local memory, thereby effectively improving the efficiency of data synchronization. It should be noted that the above application scenarios are only shown to facilitate the understanding of the ideas and principles of the present invention, and the embodiments of the present invention are not limited in this respect. On the contrary, the embodiments of the present invention can be applied to any applicable scenario. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention are described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0034] First see Figure 3 , Figure 3A flow chart of a data synchronization method provided in this embodiment. The present invention is an implementation scheme for synchronizing data between the global memory and the local memory of a computing unit with parallel computing capability or accelerated computing capability in a heterogeneous computing process. For ease of description, the computing unit is defined as a first computing unit. The first computing unit includes a global memory and multiple local memories. The global memory includes multiple storage areas. The data synchronization method is described below with the first computing unit as the execution subject. This embodiment may include the following contents:

[0035] S301: When the second computing unit sends the data blocks of the original data to be processed to the corresponding storage area, the quantity correspondence between each storage area and the local memory is determined based on the data volume of the original data to be processed and the total number of data blocks.

[0036] In this embodiment, the second computing unit and the first computing unit constitute a heterogeneous computing platform. When the second computing unit receives a user task or a system task, it will decompose the task and send the data that needs to be accelerated by the first computing unit to the first computing unit. In this embodiment, such data is defined as raw data to be processed, that is, the raw data to be processed is processed by the first computing unit. When the second computing unit sends the raw data to be processed to the global memory of the first computing unit, the global memory stores it in each storage area. In order to realize the partitioned storage of the raw data, the raw data to be processed can be divided into multiple data blocks, and each data block can be stored in one storage area. When the global memory and the local memory are synchronizing data, it is necessary to transfer the data of each storage area of ​​the global memory to the local memory, or obtain each data processing result from the local memory and then store it in the global memory. After the present invention stores the raw data to be processed in partitions, it is also necessary to further determine how many local memories each storage area can interact with at the same time. In other words, how many local memories are needed to store the data blocks stored in a storage area. This embodiment is determined by the amount of data of the raw data to be processed and the total number of data blocks. The data amount is the total amount of data contained in the raw data to be processed or the data dimension of the raw data to be processed, which does not affect the implementation of the present invention. The total number of data blocks is the total number of blocks of the data blocks divided into the raw data to be processed. The corresponding relationship of quantity is that one storage area corresponds to several local memories, such as the data interaction between one storage area and 8 local memories. At this time, the corresponding relationship of quantity is that one storage area corresponds to 8 local memories.

[0037] S302: Determine data transmission parameters between the global memory and each local memory according to communication parameters between the global memory and the data transmission interface.

[0038] It is understandable that the global memory and the local memory of the first computing unit will transmit data through the data transmission interface when performing data synchronization. The data transmission interface of this embodiment is the port between the global memory and each local memory. Due to the different communication capabilities between the global memory and the data transmission interface, the communication parameters are parameters for translating the data transmission capability, such as bit width and clock frequency, and the data transmission capability of the global memory is subject to the physical parameters of the hardware device used. In order to efficiently read data from the global memory and transmit it to the local memory, it is necessary to convert the communication parameters between the global memory and the data transmission interface. For example, the bit width of each channel of the global memory is 256 bits, and each data is composed of a 32-bit real part plus a 32-bit imaginary part, so 4 data can be transmitted per clock cycle. The clock frequency of the global memory is 450MHz. In order to make full use of the high-speed characteristics of the global memory and efficiently retrieve data from the global memory, the bit width of the data transmission interface can be set to 512 bits, and the clock frequency used is 250MHz. In order to realize data communication between the two, it is necessary to convert the clock and bit width. The data transfer parameters refer to how to read data and how to transfer data when the communication capabilities of the global memory and the data transfer interface match, including but not limited to the data reading frequency, the amount of data read each time, and the data transfer frequency.

[0039] S303: Read data matching the quantity correspondence according to the data transmission parameters, and perform corresponding read and write processing according to the address signal.

[0040] After the previous step determines how to read data and how to perform data transmission, this embodiment determines the data reading position according to whether to send data from the global memory to the local memory or to obtain the data processing result from the local memory. The data reading position is the local memory or the global memory, and then reads the corresponding data from the data reading position according to the corresponding data reading method, such as the reading frequency and the amount of data read each time, until the condition of matching the quantity correspondence is met. For example, the quantity correspondence is that 1 storage area corresponds to 8 local memory devices, then in one clock cycle, the amount of data read from the global memory device is the amount of data that can be stored in the 8 local memory devices. When obtaining the data processing result, the data processing results stored in each of the 8 local memory devices are read simultaneously in one clock cycle. It can be understood that the interface protocols of the global memory device and the local memory device are different. For example, the interface protocol of the global memory device is the advanced extensible interface protocol, and the interface protocol of the local memory device is the local interface protocol of the random access memory, so in the process of data synchronization, the interface protocol conversion is also required. The so-called conversion of the interface protocol refers to determining the location where the data is to be written or read according to the address-related signals carried by the data during the data read and write operations. The address signals include read address-related signals and write address-related signals.

[0041] In the technical solution provided in this embodiment, the original data that needs to be processed by the first computing unit is partitioned and stored, and each partition can be used as a channel for external data reading and writing. According to the data dimension of the original data and the total number of divided data blocks, the number of local memories that can correspond to a storage area can be determined, so that the original data can be transferred from each storage area of ​​the global memory to the local memory through multiple channels at the same time, effectively improving the data synchronization efficiency between the global memory and the local memory. In addition, during the data synchronization process between the global memory and the local memory, the communication capability between the global memory and the data transmission interface is taken into consideration at the same time, which is not only conducive to maximizing the efficiency of data synchronization, but also ensuring the effectiveness of data synchronization.

[0042] In the above embodiment, there is no limitation on how to read data from the global memory and write it to the local memory, and how to read the data processing result from the local memory and then write it to the global memory. Based on the above embodiment, the present invention also provides an exemplary implementation method, which may include the following contents:

[0043] For the process of reading data from the global memory and writing it to the local memory: when each data block of the original data to be processed is transferred to the corresponding local memory, multiple target data are read from the corresponding address of the global memory according to the first read address signal; the target local memory location is determined according to the first write address signal, and a write valid signal is generated; based on the target local memory location, each target data is written into the corresponding target local memory.

[0044] For the process of storing the data processing results of the local memory into the global memory: when the data processing results of the original data to be processed are transmitted to the global memory, the source local memory location is determined according to the second read address signal; based on the source local memory location, the corresponding target data processing results are read from each source local memory, and each target data processing result is combined into a data processing result; the storage address of the global memory is determined according to the second write address signal, and based on the storage address of the global memory, the data processing result is written to the corresponding location of the global memory.

[0045] In this embodiment, for the convenience of description, the local memory currently being written is defined as the target memory, and the data written to each target memory is defined as the target data. The target local memory location includes the address information of each target local memory, and the total number of target local memories is the same as the total number of target data. Since the target data is read from the storage area and written to the local memory, and the data correspondence relationship describes the quantitative relationship between the storage area and the local memory, the total number of target data and the quantitative correspondence relationship are matched. The local memory that reads the data processing result is defined as the source memory, and the data processing result read from the local memory and written to the global memory is defined as the target data processing result. The source local memory location includes the address information of each source local memory, and the total number of each source local memory matches the quantitative correspondence relationship. In the process of reading data from the global memory and writing it to the local memory, the global memory is told from which address and how long the data is read through the read address related signal, and then the global memory feeds back the corresponding data to the data transmission interface. When the data transmission interface writes the read data to the local memory, a write valid signal can be generated by pulling up the signal, telling the local memory the write position of the data through the write address signal, and writing the target data through the write data signal. When the target data processing results are read out from the local memory and written into the global memory, the read address signal is used to tell multiple source local memories the location of the data to be read, and then the read data signal is used to simultaneously take out the corresponding data from these source local memories and combine the data, and the write address-related signal is used to tell the global memory the address where the data is written, and then the write data signal is used to write the data into the global memory.

[0046] From the above, it can be seen that this embodiment can accurately read data from the global memory and accurately write it to the local memory through interface protocol conversion. It can accurately obtain data processing results from the local memory and combine the data processing results into data of the interface bit width of the global memory, and efficiently write them to the global memory, which is beneficial to improving the data synchronization efficiency between the global memory and the local memory.

[0047] The above embodiment does not limit how to determine the quantity correspondence relationship between each storage area and the local memory. Based on the above embodiment, the present invention also provides an exemplary method for determining the quantity correspondence relationship between the storage area and the local memory, which may include the following contents:

[0048] According to the data dimension of the original data to be processed, the local storage parameters of each local memory are determined; based on the local storage parameters, the number of data stored in each local memory is determined; according to the number of data stored in each storage area and the number of data stored in each local memory, the quantity correspondence between each storage area and the local memory is determined.

[0049] In this embodiment, the raw data to be processed is three-dimensional matrix data. Accordingly, the raw data to be processed can use data dimensions to represent the data volume. The local storage parameter refers to how the local memory stores data and the total number of local memories. For example, the three-dimensional matrix data includes three planes. The local storage parameter is that the data of each XZ plane is stored in a local memory, or the data of several XZ planes are stored in a local memory. After determining how to store the data, the number of local memories used can be determined according to the data volume. After determining the local storage parameter, the number of data stored in each local memory can be determined. Taking the raw data to be processed as 128*128*128 as an example, the data of each XZ plane is stored in a local memory, and 128 XZ planes use 128 local memories, that is, 128 local memories are used to store data, so the number of data stored in each local memory is 128×128. If a local memory stores data of other number of planes, the number of local memories used can be a number that divides 128, such as 16, 32, and 64. Taking the total number of data blocks as 16 as an example, the data are stored in 16 storage areas of the global memory respectively, that is, the number of data stored in each storage area is 8×128×128, and each storage area exchanges data with 8 local memories.

[0050] As can be seen from the above, this embodiment determines the quantitative correspondence between the storage area and the local memory according to the amount of original data to be processed and the storage capacity of the local memory, so that the global memory and the local memory can achieve parallel data synchronization, which is conducive to improving data synchronization efficiency.

[0051] The above embodiment does not limit how to determine the data transmission parameters. Based on the above embodiment, the present invention also provides an exemplary method for determining the data transmission parameters between the global memory and each local memory, which may include the following contents:

[0052] Obtain the bit width and clock frequency of the global memory and the data transmission interface; determine the read parameters and send parameters for reading the global memory according to the bit width and clock frequency of the global memory and the data transmission interface, and use the read parameters and send parameters as data transmission parameters.

[0053] Among them, the bit width is the width of data processed or transmitted in a computer system, that is, the number of data bits that can be processed or transmitted at one time, which can reflect the data transmission capacity. The clock frequency determines the number of basic operations or instructions that can be executed per second. This embodiment uses the bit width and clock frequency to measure the data transmission capacity. The reading parameter includes at least the number of data reads, the amount of data and the reading frequency, and the sending parameter includes the sending frequency. Exemplarily, the data reading parameter for reading the global memory can be determined according to the global bit width of the global memory and the interface bit width of the data transmission interface; the data reading parameter includes at least the number of data reads and the amount of data; according to the global clock frequency of the global memory and the interface clock frequency of the data transmission interface, the data reading frequency for reading the global memory and the data sending frequency for sending data from the global memory are determined; finally, the data reading parameter and the data reading frequency are used as the reading parameter, and the data sending frequency is used as the sending parameter. Taking the bit width of the global memory as 256 bits, the bit width of the data transmission interface as 512 bits, the clock frequency of the global memory as 450MHz, and the clock frequency of the data transmission interface as 250MHz as an example, the clock and bit width conversion is implemented using the relevant IP provided by the FPGA supplier. In simple terms, the bit width of the high bandwidth memory is 256 bits, the data loading module is 512 bits, and the data transmission interface reads 256 bits of data from the global memory twice to form a 512-bit data to the data transmission interface. The clock frequency of the global memory is nearly twice that of the data transmission interface, which reads from the global memory twice at a frequency of 450MHz, and then transmits data once to the data transmission interface at a frequency of 250MHz. If the data is not enough, it can wait until the data volume reaches the requirement before performing the corresponding operation.

[0054] As can be seen from the above, this embodiment represents the data transmission capability through bit width and clock frequency. By converting the bit width and clock frequency between the global memory and the data transmission interface, it can ensure that data is transmitted efficiently and accurately, which is conducive to improving data synchronization efficiency.

[0055] In order to ensure that the local memory can store continuous data, facilitate the first computing unit to process the data, and improve the data processing efficiency of the first computing unit, based on the above embodiment, the present invention also provides a storage method for the raw data to be processed, which may include the following contents:

[0056] The original data to be processed is three-dimensional matrix data, and the data values ​​of each data block on the target coordinate axis match the corresponding quantity relationship; when storing each data block, the storage area of ​​the global memory regards the data of the data block to be stored in the direction of the target coordinate axis as a group, and stores the data of the data block to be stored in sequence, so as to store each group of data in the corresponding local memory group in each clock cycle.

[0057] In this embodiment, the local memory group includes a plurality of local memories, and the number of local memories is the same as the number of data in each group. The original data to be processed is three-dimensional matrix data, that is, it has data in the directions of the three coordinate axes, wherein the target coordinate axis is, for example, the direction of the original data to be processed. Also, taking the data dimension of the original data to be processed as 128×128×128 as an example, Figure 4 As shown in the figure, each small cube represents 8×8×8 data. The global memory includes 16 storage areas, the target coordinate axis is the Y axis, and the raw data to be processed is divided into 16 data blocks in the Y axis direction. Each storage area of ​​the global memory stores 128×8×128 data (Z direction×Y direction×X direction), corresponding to Figure 4 A data block in the Z direction, the data of each plane, that is, the XY plane, such as Figure 5 shown.

[0058] Among them, since the data bit width of the data transmission interface is 512 bits and the bit width of each data is 64 bits, 8 data can be transmitted in each clock cycle. Therefore, when the original data is to be processed, the 8 vertical data are regarded as a group, such as Figure 5 As shown, 1-1 to 8-1 can be stored first, followed by 1-2 to 8-2, and so on for subsequent data. When the data transmission interface takes out data from each storage area, it takes out 512 bits of data, that is, 8 data, in each clock cycle, and then stores these 8 data in 8 local memories respectively. In the next cycle, another 8 data are taken and stored in 8 local memories respectively, and so on for subsequent data. When 128 clock cycles are completed, the data of one plane can be transmitted. Each storage area stores 128 planes of data, so 128 rounds of reading are required to complete the synchronization of all data in each storage area. This data storage and reading method can ensure that the data stored in the local memory is continuous in the X direction, such as Figure 51-1, 1-2, 1-3, ..., 1-128 shown. Taking the first computing unit performing the three-dimensional Fourier transform task as an example, when each sub-computing unit performs the X-dimensional Fourier transform, that is, when performing the first dimension of the three-dimensional Fourier transform, it can directly read data continuously from the local memory. After the three-dimensional Fourier transform is completed, the process of writing the data processing results from the local memory back to the global memory: the data transmission interface reads one data from the 8 local memories in each clock cycle, forms a 512-bit data and writes it back to the global memory. When the next clock cycle comes, the data is read again to form a 512-bit data and write it back to the global memory until all the data is written back to the global memory.

[0059] From the above, it can be seen that this embodiment stores the original data to be processed as a group of data in a local memory in the target time axis direction, so that when the first computing unit executes the task, it can directly read the data continuously from the local memory for processing, which is conducive to improving data processing efficiency.

[0060] It can be understood that the present invention is directed to heterogeneous computing. The above embodiment is a method flow executed by one computing unit of a heterogeneous computing platform. The following embodiment describes a method flow executed by another computing unit during data synchronization. For ease of description, the computing unit is defined as a second computing unit. Figure 6 As shown, this embodiment may include the following contents:

[0061] S601: Acquire original data to be processed, and divide the original data to be processed into multiple data blocks.

[0062] S602: Send each data block to the corresponding storage area of ​​the global memory of the first computing unit, so that the first computing unit determines the quantity correspondence between each storage area and the local memory based on the data volume of the original data to be processed and the total number of data blocks; determines the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; reads the data matching the quantity correspondence according to the data transmission parameters, and performs corresponding read and write processing according to the address signal.

[0063] The same steps as those in the above embodiment can be found in the above embodiment, which will not be described again.

[0064] From the above, it can be seen that this embodiment divides the original data to be processed into multiple data blocks through the second computing unit, and sends the data blocks to the second computing unit, so that the second computing unit can realize partitioned storage of the original data, and then realize parallel data synchronization between the global memory and the local memory, thereby effectively improving the efficiency of data synchronization.

[0065] Based on the above embodiment, the present invention further provides an exemplary data segmentation method, which may include the following contents:

[0066] Based on the total number of storage areas of the global memory of the first computing unit, the original data to be processed is divided into data blocks to obtain a plurality of data blocks; and each data block is stored in a corresponding storage area through a read-write channel of the global memory.

[0067] In this embodiment, the number of data blocks into which the raw data to be processed is divided is equal to the number of storage areas of the global memory used. Each storage area stores one data block. If 16 storage areas are used, the data is divided into 16 data blocks along the Y direction. This embodiment takes 128×128×128 three-dimensional matrix data as the raw data to be processed and the global memory includes 16 storage areas as an example for explanation. Figure 4 As shown, each small cube represents 8×8×8 data, with 16 small squares in each dimension. The second computing unit divides the three-dimensional matrix data into 16 data blocks along the Y direction, and then uses the direct memory access engine to store the raw data to be processed into 16 storage areas through a channel of the global memory, thereby completing the host sending the raw data.

[0068] It should be noted that there is no strict order of execution between the steps in the present invention. As long as they comply with the logical order, these steps can be executed simultaneously or in a certain preset order. Figure 3 and Figure 6 This is just a schematic and does not mean that this is the only execution order.

[0069] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, or by hardware, but in many cases the former is a better implementation method. The present invention also provides a corresponding device for the data synchronization method, which further makes the method more practical. The device can be described from the perspective of functional modules and hardware.

[0070] The following is an introduction to the data synchronization device provided by the present invention based on the perspective of functional modules. Figure 7 The device is applied to a first computing unit including a global memory and a plurality of local memories, the global memory of the first computing unit including a plurality of storage areas, which is used to implement the data synchronization method provided by the present invention, and the device may include:

[0071] The quantity relationship determination module 701 is used to determine the quantity correspondence relationship between each storage area and the local memory based on the data volume of the original data to be processed and the total number of data blocks when the second computing unit sends the data blocks of the original data to be processed to the corresponding storage area;

[0072] The parameter conversion module 702 is used to determine the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface;

[0073] The data reading and writing module 703 is used to read the data matching the quantity correspondence according to the data transmission parameters, and perform corresponding reading and writing processing according to the address signal.

[0074] In this embodiment, the data synchronization device may include or be divided into multiple program modules, namely, a quantity relationship determination module 701, a parameter conversion module 702, and a data reading and writing module 703. These program modules are stored in a storage medium and executed by one or more processors, thereby completing the data synchronization method disclosed in the above method embodiment. The program module referred to in this embodiment refers to a series of computer program instruction segments that can perform specific functions, which is more suitable for describing the execution process of the data synchronization device in the storage medium than the program itself. For the description of the features in the embodiment corresponding to the data synchronization device, please refer to the relevant description of the embodiment corresponding to the data synchronization method, and will not be repeated here.

[0075] Exemplarily, in some implementations of the present embodiment, the above-mentioned data reading and writing module 703 can also be used for: when each data block of the original data to be processed is transferred to the corresponding local memory, multiple target data are read from the corresponding address of the global memory according to the first read address signal; the target local memory location is determined according to the first write address signal, and a write valid signal is generated; based on the target local memory location, each target data is written into the corresponding target local memory respectively; wherein the target local memory location includes the address information of each target local memory, the total number of target local memories is the same as the total number of target data, and the total number of target data matches the quantity correspondence.

[0076] Illustratively, in some other implementations of this embodiment, the above-mentioned data read and write module 703 can also be used for: when the data processing result of the original data to be processed is transmitted to the global memory, the source local memory location is determined according to the second read address signal; based on the source local memory location, the corresponding target data processing result is read from each source local memory, and each target data processing result is combined into a data processing result; the storage address of the global memory is determined according to the second write address signal, and based on the storage address of the global memory, the data processing result is written to the corresponding location of the global memory; wherein the source local memory location includes the address information of each source local memory, and the total number of each source local memory matches the quantity correspondence.

[0077] Exemplarily, in some other embodiments of this embodiment, the above-mentioned quantity relationship determination module 701 may also be used to: determine the local storage parameters of each local memory according to the data dimension of the raw data to be processed; determine the number of data stored in each local memory based on the local storage parameters; and determine the quantity correspondence relationship between each storage area and the local memory according to the number of data stored in each storage area and the number of data stored in each local memory.

[0078] Exemplarily, in some other embodiments of this embodiment, the above-mentioned parameter conversion module 702 may also be used to: obtain the bit width and clock frequency of the global memory and the data transmission interface; determine the read parameters and send parameters for reading the global memory according to the bit width and clock frequency of the global memory and the data transmission interface, and use the read parameters and send parameters as data transmission parameters; wherein, the read parameters at least include the number of data reads, the data volume, and the read frequency, and the send parameter includes the send frequency.

[0079] As an exemplary implementation of the above embodiment, the above-mentioned parameter conversion module 702 may further be used to: determine the data read parameters for reading the global memory according to the global bit width of the global memory and the interface bit width of the data transmission interface; the data read parameters at least include the number of data reads and the data volume; determine the data read frequency for reading the global memory and the data send frequency for the global memory to send data according to the global clock frequency of the global memory and the interface clock frequency of the data transmission interface; and use the data read parameters and the data read frequency as the read parameters, and the data send frequency as the send parameter.

[0080] Exemplarily, in some other embodiments of this embodiment, the above-mentioned device may further include a raw data storage module, and this module may be used to: when the raw data to be processed is three-dimensional matrix data, the numerical values of the number of data in each data block on the target coordinate axis match the quantity correspondence relationship; when the storage area of the global memory stores each data block, each data in the direction of the target coordinate axis of the data block to be stored is used as a group, and each data of the data block to be stored is stored in sequence, so as to store each group of data in the corresponding local memory group in each clock cycle; wherein, the local memory group includes multiple local memories, and the number of local memories is the same as the number of data in each group.

[0081] Exemplarily, the data synchronization device applied to the second computing unit will be introduced from the perspective of functional modules below. Please refer to Figure 8 , and this device may include:

[0082] A data cutting module 801, configured to cut the obtained raw data to be processed into multiple data blocks.

[0083] The data sending module 802 is used to send each data block to the corresponding storage area of ​​the global memory of the first computing unit, so that the first computing unit determines the quantity correspondence between each storage area and the local memory based on the data volume of the original data to be processed and the total number of data blocks; determines the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; reads the data matching the quantity correspondence according to the data transmission parameters, and performs corresponding read and write processing according to the address signal.

[0084] Exemplarily, in some implementations of the present embodiment, the data segmentation module 801 may also be used to: perform data block segmentation on the original data to be processed based on the total number of storage areas of the global memory of the first computing unit to obtain a plurality of data blocks; and store each data block in a corresponding storage area through the read and write channels of the global memory.

[0085] The data synchronization device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from a hardware perspective. The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned data synchronization method embodiments.

[0086] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data synchronization method embodiments when running.

[0087] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0088] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above data synchronization method embodiments are implemented.

[0089] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data synchronization method embodiments are implemented.

[0090] Finally, the present invention also provides a data synchronization system, see Fig. 9The data synchronization system at least includes a first computing unit 901 and a second computing unit 902, wherein the first computing unit includes an internal memory, a plurality of computing sub-units and a data synchronization processor; wherein the internal memory includes a global memory and a plurality of local memories, each computing sub-unit corresponds to each local memory, and each computing sub-unit only interacts with data with the corresponding local memory; when the first computing unit executes a computer program, the steps of the data synchronization method described in any one of the method embodiments executed by the first computing unit are implemented; when the second computing unit executes a computer program, the steps of the data synchronization method described in any one of the method embodiments executed by the two computing units are implemented.

[0091] It can be seen from the above that the data synchronization system of this embodiment can realize parallel data synchronization between the global memory and the local memory, thereby effectively improving the efficiency of data synchronization.

[0092] The above embodiment does not impose any limitation on the data structure of the global memory. In order to improve the data synchronization efficiency, this embodiment also provides a high-performance data storage structure, which may include the following contents:

[0093] The global memory includes at least a first interconnection unit and a second interconnection unit, and the first interconnection unit and the second interconnection unit are connected to each other; the first interconnection unit includes at least a first interconnection module and a second interconnection module, and the first interconnection module and the second interconnection module are connected through the first interconnection unit; the first interconnection module includes at least a first storage area and a second storage area, and the first storage area and the second storage area are connected through the first interconnection module, and the second interconnection module includes at least a third storage area and a fourth storage area, and the third storage area and the fourth storage area are connected through the second interconnection module; the second interconnection unit includes at least a third interconnection module and a fourth interconnection module, and the fourth interconnection module and the third interconnection module are connected through the second interconnection unit; the third interconnection module includes at least a fifth storage area and a sixth storage area, and the fifth storage area and the sixth storage area are connected through the third interconnection module, and the fourth interconnection module includes at least a seventh storage area and an eighth storage area, and the seventh storage area and the eighth storage area are connected through the fourth interconnection module.

[0094] In this embodiment, the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area are independent of each other and correspond to a read-write channel respectively; there is at least one read-write channel to perform data read and write operations on any number of storage areas in the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area. That is, there are multiple independent storage areas inside the global memory, and data can be accessed independently through the corresponding channels. Fig.10As shown, the global memory includes 32 independent storage areas, every two storage areas are fully connected through the interconnection module, every two interconnection modules are fully connected through the interconnection unit, and the interconnection units are also connected. The global memory provides 32 channels to the outside, and each channel can perform data read and write operations on all storage areas. In order to improve the data reading and writing efficiency, the global memory performs read and write operations on at least the first storage area and the second storage area at the same time through at least the first read and write channel and the second read and write channel. In other words, the global memory enables 32 channels at the same time to perform parallel data read and write operations on 32 storage areas at the same time.

[0095] As can be seen from the above, this embodiment provides a global memory with large-capacity storage and efficient memory bandwidth, which achieves a shorter data path and a higher data transmission rate, reduces access delay, and is conducive to improving data reading and writing efficiency.

[0096] In the actual application process, this embodiment can also instantiate the relevant methods recorded in the above embodiments into multiple functional modules for quick application. Based on the above embodiments, any instantiation method can be used to instantiate the data synchronization method into multiple data adaptation modules and multiple data loading modules. The data adaptation module group is configured to determine the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; the data loading module group is configured to read the data matching the quantity correspondence according to the data transmission parameters, and perform corresponding read and write processing according to the address signal. For the convenience of description, multiple data adaptation modules can be defined as a data adaptation module group, and multiple data loading modules can be defined as a data loading module group. Accordingly, the data synchronization processor includes a data adaptation module group and a corresponding data loading module group. Fig.11 As shown, the number of data adaptation modules included in the data adaptation module group is the same as the number of data loading modules included in the data loading module group, and both are the same as the total number of storage areas; each data loading module includes a first data transmission interface and a second data transmission interface, each data adaptation module is connected to the corresponding data loading module through the first data transmission interface, and each data loading module is connected to the corresponding local memory device group through the second data transmission interface; wherein the local memory device group includes multiple local memories, and the number of local memories is determined based on the corresponding relationship between the number of storage areas and the local memories.

[0097] In this embodiment, in order to further improve the data synchronization efficiency, the first data transmission interface of the data loading module adopts the advanced scalable interface, and the second data transmission interface adopts the local interface of the random access memory. Accordingly, the interface between the data loading module and the data adaptation module adopts the advanced scalable interface for data interaction, and the interface between the data loading module and the local memory adopts the local interface of the random access memory for data interaction. The data bit width of the advanced scalable interface is 512 bits, and the bit width of each data is 64 bits. Therefore, each time the data loading module reads data, it contains 8 actual data, which are written into 8 local memories respectively. Fig.12 As shown in FIG. 1 , in the process of reading the raw data to be processed from the global memory and writing it into the local memory, the data loading module can inform the global memory from which address and how long the data is to be read through the read address related signal, and the global memory feeds back the corresponding data to the data loading module. Among them, the read address related signal is as follows: Fig.13 As shown, when the data loading module writes the raw data to be processed, it can generate a write valid signal by pulling up the signal, inform the local memory of the write location of the raw data to be processed through the write address related signal, and write the data through the write data signal. Fig.13 As shown in the figure, when the data processing result is read from the local memory and written into the global memory, the read address signal is used to inform the 8 local memories of the location of the data to be read, and the read data signal is used to simultaneously obtain 8 64-bit data from the 8 local memories and combine them into a 512-bit data. The write address-related signal is used to inform the global memory of the address where the data processing result is written, and finally the data is written into the global memory through the write data signal.

[0098] As can be seen from the above, this embodiment forms a data loading module through an advanced extensible interface and a random access memory local interface, and implements protocol conversion between the global memory and the local memory through the data loading module. It can accurately read data from the global memory and accurately write it to the local memory. It can accurately obtain data processing results from the local memory, and combine the data processing results into data of the interface bit width of the global memory, and efficiently write them to the global memory, which is beneficial to improving the data synchronization efficiency between the global memory and the local memory.

[0099] In order to make the technical solution of the present invention more clear to those skilled in the art, the present invention also provides an exemplary data synchronization system, such as Fig.14As shown, the first computing unit is an FPGA, and the FPGA uses a high-bandwidth memory as a global memory. The high-bandwidth memory includes 16 storage areas. Since these 16 storage areas are interconnected, all storage areas can be accessed through one channel. The host sends data to the global memory of the first computing unit through a direct memory access engine. Because the direct memory access engine has only one channel, data can only be sent to 16 storage areas through one channel, and data in different storage areas are distinguished according to the address. In order to improve the efficiency of the first computing unit in reading and writing data from the global memory, the first computing unit uses 16 channels to read and write 16 storage areas simultaneously when reading and writing data from the global memory to the local memory. The second computing unit is the CPU on the host. The parallel data synchronization process between the global memory and the local memory in the three-dimensional Fourier transform based on the data synchronization system may include the following contents:

[0100] The host runs a target upper-layer application, which is an application that can use a direct memory access engine. The target upper-layer application uses the direct memory access engine to drive the original data to be processed in the host memory and sends it to 16 storage areas through the direct memory access engine. After the original data to be processed is sent, the host sends an instruction to the first computing unit to inform the first computing unit that the data sending is completed. After receiving the instruction, the first computing unit uses 16 data loading modules to read data from the 16 storage areas in parallel and store them in 128 local memory devices respectively. After completing the data synchronization between the global memory and the local memory, each computing subunit of the first computing unit takes out data from the local memory for three-dimensional Fourier transform, and returns the data to the local memory after the calculation is completed. Each data loading module reads the calculation results from the 128 local memory devices and stores them in 16 storage areas. After completing the data synchronization between the global memory and the local memory, the first computing unit can send a task completion message to the host by way of an interrupt. The target upper-layer application again uses the access engine driver to read the data in the high-bandwidth memory back to the host memory through the direct memory access engine, completing the parallel data synchronization between the global memory and the local memory in the three-dimensional Fourier transform.

[0101] The above is a detailed introduction to a data synchronization method, device, electronic device, computer-readable storage medium and computer program product provided by the present invention. The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can refer to each other. Whether the units and algorithm steps of each example described in each disclosed embodiment are executed in electronic hardware or computer software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods for each specific application to implement the described functions, and such implementation should not be considered to exceed the scope of the present invention. Without departing from the principles of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.

Claims

1. A data synchronization method, characterized in that: Applied to a first computing unit including a global memory and a plurality of local memories, wherein the global memory of the first computing unit includes a plurality of storage areas, including: When the second computing unit sends the data blocks of the raw data to be processed to the corresponding storage area, based on the data volume of the raw data to be processed and the total number of data blocks, the quantity correspondence between each storage area and the local memory is determined; Determine the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; wherein the data transmission interface is a port between the global memory and each local memory; According to the data transmission parameters, data matching the quantity correspondence is read, and corresponding read and write processing is performed according to the address signal.

2. The data synchronization method according to claim 1, characterized in that: The reading of data matching the quantity correspondence and performing corresponding read and write processing according to the address signal includes: When each data block of the raw data to be processed is transmitted to the corresponding local memory, a plurality of target data are read from the corresponding address of the global memory according to the first read address signal; Determine the target local memory location according to the first write address signal, and generate a write valid signal; Based on the target local memory location, each target data is written into the corresponding target local memory; The target local memory locations include address information of each target local memory, the total number of target local memories is the same as the total number of target data, and the total number of target data matches the quantity correspondence.

3. The data synchronization method according to claim 1, characterized in that: The reading of data matching the quantity correspondence and performing corresponding read and write processing according to the address signal includes: When the data processing result of the raw data to be processed is transmitted to the global memory, the source local memory location is determined according to the second read address signal; Based on the source local memory locations, reading corresponding target data processing results from each source local memory, and combining each target data processing result into a data processing result; Determine a storage address of the global memory according to the second write address signal, and write the data processing result to a corresponding position of the global memory based on the storage address of the global memory; The source local memory locations include address information of each source local memory, and the total number of each source local memory matches the quantity correspondence.

4. The data synchronization method according to claim 1, characterized in that: The determining the quantity correspondence between each storage area and the local memory based on the data volume of the raw data to be processed and the total number of data blocks includes: Determining local storage parameters of each local memory according to the data dimension of the original data to be processed; Based on the local storage parameters, determine the number of data stored in each local memory; According to the number of data stored in each storage area and the number of data stored in each local memory, the quantity correspondence between each storage area and the local memory is determined.

5. The data synchronization method according to any one of claims 1 to 4, characterized in that: Determining the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface includes: Obtaining the bit width and clock frequency of the global memory and the data transmission interface; Determine, according to the bit width and clock frequency of the global memory and the data transmission interface, a read parameter and a send parameter for reading the global memory, and use the read parameter and the send parameter as the data transmission parameter; The reading parameters include at least the number of data reads, the data volume and the reading frequency, and the sending parameters include the sending frequency.

6. The data synchronization method according to claim 5, characterized in that: The step of determining the read parameters and the send parameters of the global memory according to the bit width and the clock frequency of the global memory and the data transmission interface comprises: Determine data reading parameters for reading the global memory according to the global bit width of the global memory and the interface bit width of the data transmission interface; the data reading parameters at least include the number of data reads and the amount of data; Determine a data reading frequency of the global memory and a data sending frequency of the global memory according to a global clock frequency of the global memory and an interface clock frequency of the data transmission interface; The data reading parameter and the data reading frequency are used as reading parameters, and the data sending frequency is used as the sending parameter.

7. The data synchronization method according to any one of claims 1 to 4, characterized in that: The raw data to be processed is three-dimensional matrix data, and the data value of each data block on the target coordinate axis matches the quantity correspondence relationship; After the second computing unit sends the data block of the raw data to be processed to the corresponding storage area, the method further includes: When storing each data block, the storage area of ​​the global memory sequentially stores each data of the data block to be stored in the direction of the target coordinate axis as a group, so as to store each group of data into the corresponding local memory group in each clock cycle; The local memory group includes a plurality of local memories, and the number of the local memories is the same as the number of each group of data.

8. A data synchronization method, characterized in that: Applied to the second computing unit, comprising: Acquire raw data to be processed, and divide the raw data to be processed into multiple data blocks; Each data block is sent to a corresponding storage area of ​​the global memory of the first computing unit, so that the first computing unit determines the quantity correspondence between each storage area and the local memory based on the data volume of the original data to be processed and the total number of data blocks; determines the data transmission parameters between the global memory and each local memory according to the communication parameters of the global memory and the data transmission interface; reads the data matching the quantity correspondence according to the data transmission parameters, and performs corresponding read and write processing according to the address signal.

9. The data synchronization method according to claim 8, characterized in that: The step of dividing the raw data to be processed into a plurality of data blocks comprises: Based on the total number of storage areas of the global memory of the first computing unit, dividing the original data to be processed into data blocks to obtain a plurality of data blocks; Each data block is stored in a corresponding storage area through the read and write channels of the global memory.

10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data synchronization method as claimed in any one of claims 1 to 9 when executing the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data synchronization method according to any one of claims 1 to 9 are implemented.

12. A data synchronization system, comprising at least a first computing unit and a second computing unit, wherein the first computing unit comprises an internal memory, a plurality of computing subunits and a data synchronization processor; in, The internal memory includes a global memory and a plurality of local memories, each computing subunit corresponds to each local memory, and each computing subunit only exchanges data with the corresponding local memory; When the first computing unit executes the computer program, the steps of the data synchronization method according to any one of claims 1 to 7 are implemented; when the second computing unit executes the computer program, the steps of the data synchronization method according to claim 8 or 9 are implemented.

13. The data synchronization system according to claim 12, characterized in that: The data synchronization processor includes a data adaptation module group and a corresponding data loading module group; the number of data adaptation modules included in the data adaptation module group is the same as the number of data loading modules included in the data loading module group, and both are the same as the total number of storage areas; Each data loading module includes a first data transmission interface and a second data transmission interface, each data adaptation module is connected to the corresponding data loading module through the first data transmission interface, and each data loading module is connected to the corresponding local memory group through the second data transmission interface; wherein the local memory group includes a plurality of local memories, and the number of local memories is determined based on the corresponding relationship between the number of storage areas and the local memories; The data adaptation module group is configured to determine the data transmission parameters between the global memory and each local memory according to the communication parameters between the global memory and the data transmission interface; The data loading module group is configured to read the data matching the quantity correspondence according to the data transmission parameters, and perform corresponding reading and writing processing according to the address signal.

14. The data synchronization system according to claim 12, characterized in that: The global memory comprises at least a first interconnect unit and a second interconnect unit, wherein the first interconnect unit and the second interconnect unit are connected to each other; The first interconnection unit at least includes a first interconnection module and a second interconnection module, and the first interconnection module and the second interconnection module are connected through the first interconnection unit; the first interconnection module at least includes a first storage area and a second storage area, and the first storage area and the second storage area are connected through the first interconnection module, and the second interconnection module at least includes a third storage area and a fourth storage area, and the third storage area and the fourth storage area are connected through the second interconnection module; The second interconnection unit includes at least a third interconnection module and a fourth interconnection module, and the fourth interconnection module and the third interconnection module are connected through the second interconnection unit; the third interconnection module includes at least a fifth storage area and a sixth storage area, and the fifth storage area is connected to the sixth storage area through the third interconnection module; the fourth interconnection module includes at least a seventh storage area and an eighth storage area, and the seventh storage area is connected to the eighth storage area through the fourth interconnection module; Among them, the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area are independent of each other and correspond to a read-write channel respectively; there is at least one read-write channel to perform data read and write operations on any number of storage areas among the first storage area, the second storage area, the third storage area, the fourth storage area, the fifth storage area, the sixth storage area, the seventh storage area and the eighth storage area.

15. The data synchronization system according to claim 14, characterized in that: The global memory simultaneously performs read and write operations on at least the first storage area and the second storage area through at least a first read and write channel and a second read and write channel.

Citation Information

Patent Citations

  • Multi-core fine grit synchronous DMA transmission method used for GPDSP

    CN104615557A

  • Data processing method and device, reduction server and mapping server

    CN115203133A

  • Storage and calculation architecture FPGA supporting instruction broadcast

    CN116775554A

  • Cache processing method of heterogeneous device, heterogeneous system, product, device and medium

    CN119473168A

  • Method for rapid communication within a parallel computer system, and a parallel computer system operated by the method

    US20020078322A1

Cited By

  • Data synchronization method and system, electronic device, and computer-readable storage medium

    WO2026179820A1