Method and computer program product for data synchronization among multiple acceleration cards and computing device
By adopting a unified transmission control protocol (TCP) and data synchronization nodes in distributed computing of multi-accelerator cards, the problem of low synchronization efficiency among different accelerator cards is solved, and efficient data synchronization and computing efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510257853.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-18
AI Technical Summary
In distributed computing of multi-acceleration cards, the prior art requires different synchronization mechanisms to realize data synchronization between different types of accelerator cards, resulting in low communication efficiency and inability to become a universal solution.
Data synchronization is achieved between accelerator cards by using a unified transmission control protocol (TCP). By inserting data synchronization nodes in the computing subtask, data exchange and synchronization are performed using the central processor, and data transmission is optimized using a central controller or message subscription mechanism.
It reduces the data conversion time between different communication protocols, improves data synchronization efficiency, and thus improves the overall computing efficiency.
Smart Images

Figure CN120335944A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of technical resource scheduling, and in particular, to a method, a computer program product, and a computing device for data synchronization between multiple acceleration cards. Background Art
[0002] With the development of domestic enterprise informatization and deep learning, more and more artificial intelligence technologies are being used by many enterprises and governments in actual business. The scale of calculation and the required computing power of AI (Artificial Intelligence) systems are continuously increasing, and the computing power of a single acceleration card is no longer sufficient to meet the needs of complex model training and inference. Therefore, distributed computing using multiple acceleration cards to improve computing power has become one of the mainstream problem-solving solutions.
[0003] When using multiple acceleration cards for distributed computing, when the calculation result on a certain node requires the result data of the calculation of other acceleration cards, data transmission and synchronization between the acceleration cards are required. Different types of acceleration cards (NPU, GPU, TPU, etc.) provide different data transmission and synchronization solutions to achieve data synchronization, so different synchronization mechanisms need to be implemented when using different acceleration cards for distributed computing. When using different acceleration cards for distributed computing, different customized synchronization mechanisms need to be implemented to ensure synchronization between multiple acceleration cards, which not only has low communication efficiency, but also cannot be a general solution for multi-acceleration card inference.
[0004] Therefore, a technical solution is needed that can use a unified solution to achieve data synchronization, minimize communication time as much as possible, and improve data synchronization efficiency. Summary of the Invention
[0005] The present invention aims to provide a method, a computer program product, and a computing device for data synchronization between multiple acceleration cards, which can use a unified solution to achieve data synchronization, minimize communication time as much as possible, and improve data synchronization efficiency.
[0006] According to one aspect of the present invention, there is provided a method for data synchronization between multiple acceleration cards, including:
[0007] Loading a decomposed artificial intelligence computing task and distributing it to at least two acceleration cards, each acceleration card executing the assigned computing subtask, and the at least two acceleration cards including a first acceleration card;
[0008] Adding data synchronization nodes to the computing subtasks of each acceleration card, the data synchronization nodes including a first data synchronization node for the first computing subtask of the first acceleration card;
[0009] During the execution of the first computing subtask on the first acceleration card, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards;
[0010] Continue to execute the computing subtask until the calculation ends,
[0011] Among them, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards based on the Transmission Control Protocol, and the data synchronization node is executed on the central processing unit.
[0012] According to some embodiments, adding data synchronization nodes to the computing subtasks of each acceleration card includes:
[0013] Insert a data synchronization node after the computing nodes that need to synchronize data in each computing subtask.
[0014] According to some embodiments, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards, including:
[0015] When the first acceleration card executes to the data synchronization node of the first computing subtask, call the Transmission Control Protocol to send a synchronization request.
[0016] According to some embodiments, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards, including:
[0017] Wait until the computing subtasks on the other acceleration cards execute to the corresponding data synchronization nodes, and then start data exchange and synchronization.
[0018] According to some embodiments, the method further includes: using a central control machine to implement data exchange and synchronization.
[0019] According to some embodiments, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards using a message subscription mechanism.
[0020] According to some embodiments, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computing subtasks of the other acceleration cards using a message subscription mechanism, including:
[0021] The first computing subtask subscribes to the data messages of the corresponding data synchronization nodes of the retrieval subtasks of the other acceleration cards at the first data synchronization node;
[0022] After the first computing subtask receives the data message, perform data exchange and synchronization based on the data message.
[0023] According to some embodiments, the acceleration card includes one or more of a GPU, an NPU, and a TPU.
[0024] According to another aspect of the present invention, there is provided a computer program product, including: a computer program, which when executed by a processor, implements the method described in any one of the above.
[0025] According to another aspect of the present invention, there is provided a computing device, including:
[0026] a processor; and
[0027] a memory storing a computer program, which when executed by the processor, causes the processor to execute the method described in any one of the above.
[0028] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium having stored thereon computer-readable instructions, which when executed by a processor, cause the processor to execute the method described in any one of the above.
[0029] According to the embodiments of the present invention, data synchronization of different computing subtasks in multiple acceleration cards is achieved through a unified communication control protocol (TCP). Compared with the prior art where different communication protocols are used between individual acceleration cards, the design solution of the present invention eliminates the data conversion time between different communication protocols, greatly reduces the communication time, and ensures the data synchronization efficiency on the premise of a unified rather than customized mechanism, thereby further improving the overall development and computing efficiency.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0032] Figure 1 A flowchart showing a method for data synchronization between multiple acceleration cards according to an exemplary embodiment.
[0033] Figure 2 A schematic diagram showing a method for data synchronization between multiple acceleration cards according to an exemplary embodiment.
[0034] Figure 3 A block diagram showing a computing device according to an exemplary embodiment. DETAILED DESCRIPTION
[0035] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. Like reference numerals refer to like or similar parts throughout the figures, and thus their repetitive description will be omitted.
[0036] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present invention. However, those skilled in the art will realize that the technical solutions of the present invention can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.
[0037] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0038] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all of the content and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0039] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below can be referred to as the second component without departing from the teachings of the concept of the present invention. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0040] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or reject.
[0041] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of exemplary embodiments, and the modules or processes in the accompanying drawings are not necessarily essential for implementing the present invention. Therefore, they cannot be used to limit the protection scope of the present invention.
[0042] With the development of domestic enterprise informatization and deep learning, more and more artificial intelligence technologies are being used by many enterprises and governments in their actual operations. The scale of AI system computing and the required computing power are continuously increasing, and the computing power of a single accelerator card is no longer sufficient to meet the needs of complex model training and inference. Therefore, using multiple accelerator cards for distributed computing to improve computing power has become one of the mainstream problem-solving solutions.
[0043] When using multiple accelerator cards for distributed computing, when the calculation result on a certain node requires the result data of the calculation of other accelerator cards, data transmission and synchronization between the accelerator cards are required. Different types of accelerator cards (NPU, GPU, TPU, etc.) provide different data transmission and synchronization solutions to achieve data synchronization, so different synchronization mechanisms need to be implemented when using different accelerator cards for distributed computing. When using different accelerator cards for distributed computing, different synchronization mechanisms need to be implemented to ensure synchronization between multiple accelerator cards, which not only has low communication efficiency but also cannot be a general solution for multi-accelerator card inference.
[0044] For this reason, the present invention proposes a method, a computer program product, and a computing device for data synchronization between multiple accelerator cards, which can adopt a unified solution to achieve data synchronization, minimize communication time as much as possible, and improve data synchronization efficiency. According to the embodiment, data synchronization of different computing subtasks in multiple accelerator cards is achieved through a unified communication control protocol (TCP). Compared with the prior art where different communication protocols are used between each accelerator card, the design solution of the present invention eliminates the data conversion time between different communication protocols, greatly reduces communication time, improves data synchronization efficiency, and thus further improves the overall computing efficiency.
[0045] Before describing the embodiments of the present invention, some terms or concepts related to the embodiments of the present invention are explained.
[0046] Graphics Processing Unit (GPU): Initially designed to accelerate image rendering tasks in computer graphics. Due to its highly parallel architecture, the GPU is also very suitable for accelerating scientific computing, deep learning, and other tasks that require a large amount of parallel computing. It is widely used in fields such as gaming, video processing, high-performance computing (HPC), and machine learning.
[0047] Neural Processing Unit (NPU): The NPU is designed specifically to accelerate machine learning, especially deep learning algorithms. It optimizes operations crucial to neural networks, such as matrix and vector operations. Commonly found in mobile devices, smart cameras, and other embedded systems, it provides efficient AI inference capabilities.
[0048] Tensor Processing Unit (TPU): The TPU is an application-specific integrated circuit (ASIC) developed by Google, designed specifically to accelerate machine learning workloads under the TensorFlow framework. They are particularly good at handling tensor-related mathematical operations, which are a core component of many modern machine learning models. TPUs are widely used in cloud services, such as the TPU resources provided by Google Cloud Platform, to support large-scale training and inference tasks.
[0049] The exemplary embodiments of the present invention will be described below with reference to the accompanying drawings.
[0050] Figure 1 A flowchart of a method for data synchronization between multiple acceleration cards according to an exemplary embodiment is shown.
[0051] See Figure 1 , which shows a method for data synchronization between multiple acceleration cards according to an exemplary embodiment. The method includes the following steps:
[0052] In S101, load the decomposed artificial intelligence computing task and allocate it to at least two acceleration cards. Each acceleration card executes the allocated computing subtask. The at least two acceleration cards include a first acceleration card.
[0053] According to some embodiments, a large AI computing task is divided into several smaller, parallelizable subtasks. Each subtask is assigned to at least two different acceleration cards to fully utilize hardware resources and improve computing efficiency. Among them, the acceleration cards include one or more of GPU, NPU, and TPU.
[0054] In S103, add data synchronization nodes to the computing subtasks of each acceleration card. The data synchronization nodes include a first data synchronization node for the first computing subtask of the first acceleration card.
[0055] According to some embodiments, insert data synchronization nodes into the computing subtasks that require data synchronization among the split computing subtasks. The setting of the data synchronization nodes should consider the dependencies in the computing process to ensure the correct data flow and logical order.
[0056] In S105, during the execution of the first computational subtask on the first acceleration card, the first data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the computational subtasks of the remaining acceleration cards.
[0057] According to some embodiments, when an acceleration card reaches the data synchronization node of a computational subtask, it calls the Transmission Control Protocol (TCP) to send a synchronization request. After waiting for the computational subtasks on the remaining acceleration cards to also reach the corresponding data synchronization nodes, data exchange and synchronization begin. Optionally, a central control machine (such as Redis, Kafka, MQTT, ZeroMQ, etc.) is used to achieve more efficient data exchange and synchronization management.
[0058] In S107, continue to execute the computational subtask until the calculation ends, where the first data synchronization node exchanges and synchronizes data based on the Transmission Control Protocol with the corresponding data synchronization nodes of the computational subtasks of the remaining acceleration cards, and the data synchronization node is executed on the central processing unit.
[0059] According to some embodiments, the first data synchronization node exchanges and synchronizes data based on the Transmission Control Protocol with the corresponding data synchronization nodes of the computational subtasks of the remaining acceleration cards, and the data synchronization node is executed on the central processing unit. After data synchronization is completed, each subtask continues to execute the remaining calculation part according to the predetermined process. The whole process is repeated until all computational subtasks have completed their work, marking the successful end of the entire AI computing task. Such a data synchronization method ensures that a common communication method is adopted between acceleration cards, reduces data conversion between different communication protocols, improves data synchronization efficiency, and thus helps to improve the overall AI computing execution efficiency.
[0060] Figure 2 A schematic diagram of a method for data synchronization between multiple acceleration cards according to an exemplary embodiment is shown.
[0061] See Figure 2 , the figure shows a method for data synchronization between multiple acceleration cards according to an exemplary embodiment. As can be seen from the figure, the AI computing task is decomposed into multiple computational subtasks 1 to n that can be executed in parallel and allocated to n acceleration cards. Each subtask is assigned to a different acceleration card (such as GPU, NPU, TPU, etc.), and each acceleration card executes one computational subtask to make full use of hardware resources and improve computing efficiency. Then, data synchronization nodes are inserted after the computational nodes that require data synchronization in each of the computational subtasks 1 to n. During the execution of the computational subtask, the data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the remaining computational subtasks based on the Transmission Control Protocol. After the communication ends, continue to execute the next computational subtask until the calculation ends.
[0062] According to some embodiments, when adding data synchronization nodes to each of the computing subtasks, the data synchronization nodes can be inserted after the computing nodes that need to perform data synchronization in each of the split computing subtasks. In a distributed computing environment, in order to ensure data consistency among various computing subtasks, data synchronization nodes are usually inserted after the computing nodes that need to perform data synchronization. This method can effectively coordinate parallel tasks on different acceleration cards (such as GPUs, NPUs, TPUs, etc.), ensure that they can execute in a predetermined logical order, and exchange necessary information at key points. Specifically, first, analyze the logical flow of each computing subtask to determine which computing nodes' data depends on the results of other subtasks, or in other words, where it is necessary to ensure the consistency of the states of all participating parties. For example, during the training process of a deep learning model, the parameter update of certain layers may depend on the output of the previous layer; in this case, a synchronization point needs to be set between these two layers. According to the above analysis results, data synchronization nodes are inserted after these specific computing nodes. The synchronization node can pause the execution of the current subtask at this point, wait for other related subtasks to reach the same position, and then perform necessary data exchange and state update together. The selection of the data synchronization node should take into account the dependency relationships in the computing process to ensure the correct data flow and logical order.
[0063] According to some embodiments, during the execution of the computing subtasks, the data synchronization node exchanges and synchronizes data with the corresponding data synchronization nodes of the remaining computing subtasks through the Transmission Control Protocol. Specifically, when the first acceleration card executes to the data synchronization node of the first computing subtask, it calls the Transmission Control Protocol to send a synchronization request. After waiting for the computing subtasks on the remaining acceleration cards to execute to the corresponding data synchronization nodes, data exchange and synchronization begin. In a distributed computing environment, when an acceleration card (such as a GPU, TPU, NPU, etc.) executes to the data synchronization node of its corresponding computing subtask, it will call the Transmission Control Protocol (TCP) to send a synchronization request. At this time, it will pause further computing operations and prepare to initiate a synchronization request, becoming the "requesting party". The requesting party establishes a connection with the remaining acceleration cards through TCP and sends a synchronization request containing its own identity information and current status. In practical applications, a specially designed message format is usually used to encapsulate this information, such as JSON or Protocol Buffers, for easy parsing and processing. After sending the synchronization request, the requesting party enters a waiting mode until it receives confirmation messages returned by all other acceleration cards. During this period, it can continue to execute tasks that do not affect the global result to improve resource utilization. For some high-performance computing scenarios, more complex synchronization algorithms, such as Paxos or Raft, may be adopted to enhance the fault tolerance and consistency guarantee of the system.
[0064] According to some embodiments, after waiting for the computing subtasks on the remaining acceleration cards to execute to the corresponding data synchronization nodes, data exchange and synchronization begin. When all acceleration cards have completed their respective computing subtasks and responded to the synchronization request, data exchange begins between them. Depending on the specific application scenario, it may involve various forms of data transfer such as parameter update, gradient accumulation, model weight sharing, etc. After synchronization is completed, all acceleration cards resume the normal computing process and continue to advance the subsequent computing subtasks. Such a design ensures that all acceleration cards participating in the computing can reach a consistent state at a specific time point, maintaining data consistency and the correctness of the computing logic.
[0065] According to some embodiments, the method can also adopt a central control machine to achieve data exchange and data synchronization. The central control machine uses Redis, Kafka, or MQTT to achieve data exchange and data synchronization. Among them, Redis, as a high-performance key-value storage system, not only supports in-memory data structure storage but also provides rich replication functions for data synchronization under the master-slave architecture. In the central control machine environment, one or more Redis instances can be configured as master nodes, responsible for receiving data update requests from different acceleration cards; at the same time, several slave nodes are set up to share the reading pressure and provide redundant backups. When the master node receives a new write operation, it immediately synchronizes the changes to all configured slave nodes to ensure data consistency within the entire cluster. In addition, Redis supports two modes: full synchronization and incremental synchronization. When starting for the first time, the slave node will perform full synchronization, that is, download a complete data snapshot from the master node; during normal operation, it mainly relies on incremental synchronization, only transmitting the part that has changed since the last synchronization. This design effectively reduces bandwidth occupancy and speeds up the synchronization speed.
[0066] For application scenarios that need to process large-scale real-time data streams, Kafka is also an ideal choice. As a distributed stream platform, Kafka can efficiently manage a large number of message queues, ensure message delivery in order, and has good horizontal scalability. In the central control machine solution, a Kafka cluster can be deployed to collect the calculation results or other types of data records generated by each acceleration card, and then forward them to the specified target location for subsequent analysis or persistent storage. Specifically, each acceleration card can publish the latest calculation status information to a Kafka topic when it reaches the data synchronization node. The central control machine subscribes to these topics, listens for the arrival of new messages, and triggers corresponding synchronization actions accordingly. Thanks to Kafka's high throughput characteristics and powerful fault tolerance mechanism, even in the case of network fluctuations or hardware failures, data can be ensured not to be lost and normal services can be quickly restored. In addition, with the help of the MirrorMaker tool, cross-data center data mirror replication can be easily achieved, providing additional guarantee for disaster recovery.
[0067] MQTT adopts the publish / subscribe model, allowing information to be exchanged between clients in a loosely coupled manner, which is very suitable for resource-constrained scenarios. In this case, if some acceleration cards are built on an embedded platform, a secure and reliable connection can be established with the central control machine through the MQTT protocol to upload local generated data in a timely manner or receive remote instructions. To improve synchronization efficiency, a dedicated pair of topics can be defined for each pair of acceleration cards, which are used to send requests and receive responses respectively. For example, when a certain acceleration card is ready to enter the synchronization phase, it can first publish a message to the "request" topic indicating that it is ready; while the central control machine at the other end will listen to this topic and immediately respond once it detects a new message, informing the other party whether it is suitable to start the synchronization process at present. Such an interaction method is both simple and direct, and at the same time facilitates the implementation of maintenance measures such as heartbeat detection.
[0068] According to some other embodiments, it is also possible to achieve data exchange and data synchronization without going through the central control machine. For example, the first data synchronization node and the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards adopt a message subscription mechanism for data exchange and data synchronization. The first computing subtask subscribes to the data messages of the corresponding data synchronization nodes of the retrieval subtasks of the remaining acceleration cards at the first data synchronization node; after the first computing subtask receives the data messages, it performs data exchange and data synchronization based on the data messages. For example, ZeroMQ is used. ZeroMQ is a high-performance asynchronous messaging library designed specifically for building fast and complex network applications. Compared with traditional TCP / IP programming interfaces, ZeroMQ provides a higher level of abstraction layer, simplifying the workload of developers. It supports multiple topologies, such as: point-to-point, publish / subscribe, request / reply, etc., and can meet the requirements of different types of tasks. ZeroMQ can be used to create a flexible message routing system to connect acceleration cards distributed in different geographical locations. By defining clear service interfaces and message formats, even in the face of complex and changing business logics, good maintainability and scalability can be maintained. For example, specific ports can be set to listen for heartbeat signals from each acceleration card and dynamically adjust its online status; or according to preset rules, specific types of data can be directed and forwarded to interested receivers, thus forming an efficiently operating data sharing network.
[0069] In summary, by introducing Redis, Kafka, MQTT or ZeroMQ, etc., the ability to achieve data exchange and synchronization can be significantly improved, constituting a powerful and stable infrastructure, providing solid technical support for collaboration in a distributed computing environment.
[0070] According to some embodiments, the design solution of the present invention can also be applied to the design of a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the method described in any one of the above, enabling efficient and reliable communication between each acceleration card, achieving data synchronization, and efficiently completing AI computing tasks.
[0071] According to some embodiments, the design solution of the present invention realizes data synchronization of different computing subtasks in multiple acceleration cards through a unified communication control protocol (TCP). Compared with the prior art where different communication protocols are used between each acceleration card, the design solution of the present invention eliminates the data conversion time between different communication protocols, greatly reduces the communication time, improves the data synchronization efficiency, and thus further improves the overall computing efficiency. Compared with the non-universal data synchronization means in different acceleration cards in traditional AI, the technical solution of the present invention uses a universal synchronization means to achieve synchronization between different acceleration cards, which can be quickly transplanted and improve the overall efficiency of AI computing.
[0072] Figure 3 A block diagram of a computing device according to an exemplary embodiment of the present invention is shown.
[0073] As Figure 3 shown, the computing device 30 includes a processor 12 and a memory 14. The computing device 30 may further include a bus 22, a network interface 16, and an I / O interface 18. The processor 12, the memory 14, the network interface 16, and the I / O interface 18 can communicate with each other through the bus 22.
[0074] The processor 12 may include one or more general-purpose CPUs (Central Processing Unit), microprocessors, or application-specific integrated circuits, etc., for executing relevant program instructions. According to some embodiments, the computing device 30 may further include a high-performance display adapter (GPU) 20 for accelerating the processor 12.
[0075] The memory 14 may include a machine system-readable medium in the form of a volatile memory, such as a random access memory (RAM), a read-only memory (ROM), and / or a cache memory. The memory 14 is used to store one or more programs containing instructions and data. The processor 12 can read the instructions stored in the memory 14 to execute the method according to the embodiments of the present invention described above.
[0076] The computing device 30 can also communicate with one or more networks through the network interface 16. The network interface 16 may be a wireless network interface.
[0077] The bus 22 may include an address bus, a data bus, a control bus, etc. The bus 22 provides a path for exchanging information between components.
[0078] It should be noted that, in the specific implementation process, the computing device 30 may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above devices may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figures.
[0079] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented. The computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), network storage devices, cloud storage devices, or any type of medium or device suitable for storing instructions and / or data.
[0080] The embodiments of the present invention also provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of any one of the methods described in the above method embodiments.
[0081] Those skilled in the art can clearly understand that the technical solutions of the present invention can be implemented by means of software and / or hardware. The "units" and "modules" in this specification refer to software and / or hardware that can independently complete or cooperate with other components to complete specific functions, and the hardware can be, for example, a field programmable gate array, an integrated circuit, etc.
[0082] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0083] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0084] In several embodiments provided by the present invention, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0085] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0086] In addition, each functional unit in various embodiments of the present invention can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0087] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention.
[0088] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0089] The above specifically shows and describes the exemplary embodiments of the present invention. It should be understood that the present invention is not limited to the detailed structures, setting methods, or implementation methods described here; on the contrary, the present invention is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A method for data synchronization between multiple acceleration cards, characterized in that, Including: Loading the decomposed artificial intelligence computing tasks and allocating them to at least two acceleration cards, each acceleration card executing the allocated computing subtasks, and the at least two acceleration cards including a first acceleration card; Adding data synchronization nodes to the computing subtasks of each acceleration card, and the data synchronization nodes including a first data synchronization node for the first computing subtask of the first acceleration card; During the execution of the first computing subtask of the first acceleration card, the first data synchronization node performs data exchange and data synchronization with the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards; Continuing to execute the computing subtasks until the calculation ends, wherein, the first data synchronization node performs data exchange and data synchronization with the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards based on the Transmission Control Protocol, and the data synchronization nodes are executed on the central processing unit.
2. The method according to claim 1, wherein Adding data synchronization nodes to the computing subtasks of each acceleration card includes: Inserting data synchronization nodes after the computing nodes that need to perform data synchronization in each computing subtask.
3. The method according to claim 2, wherein The first data synchronization node performing data exchange and data synchronization with the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards includes: When the first acceleration card executes to the data synchronization node of the first computing subtask, calling the Transmission Control Protocol to send a synchronization request.
4. The method according to claim 2, wherein The first data synchronization node performing data exchange and data synchronization with the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards includes: Waiting for the computing subtasks on the remaining acceleration cards to execute to the corresponding data synchronization nodes, and then starting data exchange and data synchronization.
5. The method according to claim 4, wherein Also including: Implementing data exchange and data synchronization by using a central control machine.
6. The method according to claim 5, wherein The central control machine implements data exchange and data synchronization by using redis, kafka, or MQTT.
7. The method according to claim 1, wherein The first data synchronization node and the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards perform data exchange and data synchronization by using a message subscription mechanism.
8. The method according to claim 7, wherein The first data synchronization node and the corresponding data synchronization nodes of the computing subtasks of the remaining acceleration cards performing data exchange and data synchronization by using a message subscription mechanism includes: The first computing subtask subscribes to the data messages of the corresponding data synchronization nodes of the retrieval subtasks of the remaining acceleration cards at the first data synchronization node; After the first computing subtask receives the data messages, performing data exchange and data synchronization based on the data messages.
9. A computer program product, characterized in that, Including: A computer program, which when executed by a processor implements the method according to any one of claims 1-8.
10. A computing device, characterized in that, Including: A processor; And A memory storing a computer program, which when executed by the processor enables the processor to execute the method according to any one of claims 1-8.