Topology adaptive optimization method and device based on NPU training features, and storage medium
By dynamically optimizing the device topology and cache strategy of NPU training characteristics, the problem of difficult optimization of communication topology between devices is solved, communication efficiency and training speed are improved, and it is suitable for large-scale distributed training scenarios.
Patent Information
- Application Number
- CN202510407569.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-04
AI Technical Summary
In the distributed training scenario of Ascend NPU, the topology of communication between devices is difficult to continuously optimize according to actual usage, resulting in increased communication link node length and access delay, affecting training efficiency.
By collecting NPU training features, dynamically optimize the link topology relationship and multi-level cache between devices, adjust the network interface and bandwidth according to the communication frequency and data transmission volume, establish direct or indirect connections, and configure shared memory and remote cache to achieve adaptive optimization.
Improves communication efficiency between devices, reduces remote memory access latency, and improves training speed and overall performance.
Smart Images

Figure CN120263657A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optimization of artificial intelligence devices, and particularly to a topology adaptive optimization method, device and storage medium based on NPU training features. Background Art
[0002] In the distributed training scenario of Ascend NPU, the efficiency of inter-device communication directly affects the training efficiency. HCCS provides high-performance interconnection. However, in practical applications, since the usage scenarios change over time, sometimes the communication frequency and data transmission volume between two certain devices are relatively large, but they will change after a period of time. It is difficult to continuously optimize the inter-device topology according to the actual usage situation, and the fixed inter-device topology prolongs the length of the communication link nodes between devices and increases the access latency. Summary of the Invention
[0003] The present invention provides a topology adaptive optimization method, device and storage medium based on NPU training features, aiming to solve at least one of the technical problems existing in the prior art.
[0004] The technical solution of the present invention is a topology adaptive optimization method based on NPU training features. The topology adaptive optimization method based on NPU training features is applied to a topology adaptive optimization device based on NPU training features. The topology adaptive optimization method based on NPU training features includes the following steps: S100. Collect all the hardware devices on the topology adaptive optimization device based on NPU training features and their connection methods; S200. Collect and analyze the communication characteristics between each device within a preset collection period; S300. Dynamically optimize the link topology relationship between each device based on the communication characteristics between each device; S400. Obtain high-frequency communication links and low-frequency communication links based on the communication characteristics between each device, and respectively configure multi-level caches; S500. Repeat steps S200 to S400 to dynamically optimize the link topology relationship and multi-level cache method between each device.
[0005] Further, step S100 includes: S110. Collect all the hardware devices connected to the main board and their configuration information. The hardware devices include at least one or a combination of the following: central processing unit (CPU), artificial intelligence processor array, conversion device, riser card, platform controller hub (PCH), trusted platform module (TPM), USB device, SATA device, audio device, network device, complex programmable logic device (CPLD), VGA interface, RJ45 interface, fiber module interface, and non-volatile memory host controller interface (NVMe). S120. Collect and test the physical connection methods of all the hardware connected to the main board. S130. Distinguish whether the connection between every two hardware devices is a full-connection method or a partial-connection method.
[0006] Further, in step S130, The full-connection method is a fully interconnected network topology, which is set between two NPU chips inside or between artificial intelligence processor arrays that need to frequently synchronize data, and between other two hardware devices that need to frequently synchronize data. The partial-connection method is used to reduce unnecessary communication overhead and network complexity, and is set between two hardware devices with infrequent communication between some devices.
[0007] Further, step S200 includes: S210. Based on the recordTransfer method, within a preset collection period, collect and record the communication frequency commFrequency of the communication between the source device and the target device. The source device and the target device are both one or a combination of two of the artificial intelligence processor array or other hardware devices. S220. Record the amount of data transferred each time from the source device to the target device. After a collection period ends, calculate the total data transfer amount dataTransfer within the current collection period.
[0008] Further, step S300 includes: S310. Check and ensure that there is a direct communication link between every two NPU chips in the artificial intelligence processor array. If not, re-establish the direct communication link between every two NPU chips. S320. According to the communication frequency commFrequency and the total data transfer amount dataTransfer, adjust the network interfaces and bandwidths of the source device and the target device, and adjust the communication protocol to support multi-to-multi communication. S330. When the communication frequency commFrequency from the source device to the target device exceeds a preset threshold, if the connection from the source device to the target device is indirect or there are multiple communications, establish a direct connection from the source device to the target device, and print or configure the communication link according to the optimized topology information; S340. Allocate more bandwidth between the source device and the target device for high-frequency communication, and the remaining devices are connected through multi-hop or forwarded through intermediate devices.
[0009] Further, step S400 includes: S410. Configure a shared memory in the high-frequency communication link, and the shared memory includes a local cache or a high-speed cache; S420. Configure a remote cache in the low-frequency communication link.
[0010] Further, the present invention also proposes a topology adaptive optimization device based on NPU training features for performing the topology adaptive optimization method based on NPU training features. The topology adaptive optimization device based on NPU training features includes: A main board, on which at least one central processing unit CPU is provided; An artificial intelligence processor array, which is arranged on the main board; A conversion device, which is connected to the main board and the central processing unit CPU through a PCIE interface, and the conversion device is electrically connected to the artificial intelligence processor array through a PCIE bus.
[0011] Further, at least one group of artificial intelligence processor arrays is provided. Each group of artificial intelligence processor arrays NPU includes at least two NPU chips, and each NPU chip is connected based on high-speed interconnect technology (HCCS); At least one group of conversion devices is provided. Each group of conversion devices includes two conversion boards. Each conversion board is connected to the main board through a PCIE interface respectively, and each adapter board is directly connected to two NPU chips respectively; It also includes an expansion card (Riser Card), and the expansion card (Riser Card) is connected to the main board through a PCIE slot, and the expansion card (Riser Card) is connected to the adapter board through a PCIE bus.
[0012] Further, the main board also includes: A platform controller hub (PCH), the platform controller hub (PCH) is connected to the motherboard through a PCIE slot, and there are also a trusted platform module (TPM), a USB device, a SATA device, an audio device, a network device, a complex programmable logic device (CPLD), a VGA interface, an RJ45 interface, and a fiber module interface that are directly or indirectly connected to the platform controller hub (PCH), and the platform controller hub (PCH) is connected to a central processing unit CPU; A non-volatile memory host controller interface (NVMe) for connecting to a solid state drive SSD, the non-volatile memory host controller interface (NVMe) is connected to a central processing unit CPU.
[0013] Furthermore, the present invention also proposes a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the topology adaptive optimization method based on NPU training features described above is implemented.
[0014] The beneficial effects of the present invention are: The topology adaptive optimization method, device, and storage medium based on NPU training features improve the communication efficiency between devices under different training tasks through adaptive optimization on the interconnect topology of HCCS (Huawei Cache Coherent System); through a dynamically adjustable topology configuration mechanism, it dynamically adapts to the best topology structure according to the communication characteristics of the training task, either full-connection communication or sparse communication; at the same time, combined with multi-level cache acceleration, it reduces the latency of remote memory access and improves the training speed. Description of the Drawings
[0015] Figure 1 It is the overall flowchart of the topology adaptive optimization method based on NPU training features.
[0016] Figure 2 It is the structural schematic diagram of an embodiment of the topology adaptive optimization device based on NPU training features. Detailed Embodiments
[0017] The following will clearly and completely describe the concept, specific structure, and technical effects generated by the present invention in combination with embodiments and drawings to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0018] It should be noted that, unless otherwise specified, when a feature is referred to as "fixed" or "connected" to another feature, it can be directly fixed or connected to the other feature, or indirectly fixed or connected to the other feature. In addition, the descriptions such as up, down, left, right, top, bottom, etc. used in the present invention are only relative to the mutual positional relationship of the components of the present invention in the drawings.
[0019] In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this technology belongs. The terms used in the description of this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any combination of one or more of the related listed items.
[0020] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish elements of the same type from each other. For example, without departing from the scope of this disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.
[0021] Referring to Figure 1 and Figure 2 , in some embodiments, the technical solution of the present invention is a topology adaptive optimization method based on NPU training features. The topology adaptive optimization method based on NPU training features is applied to a topology adaptive optimization device based on NPU training features. Referring to Figure 1 , the topology adaptive optimization method based on NPU training features includes the following steps: S100. Collect all the hardware devices and their connection methods on the topology adaptive optimization device based on NPU training features; S200. Collect and analyze the communication features between each device within a preset collection period; S300. Dynamically optimize the link topology relationship between each device based on the communication features between each device; S400. Obtain high-frequency communication links and low-frequency communication links based on the communication features between each device, and configure multi-level caches respectively; S500. Repeat steps S200 to S400 to dynamically optimize the link topology relationship and multi-level cache mode between each device.
[0022] The beneficial effects of the present invention are: The described topology adaptive optimization method, device, and storage medium based on NPU training features improve the communication efficiency between devices under different training tasks through adaptive optimization on the interconnect topology of HCCS (Huawei Cache Coherent System); through a dynamically adjustable topology configuration mechanism, it dynamically adapts to the best topology structure according to the communication characteristics of the training task for full-connection communication or sparse communication; at the same time, combined with multi-level cache acceleration, it reduces the latency of remote memory access and improves the training speed.
[0023] Among them, NPU (Neural Processing Unit) is a processor specifically designed for processing neural network algorithms and artificial intelligence tasks, with high parallelism and efficient computing capabilities.
[0024] Furthermore, referring to Figure 1 , step S100 includes: S110. Collect all the hardware devices connected to the motherboard and their configuration information. The hardware devices include at least one or more combinations of a central processing unit (CPU), an artificial intelligence processor array, a conversion device, a riser card, a platform controller hub (PCH), a trusted platform module (TPM), a USB device, a SATA device, an audio device, a network device, a complex programmable logic device (CPLD), a VGA interface, an RJ45 interface, a fiber module interface, and a non-volatile memory host controller interface (NVMe). S120. Collect and test the physical connection methods of all the hardware connected to the motherboard. S130. Distinguish whether the connection between every two hardware devices is a full-connection method or a partial-connection method.
[0025] Specifically, the source device and the target device are a combination of two of the hardware devices, representing the communication process of the current communication link from one device (source device) to another device (target device) among the hardware devices.
[0026] Furthermore, referring to Figure 1 , in step S130, The full-connection method is a fully interconnected network topology, which is set between two NPU chips inside or between the artificial intelligence processor arrays that need to frequently synchronize data, and between two other hardware devices that need to frequently synchronize data. The partial-connection method is to reduce unnecessary communication overhead and network complexity, and is set between two hardware devices with infrequent communication between some devices.
[0027] Specifically, the full connection and local connection have the following characteristics: Full connection: Frequent data synchronization is required between all NPU devices, forming a fully interconnected network topology. Advantages: Fast communication can be achieved between any devices, and the communication latency is low. Disadvantages: High requirements for network bandwidth and device performance, and relatively high costs.
[0028] Local connection: Communication is relatively frequent between some devices, while communication between other devices is less frequent or even non-existent. Advantages: Reducing unnecessary communication overhead and network complexity. Disadvantages: Devices with frequent communication may become performance bottlenecks.
[0029] Furthermore, referring to Figure 1 , step S200 includes: S210. Based on the recordTransfer method, within a preset acquisition period, collect and record the communication frequency commFrequency of the communication between the source device and the target device. Both the source device and the target device are one or a combination of two of the artificial intelligence processor arrays or other hardware devices; S220. Record the amount of data transferred each time from the source device to the target device. After an acquisition period ends, calculate the total data transfer amount dataTransfer within the current acquisition period.
[0030] Specifically, the communication frequency commFrequency refers to the number of times of data communication between the source device and the target device during the training process. For example, in distributed training, one device may need to send gradients or intermediate results to another device multiple times. Its function is to determine which devices have relatively frequent communication by tracking the communication frequency, so as to provide a basis for subsequent topology optimization and caching strategies.
[0031] In step S210, the code implementation is as follows: class CommunicationProfiler { public: void recordTransfer(const std::string& sourceNPU, const std::string& targetNPU, size_t dataSize) { std::pair<std::string, std::string> key = {sourceNPU,targetNPU}; dataTransfer[key] += dataSize; / / Data transfer volume commFrequency[key]++; / / Communication frequency increment } std::unordered_map<std::pair<std::string, std::string>, size_t,pair_hash> getDataTransfer() const { return dataTransfer; } std::unordered_map<std::pair<std::string, std::string>, size_t,pair_hash> getCommFrequency() const { return commFrequency; } private: struct pair_hash { template <class T1, class T2> std::size_t operator()(const std::pair<T1, T2>& pair) const { return std::hash <t1>()(pair.first) ^ std::hash <t2>()(pair.second); } }; std::unordered_map<std::pair<std::string, std::string>, size_t, pair_hash> dataTransfer; / / Data transfer volume statistics std::unordered_map<std::pair<std::string, std::string>, size_t, pair_hash> commFrequency; / / Communication frequency statistics }; Among them, the recordTransfer method records the source device, target device, and data size of the transmission and increases the communication frequency count whenever communication occurs between two NPU devices.
[0032] The data is stored in two hash tables: dataTransfer: Records the total data transfer volume from the source device to the target device, commFrequency: Records the number of communication times from the source device to the target device.
[0033] Furthermore, referring to Figure 1 , step S300 includes: S310. Check and ensure that there is a direct communication link between every two NPU chips in the artificial intelligence processor array. If not, re - establish the direct communication link between every two NPU chips; S320. Adjust the network interfaces and bandwidths of the source device and the target device according to the communication frequency commFrequency and the total data transfer volume dataTransfer, and adjust the communication protocol to support many - to - many communication; S330. When the communication frequency commFrequency from the source device to the target device exceeds the preset threshold, if the connection from the source device to the target device is indirect or there are multiple communications, establish a direct connection from the source device to the target device, and print or configure the communication link according to the optimized topology information; S340. Allocate more bandwidth between the source device and the target device with high - frequency communication, and the remaining devices are connected through multi - hop or forwarded through intermediate devices.
[0034] Specifically, (1) Full - connection communication: Establish direct connections between all devices Implementation method: Ensure that there is a direct communication link between every two NPU devices, which can minimize communication latency and improve data transmission efficiency. At the hardware level, it may be necessary to increase network interfaces and bandwidth; at the software level, the communication protocol needs to be adjusted to support many-to-many communication.
[0035] The code implementation is as follows: class TopologyOptimizer { public: TopologyOptimizer(const CommunicationProfiler& profiler) :profiler(profiler) {} void optimizeTopology() { auto commFrequency = profiler.getCommFrequency(); for (const auto& entry : commFrequency) { const auto& key = entry.first; size_t frequency = entry.second; topology[key] = (frequency > 10)? "direct" : "indirect"; } } private: const CommunicationProfiler& profiler; std::unordered_map<std::pair<std::string, std::string>, std::string, CommunicationProfiler::pair_hash> topology; }; (2)Judging the connection method according to the communication frequency: If the communication frequency exceeds a certain threshold (for example, 10 times), establish a direct connection; otherwise, use an indirect connection (multi-hop communication may be required).
[0036] Application topology: void applyTopology() const { for (const auto& entry : topology) { const auto& key = entry.first; const std::string& connection = entry.second; std::cout << (connection == "direct" ? "Direct link" : "Indirect link") << " established: " << key.first << " -> " << key.second<< std::endl; } } Print or configure communication links according to the optimized topology information.
[0037] (3)Local connection communication: Allocate more bandwidth for high-frequency communication devices, and the rest of the devices are connected through multi-hop The code implementation is as follows: void TopologyOptimizer::optimizeTopology() { auto commFrequency = profiler.getCommFrequency(); for (const auto& entry : commFrequency) { const auto& key = entry.first; size_t frequency = entry.second; / / Assume the high-frequency communication threshold is 10 if (frequency > 10) { allocateBandwidth(key, 1000); / / Allocate more bandwidth for frequent communication } else { allocateBandwidth(key, 100); / / Allocate less bandwidth for low-frequency communication } topology[key] = (frequency > 10)? "direct" : "indirect"; } } void allocateBandwidth(const std::pair<std::string, std::string>& key, size_t bandwidth) { / / Simulate bandwidth allocation logic std::cout << "Allocated bandwidth " << bandwidth << " for " << key.first << "->" << key.second << std::endl; } For high - frequency communication devices: Allocate more bandwidth resources for device pairs in high - frequency communication; For low - frequency communication devices: Save bandwidth resources through multi - hop connections (e.g., intermediate device forwarding).
[0038] Furthermore, referring to Figure 1 , step S400 includes: S410. Configure shared memory in the high - frequency communication link, and the shared memory includes local cache or cache memory; S420. Configure remote cache in the low - frequency communication link.
[0039] Specifically, add multiple - level caches between high - frequency communication devices, including a first - level cache (shared memory) and a second - level cache (remote cache); among them, the location of the first - level cache (shared memory) is the shared memory (such as local Cache or cache memory) between the source device and the target device. The purpose of the first - level cache (shared memory) is to store the most frequently communicated data to reduce direct access to the target device and improve data access speed. The location of the second - level cache (remote cache) is the cache of the remote device (such as a remote server or shared storage). The purpose of the second - level cache (remote cache) is to store low - frequency communication data as a supplement to the local cache. The first - level cache stores high - frequency communication data in the local shared memory with fast access speed, and the second - level cache stores low - frequency communication data in the remote cache and accesses it through the network. The code implementation is as follows: class CacheManager { public: void cacheData(const std::string& source, const std::string& target, const std::string& data) { localCache[{source, target}].push_back(data); / / Primary cache / / Secondary cache of relay device (assuming source and target are relay devices) addToRemoteCache(source, target, data); } std::string accessData(const std::string& source, const std::string& target) { auto key = std::make_pair(source, target); / / Try to read from the primary cache if (!localCache[key].empty()) { std::string data = localCache[key].back(); localCache[key].pop_back(); return data; } else { / / Try to read from the remote cache if (remoteCache.find(key) != remoteCache.end()) { return remoteCache[key]; } else { return ""; / / Cache miss } } } private: / / Primary cache: local shared memory std::unordered_map<std::pair<std::string, std::string>, std::vector <std::string>, pair_hash> localCache; / / Secondary cache: remote storage std::unordered_map<std::pair<std::string, std::string>, std::string, pair_hash> remoteCache; pair_hash pair_hash; void addToRemoteCache(const std::pair<std::string, std::string>& key, const std::string& data) { remoteCache[key] = data; / / More complex logic may be needed, such as storing only the latest data } }; The topology adaptive optimization method based on NPU training features also includes an integrated optimization process. First, use the CommunicationProfiler class to collect and analyze communication data (transmission volume and frequency) between devices, and store the data for subsequent optimization.
[0040] Dynamic topology optimization: Based on communication frequency and data transmission volume, establish direct connections for high-frequency communication devices and indirect connections for low-frequency communication devices. Allocate bandwidth resources to ensure that high-frequency communication devices obtain sufficient bandwidth.
[0041] Multi-level cache optimization: Add caches between high-frequency communication devices to reduce direct access to remote devices. Use a primary cache and secondary cache strategy to balance performance and cost. The code implementation is as follows: #include <iostream> #include <unordered_map> #include <string> #include <vector> #include <utility> #include <functional> / / Pair hash for unordered_map namespace std { template<> struct hash<pair<string, string>> { size_t operator()(const pair<string, string>& p) const { return hash <string>(()(p.first) ^ hash <string>()(p.second); } }; } / / Data collection: Communication feature analysis class CommunicationProfiler { public: void recordTransfer(const string& source, const string& target, size_t size) { dataTransfer[{source, target}] += size; commFrequency[{source, target}]++; } unordered_map<pair<string, string>, size_t> getDataTransfer() const { return dataTransfer; } unordered_map<pair<string, string>, size_t> getCommFrequency() const { return commFrequency; } private: unordered_map<pair<string, string>, size_t> dataTransfer; unordered_map<pair<string, string>, size_t> commFrequency; }; / / Topology optimization class TopologyOptimizer { public: TopologyOptimizer(const CommunicationProfiler& profiler) : profiler(profiler) {} void optimizeTopology() { auto freq = profiler.getCommFrequency(); for (const auto& entry : freq) { const auto& key = entry.first; size_t f = entry.second; if (f > 10) { topology[key] = "direct"; allocateBandwidth(key, 1000); } else { topology[key] = "indirect"; allocateBandwidth(key, 100); } } } void applyTopology() const { for (const auto& entry : topology) { cout << entry.second << " link: " << entry.first.first <<"->" << entry.first.second << endl; } } private: const CommunicationProfiler& profiler; unordered_map<pair<string, string>, string> topology; void allocateBandwidth(const pair<string, string>& key, size_tbandwidth) { cout << "Allocated bandwidth " << bandwidth << " for " <<key.first << "->" << key.second << endl; } }; / / Multi - level cache management class CacheManager { public: void cacheData(const string& src, const string& tgt, const string& data) { localCache[{src, tgt}].push_back(data); remoteCache[{src, tgt}] = data; / / Simplified: store latest data } string accessData(const string& src, const string& tgt) { const auto& key = make_pair(src, tgt); if (!localCache[key].empty()) { string data = localCache[key].back(); localCache[key].pop_back(); return data; } else if (remoteCache.find(key) != remoteCache.end()) { return remoteCache[key]; } return ""; } private: unordered_map<pair<string, string>, vector <string>> localCache; / / Level 1 cache unordered_map <pair<string, string> , string> remoteCache; / / Level2 cache }; / / Overall training process void trainingPipeline() { / / Communication characteristics analysis CommunicationProfiler profiler; profiler.recordTransfer("NPU1", "NPU2", 1024); profiler.recordTransfer("NPU1", "NPU3", 512); profiler.recordTransfer("NPU2", "NPU1", 2048); / / Dynamic topology optimization TopologyOptimizer optimizer(profiler); optimizer.optimizeTopology(); optimizer.applyTopology(); / / Cache optimization CacheManager cache; cache.cacheData("NPU1", "NPU2", "Batch Data 1"); cache.cacheData("NPU1", "NPU2", "Batch Data 2"); string data = cache.accessData("NPU1", "NPU2"); if (!data.empty()) { cout << "Training with cached data: " << data << endl; } } int main() { trainingPipeline(); return 0; } The results of verifying the static topology against the dynamic topology optimization through experiments are as follows: Table 1:
[0042] Referring to Table 1, the dynamic topology optimization of the topology adaptive optimization method based on NPU training features is superior in terms of communication latency, memory access latency, and total training time compared to the static topology, with less time consumption.
[0043] Furthermore, referring to Figure 2 , the present invention also proposes a topology adaptive optimization device based on NPU training features for performing the topology adaptive optimization method based on NPU training features. The topology adaptive optimization device based on NPU training features includes: A main board, on which at least one central processing unit (CPU) is provided; An artificial intelligence processor array, which is arranged on the main board; A conversion device, which is connected to the main board and the central processing unit (CPU) through a PCIE interface, and is electrically connected to the artificial intelligence processor array through a PCIE bus.
[0044] Furthermore, referring to Figure 2 , at least one group of artificial intelligence processor arrays is provided. Each group of artificial intelligence processor arrays (NPUs) includes at least two NPU chips, and each NPU chip is connected based on high-speed interconnect technology (HCCS); At least one group of conversion devices is provided. Each group of conversion devices includes two conversion boards. Each conversion board is respectively connected to the main board through a PCIE interface, and each adapter board is directly connected to two NPU chips; It further includes an expansion card (Riser Card), which is connected to the main board through a PCIE slot, and the expansion card (Riser Card) is connected to the adapter board through a PCIE bus.
[0045] Furthermore, referring to Figure 2 , the main board further includes: Platform Controller Hub (PCH), the Platform Controller Hub (PCH) is connected to the motherboard through a PCIE slot, and there are also a Trusted Platform Module (TPM), USB devices, SATA devices, audio devices, network devices, Complex Programmable Logic Device (CPLD), VGA interface, RJ45 interface, and fiber module interface that are directly or indirectly connected to the Platform Controller Hub (PCH), and the Platform Controller Hub (PCH) is connected to the Central Processing Unit CPU; Non-Volatile Memory Host Controller Interface (NVMe), used to connect to a Solid State Drive (SSD), and the Non-Volatile Memory Host Controller Interface (NVMe) is connected to the Central Processing Unit CPU.
[0046] In a specific embodiment, referring to Figure 2 , between the top layer D1~D8 is the NPU chip, and the connection between NPUs is the connection method of High-Speed Coherent Interconnect Technology (HCCS) (indicated by blue arrows); there are 4 switches (SW1~SW4) set below, and one NPU card is connected to each switch. The switch is connected to the Riser Card using PCIE, so as to connect the Riser Card to the motherboard; finally, other structures and functional components of the server are set, including two Central Processing Units (CPUs), Non-Volatile Memory Host Controller Interface (NVMe), Platform Controller Hub (PCH), Trusted Platform Module (TPM), USB devices, SATA devices, audio devices, network devices, Complex Programmable Logic Device (CPLD), VGA interface, RJ45 interface, and fiber module interface, etc.
[0047] The present invention dynamically optimizes the topology based on task communication characteristics, makes full use of High-Speed Coherent Interconnect Technology (HCCS), and improves performance. It greatly reduces communication and memory access latency and improves training efficiency. It is applicable to training scenarios of large models and large-scale distributed training, such as supercomputer centers.
[0048] Furthermore, referring to Figure 1 , the present invention also proposes a computer-readable storage medium, on which program instructions are stored, and when the program instructions are executed by a processor, the topology adaptive optimization method based on NPU training characteristics described above is implemented.
[0049] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-mentioned embodiments. As long as the same means are used to achieve the technical effects of the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure. They shall all fall within the scope of protection of the present invention. Within the scope of protection of the present invention, various different modifications and variations can be made to its technical solutions and / or implementation manners.< / string> < / string> < / string> < / functional> < / utility> < / vector> < / string> < / iostream> < / std::string>
Claims
1. A topology adaptive optimization method based on NPU training features, where the topology adaptive optimization method based on NPU training features is applied to a topology adaptive optimization device based on NPU training features, and is characterized in that The topology adaptive optimization method based on NPU training features includes the following steps: S100. Collect all the hardware devices on the topology adaptive optimization device based on NPU training features and their connection methods; S200. Collect and analyze the communication features between each device within a preset collection period; S300. Dynamically optimize the link topology relationship between each device based on the communication features between each device; S400. Obtain high-frequency communication links and low-frequency communication links based on the communication features between each device, and configure multi-level caches respectively; S500. Repeat steps S200 to S400 to dynamically optimize the link topology relationship and multi-level cache method between each device.
2. The topology adaptive optimization method based on NPU training features according to claim 1, wherein Step S100 includes: S110. Collect all the hardware devices connected to the main board and their configuration information. The hardware devices include at least one or a combination of a central processing unit, an artificial intelligence processor array, a conversion device, an expansion card, a platform controller hub, a trusted platform module, a USB device, a SATA device, an audio device, a network device, a complex programmable logic device, a VGA interface, an RJ45 interface, a fiber module interface, and a non-volatile memory host controller interface; S120. Collect and test the physical connection methods of all the hardware connected to the main board; S130. Distinguish whether the connection between every two pieces of hardware is a full connection method or a partial connection method.
3. The topology adaptive optimization method based on NPU training features according to claim 2, characterized in that In step S130, The full connection method is a fully interconnected network topology, which is set inside the artificial intelligence processor array that needs to frequently synchronize data or between two NPU chips among them, and between two other hardware devices that need to frequently synchronize data; The partial connection method is to reduce unnecessary communication overhead and network complexity, and is set between two hardware devices with infrequent communication between some devices.
4. The topology adaptive optimization method based on NPU training features according to claim 1, characterized in that Step S200 includes: S210. Based on the recordTransfer method, within a preset collection period, collect and record the analysis of the communication frequency between the source device and the target device. Both the source device and the target device are one or a combination of two of the artificial intelligence processor array or other hardware devices; S220. Record the amount of data transferred each time from the source device to the target device. When a collection period ends, calculate the total data transfer amount within the current collection period.
5. The topology adaptive optimization method based on NPU training features according to claim 1, characterized in that Step S300 includes: S310. Check and ensure that there is a direct communication link between every two NPU chips in the artificial intelligence processor array. If not, re-establish the direct communication link between every two NPU chips; S320. Adjust the network interfaces and bandwidths of the source device and the target device according to the communication frequency and the total data transfer amount, and adjust the communication protocol to support multi-to-multi communication; S330. When the communication frequency from the source device to the target device exceeds a preset threshold, if the connection from the source device to the target device is an indirect connection or multiple communications, establish a direct connection from the source device to the target device, and print or configure the communication link according to the optimized topology information; S340. Allocate more bandwidth between the source device and the target device for high-frequency communication, and the remaining devices are connected through multi-hop or forwarded through intermediate devices.
6. The topology adaptive optimization method based on NPU training features according to claim 1, characterized in that Step S400 includes: S410. Configure a shared memory on the high-frequency communication link, and the shared memory includes a local cache or a cache; S420. Configure a remote cache on the low-frequency communication link.
7. A topology adaptive optimization device based on NPU training features, which is used to execute the topology adaptive optimization method based on NPU training features as described in any one of claims 1 to 6, and is characterized in that, The topology adaptive optimization device based on NPU training features includes: A main board, on which at least one central processing unit CPU is provided; An artificial intelligence processor array, which is arranged on the main board; A conversion device, which is connected to the main board and the central processing unit CPU through a PCIE interface, and the conversion device is electrically connected to the artificial intelligence processor array through a PCIE bus.
8. The topology adaptive optimization device based on NPU training features according to claim 7, wherein At least one group of the artificial intelligence processor arrays is provided, and each group of artificial intelligence processor arrays NPU includes at least two NPU chips, and each NPU chip is connected based on a high-speed interconnection technology; At least one group of the conversion devices is provided, and each group of conversion devices includes two conversion boards, each conversion board is connected to the main board through a PCIE interface respectively, and each adapter board is directly connected to two NPU chips respectively; It further includes an expansion card, the expansion card is connected to the main board through a PCIE slot, and the expansion card is connected to the adapter board through a PCIE bus.
9. The topology adaptive optimization device based on NPU training features according to claim 7, wherein The main board further includes: A platform controller hub, which is connected to the main board through a PCIE slot, and a trusted platform module, a USB device, a SATA device, an audio device, a network device, a complex programmable logic device, a VGA interface, an RJ45 interface, and a fiber module interface that are directly or indirectly connected to the platform controller hub are also provided, and the platform controller hub is connected to the central processor; A non-volatile memory host controller interface for connecting a solid-state drive, and the non-volatile memory host controller interface is connected to the central processor.
10. A computer-readable storage medium storing program instructions thereon, characterized in that, When the program instructions are executed by the processor, the method described in any one of claims 1 to 6 is implemented.