Method, device, equipment and storage medium for planning traffic transmission paths

By building a communication traffic loop, the task equipment of the same LA group is connected first, which solves the problem of low communication efficiency of AI large-model task equipment across LA groups, realizes efficient traffic transmission path planning, and improves system performance.

CN118612131BActive Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311870093.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-08-29
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

In the prior art, AI large-model task equipment frequently crosses LA groups when communicating across LA groups, resulting in reduced communication efficiency and failure to effectively plan traffic transmission paths.

Method used

By determining the LA group to which the task equipment belongs and building a communication traffic loop, the devices of the same LA group are connected in series first to form a path closed loop to avoid unnecessary traffic transmission across LA groups.

Benefits of technology

It improves the communication efficiency between task equipment, reduces traffic across LA groups, optimizes communication paths, and improves the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118612131B_ABST
    Figure CN118612131B_ABST
Patent Text Reader

Abstract

A method, device, equipment and storage medium for planning a traffic transmission path, belonging to the field of communication technology. The technical solution provided by this application can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving. The method includes: determining the access layer switch LA group to which each of the N task devices belongs, and the N task devices jointly perform the same artificial intelligence AI large model task; according to the LA group to which the N task devices belong, a communication traffic ring composed of N task devices is obtained, and each task device belonging to the same LA group is connected in series. The above method, according to the LA group to which the N task devices belong, reasonably plans the relative positions of the N task devices during the construction of the communication traffic ring. Since the devices of the same LA group are preferentially connected in series, unnecessary cross-LA group traffic transmission is avoided during traffic transmission, thereby improving the communication efficiency between task devices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a method, apparatus, device and storage medium for planning a traffic transmission path. Background Art

[0002] Traffic transmission path planning refers to the reasonable planning of traffic transmission paths to ensure the efficiency and reliability of data transmission between task devices.

[0003] In related technologies, topology affinity is often used for traffic planning. Specifically, when assigning devices for a large AI (Artificial Intelligence) model task, it is best to assign the task to devices in the same LA (Access Layer) group. If the multiple devices involved in the large AI model task can be assigned to the same LA group, the traffic transmission path will be optimal.

[0004] The above method, due to the limited number of LA ports, can only connect a limited number of task devices. When the number of task devices exceeds the maximum number that can be connected to the LA, traffic will inevitably cross LA groups. In other words, multiple task devices will be assigned to different LA groups. However, the related art does not consider the reasonable planning of the traffic transmission paths formed by multiple task devices. As a result, when task devices communicate, traffic frequently crosses LA groups, thereby reducing communication efficiency between task devices. Summary of the Invention

[0005] The embodiments of the present application provide a method, apparatus, device, and storage medium for planning a traffic transmission path. The technical solutions provided by the embodiments of the present application are as follows:

[0006] According to one aspect of an embodiment of the present application, a method for planning a traffic transmission path is provided, the method comprising:

[0007] Determine the LA groups to which N task devices belong, where each LA group includes at least one task device and a LA connected to the at least one task device, the N task devices are divided into at least two LA groups, and the N task devices jointly execute the same AI large model task, where N is an integer greater than 1;

[0008] According to the LA groups to which the N task devices respectively belong, a communication traffic ring composed of the N task devices is obtained, wherein the communication traffic ring refers to a path closed loop formed by connecting the N task devices in sequence, which is used to describe the transmission path of the communication traffic between the N task devices, and each of the task devices belonging to the same LA group is connected in series.

[0009] According to one aspect of an embodiment of the present application, a device for planning a traffic transmission path is provided, the device comprising:

[0010] a determination module, configured to determine the LA groups to which N task devices each belong, wherein each LA group includes at least one task device and a LA connected to the at least one task device, the N task devices are divided into at least two LA groups, and the N task devices jointly execute the same artificial intelligence (AI) large model task, where N is an integer greater than 1;

[0011] An obtaining module is used to obtain a communication traffic ring composed of the N task devices according to the LA groups to which the N task devices respectively belong, wherein the communication traffic ring refers to a path closed loop formed by connecting the N task devices in sequence, which is used to describe the transmission path of the communication traffic between the N task devices, and each of the task devices belonging to the same LA group is connected in series.

[0012] According to one aspect of an embodiment of the present application, a computer device is provided, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned traffic transmission path planning method.

[0013] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned method for planning a traffic transmission path.

[0014] According to one aspect of an embodiment of the present application, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium, and a processor reads and executes the computer program from the computer-readable storage medium to implement the above-mentioned traffic transmission path planning method.

[0015] The technical solutions provided by the embodiments of the present application include at least the following beneficial effects:

[0016] During the construction of the communication traffic ring, the relative positions of the N task devices are rationally planned based on the LA groups to which they each belong. Task devices belonging to the same LA group are prioritized for serial connection, and the N task devices are then combined to form a closed path loop, serving as the final communication traffic ring. This method, by prioritizing serial connection of devices within the same LA group, avoids unnecessary cross-LA group traffic transmission, reduces cross-LA group traffic, and improves communication efficiency between task devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;

[0018] Figure 2 This is a schematic diagram of traffic transmission based on a ring topology provided by an embodiment of the present application;

[0019] Figure 3 This is a flow chart of a method for planning a traffic transmission path provided by an embodiment of the present application;

[0020] Figure 4 This is a schematic diagram of a communication traffic ring including four task devices provided by an embodiment of the present application;

[0021] Figure 5 is a schematic diagram of a communication traffic ring including four task devices provided by another embodiment of the present application;

[0022] Figure 6 This is a schematic diagram of a task device including multiple network cards provided by one embodiment of the present application;

[0023] Figure 7 This is a flow chart of a method for planning a traffic transmission path provided by another embodiment of the present application;

[0024] Figure 8 This is a schematic diagram of a method for obtaining a hash value in a cloud environment provided by an embodiment of the present application;

[0025] Figure 9 1 is a schematic diagram of experimental results comparing flow ratios across LA groups provided by an embodiment of the present application;

[0026] Figure 10 This is a schematic diagram of the AllReduce performance comparison experiment results provided by an embodiment of the present application;

[0027] Figure 11 This is a block diagram of a traffic transmission path planning device provided by one embodiment of the present application;

[0028] Figure 12 This is a structural block diagram of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0030] AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0031] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0032] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0033] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (Artificial Intelligence Generated Content, AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0034] A pre-training model, also known as a cornerstone model or large model, refers to a deep neural network (DNN) with large parameters. This model is trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, the pretrained machine learning (PTM) extracts common features from the data. Through techniques such as fine tuning, efficient parameter fine tuning (PEFT), and prompt-tuning, the model is then adapted for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be categorized by the data modality they process, including language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), and multimodal models (ViBERT, CLIP, Flamingo, Gato). Multimodal models are those that represent features from two or more data modalities. Pre-trained models are important tools for outputting artificial intelligence generated content (AIGC) and can also serve as a universal interface for connecting multiple specific task models.

[0035] The solutions provided in the embodiments of this application involve technologies such as large-scale model technology of artificial intelligence, which are specifically illustrated by the following embodiments.

[0036] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application. The implementation environment of the solution may include at least two task devices 10, at least two LAs 20, and at least one LC (Layer Convergence, aggregation layer switch) 30.

[0037] In some embodiments, the task device 10 may be an electronic device such as a PC (Personal Computer) or a server. The server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0038] In some embodiments, the task devices 10 communicate with each other based on a Ring-based (ring structure) method, and communication specifically refers to traffic (data) transmission between the task devices 10. Specifically, at least two task devices 10 involved in the AI ​​large model task are connected end to end in traffic interaction to form a ring, which is also called a communication traffic ring, that is, the traffic exchanged between the task devices is transmitted through a ring topology. For each task device, there is only one left neighbor and one right neighbor, and it will only receive data from the left neighbor and send data to the right neighbor. For example, Figure 2 As shown in the figure, a large AI model task involves four task devices: Task Device 1, Task Device 2, Task Device 3, and Task Device 4. These four task devices communicate using a ring-based method, forming a loop with traffic exchanged end to end. For example, Task Device 1 can only receive data from its left neighbor, Task Device 4, and send data to its right neighbor, Task Device 2.

[0039] In some embodiments, each task device can be connected to one or more LAs, each LA can be connected to one or more task devices, and each LA can be connected to one or more LCs, which is not limited in this application.

[0040] In some embodiments, when any two task devices 10 belong to the same LA group, the two task devices 10 can communicate with each other through the LA 20 connected thereto. Figure 1 As shown, task device 1 and task device 2 belong to the same LA group and can communicate through LA1 or LA2, which is not limited in this application.

[0041] In some embodiments, when any two task devices 10 do not belong to the same LA group, that is, when the two task devices need to transmit traffic across LA groups, the two task devices 10 can communicate with each other through the LA and LC connected to them. For example, task device 1 and task device 3 can communicate through LA1, LC1, and LA3, or through LA2, LC2, and LA3, although this application does not limit this.

[0042] In some embodiments, the technical solution proposed in this application can be applied to the parallel execution scenario of artificial intelligence models, that is, at least two task devices 10 jointly execute the same AI large model task, which can improve computing efficiency, increase system fault tolerance, shorten training time and reduce communication overhead.

[0043] Please refer to Figure 3, which shows a flow chart of a method for planning a flow transmission path provided by an embodiment of the present application. The execution subject of each step of the method may be a computer device, for example, the computer device may be Figure 1 The task device 10 in the solution implementation environment shown in FIG.

[0044] Step 310: Determine the LA groups to which each of the N task devices belongs, where each LA group includes at least one task device and a LA connected to the at least one task device. The N task devices are divided into at least two LA groups. The N task devices jointly execute the same AI large model task, where N is an integer greater than 1.

[0045] Large AI model tasks are AI algorithms or tasks that require extensive computation and processing. Typically, these tasks involve extensive data processing, computation, training, and optimization, requiring powerful computing and storage resources. To improve execution efficiency and performance, these large AI model tasks can be distributed across N task devices. This fully utilizes the computing power of each task device, speeding up task processing and reducing the burden on individual devices.

[0046] The large AI model task is distributed to N task devices for parallel execution. During parallel execution, data exchange and collaborative processing, such as message passing and data transmission, are required between the task devices. By exchanging data, sharing computation results, and synchronizing execution status among the N task devices, efficient task allocation and collaborative processing are achieved, helping to improve the execution efficiency and performance of the large AI model task.

[0047] A task device refers to a device with computing or execution capabilities that participates in the execution of AI large model tasks. LA can be used for communication between task devices. LA is part of a local area network (LAN), which provides data exchange, forwarding, and distribution functions, allowing communication and data transmission between task devices. For at least two task devices connected to the same LA, the task devices can communicate directly through the LA without the help of network devices at other levels. LA can achieve data transmission between task devices by forwarding data packets based on the MAC (Media Access Control) address of the target task device or other identification information of the target task device. The target task device here refers to the recipient of the data packet, that is, the device in the task device that receives the data packet.

[0048] LA group is a set of task devices and LAs connected to them, which can be used as a unit in the network architecture. Each LA group contains at least one task device and the LA connected to it. Figure 4 As shown, LA Group 1, for example, includes Task Device 1, Task Device 2, and access layer switches LA1, LA2, LA3, and LA4 connected thereto. LA Group 2 includes Task Device 3, Task Device 4, and access layer switches LA5, LA6, LA7, and LA8 connected thereto. Task Device 1, Task Device 2, Task Device 3, and Task Device 4 all execute the same large AI model task.

[0049] In some embodiments, LA can be configured with multiple ports, and each task device is connected to LA through one of the ports. For multiple task devices in the same LA group, when data exchange is required between task devices, they can transmit data packets through LA. LA will determine the forwarding path of the data packet based on the MAC address of the target task device and accurately deliver the data packet to the target task device. For example, please refer to Figure 4 , sub-figure (a) shows a schematic diagram of the LA group division of the four task devices, where task device 1 and task device 2 belong to LA group 1, and task device 3 and task device 4 belong to LA group 2. Sub-figure (b) shows the communication traffic ring composed of the four task devices. Assume that task device 1 wants to send data to task device 2, that is, the target task device is task device 2, first task device 1 can send a data packet to LA1, and LA1 forwards the received data packet to task device 2. LA2 can also be selected to forward the data packet, LA3 can also be selected to forward the data packet, or LA4 can be selected to forward the data packet, and this application does not limit this.

[0050] In some embodiments, for multiple task devices in different LA groups, when data exchange is required between task devices, they can transmit data packets through LA and LC. Specifically, LA first forwards the received data packet to the LC connected to it, and the LC then forwards the data packet to the target LA. The target LA can accurately transmit the data packet to the target task device through forwarding and routing functions based on the MAC address of the target task device in the data packet, thereby realizing data transmission between task devices across LA groups. Among them, LC is located between LA and the core layer switch SGLC. Its main function is to receive data packets from LA and transmit data packets from LA to SGLC or other target devices. In this application, LC transmits data packets from LA to target task devices. For example, please refer to Figure 5, sub-figure (a) shows a schematic diagram of the LA group division of 4 task devices, where task device 1 and task device 2 belong to LA group 1, and task device 3 and task device 4 belong to LA group 2. Sub-figure (b) shows a communication traffic ring composed of 4 task devices. Assume that task device 1 wants to send data to task device 3, that is, the target task device is task device 3, and task device 1 and task device 3 belong to different LA groups. First, task device 1 can send a data packet to LA1, then LA1 forwards the received data packet to LC2, LC2 sends the received data packet to LA5, and finally LA5 sends the data packet to task device 3. Similarly, LC2 can also be replaced by LC1. LA5 can also be replaced by LA6, LA7 or LA8. This application is not limited to this.

[0051] In some embodiments, for communication between two task devices that are not directly connected in a communication traffic ring, data transmission between the two task devices that are not directly connected can be achieved through an intermediate task device. For example, please refer to Figure 4 , assuming that task device 1 wants to send data to task device 3, that is, the target task device is task device 3, task device 1 and task device 3 belong to different LA groups and task device 1 and task device 3 are in Figure 4 The communication traffic loop shown is not directly connected. The intermediate task device between task device 1 and task device 3 is task device 2. First, task device 1 sends data to its task device 2 through an access layer switch (such as LA1). Then, task device 2 forwards the data to a LA (such as LA5) directly connected to task device 3 through an access layer switch (such as LA2) and an LC (such as LC1). Finally, the LA transmits the data to task device 3.

[0052] The above method, by using LA and LC for communication between task devices, can achieve efficient and low-latency data transmission, provide a faster and more stable communication environment, and thus support the collaborative execution of the same artificial intelligence (AI) large model task between task devices.

[0053] In some embodiments, the network segment information corresponding to each of the N task devices is obtained, and the network segment information corresponding to the task device refers to the network address segment to which the task device belongs; based on the network segment information corresponding to each of the N task devices, the LA group to which each of the N task devices belongs is determined, wherein task devices belonging to the same network address segment are divided into the same LA group, and task devices belonging to different network address segments are divided into different LA groups. Figure 4 As shown, the network segment information corresponding to task device 1 and task device 2 is the same, so task device 1 and task device 2 are divided into LA group 1, and the network segment information corresponding to task device 3 and task device 4 is the same, so task device 3 and task device 4 are divided into LA group 2.

[0054] A network address segment refers to a part of an IP (Internet Protocol) address, which is used to divide the address ranges of different networks. A network address segment consists of a network address and a subnet mask. Specifically, an IP address consists of 32-bit binary numbers, one part of which is used to represent the network address and the other part is used to represent the host address. The subnet mask is a 32-bit binary number used to identify the division of the network address and host address in the IP address. By applying the subnet mask to the IP address, the network address can be obtained. The main function of the network address segment is to determine whether the task device belongs to the same LA group. If the network address segments of two task devices are the same, that is, they have the same network address part, then they belong to the same LA group.

[0055] This method, which groups task devices into different LA groups based on network segment information, facilitates management and control of communication and data exchange between task devices. Task devices within the same LA group can communicate directly, while communication between different LA groups must be forwarded through an LC or other network device. This improves network efficiency and security.

[0056] In some embodiments, for each of the N task devices, the network segment information corresponding to the network card of the task device is obtained by calling the communication library as the network segment information corresponding to the task device; wherein, the network card of the task device refers to a hardware interface used to communicate with an external network or other devices, and the communication library is used to provide functions and interfaces related to network communication.

[0057] A network card (NIC) is a computer hardware device that physically connects a computer device to a network. NICs are typically connected to a link interface (LA) via a network cable, and a single LA can connect to multiple NICs. This connection between the NIC and the LA enables network communication between task devices, enabling data transmission and exchange. On a network, when task devices need to communicate, a data packet is sent to the LA via the NIC. The LA then forwards the packet to the target task device. If the target task device is not in the same LA group, the LA forwards the packet to the upper-layer device, which can be the LC. The LC checks the LA where the target task device is located and then forwards the packet to the LA where the target device is located.

[0058] Obtaining the network segment information corresponding to the task device's network card by calling the communication library means obtaining the task device's network card information by calling the functions or interfaces provided in the communication library, including obtaining the network card's name, MAC address, IP address, subnet mask, etc. This information can then be used to calculate the task device's network segment information and use it as the task device's corresponding network segment information.

[0059] The process of obtaining the network card information of the task device by calling the communication library is as follows: run the AI ​​large model task on the task device, and at the same time call the method of obtaining network card information in the communication library; obtain network card information, including network card name, MAC address, IP address, subnet mask, etc.; calculate the network segment information of the task device based on the obtained IP address and subnet mask; use the calculated task device network segment information as the task device's network segment information for subsequent task device LA group division and other operations.

[0060] The above method, by calling the communication library, can obtain the network segment information corresponding to N task devices. This method avoids the trouble of manual configuration and can automatically identify the network segment information of the task devices. Furthermore, based on the obtained network segment information, it can perform intelligent network management operations such as LA group division of task devices and traffic transmission path planning, thereby improving the efficiency and accuracy of network configuration.

[0061] In some embodiments, each task device is configured with multiple network cards, and different network cards can be connected to different LA ports. Different LAs have different network segment information, so different network cards correspond to different network segment information. For example, please refer to Figure 6 For example, each task device can be configured with eight network cards: network card 61, network card 62, network card 63, network card 64, network card 65, network card 66, network card 67, and network card 68. For each LA group, two LAs with the same grayscale value represent the same network segment information, while LAs with different grayscale values ​​represent different network segment information. For example, LA69 and LA610 have the same network segment information; for example, LA69 and LA611 have different network segment information. For two LAs in different LA groups, the same grayscale value represents different network segment information, while different grayscale values ​​represent different network segment information. For example, LA69 and LA612 have different network segment information; for example, LA69 and LA613 have different network segment information. Network card 61 and network card 62 can connect to LA69 and LA610, respectively. The network segment information corresponding to network card 61 and network card 62 is the same. The network segment information corresponding to the first network card of any task device included in LA group 1 is the same. The network segment information corresponding to the first network card of multiple task devices in different LA groups is different.

[0062] In some embodiments, based on the configuration parameters of the communication library, the network segment information corresponding to the first network card of the task device is obtained as the network segment information corresponding to the task device; wherein the first network card is a network card indicated by the configuration parameters among the multiple network cards of the task device, and the configuration parameters are used to indicate that the network segment information corresponding to the first network card is obtained from the multiple network cards.

[0063] The first network card refers to any network card in the task device. For example, please refer to Figure 6 , the first network card can be any one of the eight network cards.

[0064] The configuration parameters of the communication library refer to a set of parameters used to specify the behavior and settings of the communication library. Specifically, the configuration parameters of the communication library include network card selection parameters, which are used to indicate the network segment information of which network card to select.

[0065] Optionally, the communication library configuration parameters may also include security parameters, which are used to set the communication library's security options, such as encryption algorithms, authentication methods, and access control. They may also include performance optimization parameters, which are used to adjust the communication library's performance and resource utilization, such as buffer size, number of concurrent connections, and timeout period.

[0066] The above method can easily obtain the network segment information corresponding to any network card of the task device by modifying the configuration parameters of the communication library, thereby achieving a more flexible and scalable network configuration.

[0067] In some embodiments, please refer to Figure 7 , which shows a flow chart of a method for planning a traffic transmission path provided by another embodiment of the present application. For each task device, in the startup module, after starting the artificial intelligence AI model, the model will call the Ring-based algorithm of the communication library to allow multiple task devices to synchronize data and enable the task device to enter the initialization ring building state. In the initialization ring building module, the communication library of the task device will actively retrieve the task device network card information and try to obtain the network segment information of the task device. Entering the topology perception module, when N task devices perform the initialization ring building operation in the communication library, the bootstrap network will organize the interaction of relevant information (such as network segment information) between the task devices. The bootstrap network is used for communication between task devices, and it is a bridge for data interaction between task devices. Specifically, the bootstrap network can collect the network segment information of all task devices through the AllGather (global collection) operation. For each task device, based on the network segment information of this device and other task devices, the LA group information of this device and other task devices can be obtained.

[0068] In some embodiments, in a cloud environment, due to factors such as virtualization and security, the communication library cannot obtain the network card information of the task device, and therefore cannot directly perform network topology positioning through the network segment information of the task device. To solve this problem, the hash value corresponding to each of the N task devices can be obtained, and the hash value can be used to determine the LA group to which the task device belongs. Specifically, the hash value corresponding to each of the N task devices is obtained, and the hash value corresponding to the task device is determined according to the identification information of the task device through a hash algorithm; based on the hash value corresponding to each of the N task devices, the LA group to which each of the N task devices belongs is determined, wherein task devices corresponding to the same hash value are divided into the same LA group, and task devices corresponding to different hash values ​​are divided into different LA groups.

[0069] A hash value is a unique numerical value calculated using a hash algorithm based on the identification information of a task device. A hash algorithm is an algorithm that maps data of any size to a fixed-size value. Optionally, the hash algorithm may be MD5 (Message Digest Algorithm 5), SHA1 (Secure Hash Algorithm 1), SHA256 (Secure Hash Algorithm 256-bit), or the like. The identification information of a task device may be a unique identifier for the task device, such as a device ID (Identification), MAC address, or IP address.

[0070] In network topology positioning, the hash algorithm can map the identification information of the task device into a corresponding hash value.

[0071] The above method collects the identification information of N task devices and then calculates the corresponding hash value using a hash algorithm. Task devices with the same hash value are assigned to the same LA group, while task devices with different hash values ​​are assigned to different LA groups. This allows task devices to be effectively grouped to achieve appropriate network topology positioning without relying on their network segment information.

[0072] In some embodiments, for each of the N task devices, the communication thread of the communication library obtains the hash value corresponding to the task device from the controller; wherein the controller is used to centrally manage and schedule the communication information of the N task devices, and the communication library is used to provide functions and interfaces related to network communication. Figure 8, the communication thread of each task device's communication library initiates a request to the controller to obtain the hash value corresponding to the task device. The specific process can be: each task device's communication thread initiates a request to the controller to obtain the hash value corresponding to the device. This can be achieved through the interface provided by the communication library. After receiving the request, the controller runs a hash algorithm based on the identification information of the task device (such as device ID, IP address, etc.) to calculate the hash value corresponding to the device. The controller returns the calculated hash value to the communication thread of the requesting task device.

[0073] The communication library's communication thread is the thread running within the library responsible for handling communication operations for task devices. The controller is the component responsible for centrally managing and scheduling communication information for these N task devices. It performs hash calculations based on the task device's specific identification information (such as device ID, IP address, etc.), enabling grouping and locating of task devices. It also manages and schedules communication between task devices, ensuring orderly collaboration among them.

[0074] The above method, through the centralized management and scheduling of N task devices by a controller, ensures orderly, efficient, and stable communication between task devices and guarantees the secure transmission of data between task devices. Furthermore, by using a hash algorithm to locate task devices, task devices can be effectively divided into different LA groups, thereby improving the efficiency and reliability of data transmission between task devices. Furthermore, through the communication threads of the communication library, task devices can quickly and accurately obtain their corresponding hash values, which facilitates topological awareness and organization of task devices, further improving the efficiency and quality of collaboration between task devices.

[0075] In some embodiments, please refer to Figure 7 ,In the topology awareness module, when the bootstrap network in the ,communication library is initialized, the communication thread will obtain the ,hash value of the task device from the controller.

[0076] Step 320, based on the LA groups to which the N task devices belong, obtain a communication traffic ring composed of N task devices, wherein the communication traffic ring refers to a path closed loop formed by connecting N task devices in sequence, which is used to describe the transmission path of the communication traffic between the N task devices, and each task device belonging to the same LA group is connected in series.

[0077] Large AI models often use three collective communication primitives: AllReduce (global reduction), ReduceScatter (distributed reduction), and AllGather to synchronize and transfer data between multiple task devices. AllReduce applies a reduction operation (such as addition or multiplication) to the data values ​​to be reduced on each task device and then returns the reduced results to each task device synchronously. AllReduce is commonly used in computational scenarios such as aggregating model parameters and calculating loss functions, significantly accelerating computation. ReduceScatter applies a reduction operation to the data values ​​to be reduced on each task device and then returns the reduced results to multiple task devices. ReduceScatter is commonly used to reduce global data to each task device for parallel processing. In ReduceScatter, each task device only receives a portion of the reduced results from other task devices, rather than the entire result. This helps reduce communication link load and data transmission latency, thereby improving communication efficiency. AllGather gathers the local data from each task device to form a global view and returns it synchronously to each task device. In an AllGather operation, each task device sends its local data to all other task devices and collects the local data from all task devices. This AllGather operation aggregates the local data from each task device into a global view, enabling parallel computation and processing. These three types of aggregation primitives are implemented using a ring-based approach, combining N task devices into a traffic ring. This ring determines the transmission path for traffic between the N task devices. Different traffic transmission paths directly affect the amount of traffic across the LA group, thereby impacting communication efficiency.

[0078] Specifically, to synchronize and transfer data for large models, these collective primitives require data transmission between multiple task devices. Implementing a collective primitive based on a ring structure connects N task devices into a traffic ring, thus establishing a data communication path between the task devices. Different task devices may belong to different LA groups, which results in data being transferred across LA groups. If the amount of data transferred across groups is large, this increases communication latency and bandwidth pressure, reducing communication efficiency. Therefore, how to construct the traffic ring becomes the key to optimizing the performance of collective communication primitives.

[0079] In some embodiments, please refer to Figure 7In the LA group classification module, N task devices are divided into M groups based on their respective LA groups, with M being an integer greater than 1. In the logical concatenation module, for each group of M task devices, the individual task devices within each group are connected in series to form M task device chains. In the topology affinity ring building module, these M task device chains are sequentially connected to form a communication traffic ring consisting of N task devices. Finally, when the N task devices communicate, traffic is transmitted in a ring-like sequence.

[0080] A task device chain defines the order and direction of data transmission between task devices. It can be used to specify the direction of data transmission and reception, as well as the path along which data is transmitted. For a task device chain [task device A -> task device D], task device A sends data, and task device D receives it.

[0081] Any task device has obtained the topological location information of other task devices. For each task device, the obtained information may be: [Task device A: LA Group 1, Task device B: LA Group 2, Task device C: LA Group 2, Task device D: LA Group 1, Task device E: LA Group 3]. Based on this information, the task devices in the same LA group are clustered, and N task devices are divided into three groups: LA Group 1: [Task device A, Task device D], LA Group 2: [Task device B, Task device C], and LA Group 3: [Task device E].

[0082] In some embodiments, based on the above information, the task devices contained in each group of task devices are connected in series to obtain M task device chains. In these task device chains, the relative positions of the task devices in each group can be arranged in a variety of ways, which is not limited in this application. For example, in the process of building a communication library ring, the devices of the same LA group are logically connected in series first. For example, for LA group 1 in the above example, the task device chain can be [task device A->task device D] or [task device D->task device A]. For LA group 2 in the above example, the task device chain can be [task device B->task device C] or [task device C->task device B]. The setting of the task device chain can be flexibly configured according to specific needs and system architecture. According to the different positions and traffic transmission paths between the task devices, a suitable task device chain can be determined. By setting an appropriate task device chain, the path and direction of data transmission can be optimized, and communication efficiency and performance can be improved.

[0083] In some embodiments, the relative positions of the task devices within each group are arranged differently, and the traffic transmission paths are also different. For a task device chain [task device A -> task device D], task device A sends data and task device D receives the data. For a task device chain [task device D -> task device A], task device D sends data and task device A receives the data.

[0084] In some embodiments, M task device chains are connected in sequence to obtain a communication traffic ring consisting of N task devices. The communication traffic can be arranged in a variety of ways, which is not limited in this application. For the task device chain [task device A->task device D], the task device chain [task device B->task device C] and the task device chain [task device E], the communication traffic ring can be [task device A->task device D->task device B->task device C->task device E->task device A]. The communication traffic ring can also be [task device A->task device D->task device E->task device B->task device C->task device A]. For the task device chain [task device D->task device A], the task device chain [task device C->task device B] and the task device chain [task device E], the communication traffic ring can be [task device D->task device A->task device C->task device B->task device E->task device D]. The communication traffic ring can be [task device D->task device A->task device E->task device C->task device B->task device D]. The communication traffic ring can be [task device D->task device A->task device E->task device C->task device B->task device D]. This application does not limit this. Therefore, this application allows for multiple different arrangements of M task device chains to form a communication traffic ring consisting of N task devices. This flexibility allows the technical solution to adapt to different network topologies and environmental requirements, and provide more efficient communication transmission. During implementation, the appropriate task device chain arrangement can be selected based on specific needs to meet the communication requirements of different scenarios.

[0085] In some embodiments, if the network segment information of N task devices is obtained, the N task devices can also be divided into M groups of task devices according to the network segment information corresponding to the N task devices, based on the principle that devices with the same network segment information are classified into the same group of task devices. For example, for each task device, the information obtained can be [task device A: network segment 1, task device B: network segment 2, task device C: network segment 2, task device D: network segment 1, task device E: network segment 3]. Based on the above information, the task devices in the same network segment are clustered, and the N task devices are divided into 3 groups of task devices. Network segment 1: [device A, device D], network segment 2: [device B, device C] and network segment 3: [device E]. The present application does not limit how to divide N task devices into M groups of task devices. They can be divided according to the LA group of the N task devices, or according to the network segment information of the N task devices. The division method can be flexibly adjusted according to specific needs and scenarios.

[0086] The above method can effectively manage and optimize the communication between N task devices. This organization ensures that communication traffic is transmitted within the same LA group as much as possible, reducing cross-LA group traffic and further improving communication efficiency and reliability.

[0087] In some embodiments, the technical solution proposed in this application can greatly reduce the traffic across LA groups. For example, Figure 5 As shown, when the order of the communication traffic ring composed of task devices is task device 1->task device 3->task device 2->task device 4->task device 1, specifically, when task device 1 and task device 3 communicate, task device 1 and task device 3 belong to different LA groups, so it is necessary to transmit traffic across LA groups. For example, LA1, LC1, and LA5 can be selected for forwarding traffic; when task device 3 and task device 2 communicate, task device 3 and task device 2 belong to different LA groups, so it is necessary to transmit traffic across LA groups. For example, LA5, LC2, and LA4 can be selected for forwarding traffic; when task device 2 and task device 4 communicate, task device 2 and task device 4 belong to different LA groups, so it is necessary to transmit traffic across LA groups. For example, LA2, LC2, and LA6 can be selected for forwarding traffic; when task device 4 and task device 1 communicate, task device 4 and task device 1 belong to different LA groups, so it is necessary to transmit traffic across LA groups. For example, LA6, LC2, and LA4 can be selected for forwarding traffic. Therefore, according to the above traffic communication ring, the number of times the traffic crosses the LA group is 4 times, and the number of times the traffic is restricted within the LA group is 0 times, that is, each traffic communication is across the LA group. We can establish an indicator to measure the quality of the communication traffic ring, that is, Figure 4The communication traffic ring shown, the traffic topology affinity

[0088] According to the technical solution provided by this application, topology affinity planning is performed on the four task devices, such as Figure 5 As shown, the ring order of the task devices is changed to task device 1 -> task device 2 -> task device 3 -> task device 4 -> task device 1. When task devices 1 and 2 communicate, they belong to the same LA group, so there is no need to transmit traffic across LA groups. For example, LA1 can be selected for traffic forwarding. When task devices 2 and 3 communicate, they belong to different LA groups, so there is need to transmit traffic across LA groups. For example, LA1, LC1, and LA5 can be selected for traffic forwarding. When task devices 3 and 4 communicate, they belong to the same LA group, so there is no need to transmit traffic across LA groups. For example, LA5 can be selected for traffic forwarding. When task devices 4 and 1 communicate, they belong to different LA groups, so there is need to transmit traffic across LA groups. For example, LA8, LC1, and LA4 can be selected for traffic forwarding. Therefore, according to the communication traffic ring, the number of times the traffic crosses LA groups is 2, and the number of times the traffic is limited within the LA group is 2. Compared Figure 4 , the cross-LA group traffic drops by 50%, and the traffic topology affinity is Achieve optimal topological affinity.

[0089] In some embodiments, the technical solution provided by this application reduces the traffic across LA groups. Specific measured data, such as Figure 9 As shown, the horizontal axis represents time and the vertical axis represents the ratio of cross-LA group traffic. When the artificial intelligence AI large model task runs on the AllReduce (ring) of communication library 1, the cross-LA group traffic is as high as 91%, where AllReduce (ring) is a collective communication primitive using a ring topology structure to synchronize and transmit data between multiple task devices. By using this primitive, task devices can reduce their data and then transmit the results to other task devices through a ring path. However, without optimization, large amounts of data transmission across LA groups may result in higher communication delays and bandwidth pressure. In order to solve this problem, the present application provides a topology affinity technical solution, which optimizes the connection relationship between task devices through communication library 2 and reduces traffic transmission across LA groups. In actual tests, this technical solution reduced cross-LA group traffic by 75%. Among them, communication library 1 and communication library 2 are communication libraries that can be obtained through open source.

[0090] In some embodiments, the present application conducted comparative tests to evaluate the performance of the technical solution provided by the present application in terms of AllReduce, and demonstrated its greater stability. AllReduce is a collective communication operation in parallel computing that is used to reduce data from multiple task devices and distribute the results to all task devices. It is used to implement global reduction operations in parallel computing, such as summation and averaging. Figure 10 The test results are shown, where the horizontal axis represents the number of runs and the vertical axis represents the bus bandwidth. The test environment is a cloud environment, using a single network card and 4 task device nodes, and 200 long-term stability tests were performed between communication library 1 and communication library 2. According to the data in the figure, it can be observed that the bus bandwidth of communication library 1 fluctuates greatly, while the bandwidth of communication library 2 is close to stable with almost no jitter. This is because the communication library 2 in the technical solution provided by this application utilizes the characteristics of topological affinity to place communication traffic within the same LA group as much as possible, thereby reducing the probability of load imbalance. Load imbalance may cause some task devices to process too much data while other task devices are idle, thereby affecting the overall communication performance and efficiency. By optimizing the connection path between task devices, communication library 2 successfully reduces data transmission across LA groups, making the communication traffic more evenly distributed within each LA group. This optimized design enables communication library 2 to have more stable bus bandwidth performance during the AllReduce process, reduces communication delay, and improves data transmission efficiency between task devices. Therefore, based on the technical solution provided in this application, the comparative test results show that communication library 2 has stronger stability, can effectively reduce the load imbalance problem, and optimize the utilization of bus bandwidth, thereby improving the performance of AllReduce and overall communication efficiency.

[0091] In some embodiments, the technical solution provided by this application uses a method to reduce the probability of hash conflicts to improve performance. In the training of large model tasks, due to the lack of topological affinity conditions, there is a load imbalance on the link. By introducing relevant optimization measures at the communication level, significant performance improvements have been achieved. Specifically, at the communication level, we conducted performance index tests on large model tasks. Without using topological affinity to bypass the load imbalance link, the measured bandwidth of the 4G message size of the task is 87.51GB / s. However, through our technical solution, after using topological affinity to bypass the load imbalance link, the bandwidth was increased to 140.71GB / s, an increase of nearly 60%. This means that by reducing the load imbalance of the link, the data transmission rate is significantly improved. In addition to the performance improvements at the communication level, our technical solution also improves the performance indicators of the task. After optimization, the sample transmission rate increased by 11.4%. Through this improvement, we can process task data more efficiently.

[0092] In some embodiments, the technical solution provided by the present application allows network architects to limit traffic as much as possible within the LA group. When traffic is transmitted within the LA group, the LA convergence ratio can be improved due to its localized characteristics. The LA convergence ratio (Local Area Convergence Ratio) refers to the distribution ratio of traffic transmitted within a specific network area (within the LA group) within the area (within the LA group). By limiting the transmission of traffic within a smaller area (within the LA group), the cross-regional transmission demand can be reduced, and the delay and resource consumption caused by cross-regional (LA group) communications can be reduced. By improving the LA convergence ratio, the number of LCs and SGLCs built can be reduced, thereby achieving the purpose of reducing costs.

[0093] The technical solution provided by this application rationally plans the relative positions of N task devices based on the LA groups to which they each belong during the construction of a communication traffic ring. Task devices belonging to the same LA group are preferentially connected in series, and the N task devices are then combined to form a closed path loop, serving as the final communication traffic ring. This method, by preferentially connecting devices in the same LA group in series, avoids unnecessary cross-LA group traffic transmission during traffic transmission, reduces cross-LA group traffic, and improves communication efficiency between task devices.

[0094] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0095] Please refer to Figure 11 , which shows a block diagram of a traffic transmission path planning device provided by one embodiment of the present application. The device has the function of implementing the above-mentioned traffic transmission path planning method. The function can be implemented by hardware or by hardware executing corresponding software. The device can be a computer device or can be set in a computer device. The device 1100 can include: a determination module 1110 and an acquisition module 1120.

[0096] Determination module 1110 is used to determine the access layer switch LA group to which each of the N task devices belongs, wherein each LA group includes at least one task device and a LA connected to the at least one task device, the N task devices are divided into at least two LA groups, and the N task devices jointly execute the same artificial intelligence AI large model task, where N is an integer greater than 1.

[0097] Obtain module 1120, which is used to obtain a communication traffic ring composed of the N task devices according to the LA groups to which the N task devices respectively belong, wherein the communication traffic ring refers to a path closed loop formed by connecting the N task devices in sequence, which is used to describe the transmission path of the communication traffic between the N task devices, and each of the task devices belonging to the same LA group is connected in series.

[0098] In some embodiments, the determining module 1110 includes: a first obtaining unit and a first determining unit ( Figure 11 not shown).

[0099] The first acquiring unit is configured to acquire network segment information corresponding to each of the N task devices, where the network segment information corresponding to the task device refers to a network address segment to which the task device belongs.

[0100] The first determination unit is used to determine the LA group to which the N task devices each belong based on the network segment information corresponding to each of the N task devices, wherein task devices belonging to the same network address segment are divided into the same LA group, and task devices belonging to different network address segments are divided into different LA groups.

[0101] In some embodiments, the first acquisition unit is used to obtain the network segment information corresponding to the network card of each of the N task devices by calling the communication library as the network segment information corresponding to the task device; wherein the network card of the task device refers to a hardware interface for communicating with an external network or other devices, and the communication library is used to provide functions and interfaces related to network communication.

[0102] In some embodiments, the first acquisition unit is used to obtain the network segment information corresponding to the first network card of the task device according to the configuration parameters of the communication library, as the network segment information corresponding to the task device; wherein, the first network card is a network card indicated by the configuration parameters among the multiple network cards of the task device, and the configuration parameters are used to indicate the acquisition of the network segment information corresponding to the first network card from the multiple network cards.

[0103] In some embodiments, the determining module 1120 includes: a second obtaining unit and a second determining unit ( Figure 11 not shown).

[0104] The second acquiring unit is configured to acquire a hash value corresponding to each of the N task devices, where the hash value corresponding to the task device is determined according to identification information of the task device through a hash algorithm.

[0105] The second determination unit is used to determine the LA group to which the N task devices each belong based on the hash values ​​corresponding to the N task devices, wherein task devices corresponding to the same hash value are divided into the same LA group, and task devices corresponding to different hash values ​​are divided into different LA groups.

[0106] In some embodiments, the second acquisition unit is used to obtain the hash value corresponding to each of the N task devices from the controller through the communication thread of the communication library; wherein the controller is used to centrally manage and schedule the communication information of the N task devices, and the communication library is used to provide functions and interfaces related to network communication.

[0107] In some embodiments, the obtaining module 1120 is used to divide the N task devices into M groups of task devices according to the LA groups to which the N task devices respectively belong, and in accordance with the principle that devices with the same LA group are classified into the same group of task devices, where M is an integer greater than 1; for each group of task devices in the M groups of task devices, the individual task devices contained in each group of task devices are connected in series to obtain M task device chains; and the M task device chains are connected in sequence to obtain the communication traffic ring composed of the N task devices.

[0108] The technical solution provided by this application rationally plans the relative positions of N task devices based on the LA groups to which they each belong during the construction of a communication traffic ring. Task devices belonging to the same LA group are preferentially connected in series, and the N task devices are then combined to form a closed path loop, serving as the final communication traffic ring. This method, by preferentially connecting devices in the same LA group in series, avoids unnecessary cross-LA group traffic transmission during traffic transmission, reduces cross-LA group traffic, and improves communication efficiency between task devices.

[0109] It should be noted that the apparatus provided in the above embodiments, when implementing its functions, is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0110] Please refer to Figure 12 , which shows a structural block diagram of a computer device 1200 provided in one embodiment of the present application.

[0111] Typically, the computer device 1200 includes a processor 1210 and a memory 1220 .

[0112] The processor 1210 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1210 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1210 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1210 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1210 may also include an AI processor for processing computing operations related to machine learning.

[0113] Memory 1220 may include one or more computer-readable storage media, which may be non-transitory. Memory 1220 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1220 is used to store a computer program, which is configured to be executed by one or more processors to implement the above-mentioned traffic transmission path planning method.

[0114] Those skilled in the art will understand that Figure 12 The structure shown in the figure does not constitute a limitation on the computer device 1200, and the computer device 1200 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.

[0115] In some embodiments, a computer-readable storage medium is further provided, wherein the storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned method for planning a traffic transmission path.

[0116] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or an optical disk, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).

[0117] In some embodiments, a computer program product is also provided, which includes a computer program stored in a computer-readable storage medium, and a processor reads and executes the computer program from the computer-readable storage medium to implement the above-mentioned traffic transmission path planning method.

[0118] It should be understood that the "multiple" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. In addition, the step numbers described in this article only illustrate a possible execution sequence between the steps. In some other embodiments, the above steps may not be executed in the order of the numbers, such as two steps with different numbers are executed at the same time, or two steps with different numbers are executed in the opposite order to the diagram. The embodiments of the present application do not limit this.

[0119] The above description is merely an exemplary embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for planning a traffic transmission path, characterized in that: The method comprises: Determine the access layer switch LA group to which each of the N task devices belongs, wherein each LA group includes at least one task device and an LA connected to the at least one task device, the N task devices are divided into at least two LA groups, and the N task devices jointly execute the same artificial intelligence (AI) large model task, where N is an integer greater than 1; According to the LA groups to which the N task devices respectively belong, and in accordance with the principle that devices with the same LA group are grouped as the same task device, the N task devices are divided into M groups of task devices; or, according to the network segment information corresponding to the N task devices respectively, and in accordance with the principle that devices with the same network segment information are grouped as the same task device, the N task devices are divided into M groups of task devices, where M is an integer greater than 1; For each group of task devices in the M groups of task devices, each task device included in each group of task devices is connected in series to obtain M task device chains; The M task device chains are connected in sequence to obtain a communication traffic ring composed of the N task devices, wherein the communication traffic ring refers to a closed path loop formed by the N task devices connected in sequence, which is used to describe the transmission path of the communication traffic between the N task devices. The task devices belonging to the same LA group are connected in series, or the task devices with the same corresponding network segment information are connected in series; The quotient between the number of times the traffic in the communication traffic ring is limited within the LA group and the number of times the traffic crosses the LA group is determined as the traffic topology affinity of the communication traffic ring. The traffic topology affinity is used to measure the quality of the communication traffic ring, wherein traffic crosses the LA group means that the task device sending the traffic and the task device receiving the traffic belong to different LA groups, and traffic limited within the LA group means that the task device sending the traffic and the task device receiving the traffic belong to the same LA group.

2. The method according to claim 1, characterized in that Determining the LA group to which each of the N task devices belongs includes: Obtaining network segment information corresponding to each of the N task devices, where the network segment information corresponding to the task device refers to the network address segment to which the task device belongs; According to the network segment information corresponding to each of the N task devices, the LA groups to which the N task devices belong are determined, wherein task devices belonging to the same network address segment are divided into the same LA group, and task devices belonging to different network address segments are divided into different LA groups.

3. The method according to claim 2, characterized in that The obtaining of the network segment information corresponding to each of the N task devices includes: For each of the N task devices, obtain the network segment information corresponding to the network card of the task device by calling the communication library as the network segment information corresponding to the task device; The network card of the task device refers to a hardware interface for communicating with an external network or other devices, and the communication library is used to provide functions and interfaces related to network communication.

4. The method according to claim 3, characterized in that The acquiring, by calling a communication library, the network segment information corresponding to the network card of the task device as the network segment information corresponding to the task device includes: According to the configuration parameters of the communication library, obtain the network segment information corresponding to the first network card of the task device as the network segment information corresponding to the task device; The first network card is a network card indicated by the configuration parameter among multiple network cards of the task device, and the configuration parameter is used to instruct to obtain network segment information corresponding to the first network card from the multiple network cards.

5. The method according to claim 1, wherein Determining the LA group to which each of the N task devices belongs includes: Obtaining a hash value corresponding to each of the N task devices, where the hash value corresponding to the task device is determined based on identification information of the task device using a hash algorithm; According to the hash values ​​corresponding to the N task devices, the LA groups to which the N task devices belong are determined, wherein task devices corresponding to the same hash value are divided into the same LA group, and task devices corresponding to different hash values ​​are divided into different LA groups.

6. The method according to claim 5, characterized in that The obtaining of the hash values ​​corresponding to the N task devices includes: For each of the N task devices, obtain a hash value corresponding to the task device from the controller through a communication thread of the communication library; The controller is used to centrally manage and schedule the communication information of the N task devices, and the communication library is used to provide functions and interfaces related to network communication.

7. A flow transmission path planning device, characterized in that: The device comprises: A determination module is configured to determine the access layer switch LA group to which each of the N task devices belongs, wherein each LA group includes at least one task device and an LA connected to the at least one task device, the N task devices are divided into at least two LA groups, and the N task devices jointly execute the same artificial intelligence (AI) large model task, where N is an integer greater than 1; An obtaining module is configured to divide the N task devices into M groups of task devices according to the LA groups to which the N task devices respectively belong and the principle that devices with the same LA groups are classified into the same group of task devices; or, according to the network segment information corresponding to the N task devices respectively and the principle that devices with the same network segment information are classified into the same group of task devices, divide the N task devices into M groups of task devices, where M is an integer greater than 1; for each group of task devices in the M groups of task devices, connect the task devices contained in each group of task devices in series to obtain M task device chains; connect the M task device chains in sequence to obtain a communication traffic ring composed of the N task devices, wherein the communication traffic ring refers to a path closed loop formed by connecting the N task devices in sequence, which is used to describe the transmission path of the communication traffic between the N task devices, and the task devices belonging to the same LA group are connected in series, or the task devices with the same corresponding network segment information are connected in series; The determination module is also used to determine the quotient between the number of times the traffic in the communication traffic ring is limited within the LA group and the number of times the traffic crosses the LA group as the traffic topology affinity of the communication traffic ring, and the traffic topology affinity is used to measure the quality of the communication traffic ring, wherein traffic across the LA group means that the task device sending the traffic and the task device receiving the traffic belong to different LA groups, and traffic limited within the LA group means that the task device sending the traffic and the task device receiving the traffic belong to the same LA group.

8. The device according to claim 7, characterized in that The determination module includes a first acquisition unit and a first determination unit; The first acquiring unit is configured to acquire network segment information corresponding to each of the N task devices, where the network segment information corresponding to the task device refers to the network address segment to which the task device belongs; The first determination unit is used to determine the LA group to which the N task devices each belong based on the network segment information corresponding to each of the N task devices, wherein task devices belonging to the same network address segment are divided into the same LA group, and task devices belonging to different network address segments are divided into different LA groups.

9. The device according to claim 8, characterized in that The first acquisition unit is used to obtain the network segment information corresponding to the network card of each of the N task devices by calling the communication library, as the network segment information corresponding to the task device; wherein the network card of the task device refers to a hardware interface used to communicate with an external network or other devices, and the communication library is used to provide functions and interfaces related to network communication.

10. The device according to claim 9, characterized in that The first acquisition unit is used to obtain the network segment information corresponding to the first network card of the task device according to the configuration parameters of the communication library, as the network segment information corresponding to the task device; wherein, the first network card is a network card indicated by the configuration parameters among the multiple network cards of the task device, and the configuration parameters are used to indicate the acquisition of the network segment information corresponding to the first network card from the multiple network cards.

11. The device according to claim 7, characterized in that The determining module includes a second acquiring unit and a second determining unit; The second acquiring unit is configured to acquire a hash value corresponding to each of the N task devices, where the hash value corresponding to the task device is determined based on identification information of the task device using a hash algorithm; The second determination unit is used to determine the LA group to which each of the N task devices belongs based on the hash values ​​corresponding to each of the N task devices, wherein task devices corresponding to the same hash value are divided into the same LA group, and task devices corresponding to different hash values ​​are divided into different LA groups.

12. The device according to claim 11, characterized in that The second acquisition unit is used to obtain the hash value corresponding to each of the N task devices from the controller through the communication thread of the communication library; wherein the controller is used to centrally manage and schedule the communication information of the N task devices, and the communication library is used to provide functions and interfaces related to network communication.

13. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.

15. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a computer-readable storage medium. A processor reads and executes the computer program from the computer-readable storage medium to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Model training method, device, equipment, medium and system

    CN115879543A