Chip interconnection communication method based on performance and complexity balance adaptive algorithm
By adopting adaptive algorithms based on performance and complexity balance in high-performance computing systems, dynamically adjusting the communication path and task allocation between computing power chips, the problem of poor interconnection communication efficiency in computing power chips in high-performance computing systems is solved, and efficient communication optimization and overall performance improvement are achieved.
Patent Information
- Application Number
- CN202510171028.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In high-performance computing systems, the interconnection communication between computing power chips is poor, resulting in limited overall performance.
Adaptive algorithm based on performance and complexity balance is adopted to monitor the communication performance and computational complexity of data tasks in real time, adjust the scheduling strategy dynamically, and optimize communication paths and task allocation.
It realizes efficient optimization of interconnected communication between computing power chips, dynamically balances performance and complexity, adapts to different application scenarios and complex network environments, and improves the overall performance of high-performance computing.
Smart Images

Figure CN120111048A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a chip interconnection communication method, device, electronic device and medium based on a performance and complexity balance adaptive algorithm. Background Art
[0002] With the rapid development of artificial intelligence, deep learning, big data analysis, biological sciences and other fields, the demand for computing power continues to grow, and the requirements for high-performance computing (HPC) are also getting higher and higher. In the existing technology, systems for high-performance computing are usually composed of hundreds to thousands of computing chips, which work together through high-speed interconnection networks to achieve parallel processing of complex tasks. However, in systems for high-performance computing, the efficiency of interconnection and communication between computing chips is poor, which in turn affects the overall performance of high-performance computing. Summary of the invention
[0003] In view of the above problems, a chip interconnection communication method, device, electronic device and medium based on a performance and complexity balance adaptive algorithm are proposed to overcome the above problems or at least partially solve the above problems, including:
[0004] A chip interconnection communication method based on a performance and complexity balance adaptive algorithm is applied to a computing chip cluster system, wherein the computing chip cluster system includes multiple computing chip nodes, and the method includes:
[0005] Acquire a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes;
[0006] Determine the communication performance data and computational complexity data corresponding to the scheduling strategy, adjust the scheduling strategy according to the communication performance data and computational complexity data, and process the data task according to the adjusted scheduling strategy.
[0007] Optionally, the determining the communication performance data and the computational complexity data corresponding to the scheduling strategy includes:
[0008] Obtaining real-time monitoring data of the multiple target computing power chip nodes;
[0009] According to the real-time monitoring data, communication performance data corresponding to the scheduling strategy is determined.
[0010] Optionally, the communication performance data includes any one or more of the following: bandwidth utilization, average communication delay, and load balancing degree of task allocation.
[0011] Optionally, the determining the communication performance data and the computational complexity data corresponding to the scheduling strategy includes:
[0012] Determine the computational effort and computational power consumption corresponding to the scheduling strategy;
[0013] The computational complexity data of the scheduling strategy is determined according to the computational amount and the computational power consumption.
[0014] Optionally, adjusting the scheduling strategy according to the communication performance data and the computational complexity data includes:
[0015] Obtaining a preset communication performance threshold and a computational complexity threshold;
[0016] The scheduling strategy is adjusted according to the comparison result between the communication performance data and the communication performance threshold and the comparison result between the calculation complexity data and the calculation complexity threshold.
[0017] Optionally, adjusting the scheduling strategy according to a comparison result between the communication performance data and the communication performance threshold and a comparison result between the computational complexity data and the computational complexity threshold includes:
[0018] When the communication performance data is greater than the communication performance threshold, update the information of multiple target computing power chip nodes and the communication paths between the multiple target computing power chip nodes;
[0019] When the computational complexity data is greater than a computational complexity threshold, the scheduling algorithm is adjusted.
[0020] Optionally, it also includes:
[0021] The computational complexity threshold is adjusted according to the communication performance data.
[0022] A chip interconnection communication device based on a performance and complexity balance adaptive algorithm is applied to a computing chip cluster system, wherein the computing chip cluster system includes a plurality of computing chip nodes, and the device includes:
[0023] A scheduling strategy generation and execution module, used to obtain a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes;
[0024] The scheduling strategy adjustment and execution module determines the communication performance data and computational complexity data corresponding to the scheduling strategy, adjusts the scheduling strategy according to the communication performance data and computational complexity data, and processes the data task according to the adjusted scheduling strategy.
[0025] An electronic device comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method described in the first line when executed by the processor.
[0026] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0027] The embodiments of the present invention have the following advantages:
[0028] In an embodiment of the present invention, a data task is obtained in a computing chip cluster system, a scheduling strategy for the data task is generated, and the data task is processed according to the scheduling strategy; the scheduling strategy includes information of multiple target computing chip nodes for processing data tasks and communication path information between multiple target computing chip nodes, and then the communication performance data and computational complexity data corresponding to the scheduling strategy are determined, the scheduling strategy is adjusted according to the communication performance data and the computational complexity data, and the data task is processed according to the adjusted scheduling strategy, thereby realizing adaptive adjustment of the scheduling strategy for indicating the communication path between computing chip nodes, achieving a dynamic balance between performance and complexity to adapt to the needs of different application scenarios and complex network environments, improving the efficiency of interconnection and communication between computing chips, and thereby improving the overall performance of high-performance computing. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0030] Figure 1 is a schematic diagram of a system architecture provided by some embodiments of the present invention;
[0031] Figure 2 is a flowchart of a chip interconnection communication method based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention;
[0032] Figure 3is a flowchart of another chip interconnection communication method based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention;
[0033] Figure 4 is a schematic diagram of an example of interconnected communication of computing power chips provided by some embodiments of the present invention;
[0034] Figure 5 It is a structural block diagram of a chip interconnection communication device based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention. DETAILED DESCRIPTION
[0035] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] In the related art, a system for high-performance computing is usually composed of hundreds to thousands of computing chips (i.e., a computing chip cluster system). These chips work together through a high-speed interconnection network to achieve parallel processing of complex tasks. However, in a system for high-performance computing, the efficiency of interconnection and communication between computing chips is poor, which in turn affects the overall performance of high-performance computing, as follows:
[0037] 1. The exponential growth of computing power demand. Applications such as deep learning training tasks, weather simulation, and genetic analysis require the processing of massive amounts of data, which places extremely high demands on computing resources. For example, the training of large language models often involves tens of billions of parameters, and the amount of data exchange generated during the training process far exceeds the scale of traditional applications. This increase in demand requires not only higher performance of computing chips, but also greater bandwidth and lower latency from the Internet.
[0038] 2. The continuous expansion of system scale. As the complexity of computing tasks increases, the number of computing chips in HPC systems increases year by year. From small clusters of dozens of nodes to large supercomputers with millions of cores, the expansion of scale also brings the complexity of communication management. For example, in large-scale clusters, the connection relationship between computing chips grows exponentially, and the impact of communication bottlenecks becomes more and more significant.
[0039] 3. The popularity of heterogeneous computing has also brought new challenges to high-performance computing. In modern HPC systems, the collaborative work of heterogeneous computing chips such as CPU, GPU, DPU, and TPU has become the mainstream, but due to the large differences in the architecture of these chips, their interconnection and communication modes and requirements are also different, making it difficult for related communication technologies to meet complex heterogeneous computing scenarios.
[0040] In practical applications, the interconnection and communication between computing chips is the basis of HPC system performance, and the communication performance bottleneck has become a key issue restricting the expansion of computing power, as follows:
[0041] 1. Insufficient bandwidth: The amount of data exchange between computing chips increases rapidly as the scale of tasks increases. However, the Internet network architecture is limited by physical bandwidth and topology and cannot meet this growing demand. For example, although interconnect technologies such as PCIe and InfiniBand have powerful performance, they still have insufficient bandwidth in extremely data-intensive applications, resulting in data transmission delays and reduced throughput.
[0042] 2. Communication delay: Communication delay is determined by multiple factors such as transmission path length, network congestion, and protocol processing time. In multi-hop communication scenarios, delays accumulate rapidly as the number of nodes and topological complexity increase, especially in high-concurrency communications in large clusters. This delay not only affects the completion time of the task, but also causes load imbalance in the entire system.
[0043] 3. Unbalanced load in task distribution: In HPC systems, the loads of different computing chips may vary significantly due to uneven task distribution. Some chips are overloaded by data-intensive tasks, resulting in increased latency and error rates, while other chips may be in a state of low utilization. Unbalanced load not only wastes hardware resources, but also further aggravates the congestion of communication networks.
[0044] 4. Compatibility issues of heterogeneous communications: In heterogeneous computing scenarios, different types of computing chips (such as GPU and DPU) have different requirements for communication bandwidth, latency, and protocols. It is difficult for related communication technologies to meet these requirements at the same time, resulting in reduced system performance. For example, GPUs usually require low-latency, high-frequency communications, while DPUs are more concerned with high-throughput transmission of large amounts of data.
[0045] In related communication technologies (such as RDMA, InfiniBand, RoCE, etc.), although some improvement solutions have been proposed for the above problems, there are still obvious shortcomings, as follows:
[0046] 1. Limitations of static path scheduling: Most communication path scheduling solutions in related technologies use static configuration and cannot dynamically adjust the data flow according to the actual network status. This static scheduling method has acceptable performance in low-load scenarios, but will cause serious communication bottlenecks in high-load or dynamically changing scenarios. For example, when a path is congested, the static scheduling solution cannot effectively switch the data flow to other paths, which ultimately leads to a decrease in network utilization.
[0047] 2. Imbalance between performance and complexity: High-performance scheduling algorithms can significantly optimize communication paths, but often at the expense of complex computational processes. These algorithms require a large amount of hardware resources to support path calculation and task allocation, and the consumption of these resources directly affects the scalability and overall efficiency of the system. When the number of computing chips reaches thousands or even tens of thousands, the cost of complexity becomes unacceptable.
[0048] 3. Lack of support for heterogeneous computing: Related technologies are mainly designed for homogeneous computing environments and lack the ability to optimize heterogeneous computing scenarios. In heterogeneous scenarios, the communication requirements and performance indicators between computing chips vary significantly, and the unified scheduling scheme in related technologies cannot be effectively adapted, resulting in reduced communication performance. For example, the same task may require coordination of the CPU and GPU at the same time, but existing technologies often cannot provide the best path selection for heterogeneous tasks.
[0049] 4. Insufficient load balancing capabilities: The load balancing algorithm has limited effect in high-concurrency scenarios, and can often only allocate tasks based on static indicators, and cannot monitor and adjust the load status of computing chips in real time. This deficiency causes some chips to be overloaded for a long time, while other chip resources are idle, seriously affecting the overall efficiency of the system.
[0050] In summary, the current development of high-performance computing has put forward higher requirements for the interconnection and communication of computing chips, but there are still major deficiencies in bandwidth utilization, communication delay, load balancing and heterogeneous support.
[0051] In an embodiment of the present invention, a new communication technology is proposed, which realizes efficient optimization of the interconnection and communication of computing chips through an adaptive algorithm that balances performance and complexity, and can achieve a dynamic balance between performance and complexity to adapt to the needs of different application scenarios and complex network environments.
[0052] In the process of computing chip interconnection and communication, the selection of communication paths and task scheduling need to meet the two goals of high performance and low complexity at the same time, and these two goals usually restrict each other. Specifically, in the embodiment of the present invention, they mainly include the following:
[0053] 1. Adaptive scheduling mechanism: Based on the real-time communication status and resource occupancy, the present invention dynamically monitors system bottlenecks and performance requirements, and realizes dynamic adjustment of communication paths and data scheduling through adaptive algorithms to avoid excessive resource consumption and achieve the optimal balance between performance and complexity.
[0054] 2. Dynamic complexity constraint control: The complexity constraint mechanism aims to solve the resource consumption and performance bottleneck problems caused by high algorithm complexity in high-performance communication. In particular, in high-load scenarios, system resources are easily occupied by complex algorithms, thus affecting the overall communication efficiency and stability. This mechanism ensures that communication tasks can be executed efficiently under different system states by dynamically adjusting and optimizing algorithm complexity, while avoiding excessive consumption of computing resources.
[0055] Specifically, the complexity constraint mechanism first performs dynamic constraints based on task characteristics and communication requirements, and adaptively adjusts the upper limit of algorithm complexity by real-time monitoring of data scale, traffic distribution, and task type. In low-load scenarios, high-complexity algorithms are allowed to prioritize performance; in high-load or resource-constrained situations, the system will simplify the algorithm path and reduce computational complexity to ensure system stability and communication continuity.
[0056] To achieve refined control, the complexity constraint mechanism introduces a hierarchical complexity control strategy, which divides the algorithm complexity into three levels: basic, optimized, and enhanced. When resources are limited or the load is high, the system will adopt a basic level solution to simplify the path scheduling and calculation process, giving priority to basic communication performance; in medium-load scenarios, it will combine communication performance and computing resource allocation for moderate optimization; and in low-load scenarios, it will adopt the optimal algorithm strategy to give full play to the performance of hardware resources and improve communication efficiency. In addition, to ensure adaptability and real-time performance, the complexity constraint mechanism introduces a feedback control mechanism to collect performance indicators in the communication process in real time, such as latency, bandwidth occupancy, and resource consumption, to form a closed-loop feedback control. When performance degradation is detected or resources exceed the threshold, the system will dynamically reduce the algorithm complexity; and when resources are abundant, the complexity constraints will be appropriately relaxed to achieve a dynamic balance between performance and complexity.
[0057] On this basis, the mechanism also effectively coordinates system resource allocation and algorithm execution efficiency to avoid computing bottlenecks through joint optimization of computational complexity and resource allocation. In addition, based on the dynamic complexity model, the complexity constraint mechanism constructs a communication load, resource occupancy and complexity-performance mapping model, evaluates system status and task requirements in real time, optimizes complexity constraint strategies, and ensures that the system achieves optimal performance and resource utilization in different scenarios. In summary, the complexity constraint mechanism effectively solves the performance and resource occupancy problems caused by high-complexity algorithms through dynamic optimization and feedback control, ensuring the efficient and stable operation of the system in large-scale high-load communication scenarios.
[0058] In the specific implementation, Figure 1 , the embodiments of the present invention include the following:
[0059] Real-time monitoring module, specifically used for:
[0060] Monitor communication network bandwidth utilization, data transmission delay, task load status and other information.
[0061] Output monitoring results and provide input data for scheduling algorithms.
[0062] Adaptive scheduling algorithm module, specifically used for:
[0063] Based on real-time monitoring data, the optimal data transmission path is dynamically selected.
[0064] An adaptive algorithm is used to balance communication performance and scheduling complexity according to the system status.
[0065] Complexity constraint control module, specifically used for:
[0066] A complexity constraint mechanism is introduced in the scheduling process to dynamically control the algorithm calculation amount and reduce scheduling delays and resource consumption.
[0067] Dynamic feedback and optimization module, specifically used for:
[0068] According to communication performance indicators and system resource status, the scheduling strategy is dynamically adjusted to achieve closed-loop optimization.
[0069] Load balancing and resource management module, specifically used for:
[0070] Optimize data flow distribution, achieve load balancing among computing chips, and maximize resource utilization.
[0071] Chip interconnect communication interface, specifically used for:
[0072] Provides data transmission channels between computing chips and supports a variety of interconnection topologies.
[0073] The following further describes the embodiments of the present invention:
[0074] Reference Figure 2 , showing a step flow chart of a chip interconnection communication method based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention. The method can be applied to a computing power chip cluster system, wherein the computing power chip cluster system includes multiple computing power chip nodes, and the multiple computing power chip nodes can be interconnected and communicated with each other.
[0075] Specifically, the following steps may be included:
[0076] Step 201, obtain a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes.
[0077] Step 202, determining the communication performance data and computational complexity data corresponding to the scheduling strategy, adjusting the scheduling strategy according to the communication performance data and computational complexity data, and processing the data task according to the adjusted scheduling strategy.
[0078] In the embodiment of the present invention, a system performance evaluation model is established based on real-time monitoring data (such as bandwidth utilization, communication delay, and task load) to calculate the performance indicators of the current communication network, and the complexity constraints of the scheduling algorithm can be defined, including:
[0079] Computation threshold: Limit the number of computational steps performed by the algorithm to avoid resource overload.
[0080] Power consumption threshold: controls the power consumption during algorithm execution to adapt to the energy consumption requirements of different hardware platforms.
[0081] In practical applications, according to the real-time status of the system and application requirements, the scheduling strategy is dynamically adjusted between performance evaluation and complexity constraints to select the optimal communication path and task allocation solution.
[0082] In some embodiments of the present invention, determining the communication performance data and computational complexity data corresponding to the scheduling strategy includes: obtaining real-time monitoring data of the multiple target computing power chip nodes; and determining the communication performance data corresponding to the scheduling strategy based on the real-time monitoring data.
[0083] In some examples, the communication performance data includes any one or more of the following: bandwidth utilization, average communication delay, and load balancing of task allocation.
[0084] In actual applications, the communication data between computing chips is collected through the real-time monitoring module, including: the bandwidth utilization rate of each node, the delay of data transmission, and the load distribution of the current task.
[0085] After obtaining the real-time monitoring data, the monitoring data can be sent to the adaptive scheduling algorithm module for analysis. In the adaptive scheduling algorithm module, the performance indicators of the current system (i.e., communication performance data) are calculated, including: bandwidth utilization, average communication delay, and load balancing degree of task allocation.
[0086] In some embodiments of the present invention, communication performance data and computational complexity data corresponding to the scheduling strategy are determined. In some examples, this includes: determining the computational amount and computational power consumption corresponding to the scheduling strategy; and determining the computational complexity data of the scheduling strategy based on the computational amount and computational power consumption.
[0087] In the adaptive scheduling algorithm module, the complexity of the scheduling algorithm (i.e., computational complexity data) can also be calculated. In some examples, it includes: computational amount: the computational overhead of the current scheduling strategy, computational power consumption: the energy consumption during algorithm execution.
[0088] In some embodiments of the present invention, the scheduling strategy is adjusted based on the comparison result of the communication performance data and the communication performance threshold, and the comparison result of the computational complexity data and the computational complexity threshold, including: when the communication performance data is greater than the communication performance threshold, updating the information of multiple target computing power chip nodes and the communication paths between the multiple target computing power chip nodes; when the computational complexity data is greater than the computational complexity threshold, adjusting the scheduling algorithm.
[0089] During the system initialization phase, you can set complexity constraint parameters such as bandwidth threshold, delay threshold, computation threshold, and power consumption threshold. In actual applications, if the performance index reaches the preset standard and the complexity does not exceed the threshold, the current scheduling strategy is executed. If the complexity exceeds the threshold, the scheduling algorithm is simplified.
[0090] In some examples, the following operations are also included: Path compression: merging redundant data paths to reduce the amount of calculation. Task batching: merging multiple small tasks into a large task to reduce the number of scheduling times.
[0091] After obtaining the optimized scheduling strategy, the execution results and performance indicators are returned to the scheduling algorithm module through a dynamic feedback mechanism for adaptive adjustment.
[0092] In some embodiments of the present invention, it also includes:
[0093] The computational complexity threshold is adjusted according to the communication performance data.
[0094] In actual applications, if the performance does not meet the standard, adjust the complexity constraint parameters (i.e., the calculation complexity threshold) and re-execute the optimization. If the performance and complexity are balanced, save the current scheduling strategy, enter the next monitoring cycle, and wait for the next monitoring data input.
[0095] In an embodiment of the present invention, a data task is obtained in a computing chip cluster system, a scheduling strategy for the data task is generated, and the data task is processed according to the scheduling strategy; the scheduling strategy includes information of multiple target computing chip nodes for processing data tasks and communication path information between multiple target computing chip nodes, and then the communication performance data and computational complexity data corresponding to the scheduling strategy are determined, the scheduling strategy is adjusted according to the communication performance data and the computational complexity data, and the data task is processed according to the adjusted scheduling strategy, thereby realizing adaptive adjustment of the scheduling strategy for indicating the communication path between computing chip nodes, achieving a dynamic balance between performance and complexity to adapt to the needs of different application scenarios and complex network environments, improving the efficiency of interconnection and communication between computing chips, and thereby improving the overall performance of high-performance computing.
[0096] Reference Figure 3 , shows a flowchart of another chip interconnection communication method based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention, which is applied to a computing chip cluster system, wherein the computing chip cluster system includes multiple computing chip nodes, and specifically may include the following steps:
[0097] Step 301, obtain a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes.
[0098] Step 302: obtain real-time monitoring data of the multiple target computing power chip nodes, and determine the communication performance data corresponding to the scheduling strategy based on the real-time monitoring data.
[0099] Step 303: determine the amount of calculation and the amount of calculation power consumption corresponding to the scheduling strategy, and determine the calculation complexity data of the scheduling strategy based on the amount of calculation and the amount of calculation power consumption.
[0100] Step 304: adjust the scheduling strategy according to the communication performance data and the computational complexity data, and process the data task according to the adjusted scheduling strategy.
[0101] The following is combined with Figure 4 The present invention is exemplified as follows:
[0102] Assume that a computing chip cluster system consists of 8 computing chip nodes (N1, N2, N3, N4, N5, N6, N7, N8) and uses a grid topology for communication. The system needs to perform large-scale data task transmission, and the tasks are assigned to multiple computing chip nodes, and data path scheduling and complexity optimization are performed through adaptive algorithms.
[0103] 1. Initial state
[0104] Data task A is initiated by node N1 and targets node N8.
[0105] Data task B is initiated by node N2 and its target is node N7.
[0106] The system detects that the current communication bandwidth load is high and the delay is gradually increasing.
[0107] 2. Real-time monitoring
[0108] Bandwidth utilization of the path from N1 to N8: 80%.
[0109] Bandwidth utilization of the path from N2 to N7: 85%.
[0110] Average communication delay: 3ms.
[0111] 3. Performance and complexity evaluation
[0112] The current communication performance is not optimal and the path is obviously congested.
[0113] The scheduling algorithm has a high computational complexity and the complexity constraint is greater than the threshold.
[0114] 4. Adaptive path scheduling and optimization
[0115] Adjust the path:
[0116] New path for data task A: N1 via N3 to N8.
[0117] New path for data task B: N2 via N4 via N6 to N7.
[0118] Simplify strategy:
[0119] Adopt path compression mechanism to merge redundant nodes.
[0120] Batch processing of small data packet tasks can reduce the number of path calculations.
[0121] 5. Optimize results
[0122] Bandwidth utilization is reduced to 60%.
[0123] The average communication delay is reduced to 1.5ms.
[0124] The complexity of the scheduling algorithm is reduced and the complexity constraints are met.
[0125] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0126] Reference Figure 5 , shows a schematic diagram of the structure of a chip interconnection communication device based on a performance and complexity balance adaptive algorithm provided by some embodiments of the present invention, which is applied to a computing chip cluster system. The computing chip cluster system includes multiple computing chip nodes, which may specifically include the following modules:
[0127] The scheduling strategy generation and execution module 501 is used to obtain a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes;
[0128] The scheduling strategy adjustment and execution module 502 determines the communication performance data and computational complexity data corresponding to the scheduling strategy, adjusts the scheduling strategy according to the communication performance data and computational complexity data, and processes the data task according to the adjusted scheduling strategy.
[0129] Optionally, the determining the communication performance data and the computational complexity data corresponding to the scheduling strategy includes:
[0130] Obtaining real-time monitoring data of the multiple target computing power chip nodes;
[0131] According to the real-time monitoring data, communication performance data corresponding to the scheduling strategy is determined.
[0132] Optionally, the communication performance data includes any one or more of the following: bandwidth utilization, average communication delay, and load balancing degree of task allocation.
[0133] Optionally, the determining the communication performance data and the computational complexity data corresponding to the scheduling strategy includes:
[0134] Determine the computational effort and computational power consumption corresponding to the scheduling strategy;
[0135] The computational complexity data of the scheduling strategy is determined according to the computational amount and the computational power consumption.
[0136] Optionally, adjusting the scheduling strategy according to the communication performance data and the computational complexity data includes:
[0137] Obtaining a preset communication performance threshold and a computational complexity threshold;
[0138] The scheduling strategy is adjusted according to the comparison result between the communication performance data and the communication performance threshold and the comparison result between the calculation complexity data and the calculation complexity threshold.
[0139] Optionally, adjusting the scheduling strategy according to a comparison result between the communication performance data and the communication performance threshold and a comparison result between the computational complexity data and the computational complexity threshold includes:
[0140] When the communication performance data is greater than the communication performance threshold, update the information of multiple target computing power chip nodes and the communication paths between the multiple target computing power chip nodes;
[0141] When the computational complexity data is greater than a computational complexity threshold, the scheduling algorithm is adjusted.
[0142] Optionally, it also includes:
[0143] The computational complexity threshold is adjusted according to the communication performance data.
[0144] In an embodiment of the present invention, a data task is obtained in a computing chip cluster system, a scheduling strategy for the data task is generated, and the data task is processed according to the scheduling strategy; the scheduling strategy includes information of multiple target computing chip nodes for processing data tasks and communication path information between multiple target computing chip nodes, and then the communication performance data and computational complexity data corresponding to the scheduling strategy are determined, the scheduling strategy is adjusted according to the communication performance data and the computational complexity data, and the data task is processed according to the adjusted scheduling strategy, thereby realizing adaptive adjustment of the scheduling strategy for indicating the communication path between computing chip nodes, achieving a dynamic balance between performance and complexity to adapt to the needs of different application scenarios and complex network environments, improving the efficiency of interconnection and communication between computing chips, and thereby improving the overall performance of high-performance computing.
[0145] Some embodiments of the present invention further provide an electronic device, comprising a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the above method is implemented when the computer program is executed by the processor.
[0146] Some embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored, and the computer program implements the above method when executed by a processor.
[0147] Some embodiments of the present invention further provide a computer program product, including a computer program, which implements the above method when executed by a processor.
[0148] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0149] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0150] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0151] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0152] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0153] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0155] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0156] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the above elements.
[0157] The above is a detailed introduction to the chip interconnection communication method, device, electronic device and medium based on the performance and complexity balance adaptive algorithm. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A chip interconnection communication method based on a performance and complexity balance adaptive algorithm, characterized in that: Applied to a computing chip cluster system, the computing chip cluster system includes multiple computing chip nodes, and the method includes: Acquire a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes; Determine the communication performance data and computational complexity data corresponding to the scheduling strategy, adjust the scheduling strategy according to the communication performance data and computational complexity data, and process the data task according to the adjusted scheduling strategy.
2. The method according to claim 1, characterized in that The determining of the communication performance data and the computational complexity data corresponding to the scheduling strategy includes: Obtaining real-time monitoring data of the multiple target computing power chip nodes; According to the real-time monitoring data, communication performance data corresponding to the scheduling strategy is determined.
3. The method according to claim 2, characterized in that The communication performance data includes any one or more of the following: bandwidth utilization, average communication delay, and load balancing degree of task allocation.
4. The method according to claim 1, characterized in that: The determining of the communication performance data and the computational complexity data corresponding to the scheduling strategy includes: Determine the computational effort and computational power consumption corresponding to the scheduling strategy; The computational complexity data of the scheduling strategy is determined according to the computational amount and the computational power consumption.
5. The method according to any one of claims 1 to 4, characterized in that: The adjusting the scheduling strategy according to the communication performance data and the computational complexity data includes: Obtaining a preset communication performance threshold and a computational complexity threshold; The scheduling strategy is adjusted according to the comparison result between the communication performance data and the communication performance threshold and the comparison result between the calculation complexity data and the calculation complexity threshold.
6. The method according to claim 5, characterized in that The adjusting the scheduling strategy according to the comparison result between the communication performance data and the communication performance threshold and the comparison result between the computational complexity data and the computational complexity threshold includes: When the communication performance data is greater than the communication performance threshold, update the information of multiple target computing power chip nodes and the communication paths between the multiple target computing power chip nodes; When the computational complexity data is greater than a computational complexity threshold, the scheduling algorithm is adjusted.
7. The method according to claim 5, characterized in that Also includes: The computational complexity threshold is adjusted according to the communication performance data.
8. A chip interconnection communication device based on a performance and complexity balance adaptive algorithm, characterized in that: Applied to a computing chip cluster system, the computing chip cluster system includes multiple computing chip nodes, and the device includes: A scheduling strategy generation and execution module, used to obtain a data task, generate a scheduling strategy for the data task, and process the data task according to the scheduling strategy; the scheduling strategy includes information of multiple target computing power chip nodes for processing the data task and communication path information between the multiple target computing power chip nodes; The scheduling strategy adjustment and execution module determines the communication performance data and computational complexity data corresponding to the scheduling strategy, adjusts the scheduling strategy according to the communication performance data and computational complexity data, and processes the data task according to the adjusted scheduling strategy.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.