Switching network, online computing method, electronic equipment and storage medium
By introducing dynamic topology and central management nodes into the switching network, the packet parallelization and pipelined processing of switching nodes are solved, and the computing efficiency and scalability of the switching network are improved.
Patent Information
- Application Number
- CN202510771781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-12
AI Technical Summary
When performing on-network computing, existing switching networks have problems such as serious competition in cache resources, insufficient scalability of hardware architecture, and uncontrollable computing delays, which are difficult to meet the needs of large-scale computing tasks.
Design a switching network, including multiple switching nodes, each node includes on-network computing proxy nodes. Through the coordination of dynamic topology and central management nodes, it realizes packet parallelization and pipelined processing of computing tasks, reduces cache resource competition, improves computing efficiency, supports elastic expansion, and reduces computing delay.
It effectively solves the problems of cache resource competition and computing latency, improves the throughput and scalability of the switching network, reduces hardware costs and complexity, and realizes efficient processing of computing tasks.
Smart Images

Figure CN120474866A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to a switching network, an on-line computing method for the switching network, an electronic device, and a non-transitory computer-readable storage medium. Background Art
[0002] In-network computing (IOC) leverages network devices like network cards and switches to perform online computations simultaneously with data transmission. In data centers with distributed processing architectures, IOC can become a core technology that influences overall system balance and plays a crucial role in improving data center performance and efficiency. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a switching network, comprising: a plurality of switching nodes connected to each other to provide a communication network, wherein each of the plurality of switching nodes includes an on-line computing agent node, and the on-line computing agent node is configured to receive and execute on-line computing tasks.
[0004] For example, in a switching network provided in an embodiment of the present disclosure, the multiple switching nodes are configured as multiple switching groups; each switching group includes at least two switching nodes and is configured to aggregate the on-line computing results of each on-line computing agent node in the switching group.
[0005] For example, in a switching network provided in an embodiment of the present disclosure, the switching network further includes a central management node, wherein the central management node is configured to receive instructions for the on-line computing task and assign the on-line computing task to a target switching group among the multiple switching groups; the multiple switching groups are also communicatively connected through the central management node.
[0006] For example, in the switching network provided by an embodiment of the present disclosure, the on-line computing agent nodes of at least two switching nodes of each switching group are connected sequentially through the communication link within the switching group, and are configured to sequentially execute the on-line computing tasks and transmit the on-line computing results.
[0007] For example, in the switching network provided in one embodiment of the present disclosure, the multiple switching nodes are also configured to respectively communicate with and connect to multiple data processing units; each of the on-line computing agent nodes is further configured to, in response to executing the on-line computing task, initiate an on-line computing data request or return an on-line computing result to the corresponding data processing unit to which the communication connection is connected.
[0008] For example, in the switching network provided in an embodiment of the present disclosure, the on-line computing agent node is further configured to receive the on-line computing data provided by the corresponding data processing unit, and use the on-line computing data to execute the on-line computing task.
[0009] For example, in a switching network provided in an embodiment of the present disclosure, the multiple switching nodes include a first switching node and a second switching node, the first switching node includes a first on-line computing agent node, the second switching node includes a second on-line computing agent node, and the second on-line computing agent node is configured to receive intermediate computing results from the first on-line computing agent node and use the intermediate computing results to perform the on-line computing task.
[0010] For example, in a switching network provided in an embodiment of the present disclosure, each of the multiple switching nodes includes a first cache unit, the first cache unit includes a first storage space and a second storage space, the first storage space is configured to store communication data transmitted by the communication network, and the second storage space is configured to store computing data of the on-line computing task.
[0011] For example, in the switching network provided in an embodiment of the present disclosure, the on-line computing agent node of each of the multiple switching nodes is further configured to initiate an on-line computing data request to the corresponding data processing unit based on the capacity of the second storage space of the switching node.
[0012] For example, in the switching network provided in an embodiment of the present disclosure, each of the online computing agent nodes is further configured to receive at least part of the code of the online computing software program corresponding to the online computing task to execute the online computing task.
[0013] At least one embodiment of the present disclosure provides a method for on-line computing of a switching network, the method comprising: in response to respectively setting up on-line computing agent nodes on a plurality of switching nodes of the switching network, receiving and executing on-line computing tasks through the on-line computing agent nodes of the plurality of switching nodes, wherein the plurality of switching nodes are communicatively connected to each other to provide a communication network.
[0014] At least one embodiment of the present disclosure provides an electronic device, including the switching network provided by any embodiment of the present disclosure.
[0015] At least one embodiment of the present disclosure provides an electronic device, comprising the switching network provided by any embodiment of the present disclosure and the multiple data processing units, wherein each of the multiple data processing units is configured to respond to the network computing data request or receive the network computing result.
[0016] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium for non-temporarily storing computer-readable instructions. When the computer-readable instructions are executed by at least one processor, the method for on-line computing for a switching network provided by any embodiment of the present disclosure is implemented.
[0017] At least one embodiment of the present disclosure provides an electronic device, comprising: at least one processor and a memory, wherein the memory stores at least one computer program, and when the at least one computer program is executed by the at least one processor, the method for on-line computing for a switching network provided in any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0019] Figure 1 A schematic block diagram of a switching network is shown;
[0020] Figure 2 shows a schematic block diagram of another switching network;
[0021] Figure 3 shows a schematic block diagram of yet another switching network;
[0022] Figure 4 A schematic block diagram of a switching network provided by at least one embodiment of the present disclosure is shown;
[0023] Figure 5 A schematic block diagram of another switching network provided by at least one embodiment of the present disclosure is shown;
[0024] Figure 6 A schematic block diagram of another switching network provided by at least one embodiment of the present disclosure is shown;
[0025] Figure 7 A schematic block diagram of another switching network provided by at least one embodiment of the present disclosure is shown;
[0026] Figure 8 A schematic block diagram of an application of a switching network provided by at least one embodiment of the present disclosure is shown;
[0027] Figure 9 A schematic block diagram of another switching network application provided by at least one embodiment of the present disclosure is shown;
[0028] Figure 10 A flowchart of a method for on-line computing in a switching network provided by at least one embodiment of the present disclosure is shown;
[0029] Figure 11 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown;
[0030] Figure 12 A schematic block diagram showing another electronic device provided by at least one embodiment of the present disclosure; and
[0031] Figure 13 A schematic diagram of a computer-readable storage medium provided by at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0032] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0033] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0034] In-Network Computing (INC) is a computing logic unit embedded in network devices, used, for example, to perform data aggregation operations. INC embeds computing tasks into the data transmission process of network devices (such as switches and network interface cards), allowing network devices to not only handle data transmission but also directly perform some computing operations.
[0035] The types of INC data involved in online computing can include various types, and their characteristics are that they need to be intermediately calculated through network devices rather than just end-to-end transmission.
[0036] For example, AllReduce is a collective communication operation that aggregates data from multiple computing nodes and distributes it to all nodes. AllReduce can be used for parameter synchronization in distributed AI training. In-network computing can serve parameter synchronization (such as AllReduce operations) in AI (Artificial Intelligence) training or large-scale parallel computing tasks in high-performance computing (HPC).
[0037] Figure 1 A schematic block diagram of a switching network is shown.
[0038] like Figure 1 As shown, the switching network includes multiple ( Figure 1 40 switches are shown as switching nodes. These switching nodes are connected to each other through a certain network structure (such as a star network, a ring network, etc.). Each switch includes a buffer unit (BUFFER) for buffering the transmitted information (such as data or instructions, etc.). Multiple data processing units (DCUs) communicate with each other through the switching network. Figure 1 40 data processing units are shown as an example in FIG. 4 . For example, each data processing unit is connected to a switching network through a corresponding switch, thereby sending or receiving information during operation.
[0039] For example, during work, each data processing unit (DCU) used for distributed in-network computing actively initiates calculation data. The DCU can actively send INC data (INC DAT) to the switch. The switch's cache unit caches the INC data of all DCUs and performs calculations (such as aggregation operations) after all data arrives. After completion, the cache is released and the results are forwarded.
[0040] For example, a DCU can push INC data to the switch's cache. The switch then checks to see if it has received data from all relevant DCUs. Once the data is complete, the DCU performs calculations and sends the results to the destination node. For example, the cache can also store other data (OTHER DAT) for other purposes. For example, an in-network calculation software program (INC PROGRAM) can control each DCU to perform corresponding INC calculations. For example, the switch's internal INC calculations can depend on whether INC data from the DCUs has been received. If INC data stored in DCU0 and DCU1 is needed for INC calculations, and DCU1 data is not received, DCU0's data is stored in the cache until it is received. When data from both DCU0 and DCU1 is received, the corresponding data is read for calculations, freeing up storage space in the buffer.
[0041] However, the inventors of this disclosure noted that due to the asynchronous nature of DCU data transmission (e.g., network congestion causing some data delays), switches must cache already received INC data for a long time while waiting for other data to become available. This results in significant cache resource usage and impacts the transmission performance of non-INC services. Furthermore, in collective communication scenarios such as AllReduce, if the number of participating DCUs increases (e.g., 40 DCUs), cache resource contention is further exacerbated.
[0042] Figure 2 A schematic block diagram of another switching network is shown.
[0043] like Figure 2 As shown, the switching network comprises multiple switching nodes and adopts a tree-like topology (e.g., a multi-stage pipeline). Each switching node includes a computing unit (e.g., an arithmetic logic unit (ALU)). Each DCU accesses the switching network through a corresponding switching node. For example, in the tree-like on-network computing architecture described above, computing tasks are processed hierarchically by the switching nodes. For example, a switch with 16 DCU ports divides computing into four levels. At each level, aggregation operations are performed on two adjacent switching nodes using computing units, with the results aggregated upwards through each level.
[0044] However, the inventors of the present disclosure have noted that a tree-structured switching network requires computing units to be distributed hierarchically within the switching network (e.g., a network-on-chip (NoC) within a chip) for on-network computing. This results in high complexity in back-end layout and routing, and as the number of computing levels increases, the node spacing increases, making it difficult to meet timing closure requirements. Furthermore, the fixed hierarchical design of the tree architecture is difficult to adapt to dynamically changing computing task requirements. For example, the tree structure requires computing units to be laid out hierarchically within the chip, resulting in increased physical wiring length and difficulty in signal timing convergence. Especially when the number of computing levels exceeds four, it may also exacerbate the clock skew problem.
[0045] Figure 3 A schematic block diagram of yet another switching network is shown.
[0046] like Figure 3 As shown, the switching network includes N agent nodes (e.g., INCA). The figure illustrates 16 INCAs (e.g., INC-A0 through INC-A15). Each INCA corresponds to a data processing unit (DCU), and the figure illustrates 16 DCUs (e.g., DCU0 through DCU15). For example, INC-A0 corresponds to DCU0, INC-A1 to DCU1, and so on. INC-A15 corresponds to DCU15. All INCA nodes in the switching network are connected sequentially via bidirectional or unidirectional links, forming a ring-shaped data path.
[0047] For example, in a ring-structured in-network computing architecture, a ring data path can be built inside the switch. Each agent node (such as INC-A) can be configured to aggregate the calculation results of its corresponding DCU and the output data of the previous-level agent node, and achieve full calculation through loop transmission.
[0048] For example, INC-A1 aggregates the calculation result D0 of DCU0 and the calculation result D1 of DCU1, and provides the aggregated result (e.g., reduce (D0, D1)) to INC-A2; INC-A2 aggregates the calculation result D2 of DCU2 and provides the aggregated result (e.g., reduce (D0, D1, D2)) to INC-A3, and so on, until INC-A15 aggregates the calculation results (e.g., reduce (D0, D1, D2…D15)).
[0049] However, the inventors of this disclosure noted that computational latency in a ring architecture increases linearly with scale. Data must traverse all nodes to complete computation. As the number of nodes increases (e.g., 16 DCUs), computational latency increases significantly. Furthermore, the serial nature of the ring path can cause some computing units to remain idle, preventing full utilization of parallel computing capabilities.
[0050] Embodiments of the present disclosure provide a switching network, an on-line computing method for the switching network, an electronic device, and a non-transitory computer-readable storage medium.
[0051] The switching network includes a plurality of switching nodes connected to each other to provide a communication network, wherein each of the plurality of switching nodes includes an on-line computing agent node configured to receive and execute an on-line computing task.
[0052] In at least one embodiment of the present disclosure, the switching network can reduce cache resource competition and improve the coexistence performance of INC and non-INC services. It can also break through the physical limitations of tree and ring architectures and support elastic expansion. It can also significantly reduce computing latency through parallelization and pipeline design.
[0053] For example, in at least one embodiment of the present disclosure, a switching network can solve the problem of inefficient cache resources, such as cache occupancy caused by asynchronous data arrival, which affects the overall switching network throughput.
[0054] For example, in at least one embodiment of the present disclosure, the switching network can also solve the problem of insufficient scalability of the hardware architecture, such as the physical limitations of the tree and ring structures making it difficult to support large-scale computing tasks.
[0055] For example, in at least one embodiment of the present disclosure, the switching network can also solve the problem of uncontrollable computing delay and the inability of the architecture to balance the contradiction between computing efficiency and node scale.
[0056] Figure 4 A schematic block diagram of a switching network provided by at least one embodiment of the present disclosure is shown.
[0057] In some embodiments of the present disclosure, the switching network 1000 includes a plurality of switching nodes connected to each other to provide a communication network. Each of the plurality of switching nodes includes an on-line computing agent node configured to receive and execute on-line computing tasks.
[0058] like Figure 4As shown, the multiple switching nodes include switching node 110, switching node 120, switching node 130, etc. The number of switching nodes can be dynamically increased or decreased according to actual needs (for example, expanded from 3 to N), and the network topology automatically adapts to node changes to achieve flexible allocation of computing and communication resources. The embodiments of the present disclosure do not limit the number of switching nodes.
[0059] For example, a switching node may comprise a node in a Network-on-Chip (NOC), enabling high-speed data processing through high-density interconnection within the chip. For example, a switching node may also comprise a switch or gateway, integrating INC proxy node functionality through hardware (e.g., integrated dedicated computing chips) or software-defined methods to provide on-network computing capabilities. The embodiments of this disclosure do not limit the implementation of switching nodes.
[0060] For example, the switching network 1000 may include a network-on-chip (NOC), which is suitable for low-latency computing and communication scenarios at the chip level. For example, the switching network 1000 may also include a distributed network composed of multiple interconnected on-chip computing agent nodes (e.g., switches or gateways) that support INC functions.
[0061] For example, each in-network computing proxy node may include a processing unit (e.g., an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA), which may have a built-in computing unit such as an arithmetic logic unit (ALU)) and a cache unit (e.g., SRAM, DRAM, etc.). The processing unit may be configured to execute INC computing tasks, and the cache unit may be configured to store temporary data (e.g., pending INC data, intermediate computation results, etc.), enabling temporary storage and smooth transmission of data before and after computation. For example, switching node 110 and its INC proxy node 111 may process INC data using an internal processing unit, while the cache unit may store raw INC data, working in conjunction with the processing unit to achieve an efficient process for data acquisition, computation, and forwarding.
[0062] For example, the in-network computing agent node can also be at least partially implemented in software, and the switching node can have the corresponding code running capabilities. For example, each switching node can integrate a processor (such as a lightweight processor) or a programmable logic module to build a software running environment, which can run a matching operating system. At this time, the INC agent node can be expressed as a software program to monitor the data flow flowing through the corresponding switching node, identify the data packet carrying the INC task identifier (for example, a specific message header tag), parse the INC task instruction, and dynamically load and execute the computing logic code based on the running capability of the switching node. For example, Figure 4As shown, INC proxy node 111 of switch node 110 can monitor inputs through software code. After identifying INC data D0, it processes D0 according to the parsed instructions using the switch node's operational capabilities, then passes the intermediate results to INC proxy node 121 of switch node 120. INC proxy node 121 can then combine new data D1 with software to perform intermediate computations (e.g., reduce (D0, D1)). Finally, INC proxy node 131 of switch node 130 incorporates data D2 to complete the final computation (e.g., reduce (D0, D1, D2)). This allows INC proxy nodes to flexibly adapt to tasks, leveraging the switch node's underlying operational capabilities to achieve dynamic configuration and efficient execution of data processing, while further reducing reliance on high-performance processors and lowering hardware cost and complexity.
[0063] For example, the switching network 1000 may be composed of a plurality of switching nodes (such as switching nodes 110 , 120 , 130 , etc.) interconnected to form a distributed communication network.
[0064] For example, each switching node includes a corresponding INC agent node.
[0065] For example, the switching node 110 includes a corresponding INC proxy node 111 , the switching node 120 includes a corresponding INC proxy node 121 , and the switching node 130 includes a corresponding INC proxy node 130 .
[0066] For example, the INC agent node deployed in each switching node can be directly embedded in the data forwarding path, intercepting INC data packets flowing through the node in real time and executing predefined computing tasks (such as aggregation, reduction, or custom operators).
[0067] For example, an INC proxy node can include a dedicated computing unit (e.g., an arithmetic logic unit) integrated into the corresponding switching node and can be configured to perform distributed computing tasks (such as data aggregation, reduction, and filtering) in real time during network data transmission. INC proxy nodes can be configured to offload computing tasks from traditional computing nodes (such as CPUs, GPUs, or DCUs) to network devices through deep collaboration with network forwarding paths, reducing redundant data transmission between computing nodes and network devices, thereby lowering communication latency and improving the overall efficiency of the switching network.
[0068] For example, multiple switching nodes in the switching network 1000 may be divided into multiple switching stages, where each stage may represent a logical level in a data processing link.
[0069] For example, examples of switching stages include: switching node 110 (stage 1) can be configured to receive raw INC data from a data source (such as a DCU); switching node 120 (stage 2) can be configured to perform intermediate calculations (such as local aggregation) on the output of stage 1; switching node 130 (stage 3) can be configured to complete the final calculation and distribute the results to the destination node.
[0070] For example, the switching nodes at each stage work together in a pipelined manner. For example, after completing preliminary calculations, INC proxy node 111 in stage 1 passes the intermediate results to INC proxy node 121 in stage 2. Ultimately, INC proxy node 131 in stage 3 outputs the global calculation results. This enables hierarchical processing of computing tasks, reduces the load on individual nodes, and optimizes latency.
[0071] For example, Figure 4 As shown, the switching node 110 (stage 1) can receive the original INC data D0 from the data source (such as DCU0 corresponding to the switching node 110), and its INC proxy node 111 performs preliminary processing on D0 and then passes it to the switching node 120 (stage 2); after receiving D0, the INC proxy node 121 of the switching node 120 combines the new INC data D1 (for example, provided by DCU1 corresponding to the switching node 120) to perform intermediate calculations (such as reduction (D0, D1)); then the intermediate results are passed to the switching node 130 (stage 3), and its INC proxy node 131 then incorporates the INC data D2 (for example, provided by DCU2 corresponding to the switching node 130) to complete the final calculation reduction (D0, D1, D2), and distributes the results to the destination node, for example, to the next INC proxy node (for example, INC proxy node 141, not shown in the figure).
[0072] For example, each INC agent node can be configured to monitor the data flow passing through its corresponding switching node, identify data packets carrying INC task identifiers (such as specific message headers or metadata tags), parse INC task instructions (such as aggregation operation types and the list of DCUs involved in the calculation), and dynamically load the corresponding calculation logic.
[0073] For example, each INC proxy node can also be configured to perform calculations on INC data (such as summation, maximum value screening, etc., which are not limited in the embodiments of the present disclosure) and temporarily store the intermediate results in the corresponding cache unit.
[0074] For example, each INC proxy node can also be configured to retain necessary data only when the neighboring nodes are not ready, and release resources immediately after completion to avoid long-term occupancy.
[0075] For example, each INC proxy node can be configured to coordinate computation progress with other INC proxy nodes through a state synchronization protocol (e.g., a token-based handshake mechanism). For example, if downstream node congestion is detected (e.g., switch node 120 is overloaded), the data forwarding path can be dynamically adjusted (e.g., bypassing phase 2 and directly forwarding to phase 3), achieving adaptive computation routing.
[0076] In at least one embodiment of the present disclosure, by sharing computing tasks among nodes in multiple stages, the serial delay problem of the ring architecture can be solved, and the chip wiring complexity can be reduced.
[0077] Figure 5 A schematic block diagram of another switching network provided by at least one embodiment of the present disclosure is shown.
[0078] In some embodiments of the present disclosure, multiple switching nodes can be configured as multiple switching groups; each switching group includes at least two switching nodes, and each switching group can be configured to aggregate the on-line computing results of each on-line computing agent node within the switching group.
[0079] like Figure 5 As shown, in an embodiment of the present disclosure, multiple switching nodes in the switching network 1000 can be configured as multiple switching groups, each switching group includes at least two switching nodes, and each switching group can be configured to aggregate the calculation results of each INC agent node in the group.
[0080] For example, each switching group includes at least two switching nodes.
[0081] For example, the switch nodes included in each switch group can support intra-group parallel computing and redundant backup.
[0082] like Figure 5 As shown, for example, switching group A may include switching node 110 and switching node 120, and switching group B may include switching node 130 and switching node 140, etc.
[0083] For example, within a switching network, the division of switching groups can be adjusted based on network load, task type, or topological relationship (for example, nodes that communicate frequently can be grouped into the same group based on task relevance).
[0084] For example, the INC agent node 111 and the INC agent node 121 corresponding to each switching node in the switching group (such as the switching node 110 and the switching node 120 of the switching group A) can be configured to receive INC data from the upstream switching node, DCU or the central management node (to be described below) and perform local calculations.
[0085] In at least one embodiment of the present disclosure, processing parallelism is achieved through the grouping form of switch groups in a switching network, and large-scale computing tasks can be decomposed into multiple switch groups, thereby solving the serial delay problem of the ring architecture.
[0086] Figure 6 A schematic block diagram of another switching network provided by at least one embodiment of the present disclosure is shown.
[0087] like Figure 6 As shown, in some embodiments of the present disclosure, the switching network 1000 further includes a central management node 200 in addition to the switching nodes 110 - 140 .
[0088] The central management node 200 can be configured to receive instructions for online computing tasks and assign the online computing tasks to target switching groups among the multiple switching groups. For example, the multiple switching groups are also communicatively connected to each other via the central management node. For example, the central management node 200 can be implemented in the same manner as the switching nodes, and will not be further described here.
[0089] Here, the central management node 200 is the global control center of the switching network 1000 and can be configured to coordinate INC task allocation, resource scheduling and status monitoring of multiple switching groups.
[0090] For example, the central management node 200 can be configured to receive and parse INC tasks. For example, the central management node can be configured to receive online computing tasks (such as AllReduce and Broadcast) issued by an external system (such as a distributed training framework) and parse the task requirements (such as data type, participating DCU list, and computing type).
[0091] For example, the central management node may be configured to perform dynamic task allocation. For example, the central management node may be configured to split and allocate INC tasks to target switching groups based on network topology, switching group load, and / or task priority.
[0092] For example, the central management node can be configured to coordinate cross-group communications. For example, the central management node can be configured to manage data routing and synchronization between exchange groups to ensure global computing consistency.
[0093] For example, the central management node can collect the load status (such as cache usage, link bandwidth), topological location and computing power of each switching group in real time, and select the target switching group and / or split the task.
[0094] For example, target switch group selection can include proximity and / or load balancing. For example, proximity can include prioritizing switch groups directly connected to the data source DCU (e.g., DCU0-DCU3 connected to switch group A). For example, with load balancing, if a switch group is overloaded (e.g., the cache utilization rate of switch group B is greater than 80%), the task is assigned to an idle group (e.g., switch group C).
[0095] For example, a task splitting strategy can include splitting global tasks into local aggregation (performed in parallel by multiple exchange groups) and global aggregation (completed by the top exchange group). For example, exchange groups A-D process the local summation of DCUs 0-3, 4-7, 8-11, and 12-15 respectively; the central management node can aggregate the intermediate results of exchange groups A-D.
[0096] For example, the central management node may be configured to send task instructions to the target exchange group.
[0097] For example, a task instruction may include: calculation type (such as summation), data source address, result return path, dependency (such as exchange group A needs to wait for the intermediate result of exchange group B) or timeout constraints and retry strategies, etc. The embodiments of the present disclosure do not limit the composition of task instructions.
[0098] For example, the central management node can be configured to select low-latency or high-bandwidth paths for data flows based on real-time network status. For example, if an intermediate result from switch group A needs to be delivered to switch group X, it will be transferred through switch group Y if the direct link is congested.
[0099] For example, if a switching group goes offline (e.g., switching group C goes down), the central management node migrates its tasks to the standby switching group and updates the routing table.
[0100] In at least one embodiment of the present disclosure, centralized scheduling by a central management node can maximize the utilization of the computing power of the switching group, further reduce local hot spots, and thus overcome the problem of uneven load.
[0101] In some embodiments of the present disclosure, the on-line computing agent nodes of at least two switching nodes of each switching group are connected sequentially through a communication link within the switching group, and are configured to sequentially execute on-line computing tasks and transmit on-line computing results.
[0102] For example, at least two switch nodes within each switch group can be connected sequentially via a communication link within the group, forming a chain or ring data path. The on-network computing agent nodes included in each switch node can process and transmit data along this link in sequence, implementing phased execution of tasks.
[0103] For example, the connection mode of at least two switching nodes in each switching group may include a unidirectional chain or a bidirectional link.
[0104] For example, for a one-way chain, the on-network computing agent nodes included in each switching node can transmit data in a fixed direction (such as INC agent node 111 →INC agent node 121 →INC agent node 131...).
[0105] For example, for a bidirectional link, bidirectional communication can be supported between the in-network computing agent nodes included in each switching node (such as INC agent node 111↔INC agent node 121↔INC agent node 131...), and the transmission direction can be dynamically selected according to the load.
[0106] For example, links within a switching group can use low-latency direct connection channels (such as on-chip networks NoC or dedicated SerDes links) with bandwidth matching the computing task requirements (for example, 100Gbps / link).
[0107] For example, the entry node of the exchange group (such as INC agent node 111) receives INC task instructions (such as AllReduce aggregation) and parses task parameters (such as operation type and participating DCU list). For example, the central management node can split the global task into multiple sequential subtasks and assign them to each INC agent node in the group. Exemplarily, INC agent node 111 can be configured to receive raw data and perform first-level aggregation; INC agent node 121 can be configured to receive the intermediate results of node A and perform second-level aggregation; INC agent node 131 can be configured to complete the final aggregation and return the results to SM or other exchange groups.
[0108] For example, each INC agent node can be configured to first process the data of the locally bound DCU (such as INC agent node 121 receiving data of DCU1) and aggregate it with the intermediate results passed upstream (such as INC agent node 111).
[0109] For example, after completing local computation, each INC agent node can immediately encapsulate the intermediate results into a task data packet and push it to the next INC agent node through the intra-group link.
[0110] For example, the data packet format may include a packet header, a payload, or a check code, etc., which is not limited in the embodiments of the present disclosure.
[0111] For example, the packet header may include a task ID, an operation code (OP_AGGREGATE), or a sequence number, etc., which is not limited in the embodiments of the present disclosure.
[0112] For example, the load may include intermediate result data blocks, etc., which is not limited in the embodiments of the present disclosure.
[0113] For example, the check code may include a CRC or a hash value, etc., which is not limited in the embodiments of the present disclosure, and is used for integrity verification.
[0114] For example, the end node in the exchange group (such as INC agent node 131) can return the final result to the entry node or the central management node to complete a calculation cycle.
[0115] For example, if each exchange group has a ring topology, the calculation results of each exchange group can be circulated and transmitted multiple times within the exchange group to achieve iterative calculation (such as multiple reduction optimization).
[0116] In at least one embodiment of the present disclosure, through pipelined processing, multiple INC proxy nodes can execute tasks at different stages in parallel, further reducing overall latency. Furthermore, each INC proxy node's intermediate results can be passed to downstream nodes, and each INC proxy node's local cache only temporarily stores the currently processed data, avoiding long-term occupancy and further reducing cache contention.
[0117] Figure 7 A schematic block diagram of yet another switching network provided by at least one embodiment of the present disclosure is shown.
[0118] In some embodiments of the present disclosure, multiple switching nodes are also configured to communicate and connect to multiple data processing units respectively; each on-line computing agent node is further configured to initiate an on-line computing data request or return an on-line computing result to the corresponding data processing unit to which it is communicated in response to executing an on-line computing task.
[0119] like Figure 7 As shown, for example, each switching node of switching network 1000 can be configured to establish a corresponding communication connection with a data processing unit (e.g., DCU), forming a collaborative architecture of distributed data processing units and network devices. For example, switching node 110 ↔ DCU0, switching node 120 ↔ DCU1, switching node 130 ↔ DCU2, and so on, forming a corresponding connection relationship between switching nodes and DCUs.
[0120] For example, each switching node can be connected to the corresponding data processing unit (e.g., DCU) through an exclusive communication link (e.g., PCIe channel or customized high-speed interface, etc.) to achieve low-latency and high-bandwidth data transmission.
[0121] For example, each data processing unit (eg, DCU) includes a corresponding cache unit that can be configured to store corresponding INC data and other data.
[0122] For example, the INC agent node of each switching node can be configured to pull or receive original computing data from the cache unit of the corresponding data processing unit (such as DCU), and can also receive intermediate computing results from the upstream switching node (such as the previous level INC agent node).
[0123] For example, through chain cascading, all INC agent nodes in the exchange group can form a computing pipeline to achieve distributed and balanced load of computing tasks.
[0124] In one possible implementation, upon receiving an INC computing task, all INC agent nodes in the exchange group may (eg, in parallel) initiate data requests to their respective corresponding DCUs (eg, INC-A110 sends an INC_DATA_REQUEST message to DCU0).
[0125] For example, the request message may include an operation code (OP_PULL_DATA), the memory address range of the required data (e.g., the gradient buffer address of DCU0 may include 0x1000-0x2000), or a task identifier (e.g., Task ID), etc., although the embodiments of the present disclosure do not impose any restrictions on this. For example, after receiving the request, the DCU may encapsulate the specified data into an INC_DATA_RESPONSE message and return it to the corresponding INC proxy node (e.g., DCU0 → INC-A110) via a direct link.
[0126] For example, the INC agent node can be configured to receive local data, previous-level data, etc.
[0127] For example, local data may include raw data obtained from the DCU to which the INC agent node is bound.
[0128] For example, the previous-stage data may include the intermediate results delivered by the upstream INC agent node (such as INC-A110→INC-A120).
[0129] In at least one embodiment of the present disclosure, each switching node in the switching network can process two paths of data (eg, local data and previous-stage data), which can further reduce computing delay and wiring complexity.
[0130] In some embodiments of the present disclosure, the on-line computing agent node may be further configured to receive on-line computing data provided by the corresponding data processing unit, and execute an on-line computing task using the on-line computing data.
[0131] For example, an INC proxy node can actively pull in online computing data (INC data) from its corresponding data processing unit (DCU). For example, when an INC proxy node receives a task instruction from a central management node, or detects that downstream computing depends on its bound DCU data, it can proactively initiate a data request to its corresponding DCU. Upon receiving the data request, the DCU can return the corresponding INC data to the INC proxy node via a direct link.
[0132] For example, in the case of real-time streaming computing or tasks with strong timing dependencies between the DCU and proxy nodes, the INC proxy node can passively receive the online computing data (INC data) provided by the corresponding data processing unit (DCU). For example, the DCU can actively push the corresponding INC data to the corresponding INC proxy node based on a preset policy (such as periodic updates or event triggers).
[0133] For example, the specific implementation of the on-line computing agent node using the on-line computing data to perform the on-line computing task can refer to the relevant description of the above embodiment, which will not be repeated here.
[0134] In at least one embodiment of the present disclosure, direct links and zero-copy transmission are used between the network computing agent node and the corresponding data processing unit, which can significantly reduce the processing delay of data from the DCU to the corresponding INC agent node and then to the downstream INC agent node.
[0135] Here, zero-copy transmission may include that INC data does not need to be copied to an intermediate buffer during the transmission process, and can be directly transferred from the cache of the data processing unit to the cache of the INC proxy node.
[0136] In some embodiments of the present disclosure, the plurality of switch nodes include a first switch node and a second switch node, the first switch node includes a first on-line computing agent node, and the second switch node includes a second on-line computing agent node. The second on-line computing agent node is configured to receive intermediate computing results from the first on-line computing agent node and perform an on-line computing task using the intermediate computing results.
[0137] For example, the switching network 1000 includes at least two switching nodes, ie, a first switching node and a second switching node.
[0138] For example, the first switching node and the second switching node may be two adjacent switching nodes in a switching group.
[0139] For example, the first switching node may include a first on-net computing agent node, which may be configured to perform primary computing and generate intermediate results.
[0140] For example, the second switching node may include a second on-line computing agent node, which may be configured to receive the intermediate computing result generated by the first on-line computing agent node and complete subsequent on-line computing tasks.
[0141] For example, independent virtual lanes can be allocated for intermediate results to isolate them from communication traffic, further reducing bandwidth and deterministic latency.
[0142] For example, the first on-line computing agent node may be configured to receive INC data provided by a corresponding data processing unit, perform on-line computing operations on the received INC data to generate an intermediate result, and provide the intermediate result to a second on-line computing agent node.
[0143] For example, the second on-line computing proxy node can aggregate the intermediate result with local data (e.g., from the DCU corresponding to the second on-line computing proxy node) or other inputs. For example, if a subsequent computation phase occurs, the second on-line computing proxy node can provide the aggregated result (i.e., the new intermediate result) to a third on-line computing proxy node. For example, if it is at the end of a computation chain in a switching group, the second on-line computing proxy node can provide the final result to the central management node.
[0144] In some embodiments of the present disclosure, each of the multiple switching nodes includes a first cache unit, the first cache unit includes a first storage space and a second storage space, the first storage space is configured to store communication data transmitted by the communication network, and the second storage space is configured to store computing data of online computing tasks.
[0145] For example, each switching node includes a corresponding cache unit (eg, a first cache unit).
[0146] For example, the first cache unit may be a cache unit corresponding to an on-network computing agent node included in each switching node.
[0147] For example, the first unit can be divided into multiple storage spaces, such as a first storage space and a second storage space, through physical or logical partitioning. For example, the first storage space and the second storage space can be used to store communication data and computing data of online computing tasks, respectively, to achieve resource isolation and performance optimization.
[0148] For example, the first storage space and the second storage space may have different corresponding address ranges.
[0149] For example, independent read and write bandwidth can be allocated to the first and second storage spaces. For example, the first storage space can be configured to occupy 40% of the total bandwidth, and the second storage space can be configured to occupy 60% of the total bandwidth to prioritize INC task requirements.
[0150] For example, when bandwidth demands conflict, the allocation ratio can be dynamically adjusted based on task priority. For example, when the INC task is marked as high priority, the bandwidth allocated to communication data can be temporarily preempted. For example, when the INC task is at its peak, the bandwidth quota for the secondary storage space can be temporarily increased to 70%, while the bandwidth allocated to communication data can be reduced to 30%.
[0151] In at least one embodiment of the present disclosure, communication data and computation data are stored independently, for example, physically or logically isolated, to further reduce throughput degradation or latency jitter caused by resource contention. In at least one embodiment of the present disclosure, the second storage space can be optimized specifically for INC tasks, supporting zero-copy access and hardware-accelerated computing, significantly reducing aggregation operation latency.
[0152] In some embodiments of the present disclosure, the on-line computing agent node of each of the plurality of switching nodes is further configured to initiate an on-line computing data request to the corresponding data processing unit according to the capacity of the second storage space of the switching node.
[0153] For example, each on-line computing agent node can dynamically initiate or suspend INC data requests based on the usage status of the second storage space in the switching node where it is located and the preset capacity threshold policy, so that INC task data will not occupy cache resources indefinitely.
[0154] For example, the capacity threshold settings may include a high watermark (HWM) and a low watermark (LWM).
[0155] For example, for the high watermark, when the usage of the secondary storage space reaches HWM (e.g., 80% of the total capacity), the INC data request instruction can be suspended. For example, for the low watermark, when the usage drops back to LWM (e.g., 30% of the total capacity), the INC data request can be resumed.
[0156] For example, dynamic waterline adjustments can be made, such as dynamically adjusting HWM / LWM based on historical load and task priority (for example, high-priority tasks can temporarily increase HWM to 90%).
[0157] For example, the remaining capacity of the second storage space may be read periodically (eg, sampled once every 10 μs).
[0158] For example, after the in-network computing agent node completes the corresponding INC computing task, the corresponding data block in the corresponding second storage space can be released immediately, for example, it can be marked as an available area.
[0159] For example, each online computing agent node may periodically report the usage rate of the second storage space to the central management node.
[0160] For example, if the second storage space corresponding to an online computing agent node is continuously congested (such as the number of HWM triggers exceeds the threshold), the central management node can migrate some INC tasks to an online computing agent node with low load.
[0161] For example, if the second storage space corresponding to a certain online computing agent node reaches the HWM, a backpressure signal may be sent to its upstream node to request the upstream node to suspend or slow down the sending of INC data.
[0162] For example, backpressure signals can propagate backward along the computation links within the switch group.
[0163] In at least one embodiment of the present disclosure, the INC data occupancy ratio can be limited by a capacity threshold, thereby achieving efficient utilization of the network computing agent node resources.
[0164] In some embodiments of the present disclosure, each online computing agent node is further configured to receive at least a portion of the code of the online computing software program corresponding to the online computing task to execute the online computing task.
[0165] For example, software programs for in-network computing (INC) tasks may be deployed on various INC agent nodes of the switching network 1000 in a distributed manner.
[0166] For example, each on-line computing agent node may be configured to receive and execute a portion of code corresponding to a task based on a corresponding on-line computing software program, thereby achieving collaborative computing within the exchange group.
[0167] For example, an INC task software program can be split into multiple code shards, each of which corresponds to a computing stage (such as data preprocessing, local aggregation, and result distribution).
[0168] For example, a code shard can be bound to a data source (e.g., a code shard that processes INC data in DCU0 can be deployed to the INC proxy node corresponding to DCU0).
[0169] For example, each code shard may include metadata identifying its functional type, dependencies, and applicable data scope.
[0170] In at least one embodiment of the present disclosure, sinking computing logic to an on-network computing proxy node can reduce end-to-end latency. Furthermore, dynamic code loading can support flexible task splitting and combination to accommodate INC task requirements of varying scales. Compared to the technical solution of distributing INC programs across various DCUs and having the DCUs push data to the switching network, the uneven amount of INC data across different DCUs can cause certain caches in each switching node to be heavily occupied by INC data, thereby affecting non-INC data.
[0171] Figure 8 A schematic block diagram of an application of a switching network provided by at least one embodiment of the present disclosure is shown.
[0172] like Figure 8 As shown, the distributed computing system architecture based on the switching network includes multiple data processing units (for example, Figure 8 The example shows 40 data processing units (DCU0 to DCU39). Each DCU is configured with a corresponding in-network computing agent node (INC AGENT). Each INC agent node can implement data exchange and computing collaboration through a switching network.
[0173] For example, each DCU can communicate with the INC agent node through an interactive interface. Figure 8 In the example, Read can indicate that the INC agent node reads INC data (INC DAT) from the DCU, and Response can indicate that the DCU feeds back the data reading status or the processing result of a non-INC task.
[0174] For example, the deployment location of the INC agent node can be independent of the DCU.
[0175] For example, the INC agent node can be configured on a switch or network node instead of a DCU.
[0176] For example, an in-network computing software program (INC PROGRAM) may store logic programs for INC computing tasks. For example, it may define full-reduce computing rules (such as summation and averaging) and data scheduling strategies, although the embodiments of the present disclosure are not limited thereto.
[0177] For example, the INC program can be deployed on each INC agent node.
[0178] For example, the INC data cache space in the cache unit may include a cache area reserved exclusively for INC tasks, and the capacity is configurable.
[0179] For example, the INC proxy node can dynamically decide whether to issue an INC data request to the DCU by monitoring the cache occupancy in real time. For example, when the remaining space in the cached INC DAT falls below a threshold, the request is suspended to avoid squeezing out non-INC task resources.
[0180] For example, INC DAT is stored by DCU and can be retrieved by INC agent node through the “Read” interface.
[0181] For example, other data (OTHER DAT) can also be stored by the DCU and can be used for non-INC tasks.
[0182] For example, each INC proxy node can serve as a computing node and can temporarily store the INC data read from the corresponding DCU in the local cache; for example, the INC proxy node can merge the corresponding data through the switching network (such as the local reduction result of the calculation result of the previous INC proxy node and the calculation result of the current INC proxy node, and then pass it to the next node for merging), and finally complete the global full reduction (such as the final result is obtained after the INC data of all DCUs are aggregated through multiple layers).
[0183] For example, during the INC data application phase, the INC proxy node can detect the local cache space in real time and initiate a "Read" request only when the remaining space is sufficient to accommodate the new data, thereby avoiding the non-INC task cache being squeezed out due to excessive accumulation of INC data (for example, when the cache occupancy rate exceeds 80%, the application is suspended until space is released).
[0184] In at least one embodiment of the present disclosure, INC computations can be offloaded from DCUs to network proxy nodes, decoupling computation and storage. The dedicated cache space and dynamic application mechanism of the INC proxy nodes ensure that INC data processing does not interfere with non-INC tasks. This switching network is suitable for scenarios requiring high resource isolation and distributed computing efficiency, such as large-scale data aggregation and gradient reduction in distributed training. Intelligent proxy nodes at the network layer achieve efficient and reliable data processing.
[0185] Figure 9 A schematic block diagram of another switching network application provided by at least one embodiment of the present disclosure is shown.
[0186] like Figure 9 The switching network shown includes a "petal" type architecture, and the total number of nodes in the switching network = the number of "petals" × the number of nodes on each "petal".
[0187] For example, a petal may include a switch group, and each switch group may be configured to aggregate the on-line computing results of each on-line computing agent node within the switch group.
[0188] For example, each switching group may include at least two switching nodes (e.g. Figure 9 INC-A nodes as shown in INC-A0 to INC-A3, INC-A4 to INC-A7, INC-A8 to INC-A11, INC-A12 to INC-A15, etc. Each INC-A is connected to a corresponding DCU (e.g., DCU0 to DCU3, DCU4 to DCU7, DCU8 to DCU11, DCU12 to DCU15, etc.).
[0189] For example, INC-A0 can be configured to receive data from DCU0, and INC-A1 can be configured to receive data from DCU1. Nodes within the switching group first perform local calculations and aggregation to reduce the amount of data transmission.
[0190] For example, if the number of switching groups is N and the number of switching nodes in each switching group is M, then the total number of nodes in the switching network is (N×M), where N and M are positive integers.
[0191] For example, multiple exchange groups can establish communication connections through the central management node to achieve cross-group data interaction and collaboration.
[0192] For example, the central management node can be configured to receive instructions for in-network computing tasks and assign INC tasks to target switch groups among multiple switch groups based on system policies. For example, when receiving a computing task, the central management node can determine which switch group will undertake the task based on factors such as the load of each switch group and node performance.
[0193] For example, the central management node can be configured to receive INC computation task instructions, parse them, and assign the INC computation tasks to the target exchange group (e.g., the exchange group containing INC-A0 through INC-A3). For example, the INC-A nodes (e.g., INC-A0 and INC-A1) of the target exchange group retrieve the corresponding INC data from the connected DCUs (DCU0 and DCU1) to perform preliminary computations and aggregation. If collaboration with other exchange groups is required (e.g., INC-A8 through INC-A11), communication links can be established through the central management node to transmit intermediate results, completing cross-group data fusion and final computation.
[0194] In at least one embodiment of the present disclosure, the switching network architecture can achieve efficient allocation and execution of on-line computing tasks through centralized management of the central management node and distributed processing of the switching group, while optimizing system performance by utilizing the characteristics of the "petal" architecture, and can be applied to distributed computing scenarios of large-scale nodes.
[0195] For example, in the embodiments of the present disclosure, the switching nodes, on-line computing agent nodes, central management nodes, data processing units, or cache units may be hardware, software, firmware, or any feasible combination thereof. For example, the switching nodes, on-line computing agent nodes, central management nodes, data processing units, or cache units may be dedicated or general-purpose circuits, chips, or devices, or may be a combination of a processor and memory. The embodiments of the present disclosure do not limit the specific implementation of each of the above modules.
[0196] It should be noted that in the embodiments of the present disclosure, the various modules of the switching network may correspond to the various steps of the method for on-line computing in a switching network provided by the present disclosure (described below). The components and structures of the switching network shown in the above embodiments are merely exemplary and non-restrictive. The switching network may also include other components and structures as needed.
[0197] Figure 10 A flowchart of a method for on-line computing in a switching network provided by at least one embodiment of the present disclosure is shown.
[0198] like Figure 10 As shown, the method for on-line computing in a switching network includes step S300.
[0199] The method for on-line computing in a switching network can be applied to the switching network provided by any embodiment of the present disclosure, for example.
[0200] Step S300: In response to network computing agent nodes being respectively set on a plurality of switching nodes of a switching network, the network computing agent nodes of the plurality of switching nodes receive and execute network computing tasks.
[0201] For example, a plurality of switching nodes are communicatively coupled to one another to provide a communication network.
[0202] In some embodiments of the present disclosure, the method for on-line computing in a switching network further includes step S310: aggregating on-line computing results of each on-line computing agent node in the switching group.
[0203] Here, the multiple switching nodes in the switching network may be configured into multiple switching groups, each switching group including at least two switching nodes.
[0204] In some embodiments of the present disclosure, the method for on-line computing in a switching network further includes step S320: receiving an instruction for an on-line computing task, and allocating the on-line computing task to a target switching group among the plurality of switching groups.
[0205] Here, the switching network may further include a central management node, and the plurality of switching groups may be communicatively connected via the central management node.
[0206] In some embodiments of the present disclosure, the method for on-line computing of a switching network further includes step S330: in response to the on-line computing agent nodes of at least two switching nodes of each switching group being connected in sequence through the communication link within the switching group, the on-line computing tasks are sequentially executed and the on-line computing results are transmitted.
[0207] In some embodiments of the present disclosure, the method for on-line computing of a switching network further includes step S340: in response to executing an on-line computing task, initiating an on-line computing data request or returning an on-line computing result to a corresponding data processing unit of the communication connection.
[0208] Here, the plurality of switching nodes are further configured to be communicatively connected to the plurality of data processing units respectively.
[0209] In some embodiments of the present disclosure, in the method for on-line computing in a switching network, step S340 further includes steps S341 to S342.
[0210] Step S341: receiving online computing data provided by the corresponding data processing unit.
[0211] Step S342: Use the online computing data to execute the online computing task.
[0212] In some embodiments of the present disclosure, the method for on-line computing in a switching network further includes step S350.
[0213] Step S350: Receive the intermediate computing result from the first on-line computing agent node, and use the intermediate computing result to execute the on-line computing task.
[0214] Here, the multiple switching nodes include a first switching node and a second switching node. The first switching node includes a first on-line computing agent node, and the second switching node includes a second on-line computing agent node. Step S350 can be applied to the second on-line computing agent node, for example.
[0215] In some embodiments of the present disclosure, the method for on-line computing in a switching network further includes step S360.
[0216] Step S360: Initiate an on-line computing data request to the corresponding data processing unit according to the capacity of the second storage space of the switching node.
[0217] For example, each of the multiple switching nodes includes a first cache unit, the first cache unit includes a first storage space and a second storage space, the first storage space is configured to store communication data transmitted by the communication network, and the second storage space is configured to store computing data of network computing tasks.
[0218] In some embodiments of the present disclosure, the method for on-line computing in a switching network further includes step S370.
[0219] Step S370: Receive at least a portion of the code of the online computing software program corresponding to the online computing task to execute the online computing task.
[0220] It should be noted that the relevant content of the functions or beneficial effects of each step in the method for on-line computing of a switching network provided in any embodiment of the present disclosure can be referred to the relevant description of the switching network provided in any embodiment of the present disclosure, and will not be repeated here.
[0221] Figure 11 A schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0222] At least some embodiments of the present disclosure also provide an electronic device, such as Figure 11 As shown, the electronic device 400 includes a switching network provided by any embodiment of the present disclosure. For example, the switching network can be implemented as a sender and / or receiver of instructions / data to transmit data or instructions, such as for internal memory or external memory.
[0223] For example, the processor 401 may be a central processing unit (CPU), a graphics processing unit (GPU), or other forms of processing units with data processing capabilities and / or program execution capabilities; for example, the central processing unit (CPU) may be RISC, X86, or ARM architecture, etc.
[0224] The electronic device 400 in the embodiment of the present disclosure may include mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as any devices such as digital TVs, desktop computers, servers, etc., and may also be a combination of any data processing devices and hardware, which is not limited by the embodiment of the present disclosure.
[0225] Figure 11 The electronic device 400 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0226] For example, Figure 11As shown, in some examples, processor 401 can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 402 or programs loaded from storage device 408 into random access memory (RAM) 403, for example, executing one or more steps of the method for on-network computing based on a switching network according to at least one embodiment of the present disclosure. RAM 403 also stores various programs and data required for the operation of the computer system. Processor 401, ROM 402, and RAM 403 are connected to each other via a communication channel 404. An input / output (I / O) interface 405 is also connected to communication channel 404.
[0227] For example, the following components can be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc., for example, the storage controller of the storage device 408 includes the above-mentioned switching network; and a communication device 409 including, for example, a network interface card such as a LAN card, a modem, etc. The communication device 409 can allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data, and perform communication processing via a network such as the Internet. The drive 410 is also connected to the I / O interface 405 as needed. Removable media 411, such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive 410 as needed, so that the computer program read therefrom can be installed into the storage device 408 as needed. Although Figure 11 The electronic device 400 is shown as including various devices, but it should be understood that it is not required to implement or include all of the devices shown. More or fewer devices may be implemented or included instead.
[0228] For example, the electronic device 400 may further include a peripheral interface (not shown in the figure), etc. The peripheral interface may be various types of interfaces, such as a USB interface, a lightning interface, etc. The communication device 409 may communicate with a network and other devices through wireless communication, such as the Internet, an intranet, and / or a wireless network such as a cellular telephone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). Wireless communications may use any of a variety of communication standards, protocols, and technologies, including, but not limited to, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0229] It should be noted that, in the embodiment of the present disclosure, the specific functions and technical effects of the electronic device 400 can be referred to, for example, the relevant description of the switching network in the embodiment of the present disclosure, and will not be repeated here.
[0230] Figure 12 A schematic block diagram of another electronic device provided by at least one embodiment of the present disclosure is shown.
[0231] At least one embodiment of the present disclosure further provides an electronic device, such as Figure 12 As shown, the electronic device 500 includes at least one processor 510 and at least one memory 520 .
[0232] For example, the memory 520 can be used to non-transitorily store computer-readable instructions (e.g., one or more computer program modules). The processor 510 can be used to execute the computer-readable instructions. When the computer-readable instructions are executed by the processor 510, one or more steps of the above-described method for on-network computing in a switching network can be performed. The memory 520 and the processor 510 can be interconnected via a bus or link, using wired, wireless, and / or other forms of communication media, etc., and the embodiments of the present disclosure are not limited in this regard.
[0233] For example, the processor 510 may be a central processing unit (CPU), a graphics processing unit (GPU), or other processing units with data processing capabilities and / or program execution capabilities. For example, the central processing unit (CPU) may be a RISC, X86, or ARM architecture. The processor 510 may be a general-purpose processor or a dedicated processor, and may control other components in the electronic device 500 to perform desired functions.
[0234] For example, memory 520 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, and the like. One or more computer program modules may be stored on the computer-readable storage medium, and the processor 510 may execute one or more computer program modules to implement the various functions of the electronic device 500. The computer-readable storage medium may also store various applications and data, as well as data used and / or generated by the applications.
[0235] At least one embodiment of the present disclosure further provides a non-transitory computer-readable storage medium for non-temporarily storing computer-readable instructions, which, when executed by at least one processor, can implement the above-mentioned method for on-line computing in a switching network.
[0236] Figure 13 A schematic diagram of a computer-readable storage medium provided by at least one embodiment of the present disclosure is shown.
[0237] like Figure 13 As shown, the computer-readable storage medium 600 is used to store computer-readable instructions 610. For example, when the computer-readable instructions 610 are executed by at least one processor, one or more steps in the method for on-line computing in a switching network described above can be performed.
[0238] For example, the computer-readable storage medium 600 can be applied to the electronic device 400 or the electronic device 500. For example, the relevant description of the non-volatile computer-readable storage medium 600 can also be referred to Figure 11 The storage device 408 in the electronic device 400 is shown as well as Figure 12 The corresponding description of the memory 520 in the electronic device 500 is not repeated here.
[0239] It should be noted that, in the embodiments of the present disclosure, the specific functions and technical effects of the computer-readable storage medium 600 can be referred to the above description of the switching network and / or the method for on-line computing for the switching network, and will not be repeated here.
[0240] There are a few points to note:
[0241] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0242] (2) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0243] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A switching network comprising: A plurality of switching nodes are connected to each other to provide a communication network, wherein each of the plurality of switching nodes includes a network computing agent node, The on-line computing agent node is configured to receive and execute on-line computing tasks.
2. The switching network according to claim 1, wherein: The plurality of switching nodes are configured into a plurality of switching groups; Each switching group includes at least two switching nodes and is configured to aggregate the on-line computing results of the on-line computing agent nodes in the switching group.
3. The switching network according to claim 2, further comprising a central management node, in, The central management node is configured to receive an instruction for the online computing task and allocate the online computing task to a target switching group among the multiple switching groups; The multiple switching groups are also communicatively connected via the central management node.
4. The switching network according to claim 2, wherein: The on-line computing agent nodes of at least two switching nodes of each switching group are sequentially connected via a communication link within the switching group, and are configured to sequentially execute the on-line computing tasks and transmit the on-line computing results.
5. The switching network according to any one of claims 1 to 4, wherein: The plurality of switching nodes are further configured to be communicatively connected to the plurality of data processing units respectively; Each of the online computing agent nodes is further configured to, in response to executing the online computing task, initiate an online computing data request or return an online computing result to a corresponding data processing unit connected thereto.
6. The switching network according to claim 5, wherein: The on-line computing agent node is further configured to: receiving the online computing data provided by the corresponding data processing unit, and The online computing task is executed using the online computing data.
7. The switching network according to claim 5, wherein: The plurality of switching nodes include a first switching node and a second switching node, wherein the first switching node includes a first on-line computing agent node, and the second switching node includes a second on-line computing agent node. The second on-line computing agent node is configured to receive an intermediate computing result from the first on-line computing agent node and execute the on-line computing task using the intermediate computing result.
8. The switching network according to claim 5, wherein: Each of the plurality of switching nodes comprises a first cache unit, The first cache unit includes a first storage space and a second storage space. The first storage space is configured to store communication data transmitted by the communication network, and the second storage space is configured to store computing data of the on-line computing task.
9. The switching network of claim 8, wherein: The on-line computing agent node of each of the plurality of switching nodes is further configured to initiate an on-line computing data request to the corresponding data processing unit according to the capacity of the second storage space of the switching node.
10. The switching network according to any one of claims 1 to 4, wherein: Each of the online computing agent nodes is further configured to receive at least a portion of the code of the online computing software program corresponding to the online computing task to execute the online computing task.
11. A method for on-line computing in a switching network, comprising: In response to respectively setting network computing agent nodes on a plurality of switching nodes of the switching network, receiving and executing network computing tasks through the network computing agent nodes of the plurality of switching nodes, The plurality of switching nodes are communicatively connected to each other to provide a communication network.
12. An electronic device comprising: The switching network according to any one of claims 1 to 10.
13. An electronic device comprising: The switching network according to any one of claims 5 to 9; as well as the plurality of data processing units, Wherein, each of the plurality of data processing units is configured to respond to the network computing data request or receive the network computing result.
14. A non-transitory computer-readable storage medium for non-transitory storage of computer-readable instructions, which, when executed by at least one processor, implements the method for on-network computing in a switching network according to claim 11.
15. An electronic device comprising: at least one processor; as well as A memory, wherein the memory stores at least one computer program, and when the at least one computer program is executed by the at least one processor, the method for on-line computing in a switching network according to claim 11 is implemented.
Citation Information
Cited By
Data communication method and device, electronic equipment and storage medium
CN121509528A