Data processing method and switch in asymmetric collective communication

By using data packets carrying message type fields and virtual addresses in asymmetric aggregate communication, switches dynamically manage memory access, solving the problem of memory waste and improving processor memory utilization.

CN122633371APending Publication Date: 2026-08-25LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610525681.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In asymmetric aggregate communication, existing technologies use static pre-allocation of video memory, which leads to memory waste in the XPU and causes serious waste of memory space.

Method used

By including the message type field and virtual address in the data packet, the switch determines the physical address for memory access based on the processor identifier, avoiding the need to pre-allocate storage space in the processor memory area and achieving dynamic memory management.

Benefits of technology

It effectively solves the problem of redundant and idle memory and improves the processor's memory utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633371A_ABST
    Figure CN122633371A_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and an exchange in asymmetric set communication. The method applied to the exchange comprises the following steps: receiving a first data packet sent by a first processor; the first data packet comprises at least a packet type field, a processor identifier corresponding to the first processor and a virtual address; the packet type field is located in a packet header of the first data packet; the virtual address corresponds to a memory access physical address in a processor; the packet type field represents a data processing stage of the set communication; determining a second processor corresponding to the first processor according to the processor identifier corresponding to the first processor; and sending a second data packet to the second processor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data processing method and switch in asymmetric set communication. Background Technology

[0002] In asymmetric aggregate communication, a static pre-allocation of video memory is used to allocate memory regions for other XPUs in the same communication group within each heterogeneous computing unit (XPU) of the communication group. However, this approach leads to memory waste in the XPUs. Summary of the Invention

[0003] In view of this, this application provides a data processing method and a switch for asymmetric aggregate communication, as follows:

[0004] A data processing method in asymmetric aggregate communication, applied to a switch, the method comprising:

[0005] Receive a first data packet sent by a first processor; the first data packet includes at least a packet type field, a processor identifier corresponding to the first processor, and a virtual address; the packet type field is located in the packet header of the first data packet;

[0006] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication;

[0007] The second processor corresponding to the first processor is determined according to the processor identifier corresponding to the first processor;

[0008] The second data packet is sent to the second processor.

[0009] Optionally, in the above method, if the message type field represents the data distribution stage, the second data message may also include first target data;

[0010] The first target data is written to the first physical address in the memory region of the second processor according to the virtual address.

[0011] Optionally, in the above method, if the message type field represents the data merging stage, the first data message may also include second target data;

[0012] The method further includes:

[0013] The second target data in all the first data packets sent by multiple first processors is aggregated to obtain an aggregation result;

[0014] Based on the aggregation result, the second data packet is generated; the aggregation result is written to the corresponding second physical address in the memory region of the second processor according to the virtual address.

[0015] Optionally, in the above method, where the message type field represents the data merging stage, the first data message further includes a transaction identifier field; the transaction identifier field is used to record the target transaction identifier corresponding to the data distribution stage corresponding to the data merging stage.

[0016] The target transaction identifier and the virtual address are used to indicate reading data from the corresponding third physical address in the memory region of the second processor.

[0017] Optionally, the above method involves determining the second processor corresponding to the first processor based on the processor identifier corresponding to the first processor, including:

[0018] Based on the processor identifier corresponding to the first processor, query the processor identifier corresponding to the second processor that is in the same set of communication groups as the first processor in the communication group configuration information registered in the switch.

[0019] The processor identifier corresponding to the second processor is used to send the second data packet to the corresponding second processor.

[0020] Optionally, in the above method, the virtual address includes: a base address and an address offset;

[0021] Among them, processors belonging to the same communication group have the same base address; the address offset and the access count value recorded in the processor are used to determine the physical address of memory access in the processor corresponding to the virtual address.

[0022] A data processing method in asymmetric set communication, applied to a second processor, the method comprising:

[0023] The receiver receives a third data packet sent by the switch; the third data packet includes at least a packet type field, a processor identifier of the first processor, and a virtual address; the packet type field is located in the header of the third data packet.

[0024] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication; the processor identifier of the first processor matches the second processor;

[0025] Data is accessed at the physical address corresponding to the virtual address in the memory region of the second processor.

[0026] Optionally, in the above method, the virtual address includes: a base address and an address offset;

[0027] Specifically, data access is performed at the physical address corresponding to the virtual address in the memory region of the second processor, including:

[0028] When the message field type represents the data merging stage, the corresponding first offset is queried in the address mapping table of the second processor according to the processor identifier of the first processor;

[0029] The third physical address is determined based on the first offset and the base address;

[0030] Read the second target data from the third physical address corresponding to the memory region in the second processor.

[0031] Optionally, the third data message further includes a transaction identifier field; the transaction identifier field is used to record the target transaction identifier corresponding to the data distribution stage corresponding to the data merging stage.

[0032] Specifically, in the address mapping table of the second processor, the first offset is queried according to the processor identifier corresponding to the first processor, including:

[0033] In the address mapping table of the second processor, the corresponding first offset is queried according to the processor identifier corresponding to the first processor and the target transaction identifier.

[0034] Optionally, in the above method, the virtual address includes: a base address and an address offset;

[0035] Specifically, data access is performed at the physical address corresponding to the virtual address in the memory region of the second processor, including:

[0036] When the message field type represents the data distribution phase, the second offset is determined based on the address offset and the access count value recorded by the second processor;

[0037] The first physical address is determined based on the base address and the second offset;

[0038] Write the first target data in the third data packet into the first physical address corresponding to the memory region in the second processor.

[0039] Optionally, the above method may further include:

[0040] Add the second offset to the address mapping table according to the processor identifier corresponding to the first processor;

[0041] Increment the access count value recorded by the second processor by 1.

[0042] Optionally, the third data message may further include a calculated anomaly field.

[0043] The method further includes:

[0044] If the calculation anomaly field indicates an anomaly in the data calculation of the second processor, the operation of incrementing the access count value by 1 is revoked, and the operation of adding the second offset to the address mapping table is also revoked.

[0045] A switch, comprising:

[0046] A communication module is configured to receive a first data packet sent by a first processor; the first data packet includes at least a packet type field, a processor identifier corresponding to the first processor, and a virtual address; the packet type field is located in the header of the first data packet;

[0047] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication;

[0048] The controller is used to determine the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor;

[0049] The communication module is also used to send the second data message to the second processor.

[0050] As can be seen from the above technical solution, in the data processing method and switch for asymmetric aggregated communication disclosed in this application, the data packet sent by the first processor to the switch contains a message type field and a virtual address. The message type field represents the data processing stage of the aggregated communication, such as the data distribution stage or the data merging stage, and the virtual address represents the physical address for memory access. Thus, after the switch sends the data packet to the processor according to the processor identifier in the data packet, the processor can access the memory physical address represented by the virtual address according to the data processing stage represented by the message type field. It is evident that in this application, the data packet sent by the processor to the switch carries a message type field and a virtual address, thereby specifying the physical address for memory access by the processor according to the data processing stage. Therefore, in asymmetric aggregated communication, it is not necessary to pre-allocate storage space for other processors in the processor's memory area, effectively solving the technical defects of memory redundancy and idle space waste in aggregated communication, thus avoiding space waste in the processor's memory area and improving processor memory utilization. Attached Figure Description

[0051] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 A flowchart illustrating a data processing method in asymmetric communication provided in this application embodiment;

[0053] Figure 2 This is a diagram of the architecture of a scale-up network.

[0054] Figure 3 This is a schematic diagram illustrating the interaction between the processor and the switch in an embodiment of this application;

[0055] Figure 4 This is another schematic diagram illustrating the interaction between the processor and the switch in an embodiment of this application;

[0056] Figure 5 This is another schematic diagram illustrating the interaction between the processor and the switch in an embodiment of this application;

[0057] Figure 6 A flowchart illustrating another data processing method in asymmetric set communication provided in this application embodiment;

[0058] Figure 7 A partial flowchart of another data processing method in asymmetric communication provided in this application embodiment;

[0059] Figure 8 This is an example diagram of the address mapping table corresponding to the memory regions in the processor in this application embodiment;

[0060] Figure 9 This is another example diagram of the address mapping table corresponding to the memory region in the processor in this application embodiment;

[0061] Figure 10 Another part of the flowchart of a data processing method in asymmetric communication provided in the embodiments of this application;

[0062] Figure 11 This application provides a schematic diagram of the structure of a data processing device in asymmetric set communication.

[0063] Figure 12 Another schematic diagram of a data processing device in asymmetric set communication provided in this application embodiment;

[0064] Figure 13A schematic diagram of the structure of a data processing device in another asymmetric set communication provided in this application embodiment;

[0065] Figure 14 A schematic diagram of the structure of a switch provided in an embodiment of this application;

[0066] Figure 15 This is a schematic diagram of the ECH standard header structure for the application of this application in scale-up networks.

[0067] Figure 16 This is a flowchart illustrating the dynamic data redirection and writing process during the Dispatch phase in scenarios applicable to scale-up networks in this application.

[0068] Figure 17 This application is applicable to the scenario of scale-up networks, and includes a block diagram of the receiving-side memory address interception and redirection logic.

[0069] Figure 18 This is an example diagram illustrating the data backtracking and aggregation process based on mapping in the Combine phase, which is applicable to the scale-up network scenario of this application. Detailed Implementation

[0070] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0071] refer to Figure 1 The diagram shown illustrates a flowchart of a data processing method in asymmetric communication provided in this application. This method can be applied to switches with data forwarding capabilities, such as switching chips that integrate computing engines and support the In-Fabric Extended Computation (IFEC) protocol. For example, Figure 2As shown, the switch can be a core computing switch or an edge switch in a scale-up intelligent computing network. The intelligent computing network interacts with a computing cluster composed of multiple processors, such as XPUs (Graphics Processing Units), GPUs (Graphics Processing Units), NPUs (Neural Processing Units), TPUs (Tensor Processing Units), and DPUs (Data Processing Units). These processors can be used in scenarios such as inference for artificial intelligence models, and they collaborate to complete a distributed computing task. The technical solution in this embodiment is mainly used to improve the memory utilization of the processors.

[0072] Specifically, the method in this embodiment may include the following steps:

[0073] Step 101: Receive the first data packet sent by the first processor.

[0074] The first data packet includes at least a packet type field, a processor identifier corresponding to the first processor, and a virtual address. The packet type field is located in the packet header of the first data packet, such as the Expansion Computing Header (ECH). The virtual address corresponds to the physical address for memory access in the processor. The packet type field represents the data processing stage of the aggregated communication. The data processing stage can be a data distribution stage or a data merging stage (also known as a data reduction stage). In the data distribution stage, the processor writes data to the memory areas of other processors through the switch. In the data merging stage, the switch aggregates the data sent by the processors and then forwards it to other processors.

[0075] For example, in this embodiment, the first data packet can be a communication packet transmitted based on IFEC, including at least a packet type field, a processor identifier, and a virtual address. The processor identifier is used to uniquely identify the processor to which it belongs, such as a communication node rank identifier (e.g., Rank number) or a processor physical unique identifier (e.g., physical ID). The packet header is the beginning of the data packet and contains control information such as routing, protocol type, and length. For example, the packet type field can be the Type field in the packet header. If the Type field is Dispatch, it indicates the data distribution stage; if the Type field is Combine, it indicates the data merging stage.

[0076] In this embodiment, the first data packet sent by the first processor can be received through the data transmission channel between the switch and the first processor.

[0077] Step 102: Determine the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor.

[0078] The switch stores communication group configuration information, which includes a Group ID corresponding to at least one aggregated communication group and member information of the communication group members. Each communication group member is a processor within the aggregated communication group, and the member information includes at least the processor identifier of the processor within the aggregated communication group. For example, the switch stores the processor identifiers of processors A, B, and C in communication group 1, and the processor identifiers of processors D, E, and F in communication group 2. Based on this, in this embodiment, the second processor corresponding to the first processor can be determined according to the pre-stored communication group configuration information in the switch. The second processor and the first processor belong to the same aggregated communication group.

[0079] Specifically, in this embodiment, when a collective communication group is formed, the processor identifiers of each processor included in the collective communication group can be registered in the switch to update the communication group configuration information in the switch.

[0080] Step 103: Send the second data packet to the second processor.

[0081] In this embodiment, the second data packet can be sent to the second processor through the data transmission channel between the switch and the second processor.

[0082] It should be noted that the second data packet is obtained based on the first data packet.

[0083] In one scenario, where the message type field represents the data distribution phase (such as the Dispatch phase), the second data message includes the first target data and virtual address carried by the first data message, so that the second processor can write the first target data into the memory area of ​​the second processor according to the virtual address.

[0084] In another case, where the message type field represents the data merging stage (such as the Combine stage), the second data message includes the aggregation result obtained by data aggregation of the second target data carried by the first data message and the virtual address, so that the second processor can write the aggregation result into the memory area of ​​the second processor according to the virtual address.

[0085] In another scenario, where the message type field represents the data merging phase, the second data message may also include the transaction identifier field carried by the first data message. The transaction identifier field is used to record the target transaction identifier corresponding to the data distribution phase corresponding to the data merging phase, such as the Instruction ID of the Dispatch phase, so that the second processor can read the corresponding data from the memory area in the second processor according to the virtual address and the target transaction identifier.

[0086] For example, such as Figure 3 As shown, the first processor can send a first data packet to the switch as the sender in the data distribution phase (Dispatch phase). The first data packet includes at least first target data, a virtual address, and the processor identifier of the first processor. The second processor can receive a second data packet from the switch as the receiver in the data distribution phase. The second data packet includes at least first target data, a virtual address, and the processor identifier of the first processor, so that the second processor can write the first target data into the memory area of ​​the second processor according to the virtual address and the processor identifier of the first processor.

[0087] For example, such as Figure 4 As shown, the first processor can send a first data packet to the switch as the receiving side in the data merging stage. The first data packet includes at least the processor identifier and virtual address of the first processor (and may also include the target transaction identifier recorded in the transaction identifier field). The second processor can receive a second data packet from the switch as the sending side in the data merging stage. The second data packet includes at least the processor identifier and virtual address of the first processor (and may also include the target transaction identifier recorded in the transaction identifier field), so that the second processor can read the corresponding data from the memory area in the second processor according to the virtual address and the processor identifier of the first processor (and may also include the target transaction identifier recorded in the transaction identifier field).

[0088] For example, such as Figure 5 As shown, the first processor can send a first data packet to the switch as the sending side during the data merging stage. The first data packet includes at least a virtual address, second target data, and the processor identifier of the first processor. The second processor can receive a second data packet from the switch as the receiving side during the data merging stage. The second data packet includes at least a virtual address, an aggregation result (obtained by the switch aggregating the second target data carried in the first data packet), and the processor identifier of the first processor, so that the second processor can write the aggregation result into the memory area of ​​the second processor according to the virtual address and the processor identifier of the first processor.

[0089] As can be seen from the above technical solution, in the data processing method for asymmetric aggregate communication provided in this application embodiment, the data packet sent by the processor to the switch carries a packet type field and a virtual address, thereby specifying the physical address for the processor to access memory according to the data processing stage. Thus, in asymmetric aggregate communication, it is not necessary to pre-allocate storage space for other processors in the memory area of ​​the processor, effectively solving the technical defects of memory redundancy and idle space waste in aggregate communication, thereby avoiding the waste of space in the memory area of ​​the processor and improving the processor memory utilization rate.

[0090] In one implementation, where the message type field represents the data distribution phase (such as the Dispatch phase), the first data packet also includes first target data, which is the data that the first processor needs to distribute to the second processor. For example, the first processor, acting as the initiator of the Dispatch phase, sends the first data packet to the switch. After the first data packet is sent to the switch as a write request, the switch sends a second data packet to the second processor instructing the writing of the first target data.

[0091] Accordingly, the second data packet sent by the switch to the second processor contains not only the virtual address and the processor identifier of the first processor, but also the first target data. Based on this, after the second data packet is sent by the switch to the second processor, the second processor writes the first target data in the second data packet into the corresponding first physical address in the memory area of ​​the second processor according to the virtual address and the processor identifier of the first processor.

[0092] The second processor locates the corresponding first physical address in the memory region of the second processor according to the virtual address and the processor identifier of the first processor. Accordingly, the second processor writes the first target data in the second data packet to the located first physical address.

[0093] As can be seen, in this embodiment, during the data distribution phase between the processor and the switch, the virtual address and the processor identifier of the first processor are carried in the data packet. This enables the processor to locate the mapped physical address according to the virtual address and the processor identifier of the first processor, thereby realizing data distribution from the first processor to the second processor. In this process, by mapping the physical address to the virtual address, it is not necessary to pre-allocate storage space in the memory area of ​​the processor. Therefore, in asymmetric set communication, it is possible to avoid wasting space in the memory area of ​​the processor, thereby improving the processor memory utilization.

[0094] In one implementation, where the message type field represents a data merging stage (such as the Combine stage), the first data message also includes second target data. For example, the first processor, acting as the initiator of the Combine stage, sends a first data message carrying the second target data to the switch. Based on this, in this embodiment, the following processing is further performed on the switch:

[0095] The second target data in all first data packets sent by multiple first processors is aggregated to obtain an aggregation result. Then, a second data packet is generated based on the aggregation result.

[0096] The second data packet includes the aggregation result, as well as the virtual address and the processor identifier of the first processor. Based on this, the aggregation result is written to the corresponding second physical address in the memory region of the second processor according to the virtual address and the processor identifier of the first processor.

[0097] It should be noted that aggregation processing on the switch refers to the processing of performing at least one calculation operation on second target data from multiple first processors. The calculation operation can be summation (Sum), averaging, finding the maximum value (Max), finding the minimum value (Min), etc. The calculation operation can support calculations of data types such as INT, Float, and BFloat.

[0098] Accordingly, after the switch sends the second data packet containing the aggregation result to the second processor, the second processor can write the aggregation result in the second data packet into the corresponding second physical address in the memory area of ​​the second processor according to the virtual address and the processor identifier of the first processor.

[0099] The second processor locates the corresponding second physical address in the memory region of the second processor according to the virtual address and the processor identifier of the first processor. Accordingly, the second processor writes the aggregation result in the second data packet to the located second physical address.

[0100] As can be seen, in this embodiment, when the processor and the switch are in the data merging stage, the data packet carries a virtual address and the processor identifier of the first processor. This allows the processor to locate the mapped physical address according to the virtual address and the processor identifier of the first processor. After the switch performs aggregation processing, the aggregated data is sent to the second processor. In this process, by mapping the physical address to the virtual address, it is not necessary to pre-allocate storage space in the memory area of ​​the processor. Therefore, in asymmetric aggregate communication, it is possible to avoid wasting space in the memory area of ​​the processor, thereby improving the processor memory utilization.

[0101] In one implementation, where the message type field represents a data merging phase (such as the Combine phase), the first data message does not contain the target data. Optionally, the first data message may contain a data read identifier. For example, a first processor, acting as the receiving side in the Combine phase, sends a first data message to the switch. After the first data message is sent to the switch as a read request, the switch, according to the instructions in the first data message, requests the corresponding data to be read from a second processor, acting as the sending side.

[0102] Correspondingly, the second data packet sent by the switch to the second processor is the same as the first data packet, containing a virtual address and the processor identifier of the first processor. Based on this, after the second data packet is sent to the second processor by the switch, the second processor reads the corresponding data from the third physical address in the memory region of the second processor according to the virtual address and the processor identifier of the first processor.

[0103] The second processor locates the corresponding third physical address in the memory region of the second processor according to the virtual address and the processor identifier of the first processor. Accordingly, the second processor reads the corresponding data from the located third physical address and encapsulates the data into a data packet and sends it to the switch. The switch then aggregates the data and encapsulates the aggregation result into a data packet and sends it to the first processor.

[0104] As can be seen, in this embodiment, when the processor and the switch are in the data merging stage, the data packet carries a virtual address and the processor identifier of the first processor. This allows the processor to locate the mapped physical address according to the virtual address and the processor identifier of the first processor, thereby enabling data reading from the second processor to the switch. In this process, by mapping the physical address to the virtual address, it is not necessary to pre-allocate storage space in the memory area of ​​the processor. Therefore, in asymmetric set communication, it is possible to avoid wasting space in the memory area of ​​the processor, thereby improving the processor memory utilization.

[0105] Furthermore, when the message type field represents the data merging stage, the first data message may also include a transaction identifier field. This transaction identifier field records the target transaction identifier (Instruction ID) corresponding to the data distribution stage corresponding to the data merging stage. The target transaction identifier establishes the association between the data merging stage and the data distribution stage, indicating that the data merging stage and the data distribution stage are two stages processing the same transaction. Based on this, in this embodiment, the transaction identifier field, the virtual address, and the processor identifier of the first processor are used to instruct the switch to read the corresponding data from the second processor.

[0106] Correspondingly, the second data packet sent by the switch to the second processor is the same as the first data packet, containing a virtual address, a target transaction identifier, and the processor identifier of the first processor. Based on this, after the second data packet is sent by the switch to the second processor, the second processor reads the corresponding data from the third physical address in the memory region of the second processor according to the virtual address, the target transaction identifier, and the processor identifier of the first processor.

[0107] The second processor locates the corresponding third physical address in the memory region of the second processor according to the virtual address, the target transaction identifier and the processor identifier of the first processor. Accordingly, the second processor reads the corresponding data from the located third physical address and encapsulates the data into a data packet and sends it to the switch. The switch then aggregates the data and encapsulates the aggregation result into a data packet and sends it to the first processor.

[0108] As can be seen, in this embodiment, when the processor and the switch are in the data merging stage, the data packet carries a virtual address, a target transaction identifier, and a processor identifier of the first processor. This allows the processor to locate the mapped physical address according to the virtual address, the target transaction identifier, and the processor identifier of the first processor, thereby enabling data reading from the second processor to the switch. In this process, by mapping the physical address to the virtual address, it is not necessary to pre-allocate storage space in the processor's memory area. Therefore, in asymmetric aggregate communication, it is possible to avoid wasting space in the processor's memory area, thereby improving the processor's memory utilization.

[0109] In one implementation, when determining the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor in step 102, it can be achieved in the following way:

[0110] Based on the processor identifier corresponding to the first processor, query the processor identifier corresponding to the second processor that is in the same set of communication groups as the first processor in the communication group configuration information registered in the switch.

[0111] The processor identifier corresponding to the second processor is used to send the second data packet to the corresponding second processor. That is, the processor identifier corresponding to the second processor is found in the communication group configuration information registered in the switch, and the second data packet is sent to the second processor according to the found processor identifier.

[0112] The communication group configuration information may include the processor identifiers of each processor included in the communication group. For example, communication group 1 and communication group 2 are registered in the switch. The communication group configuration information of communication group 1 includes Rank numbers A, B, and C, indicating that communication group 1 includes processors A, B, and C. The communication group configuration information of communication group 2 includes Rank numbers D, E, and F, indicating that communication group 2 includes processors D, E, and F. Based on this, in this embodiment, the processor identifiers of processors B and C can be found in the communication group configuration information according to processor A, thereby sending the second data packet to processors B and C respectively.

[0113] In one implementation, a virtual address may include a base address and an address offset.

[0114] The base address can be a pre-set address. Processors in the same communication group have the same base address, which can be represented by V. G This is indicated by the address offset and the access count value recorded in the processor, which are used to determine the physical address of the memory access corresponding to the virtual address in the processor.

[0115] It should be noted that the access count value represents the number of times data is written to the memory area in the processor, which can be represented by T. Each time data is written, the access count value corresponding to the processor is incremented by 1. Based on this, data can be written to the memory area in the processor in the order of physical addresses.

[0116] Specifically, the address offset at least represents the size of the currently accessed data block, which can be represented by X. Furthermore, the address offset can also carry a tensor index, represented by K. The tensor index represents the sequence number of the currently accessed data block after the original data (tensor data) to be transferred by the processor has been segmented. For example, a virtual address can be represented by V. G +K×X means that, based on the virtual address, the access count value T can be used to replace K in the virtual address, so the physical address of the memory access determined by the virtual address can be V. G +T×X. A mapping relationship is established between this memory access physical address and the processor identifier (and target transaction identifier) ​​of the processor to which the corresponding data belongs, so that the corresponding data can be read in the memory area according to the mapping relationship.

[0117] As can be seen, in this embodiment, the access count value can be used to represent the number of times data is written to the memory area in the processor. Each time data is written, the access count value corresponding to the processor is incremented by 1, so that data can be written in the memory area in the processor in the order of physical addresses, which can avoid the generation of space fragmentation in the memory area and thus improve the space utilization of the memory area.

[0118] refer to Figure 6 This is a flowchart illustrating the implementation of a data processing method in asymmetric set communication provided in this application embodiment. This method can be applied to... Figure 2 Any processor shown, such as Figure 3 , Figure 4 , Figure 5 The second processor is shown. The technical solution in this embodiment is mainly used to improve the memory utilization of the processor.

[0119] Specifically, the method in this embodiment may include the following steps:

[0120] Step 601: Receive the third data packet sent by the switch.

[0121] The third data packet includes at least a packet type field, the processor identifier of the first processor, and a virtual address. Combined with... Figure 1 The technical solution shown in this embodiment is the third data packet received by the second processor from the switch.

[0122] It should be noted that the message type field in the third data packet is located in the packet header. The virtual address corresponds to the physical address for memory access in the processor. The message type field represents the data processing stage of the aggregated communication, such as the data distribution stage or the data merging stage. The processor identifier of the first processor matches the second processor; that is, the second processor is determined on the switch according to the processor identifier of the first processor, so that the third data packet can be sent to the second processor.

[0123] Step 602: Access the data at the corresponding physical address in the memory region of the second processor according to the virtual address.

[0124] Specifically, in this embodiment, the corresponding physical address can be located in the memory region of the second processor according to the virtual address and the processor identifier of the first processor to achieve data access.

[0125] It should be noted that, within the memory region of the second processor, a mapping relationship between virtual addresses and physical addresses is established using the processor identifier of the first processor. Furthermore, to distinguish the data processing stage of the aggregated communication in which the first processor is located, a mapping relationship between virtual addresses and physical addresses is established within the memory region of the second processor using the first processor, its processor identifier, and the target transaction identifier. In this embodiment, data can be read or written according to the mapping relationship, based on the virtual address and the corresponding physical address in the memory region of the second processor corresponding to the processor identifier of the first processor.

[0126] As can be seen from the above technical solution, in the data processing method for asymmetric aggregate communication provided by the embodiments of this application, the data packet sent by the switch to the processor carries a packet type field, the processor identifier of the first processor, and a virtual address. The processor can access the physical address of memory based on the virtual address and the processor identifier according to the corresponding data processing stage. Thus, in asymmetric aggregate communication, it is not necessary to pre-allocate storage space for other processors in the memory area of ​​the processor, which effectively solves the technical defects of memory redundancy and idle space waste in aggregate communication, thereby avoiding the waste of space in the memory area of ​​the processor and improving the processor memory utilization rate.

[0127] In one implementation, a virtual address may include a base address and an address offset.

[0128] The base address can be a pre-set address. Processors in the same communication group have the same base address, which can be represented by V. G This is indicated by the address offset and the access count value recorded in the processor, which are used to determine the physical address of the memory access corresponding to the virtual address in the processor.

[0129] Based on this, in step 602, when accessing data according to the virtual address and the corresponding physical address in the memory region of the second processor, it can be achieved in the following way, such as... Figure 7 As shown:

[0130] Step 701: In the case of the message field type representing the data merging stage, query the corresponding first offset in the address mapping table of the second processor according to the processor identifier of the first processor.

[0131] In this embodiment, an address mapping table is established in the memory area of ​​the second processor (such as the video memory of the XPU). The address mapping table includes multiple mapping relationships, which are mapping relationships between processor identifiers and address offsets. Based on this, the corresponding address offset, i.e., the first offset, can be found according to the processor identifier of the first processor in the third data packet.

[0132] The first offset can be the product of the access count value T and the data block size X, i.e., T×X.

[0133] For example, such as Figure 8 As shown, in the address mapping table of the memory region of the second processor, there are two mapping relationships: Sender_ID1 represents processor 1, and T1×X is the address offset of the data written by processor 1 in the T1th time; Sender_ID2 represents processor 2, and T2×X is the address offset of the data written by processor 2 in the T2th time. Based on this, according to the processor identifier of the first processor, such as processor 1, the corresponding first offset T1×X is found in the address mapping table.

[0134] Step 702: Determine the third physical address based on the first offset and the base address.

[0135] The base address is the address set for each processor in the same communication group, such as V. G Based on this, the third physical address can be the base address plus the first offset, i.e., V. G +T×X.

[0136] Step 703: Read the second target data from the third physical address corresponding to the memory region in the second processor.

[0137] Specifically, in this embodiment, the third physical address, V, is located in the memory region of the second processor. G +T×X, thereby reading the second target data from the third physical address.

[0138] Furthermore, the second processor, acting as the sender in the data merging phase, can generate a new data packet based on the second target data, such as the first data packet mentioned above. The first data packet is then sent to the switch, which aggregates the second target data in all the first data packets to obtain the aggregation result. Based on the aggregation result, a new data packet is generated, namely the second data packet mentioned above. The second data packet is then sent to the first processor, acting as the receiver in the data merging phase, to achieve data merging (such as data protocol).

[0139] Furthermore, in this embodiment, the third data packet also includes a transaction identifier field. This field records the target transaction identifier corresponding to the data distribution stage, which is the identifier connecting the data distribution stage and the data merging stage (essentially a unique key), such as an Instruction ID. Correspondingly, in the address mapping table of the second processor, the mapping relationship is between the processor identifier, the target transaction identifier, and the address offset. In the address mapping table, the target transaction identifier distinguishes the address offsets corresponding to different communication transactions of the same first processor during the data distribution stage.

[0140] Based on this, when querying the corresponding first offset according to the processor identifier corresponding to the first processor in step 701, the corresponding first offset can be queried in the address mapping table of the second processor according to the processor identifier corresponding to the first processor and the target transaction identifier.

[0141] For example, such as Figure 9 As shown, in the address mapping table of the memory region of the second processor, there are three mapping relationships: Sender_ID1 represents processor 1, Instruction ID1 represents data communication transaction 1, Instruction ID2 represents data communication transaction 2, and Instruction ID3 represents data communication transaction 3. T1×X is the address offset of the data written by processor 1 in the T1th operation, T2×X is the address offset of the data written by processor 1 in the T2th operation, Sender_ID represents processor 2, and T3×X is the address offset of the data written by processor 2 in the T3th operation. Based on this, according to the processor identifier of the first processor, such as processor 1, and the target transaction identifier, such as Instruction ID2, the corresponding first offset T2×X is looked up in the address mapping table.

[0142] In one implementation, a virtual address may include a base address and an address offset.

[0143] The base address can be a pre-set address. Processors in the same communication group have the same base address, which can be represented by V. G The address offset represents at least the size X of the data block currently being accessed. The address offset, along with the access count value recorded in the processor, is used to determine the physical address of the memory access in the processor corresponding to the virtual address.

[0144] Based on this, in step 602, when accessing data according to the virtual address and the corresponding physical address in the memory region of the second processor, it can be achieved in the following way, such as... Figure 10 As shown:

[0145] Step 1001: In the case where the message field type represents the data distribution stage (such as the Dispatch stage), determine the second offset based on the address offset and the access count value recorded by the second processor.

[0146] The second offset can be the product of the data block size X in the address offset of the virtual address and the access count value T recorded by the second processor, i.e., T×X. For example, in this embodiment, the tensor index is removed from the address offset in the virtual address to obtain the data block size X, and the current access count value T is queried in the second processor. Then, the data block size X is multiplied by the access count value T to obtain the second offset.

[0147] Step 1002: Determine the first physical address based on the base address and the second offset.

[0148] In this embodiment, the base address can be added to the second offset to obtain the first physical address, i.e., V. G +T×X.

[0149] Step 1003: Write the first target data in the third data packet into the corresponding first physical address in the memory area of ​​the second processor.

[0150] As can be seen, in this embodiment, the access count value recorded in the second processor is mapped to the latest written physical address of the memory region in the second processor. Therefore, each time data is written to the memory region in the second processor, it is written in address order, which can avoid memory fragmentation and improve the space utilization of the memory region.

[0151] Furthermore, after step 1003, in this embodiment, the second offset can be added to the address mapping table according to the processor identifier corresponding to the first processor, and the access count value recorded by the second processor can be incremented by 1.

[0152] In this embodiment, a mapping relationship can be generated according to the processor identifier and the second offset corresponding to the first processor. The mapping relationship is added to the address mapping table. The updated address mapping table and the access count value incremented by 1 can be used to read the relevant data corresponding to the first processor from the memory area in the second processor.

[0153] Based on the above implementation scheme, the third data packet in this embodiment may also include a calculation exception field, such as a Status field. The calculation exception field is used to record whether there is an exception in the data calculation in the second processor. Based on this, in this embodiment, when the calculation exception field indicates that there is an exception in the data calculation in the second processor, the operation of incrementing the access count value by 1 can be canceled (i.e., the auto-increment operation of the access count value is rolled back), and the operation of adding the second offset to the address mapping table can be canceled (i.e., the corresponding mapping relationship in the address mapping table is deleted).

[0154] Therefore, in this embodiment, when data calculation is abnormal, the increment of the access count value is canceled and the corresponding mapping relationship in the address mapping table is cleared to prevent data leakage in the memory area.

[0155] Additionally, the third data packet in this embodiment may also include a data type field, such as a Data Type field, which records the precision of the target data. Based on this, this embodiment can dynamically adjust the data block size X according to the data type field. For example, the FP32 type corresponds to a 4-byte offset, while the INT8 type corresponds to a 1-byte offset.

[0156] refer to Figure 11 This is a schematic diagram of a data processing device in asymmetric set communication provided in an embodiment of this application. This device can be applied to... Figure 2 The switch shown in this embodiment may include the following structure:

[0157] The message receiving unit 1101 is used to receive a first data message sent by a first processor; the first data message includes at least a message type field, a processor identifier corresponding to the first processor, and a virtual address; the message type field is located in the message header of the first data message;

[0158] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication;

[0159] The processor determining unit 1102 is used to determine the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor;

[0160] The message sending unit 1103 is used to send the second data message to the second processor.

[0161] As can be seen from the above technical solution, in the data processing device for asymmetric aggregate communication provided in this application embodiment, the data packet sent by the processor to the switch carries a packet type field and a virtual address, thereby specifying the physical address for the processor to access memory according to the data processing stage. Thus, in asymmetric aggregate communication, it is not necessary to pre-allocate storage space for other processors in the memory area of ​​the processor, effectively solving the technical defects of memory redundancy and idle space waste in aggregate communication, thereby avoiding the waste of space in the memory area of ​​the processor and improving the processor memory utilization rate.

[0162] In one implementation, where the message type field represents the data distribution stage, the second data message also includes first target data;

[0163] The first target data is written to the first physical address in the memory region of the second processor according to the virtual address.

[0164] In one implementation, where the message type field represents the data merging stage, the first data message also includes second target data.

[0165] The device in this embodiment may further include the following structures, such as... Figure 12 As shown:

[0166] The data aggregation unit 1104 is used to aggregate the second target data in all the first data packets sent by the multiple first processors to obtain an aggregation result; generate the second data packet according to the aggregation result; and write the aggregation result into the corresponding second physical address in the memory region of the second processor according to the virtual address.

[0167] In one implementation, where the message type field represents the data merging stage, the first data message further includes a transaction identifier field; the transaction identifier field is used to record the target transaction identifier corresponding to the data distribution stage corresponding to the data merging stage.

[0168] The target transaction identifier and the virtual address are used to indicate reading data from the corresponding third physical address in the memory region of the second processor.

[0169] In one implementation, the processor determining unit 1102 is specifically used to: query the processor identifier of a second processor that is in the same set of communication groups as the first processor in the communication group configuration information registered in the switch, according to the processor identifier corresponding to the first processor; wherein, the processor identifier corresponding to the second processor is used to send the second data packet to the corresponding second processor.

[0170] In one implementation, the virtual address includes: a base address and an address offset;

[0171] Among them, processors belonging to the same communication group have the same base address; the address offset and the access count value recorded in the processor are used to determine the physical address of memory access in the processor corresponding to the virtual address.

[0172] It should be noted that the specific implementation of each unit in this embodiment can be referred to the corresponding content above, and will not be described in detail here.

[0173] refer to Figure 13 This is a schematic diagram of a data processing device in asymmetric set communication provided in an embodiment of this application. This device can be applied to... Figure 2 Any processor in, such as Figure 3 , Figure 4 and Figure 5 The second processor in the device may include the following units:

[0174] The message receiving unit 1301 is used to receive a third data message sent by the switch; the third data message includes at least a message type field, a processor identifier of the first processor, and a virtual address; the message type field is located in the message header of the third data message;

[0175] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication; the processor identifier of the first processor matches the second processor;

[0176] The data access unit 1302 is used to access data according to the virtual address and the corresponding physical address in the memory region of the second processor.

[0177] As can be seen from the above technical solution, in the data processing device for asymmetric aggregate communication provided in this application embodiment, the data packet sent by the switch to the processor carries a message type field, the processor identifier of the first processor, and a virtual address. The processor can access the physical address of memory based on the virtual address and the processor identifier according to the corresponding data processing stage. Thus, in asymmetric aggregate communication, it is not necessary to pre-allocate storage space for other processors in the memory area of ​​the processor, which effectively solves the technical defects of memory redundancy and idle space waste in aggregate communication, thereby avoiding the waste of space in the memory area of ​​the processor and improving the processor memory utilization rate.

[0178] In one implementation, the virtual address includes: a base address and an address offset;

[0179] Based on this, the data access unit 1302 is specifically configured to: when the message field type represents the data merging stage, query the corresponding first offset in the address mapping table of the second processor according to the processor identifier of the first processor; determine the third physical address according to the first offset and the base address; and read the second target data from the corresponding third physical address in the memory area of ​​the second processor.

[0180] The third data message also includes a transaction identifier field; the transaction identifier field is used to record the target transaction identifier corresponding to the data distribution stage corresponding to the data merging stage.

[0181] Based on this, when the data access unit 1302 queries the corresponding first offset in the address mapping table of the second processor according to the processor identifier corresponding to the first processor, it is specifically used to: query the corresponding first offset in the address mapping table of the second processor according to the processor identifier corresponding to the first processor and the target transaction identifier.

[0182] In one implementation, the virtual address includes: a base address and an address offset;

[0183] Specifically, the data access unit 1302 is used to: determine a second offset based on the address offset and the access count value recorded by the second processor when the message field type represents the data distribution stage; determine a first physical address based on the base address and the second offset; and write the first target data in the third data message into the first physical address corresponding to the memory area in the second processor.

[0184] Furthermore, the data access unit 1302 is also configured to: add the second offset to the address mapping table according to the processor identifier corresponding to the first processor; and increment the access count value recorded by the second processor by 1.

[0185] In one implementation, the third data message further includes a calculation exception field;

[0186] The data access unit 1302 is further configured to: cancel the operation of incrementing the access count value by 1 and cancel the operation of adding the second offset to the address mapping table when the calculation exception field indicates that there is an exception in the data calculation in the second processor.

[0187] It should be noted that the specific implementation of each unit in this embodiment can be referred to the corresponding content above, and will not be described in detail here.

[0188] refer to Figure 14 This is a schematic diagram of the structure of a switch provided in an embodiment of this application. The switch may include the following structure:

[0189] The communication module 1401 is used to receive a first data packet sent by a first processor; the first data packet includes at least a message type field, a processor identifier corresponding to the first processor, and a virtual address; the message type field is located in the message header of the first data packet;

[0190] Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication;

[0191] The controller 1402 is used to determine the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor;

[0192] The communication module is also used to send the second data message to the second processor.

[0193] As can be seen from the above technical solutions, in the switch provided by the embodiments of this application, the data packets sent by the processor to the switch carry a message type field and a virtual address, thereby specifying the physical address for the processor to access memory according to the data processing stage. Thus, in asymmetric aggregate communication, it is not necessary to pre-allocate storage space for other processors in the memory area of ​​the processor, effectively solving the technical defects of memory redundancy and idle space waste in aggregate communication, thereby avoiding the waste of space in the memory area of ​​the processor and improving the processor memory utilization rate.

[0194] Taking scale-up networks as an example, the technical solution of this application is illustrated below:

[0195] In scale-up networks, based on the correspondence between the memory addresses of the input and output XPUs in aggregate communication, aggregate communication can be classified into symmetric aggregate communication and asymmetric aggregate communication.

[0196] For asymmetric set communication, such as Figure 2 As shown, taking the Dispatch and Combine phases under MoE (Mixture of Experts) as an example, in the Dispatch phase, each piece of data is copied top-K times and sent to at most top-K XPUs. Here, the memory address of the same data on different XPUs depends on the amount of data sent from other XPUs to that XPU and is dynamically determined. Traditional static pre-allocation of video memory can lead to huge video memory waste and also cause discontinuous memory layout, affecting the execution efficiency of XPU kernel functions.

[0197] To address the aforementioned technical deficiencies, this application proposes the following technical solution:

[0198] 1. Address redirection mechanism for end-to-end network collaboration: A collaborative paradigm based on the extended computing header (ECH) is established between the switch with computing capabilities and the XPU. When the switch performs asymmetric data distribution (Dispatch / Push), the receiving XPU does not write to the packet according to the preset address, but intercepts it through the internally integrated address redirection unit.

[0199] 2. Dynamic continuous arrangement based on atomic counters: The receiving-side XPU maintains a local atomic counter. Map all arriving asynchronous data streams to V G In a continuous video memory space based on +T×X, zero-copy "compact" storage of data is achieved (i.e., zero-copy continuous compact video memory arrangement).

[0200] 3. Bidirectional backtracking logic based on state mapping: An address mapping table is established on the receiving side from "sender information and original offset" to "local actual offset". In the subsequent protocol (Combine / PULL) phase, this table is used to reverse-index the real physical address, ensuring that the switch can accurately extract data from each node and complete the on-network protocol operation.

[0201] Based on this, the technical solution of this application has the following advantages:

[0202] 1. Maximize video memory utilization: The receiving side does not need to reserve static cache space for each potential sender, but instead dynamically increases according to the actual amount of data arriving, completely eliminating video memory fragmentation.

[0203] 2. Zero-copy continuous arrangement: Data is arranged continuously in the receiving side video memory in the order of arrival, so that subsequent calculation operators T can be directly read in vectorized form without the need for expensive memory defragmentation operations.

[0204] 3. High reliability feedback: In conjunction with the Status field in ECH, when an error occurs or a transaction times out during the mapping process, the switch can promptly notify the source end to retransmit or switch paths through the Bitmap field.

[0205] Specifically, the technical solution of this application divides asymmetric collection communication (taking Dispatch as an example) into two key stages: "dynamic redirection writing" and "state mapping backtracking". The following describes the establishment of the collaborative view:

[0206] 1. Before the aggregated communication begins, all participating nodes (i.e., XPUs) determine a globally consistent virtual base address through the control plane (such as a network management interface based on remote procedure calls, abbreviated as GNMI-gRPC, short for Generic Remote Procedure Call Network Management Interface) or a negotiation protocol. .

[0207] 2. Each node's XPU initializes its local atom counter T to 0.

[0208] 3. Register the group ID of the aggregated communication group in the switch and configure the member information of the aggregated communication group, i.e., the communication group configuration information.

[0209] like Figure 15 The diagram shown is a schematic of the ECH standard header structure in this application, wherein the relevant key fields of ECH are as follows:

[0210] (1) Instruction ID (target transaction identifier, i.e., transaction ID): The "unique key" connecting the two asynchronous phases of Dispatch (push) and Combine (pull), serving as the primary key of the address mapping table. One Sender_ID can correspond to multiple transaction IDs. When the receiving XPU records the mapping relationship in the address mapping table, it must bind the Instruction ID so that when a subsequent Combine request arrives, the corresponding local memory address offset can be quickly retrieved by matching the ID.

[0211] (2) Type (message type field): distinguishes different communication semantics (such as Dispatch, Combine). When the Dispatch bit is set, the atomic counter T used to record the access count value is automatically queried; when the Combine bit is set, the resource reservation on the switch side is triggered.

[0212] (3) Status (Status field and exception bit): Handles exceptions during the calculation process. When this field reports an error, the receiving side has an automatic rollback mechanism, that is, it cancels the atomic increment of the counter T and cleans up the invalid entries in the mapping table to prevent memory leaks.

[0213] (4) Data Type: Defines the precision of the reduction calculation (e.g., FP32, BF16, etc.). The address redirection unit dynamically calculates the address step size X based on the Data Type. For example, FP32 corresponds to a 4-byte offset, while INT8 corresponds to a 1-byte offset.

[0214] The following explains the Dispatch and Combine phases:

[0215] refer to Figure 16 Here is a flowchart illustrating the process of dynamically redirecting and writing data during the Dispatch phase:

[0216] 1. Message transmission: The sending-side XPU transmits the tensor data to be distributed (i.e., the data to be written to the XPU). The target data is encapsulated into a data packet with ECH and a corresponding write request is sent to the switch. This request includes the target data to be written, the XPU encoding, and the virtual address. Where K is the tensor index and X is the data block step size (i.e., the data block size).

[0217] 2. Network distribution: The switch parses the multicast header in the ECH (such as id16 of 256MH or XPU) and broadcasts the data packet (containing the target data) to multiple target node XPUs.

[0218] 3. Interception and Redirection: After the receiving XPU receives a data packet, the address resolution unit does not follow the address specified in the packet. Instead of writing, it retrieves the current atomic counter T from the local memory and calculates the actual write address (i.e., the physical memory access address): V G +T×X.

[0219] 4. Mapping relationship storage: X Store the following mapping entries in the local mapping table: X, such as Figure 17 The diagram shown is a block diagram of the receiving side memory address interception and redirection logic.

[0220] 5. Atomic Update: After the write is completed, the atomic counter T performs an atomic increment operation (atomic increment is a thread-safe atomic operation) to reserve continuous address space for the next inbound tensor.

[0221] like Figure 18 The diagram shown is an example of the data backtracking and aggregation process based on mapping in the Combine phase:

[0222] 1. Read Request Distribution: When the system needs to combine data dispatched from the dispatcher, the receiving side (the original dispatcher) initiates a read request with PULL semantics, which includes XPU encoding and virtual address V. G +K×X.

[0223] 2. Reverse Address Lookup: After receiving a read request (a read request broadcast after the switch parses the XPU encoding), the sending side (the receiving side of the original Dispatch) uses the transaction ID and original offset in the data packet to search the mapping table and obtain the corresponding local offset T×X, thereby locating the actual physical address of the video memory using V. G +T×X.

[0224] 3. On-network protocol processing: Each node sends the target data with the real address to the switch. After the switch completes the protocol operation such as aggregation in the buffer, it returns the final result (i.e. the aggregation result) to the receiving side at once, completing the entire asymmetric communication closed loop.

[0225] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0226] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0227] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0228] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A data processing method in asymmetric aggregate communication, applied to a switch, the method comprising: Receive the first data packet sent by the first processor; The first data packet includes at least a packet type field, a processor identifier corresponding to the first processor, and a virtual address; the packet type field is located in the packet header of the first data packet; Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication; The second processor corresponding to the first processor is determined according to the processor identifier corresponding to the first processor; The second data packet is sent to the second processor.

2. The method according to claim 1, wherein when the message type field represents the data distribution stage, the second data message further includes first target data; in, The first target data is written to the first physical address in the memory region of the second processor according to the virtual address.

3. The method according to claim 1, wherein when the message type field represents the data merging stage, the first data message further includes second target data; in, The method further includes: The second target data in all the first data packets sent by multiple first processors is aggregated to obtain an aggregation result; Based on the aggregation result, the second data packet is generated; the aggregation result is written to the corresponding second physical address in the memory region of the second processor according to the virtual address.

4. The method according to claim 1, wherein when the message type field represents a data merging stage, the first data message further includes a transaction identifier field; the transaction identifier field is used to record the target transaction identifier corresponding to the data distribution stage corresponding to the data merging stage; in, The target transaction identifier and the virtual address are used to indicate reading data from the corresponding third physical address in the memory region of the second processor.

5. The method according to claim 1, wherein the virtual address comprises: Base address and address offset; Among them, processors belonging to the same communication group have the same base address; the address offset and the access count value recorded in the processor are used to determine the physical address of memory access in the processor corresponding to the virtual address.

6. A data processing method in asymmetric set communication, applied to a second processor, the method comprising: Receive third data packets sent by the switch; The third data message includes at least a message type field, a processor identifier of the first processor, and a virtual address; The message type field is located in the message header of the third data message; Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication; the processor identifier of the first processor matches the second processor; Data is accessed at the physical address corresponding to the virtual address in the memory region of the second processor.

7. The method according to claim 6, wherein the virtual address comprises: Base address and address offset; Specifically, data access is performed at the physical address corresponding to the virtual address in the memory region of the second processor, including: When the message field type represents the data merging stage, the corresponding first offset is queried in the address mapping table of the second processor according to the processor identifier of the first processor; The third physical address is determined based on the first offset and the base address; Read the second target data from the third physical address corresponding to the memory region in the second processor.

8. The method according to claim 6, wherein the virtual address comprises: Base address and address offset; Specifically, data access is performed at the physical address corresponding to the virtual address in the memory region of the second processor, including: When the message field type represents the data distribution phase, the second offset is determined based on the address offset and the access count value recorded by the second processor; The first physical address is determined based on the base address and the second offset; Write the first target data in the third data packet into the first physical address corresponding to the memory region in the second processor.

9. The method according to claim 8, further comprising: Add the second offset to the address mapping table according to the processor identifier corresponding to the first processor; Increment the access count value recorded by the second processor by 1.

10. A switch, comprising: The communication module is used to receive the first data packet sent by the first processor; The first data packet includes at least a packet type field, a processor identifier corresponding to the first processor, and a virtual address; the packet type field is located in the packet header of the first data packet; Wherein, the virtual address corresponds to the physical address for memory access in the processor; the message type field characterizes the data processing stage of the aggregate communication; The controller is used to determine the second processor corresponding to the first processor according to the processor identifier corresponding to the first processor; The communication module is also used to send the second data message to the second processor.