Data processing method and device, electronic equipment and storage medium
By discarding the buffer request data packet of the receiving agent node in the on-chip network and generating an internal retry response, network congestion and lockup problems are solved, high-priority data packets are transmitted quickly, and network performance and reliability are improved.
Patent Information
- Application Number
- CN202410310864.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-19
AI Technical Summary
In on-chip networks, unbalanced network loads lead to congestion or even lockup. Existing technologies make it difficult to avoid request channel congestion and ensure that high-priority request packets are quickly transmitted to their destinations.
By discarding the request data packet in the receiving buffer and generating an internal retry response in the receiving agent node, it ensures that there are available buffer units in the buffer, prevents network congestion or lockup, and allows high-priority data packets to be transmitted quickly.
It effectively prevents on-chip network congestion and locking, ensures that high-priority data packets can quickly reach their destination even in congested scenarios, and improves network transmission efficiency and reliability.
Smart Images

Figure CN120675950A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a data processing method, a data processing device, an electronic device, and a non-transitory storage medium. Background Art
[0002] In computer networks and on-chip networks, network congestion or deadlock may occur when the data traffic in the network exceeds the processing capacity of network devices (such as routers and switches), resulting in degraded network performance, increased data transmission latency, and even data loss. Network congestion can negatively impact data transmission and communication within the network, reducing network efficiency and reliability, leading to decreased throughput and increased transmission latency. However, application scenarios with higher throughput requirements are more prone to network congestion or deadlock. Preventing network congestion and deadlock is a critical technology in computer networks and on-chip networks. Summary of the Invention
[0003] At least one embodiment of the present disclosure provides a data processing method, which includes: in response to receiving a first request data packet, determining the number of available buffer units in a receiving request buffer in a receiving agent node for receiving the first request data packet; and in response to the number of available buffer units being less than or equal to a first threshold, discarding at least one request data packet stored in the receiving request buffer.
[0004] At least one embodiment of the present disclosure also provides a data processing device, which includes a request agent node and a receiving agent node, the request agent node is configured to initiate a first request data packet, and the receiving agent node is configured to determine the number of available buffer units in a receiving request buffer for receiving the first request data packet in response to receiving the first request data packet; and discard at least one request data packet stored in the receiving request buffer in response to the number of available buffer units being less than or equal to a first threshold.
[0005] At least one embodiment of the present disclosure also provides an electronic device, comprising a memory and a processor, wherein the memory is configured to store computer-executable instructions; and the processor is configured to execute the computer-executable instructions, wherein the computer-executable instructions, when executed by the processor, implement the method described in any of the above embodiments.
[0006] At least one embodiment of the present disclosure further provides a non-transitory storage medium that non-transitorily stores computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the method described in any of the above embodiments is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.
[0008] Figure 1 A schematic diagram of a system on chip (SOC) including a network on chip (NOC) is shown.
[0009] Figure 2 A flow chart of a data processing method provided by at least one embodiment of the present disclosure is shown.
[0010] Figure 3 A schematic diagram showing the working principle of a request proxy node provided by at least one embodiment of the present disclosure is shown.
[0011] Figure 4 A schematic diagram showing the working principle of a receiving agent node provided by at least one embodiment of the present disclosure is shown.
[0012] Figure 5 A schematic diagram of a data processing device provided by at least one embodiment of the present disclosure is shown.
[0013] Figure 6 A schematic diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0014] Figure 7 A schematic diagram of a non-transitory storage medium provided by at least one embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0015] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0016] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by persons of ordinary skill in the field to which this disclosure belongs. The words "first", "second" and similar terms used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0017] In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and components.
[0018] A system-on-chip (SoC) consists of a processor core, memory, input / output interfaces, controllers, and other related functional modules. The Network-on-Chip (NoC) is a crucial component within the SoC. It is a network structure implemented on the chip that connects the various functional modules and processing units within the chip. The NoC is typically used for data transmission and communication within the chip, enabling efficient data exchange and communication. The design and implementation of the NoC has a significant impact on the performance and efficiency of the SoC, as it directly affects the communication speed and data transmission efficiency between the various functional modules.
[0019] With the rapid development of very large-scale integrated circuit (VLSI) technology, the size of processor chips has continued to increase. Multiple processor cores are integrated within a single chip, and on-chip (NOC) networks (NOCs) are used to replace buses to provide data transmission services between processor cores. Therefore, NOCs play a crucial role. However, due to the limited resources on a chip and the local nature of large amounts of data transmission, NOC load imbalances are common in practical applications, leading to network congestion and even lockup.
[0020] Coherent hub interface (CHI)-based interconnects are widely used in chip design. The CHI interconnect provides a comprehensive set of layered specifications for building systems of varying scale, such as system-on-chip (SoC), consisting of multiple components communicating via a scalable coherent interface and on-chip interconnect. For example, ARM's Coherent Mesh Network (CMN) utilizes the CHI architecture specification. CMN communicates based on the fundamental data unit transmitted within the SoC (e.g., flow control unit or flit), with four types of channels: request channel, response channel, snoop channel, and data channel. CMN's flow control utilizes a protocol-layer retry mechanism to allocate bandwidth and resources. Furthermore, CMN integrates end-to-end Quality of Service (QoS) functionality. When a slave node lacks sufficient resources to receive a request, a protocol retry mechanism is provided to indicate resource availability, preventing congestion in the request channel. Slave nodes are responsible for determining and recording the type of protocol credit required to process a request. Another traditional approach to network interconnection is to employ different routing schemes to minimize network congestion. Although these routing algorithms provide an alternative, they are relatively complex to implement, cannot guarantee deadlock-freeness, and cannot provide a path for delivering high-priority request packets to their destination.
[0021] Therefore, the inventors have noticed that there is a need for a technology that avoids the congestion of the request channel of the on-chip network, ensures that the on-chip network is deadlock-free, and provides a path for quickly transmitting high-priority request data packets to the destination.
[0022] At least one embodiment of the present disclosure provides a data processing method. For example, the present disclosure provides a data processing method including: in response to receiving a first request data packet, determining the number of available buffer units in a receive request buffer for receiving the first request data packet in a receiving agent node; and in response to the number of available buffer units being less than or equal to a first threshold, discarding at least one request data packet stored in the receive request buffer.
[0023] In the above-mentioned data processing method provided in the present disclosure, an internal retry mechanism is provided, which discards at least one request data packet in the buffer for receiving or storing requests of the receiving proxy node and generates an internal retry response to the requesting proxy node for the discarded request data packet. After detecting the internal retry response, the requesting proxy node resends the discarded request data packet to the receiving proxy node, thereby ensuring that there is at least one available buffer unit (or available credit) in the buffer for receiving or storing requests of the receiving proxy node, thereby preventing congestion or locking of the interconnected on-chip network, and also ensuring that even in congested scenarios, a path can always be allowed to quickly transmit high-priority request data packets to the destination.
[0024] At least one embodiment of the present disclosure further provides a data processing device, comprising a request agent node and a receiving agent node. The request agent node is configured to initiate a first request packet, and the receiving agent node is configured to: in response to receiving the first request packet, determine the number of available buffer units in a receive request buffer for receiving the first request packet; and in response to the number of available buffer units being less than or equal to a first threshold, discard at least one request packet stored in the receive request buffer.
[0025] The technical effects of the data processing device of the above embodiment of the present disclosure are the same as the technical effects of the above data processing method, and therefore are not described in detail.
[0026] The various embodiments of the present disclosure will be described below with reference to specific examples.
[0027] Figure 1 FIG. 1 shows a schematic diagram of a system on chip (SOC) including a network on chip (NOC). Figure 1 In the SOC shown in FIG, a master device 103 communicates with a slave device 104 via a NOC. The NOC includes a request agent node 101, an interconnect structure, a receiving agent node 102, and a consistency node (HN) 105. The interconnect structure includes multiple routers 106 and connecting lines connecting the multiple routers 106 to enable communication between the routers. The multiple routers 106 can also communicate with each other via a wireless network.
[0028] It should be noted that although the present disclosure Figure 1 The SOC is used as an example for illustration, but this is only for the convenience of explaining the technology of the present disclosure clearly. For example, Figure 1 In fact, it can be a larger computer system. Figure 1 The scope of application is not specifically limited, and any object to which the technical ideas of the present disclosure can be applied is within the scope of protection of the present disclosure.
[0029] For example, the master device 103 can be called a request node (RN), and the master device 103 can be a different host or processor (or core). The request agent node 101 can also be called an initiator agent (IA), and the request agent node 101 can act as a communication bridge between the master device 103 and the interconnection structure, and can receive requests from the master device 103 and respond to the requests. For example, the request agent node 101 can be responsible for processing tasks such as communication requests, routing scheduling and data transmission to ensure that data can be effectively transmitted in the network. The router 106 can be used to connect the various components on the NOC and provide routing functions between the various components. For example, the router 106 can be used as a communication bridge between the request agent node 101, the receiving agent node 102 and the consistency node 105, and provide communication routing functions between the request agent node 101, the receiving agent node 102 and the consistency node 105.
[0030] For example, the consistency node 105 may also be referred to as a home node (HN). The consistency node 105 may be responsible for maintaining the consistency of the cache in the system (e.g., a system on a chip). The consistency node 105 may act as a receiving agent node when receiving a request data packet (or flit) from the request agent node 101, and may act as a request agent node when receiving a request data packet from the receiving agent node 102. The consistency node 105 may refer to a node responsible for maintaining data consistency in a distributed system or a multi-processor (multi-core) system to ensure data access consistency between multiple processors or storage units to avoid problems caused by data inconsistency. The consistency node 105 may handle tasks such as cache consistency protocols, memory access requests, and data updates to ensure data consistency in the system.
[0031] The receiving agent node 102 can also be called a target agent (TA). The receiving agent node 102 can act as a communication bridge between the interconnect structure and the slave device 104, and can receive requests from the consistency node 105 and the request agent node 101 and respond to the request. For example, the receiving agent node 102 can be used to receive and process communication requests from the request agent node 101 in the on-chip network, and is responsible for receiving data, processing and responding, and coordinating with the request agent node 101 to ensure that data can be effectively transmitted and processed in the network. The slave device 104 can also be called a receiving node (SN). For example, the slave device 104 can be a storage (e.g., memory, cache, etc.) or a device (e.g., various communication devices).
[0032] For example, the receiving agent node 102 and the requesting agent node 101 communicate through the interconnection structure. The requesting agent node 101 is located between the interconnection structure and the master device 103 that initiates the request data packet and is used to send the request data packet to the receiving agent node. The receiving agent node 102 is located between the interconnection structure and the slave device 104 that receives the request data packet and is used to receive the request data packet from the master device 103 from the interconnection structure.
[0033] It should be noted that Figure 1 The mesh structure of the NOC shown in the figure is merely exemplary. The NOC structure may also be of various other types, depending on specific applications and design requirements. For example, the NOC may also include, but is not limited to, a tree structure in which nodes are connected in a tree-like hierarchy, an adaptive structure in which the connections and routes between nodes are dynamically adjusted based on the communication load and system status, or a custom structure created by the designer. Figure 1 The number of various components in the SOC shown in FIG1 (e.g., master device 103, slave device 104, request agent node 101, interconnect structure, receiving agent node 102, consistency node 105, etc.) can be adjusted according to actual needs. Figure 1 Although introduced as SOC, the technology disclosed in this disclosure is not limited to SOC. Figure 1 The schematic diagram shown in FIG may also be a distributed system including multiple hosts.
[0034] Figure 2 FIG. 1 is a flow chart showing a data processing method provided by at least one embodiment of the present disclosure. Figure 2 As shown, in some embodiments of the present disclosure, the data processing method includes the following steps S101-S102.
[0035] Step S101 : in response to receiving a first request data packet, determining the number of available buffer units in a receiving request buffer in a receiving proxy node for receiving the first request data packet.
[0036] Step S102: In response to the number of available buffer units being less than or equal to a first threshold, discarding at least one request data packet stored in the received request buffer.
[0037] Figure 3 A schematic diagram showing the working principle of a request proxy node provided by at least one embodiment of the present disclosure is shown. Figure 4 A schematic diagram showing the working principle of a receiving agent node provided by at least one embodiment of the present disclosure is shown.
[0038] like Figure 3 The request proxy node shown can be Figure 1 The request agent node 101 or consistency node 105 shown in Figure 4 The receiving agent node shown can be Figure 1 The receiving agent node 102 or the consistency node 105 shown in FIG. For example, the consistency node 105 may include Figure 3 The request agent node shown in Figure 4 For example, the consistency node 105 can implement the function of the request proxy node with the request proxy node functional module included therein, and implement the function of the receiving proxy node with the receiving proxy node functional module included therein.
[0039] When a core or processing unit needs to communicate or transmit data, it sends a request (e.g., a request data packet or a bus transmission request (ftr flit, Fabric TransmitRequest flit)) to the request agent node. For example, the request may include information such as the data source address, the address of the target node, the data type, and the transmission method. For example, Figure 3 As shown in FIG, after the request agent node receives the request data packet from the master device, the request data packet may be stored in the cache 305 and then sent to the master device. Figure 4 For example, the request proxy node may also back up the received request data packet in its request table 304, and the backup data may not be cleared until a response to the request data packet returned from the receiving proxy node indicates that the backup data can be released. For example, after the request proxy node sends a request data packet to the receiving proxy node for processing, if it receives a response from the receiving proxy node indicating that the request data packet has been received and processed (for example, Figure 3 frp flit in the request table 304), the request proxy node may clear or release the data corresponding to the request data packet backed up in the request table 304.
[0040] For example, when the requesting agent node receives a response to the request packet from the receiving agent node (e.g., Figure 3 After receiving the response data from the master device, the request agent node can store the response data in the cache 301, and then decode the response through the response decoder 302 to obtain the data information in the response. For example, the request agent node can return the response flit (for example Figure 3 The normal response flit in the request packet is used to feedback the processing status of the request data packet to the master device.
[0041] like Figure 4 The receiving agent node shown can be Figure 1 For example, in response to receiving a request packet (e.g., Figure 4 frr flit in the receiving proxy node), the receiving proxy node determines the number of available buffer units in the receiving request buffer in the receiving proxy node for receiving the request data packet. For example, the available buffer unit can be the size of the available space in the buffer for receiving or storing the request data packet received from the request proxy node (for example, indicating the number of flits that can be stored). For example, the receiving request buffer includes a first receiving request buffer (for example, cache 401) and a second receiving request buffer (for example, cache 402), the first receiving request buffer is used to receive the request data packet and cache at least part of the received request data packet in the second receiving request buffer. It should be noted that, although Figure 4 The receiving request buffer specifically shows the case where it includes two caches, cache 401 and cache 402, but the present disclosure is not limited to this. The receiving request buffer may also include only one cache or more than two caches. The number of caches included in the receiving request buffer can be set according to actual needs.
[0042] For example, the storage capacity of the first receive request buffer can be relatively small, such as being able to store two or three flits. The present disclosure does not impose any restrictions on the storage capacity of the first receive request buffer, and the storage capacity of the first receive request buffer can also be a size capable of storing any number of flits other than two or three flits. For example, the storage capacity of the first receive request buffer can be smaller than the storage capacity of the second receive request buffer to ensure the speed of reading cached request packets in the first receive request buffer. The larger storage capacity of the second receive request buffer can also enable receiving or storing more request packets or flits from the request proxy node, thereby alleviating or preventing interconnect back pressure in the NOC.
[0043] For example, in response to the number of available buffer units being less than or equal to a first threshold, the receiving agent node discards at least one request data packet stored in the receive request buffer. For example, the following description assumes that the first threshold is 1, but the present disclosure is not limited thereto. The first threshold may also be any other suitable value, such as 2, 3, or a larger or smaller value.
[0044] For example, in response to the number of available buffer units in the first receive request buffer being less than or equal to a first threshold and no free space in the second receive request buffer, at least one request data packet stored in the second receive request buffer is discarded.
[0045] For example, in response to the number of available buffer units being less than or equal to 1, the receiving proxy node discards at least one request data packet stored in the receive request buffer. For example, when the receiving proxy node receives a request data packet sent by the requesting proxy node, if the receiving proxy node determines that the number of available buffer units or the number of data packets that can be stored in cache 401 is 1 (for example, a number of 1 indicates that cache 401 is in a near-full state, but the present disclosure does not impose any particular limitation on the value of the number of available buffer units or the number of data packets that can be stored in cache 401 when it is in a near-full state, and the value may be any other suitable value other than 1, for example, 2 or 3 or greater), and there is no more available buffer units or space for data packets in cache 402 (this indicates that cache 402 is in a full state or is completely full), then the receiving proxy node discards the at least one request data packet stored in cache 402. For example, the receiving proxy node sends an internal retry signal to cache 402 via internal retry generation unit 403 to notify cache 402 to discard the at least one request data packet. For example, the receiving agent node sends one or more data packets from the cache 402 to the internal retry generation unit 403 so that the cache 402 has an available buffer unit to receive or store the request data packet from the cache 401, thereby ensuring that the cache 401 normally maintains at least one available buffer unit.
[0046] For example, after cache 402 sends the request data packet to the destination logic for processing through the normal request flit processing path, if the destination logic is in a full state (i.e., there are no idle logical units to receive and process the request data packet from cache 402), the destination logic can feedback an OCN (outstanding) full state to the internal retry generation unit 403. For example, when the internal retry generation unit 403 determines that the destination logic's OCN is full, cache 402 is full, and cache 401 is almost full, the internal retry generation unit 403 sends an internal retry signal to cache 402 to notify cache 402 to abandon at least one request data packet.
[0047] For example, if cache 401 normally maintains at least one available buffer unit, the receiving proxy node can receive any request data packet from the requesting proxy node at any time, and if cache 401 is almost full, it can cache the received or cached request data packet in cache 402. For example, if cache 402 needs to receive a request data packet from cache 401 and cache 402 is full, the receiving proxy node can discard at least one request data packet in cache 402 to make buffer space for the request data packet from cache 401 that needs to be received or cached.
[0048] For example, cache 402 can send the cached request data packet to the destination of the response for processing. For example, as shown in Figure 4, cache 402 can send the cached request data packet to the destination or corresponding logic for processing through a conventional request flit processing path.
[0049] For example, since the cache 401 can normally maintain at least one available buffer unit, the receiving proxy node can receive the request data packet from the requesting proxy node at any time. Figure 4 As shown in , when cache 401 receives a request data packet with a high priority from a request proxy node, the high priority request data packet does not need to be cached in cache 402, but the high priority request data packet is sent to the destination for processing via the high priority flit path. For example, the storage space of cache 401 can be very small. For example, the storage space or the number of available buffer units of cache 401 can be 2 or 3, but the embodiments of the present disclosure are not limited thereto, and the storage space or the number of available buffer units of cache 401 can be set to any other value as needed. Here, since the storage space of cache 401 is very small, the speed of reading the request data packet (e.g., a high priority request data packet) from cache 401 is very fast, thereby allowing the high priority request data packet to be quickly transmitted to the path of the destination for processing.
[0050] For example, the number of available buffer units can be the minimum storage unit in the cache. For example, one minimum storage unit can store one request data packet or flit. For example, a flit path with high priority or criticality will always be non-blocking, thereby ensuring that a high priority or critical flit will never be blocked (i.e., it can definitely be received by the receiving agent node and sent to the destination for processing). For example, a high priority flit can be a flit related to the read / write operation of the CSR (control status register). It should be noted that the present disclosure does not impose any specific restrictions on high priority. In actual circumstances, whether a request data packet belongs to a priority request data packet can be flexibly set according to different standards.
[0051] For example, in response to the number of available buffer units being less than or equal to a first threshold, at least one request data packet stored earliest in the received request buffer is discarded. For example, in response to the number of available buffer units being less than or equal to the first threshold, at least one request data packet stored earliest in cache 402 is discarded. For example, if cache 402 is a FIFO (First In, First Out) queue, when new data enters cache 402, it is added to the end of the queue, and when data needs to be deleted or discarded, it is always deleted or discarded starting from the head of the queue. For example, when at least one request data packet stored in the received request buffer needs to be discarded, the data packet that entered cache 402 earliest is discarded first. It should be noted that the present disclosure does not impose any restrictions on the type of cache 402, as long as it has storage functionality. For example, when at least one request data packet stored in the received request buffer needs to be discarded, the discarded request data packet can be any location in cache 402, and data in cache 402 can be discarded in any predetermined order or priority. The present disclosure does not impose any specific restrictions on this, as long as cache 402 has the function of discarding stored data packets.
[0052] For example, in response to at least one request data packet stored in the receive request buffer being discarded, a retransmission request data packet is generated based on the discarded request data packet and the retransmission request data packet is sent to the request proxy node. For example, after receiving at least one request data packet discarded by cache 402, internal retry generation unit 403 swaps the source address and destination address of each of the at least one discarded request data packet to generate a retransmission request data packet. For example, internal retry generation unit 403 sends the generated retransmission request data packet to the request proxy node based on the swapped destination address. For example, internal retry generation unit 403 first caches the generated retransmission request data packet in cache 404 based on the swapped destination address, and the receiving proxy node then sends the retransmission request data packet to the request proxy node via cache 404.
[0053] For example, in a case where the data to be sent from the receiving proxy node to the requesting proxy node includes a retransmission request data packet and a response data packet, the retransmission request data packet is preferentially sent to the requesting proxy node relative to the response data packet, wherein the response data packet is the response data of the receiving proxy node to the received request data packet. For example, in a case where the data to be sent to the requesting proxy node cached in the cache 404 includes a retransmission request data packet and a response data packet, the retransmission request data packet is preferentially sent to the requesting proxy node relative to the response data packet. For example, the response data packet can be the response data of the receiving proxy node to the received request data packet. For example, the response data packet can be the response data of the receiving proxy node to the requesting proxy node of the processing result of the request data. For example, preferentially sending the retransmission request data packet can ensure that the discarded request packet is quickly retransmitted and processed, so as to minimize the impact on the data processing flow of the master device.
[0054] For example, after the request agent node receives the retransmission request data packet, it generates a retry request data packet based on the retransmission request data packet and sends the retry request data packet to the receiving agent node for processing. Figure 3 As shown, when the requesting agent node receives a retransmission request packet (e.g., a frp flit) from the receiving agent node, it first buffers the packet in the buffer 301 and then sends the packet to the response decoder 302 for decoding. For example, when the response decoder 302 determines that the data type of the request packet is a retransmission request packet, it sends the retransmission request packet to the request retry unit 303.
[0055] For example, the request retry unit 303 exchanges the source address and the destination address in the retransmission request data packet to generate a retry request data packet. For example, after receiving the retransmission request data packet, the request retry unit 303 exchanges the source address and the destination address of the retransmission request data packet to generate a retry request data packet and notifies the request table 304 to resend the request data packet corresponding to the retransmission request data packet to the receiving proxy node, so as to send the request data packet discarded by the receiving proxy node to the receiving proxy node for reprocessing. Alternatively, in other embodiments, the receiving proxy node can generate a retransmission request data packet based on the identification information (e.g., an identification number) of the discarded request data packet and send the retransmission request data packet to the request proxy node. After receiving the retransmission request data packet, the request proxy node searches for the request data packet corresponding to the identification information in the retransmission request data packet from the stored request data packets and retransmits it.
[0056] It should be noted that although Figure 3 and Figure 4The present disclosure specifically describes how different units or modules in the request proxy node and the receiving proxy node perform different operations. However, this description is intended only to facilitate those skilled in the art in understanding the technical concepts of the present disclosure. For example, in actual scenarios, an operation step implemented by each of the request proxy node and the receiving proxy node may actually require the cooperation of more functional units or modules, and multiple operation steps implemented by each of the request proxy node and the receiving proxy node may actually be completed by a single functional unit or module. The present disclosure does not impose any restrictions on the specific structures of the request proxy node and the receiving proxy node, nor on which parts thereof perform which operations, as long as they can each implement the technical concepts of the various method steps provided in the present disclosure.
[0057] For example, if the data to be sent from the requesting proxy node to the receiving proxy node includes a retry request packet and other request packets, the retry request packet is sent to the receiving proxy node with priority over the other request packets, where the other request packets include request packets received from the requesting proxy node from the master device. For example, request table 304 in the requesting proxy node sends the retry request packet to cache 305, which then sends the retry request packet to the receiving proxy node via cache 305. For example, if the data to be sent to the receiving proxy node cached in cache 305 includes a retry request packet and other request packets, the retry request packet is sent to the receiving proxy node with priority over the other request packets. For example, the other request packets may include request packets received from the master device by the requesting proxy node. For example, sending the retry request packet with priority can ensure that discarded request packets are quickly sent to the receiving proxy node for processing, thereby minimizing the impact on the data processing flow of the master device.
[0058] In the above-mentioned data processing method provided in the present disclosure, an internal retry mechanism is provided, which discards at least one request data packet in the buffer for receiving or storing requests of the receiving proxy node and generates an internal retry response to the requesting proxy node for the discarded request data packet. After detecting the internal retry response, the requesting proxy node resends the discarded request data packet to the receiving proxy node, thereby ensuring that there is at least one available buffer unit (or available credit) in the buffer for receiving or storing requests of the receiving proxy node, thereby preventing congestion or locking of the interconnected on-chip network, and also ensuring that even in congested scenarios, a path can always be allowed to quickly transmit high-priority request data packets to the destination.
[0059] Figure 5A schematic diagram of a data processing device 50 provided by at least one embodiment of the present disclosure is shown. The data processing device 50 may include a request agent node 501 and a receiving agent node 502. For example, the request agent node 501 may be configured to initiate a first request packet. For example, the receiving agent node 502 may be configured to, in response to receiving the first request packet, determine the number of available buffer units in a receive request buffer for receiving the first request packet; and, in response to the number of available buffer units being less than or equal to a first threshold, discard at least one request packet stored in the receive request buffer.
[0060] For example, the receiving agent node 502 may be configured to discard at least one request data packet stored earliest in the receiving request buffer in response to the number of available buffer units being less than or equal to a first threshold.
[0061] For example, the receive request buffer may include a first receive request buffer and a second receive request buffer, wherein the first receive request buffer is configured to receive request packets and cache at least a portion of the received request packets in the second receive request buffer. For example, the receiving agent node 502 may be configured to discard at least one request packet stored in the second receive request buffer in response to the number of available buffer units in the first receive request buffer being less than or equal to a first threshold and no free space in the second receive request buffer. For example, the first threshold may be 1.
[0062] For example, receiving proxy node 502 is further configured to: in response to at least one request data packet stored in the receive request buffer being discarded, generate a retransmission request data packet based on the discarded request data packet and send the retransmission request data packet to request proxy node 501. For example, request proxy node 501 may also be configured to: upon receiving the retransmission request data packet, generate a retry request data packet based on the retransmission request data packet and send the retry request data packet to receiving proxy node 502 for processing. For example, receiving proxy node 502 and request proxy node 501 communicate via an interconnect structure. Request proxy node 501 is located between the interconnect structure and a master device that initiates a request data packet and is configured to send the request data packet to receiving proxy node 502. Receiving proxy node 502 is located between the interconnect structure and a slave device that receives the request data packet and is configured to receive the request data packet from the master device via the interconnect structure.
[0063] For example, the receiving proxy node 502 may be configured to, when the data to be sent to the requesting proxy node 501 includes a retransmission request data packet and a response data packet, send the retransmission request data packet to the requesting proxy node 501 in priority over the response data packet. For example, the response data packet may be response data of the receiving proxy node 502 to the received request data packet.
[0064] For example, the request agent node 501 may be configured to, when the data to be sent to the receiving agent node 502 includes a retry request packet and other request packets, prioritize the retry request packet over the other request packets when sending the retry request packet to the receiving agent node 502. For example, the other request packets may include request packets received by the request agent node 501 from the master device.
[0065] For example, the receiving agent node 502 may be configured to swap the source address and the destination address in the discarded request data packet to generate a retransmission request data packet.
[0066] For example, the request agent node 501 may be configured to: swap the source address and the destination address in the retransmission request data packet to generate a retry request data packet.
[0067] It should be noted that Figure 5 The components and structure of the data processing device 50 shown are merely exemplary and non-limiting. The data processing device 50 may further include other components and structures as needed. The data processing device 50 may include more or fewer nodes or units, and the connections between the nodes or units are not limited and may be determined based on actual needs. For example, the data processing device 50 may further include a requesting node and a receiving node.
[0068] In the above-mentioned data processing device provided by the present disclosure, an internal retry mechanism is provided, which discards at least one request data packet in the buffer for receiving or storing requests of the receiving proxy node and generates an internal retry response to the requesting proxy node for the discarded request data packet. After the requesting proxy node detects the internal retry response, it resends the discarded request data packet to the receiving proxy node. It can ensure that there is at least one available buffer unit (or available credit) in the buffer for receiving or storing requests of the receiving proxy node, thereby preventing congestion or locking of the interconnected on-chip network, and can also ensure that even in congested scenarios, a path can always be allowed to quickly transmit high-priority request data packets to the destination.
[0069] At least some embodiments of the present disclosure further provide an electronic device comprising a processor and a memory. For example, the memory is configured to store computer-executable instructions. For example, the processor is configured to execute the computer-executable instructions. For example, when the computer-executable instructions are executed by the processor, the data processing method provided in at least one embodiment of the present disclosure is implemented.
[0070] Figure 6 A schematic diagram of an electronic device provided by at least one embodiment of the present disclosure is shown.
[0071] like Figure 6As shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601 and a memory 602 , and the processor 601 and the memory 602 may be interconnected via a bus 603 .
[0072] The processor 601 can perform various actions and processes according to the program or code stored in the memory 602. Specifically, the processor 601 can be an integrated circuit chip with signal processing capabilities. For example, the above-mentioned processor 601 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the various methods and steps disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture, an ARM architecture, a RISC-V architecture, etc.
[0073] The memory 602 is used to non-temporarily store computer-executable instructions, and the processor 601 is used to run the computer-executable instructions. When the computer-executable instructions are executed by the processor 601, the data processing method provided by at least one embodiment of the present disclosure is implemented.
[0074] For example, memory 602 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. Non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0075] In the above-mentioned electronic device provided in the present disclosure, an internal retry mechanism is provided, which discards at least one request data packet in the buffer for receiving or storing requests of the receiving proxy node and generates an internal retry response to the requesting proxy node for the discarded request data packet. After the requesting proxy node detects the internal retry response, it resends the discarded request data packet to the receiving proxy node. This can ensure that there is at least one available buffer unit (or available credit) in the buffer for receiving or storing requests of the receiving proxy node, thereby preventing congestion or locking of the interconnected on-chip network, and can also ensure that even in congested scenarios, a path can always be allowed to quickly transmit high-priority request data packets to the destination.
[0076] At least one embodiment of the present disclosure further provides a non-transitory storage medium that non-transitorily stores computer-executable instructions. For example, when the computer-executable instructions are executed by a processor, the data processing method provided by at least one embodiment of the present disclosure is implemented.
[0077] Figure 7 Schematic diagram of a non-transitory storage medium provided by some embodiments of the present disclosure. Figure 7 As shown, the non-transitory storage medium 700 can non-transitory store computer-executable instructions 710 , which implement the data processing method provided by any embodiment of the present disclosure when executed by a computer.
[0078] Similarly, the non-transitory storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0079] The technical effects of the above-mentioned non-transitory storage medium are the same as the technical effects of the above-mentioned data processing method, and will not be repeated here.
[0080] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the module, program segment, or part of the code contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0081] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0082] In addition to the above exemplary description, the following points need to be explained for this disclosure:
[0083] (1) The drawings of the embodiments of the present disclosure only relate to the structures related to the embodiments of the present disclosure. Other structures may refer to conventional designs.
[0084] (2) For the sake of clarity, the thickness and size of layers or structures in the drawings used to describe the embodiments of the present disclosure are exaggerated. It will be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly on" or "under" the other element, or intervening elements may be present.
[0085] (3) In the absence of conflict, the embodiments of the present disclosure and the features therein may be combined with each other to form new embodiments.
[0086] The above description is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be based on the protection scope of the claims.
Claims
1. A data processing method, comprising: In response to receiving a first request data packet, determining a number of available buffer units in a receive request buffer in a receiving agent node for receiving the first request data packet; as well as In response to the number of available buffer units being less than or equal to a first threshold, at least one request data packet stored in the receive request buffer is discarded.
2. The data processing method according to claim 1, wherein: The step of discarding at least one request data packet stored in the receive request buffer in response to the number of available buffer units being less than or equal to a first threshold comprises: In response to the number of available buffer units being less than or equal to a first threshold, at least one request data packet stored earliest in the receive request buffer is discarded.
3. The data processing method according to claim 1, wherein: The receiving request buffer includes a first receiving request buffer and a second receiving request buffer, the first receiving request buffer being used to receive the first request data packet and cache at least part of the received request data packets in the second receiving request buffer, and discarding at least one request data packet stored in the receiving request buffer in response to the number of available buffer units being less than or equal to a first threshold, comprising: In response to the number of available buffer units in the first receive request buffer being less than or equal to a first threshold and no free space existing in the second receive request buffer, at least one request data packet stored in the second receive request buffer is discarded.
4. The data processing method according to claim 1, wherein: The first threshold is 1.
5. The data processing method according to claim 1, further comprising: In response to at least one request data packet stored in the received request buffer being discarded, generating a retransmission request data packet based on the discarded request data packet and sending the retransmission request data packet to a request proxy node; When the request proxy node receives the retransmission request data packet, it generates a retry request data packet based on the retransmission request data packet and sends the retry request data packet to the receiving proxy node for processing, wherein the receiving proxy node and the request proxy node communicate with each other through an interconnection structure, the request proxy node is located between the interconnection structure and the master device that initiates the first request data packet and is used to send a request data packet to the receiving proxy node, and the receiving proxy node is located between the interconnection structure and the slave device that receives the first request data packet and is used to receive the request data packet from the master device from the interconnection structure.
6. The data processing method according to claim 5, wherein: In response to at least one request data packet stored in the received request buffer being discarded, generating a retransmission request data packet based on the discarded request data packet and sending the retransmission request data packet to the request proxy node, comprising: In the case that the data to be sent by the receiving proxy node to the requesting proxy node includes a retransmission request data packet and a response data packet, the retransmission request data packet is sent to the requesting proxy node in priority to the response data packet, wherein the response data packet is the response data of the receiving proxy node to the received request data packet.
7. The data processing method according to claim 5, wherein: The method further comprises: generating a retry request packet based on the retry request packet after the request proxy node receives the retry request packet and sending the retry request packet to the receiving proxy node for processing. In a case where the data to be sent by the requesting proxy node to the receiving proxy node includes a retry request data packet and other request data packets, the retry request data packet is sent to the receiving proxy node in priority to the other request data packets, wherein the other request data packets include the request data packets received by the requesting proxy node from the master device.
8. The data processing method according to claim 5, wherein: The generating of a retransmission request data packet based on the discarded request data packet comprises: The source address and the destination address in the discarded request data packet are swapped to generate the retransmission request data packet.
9. The data processing method according to claim 5, wherein: Generating a retry request data packet based on the retransmission request data packet includes: The source address and the destination address in the retransmission request data packet are exchanged to generate the retry request data packet.
10. A data processing device comprising: a request proxy node configured to initiate a first request data packet; a receiving agent node configured to, in response to receiving the first request data packet, determine a number of available buffer units in a receive request buffer for receiving the first request data packet; And in response to the number of available buffer units being less than or equal to a first threshold, discarding at least one request data packet stored in the receive request buffer.
11. The data processing apparatus according to claim 10, wherein: The receiving agent node is configured to discard at least one request data packet stored earliest in the receiving request buffer in response to the number of available buffer units being less than or equal to a first threshold.
12. The data processing apparatus according to claim 10, wherein: The receiving request buffer includes a first receiving request buffer and a second receiving request buffer, wherein the first receiving request buffer is used to receive the first request data packet and cache at least part of the received request data packet in the second receiving request buffer, and The receiving agent node is configured to discard at least one request data packet stored in the second receiving request buffer in response to the number of available buffer units in the first receiving request buffer being less than or equal to a first threshold and there being no free space in the second receiving request buffer.
13. The data processing apparatus according to claim 10, wherein: The first threshold is 1.
14. An electronic device comprising: a memory configured to store computer-executable instructions; as well as a processor configured to execute the computer-executable instructions, When the computer executable instructions are executed by the processor, the method according to any one of claims 1 to 9 is implemented.
15. A non-transitory storage medium that non-transitory stores computer-executable instructions, wherein: When the computer-executable instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Cited By
Data processing method and apparatus, electronic device, and storage medium
EP4750143A1
Data processing method and apparatus, electronic device, and storage medium
WO2025195330A1