Inter-core communication system and chip circuit
By designing a distributed inter-core communication system in SOC, the message cache module and the message routing module are used to achieve efficient message delivery between processor cores, which solves the problem of low inter-core communication efficiency in heterogeneous systems, improves the operation efficiency of SOC and has good scalability.
Patent Information
- Application Number
- CN202510174287.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
In heterogeneous systems, the communication efficiency between processor cores is low, resulting in limited SOC operation efficiency, and the existing shared memory technology has problems such as read and write delay and cumbersome operation processes.
A distributed architecture-based inter-core communication system is designed, and message transmission between processor cores is realized by introducing a message cache module and a message routing module in each subsystem. The message cache module and the processor core are connected through a bus. The transmission link of the message passing from the processor core to the message cache module and from the message cache module to the processor core is short, which reduces the bus transmission pressure. The message routing module is responsible for routing messages from the subsystem where the producer core is located to the subsystem where the consumer core is located.
Through the distributed architecture, the delay of inter-core communication is reduced, the operation efficiency of SOC is improved, the load on the processor is reduced, and the scalability is good.
Smart Images

Figure CN120104560A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of system-on-chip (SOC), and in particular to an inter-core communication system and a chip circuit. Background Art
[0002] With the rapid development of semiconductor technology, the integration scale of SOC has shown a trend of continuous expansion, and its functions have become more and more diverse. At present, in order to meet a variety of complex application scenarios and diversified needs, SOC chips are often divided into multiple subsystems according to their functions. Each subsystem is equipped with a processor and a corresponding processor core, and each subsystem cooperates with each other to form a heterogeneous system. In such a heterogeneous system, the communication efficiency between each processor core becomes a key factor affecting the operating efficiency of the entire SOC. Summary of the invention
[0003] One aspect of the present application provides an inter-core communication system, comprising: one or more subsystems, each subsystem comprising a message cache module and one or more processor cores each connected to the message cache module via a bus, each processor core being a producer core or a consumer core, the producer core being configured to write a delivery message to the message cache module integrated in the subsystem where it is located, and the consumer core being configured to read the delivery message from the message cache module integrated in the subsystem where it is located; and one or more message routing modules, configured to: obtain a delivery message written by the production core to the message cache module integrated in the subsystem where the producer core is located, and route the obtained delivery message to the message cache module integrated in the subsystem where the consumer core is located.
[0004] In some embodiments, each of the message cache modules includes: one or more cache processing units, which correspond to one or more processing cores in the subsystem where the message cache module is located, each cache processing unit includes a sending buffer and a receiving buffer, and an internal routing unit, configured to: obtain the delivery message, and route the delivery message to the receiving buffer of the cache processing unit included in the message cache module where the internal routing unit is located; wherein the consumer core is also configured to: read the delivery message from the receiving buffer in the corresponding cache processing unit.
[0005] In some embodiments, the message cache module is also configured to: write the delivery message into the sending buffer in the cache processing unit corresponding to the producer core, and the sending buffer is configured to pass the delivery message to the internal routing unit; the internal routing unit is also configured to: obtain the delivery message from the sending buffer in the cache processing unit, and when the producer core and the consumer core are located in the same subsystem, route the obtained delivery message to the receiving buffer in the cache processing unit corresponding to the consumer core; when the producer core and the consumer core are located in different subsystems, route the obtained delivery message to the message routing module; the message routing module is also configured to route the obtained delivery message to the internal routing unit in the message cache module in the subsystem where the consumer core is located.
[0006] In some embodiments, the internal routing unit is further configured to obtain the delivery message from the message routing module and route the delivery message to a receiving buffer in a cache processing unit corresponding to the consumer core.
[0007] In some embodiments, the cache processing unit is also configured to, in response to the number of delivery messages in the receiving cache in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal to a processor core corresponding to the cache processing unit, so that the processor core acts as the consumer core to read the delivery message from the receiving cache.
[0008] In some embodiments, the delivery message includes: a producer core identification bit for identifying the producer core from which the delivery message comes, and a consumer core identification bit for identifying the consumer core to which the delivery message is to be sent.
[0009] In some embodiments, the internal routing unit is further configured to discard the delivery message when both the producer core indicated by the producer core identification bit in the delivery message and the consumer core indicated by the consumer identification bit in the delivery message are not in the subsystem where the internal routing unit is located.
[0010] In some embodiments, the producer core identification bit is also used to identify the cache processing module included in the subsystem where the producer core is located and the cache processing unit corresponding to the producer core, and the consumer core identification bit is also used to identify the cache processing module included in the subsystem where the consumer core is located and the cache processing unit corresponding to the consumer core.
[0011] In some embodiments, the cache processing unit is further configured to convert the format of the transmitted message in the receiving cache in the cache processing unit according to the width of the bus connecting the message cache module where the cache processing unit is located and the one or more processor cores.
[0012] In some embodiments, the cache processing unit also includes: one or more message filters, each message filter corresponds to a different cache address segment in the receiving cache; the one or more message filters are each configured to: receive the delivery message from the internal routing unit, filter the received delivery message based on their own filtering rules, and write the filtered delivery message into the corresponding cache address segment.
[0013] In some embodiments, the cache processing unit is further configured to: receive configuration parameters, and set the message filters actually enabled, the filtering rules of each message filter, and the range of each cache address segment based on the configuration parameters.
[0014] In some embodiments, each of the message filters includes one or more filtering units, and each filtering unit filters different parts of the transmitted message based on its own filtering algorithm to obtain a unit filtering result of each filtering unit; the cache processing unit is also configured to: set the filtering unit actually enabled in each message filter based on the configuration parameters, and apply the rules of the unit filtering results of each filtering unit in the message filter, thereby setting the filtering rules of each message filter.
[0015] In some embodiments, the cache processing unit is further configured to: in response to the number of transfer messages in any cache address segment of the receiving cache in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal corresponding to the cache address segment to the processor core corresponding to the cache processing unit, so that the processor core acts as the consumer core to read the transfer message from the cache address segment corresponding to the interrupt signal in the receiving cache; interrupt signals corresponding to different cache address segments have different interrupt priorities.
[0016] In some embodiments, the cache processing unit is further configured to: calculate a send verification code based on the delivery message written into the sending buffer, and attach the send verification code to the delivery message; calculate a receive verification code based on the delivery message written into the receiving buffer, and compare the receive verification code with the send verification code attached to the delivery message to determine whether a transmission error occurs in the delivery message.
[0017] In some embodiments, multiple subsystems are divided into one or more subsystem clusters according to the functions they implement and their locations. Each subsystem cluster corresponds to a message routing module, and each subsystem is connected to the message routing module corresponding to the subsystem cluster to which it belongs.
[0018] Another aspect of the present application provides a chip circuit, including an inter-core communication system according to an embodiment of the present application.
[0019] Another aspect of the present application provides a chip circuit, comprising: one or more processor cores, transmitting transfer messages between the one or more processor cores; and one or more cache processing units, corresponding one-to-one to the one or more processor cores, each cache processing unit comprising: a parameter register, used to store configuration parameters for configuring the cache processing unit; a sending buffer, used to cache transfer messages to be sent by the processor core corresponding to the cache processing unit; and a receiving buffer, used to cache transfer messages to be received by the processor core corresponding to the cache processing unit; wherein the address space in each cache processing unit is divided into two parts, one part is used to access the parameter register, and the other part is used to access the sending buffer and the receiving buffer.
[0020] According to the inter-core communication system and chip circuit of the present application, each subsystem includes a message cache module as a transmission node for transmitting messages, and the transmission nodes are distributedly integrated in each subsystem. The transmission link for transmitting messages from the processor core to the message cache module and from the message cache module to the processor core is short, and the processor core in the subsystem can access the transmission node quickly. The processor core is connected to the message cache module included in the subsystem in which it is located via a bus. The number of processor cores connected on the same bus is small, which reduces the pressure of bus data transmission and avoids bus congestion to a certain extent. Each message cache module is connected to the corresponding message routing module, and the message transmission communication between each message cache module is realized via the message routing module, thereby realizing inter-core communication between multiple processor cores. According to the inter-core communication system of the present application, the processor core has a small delay when performing inter-core communication, thereby reducing the load of the processor in each subsystem, and the inter-core communication system adopts a distributed architecture with good scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram showing an inter-core communication system according to an embodiment of the present application;
[0022] Figure 2 A schematic diagram showing the configuration of a subsystem in an inter-core communication system according to an embodiment of the present application;
[0023] Figure 3 A schematic diagram showing the working logic of the module for routing unicast messages;
[0024] Figure 4 A schematic diagram showing the working logic for routing broadcast delivery messages;
[0025] Figure 5 A schematic diagram showing the configuration of a message cache module according to an embodiment of the present application;
[0026] Figure 6 A schematic diagram showing the format of a transmission message according to an embodiment of the present application.
[0027] Figure 7 A schematic diagram showing the configuration of a cache processing unit according to an embodiment of the present application;
[0028] Figure 8 A schematic diagram showing an example of a corresponding relationship between a message filter and a cache address segment in a receiving cache according to an embodiment of the present application;
[0029] Fig. 9 A schematic diagram showing the communication process of transmitting messages in the inter-core communication system according to the present application. DETAILED DESCRIPTION
[0030] In existing SOCs, shared memory technology is usually used to achieve communication between multiple processor cores distributed in multiple subsystems. Shared memory is a memory area that can be directly accessed by multiple processor cores (such as CPUs). In a multi-core system, processes or threads running on different processor cores often need to exchange data frequently. Multiple processor cores share the same system bus to access shared memory and transfer data through shared memory.
[0031] In shared memory technology, bus mutual exclusion operations and atomic operations are relied on to lock and release shared memory. Data entering and exiting shared memory also needs to go through complex processes such as memory controller scheduling and row and column address decoding, which increases read and write delays, cumbersome operation processes, and large delays, resulting in low efficiency. As the chip scale becomes larger and larger, the SOC architecture becomes increasingly complex, and the paths for each processor core to access shared memory are longer, resulting in significant access delays. Since the data request is quite time-consuming in the bus transmission process, the processor has to be in a waiting state for a long time, which greatly affects the overall performance of the system and has poor scalability.
[0032] The present application provides an inter-core communication system based on a distributed architecture, which effectively reduces the delay during inter-core communication and has good scalability.
[0033] Figure 1A schematic diagram of an inter-core communication system according to an embodiment of the present application is shown. The inter-core communication system is used for communication between multiple processor cores, such as message passing between multiple processor cores. The communication between multiple processor cores can be communication between multiple processor cores in the same processor, or it can be communication between processor cores in different processors. The communication between multiple processor cores can be communication between multiple processor cores in the same subsystem, or it can be communication between multiple processor cores in different subsystems. The inter-core communication system according to the present application can be applied to a SOC. In the present application, the transmission of messages between multiple processor cores is transmission via a hardware link.
[0034] According to an embodiment of the present application, the inter-core communication system includes one or more subsystem cores and one or more message routing modules.
[0035] Figure 1 As an example, it is shown that the inter-core communication system includes multiple subsystem AIs and multiple message routing modules. It is easy to understand that the number of subsystems in the inter-core communication system of the present application can be only one, and the number of message routing modules can also be only one. Subsystem AI is the various subsystems in the SOC, which are divided according to the implemented functions, that is, the circuits and components in the chip are classified into one or more subsystems according to their functions and roles, and each subsystem has independent functional logic. For example, the various subsystems can be the security subsystem, protection subsystem, media subsystem, peripheral control subsystem, etc. in the SOC.
[0036] The message routing module may be a module located outside each subsystem AI in the SOC system, and may be understood as a logic circuit for implementing functions such as forwarding and routing of transmitted messages. The specific configuration of the message routing module will be described later.
[0037] Each subsystem includes a message buffer module and one or more processor cores each connected to the message buffer module via a bus.
[0038] The multiple subsystems AI respectively include a message cache module AI, and each subsystem is integrated with a message cache module as a transmission node (Node) for transmitting messages. The message cache module can be understood as a logic circuit for implementing functions such as decoding, caching, and forwarding of transmitted messages. Exemplarily, each subsystem includes a message cache module. The specific configuration of the message cache module will be described later.
[0039] Figure 2 A schematic diagram showing the configuration of subsystems in an inter-core communication system according to an embodiment of the present application. Figure 2 Only the configuration of one subsystem is shown. It is easy to understand that each subsystem in the inter-core communication system of the present application has a similar configuration.
[0040] like Figure 2 As shown, the subsystem includes a processor circuit block and a peripheral circuit block. The peripheral circuit block includes a circuit component for realizing the function of the subsystem, and the peripheral circuit blocks of each subsystem may be different. The processor circuit block integrates a series of key components and circuits closely related to the processor. One or more processors are integrated in the processor circuit block. Each processor may be a single-core processor or a multi-core processor, that is, each processor includes one or more processor cores. In the present application, the message cache module is integrated in the processor circuit block. In the processor circuit block, all processor cores in the processor circuit block are connected to the message cache module via a bus, and the message can be transmitted between the message cache module and the processor core via the bus. The transmission of the message in the SOC is transmitted via a hardware link. The message cache module is integrated into the processor circuit block of each subsystem, and the transmission link of the message from the processor core to the message cache module and from the message cache module to the processor core is shorter, and the processor core in the subsystem can access the message cache module faster. In the present application, the bus is, for example, an AHB (Advanced High Performance Bus) bus.
[0041] Each message cache module corresponds to all processor cores in the subsystem where the message cache module is located, and all processor cores in each subsystem are connected to the corresponding message cache module (i.e., the message cache module integrated in the subsystem where the processor core is located) via a bus. In each subsystem, there is an independent bus, and the bus in each subsystem only needs to connect the processor core in the subsystem, without having to connect all processor cores in all subsystems in the inter-core communication system with the same bus, thereby reducing the pressure on the bus to transmit data in each subsystem, avoiding bus congestion to a certain extent, and reducing the waiting time for the processor to send data.
[0042] Each processor core is either a producer core or a consumer core.
[0043] The producer core is responsible for generating the processor core of data or resources, and the consumer core is the processor core responsible for using or processing these data or resources. In the inter-core communication system according to the present application, the producer core generates and sends the transfer message, and the consumer core receives and uses the transfer message. It is easy to understand that the producer core and the consumer core are both processor cores in the inter-core communication system, which is only a role division in the inter-core communication process. A processor core in the inter-core communication system is used as a producer core when generating and sending transfer messages, and as a consumer core when receiving and using transfer messages. The recipient of the transfer message can be one or more consumer cores, that is, a transfer message generated by the producer core can be transmitted to one or more consumer cores.
[0044] The producer core and the consumer core may be in the same subsystem or in different subsystems. The producer core and the consumer core may be in the same processor or in different processors. In some examples, each subsystem includes at least one of the producer core and the consumer core.
[0045] The producer core is configured to write the delivery message to the message cache module integrated in the subsystem where the producer core is located, and the consumer core is configured to read the delivery message from the message cache module integrated in the subsystem where the consumer core is located.
[0046] For example, when the producer core is located in subsystem A, the producer core is configured to write the delivery message to the message cache module A connected to it via the bus; when the consumer core is located in subsystem D, the consumer core is configured to read the delivery message from the message cache module D connected to it via the bus, that is, the message cache module D transmits the delivery message to the consumer core connected to it via the bus.
[0047] One or more message routing modules are each connected to a message cache module in at least one of its corresponding subsystems. Each message routing module corresponds to at least one subsystem in the inter-core communication system, and therefore each message routing module corresponds to at least one message cache module in the inter-core communication system, and each message routing module is connected to at least one message cache module corresponding to it. It is easy to understand that each subsystem and each message cache module has a corresponding message routing module, and only corresponds to one message routing module. When the inter-core communication system includes multiple message routing modules, each message routing module is interconnected, and the transmitted message can be transmitted between multiple message routing modules.
[0048] For example, Figure 1 As shown, message routing modules X, Y, and Z are interconnected, message routing module X is connected to message cache modules A, B, and C, message routing module Y is connected to message cache modules D, E, F, and G, and message routing module Z is connected to message cache modules H and I. The above "connection" refers to connection through hardware links. Messages can be transmitted between a message cache module and its corresponding message routing module, and between multiple message routing modules through hardware links.
[0049] It should be understood that the inter-core communication system may include only one message routing module, in which case the message cache modules in all subsystems are connected to the message routing module.
[0050] The message routing module is configured to obtain a delivery message written by a production core to a message cache module integrated in a subsystem where a producer core is located, and route the obtained delivery message to the message cache module integrated in a subsystem where a consumer core is located.
[0051] For example, when the producer core is located in subsystem A and the consumer core is located in subsystem D, the producer core writes a transfer message to the message cache module A connected thereto via the bus, and the message routing module X obtains the transfer message written by the producer core in the message cache module A via the hardware link, that is, the message cache module A transmits the transfer message written by the producer core to the message routing module X, the message routing module X transmits the transfer message to the message routing module Y, the message routing module Y routes the transfer message to the message cache module D, and the consumer core reads the transfer message from the message cache module D connected thereto via the bus, that is, the message cache module D transmits the transfer message to the consumer core connected thereto via the bus. Thus, the transmission of the transfer message from the producer core to the consumer core is realized.
[0052] The following reference Figure 3 and Figure 4 Describes the working logic of the message routing module for routing delivered messages.
[0053] Figure 3 A schematic diagram showing the working logic of the module for routing unicast delivery messages. Unicast delivery messages mean that there is only one destination for receiving a delivery message. The message routing module obtains (receives) a delivery message and sends a request to a port (i.e., an output port) connected to the corresponding destination (e.g., a message cache module integrated in the subsystem where the consumer core is located) based on the delivery message. Each delivery message corresponds to a request. Each port is provided with a port arbitrator. When there are multiple requests to send delivery messages to the same port at the same time, the port arbitrator arbitrates the multiple requests to determine the order in which the multiple delivery messages are output via the port. Thus, the delivery message is output (transmitted) to the corresponding destination.
[0054] Figure 4A schematic diagram of the working logic for routing broadcast delivery messages is shown. Broadcast delivery messages mean that there are multiple destinations for receiving a delivery message, and the delivery message is to be delivered to all processor cores in the inter-core communication system except the processor core that sends the delivery message, that is, in the case of broadcasting, all processor cores in the inter-core communication system except the producer core are consumer cores, and there may be multiple consumer cores in one transmission of the delivery message. The message routing module obtains (receives) the broadcast delivery message and issues a broadcast request based on the delivery message. When the message routing module obtains (receives) the broadcast delivery message, it issues a broadcast request and temporarily suspends all other requests, temporarily suspends the arbitration function of the port arbitrator of the output port until all output ports are idle, and the message routing module transmits the broadcast delivery message to all output ports at the same time until the broadcast delivery message is routed to each destination via each output port, and the message routing module restores the arbitration function of the port arbitrator of the output port. When multiple broadcast delivery messages are transmitted to the message delivery module at the same time, the broadcast arbitrator arbitrates the multiple broadcast requests to determine the order in which each broadcast request is processed. In the present application, the arbitration algorithm of each arbitrator may adopt the Round Robin arbitration algorithm.
[0055] According to the above working logic, the delivery message is routed by the message routing module and is routed to the destination of the delivery message via the corresponding port, for example, to the message cache module integrated in the subsystem where the consumer core is located.
[0056] According to the inter-core communication system of the present application, each subsystem includes a message cache module as a transmission node for transmitting messages, and the transmission nodes are distributedly integrated in each subsystem. The transmission link for transmitting messages from the processor core to the message cache module and from the message cache module to the processor core is short, and the processor core in the subsystem can access the transmission node quickly. The processor core is connected to the message cache module included in the subsystem in which it is located via a bus. The number of processor cores connected on the same bus is small, which reduces the pressure of bus data transmission and avoids bus congestion to a certain extent. Each message cache module is connected to the corresponding message routing module, and the message transmission communication between each message cache module is realized via the message routing module, thereby realizing inter-core communication between multiple processor cores. According to the inter-core communication system of the present application, the processor core has a small delay when performing inter-core communication, thereby reducing the load of the processor in each subsystem, and the inter-core communication system adopts a distributed architecture with good scalability.
[0057] In some embodiments, each message cache module includes: one or more cache processing units, which correspond to one or more processing cores in the subsystem where the message cache module is located, each cache processing unit includes a sending buffer and a receiving buffer, and an internal routing unit, configured to: obtain the delivery message, and route the delivery message to the receiving buffer of the cache processing unit included in the message cache module where the internal routing unit is located; wherein the consumer core is also configured to: read the delivery message from the receiving buffer in its corresponding cache processing unit.
[0058] Figure 5 A schematic diagram showing the configuration of a message cache module according to an embodiment of the present application is shown. The message cache module includes one or more cache processing units, and in each subsystem, one or more cache processing units correspond one-to-one to one or more processor cores in the subsystem where the message cache module is located, that is, one-to-one to the processor core corresponding to the message cache module. For example, when subsystem A includes 5 processor cores, the message cache module A includes at least 5 cache processing units, and the 5 cache processing units correspond one-to-one to the 5 processor cores in subsystem A. It is easy to understand that if the message cache module A includes more than 5 cache processing units, some of the cache processing units may be in an unenabled state. Each cache processing unit includes a sending buffer and a receiving buffer, for example, including a sending buffer core and a receiving buffer, which are respectively used to cache the transfer messages to be sent and received by the processor core corresponding to the cache processing unit. The cache processing unit is a sub-logic circuit in the message cache module, which can be understood as an endpoint in a transmission node. The specific configuration of the cache processing unit will be described in detail later. Exemplarily, the sending buffer and the receiving buffer may be FIFO (First Input First Output) buffers.
[0059] The internal routing unit is used to route the transfer message within the subsystem. Exemplarily, a message cache module includes an internal routing unit. The internal routing unit can obtain the transfer message in the sending buffer in the cache processing unit, that is, obtain the transfer message transmitted from the sending buffer to the internal routing unit, and can also obtain the transfer message routed to the internal routing unit by the message routing module. The internal routing unit corresponds to all cache processing units in the cache processing module where it is located. The internal routing unit is connected to all cache processing units in the cache processing module where it is located. The above "connection" refers to connection through a hardware link. The transfer message can be communicated between the cache processing unit and its corresponding internal routing unit through a hardware link. As described above, the message cache module is connected to its corresponding message routing module, specifically, the message cache module is connected to its corresponding message routing module via its internal internal routing unit, that is, the transfer message can be communicated between the internal routing unit and its corresponding message routing module through a hardware link. Each message routing module corresponds to the internal routing unit included in one or more subsystems corresponding to it.
[0060] The internal routing unit is a sub-logic circuit in the message cache module. The working logic of the internal routing unit is the same as that of the message routing module. For details, please refer to the description of the working logic of the message routing module and Figure 3 and Figure 4 .
[0061] In some examples, when the message cache module includes only one cache processing unit, the cache processing unit directly interacts with the message routing module connected to the message cache module, that is, transmits the message without passing through the internal routing unit. Further, in this case, the message cache module may not include the internal routing unit.
[0062] The internal routing unit can route the delivery message to the receiving buffer of the cache processing unit included in the message cache module where the internal routing unit is located, that is, the internal routing unit can route the delivery message to the receiving buffer in the corresponding one or more cache processing units, so that the delivery message delivered from the producer core can be cached in the receiving buffer. When the memory space of the receiving buffer has reached the upper limit of occupancy, for example, when it is fully occupied, the delivery message to be routed to the receiving buffer will be discarded, thereby preventing data congestion and logical abnormalities in the inter-core communication system. That is, the delivery message that was originally to be transmitted to the receiving buffer of the cache processing unit corresponding to the consumer core is discarded due to insufficient memory space in the receiving buffer, and therefore cannot be successfully transmitted to the consumer core. Accordingly, in order to prevent the delivery message from being discarded due to insufficient memory space in the receiving buffer, the software can monitor whether the delivery message is received by the consumer core, and set a corresponding timeout retransmission mechanism.
[0063] The consumer core reads the transfer message from the receiving buffer in the corresponding cache processing unit. The transfer message is generated and sent by the producer core, and is routed to the receiving buffer in the cache processing unit corresponding to the consumer core via the internal routing unit, so that the consumer core can obtain the transfer message. Thus, the transfer message is transmitted from the producer core to the consumer core. In addition, by setting the internal routing unit, the external links and interfaces of each subsystem are reduced.
[0064] In some embodiments, the message cache module is also configured to: write the delivery message to the sending buffer in the cache processing unit corresponding to the producer core, and the sending buffer is configured to pass the delivery message to the internal routing unit; the internal routing unit is also configured to: obtain the delivery message from the sending buffer in the cache processing unit, and when the producer core and the consumer core are located in the same subsystem, route the obtained delivery message to the receiving buffer in the cache processing unit corresponding to the consumer core; when the producer core and the consumer core are located in different subsystems, route the obtained delivery message to the message routing module; the message routing module is also configured to route the obtained delivery message to the internal routing unit in the message cache module integrated in the subsystem where the consumer core is located.
[0065] The message cache module receives the transfer message written by the corresponding producer core via the bus. The bus interface and address decoding logic in the message cache module process the written transfer message to identify the producer core from which it comes, thereby writing the transfer message into the sending buffer in the cache processing unit corresponding to the producer core. The sending buffer transfers the written transfer message to the internal routing unit corresponding to the cache processing unit where it is located, that is, the sending buffer transfers the written transfer message to the internal routing unit included in the message cache module where it is located.
[0066] Based on the delivery message transmitted by the sending buffer, the internal routing unit identifies the producer core from which the delivery message comes, and identifies the destination of the delivery message, that is, identifies which processor core or cores in the inter-core communication system the consumer core is. When the producer core and the consumer core are located in the same subsystem, the delivery message can be transmitted within the subsystem where the two are located. At this time, the internal routing unit directly routes the delivery message transmitted from the sending buffer to the receiving buffer in the cache processing unit corresponding to the consumer core. When the producer core and the consumer core are located in different subsystems, the internal routing unit routes the delivery message transmitted from the sending buffer to the message routing module connected thereto.
[0067] The message routing module routes the transfer message obtained from the internal routing unit in the subsystem where the producer core is located to the internal routing unit in the subsystem where the consumer core is located. For example, when the producer core is located in subsystem A and the consumer core is located in subsystem B, the internal routing unit in the message cache module A routes the transfer message to the message routing module X, and the message routing module X routes the transfer message to the internal routing unit in the message cache module B. When the inter-core communication system includes multiple message routing modules, the transfer message can be transmitted between multiple message routing modules and finally routed to the internal routing unit in the message cache module in the subsystem where the consumer core is located. As described above, the internal routing unit can route the transfer message to the receiving buffer of the corresponding cache processing unit, so that the consumer core can obtain the transfer message. In this embodiment, the internal routing unit can obtain the transfer message from the sending buffer as the process of the message cache module sending the transfer message, and route the transfer message according to the actual situation. Thus, the transmission of the transfer message from the producer core to the consumer core is realized.
[0068] In some embodiments, the internal routing unit is further configured to obtain a delivery message from the message routing module and route the delivery message to a receiving buffer in a cache processing unit corresponding to the consumer core.
[0069] The internal routing unit identifies the destination of the transfer message based on the transfer message transmitted by the message routing module, and routes the transfer message to the receiving buffer in the cache processing unit corresponding to the consumer core, so that the consumer core can obtain the transfer message from the receiving buffer in the corresponding cache processing unit. In this embodiment, the internal routing unit can obtain the transfer message from the message routing module as the process of the message cache module receiving the transfer message, and route the transfer message to the receiving buffer according to the actual situation. Thus, the transmission of the transfer message from the producer core to the consumer core is realized.
[0070] In some embodiments, the cache processing unit is also configured to, in response to the number of delivery messages in the receiving buffer in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal to a processor core corresponding to the cache processing unit, so that the processor core reads the delivery message from the receiving buffer as a consumer core.
[0071] The number of transfer messages in the receiving buffer in the cache processing unit is greater than or equal to the predetermined number, which means that the transfer messages generated by the producer and transmitted to the receiving buffer in the cache processing unit corresponding to the consumer core are cached in the receiving buffer, and the number of transfer messages cached in the receiving buffer is greater than or equal to the predetermined number. At this time, the cache processing unit sends an interrupt signal to the consumer core to notify the consumer core to read the transfer message from the receiving buffer of the cache processing unit, more specifically, to read the transfer message via the bus connected to the message cache module where the consumer core and the cache processing unit are located. As described above, in each subsystem, one or more cache processing units correspond to one or more processor cores in the subsystem where the message cache module is located, that is, there is a one-to-one correspondence between the cache processing unit and the processor core. When the number of transfer messages in the receiving buffer in the cache processing unit is greater than or equal to the predetermined number, it can be determined which processor core in the inter-core communication system is the corresponding processor core, and the processor core, as the consumer core, will read the transfer message in the receiving buffer.
[0072] The predetermined number can be flexibly set according to actual conditions. The predetermined number can be set according to the number of delivery messages in the receiving buffer. For example, the predetermined number can be set to 1, so that once a delivery message is transmitted to the receiving buffer, the corresponding processor core is immediately notified to read the delivery message as a consumer core. The predetermined number can also be set to multiple, such as 10, so that when the delivery messages cached in the receiving buffer are greater than or equal to 10, the corresponding processor core is notified to read the delivery message from the receiving buffer as a consumer core. The number of delivery messages read by the consumer core from the receiving buffer each time is determined according to the actual application and is not limited here. The predetermined number can be set by setting the waterline of the receiving buffer. When the number of delivery messages in the receiving buffer is higher than the waterline, the cache processing unit where the receiving buffer is located sends an interrupt signal to the corresponding consumer core. As a result, the receiver core can obtain the delivery message sent by the producer core in a timely manner.
[0073] In some embodiments, the delivery message includes: a producer core identification bit for identifying the producer core from which the delivery message comes, and a consumer core identification bit for identifying the consumer core to which the delivery message is to be sent.
[0074] Exemplarily, a delivery message may include a message header and a payload. The message header usually carries key meta-information, such as the source, destination, and type of the delivery message. When a delivery message is sent or received, the message header is sent or received first, allowing the receiving end to know in advance how to process subsequent data. The payload is the content body in the delivery message.
[0075] In some examples, the transfer message may include only a message header but not a payload, that is, the length of the payload portion of the transfer message is 0. This means that the body of the transfer message does not contain specific data content to be transmitted, but only a message header portion. Such a message can serve as a simple notification and synchronization between processor cores. For example, when the first processor core executes to a certain stage and needs to notify the second processor core to start the next step, it can send such a transfer message with no payload but only a message header. After receiving the transfer message, the second processor core knows that the corresponding "event" has been triggered, and then takes corresponding actions. In this way, such a transfer message can be used to replace the function of Event in the current inter-process communication (IPC) mechanism.
[0076] The producer core identification bit and the consumer core identification bit may be included in the message header. When the producer core generates a transfer message, the producer core identification bit and the consumer core identification bit of the transfer message are determined; when the message cache module receives the transfer message written by the producer core, it identifies the sending buffer in which cache processing unit the transfer message is to be written according to the producer core identification bit included in the transfer message through its internal bus interface and address decoding logic; when the internal routing unit obtains the transfer message from the sending buffer, it determines to which destination the transfer message is to be routed according to the consumer core identification bit included in the transfer message, that is, to which receiving buffer in which cache processing unit the transfer message is to be routed or to the message routing module; when the message routing module obtains the transfer message from the internal routing unit, it determines to which subsystem of the message cache module in the inter-core communication system the transfer message is to be routed according to the consumer core identification bit included in the transfer message; when the internal routing unit obtains the transfer message from the message routing module, it determines to which receiving buffer in which cache processing unit the transfer message is to be routed according to the consumer core identification bit included in the transfer message. In short, during the transmission process of the delivery message, the producer core from which the delivery message comes and the consumer core to which the delivery message is to be transmitted are determined according to the producer core identification bit and the consumer core identification bit included in the delivery message.
[0077] In some embodiments, the internal routing unit is further configured to discard the delivery message when both the producer core indicated by the producer core identification bit in the delivery message and the consumer core indicated by the consumer identification bit in the delivery message are not in the subsystem where the internal routing unit is located.
[0078] When the internal routing unit obtains the transfer message, it can also determine whether the producer core identification bit and the consumer core identification bit of the transfer message indicate the processor core in the subsystem where the internal routing unit is located. When the producer core identification bit and the consumer core identification bit do not indicate the processor core in the subsystem where the internal routing unit is located, that is, when the producer core indicated by the producer core identification bit and the consumer core indicated by the consumer identification bit are not in the subsystem where the internal routing unit is located, the transfer message is discarded. This situation is usually caused by a software error that leads to an erroneous transfer message. The erroneous transfer message is discarded to avoid deadlock in the inter-core communication system.
[0079] In some embodiments, the producer core identification bit is also used to identify the cache processing module included in the subsystem where the producer core is located and the cache processing unit corresponding to the producer core, and the consumer core identification bit is also used to identify the cache processing module included in the subsystem where the consumer core is located and the cache processing unit corresponding to the consumer core.
[0080] The producer core identification bit may include two parts, the first part is used to identify the cache processing module included in the subsystem where the producer core is located, and the second part is used to identify the cache processing unit corresponding to the producer core; accordingly, the consumer core identification bit includes two parts, the first part is used to identify the cache processing module included in the subsystem where the consumer core is located, and the second part is used to identify the cache processing unit corresponding to the consumer core. Since the cache processing module, the cache processing unit and the processor core have a corresponding relationship, the processor core can be uniquely identified by determining the cache processing module and the cache processing unit, thereby determining the producer core identification bit and the consumer core identification bit.
[0081] The following combination Figure 6 Describes in detail the format of the delivered message. Figure 6 A schematic diagram showing the format of a transmission message according to an embodiment of the present application.
[0082] like Figure 6 As shown, the message header of the transfer message includes a producer core identification bit PID, a consumer core identification bit CID, a length field, and a reserved field. The transfer message also includes a payload. The width of the transfer message is consistent with the width of the bus connecting the processor core to its corresponding message cache module. In practical applications, the bus width is usually 32 bits or 64 bits, so the width of the transfer message is usually 32 bits or 64 bits.
[0083] exist Figure 6In the example shown, the producer core identification bit PID and the consumer core identification bit CID are 8 bits each, and each is evenly divided into two parts, the first part is the high-order 4 bits, which are used to identify the message cache module, and the second part is the low-order 4 bits, which are used to identify the cache processing unit in the message cache module. The producer core identification bit PID and the consumer core identification bit CID can both uniquely identify the processor core in the inter-core communication system. In particular, when the consumer core identification bit CID is all 0, it is used as a function extension, and when the consumer core identification bit CID is all 1, it indicates that the transfer message is a broadcast, that is, the transfer message is to be delivered to all processor cores in the inter-core communication system except the processor core that sends the transfer message, that is, all processor cores except the processor core that sends the transfer message are used as consumer cores.
[0084] The length field is used to indicate the length of the payload in the transfer message. For a 32-bit width transfer message, the length of the payload can range from 0 to 9, and for a 64-bit width transfer message, the length of the payload can range from 0 to 4; in this case, the maximum size of a transfer message is 320 bits.
[0085] The reserved field can be defined by software according to actual application. For a 32-bit width transfer message, the length of the reserved field is 12 bits. For a 64-bit width transfer message, the length of the reserved field is 44 bits. By setting the reserved field, the function and application of the transfer message can be expanded. For example, in the case where the above-mentioned transfer message is used to replace the function of Event in the IPC mechanism, through the reserved field, the transfer message that only includes the message header can pass some additional information.
[0086] In some embodiments, the cache processing unit is further configured to convert the format of the transmitted message in the receiving buffer in the cache processing unit according to the width of the bus connecting the message cache module where the cache processing unit is located and one or more processor cores.
[0087] In different subsystems, the width of the bus connecting the message cache module to one or more processor cores may be different. It is necessary to ensure that the width of the transfer message is consistent with the width of the bus so that the transfer message can be effectively transmitted on the bus. For example, in subsystem A, the width of the bus connecting the processor core to the message cache module is 64 bits, while in subsystem D, the width of the bus connecting the processor core to the message cache module is 32 bits. If the transfer message is sent from the processor core in subsystem A to the processor core in subsystem D, the transfer message is 64 bits when the producer core generates and writes it to the message cache module in subsystem A. In this case, the format of the transfer message needs to be converted to adapt to the bus width in subsystem D, so that the consumer core in subsystem D can obtain the transfer message via a 32-bit bus. The cache processing unit converts the width of the transfer message in the receiving buffer into the width of the bus connected to the message cache module where the cache processing unit is located and the corresponding one or more processor cores.
[0088] As mentioned above, in practical applications, the bus width is usually 32 bits or 64 bits, and the width of the message transmission is correspondingly 32 bits or 64 bits. The following is an example to illustrate the conversion between 32 bits and 64 bits of the message transmission.
[0089] A transfer message with a width of 64 bits is converted into a transfer message with a width of 32 bits: the low-order 32 bits of the message header in the transfer message with a width of 64 bits are used as the message header in the transfer message with a width of 32 bits. It is easy to understand that the message header includes a 12-bit reserved field; the high-order 32 bits of the message header in the transfer message with a width of 64 bits, that is, the high-order 32 bits of the reserved field in the message header in the transfer message with a width of 64 bits, are used as a 32-bit payload; the payload in each 64-bit transfer message is converted into two 32-bit payloads according to its high-order 32 bits and low-order 32 bits respectively.
[0090] Convert a 32-bit transfer message to a 64-bit transfer message: the message header of the 32-bit transfer message is used as the low-order 32 bits of the 64-bit message header; a payload of the 32-bit transfer message is used as the high-order 32 bits of the 64-bit message header; if the 32-bit transfer message does not contain a payload, the high-order 32 bits of the converted 64-bit message header are all 0; the payloads of every two 32-bit transfer messages are converted to one 64-bit payload; if there is only one 32-bit payload left in the transfer message, the high-order 32 bits of the last 64-bit payload are all 0.
[0091] According to this embodiment, even when the transfer message is produced and consumed in systems with different bus widths, effective transmission of the transfer message can be guaranteed.
[0092] In some embodiments, the cache processing unit also includes: one or more message filters, each message filter corresponds to a different cache address segment in the receiving cache; the one or more message filters are each configured to: receive a delivery message from the internal routing unit, filter the received delivery message based on their own filtering rules, and write the filtered message into the corresponding cache address segment.
[0093] Figure 7 A schematic diagram showing the configuration of a cache processing unit according to an embodiment of the present application. Figure 8 A schematic diagram showing an example of the correspondence between a message filter and a cache address segment in a receiving buffer according to an embodiment of the present application.
[0094] like Figure 7 As shown, the cache processing unit includes one or more message filters. Figure 8 As shown, each message filter corresponds to a different cache address segment in the receiving cache in the cache processing unit where the message filter is located. Each message filter is configured with a predetermined filtering rule. The message filter receives the transfer message from the internal routing unit and filters the received transfer message. Filtering the transfer message means selecting (retaining) the transfer message that meets the predetermined filtering rule in the received message. The message filter then writes the filtered transfer message (that is, the transfer message that meets the predetermined rule of the message filter) into the cache address segment corresponding to the message filter, that is, the filtered transfer message is cached in the corresponding cache address segment. The message filter can be implemented by a logic circuit.
[0095] The filtering rule may be to filter one or more fields of the message header of the transmitted message, such as filtering the producer core identification bit field, length field or reserved field within a certain range. The filtering rule may also be to filter the payload of the transmitted message, such as filtering the payload that meets the predetermined conditions. It is easy to understand that the filtering rule of a message filter can apply one or more filtering conditions at the same time.
[0096] In some examples, such as Figure 8 As shown, the receiving buffer is also provided with a default address segment, and the delivery message that is not selected by any one of the one or more message filters in the cache processing unit will be written into the default address segment. For example, the cache processing unit also includes a default filter, and the default filter corresponds to the default address segment. The delivery message that is not selected by any one of the one or more message filters in the cache processing unit can enter the default filter and be written into the default address segment.
[0097] In some examples, by setting filtering rules of one or more message filters, the cache processing unit can also decide whether to receive a specific delivery message. Exemplarily, a delivery message that is not selected by any of the one or more message filters in the cache processing unit can be discarded and not written to the receiving buffer. In this way, the cache processing unit can discard the delivery message sent from a specific producer core, so that the consumer core corresponding to the cache processing unit does not receive the delivery message sent from the specific producer core.
[0098] In the above example, the cache processing unit can be set to set whether the delivery message that is not selected by any of the one or more message filters in the cache processing unit enters the default filter or is discarded. The setting of the cache processing unit is achieved by applying the configuration parameters of the cache processing unit to the cache processing unit to modify the circuit structure and logic of its actual operation.
[0099] In some examples, when the memory space in any cache address segment in the receiving buffer has reached the upper limit of occupancy, for example, when it is fully occupied, no more transfer messages will be written to the cache address segment, and the transfer messages to be written to the cache address segment will be discarded, thereby preventing data congestion and logical anomalies in the inter-core communication system. That is, the transfer message that was originally to be cached in the cache address segment is discarded due to insufficient memory space in the receiving buffer, and therefore cannot be successfully transmitted to the consumer core. Accordingly, in order to prevent the transfer message from being discarded due to insufficient memory space in the receiving buffer, the software can monitor whether the transfer message is received by the consumer core, and set a corresponding timeout retransmission mechanism.
[0100] According to this embodiment, a specific delivery message can be distinguished from other delivery messages, so as to facilitate management and use by consumers.
[0101] In some embodiments, the cache processing unit is further configured to: receive configuration parameters, and set actually enabled message filters, filtering rules of each message filter, and ranges of each cache address segment based on the configuration parameters.
[0102] See again Figure 7 The cache processing unit may also include one or more parameter registers. The parameter register is, for example, a control status register, which receives and stores configuration parameters for configuring the cache processing unit in which it is located. The cache processing unit modifies its actual working circuit structure and logic according to the configuration parameters stored in the parameter register, thereby realizing the setting of the cache processing unit. Figure 7 Only parameter registers included in the cache processing unit are shown, however, in some examples, some of the one or more parameter registers may be located within the message filter to configure the message filter in which they are located. More specifically, each message filter includes a parameter register.
[0103] The address space of each cache processing unit is divided into two parts. One part is used to access one or more parameter registers. When the processor core writes configuration parameters to the cache processing unit, the configuration parameters are written to the parameter registers by addressing this part of the address space; the other part is used to access the send buffer and the receive buffer. When the processor core wants to access the send buffer or the receive buffer, the send buffer or the receive buffer is accessed by addressing this part of the address space. In this example, in order to flexibly support page tables of different sizes and address attributes, the two parts of the address space each occupy 64K, and one cache processing unit occupies 128K of address space.
[0104] The filtering rules of each message filter, which message filter or filters in the cache processing unit are actually enabled, and the range of each cache address segment (e.g., the start and end addresses of each cache address segment) can be set according to actual needs. For example, if the capacity of a certain type of transmission message is expected to be large, the cache address segment to be written can be set to a larger address segment range. In other words, the circuit structure and logic of the actual operation of the cache processing unit can be configured according to actual needs. Each enabled message filter has its corresponding cache address segment, and the address ranges of the cache address segments corresponding to each enabled message filter do not overlap.
[0105] It is easy to understand that n-1 enabled message filters correspond to n-1 cache address segments in the receiving buffer, and there is another default address segment in the receiving buffer for writing the delivery messages that are not selected by the n-1 enabled message filters. Therefore, in this case, the receiving buffer includes n cache address segments, and each cache address segment is continuous and non-overlapping. Correspondingly, the delivery messages are divided into n types. One cache address segment corresponds to one type of delivery message.
[0106] like Figure 8 As shown, 7 message filters are integrated in a cache processing unit, of which 2 message filters (message filter 1 and message filter 2) are actually enabled. The filtering rules of message filter 1 and message filter 2 are to select the delivery messages with producer identification bit PID 2 and producer identification bit PID 3-8 respectively. The delivery messages selected by these two message filters will be written into cache address segment 1 and cache address segment 2 in the receiving buffer respectively. The delivery messages that are not selected by any enabled message filter will be uniformly written into cache address segment 3 as the default address segment.
[0107] According to this embodiment, the cache processing unit can be configured according to actual needs, so that the filtering of the transmitted messages can be flexibly applied to different application scenarios.
[0108] In some embodiments, each message filter includes one or more filtering units, and each filtering unit filters different parts of the transmitted message based on its own filtering algorithm to obtain the unit filtering results of each filtering unit; the cache processing unit is also configured to: set the filtering units actually enabled in each message filter based on the configuration parameters, and apply the rules of the unit filtering results of each filtering unit in the message filter, thereby setting the filtering rules of each message filter.
[0109] Exemplarily, each message filter includes one or more message header filtering units for filtering message headers of transmitted messages and one or more payload filtering units for filtering payloads of transmitted messages, each filtering unit is configured with a corresponding filtering algorithm, and the filtering algorithm acts on the message header and the corresponding payload of the transmitted message respectively. Each filtering unit can be individually set to be enabled (enabled) or disabled (closed). The cache processing unit sets which filtering unit or filtering units in a message filter are actually enabled according to the received configuration parameters.
[0110] The rule for applying the unit filtering results of each filtering unit in the message filter refers to the rule for how to apply the unit filtering results to obtain the final filtering result of the message filter. For example, the results of each unit filtering unit can be configured to be AND or OR to generate the final filtering result. The cache processing unit sets the rule for applying the unit filtering results of each filtering unit in the message filter according to the received configuration parameters.
[0111] According to this embodiment, by setting the filtering unit, the circuit structure and logic of the actual operation of the message filter can be easily set, so as to flexibly set the filtering rules of each message filter.
[0112] In some embodiments, the cache processing unit is further configured to: in response to the number of transfer messages in any cache address segment of a receiving cache in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal corresponding to the cache address segment to the processor core corresponding to the cache processing unit, so that the processor core acts as a consumer core to read the transfer message from the cache address segment corresponding to the interrupt signal in the receiving cache; interrupt signals corresponding to different cache address segments have different interrupt priorities.
[0113] In this embodiment, each cache address segment has an independent interrupt mechanism, that is, each type of transfer message has an independent interrupt mechanism. If the number of transfer messages in a certain cache address segment of the receiving cache in the cache processing unit is greater than or equal to the predetermined number, it means that the specific type of transfer message is transmitted to the receiving cache, and the number of the specific type of transfer messages in the receiving cache is greater than or equal to the predetermined number. At this time, the cache processing unit sends an interrupt signal to the consumer core to notify the consumer core to read the transfer message from the cache address segment.
[0114] Interrupt signals corresponding to different cache address segments have different interrupt priorities, which means that the consumer core processes interrupt signals corresponding to different cache address segments with different priorities. The priority of the interrupt signal can be configured by the interrupt controller, that is, the interrupt controller sets different priorities for interrupt signals corresponding to different cache address segments. It is easy to understand that the interrupt controller can also set the interrupt signals corresponding to multiple cache address segments to have the same priority.
[0115] For example, when the number of transfer messages cached in cache address segment 1 in the receiving cache is greater than or equal to a predetermined number, regardless of the number of transfer messages in cache address segment 2 or 3, an interrupt signal corresponding to the cache address segment 1 is sent to the processor core corresponding to the cache processing unit. In response to receiving the interrupt signal, the processor core reads the transfer message from the cache address segment 1 of the receiving cache as a consumer core. In addition, the interrupt controller can set the priority of the interrupt signal corresponding to cache address segment 1 to be greater than the priority of the interrupt signal corresponding to cache address segment 2, and set the priority of the interrupt signal corresponding to cache address segment 2 to be greater than the priority of the interrupt signal corresponding to cache address segment 3, so that when the processor core receives multiple interrupt signals at the same time, the interrupt signal with a higher priority is processed first. By setting different priorities for interrupt signals corresponding to different cache address segments, specific transfer messages can be distinguished from other transfer messages, and the priorities of these specific transfer messages can be increased, so that the consumer core can process these specific transfer messages first.
[0116] According to this embodiment, the processor core as the consumer core can easily distinguish different types of delivery messages and set different priorities for the different types of delivery messages, thereby processing the delivery messages as needed.
[0117] In some embodiments, the cache processing unit is further configured to: calculate a send verification code based on the delivery message written into the send buffer, and attach the send verification code to the delivery message; calculate a receive verification code based on the delivery message written into the receive buffer, and compare the receive verification code with the send verification code attached to the delivery message to determine whether a transmission error occurs in the delivery message.
[0118] As mentioned above, the message cache module in the subsystem where the producer core is located writes the transfer message into the sending buffer in the cache processing unit corresponding to the producer core. At this time, the cache processing unit receives the writing of the transfer message from the message header of the transfer message until the reception of the transfer message is completed. In the process of writing the transfer message into the sending buffer, the cache processing unit calculates the sending check code and attaches the sending check code to the transfer message so that the sending check code is transmitted together with the transfer message. The transfer message and the attached sending check code are finally transmitted to the cache processing unit corresponding to the consumer core, and are written to the receiving buffer in the cache processing unit corresponding to the consumer core.
[0119] The cache processing unit corresponding to the consumer core receives the delivery message with the send check code attached. More specifically, the delivery message with the send check code attached is written into the receive buffer in the cache processing unit corresponding to the consumer core. At this time, the cache processing unit calculates the receive check code based on the received delivery message, and compares the calculated receive check code with the send check code attached to the received delivery message. If the send check code is consistent with the receive check code, it can be considered that the delivery message is correct; if the send check code is inconsistent with the receive check code, it is considered that a transmission error occurs in the delivery message. If a transmission error occurs, the producer core can be notified to retransmit.
[0120] The check code can be ECC (Error-Correcting Code), which is a coding technology used for data transmission. When the message is transmitted in the hardware link of the SOC, it is easy to be affected by noise interference, attenuation, etc., resulting in transmission errors. By calculating the ECC code, the bit error rate can be reduced and the communication reliability can be improved.
[0121] According to this embodiment, the reliability of message delivery during the transmission process of inter-core communication can be guaranteed.
[0122] In some embodiments, multiple subsystems are divided into one or more subsystem clusters according to the functions they implement and their locations. Each subsystem cluster corresponds to a message routing module, and each subsystem is connected to the message routing module corresponding to the subsystem cluster to which it belongs.
[0123] See again Figure 1 ,exist Figure 1 In the example, subsystems A, B, and C form a subsystem cluster, corresponding to the message routing module X, and subsystems A, B, and C are all connected to the message routing module X; subsystems D, E, F, and G form a subsystem cluster, corresponding to the message routing module Y, and subsystems D, E, F, and G are all connected to the message routing module Y; subsystems H and I form a subsystem cluster, corresponding to the message routing module Z, and subsystems H and I are all connected to the message routing module Z.
[0124] It is easy to understand that when the message is transmitted between different subsystem clusters, that is, when the producer core and the consumer core are not located in the same subsystem cluster, the message is transmitted from the message routing module corresponding to the subsystem where the producer core is located, via the connection between multiple message routing modules, to the message routing module corresponding to the subsystem where the consumer core is located.
[0125] In the design of SOC, subsystems with closely related functions are designed to be close to each other in the SOC to reduce the transmission path of signals or data. For example, the GPU subsystem responsible for graphics processing, the ISP subsystem responsible for image signal processing, and the codec subsystem responsible for video encoding and decoding are all closely connected to vision-related tasks, and these subsystems can be located close to each other. When shooting a video, the ISP subsystem captures and preliminarily processes the image signal, and then it can be quickly transmitted to the codec subsystem for compression into a suitable format, or passed to the GPU subsystem for special effects rendering, reducing the loss and delay of data transmission across long distances.
[0126] Furthermore, the message routing module may be arranged close to each subsystem in the subsystem cluster to which it corresponds, thereby reducing the transmission distance of the message between the subsystem and the message routing module.
[0127] According to this embodiment, multiple subsystems are divided into one or more subsystem clusters according to the functions they implement and their locations. Subsystems with closely related functions and close locations are divided into the same subsystem cluster, which improves the efficiency of collaborative work between subsystems. In addition, different subsystem clusters have different requirements for resources (such as electricity and bandwidth) during operation. The layout of subsystem clusters and message routing modules can be adjusted according to actual needs. By dividing different subsystem clusters, the SOC can accurately control resource allocation according to the real-time needs of each subsystem cluster, and flexibly open or close some subsystem clusters to save power consumption. In addition, each subsystem cluster corresponds to a message routing module, and each subsystem is connected to the message routing module corresponding to the subsystem cluster to which it belongs. In this way, the transmission path of the message between each message cache module and the message routing module is shortened as much as possible, reducing the risk of electromagnetic interference and improving the transmission efficiency of the message.
[0128] In some embodiments, a predetermined watermark may be set for the sending buffer in the cache processing unit. When the amount of transfer messages in the sending buffer is lower than the watermark, the cache processing unit may send an interrupt signal to its corresponding processor core to notify the processor core that it can continue to send transfer messages as a producer core. The amount of transfer messages in the sending buffer being lower than the predetermined watermark means that there is a large amount of free capacity in the sending buffer, and the transfer messages are not likely to overflow. At this time, the producer core may be further reminded to continue sending transfer messages. The predetermined watermark may be a predetermined percentage of the overall capacity of the sending buffer.
[0129] According to the inter-core communication system of the present application, the transmission process of the message is transparent to the software running on the processor core, which means that the software-level application running on the processor core does not need to care about the transmission process of the message, and can directly obtain the transmission result of the message.
[0130] Fig. 9 A schematic diagram showing the communication process of the inter-core communication system according to the present application for transmitting messages. Fig. 9 The communication process of transmitting messages according to the inter-core communication system of the present application is described. Fig. 9 The arrows in the figure represent the transmission routes for delivering messages.
[0131] In step 10, the producer core generates a delivery message.
[0132] In step 20, the producer core writes the transfer message to a message cache module in the subsystem where the producer core is located via a bus, and the message cache module receives the transfer message via its bus interface.
[0133] In step 30, the message cache module processes the received transfer message via the address decoding logic to identify the producer core from which it comes, and writes the transfer message into the sending buffer in the cache processing unit corresponding to the producer core. At the same time, the cache processing unit corresponding to the producer core calculates the sending check code based on the written transfer message, and attaches the sending check code to the transfer message.
[0134] In step 40, the sending buffer in the cache processing unit transmits the transfer message with the sending verification code attached to the internal routing unit in the message cache module where the cache processing unit is located, and the internal routing unit receives the transfer message.
[0135] In step 50, the internal routing unit determines the producer core from which the delivery message comes and the consumer core to which the delivery message is to be delivered based on the delivery message. When the producer core and the consumer core are located in different subsystems, the communication process is transferred to step 60; when the producer core and the consumer core are located in the same subsystem, the communication process is transferred to step 100. It is easy to understand that in this case, the cache processing unit corresponding to the consumer core and the cache processing unit corresponding to the producer core are located in the same message cache module.
[0136] In step 60, the acquired delivery message is routed to a message routing module connected to the internal routing unit, that is, routed to a message routing module connected to the subsystem where the producer core is located. The message routing module receives the delivery message.
[0137] In step 70, the message routing module determines whether the consumer core to which the message is to be transmitted is located in the subsystem cluster corresponding to the message routing module. When the consumer core is not located in the subsystem cluster corresponding to the message routing module, the communication process transfers to step 80, and when the consumer core is located in the subsystem cluster corresponding to the message routing module, the communication process transfers to step 90.
[0138] In step 80, the message routing module routes the delivery message to the message routing module connected to the subsystem where the consumer core is located.
[0139] In step 90 , the message routing module routes the delivery message to an internal routing unit in a message buffer module of a subsystem where the consumer core is located, and the internal routing unit receives the delivery message.
[0140] In step 100, the internal routing unit transmits the delivery message to the cache processing unit corresponding to the consumer core. The cache processing unit receives the delivery message, filters the delivery message through one or more message filters, and writes the filtered delivery message into the corresponding address segment in the receiving buffer. The cache processing unit sends an interrupt signal to the consumer core.
[0141] In step 110 , the consumer core obtains the delivery message from the receiving buffer in the corresponding buffer processing unit.
[0142] exist Fig. 9 In the example, step 40 is shown with a dotted line because if the message cache module in the subsystem where the producer core is located only includes one cache processing unit, the cache processing unit can directly transmit the delivery message to the message routing module without passing through the internal routing unit. Similarly, step 90 is shown with a dotted line because if the message cache module in the subsystem where the receiver core is located only includes one cache processing unit, the message routing module can directly transmit the delivery message to the cache processing unit without passing through the internal routing unit.
[0143] Another aspect of the present application provides a chip circuit, comprising the inter-core communication system in the above-described embodiment. Exemplarily, the chip circuit may be a SOC chip.
[0144] Another aspect of the present application provides a chip circuit, including: one or more processor cores, transmitting transfer messages between the one or more processor cores; and one or more cache processing units, corresponding one to one with the one or more processor cores, each cache processing unit including: a parameter register, used to store configuration parameters for configuring the cache processing unit; a sending buffer, used to cache transfer messages to be sent by the processor core corresponding to the cache processing unit; and a receiving buffer, used to cache transfer messages to be received by the processor core corresponding to the cache processing unit; wherein the address space in each cache processing unit is divided into two parts, one part is used to access the parameter register, and the other part is used to access the sending buffer and the receiving buffer.
[0145] The functions and features of the processor core and cache processing unit described in the foregoing embodiments can be applied to the processor core and cache processing unit according to this aspect, and will not be described in detail here.
[0146] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0147] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. Inter-core communication system, including: One or more subsystems, each subsystem comprising a message cache module and one or more processor cores connected to the message cache module via a bus, each processor core being a producer core or a consumer core, the producer core being configured to write a delivery message to the message cache module integrated in the subsystem where the producer core is located, and the consumer core being configured to read the delivery message from the message cache module integrated in the subsystem where the consumer core is located; as well as One or more message routing modules are configured to: obtain the delivery message written by the production core to the message cache module integrated in the subsystem where the producer core is located, and route the obtained delivery message to the message cache module integrated in the subsystem where the consumer core is located.
2. The inter-core communication system according to claim 1, wherein: Each of the message cache modules comprises: One or more cache processing units, respectively corresponding to one or more processing cores in the subsystem where the message cache module is located, each cache processing unit includes a sending buffer and a receiving buffer, and An internal routing unit is configured to: obtain the transfer message and route the transfer message to a receiving buffer of a buffer processing unit included in a message buffer module where the internal routing unit is located; The consumer core is further configured to read the delivery message from a receiving buffer in its corresponding cache processing unit.
3. The inter-core communication system according to claim 2, wherein: The message cache module is further configured to: write the delivery message into a sending buffer in a cache processing unit corresponding to the producer core, and the sending buffer is configured to deliver the delivery message to the internal routing unit; The internal routing unit is further configured to: obtain the transfer message from the sending buffer in the cache processing unit, and when the producer core and the consumer core are located in the same subsystem, route the obtained transfer message to the receiving buffer in the cache processing unit corresponding to the consumer core; when the producer core and the consumer core are located in different subsystems, route the obtained transfer message to the message routing module; The message routing module is also configured to route the acquired delivery message to an internal routing unit in a message cache module in the subsystem where the consumer core is located.
4. The inter-core communication system according to claim 3, wherein: The internal routing unit is also configured to obtain the delivery message from the message routing module and route the delivery message to the receiving buffer in the cache processing unit corresponding to the consumer core.
5. The inter-core communication system according to claim 2, wherein: The cache processing unit is also configured to, in response to the number of delivery messages in the receiving cache in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal to the processor core corresponding to the cache processing unit, so that the processor core acts as the consumer core to read the delivery message from the receiving cache.
6. The inter-core communication system according to claim 2, wherein: The delivery message includes: a producer core identification bit for identifying the producer core from which the delivery message comes, and a consumer core identification bit for identifying the consumer core to which the delivery message is to be sent.
7. The inter-core communication system according to claim 6, wherein: The internal routing unit is further configured as: When the producer core indicated by the producer core identification bit in the transfer message and the consumer core indicated by the consumer identification bit in the transfer message are both not in the subsystem where the internal routing unit is located, the transfer message is discarded.
8. The inter-core communication system according to claim 6, wherein: The producer core identification bit is also used to identify the cache processing module included in the subsystem where the producer core is located and the cache processing unit corresponding to the producer core. The consumer core identification bit is also used to identify the cache processing module included in the subsystem where the consumer core is located and the cache processing unit corresponding to the consumer core.
9. The inter-core communication system according to claim 2, wherein: The cache processing unit is further configured to convert the format of the transmitted message in the receiving buffer in the cache processing unit according to the width of the bus connecting the message cache module where the cache processing unit is located and the one or more processor cores.
10. The inter-core communication system according to any one of claims 2 to 9, wherein: The cache processing unit further includes: one or more message filters, each message filter corresponding to a different cache address segment in the receiving cache; The one or more message filters are each configured to: receive the transfer message from the internal routing unit, filter the received transfer message based on respective filtering rules, and write the filtered transfer message into a corresponding cache address segment.
11. The inter-core communication system according to claim 10, wherein: The cache processing unit is further configured to: receive configuration parameters, and set the message filters actually enabled, the filtering rules of each message filter, and the range of each cache address segment based on the configuration parameters.
12. The inter-core communication system according to claim 11, wherein: Each of the message filters includes one or more filtering units, each filtering unit filters different parts of the transmitted message based on its own filtering algorithm to obtain a unit filtering result of each filtering unit; The cache processing unit is further configured to: set the filtering unit actually enabled in each message filter based on the configuration parameters, and apply the rules of the unit filtering results of each filtering unit in the message filter, thereby setting the filtering rules of each message filter.
13. The inter-core communication system according to claim 10, wherein: The cache processing unit is further configured to: in response to the number of delivery messages in any cache address segment of the receiving cache in the cache processing unit being greater than or equal to a predetermined number, send an interrupt signal corresponding to the cache address segment to a processor core corresponding to the cache processing unit, so that the processor core, as the consumer core, reads the delivery message from the cache address segment corresponding to the interrupt signal in the receiving cache; Interrupt signals corresponding to different cache address segments have different interrupt priorities.
14. The inter-core communication system according to claim 2 or 3, wherein: The cache processing unit is further configured as: calculating a sending check code based on the delivery message written into the sending buffer, and appending the sending check code to the delivery message; A reception check code is calculated based on the delivery message written into the reception buffer, and the reception check code is compared with a transmission check code attached to the delivery message to determine whether a transmission error occurs in the delivery message.
15. The inter-core communication system according to claim 1, wherein: The multiple subsystems are divided into one or more subsystem clusters according to the functions they implement and their locations. Each subsystem cluster corresponds to a message routing module, and each subsystem is connected to the message routing module corresponding to the subsystem cluster to which it belongs.
16. A chip circuit comprising the inter-core communication system according to any one of claims 1 to 15.
17. Chip circuit, including: One or more processor cores, transmitting and receiving messages between the one or more processor cores; and One or more cache processing units, corresponding to the one or more processor cores respectively, each cache processing unit comprising: A parameter register, used to store configuration parameters for configuring the cache processing unit; A sending buffer, used for caching a transfer message to be sent by a processor core corresponding to the cache processing unit where the sending buffer is located; and A receiving buffer, used for caching a transfer message to be received by a processor core corresponding to the cache processing unit where the receiving buffer is located; The address space in each cache processing unit is divided into two parts, one part is used to access the parameter register, and the other part is used to access the sending buffer and the receiving buffer.