On-chip inter-core communication system and method
By designing an on-chip inter-core communication system in a multi-core system, using shared memory and interrupt control units, efficient inter-core data synchronization is achieved, solving the problem of low inter-core communication performance in the existing technology, and improving the overall performance and efficiency of the system.
Patent Information
- Application Number
- CN202411748811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing inter-core communication solutions have degraded performance under high load conditions, low real-time performance, and low space utilization of storage resources, so they cannot effectively support efficient parallel computing of multi-core systems.
A on-chip inter-core communication system is designed, using shared memory and interrupt control unit to achieve inter-core data synchronization through data frame lookup tables and ring buffers, avoiding dependence on dedicated communication buses or Noc hardware integrated within the chip.
It reduces the delay in data writing/reading, improves the response speed of inter-core communication and data synchronization quality, and does not rely on high-cost hardware support, effectively utilizes shared memory space.
Smart Images

Figure CN119988306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to an inter-core communication system and method of a multi-core chip. Background Art
[0002] In modern computer architecture, especially in the fields of high-performance computing, server applications and mobile devices, multi-core processors have gradually become mainstream. As the performance improvement of a single processor core gradually encounters physical limits, improving computing power by integrating multiple processing cores has become an effective strategy. However, one of the core advantages of a multi-core system is that it can execute multiple tasks or threads in parallel, which requires efficient data exchange and communication mechanisms between different cores to ensure the overall performance and efficiency of the system.
[0003] Most existing inter-core communication solutions are implemented using shared memory, message queues, or dedicated communication buses or networks (Network-on-Chip, Noc) integrated inside the chip. Although the shared memory method is simple to implement, it will lead to performance degradation under high load conditions; the message queue has low real-time performance and is not suitable for applications with high real-time requirements, and the space utilization rate of storage resources is low; although the dedicated communication bus or network (Noc) integrated inside the chip can provide higher bandwidth and lower response delay, due to the high complexity and cost requirements of this method on chip design, common multi-core chips may not have such hardware support. Summary of the invention
[0004] In order to solve the above technical problems existing in the background technology, the present invention provides an on-chip inter-core communication system and method, which utilizes the advantage of the processor cores in a multi-core processor to quickly access the on-chip shared memory to reduce the delay of data writing / reading.
[0005] The technical solution of the present invention is: the present invention is an on-chip inter-core communication system, which is special in that: the on-chip inter-core communication system includes multiple processor cores, shared memory and interrupt control unit, the shared memory is connected to the multiple processor cores, the multiple processor cores are respectively connected to the interrupt control unit, the multiple processor cores are divided into data sending cores and data receiving cores, which are responsible for executing tasks and inter-core communication; the shared memory is divided into a data frame lookup table and a ring buffer; the data frame lookup table records the data frame number, data storage address, data update status, current available status of the ring buffer, and data locking status information, wherein the data frame number is a unique identifier of the data frame; the data storage address is used to point to a specific location in the ring buffer; the data update status is used to indicate whether the data is valid; the ring buffer space status pointer, including a head pointer and a tail pointer, is used to calculate the current available space indicating the ring buffer; the data locking status is used to prevent conflicts during data transmission; the ring buffer is used for cyclic storage of data frames, and the length is designed according to the actual needs of the system to ensure that there is enough space to store data, and the memory space is not divided in advance; the interrupt control unit is used to process inter-core interrupts.
[0006] Furthermore, each processor core has its own data frame lookup table and a ring buffer.
[0007] A communication method of the above-mentioned on-chip inter-core communication system is special in that the method comprises the following steps:
[0008] 1) Establishing a data frame lookup table between processor cores, the lookup table records the data frame number, data storage address, data update status, current available status of the ring buffer, and data lock status information; the data lock status information is used to prevent conflicts during data transmission;
[0009] 2) Create a circular buffer for storing data frames, and the data frames are stored in a circular manner in a way that the frames are connected head to tail; in the circular buffer, the data storage area is not pre-divided into memory, but is continuously stored in a circular manner according to the actual data frame size;
[0010] 3) Before writing data, the data sending core first checks the available space in the ring buffer and confirms that there is enough space to store the data to be sent; if there is insufficient space, it returns and waits; if there is sufficient space, it proceeds to step 4);
[0011] 4) Create a data frame lookup table belonging to the data frame in the shared memory;
[0012] 5) The data sending core writes the data to be sent into the ring buffer and updates the relevant information in the data frame lookup table, including the data frame number, data storage address, data update status and update tail pointer;
[0013] 6) The data sending core notifies the data receiving core that the data is in place through an inter-core interrupt, or the data receiving core actively queries the data update status in the ring buffer;
[0014] 7) The data receiving core obtains the information of the new data by accessing the data frame lookup table, including the data frame number, data storage address and data length, and reads the corresponding data from the ring buffer.
[0015] Furthermore, the method for determining whether the available space in the ring buffer is sufficient in step 3) is as follows: the number of bytes of remaining space in the ring buffer = head pointer - tail pointer. When the amount of data to be written is ≤ the number of bytes of remaining space in the ring buffer, it is considered that there is enough space, otherwise it is considered that there is not enough space.
[0016] Furthermore, the data frame lookup table belonging to the data frame created in step 4) is specifically:
[0017] Frame number +1
[0018] Frame start address = current tail pointer
[0019] Frame length = number of bytes of data to be written
[0020] Define frame lock state as locked
[0021] The frame status is defined as "data updated".
[0022] Furthermore, the specific steps of step 5) are: according to the data frame lookup table information definition of the data frame, write the data to be sent into the address recorded in the frame start address in the ring buffer according to the frame length, change the frame lock state to unlocked, and update the tail pointer.
[0023] Furthermore, the specific steps of step 6) are: the data receiving core checks whether the corresponding data sending core has an interrupt, or the data receiving core queries whether a new frame number is generated, and the data status is updated, and the frame is in an unlocked state. If not, return and wait. If so, go to step 7).
[0024] Furthermore, the specific steps of step 7) are: searching the data frame lookup table belonging to the data frame in the shared memory, searching the frame start address, frame length, defining the frame lock status, reading the frame length data from the frame start address according to the data lookup table information definition, changing the frame status to "processed", and updating the head pointer.
[0025] The present invention provides an on-chip inter-core communication system and method. Based on a shared memory design, inter-core data synchronization is achieved through an inter-core interrupt or a flag query method. The storage and indexing of data frames are mainly achieved through a data frame lookup table and a ring buffer. Different from the common shared memory method to achieve inter-core communication and divide the data storage space into a number of memory blocks of the same size, the present invention is designed to store data in a ring buffer. The data frame storage frames are connected head to tail, which can maximize the full use of the data storage area space in the shared memory. The present invention allocates the shared memory space, which is generally divided into two parts: a data frame lookup table and a ring buffer. Each core has a data frame lookup table and a ring buffer memory. The data frame lookup table has a fixed format, including the processor core write / read data behavior, data frame number, data storage address, data update status, the current available state of the ring buffer (head pointer, tail pointer), data lock status, etc. When used, multiple lookup table queues are preset (can be changed according to actual usage and actual chip conditions) for cyclic use, and each lookup table corresponds to a data frame; the ring buffer does not perform memory segmentation in advance, and is continuously cyclically stored according to the actual data frame size during use. Before the data sending core writes data to the corresponding ring buffer, it first checks the space availability of the ring buffer, and confirms that there is enough space to store the data sent by the data sending core before performing subsequent operations. After the data sending core writes the data to the corresponding ring buffer, the data sending core can use the inter-core interrupt method to notify the data receiving core that the data is in place, or in applications that do not need or cannot use interrupts, the data receiving core can also actively query whether there is new data in the ring buffer that is ready to be updated. If the data receiving core knows that there is data ready to be updated in the current ring buffer, it can obtain the data information written to the ring buffer by the data sending core by accessing the data frame lookup table, including the data frame number, data frame starting address, data frame length, lock status, update status, etc. According to the above information, the data receiving core can read the data in the ring data buffer and update the data status to old data, thereby achieving the purpose of inter-core communication.
[0026] The present invention does not rely on a dedicated communication bus or Noc hardware support integrated inside the chip, nor does it need to divide the shared memory into several memory blocks of fixed length, thereby avoiding the problem of reduced storage space utilization and inability to cope with large amounts of data communication between cores. On the basis of making full use of the shared memory space, the present invention ensures the quality of data synchronization and improves the response speed of inter-core communication to a certain extent.
[0027] The present invention is applicable to multi-core processor systems in the fields of high performance computing, server applications and mobile devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of the architecture of the on-chip inter-core communication system of the present invention;
[0029] Figure 2 A schematic diagram of a data frame lookup table frame structure and a ring buffer of the present invention;
[0030] Figure 3 A flowchart for implementing the on-chip inter-core communication method of the present invention;
[0031] Figure 4 A data sending flow chart of the data sending core of the present invention;
[0032] Figure 5 This is a flow chart of data receiving by the data receiving core of the present invention. DETAILED DESCRIPTION
[0033] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] See also Figure 1 The specific embodiment structure of the on-chip inter-core communication system of the present invention aims to improve the efficiency and reliability of data transmission between cores in a multi-core processor system, and optimizes the data reading and writing process by adopting a shared memory and a ring buffer.
[0035] The system of the present invention mainly includes multiple processor cores (such as core 1, core 2, core 3, core 4...core n), shared memory and interrupt control unit. Among them, the multiple processor cores are respectively divided into data sending cores and data receiving cores, which are responsible for executing tasks and inter-core communication; the shared memory is divided into a data frame lookup table and a ring buffer; the interrupt control unit is used to handle inter-core interrupts to ensure the timeliness of data transmission. The shared memory structure of the present invention includes two main parts: a data frame lookup table and a ring buffer.
[0036] See also Figure 2 , the data frame lookup table and the ring buffer of the present invention are as follows:
[0037] The data frame number in the data frame lookup table is the unique identifier of the data frame; the data storage address is used to point to the specific location in the ring buffer; the data update status is used to indicate whether the data is valid; the ring buffer space status pointer, including the head pointer and the tail pointer, is used to calculate and indicate the current available space of the ring buffer; the data lock status is used to prevent conflicts during data transmission.
[0038] The ring buffer is used for cyclic storage of data frames. The length is designed according to the actual needs of the system to ensure that there is enough space to store data. The memory space is not divided in advance, which ensures the flexibility of memory use and avoids the waste of storage resources caused by pre-dividing the space.
[0039] See also Figure 3 The specific method flow of the present invention is as follows:
[0040] 1) establishing a data frame lookup table 5 between processor cores, the data lookup table 5 recording the data frame number, data storage address, data update status, current available status of the ring buffer, and data locking status information; the data locking status information is used to prevent conflicts during data transmission;
[0041] 2) creating a circular buffer 8 for storing data frames, wherein the data frames are stored cyclically in a manner that the frames are connected end to end; in the circular buffer 8, the data storage area is not pre-divided into memory, but is continuously stored cyclically according to the actual data frame size;
[0042] 3) Before writing data, the data sending core 1 (such as core 1) first checks the available space of the ring buffer 8 and confirms that there is enough space to store the data to be sent; if the space is insufficient, it returns and waits; if the space is sufficient, it proceeds to step 4);
[0043] 4) creating a data frame lookup table 5 belonging to the data frame in the shared memory 3;
[0044] 5) The data transmission core 1 writes the data to be transmitted into the ring buffer 8, and updates the relevant information in the data frame lookup table 5, including the data frame number, the data storage address, the data update status and the update tail pointer;
[0045] 6) The data sending core 1 notifies the data receiving core 2 (such as core 2) that the data is in place through an inter-core interrupt, or the data receiving core 2 actively queries the data update status in the ring buffer;
[0046] 7) The data receiving core 2 obtains the information of the new data, including the data frame number, data storage address and data length, by accessing the data frame lookup table 5, and reads the corresponding data from the ring buffer 8.
[0047] After writing data into the ring buffer 8, the data sending core 1 can actively or passively notify the data receiving core 2 to receive new data.
[0048] The specific implementation process of the present invention includes two parts: data sending core process and data receiving core process. The specific process is as follows:
[0049] See also Figure 4 , the data transmission core process of the present invention is as follows:
[0050] The data transmission core 1 (such as core 1) first checks whether the available space in the ring buffer 8 is sufficient. If the space is insufficient, it returns and waits. If the space is sufficient, a data frame lookup table 5 belonging to the data frame is created in the shared memory 3; wherein: the method for judging whether the available space in the ring buffer 8 is sufficient is as follows: the number of bytes of the remaining space in the ring buffer = the head pointer 6 - the tail pointer 7. When the amount of data to be written is ≤ the number of bytes of the remaining space in the ring buffer, it is considered that the space is sufficient, otherwise it is considered that the space is insufficient. The data frame lookup table 5 belonging to the data frame is created specifically as follows: frame number + 1, frame start address = current tail pointer, frame length = number of bytes of data to be written, frame lock state is defined as locked, and frame state is defined as "data updated". The data transmission core 1 writes the data to be sent into the ring buffer 8 in the shared memory 3, and updates the relevant information in the data frame lookup table 5. The specific operation includes recording the data frame number, data storage address and data update status in the data frame lookup table 5, and updating the tail pointer 7 to point to the position of the new data.
[0051] Then the data sending core 1 sends an interrupt signal to the data receiving core 2 (such as core 2) through the interrupt control unit 4 to notify it that there is new data to read. The data receiving core 2 can also determine whether the data frame to be received is ready for update by querying whether the data frame number is updated and querying the data update status to determine whether it is new data. If the data receiving core 2 detects the interrupt signal initiated by the data sending core 1 or actively queries to know that the data is ready for update, it can enter the data receiving process.
[0052] See also Figure 5 The data receiving core process of the present invention is as follows:
[0053] The data receiving core 2 (core 2) checks whether the corresponding data sending core 1 has an interrupt, or the data receiving core 2 inquires whether a new frame number is generated, and the data status is updated, and the frame is in an unlocked state. If not, it returns and waits. If so, the data receiving core 2 searches the data frame lookup table 5 belonging to the data frame in the shared memory 3, searches for the frame start address, frame length, and defines the frame lock state. According to the definition of the data lookup table 5 information, the frame length data is read from the frame start address, the frame state is changed to "processed", and the head pointer 6 is updated. The data receiving core 2 reads the corresponding data (including the data frame number and storage address, as well as the data frame length, data update status, etc.) from the ring buffer 8 in the shared memory 3 according to the acquired data frame lookup table 5 information. After completing the data reading, the data receiving core 2 updates the data status in the data frame lookup table 5 and marks the data as "processed". Then the data receiving core 2 performs the data processing task to complete the whole process of inter-core data transmission.
[0054] The above are only specific embodiments disclosed in the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.
[0055] The content of the present invention and the technical content not specifically described in the above embodiments are the same as the prior art.
[0056] The present invention is not limited to the above embodiments, and all of the contents of the present invention can be implemented and have the above good effects.
Claims
1. An on-chip inter-core communication system, characterized in that: The on-chip inter-core communication system comprises a plurality of processor cores, a shared memory and an interrupt control unit, wherein the shared memory is connected to the plurality of processor cores, the plurality of processor cores are respectively connected to the interrupt control unit, the plurality of processor cores are divided into a data sending core and a data receiving core, and are responsible for executing tasks and inter-core communication; the shared memory is internally divided into a data frame lookup table and a ring buffer; The data frame lookup table records the data frame number, data storage address, data update status, current available status of the ring buffer, and data locking status information, wherein the data frame number is a unique identifier of the data frame; The data storage address is used to point to a specific location in the ring buffer; the data update status is used to indicate whether the data is valid; the ring buffer space status pointer, including the head pointer and the tail pointer, is used to calculate and indicate the current available space in the ring buffer; The data lock state is used to prevent conflicts during data transmission; the ring buffer is used for cyclic storage of data frames, and the length is designed according to the actual needs of the system to ensure that there is enough space to store data, and the memory space is not divided in advance; The interrupt control unit is used to process inter-core interrupts.
2. The on-chip inter-core communication system according to claim 1, characterized in that: Each processor core has its own data frame lookup table and a ring buffer.
3. A communication method of the on-chip inter-core communication system as claimed in claim 1, characterized in that: The method comprises the following steps: 1) establishing a data frame lookup table between processor cores, the data lookup table recording data frame number, data storage address, data update status, current available status of the ring buffer, and data lock status information; the data lock status information is used to prevent conflicts during data transmission; 2) Create a circular buffer for storing data frames, and the data frames are stored in a circular manner in a way that the frames are connected head to tail; in the circular buffer, the data storage area is not pre-divided into memory, but is continuously stored in a circular manner according to the actual data frame size; 3) Before writing data, the data sending core first checks the available space in the ring buffer and confirms that there is enough space to store the data to be sent; if there is insufficient space, it returns and waits; if there is sufficient space, it proceeds to step 4); 4) Create a data frame lookup table belonging to the data frame in the shared memory; 5) The data sending core writes the data to be sent into the ring buffer and updates the relevant information in the data frame lookup table, including the data frame number, data storage address, data update status and update tail pointer; 6) The data sending core notifies the data receiving core that the data is in place through an inter-core interrupt, or the data receiving core actively queries the data update status in the ring buffer; 7) The data receiving core obtains the information of the new data by accessing the data frame lookup table, including the data frame number, data storage address and data length, and reads the corresponding data from the ring buffer.
4. The on-chip inter-core communication method according to claim 3, characterized in that: The method for judging whether the available space in the ring buffer is sufficient in the step 3) is as follows: the number of bytes of the remaining space in the ring buffer = the head pointer - the tail pointer. When the amount of data to be written is ≤ the number of bytes of the remaining space in the ring buffer, it is considered that there is enough space, otherwise it is considered that there is not enough space.
5. The on-chip inter-core communication method according to claim 4, characterized in that: The data frame lookup table belonging to the data frame created in step 4) is specifically as follows: Frame number +1 Frame start address = current tail pointer Frame length = number of bytes of data to be written Define frame lock state as locked The frame status is defined as "data updated".
6. The on-chip inter-core communication method according to claim 5, characterized in that: The specific steps of step 5) are: according to the definition of the data frame lookup table information of the data frame, write the data to be sent into the address recorded in the frame start address in the ring buffer according to the frame length, change the frame lock state to unlocked, and update the tail pointer.
7. The on-chip inter-core communication method according to claim 6, characterized in that: The specific steps of step 6) are: the data receiving core checks whether the corresponding data sending core has an interrupt, or the data receiving core queries whether a new frame number is generated, and the data status is updated, and the frame is in an unlocked state. If not, return and wait. If so, enter step 7).
8. The on-chip inter-core communication method according to claim 7, characterized in that: The specific steps of step 7) are: searching the data frame lookup table belonging to the data frame in the shared memory, searching the frame start address, frame length, defining the frame lock state, reading the frame length data from the frame start address according to the definition of the data lookup table information, changing the frame state to "processed", and updating the head pointer.
Citation Information
Patent Citations
Multi-core processor interactive bus design method based on shared memory
CN108959149A
Inter-core communication method and device of multi-core processor
CN110825690A
Guidance information inter-kernel interaction method and system based on shared memory
CN113094324A
Dual-core communication method and electronic equipment
CN113760559A
Deterministic inter-kernel communication method and system and readable storage medium
CN117112255A
Cited By
Inter-core communication method and system based on hardware management and electronic equipment
CN120821586A
A method, system and electronic device for inter-core communication based on hardware management
CN120821586B