Inter-process communication method based on CPU cache

By transmitting data in batches on the CPU's L3 cache and processing it in parallel, the problem of slow inter-process communication is solved, achieving high-speed data transmission and memory saving, and making it suitable for inter-process communication on multiple platforms.

CN120950274APending Publication Date: 2025-11-14EASY THINKING HANGZHOU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511063047.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing inter-process communication methods suffer from slow transmission speeds when dealing with large amounts of data, especially those based on shared memory, which are inefficient due to the need to copy data from the CPU cache to physical memory.

Method used

By transmitting data in batches to the CPU's L3 cache and reading data directly from the cache during the receiving process, multi-threaded parallel processing is used to reduce the need for physical memory reads, and CPU core binding is combined to optimize the data transmission path.

Benefits of technology

It significantly improves the speed of inter-process communication, with data transfer speed approximately twice that of traditional methods, while reducing memory usage. It is applicable to multiple platforms and has universality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950274A_ABST
    Figure CN120950274A_ABST
Patent Text Reader

Abstract

The invention provides an inter-process communication method based on CPU (Central Processing Unit) cache, which comprises the following steps: a write-in control thread obtains the data volume of data to be transmitted, and batches the data to be transmitted by taking a preset single-batch transmission data volume A as a unit; the value of the single-batch transmission data volume A is smaller than the three-level cache capacity of the CPU; data are transmitted in batches by the following steps: a write-in control thread divides a single batch of data into a plurality of parts, and the parts are written into a shared memory in parallel by each channel write thread; temporarily storing the data in a three-level cache of the CPU at the moment of writing the data into the shared memory; each channel read thread reads the corresponding data and then stores the data in the local memory R, and after data transmission is completed, the receiving process reads the data in the local memory R to complete inter-process communication; according to the method, data is temporarily stored in a three-level cache of a CPU (Central Processing Unit) in a mode of transmitting the data in batches, and a receiving process directly reads the data in the three-level cache of the CPU; the data transmission is accelerated, and the inter-process communication speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of inter-process communication, and more specifically to an inter-process communication method based on CPU cache. Background Technology

[0002] Inter-process communication (IPC) refers to the mechanisms for exchanging data and transferring information between different processes. Currently, the methods of inter-process communication include:

[0003] 1) Inter-process communication based on the TCP protocol stack involves multiple layers of disassembly / reassembly / copying, resulting in slow data transmission speed.

[0004] 2) Inter-process communication based on shared memory: This is currently the mainstream inter-process communication method. It allows two or more processes to access the same physical memory, thus enabling data transfer through shared memory. Compared to method 1), this method significantly improves transmission speed. However, its transmission process involves two copies:

[0005] The sending end reads the sending buffer and writes it to shared memory;

[0006] The receiving end reads from shared memory and writes to the receive buffer.

[0007] When the amount of data to be transmitted is large, the shared memory written by the sending end will be gradually replaced in the CPU cache, so that the receiving end can only read data from physical memory. Therefore, inter-process communication based on shared memory also suffers from slow data transmission speed. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention provides an inter-process communication method based on CPU cache. This method temporarily stores data in the CPU's L3 cache by transmitting data in batches, allowing the receiving process to directly read the data from the CPU's L3 cache. This accelerates data transmission and improves the speed of inter-process communication.

[0009] The technical solution is as follows:

[0010] An inter-process communication method based on CPU cache is proposed, which realizes communication between the sending process and the receiving process in the following way;

[0011] First, initialization is performed: the write control thread in the sending process specifies the number of channels, creates shared memory, and starts the write thread for each channel; the read control thread in the receiving process allocates local memory R, ​​specifies the number of channels, connects to the shared memory, and starts the read thread for each channel; initialization is complete.

[0012] The write control thread obtains the amount of data to be transmitted, and divides the data to be transmitted into batches based on a preset single batch transmission data amount A, and obtains the total batch. The value of the single batch transmission data amount A is less than the CPU L3 cache capacity.

[0013] Use the following steps to transfer data in batches:

[0014] S1. The write control thread divides the current batch of data into multiple parts, which are written to the shared memory in parallel by the write threads of each channel. At the moment the data is written to the shared memory, it is temporarily stored in the CPU L3 cache.

[0015] When all data in a single batch is written to the shared memory, each channel write thread notifies each channel read thread to send the end of the write. At the same time, each channel read thread reads the corresponding data and then stores the read data into the local memory R. When the storage is finished, each channel read thread notifies each channel write thread, and each channel write thread then notifies the write control thread that the single batch of data has been sent.

[0016] S2. The write control thread determines whether the batch of data sent is equal to the total batch. If so, the write control thread notifies the read control thread that the data transmission is complete and executes step S3. If not, it continues to execute step S1 to transmit the next batch of data.

[0017] S3. The read control thread sends the address of local memory R to the receiving process. The receiving process reads data from local memory R. After reading is complete, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.

[0018] Preferably, during initialization, the shared memory created by the control thread includes the shared memory S corresponding to each channel. j Where j represents the j-th channel, j = 0, 1, 2, ..., G-1; G represents the number of channels, and S represents the number of channels in a single shared memory. j The cache size is a preset value Q; where Q × G equals the amount of data A transmitted in a single batch.

[0019] The read control thread in the receiving process specifies the number of channels G, and also specifies the buffer size Q for each channel, and allocates shared memory S. j Each is connected to a corresponding channel read thread.

[0020] Furthermore, in step S1, the write control thread divides the current batch of single-batch data into multiple parts in units of Q, and each channel write thread determines the data to be transmitted in the local memory of the sending process according to the starting address w plus the offset.

[0021] Each channel's write thread then writes the predetermined data in parallel to its corresponding shared memory S. j ;

[0022] Wherein, the starting address w is the address of the data to be transmitted in the local memory of the sending process, which is provided by the business layer of the sending process through the copy interface of the write control thread; the offset is used to distinguish the batch in which the data is located and the transmission channel.

[0023] Furthermore, in step S1, each channel read thread reads the corresponding data and then stores the read data into the following location in local memory R: the starting address of local memory R + offset, where the offset is used to distinguish the batch and transmission channel of each piece of data.

[0024] Furthermore, the offset is equal to the amount of data transmitted in a single batch, A×i+Q×j, where i represents the current batch being transmitted, i=0,1,2……M-1, and M is the total number of batches.

[0025] Preferably, the value of the data volume A in a single batch is less than one-third of the CPU's L3 cache capacity.

[0026] Preferably, the initialization step further includes:

[0027] The write control thread and the read control thread are each bound to the same CPU core;

[0028] Each channel's write thread is bound to a different CPU core, while the read and write threads of the same channel are bound to the same CPU core.

[0029] Furthermore, during initialization, the shared memory created by the write control thread also includes shared memory I; the read control thread in the receiving process is attached to shared memory I.

[0030] When data transmission begins, the write control thread obtains the amount of data N to be transmitted and stores the amount of data N into the shared memory I;

[0031] The write control thread notifies the read control thread that the transfer has started. The read control thread reads the data amount N from shared memory I and sends a transfer start confirmation signal to the write control thread.

[0032] Furthermore, in step S3, the read control thread sends the address of the local memory R and the amount of data N to the receiving process. The receiving process reads the data from the local memory R according to the address and the amount of data N. After the reading is completed, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.

[0033] This method has the following characteristics:

[0034] ① High transmission speed:

[0035] This method batches data, enabling on-demand storage and retrieval on the CPU's L3 cache. Therefore, the receiving process does not need to read physical memory, significantly improving speed. Furthermore, by binding CPU cores during initialization, the time for a single thread switch-out and switch-back on the same core is only about 5µs. Read and write threads on the same channel are bound to the same core, and L2 cache acceleration can be utilized for switching and interaction. Simultaneously, multiple pairs of threads running on multiple cores can communicate in parallel using the L3 cache at high speed, further achieving high-speed signal exchange and data transmission. Under the same conditions, the data transmission speed of this method is approximately twice that of traditional shared memory-based IPC data transmission.

[0036] ②Low memory usage:

[0037] Because this method transmits data in batches, the shared memory required is significantly less than that of traditional methods. For example, with 8 concurrent connections and a single-channel transmission limit of 100MB, traditional shared memory-based IPC uses 800MB of shared memory, while this method uses approximately 16MB.

[0038] ③ High versatility:

[0039] This method transmits data without requiring byte alignment; it is universally applicable to platforms such as Windows and ARM embedded systems. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the data batch transmission process in a specific implementation method. Detailed Implementation

[0041] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] An inter-process communication method based on CPU cache is proposed, which realizes communication between the sending process and the receiving process in the following way;

[0043] First, initialization is performed: the write control thread in the sending process specifies the number of channels, creates shared memory, and starts the write thread for each channel; the read control thread in the receiving process allocates local memory R, ​​specifies the number of channels, connects to the shared memory, and starts the read thread for each channel; initialization is complete.

[0044] The write control thread obtains the amount of data to be transmitted, and divides the data to be transmitted into batches based on a preset batch transmission data amount A, and obtains the total batch. The value of the batch transmission data amount A is less than the CPU L3 cache capacity.

[0045] Use the following steps to transfer data in batches:

[0046] S1. The write control thread divides the current batch of data into multiple parts, which are written to the shared memory in parallel by the write threads of each channel. At the moment the data is written to the shared memory, it is temporarily stored in the CPU L3 cache.

[0047] When all data in a single batch is written to the shared memory, each channel write thread notifies each channel read thread to send the end of the write. At the same time, each channel read thread reads the corresponding data and then stores the read data into the local memory R. When the storage is finished, each channel read thread notifies each channel write thread, and each channel write thread then notifies the write control thread that the single batch of data has been sent.

[0048] S2. The write control thread determines whether the batch of data sent is equal to the total batch. If so, the write control thread notifies the read control thread that the data transmission is complete and executes step S3. If not, it continues to execute step S1 to transmit the next batch of data.

[0049] S3. The read control thread sends the address of local memory R to the receiving process. The receiving process reads data from local memory R. After reading is complete, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.

[0050] Specifically, the read control thread sends the address of local memory R to the business layer in the receiving process, and the business layer reads data from local memory R.

[0051] More specifically, during initialization, the shared memory created by the control thread includes the shared memory S corresponding to each channel. j Where j represents the j-th channel, j = 0, 1, 2, ..., G-1; G represents the number of channels, and S represents the number of channels in a single shared memory. j The buffer size is a preset value Q; where Q × G equals the amount of data transmitted in a single batch A;

[0052] The read control thread in the receiving process specifies the number of channels G, and also specifies the buffer size Q for each channel, and allocates shared memory S. j Each is connected to a corresponding channel read thread.

[0053] In practice, to ensure the correspondence between data writing and data reading, a single batch of data is located using an address + offset method. Specifically:

[0054] In step S1, the write control thread divides the current batch of single-batch data into multiple parts in units of Q. Each channel write thread determines the data to be transmitted in the local memory of the sending process according to the starting address w plus the offset.

[0055] Each channel's write thread then writes the predetermined data in parallel to its corresponding shared memory S. j ;

[0056] Here, the starting address w is the address of the data to be transmitted in the local memory of the sending process, which is provided by the business layer of the sending process through the copy interface of the write control thread; the offset is used to distinguish the batch of data and the transmission channel.

[0057] In step S1, each channel read thread reads the corresponding data and then stores the read data into the following location in local memory R: the starting address of local memory R + offset. The offset is used to distinguish the batch and transmission channel of each piece of data.

[0058] Specifically, the offset = the amount of data transmitted in a single batch A×i + Q×j, where i represents the current batch of transmission, i = 0, 1, 2...M-1, M is the total number of batches, and the total number of batches M = ceil(N / A), where ceil represents the rounding up function; N is the amount of data to be transmitted; j represents the j-th channel, j = 0, 1, 2...G-1.

[0059] During initialization, the shared memory created by the write control thread includes shared memory I; the read control thread in the receiving process is attached to shared memory I.

[0060] When data transmission begins, the write control thread obtains the amount of data N to be transmitted and stores the amount of data N into shared memory I;

[0061] The write control thread notifies the read control thread that the transfer has started. The read control thread reads the data amount N from shared memory I and sends a transfer start confirmation signal to the write control thread.

[0062] In step S3, the read control thread sends the address of the local memory R and the amount of data N to the receiving process (the business layer in the receiving process). The receiving process (the business layer in the receiving process) reads the data from the local memory R according to the address and the amount of data N. After the reading is completed, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.

[0063] To further improve transmission efficiency, as a preferred implementation method, the value of the data volume A transmitted in a single batch is less than one-third of the CPU's L3 cache capacity.

[0064] The following is a comparative experimental analysis of the different values ​​of the data volume A in a single batch transmission:

[0065] Taking the CPU model i5-8500-6C6T, CPU frequency 3.9GHz, number of channels 4, and L3 capacity 9M as an example, the comparative experimental data of different values ​​of single batch data transfer A are shown in the table below:

[0066] Single batch data transmission volume A 512KB 2048KB 2560KB 3072KB 4096KB Overall transmission speed 6.4GB / s 7.2GB / s 7.0GB / s 6.6GB / s 5.6GB / s

[0067] Through comparative experiments, it was found that when the value of the single batch data transmission volume A is less than one-third of the L3 capacity, the overall data transmission speed shows an upward trend. When the value of the single batch data transmission volume A exceeds one-third of the L3 capacity, the overall data transmission speed shows a downward trend. Therefore, the preferred value of the single batch data transmission volume A is less than one-third of the L3 capacity.

[0068] Furthermore, since the time it takes for a thread on the same CPU core to switch out and back is only about 5us, pre-binding the CPU core can further improve data transfer efficiency.

[0069] Therefore, a more preferred implementation is as follows: the initialization step further includes:

[0070] The write control thread and the read control thread are each bound to the same CPU core;

[0071] Each channel's write thread is bound to a different CPU core, while the read and write threads of the same channel are bound to the same CPU core.

[0072] The following is a comparative experimental analysis of pre-bound and unbound CPU cores:

[0073] With CPU model i5-8500-6C6T, CPU frequency 3.9GHz, number of channels 1, and a single shared memory S j Taking a cache size Q = 1536KB, an L3 capacity of 9M, and a single batch data transfer volume A of 1536KB as an example, a comparative experiment was conducted with the CPU cores pre-locked and unlocked. The results are shown in the table below:

[0074]

[0075] Comparative experiments revealed that pre-locking the CPU core during the initialization process can improve the overall data transmission speed.

[0076] To facilitate understanding of the technical solution of this application, the following exemplary description is provided:

[0077] An inter-process communication method based on CPU cache, first performing initialization:

[0078] The write control thread in the sending process creates shared memory I, specifies the number of channels G, and creates shared memory S corresponding to each channel. j j represents the j-th channel, j = 0, 1, 2...G-1; a single shared memory S j The cache size is a preset value Q; where Q×G represents the amount of data transferred in a single batch A, and the value of the amount of data transferred in a single batch A is less than the CPU L3 cache capacity.

[0079] Start the write thread for each channel;

[0080] Meanwhile, the read control thread in the receiving process is connected to shared memory I; local memory R is allocated to store the data to be transmitted; the number of channels G is specified, and the buffer size Q for each channel is also specified, and each channel is connected to shared memory S respectively. j ;

[0081] Start the read thread for each channel;

[0082] Initialization complete;

[0083] The write control thread obtains the data volume N to be transmitted, stores the data volume N in shared memory I, and divides the data to be transmitted into batches according to the data volume N and the single batch transmission data volume A, using the preset single batch transmission data volume A as the unit, and obtains the total batch.

[0084] The write control thread notifies the read control thread that the transmission has started. The read control thread reads the data amount N from shared memory I and sends a transmission start confirmation signal to the write thread.

[0085] For ease of explanation: such as Figure 1 As shown, taking the data volume to be transmitted N=64B, the data volume of a single batch A=32B, the number of channels G=2 (channel 0, channel 1), the buffer size Q=16B, and the total number of batches M=64 / 32=2 as an example;

[0086] Use the following steps to transfer data in batches:

[0087] S1. The write control thread divides the current batch of data (32B) to be transmitted into two parts in units of Q=16B. The two channel write threads determine the data to be transmitted in the local memory of the sending process according to the method of starting address w + batch transmission data A×i + Q×j.

[0088] The two-channel write threads then write the determined data in parallel to the corresponding shared memory S0 and S1 respectively;

[0089] Data is written to shared memory S j For a moment, it is temporarily stored in the CPU's L3 cache;

[0090] When a single batch of data is all written to shared memory S j At the same time, the two channel write threads send write end signals to the two channel read threads respectively. At the same time, the two channel read threads read data (channel 0 read thread reads data in shared memory S0, and channel 1 read thread reads data in shared memory S1). Then, the read data is stored in the following location in local memory R: the starting address of local memory R + the amount of data transferred in a single batch A×i + Q×j.

[0091] After a batch of data is stored in local memory R, ​​each channel read thread sends a read completion signal to each channel write thread; after receiving the read completion signal, each channel write thread sends a batch of data sending completion signal to the write control thread.

[0092] S2. The write control thread determines whether the batch of data sent is equal to the total batch 2. If yes, the write control thread notifies the read control thread that the data transmission is complete and executes step S3. If no, it continues to execute step S1 to transmit the next batch of data.

[0093] S3. The write control thread notifies the read control thread that the transmission is complete. The read control thread sends the address of the local memory R and the amount of data N=64B to the business layer in the receiving process. The business layer reads the data from the local memory R according to the address and the amount of data N=64B. After reading is completed, the read control thread sends a transmission completion confirmation signal to the write control thread to complete the inter-process communication.

[0094] Specifically, the data to be transmitted is divided into batches based on a preset batch transmission data size A. The data size of the last batch is NA × floor(N / A), where floor represents the floor function.

[0095] If the data volume of the last batch is less than the data volume A of a single batch (data is insufficient for the entire batch);

[0096] In step S1, the write control thread divides the last batch of data into portions Q. If the data size Q' of the j-th portion is less than Q, then the channel j write thread only writes the data size Q' to the shared memory, and the channel j read thread only reads the data size Q'. If the data size of the j-th portion is 0, then the channel j write thread does not need to execute.

[0097] For example, if the last batch of data is 10 bytes and Q is 7 bytes, then the first batch of data is 7 bytes and the second batch of data is 3 bytes. At this time, the channel 0 write thread determines the 7 bytes of data to be transmitted in the sending process's local memory according to the starting address w + A × (M-1) + Q × 0; it writes the determined 7 bytes of data in parallel to the corresponding shared memory S0. The channel 1 write thread determines the 3 bytes of data to be transmitted in the sending process's local memory according to the starting address w + A × (M-1) + Q × 1; it writes the determined 3 bytes of data in parallel to the corresponding shared memory S1. The channel 0 read thread reads the 7 bytes of data from shared memory S0 and stores the read data in the following location in local memory R: the starting address of local memory R + A × (M-1) + Q × 0. The channel 1 read thread reads the 3 bytes of data from shared memory S1 and stores the read data in the following location in local memory R: the starting address of local memory R + A × (M-1) + Q × 1.

[0098] For example: if the last batch of data is 4B and Q is 7B, then the first batch of data is 4B and the second batch of data is 0. At this time, the channel 0 write thread determines the 4B of data to be transmitted in the local memory of the sending process according to the starting address w + A × (M-1) + Q × 0. The determined 4B of data is written in parallel to the corresponding shared memory S0. The channel 0 read thread reads the 4B of data in the shared memory S0 and stores the read data in the following location in the local memory R: the starting address of the local memory R + A × (M-1) + Q × 0.

[0099] The channel 1 write thread does not need to write to shared memory S1 and does not need to execute.

[0100] This solution divides the data into batches, enabling instant storage and retrieval on the CPU's L3 cache. As a result, the receiving process does not need to read physical memory, significantly improving speed.

[0101] The foregoing description of specific exemplary embodiments of the present invention is for illustrative and descriptive purposes. It is not intended to be exhaustive, nor to limit the invention to the precise forms disclosed; obviously, many changes and variations are possible in accordance with the foregoing teachings. The exemplary embodiments were chosen and described to explain the specific principles of the invention and its practical application, thereby enabling others skilled in the art to implement and utilize various exemplary embodiments of the invention, as well as their different alternatives and modifications. The scope of the invention is intended to be defined by the appended claims and their equivalents.

Claims

1. A method for inter-process communication based on CPU cache, characterized in that, The following methods can be used to achieve communication between the sending process and the receiving process; First, initialization is performed: the write control thread in the sending process specifies the number of channels, creates shared memory, and starts the write thread for each channel; the read control thread in the receiving process allocates local memory R, ​​specifies the number of channels, connects to the shared memory, and starts the read thread for each channel. Initialization complete; The write control thread obtains the amount of data to be transmitted, and divides the data to be transmitted into batches based on a preset single batch transmission data amount A, and obtains the total batch. The value of the single batch transmission data amount A is less than the CPU L3 cache capacity. Use the following steps to transfer data in batches: S1. The write control thread divides the current batch of data into multiple parts, which are written to the shared memory in parallel by the write threads of each channel. At the moment the data is written to the shared memory, it is temporarily stored in the CPU L3 cache. When all data in a single batch is written to the shared memory, each channel write thread notifies each channel read thread to send the end of the write. At the same time, each channel read thread reads the corresponding data and then stores the read data into the local memory R. When the storage is finished, each channel read thread notifies each channel write thread, and each channel write thread then notifies the write control thread that the single batch of data has been sent. S2. The write control thread determines whether the batch of data sent is equal to the total batch. If so, the write control thread notifies the read control thread that the data transmission is complete and executes step S3. If not, it continues to execute step S1 to transmit the next batch of data. S3. The read control thread sends the address of local memory R to the receiving process. The receiving process reads data from local memory R. After reading is complete, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.

2. The inter-process communication method based on CPU cache as described in claim 1, characterized in that: During initialization, the shared memory created by the control thread is written to include the shared memory S corresponding to each channel. j Where j represents the j-th channel, j = 0, 1, 2, ..., G-1; G represents the number of channels, and S represents the number of channels in a single shared memory. j The cache size is a preset value Q; where Q × G equals the amount of data A transmitted in a single batch. The read control thread in the receiving process specifies the number of channels G, and also specifies the buffer size Q for each channel, and allocates shared memory S. j Each is connected to a corresponding channel read thread.

3. The inter-process communication method based on CPU cache as described in claim 2, characterized in that: In step S1, the write control thread divides the current batch of single-batch data into multiple parts in units of Q. Each channel write thread determines the data to be transmitted in the local memory of the sending process according to the starting address w plus the offset. Each channel's write thread then writes the predetermined data in parallel to its corresponding shared memory S. j ; Wherein, the starting address w is the address of the data to be transmitted in the local memory of the sending process, which is provided by the business layer of the sending process through the copy interface of the write control thread; the offset is used to distinguish the batch in which the data is located and the transmission channel.

4. The inter-process communication method based on CPU cache as described in claim 2, characterized in that: In step S1, each channel read thread reads the corresponding data and then stores the read data into the following location in local memory R: the starting address of local memory R + offset, where the offset is used to distinguish the batch and transmission channel of each piece of data.

5. The inter-process communication method based on CPU cache as described in claim 3 or 4, characterized in that: The offset is equal to the amount of data transmitted in a single batch, A×i+Q×j, where i represents the current batch being transmitted, i = 0, 1, 2...M-1, and M is the total number of batches.

6. The inter-process communication method based on CPU cache as described in claim 1, characterized in that: The value of the data volume A in a single batch is less than one-third of the CPU's L3 cache capacity.

7. The inter-process communication method based on CPU cache as described in claim 1, characterized in that: The initialization steps also include: The write control thread and the read control thread are each bound to the same CPU core; Each channel's write thread is bound to a different CPU core, while the read and write threads of the same channel are bound to the same CPU core.

8. The inter-process communication method based on CPU cache as described in claim 1, characterized in that: During initialization, the shared memory created by the write control thread also includes shared memory I; the read control thread in the receiving process is attached to shared memory I. When data transmission begins, the write control thread obtains the amount of data N to be transmitted and stores the amount of data N into the shared memory I; The write control thread notifies the read control thread that the transfer has started. The read control thread reads the data amount N from shared memory I and sends a transfer start confirmation signal to the write control thread.

9. The inter-process communication method based on CPU cache as described in claim 8, characterized in that: In step S3, the read control thread sends the address of the local memory R and the amount of data N to the receiving process. The receiving process reads the data from the local memory R according to the address and the amount of data N. After the reading is completed, the read control thread notifies the write control thread that the reading is finished, thus completing the inter-process communication.