Shared memory communication method and device, intelligent terminal and storage medium

By dividing the memory space into message body and terminal cluster areas, and using a two-level memory segmentation strategy algorithm and a chainless list, the problem of increasing the number of tasks and dependence on other media in shared memory communication is solved, and efficient multi-task parallel communication is achieved.

CN120560873APending Publication Date: 2025-08-29SHENZHEN YUNYU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455234.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

There are problems in existing shared memory communication technologies that increase the number of tasks and dependence on other transmission media, especially when multitasking is in parallel, slice splitting operations lead to increased performance expenses and cannot be used in scenarios where only shared memory is available.

Method used

The memory space is divided into a message body area and a terminal cluster area. The memory two-level memory segmentation strategy algorithm is used to allocate memory to the message data, and message delivery is realized through a lock-free list, eliminating dependence on other media, and allowing variable-length message transmission.

Benefits of technology

Improve communication performance, reduce performance expenses for slicing and reorganization, reduce memory waste of multi-party communication, increase parallelism and adaptability, and avoid dependence on other media.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560873A_ABST
    Figure CN120560873A_ABST
Patent Text Reader

Abstract

The invention discloses a shared memory communication method. The method comprises the following steps: dividing a memory space into a message body region and a terminal cluster region; obtaining message data, and distributing memory for the message data through a memory two-stage segmentation strategy algorithm in a message body region to obtain an idle memory; obtaining a message body allocated to the sender according to the message data and the idle memory; and in the terminal cluster area, transmitting the message body allocated to the sender to the receiver task. According to the method and the device, the message queue is placed in the shared memory in the form of the terminal, so that the dependence on other media is eliminated, and the adaptability is improved. A lock-free data structure and a memory two-stage segmentation strategy algorithm are combined, so that the frequent performance expenditure of slicing and recombination is avoided, and the memory communication efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of memory management, and in particular to a shared memory communication method, device, intelligent terminal and storage medium. Background Art

[0002] As single-core computer processor performance reaches its peak, processor manufacturers aim to increase the number of cores to achieve performance gains through parallelism. However, when running multiple tasks in parallel, inter-task communication becomes a major performance bottleneck, necessitating shared memory communication technology. Existing shared memory communication mechanisms fall into two categories: those that require no additional media, have a fixed message storage space, and both parties can update memory usage; and those that require an additional media, guarantee concurrency safety, and have a variable message storage space. The first fixed-memory solution typically requires slicing and splitting messages of varying sizes during transmission, and the receiver must reassemble the fragmented messages to receive the complete message. This limitation increases the number of tasks and, in some cases, requires transmitting messages larger than the memory allocated for them. The second method, which requires an additional media, suffers from a core drawback: the sender and receiver cannot simultaneously update memory usage, resulting in communication conflicts. Furthermore, its reliance on an additional transmission medium makes this solution unsuitable for scenarios where only shared memory is available.

[0003] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a shared memory communication method in response to the above-mentioned defects of the prior art, aiming to solve the problems of increased number of tasks caused by slicing and dependence on other transmission media in the prior art.

[0005] The technical solutions adopted by the present invention to solve the technical problems are as follows:

[0006] In a first aspect, the present invention provides a shared memory communication method, wherein the method comprises:

[0007] Divide the memory space into a message body area and a terminal cluster area;

[0008] Acquire message data, and allocate memory for the message data in the message body area using a two-level memory segmentation strategy algorithm to obtain free memory;

[0009] Obtaining a message body allocated to the sender based on the message data and free memory;

[0010] In the terminal cluster area, the message body assigned to the sender is communicated to the receiver task that performs the dequeue operation.

[0011] In one implementation, dividing the memory space into a message body area and a terminal cluster area includes:

[0012] Dividing the memory space into a message body area, wherein the message body area includes a message body and a dormant message body; wherein the message body is used to manage allocated memory, and the dormant message body is used to manage free memory;

[0013] A terminal cluster area is divided in the memory space, and a plurality of terminal clusters are established in the terminal cluster area, wherein the terminal cluster is a first-in-first-out queue without a lock link list.

[0014] In one implementation, dividing the memory space into terminal cluster areas and establishing a plurality of terminal clusters in the terminal cluster areas includes:

[0015] Several terminal clusters are established in the terminal cluster area, and a terminal node substitute is paired with each terminal in the terminal cluster area, wherein each terminal cluster includes a terminal identifier, a terminal quantity, an insertion subscript and a terminal pair; wherein the terminal identifier is used to identify and search the terminal cluster, the terminal quantity is the number of available terminals in the terminal cluster, the insertion subscript is used to locate the terminal selected for insertion, and the terminal pair is a paired terminal and a terminal node substitute.

[0016] In one implementation, allocating memory for the message data in the message body area by a two-level memory segmentation strategy algorithm to obtain free memory includes:

[0017] In the dormant message body in the message body area, starting from the end address, memory is allocated to the message data through a two-level memory segmentation strategy algorithm to obtain free memory, wherein the end address is the address of the dormant message body adjacent to the terminal cluster area.

[0018] In one implementation, the task of transferring the message body allocated to the sender to the receiver in the terminal cluster area includes:

[0019] Obtaining the number of receiver tasks and the number of terminals, wherein the number of terminals can be dynamically adjusted;

[0020] If the number of receiver tasks is greater than or equal to the number of terminals, each receiver task is assigned to a terminal. After the message body assigned to the sender is enqueued, the message body assigned to the sender is passed to the receiver task, and the terminal corresponding to the receiver task performs the dequeuing operation.

[0021] If the number of the receiver tasks is less than the number of terminals, the terminal is scrolled and selected to perform a dequeue operation on the message body assigned to the sender.

[0022] In one implementation, the method further includes:

[0023] Dividing the memory space into a directory area, and creating a memory two-level partitioning strategy node in the directory area, wherein the memory two-level partitioning strategy node is a data structure based on a bidirectional lock-free linked list;

[0024] The state of the memory area is stored in the directory area, and the memory linked list of the memory two-level segmentation strategy algorithm is managed; wherein the state of the memory area includes uninitialized, initializing and initialized.

[0025] In one implementation, storing the state of the memory area through the directory area and managing the memory linked list of the two-level memory segmentation strategy algorithm includes:

[0026] Storing the state of the memory area at the starting address of the directory area;

[0027] Establishing a memory linked list of a two-level memory segmentation strategy algorithm, and locating a bitmap of message data according to the memory linked list of the two-level memory segmentation strategy algorithm;

[0028] Obtaining required message space according to the bitmap of the message data;

[0029] According to the required message space, a dormant message body holding the required message space is obtained;

[0030] In the dormant message body holding the required message space, a memory two-level partitioning strategy node is arbitrarily selected for dequeuing and enqueuing operations, and the memory two-level partitioning strategy node is used as the head of the memory linked list of the memory two-level partitioning strategy algorithm.

[0031] In a second aspect, an embodiment of the present invention further provides a shared memory communication device, wherein the device includes:

[0032] A region partitioning module is used to divide the memory space into a message body region and a terminal cluster region;

[0033] A memory allocation module is used to obtain message data and allocate memory for the message data in the message body area by using a two-level memory segmentation strategy algorithm to obtain free memory;

[0034] A message body acquisition module, configured to obtain a message body allocated to the sender based on the message data and free memory;

[0035] The memory communication module is used to transmit the message body assigned to the sender to the receiver task in the terminal cluster area.

[0036] In a third aspect, an embodiment of the present invention further provides an intelligent terminal, wherein the intelligent terminal includes a memory, a processor, and a shared memory communication program stored in the memory and runnable on the processor, and when the processor executes the shared memory communication program, the steps of the shared memory communication method as described in any one of the above items are implemented.

[0037] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a shared memory communication program is stored on the computer-readable storage medium, and when the shared memory communication program is executed by a processor, the steps of the shared memory communication method as described in any one of the above items are implemented.

[0038] Beneficial effects: Compared with the prior art, the present invention provides a shared memory communication method, which first divides the memory space into a message body area and a terminal cluster area. By dividing the shared memory area, the dependence on other media is eliminated and the adaptability is increased. By setting the terminal cluster area, the concurrency conflict caused by the use of a single message queue when multiple receiver tasks perform dequeue operations is weakened, thereby increasing parallelism. Then, the message data is obtained, and in the message body area, memory is allocated to the message data according to the memory two-level segmentation strategy algorithm to obtain the message body assigned to the sender. By introducing the memory two-level segmentation strategy algorithm memory management algorithm, messages of variable length are allowed, the regular performance overhead of slicing and reorganization is avoided, and the communication performance is improved. Finally, in the terminal cluster area, the message body assigned to the sender is passed to the receiving task. Since different senders can use each dormant message body in the same message body area, the memory waste of multi-party communication is greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 It is a flow chart of a shared memory communication method provided by an embodiment of the present invention.

[0041] Figure 2 This is a schematic diagram of two copies provided by an embodiment of the present invention.

[0042] Figure 3 This is a schematic diagram of a single copy provided by an embodiment of the present invention.

[0043] Figure 4 This is a schematic diagram of zero copy provided by an embodiment of the present invention.

[0044] Figure 5 This is a schematic diagram of slice splitting and transmission provided by an embodiment of the present invention.

[0045] Figure 6 It is a schematic diagram of the ring array structure provided by an embodiment of the present invention.

[0046] Figure 7 This is a schematic diagram of XSIM memory partitioning provided by an embodiment of the present invention.

[0047] Figure 8 This is a schematic diagram of bitmap-node pairing provided by an embodiment of the present invention.

[0048] Figure 9 This is a schematic diagram of a terminal cluster provided by an embodiment of the present invention.

[0049] Figure 10 This is a schematic diagram of rolling enqueue provided by an embodiment of the present invention.

[0050] Figure 11 This is a schematic diagram of static terminal allocation provided by an embodiment of the present invention.

[0051] Figure 12 This is a schematic diagram of the rolling dequeue process provided by an embodiment of the present invention.

[0052] Figure 13 It is a schematic diagram of the terminal reduction process provided by an embodiment of the present invention.

[0053] Figure 14 It is a schematic diagram of the structure of a dormant message body provided by an embodiment of the present invention.

[0054] Figure 15 It is a schematic diagram of the message body structure provided by an embodiment of the present invention.

[0055] Figure 16 This is a principle block diagram of a shared memory communication device provided by an embodiment of the present invention.

[0056] Figure 17 This is a block diagram of the internal structure of the smart terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0058] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0059] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0060] As the performance of single-core computer processors reaches its peak, processor manufacturers aim to achieve performance gains through parallelism by increasing the number of cores. However, when running multiple tasks in parallel, inter-task communication becomes a major performance bottleneck. To facilitate faster inter-task communication between different tasks within a single machine, random access memory (RAM), which is shared access by all processor cores, is used as a medium for message transmission. Similar to transmitting sound messages through air by generating sound waves, different tasks can store messages in shared memory, which is then read by the receiving task. This is shared memory communication technology. Unlike shared memory for data sharing, shared memory communication prioritizes immediacy and one-time content. Therefore, individual messages can be released immediately after transmission, eliminating the need for version control and synchronization. At the same time, in a producer-consumer model, it is desirable to minimize consumer latency and ensure that messages are delivered as quickly as possible.

[0061] For shared memory communication technology, there are three possible operations based on the number of copies:

[0062] (1) Data is copied twice. The first copy is from the sender's memory to the shared memory, and the second copy is from the shared memory to the receiver's memory. Figure 2 shown.

[0063] (2) One-time data copy, the first copy is from the sender's memory to the shared memory, and the receiver directly operates the shared memory, as shown in the attached Figure 3 shown.

[0064] (3) Zero data copy: the sender directly writes data to the shared memory, and the receiver directly operates the shared memory, as shown in the following example. Figure 4 shown.

[0065] In the shared memory communication mechanism, it is based on whether the message data storage space involved in the communication is variable, whether it is lock-free, whether it needs to use other media other than shared memory, and whether both communicating parties can update the memory usage status. Currently, there is no need for other media, the message storage space is fixed, and both parties can update the memory usage status. The mainstream solutions include Fastbox, Lamport Queue, Fast Forward and Fast FIFO Queue, among which MPMC Queue is the most efficient implementation solution and has mature application cases in the quantitative industry. MPMC Queue allows multiple producers and consumers to add or remove messages in a circular queue at the same time, and is lock-free based on the replacement instruction (Compare and Swap, CAS). However, this type of fixed space solution requires slicing and splitting operations when messages of different sizes need to be transmitted, and the receiver needs to reassemble the split messages to finally obtain the complete message, as shown in the attached figure. Figure 5 As shown. Therefore, in application scenarios where the message size fluctuates greatly, a large amount of useless recurring performance overhead will be spent on slicing and reassembling, and two copies will be unavoidable. Another drawback is that this design cannot allow multi-party communication to share the same memory, that is, communication between A and B and communication between B and A require separate memory planning and setting up separate ring arrays. This limitation will result in some memory being idle for a long time due to lack of communication as the number of tasks requiring multi-party communication increases. At the same time, some communications will need to transmit messages that are larger than the memory they are divided into, as shown in the attached figure. Figure 6 shown.

[0066] For solutions that require other media, and the media needs to ensure concurrency safety and variable message storage space, the more mature one is the transmission solution used in OPTEE, which relies on the Secure Monitor Call (SMC) to transmit the memory address allocated by the sender's memory manager to store the message content to the receiver. After the message is transmitted, the sender needs to release the memory through the memory manager, as shown in the attached Figure 7As shown in the figure. Since this model only has a single producer and consumer, and the producer and consumer do not access shared memory simultaneously, the overall implementation does not need to intentionally maintain concurrency safety. The core flaw of this solution is that the sender and receiver cannot simultaneously update the memory usage status. As a result, for messages that do not require a response from the receiver, the sender still needs to wait for the receiver to complete before releasing the memory and continuing the task. Furthermore, the reliance on other transmission media makes this solution unusable in scenarios where only shared memory is available.

[0067] This paper proposes an enhanced shared-memory intercommunication mechanism (XSIM). This mechanism allows multi-terminal communication that only requires shared memory. It is primarily suitable for cloud server scenarios with multiple virtual machines, where different virtual machines proxy work to each other. By introducing a lock-free linked list and combining it with the Two-Level Segregated Fit (TLSF) algorithm, it combines the advantages of the two mainstream solutions mentioned above.

[0068] Exemplary Methods

[0069] This embodiment provides a shared memory communication method, which can be applied to memory. Figure 1 As shown, the method includes the following steps:

[0070] Step S100: Divide the memory space into a message body area and a terminal cluster area;

[0071] In one implementation, step S100 in this embodiment includes the following steps:

[0072] Step S101: Divide the memory space into a message body area, wherein the message body area includes a message body and a dormant message body; wherein the message body is used to manage allocated memory, and the dormant message body is used to manage free memory;

[0073] In this embodiment, as shown in the attached Figure 8 As shown in the figure, XSIM divides the memory into three areas: message body area, terminal cluster area and directory area.

[0074] Specifically, the Message Construct (MC) area is used to store message data used during communication. Each message's allocated storage space does not need to be consistent, and the allocated storage space is released after the message ends. The Message Construct manages each segment of memory allocated to the sender, while the Dormant Message Construct (DMC) manages each segment of free memory. Initially, only one Dormant Message Construct manages the entire Message Construct area.

[0075] Step S102: dividing the memory space into a terminal cluster area, and establishing a plurality of terminal clusters in the terminal cluster area, wherein the terminal cluster is a first-in-first-out queue without a lock link list;

[0076] Specifically, in the Terminal Cluster area, each terminal is a FIFO (first-in, first-out) lock-free list responsible for conveying the message body assigned to the sender to a receiver task currently dequeuing the message at the current terminal. A receiver or sender can have multiple tasks, and a terminal cluster can also have multiple terminals. Terminals and receiver or sender tasks can have a corresponding relationship or be dynamically rotated. When new receivers and senders appear, the Terminal Cluster area will need to add a new terminal cluster, and therefore the required memory space will be obtained by modifying the size of the complete unallocated memory adjacent to the message body area.

[0077] In one implementation, step S102 in this embodiment includes the following steps:

[0078] Step S1021: Establish a number of terminal clusters in the terminal cluster area, and pair a terminal node avatar for each terminal in the terminal cluster area, wherein each terminal cluster includes a terminal identifier, a terminal quantity, an insertion subscript, and a terminal pair; wherein the terminal identifier is used to identify and search for the terminal cluster, the terminal quantity is the number of available terminals in the terminal cluster, the insertion subscript is used to locate the terminal selected for insertion, and the terminal pair is a paired terminal and a terminal node avatar.

[0079] Specifically, the terminal cluster area contains multiple terminal clusters, and each terminal cluster is responsible for communication between two types (not just two) of individuals. The establishment of the terminal cluster is intended to meet both one-to-one directional communication and many-to-many broadcast communication through a structure. For example, the read-write terminal cluster allows various individuals with read and write needs to join as senders, and also allows individuals capable of meeting read and write needs to join as receivers. Since the terminal is lock-free, any individual can join or leave at any time. Since the terminal itself is a one-way lock-free linked list, its cost is much higher than the MPMC Queue solution when encountering concurrency conflicts. Therefore, in order to increase parallelism and reduce conflicts, multiple terminals are used so that multiple tasks on the receiver can perform dequeue operations at the same time and reduce cache interference between each other.

[0080] Specifically, as attached Figure 9 As shown, each terminal cluster includes:

[0081] Terminal Identifier (optional): To facilitate searching and establishing correspondences, each terminal cluster is preceded by a terminal identifier. The terminal identifier is set by the individual who establishes the terminal cluster and indicates which two types of individuals the terminal cluster is responsible for communicating with. A new recipient can iterate through all terminal clusters and send a query message to see if there is an existing terminal cluster they can join. If no suitable terminal cluster exists, a new terminal cluster can be created by obtaining free memory in the message body area.

[0082] Terminal quantity: The number of available terminals in the terminal cluster. This number is variable, but has an upper limit. When the terminal cluster is initialized, the terminals are initialized according to the maximum upper limit, but the actual number of enabled terminals depends on the number of terminals in the terminal cluster.

[0083] Insertion index: When the sender wants to add a message to a terminal, it performs an atomic Fetch and Add (FAA) operation on the insertion index and obtains the actual index by modulo the number of terminals to locate the terminal to be inserted in the terminal array.

[0084] Terminal Pairs: For speed and concurrency safety considerations, terminals are implemented using a lock-free, one-way linked list. The algorithm requires at least one node in the linked list. Since the linked list may be empty in real life, each terminal is paired with a surrogate terminal node, which can be inserted when the linked list is empty to meet the algorithm's requirements.

[0085] Step S200: Acquire message data, and allocate memory for the message data in the message body area using a two-level memory segmentation strategy algorithm to obtain free memory;

[0086] Specifically, in this embodiment, the memory allocation uses the TLSF memory allocation algorithm to allocate the memory starting from the end address, so that the complete and unallocated free memory is close to the next area, that is, the terminal cluster area.

[0087] Specifically, TLSF (Two-Level Segregated Fit) is a memory allocation algorithm designed to manage memory fragmentation during dynamic memory allocation and release. The TLSF algorithm employs a two-level segregated fit strategy, dividing memory blocks into multiple levels based on size. Within each level, a segregated fit approach is used to manage memory blocks, thereby improving memory allocation efficiency. Physical memory is divided into memory blocks of varying sizes, each with a corresponding bitmap representing its free state. Memory blocks are divided into multiple levels, each with a corresponding linked list of memory blocks of the same size. When memory allocation is required, the TLSF algorithm first finds a memory block of the appropriate size, removes it from the linked list, and returns it to the user. When releasing memory, the TLSF algorithm inserts the block into a linked list of the corresponding size and merges it as needed to reduce memory fragmentation. By dividing memory into multiple levels and employing a segregated fit approach to manage memory blocks, the TLSF algorithm effectively reduces memory fragmentation. Specifically, the TLSF algorithm performs memory block merging operations during memory allocation and release to maximize the use of released memory blocks, thereby reducing memory fragmentation.

[0088] In one implementation, step S200 in this embodiment includes the following steps:

[0089] Step S201: In the dormant message body in the message body area, starting from the end address, memory is allocated for the message data using a two-level memory segmentation strategy algorithm to obtain the free memory, wherein the end address is the address of the dormant message body adjacent to the terminal cluster area.

[0090] Step S300: Obtain a message body allocated to the sender based on the message data and free memory;

[0091] Step S400: In the terminal cluster area, the message body allocated to the sender is conveyed to the receiver task that performs a dequeue operation.

[0092] In one implementation, step S400 in this embodiment includes the following steps:

[0093] Step S401: Acquire the number of receiver tasks and the number of terminals, where the number of terminals can be dynamically adjusted;

[0094] Specifically, the number of terminals can be adjusted dynamically. To increase the number of terminals, only the number of terminals needs to be modified, and no CAS operation is required. When each task on the receiving side detects that the number of terminals in the terminal cluster is greater than the number of terminals it maintains, it will update its own number of terminals to the latest value. Figure 12 The example shown is a reduction from six terminals to four terminals, which can be divided into two steps:

[0095] 1. An individual modifies the number of terminals in a terminal cluster. Subsequent senders will not be able to insert messages into the terminal to be closed because they use this value for modulo operations.

[0096] 2. When each receiver task detects that the number of terminals in the terminal cluster is less than the number of terminals it maintains, it will reduce the number of terminals it maintains by one when the terminal with the largest subscript is empty. The receiver task will keep rolling, gradually reducing the number of terminals it maintains in the process. There are message bodies and dormant message bodies in the message body area, which are distinguished by inserted linked lists: the message body (MC) in the terminal and the dormant message body (DMC) in the directory. In order to use the TLSF algorithm to manage DMC, in addition to a TLSF node used to link to the TLSF node in the directory, header data is also added. The content of DMC is as shown in the attached Figure 13 Shown, including:

[0097] Header data: The lowest bit in the first eight bytes, which is one bit, is used to mark the available space in the segment as allocated.

[0098] Guard bit: The second lowest bit in the first eight bytes, one bit, is used to mark whether the available space is the last available space in the shared memory, that is, the available space closest to the terminal cluster.

[0099] Size: The third to 63rd bits of the first eight bytes record the size of the DMC, which includes the header data, TLSF node, and available space.

[0100] The starting address of the previous physical memory segment: The second eight bytes are used to connect the DMCs before and after the TLSF algorithm releases the MC.

[0101] TLSF node:

[0102] Delete bit: The lowest bit in the first and second eight bytes, one bit, is used to mark whether the node is being removed from the linked list.

[0103] Previous / next node starting address: both are eight bytes, used to locate the previous and next TLSF nodes.

[0104] Available space: The actual available space for the sender and receiver.

[0105] MC is almost the same as DMC, except that the deletion bit is replaced by the substitute bit, which is used to distinguish between ordinary terminal nodes and substitute terminal nodes in the terminal cluster. Figure 14 shown.

[0106] Step S402: If the number of receiver tasks is greater than or equal to the number of terminals, each receiver task is assigned to a terminal. After the message body assigned to the sender is enqueued, the message body assigned to the sender is transferred to the receiver task, and the terminal corresponding to the receiver task performs a dequeue operation.

[0107] Step S403: If the number of the receiving tasks is less than the number of terminals, the terminal is scrolled to select and perform a dequeue operation on the message body assigned to the sender.

[0108] In this embodiment, although the sender can only insert the message body into the terminal in the terminal cluster in a rolling manner, the relationship between the receiver and the terminal can be that each receiver task scrolls to select a terminal to perform a dequeue operation, or each task corresponds to one terminal.

[0109] Specifically, as attached Figure 11 As shown in the figure, when the number of receiving tasks is equal to or greater than the number of terminals, a one-to-one correspondence between tasks and terminals can be established. This has the advantage of reducing some latency: after the sender enqueues a task, if there is a corresponding receiving task, the dequeue operation can be performed immediately. However, if the receiving terminal is busy, the task must wait.

[0110] Specifically, when the number of tasks on the receiving side is less than the number of terminals, or when the number of terminals and tasks needs to be dynamically adjusted during communication, the terminals must be selected in a rolling manner, as shown in the attached figure. Figure 10 As shown in the figure, each receiving task maintains an incrementing index and the number of terminals, selecting them in the same manner as the sender. While this increases the overhead of selecting terminals, it ensures that messages in all terminals can be discovered and processed by other tasks even if a single task is blocked.

[0111] In one implementation, the method of this embodiment further includes the following steps:

[0112] Step S500: Divide the memory space into a directory area, and create a memory two-level partitioning strategy node in the directory area, wherein the memory two-level partitioning strategy node is a data structure based on a bidirectional lock-free linked list.

[0113] Specifically, the catalog area is responsible for storing the current status of the entire memory area and the data and lock-free linked lists used by the TLSF algorithm to manage the message body area.

[0114] Step S600: storing the state of the memory area through the directory area and managing the memory linked list of the memory two-level segmentation strategy algorithm; wherein the state of the memory area includes uninitialized, initializing and initialized.

[0115] Specifically, the directory is placed at the starting address so that different individuals can understand whether the memory is initialized and perform allocation and release operations. The current state of the XSIM shared memory is located at the starting address. It acts as a lock for the entire memory segment, allowing only one individual to perform initialization operations. The different memory states are as follows:

[0116] Uninitialized: The individual reading this state will be responsible for initializing the various areas in the memory and will immediately execute the CAS instruction to change the state to Initializing. If the CAS instruction fails, that is, another individual has already started initialization, the memory state will be read again.

[0117] Initializing: The instance will poll the memory status until the status changes to Initialized.

[0118] Initialized: Individuals can continue to access and use initialized areas in shared memory.

[0119] Although memory initialization is not lock-free, the sender and receiver typically perform initialization before the actual task begins, and communication after the actual task begins only involves lock-free operations. In addition to the state, the directory also includes the quick lookup tables and linked lists required by the TLSF algorithm to manage the message body.

[0120] In one implementation, step S500 in this embodiment includes the following steps:

[0121] Step S501: storing the state of the memory area at the starting address of the directory area;

[0122] Step S502: establishing a memory linked list of a two-level memory segmentation strategy algorithm, and locating a bitmap of the message data according to the memory linked list of the two-level memory segmentation strategy algorithm;

[0123] Step S503: Obtain the required message space according to the bitmap of the message data;

[0124] Step S504: Obtain a dormant message body holding the required message space according to the required message space;

[0125] Step S505: arbitrarily select a memory two-level partitioning strategy node from the dormant message body holding the required message space to perform dequeue and enqueue operations, and use the memory two-level partitioning strategy node as the head of the memory linked list of the memory two-level partitioning strategy algorithm.

[0126] Specifically, as attached Figure 15 As shown, the TLSF algorithm is a bitmap-based memory allocation algorithm that divides physical memory into multiple memory blocks of varying sizes. Each memory block has a corresponding bitmap to represent the idle state of the block. In this embodiment, the functions and structures of the first-order bitmap and the second-order bitmap are basically consistent with those in the TLSF algorithm. However, since the bitmap is not locked during communication, it is possible that the bitmap is inconsistent with the actual situation. Therefore, the TLSF algorithm will re-execute when it fails, and after a certain number of repetitions, it will determine that the maximum free space cannot meet the demand and will perform a slicing operation. Although this will lose the constant complexity time upper limit of allocation and release in the TLSF algorithm, under normal circumstances, the actual impact is limited due to the small number of modifications to the bitmap.

[0127] After the TLSF algorithm locates the appropriate space interval through the bitmap, it performs a DMC dequeue operation by accessing the TLSF node to obtain the DMC holding the required message space. Because the TLSF node is a data structure based on a bidirectional lock-free linked list, and the DMC also contains TLSF nodes, dequeue and enqueue operations can be performed on any TLSF node. However, the TLSF node in the directory is created during initialization and is considered the head of the DMC linked list.

[0128] In one implementation, XSIM, without relying on other media, achieved a 20% performance improvement over MPMCQueue in iozone simulations when transferring large files, such as 128KB. Furthermore, thanks to variable message lengths, different message types can be matched with different segmentation strategies to achieve the most efficient data transmission.

[0129] In summary, XSIM first introduces the TLSF memory management algorithm, which allows messages of variable length and avoids the recurring performance overhead of slicing and reassembly. Since the message space is variable, zero copying becomes possible. By setting up terminals, different senders can use the various DMCs in the same message body area, greatly reducing memory waste in multi-party communications. Secondly, by using lock-free data structures, both communicating parties can use the TLSF algorithm to allocate and release memory at the same time. The sender does not need to wait for the receiver, and the receiver will automatically release the memory after the message ends. By placing the message queue in shared memory in the form of a terminal, the dependence on other media is eliminated and adaptability is increased. By setting up terminal clusters, the concurrency conflicts caused by the use of a single message queue when multiple receiver tasks perform dequeue operations can be reduced to a certain extent, thereby increasing parallelism.

[0130] Exemplary devices

[0131] like Figure 16 As shown in , this embodiment further provides a shared memory communication device, the device comprising:

[0132] The area division module 10 is used to divide the memory space into a message body area and a terminal cluster area;

[0133] A memory allocation module 20 is configured to obtain message data and allocate memory for the message data in the message body area using a two-level memory segmentation strategy algorithm to obtain free memory;

[0134] A message body acquisition module 30 is used to obtain a message body allocated to the sender based on the message data and free memory;

[0135] The memory communication module 40 is used to transmit the message body assigned to the sender to the receiver task in the terminal cluster area.

[0136] In one implementation, the region division module 10 includes:

[0137] A message body area division unit, configured to divide the memory space into a message body area, wherein the message body area includes a message body and a dormant message body; wherein the message body is used to manage allocated memory, and the dormant message body is used to manage free memory;

[0138] a terminal cluster area division unit, configured to divide the memory space into terminal cluster areas and establish a plurality of terminal clusters in the terminal cluster areas, wherein the terminal clusters are first-in-first-out queues without lock links;

[0139] The directory area division unit is used to divide the directory area on the memory space and create a memory two-level division strategy node on the directory area, wherein the memory two-level division strategy node is a data structure based on a bidirectional lock-free linked list.

[0140] In one implementation, the terminal cluster area division unit includes:

[0141] The terminal cluster area division subunit is used to establish several terminal clusters in the terminal cluster area and pair a terminal node substitute for each terminal in the terminal cluster area, wherein each terminal cluster includes a terminal identifier, a terminal quantity, an insertion subscript and a terminal pair; wherein the terminal identifier is used to identify and search for the terminal cluster, the terminal quantity is the number of available terminals in the terminal cluster, the insertion subscript is used to locate the terminal selected for insertion, and the terminal pair is a paired terminal and a terminal node substitute.

[0142] In one implementation, the region division module 20 includes:

[0143] The free memory acquisition unit is used to allocate memory for the message data in the dormant message body of the message body area, starting from the end address, through a two-level memory segmentation strategy algorithm to obtain the free memory, wherein the end address is the address of the dormant message body adjacent to the terminal cluster area.

[0144] In one implementation, the region division module 40 includes:

[0145] A terminal quantity adjustment unit, configured to obtain the number of receiver tasks and the number of terminals, wherein the number of terminals can be dynamically adjusted;

[0146] A first dequeue operation unit is configured to, if the number of receiver tasks is greater than or equal to the number of terminals, assign each receiver task to a terminal, and when the message body assigned to the sender is enqueued, transfer the message body assigned to the sender to the receiver task, and perform a dequeue operation on the terminal corresponding to the receiver task;

[0147] The second dequeue operation unit is configured to scroll and select a terminal to perform a dequeue operation on the message body allocated to the sender if the number of the receiver tasks is less than the number of terminals.

[0148] In one implementation, the apparatus further includes:

[0149] The directory area partitioning module 50 partitions the directory area in the memory space and creates a memory two-level partitioning strategy node in the directory area, wherein the memory two-level partitioning strategy node is a data structure based on a bidirectional lock-free linked list.

[0150] The directory area function module 60 is used to store the state of the memory area through the directory area and manage the memory linked list of the memory two-level segmentation strategy algorithm; wherein the state of the memory area includes uninitialized, initializing and initialized.

[0151] In one implementation, the directory area function module 60 includes:

[0152] a state management unit, configured to store the state of the memory area at a starting address of the directory area;

[0153] A memory linked list unit, used to establish a memory linked list of a memory two-level segmentation strategy algorithm, and locate a bitmap of message data according to the memory linked list of the memory two-level segmentation strategy algorithm;

[0154] a message space positioning unit, configured to obtain a required message space according to the bitmap of the message data;

[0155] a dormant message body partitioning unit, configured to obtain a dormant message body holding the required message space according to the required message space;

[0156] The linked list maintenance unit is used to arbitrarily select a memory two-level segmentation strategy node in the dormant message body holding the required message space to perform dequeue and enqueue operations, and use the memory two-level segmentation strategy node as the linked list head of the memory linked list of the memory two-level segmentation strategy algorithm.

[0157] Based on the above embodiment, the present invention also provides an intelligent terminal, whose principle block diagram can be shown as follows: Figure 17 As shown. The intelligent terminal includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected via a system bus. The processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a shared memory communication method is implemented. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor of the intelligent terminal is pre-set inside the intelligent terminal to detect the operating temperature of the internal device.

[0158] Those skilled in the art will understand that Figure 17 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present invention and does not constitute a limitation on the smart terminal to which the solution of the present invention is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0159] In one embodiment, a smart terminal is provided. The smart terminal includes a memory, a processor, and a shared memory communication program stored in the memory and executable on the processor. When the processor executes the shared memory communication program, the following operating instructions are implemented:

[0160] Divide the memory space into a message body area and a terminal cluster area;

[0161] Acquire message data, and allocate memory for the message data in the message body area using a two-level memory segmentation strategy algorithm to obtain free memory;

[0162] Obtaining a message body allocated to the sender based on the message data and free memory;

[0163] In the terminal cluster area, the message body assigned to the sender is communicated to the receiver task that performs the dequeue operation.

[0164] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0165] In summary, the present invention discloses a shared memory communication method, which includes: dividing the memory space into a message body area, a terminal cluster area, and a directory area; obtaining message data, and allocating memory for the message data in the message body area through a two-level memory segmentation strategy algorithm to obtain free memory; obtaining the message body allocated to the sender based on the message data and the free memory; and delivering the message body allocated to the sender to the receiving task in the terminal cluster area. The present invention eliminates dependence on other media and increases adaptability by placing the message queue in the form of a terminal in the shared memory. Combining a lock-free data structure with a two-level memory segmentation strategy algorithm avoids the regular performance overhead of slicing and reorganization, thereby improving the efficiency of memory communication.

[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A shared memory communication method, characterized in that: The method comprises: Divide the memory space into a message body area and a terminal cluster area; Acquire message data, and allocate memory for the message data in the message body area using a two-level memory segmentation strategy algorithm to obtain free memory; Obtaining a message body allocated to the sender based on the message data and free memory; In the terminal cluster area, the message body assigned to the sender is communicated to the receiver task that performs the dequeue operation.

2. The shared memory communication method according to claim 1, wherein: The dividing of the memory space into a message body area and a terminal cluster area includes: Dividing the memory space into a message body area, wherein the message body area includes a message body and a dormant message body; wherein the message body is used to manage allocated memory, and the dormant message body is used to manage free memory; A terminal cluster area is divided in the memory space, and a plurality of terminal clusters are established in the terminal cluster area, wherein the terminal cluster is a first-in-first-out queue without a lock link list.

3. The shared memory communication method according to claim 1, wherein: The dividing the terminal cluster area in the memory space and establishing a plurality of terminal clusters in the terminal cluster area includes: Several terminal clusters are established in the terminal cluster area, and a terminal node substitute is paired with each terminal in the terminal cluster area, wherein each terminal cluster includes a terminal identifier, a terminal quantity, an insertion subscript and a terminal pair; wherein the terminal identifier is used to identify and search the terminal cluster, the terminal quantity is the number of available terminals in the terminal cluster, the insertion subscript is used to locate the terminal selected for insertion, and the terminal pair is a paired terminal and a terminal node substitute.

4. The shared memory communication method according to claim 1, wherein: In the message body area, allocating memory for the message data by a two-level memory segmentation strategy algorithm to obtain free memory includes: In the dormant message body in the message body area, starting from the end address, memory is allocated to the message data through a two-level memory segmentation strategy algorithm to obtain free memory, wherein the end address is the address of the dormant message body adjacent to the terminal cluster area.

5. The shared memory communication method according to claim 1, wherein: The task of delivering the message body assigned to the sender to the receiver in the terminal cluster area includes: Obtaining the number of receiver tasks and the number of terminals, wherein the number of terminals can be dynamically adjusted; If the number of receiver tasks is greater than or equal to the number of terminals, each receiver task is assigned to a terminal. After the message body assigned to the sender is enqueued, the message body assigned to the sender is passed to the receiver task, and the terminal corresponding to the receiver task performs the dequeuing operation. If the number of the receiver tasks is less than the number of terminals, the terminal is scrolled and selected to perform a dequeue operation on the message body assigned to the sender.

6. The shared memory communication method according to claim 1, wherein: The method further comprises: Dividing the memory space into a directory area, and creating a memory two-level partitioning strategy node in the directory area, wherein the memory two-level partitioning strategy node is a data structure based on a bidirectional lock-free linked list; The state of the memory area is stored in the directory area, and the memory linked list of the memory two-level segmentation strategy algorithm is managed; wherein the state of the memory area includes uninitialized, initializing and initialized.

7. The shared memory communication method according to claim 6, wherein: The memory linked list of the two-level memory segmentation strategy algorithm is managed by storing the state of the memory area in the directory area, and includes: Storing the state of the memory area at the starting address of the directory area; Establishing a memory linked list of a two-level memory segmentation strategy algorithm, and locating a bitmap of message data according to the memory linked list of the two-level memory segmentation strategy algorithm; Obtaining required message space according to the bitmap of the message data; According to the required message space, a dormant message body holding the required message space is obtained; In the dormant message body holding the required message space, a memory two-level partitioning strategy node is arbitrarily selected for dequeuing and enqueuing operations, and the memory two-level partitioning strategy node is used as the head of the memory linked list of the memory two-level partitioning strategy algorithm.

8. A shared memory communication device, characterized in that: The device comprises: A region partitioning module is used to divide the memory space into a message body region and a terminal cluster region; A memory allocation module is used to obtain message data and allocate memory for the message data in the message body area by using a two-level memory segmentation strategy algorithm to obtain free memory; A message body acquisition module, configured to obtain a message body allocated to the sender based on the message data and free memory; The memory communication module is used to transmit the message body assigned to the sender to the receiver task in the terminal cluster area.

9. An intelligent terminal, characterized in that: The intelligent terminal includes a memory, a processor, and a shared memory communication program stored in the memory and executable on the processor. When the processor executes the shared memory communication program, the steps of the shared memory communication method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a shared memory communication program, and when the shared memory communication program is executed by the processor, the steps of the shared memory communication method according to any one of claims 1 to 7 are implemented.