Coroutine-based memory management method and device
By adopting a coroutine-based memory management method in high concurrency scenarios, the memory allocation area and cache are shared between coroutines, the problems of lock conflicts and high memory usage are solved, and performance improvement and memory optimization are achieved.
Patent Information
- Application Number
- CN202310942821.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Existing memory allocators suffer performance losses due to lock conflicts in high concurrency scenarios, and the independence of thread-local caches leads to an increase in overall memory footprint.
The memory management method based on coroutines is adopted. By starting the coroutine scheduling thread, the memory allocation area is configured for each coroutine scheduling thread, and the coroutine and memory allocation area are bound according to the preset load balancing algorithm to realize memory sharing and cache reuse between coroutines.
It effectively reduces lock conflicts caused by memory management, improves the system's concurrency performance, and reduces the overall memory usage by sharing thread local cache areas.
Smart Images

Figure CN119166320B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, specifically to the field of memory management technology, and in particular to a memory management method and device based on coroutines. Background Art
[0002] In the prior art, in order to support concurrency, most memory allocators adopt the so-called "arena" mechanism to share the load of concurrent memory requests from multiple threads. The memory allocator reduces lock conflicts and solves concurrent performance issues by "binding" different threads to the same arena and opening multiple arenas in one process. In addition, most memory allocators also provide a thread-local cache mechanism.
[0003] However, when the number of application threads reaches a certain scale, or when memory allocation / release (i.e., malloc / free library function calls) in concurrent tasks are too frequent, the performance loss caused by this lock conflict is still not negligible. In addition, in order to optimize performance, most memory allocators have introduced thread local cache functions; when there are a large number of threads in the system, and each of them has an independent local cache, the overall memory usage rate increases significantly. Summary of the invention
[0004] Embodiments of the present application provide a memory management method, apparatus, device, and storage medium based on coroutines.
[0005] According to the first aspect, an embodiment of the present application provides a coroutine-based memory management method, the method comprising: starting at least one coroutine scheduling thread; configuring at least one memory allocation area for each coroutine scheduling thread; starting a coroutine, and when submitting the coroutine to a coroutine scheduling thread among at least one coroutine scheduling threads, binding the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0006] According to the second aspect, an embodiment of the present application provides a coroutine-based memory management device, which includes: a startup module, configured to start at least one coroutine scheduling thread; a configuration module, configured to configure at least one memory allocation area for each coroutine scheduling thread; a binding module, configured to start the coroutine and, when submitting the coroutine to one of the at least one coroutine scheduling threads, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0007] According to the third aspect, an embodiment of the present application provides an electronic device, which includes one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement a coroutine-based memory management method as in any embodiment of the first aspect.
[0008] According to a fourth aspect, an embodiment of the present application provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements a coroutine-based memory management method as in any embodiment of the first aspect.
[0009] The present application starts at least one coroutine scheduling thread; configures at least one memory allocation area for each coroutine scheduling thread; starts the coroutine, and when submitting the coroutine to one of the at least one coroutine scheduling threads, binds the coroutine to a memory allocation area corresponding to the coroutine scheduling thread submitted by the coroutine according to a preset load balancing algorithm, making full use of the characteristics of coroutine multi-tasking scheduling and time-sharing multiplexing of coroutine scheduling thread time slices. When coroutines belonging to the same coroutine scheduling thread share a memory allocation area (arena), there is no need to lock when applying / releasing memory, thereby minimizing the possibility of lock conflicts caused by memory management. At the same time, multiple coroutines share a thread local cache area, namely tcache. The memory released to tcache by the previous coroutine can be reused if memory is to be allocated when scheduling the next coroutine. Compared with the existing solution in which each thread has its own independent thread local cache area, this tcache cannot be shared with other threads, effectively reducing the overall memory usage of the system and improving the concurrency performance of the system.
[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0012] Figure 2 is a flowchart of an embodiment of a coroutine-based memory management method according to the present application;
[0013] Figure 3 is an architectural diagram of an embodiment of a coroutine-based memory management method according to the present application;
[0014] Figure 4a is a flowchart of an embodiment of a coroutine-based memory management method according to the present application;
[0015] Figure 4b It is a flowchart of an application scenario of the memory management method based on coroutines according to the present application;
[0016] Figure 4c is a flowchart of another application scenario of the coroutine-based memory management method according to the present application;
[0017] Figure 5 is a flowchart of an embodiment of a coroutine-based memory management device according to the present application;
[0018] Figure 6 It is a structural diagram of a computer system suitable for implementing a server of an embodiment of the present application. DETAILED DESCRIPTION
[0019] The following is a description of exemplary embodiments of the present application in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.
[0020] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0021] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of the coroutine-based memory management method of the present application can be applied.
[0022] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0023] The terminal devices 101, 102, 103 interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications, such as storage applications, communication applications, etc., may be installed on the terminal devices 101, 102, 103.
[0024] The terminal devices 101, 102, and 103 may be hardware. When the terminal devices 101, 102, and 103 are hardware, they may be various electronic devices with display screens, including but not limited to mobile phones and laptop computers.
[0025] The server 105 may include a memory allocator that provides the following functions, for example, starting at least one coroutine scheduling thread; configuring at least one memory allocation area for each coroutine scheduling thread; starting a coroutine, and when submitting the coroutine to one of the at least one coroutine scheduling threads, binding the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0026] It should be noted that when the server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server.
[0027] It should be noted that the memory management method based on coroutines provided in the embodiments of the present disclosure can be executed by the memory allocator in the server 105, or by the memory allocator in the terminal devices 101, 102, 103. Accordingly, the various parts (such as various units, sub-units, modules, sub-modules) included in the memory management device based on coroutines can all be set in the server 105, or all be set in the terminal devices 101, 102, 103.
[0028] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0029] Figure 2 The process 200 of an embodiment of a memory management method based on coroutines that can be applied to the present application is shown. In this embodiment, the memory management method based on coroutines includes the following steps:
[0030] Step 201, start at least one coroutine scheduling thread.
[0031] In this embodiment, current operating systems all provide system calls for memory application / release to meet the memory usage requirements of applications. Taking Linux as an example, it provides sbrk / brk system calls for allocating memory in the heap area and mmap system calls for allocating memory in the mapping area (generally sbrk / brk is used to allocate small-sized memory, while mmap is often used to allocate large blocks of memory; in addition to Linux, other operating systems also have similar system calls).
[0032] However, if system calls are used every time memory is allocated / released, the performance of the application will be greatly affected; therefore, various operating system distributions provide user-mode memory allocators to improve performance, such as the default allocator ptmalloc of GNULinux, jemalloc of FreeBSD, and tcmalloc, an open source memory allocator contributed by Google.
[0033] In order to support concurrency, most memory allocators adopt the so-called "arena" mechanism to share the load of concurrent memory requests from multiple threads. An "arena" is a memory allocation area. Although existing memory allocators all use a multi-memory allocation area (arena) mechanism to reduce concurrency conflicts; however, when the number of threads in an application reaches a certain scale, or memory allocation / release (i.e., malloc / free library function calls) in concurrent tasks is too frequent, the performance loss caused by this lock conflict is still very huge.
[0034] To overcome the above problems, the execution entity (such as Figure 1 The memory allocator in the server 105 or the terminal device 101, 102, 103 shown in FIG. 1 first starts at least one, for example, 5, 10, etc., coroutine scheduling thread
[0035] Among them, the coroutine scheduling thread is used to schedule at least one coroutine. The operation of the coroutine is dependent on the coroutine scheduling thread to which it belongs. After the coroutine gives up through yield, the coroutine scheduling thread can schedule other coroutines to run, that is, the coroutine scheduling thread's time slices are reused between coroutines. Therefore, although the coroutines on the same coroutine scheduling thread are nominally executed "in parallel", its time-sharing multiplexing mechanism ensures that any two of these coroutines will not actually run at the same time (that is, the running time slices will not overlap). Based on this premise, as long as the coroutine's access to the cache is guaranteed to be "atomic" (that is, the yield operation will not be performed when the cache operation is not completed), shared cache between coroutines can be achieved.
[0036] Step 202: configure at least one memory allocation area for each coroutine scheduling thread.
[0037] In this embodiment, the execution entity may configure one or more memory allocation areas for each coroutine scheduling thread.
[0038] Each memory allocation area in the at least one memory allocation area corresponds to a thread local cache area, and the thread local cache areas corresponding to the memory allocation areas are different.
[0039] Here, the thread local cache area, namely tcache, belongs to the memory allocation area, namely arena, and they have a one-to-one correspondence. In other words, the cache in tcache is part of the memory in arena. Generally speaking, the memory scale allocated by the program is relatively small, and tcache can be considered as a small-scale cache. In most cases, memory allocation requests can be satisfied by tcache. It is necessary to apply to arena only when applying for large-sized memory space; in addition, when there is no memory space of suitable size in tcache, a large block is first applied from arena, and then it is cut into small blocks of memory for use. Therefore, tcache is an acceleration mechanism for small-scale memory allocation and reuse in an arena.
[0040] Step 203, start the coroutine, and when submitting the coroutine to a coroutine scheduling thread in at least one coroutine scheduling thread, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0041] In this embodiment, the execution entity can start one or more coroutines. For each coroutine, when submitting the coroutine to a target coroutine scheduling thread in at least one coroutine scheduling thread, the coroutine is bound to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0042] When the number of coroutines bound to the memory allocation area is greater than 1, a group of coroutines, namely the "coroutine group", is naturally generated.
[0043] After the above binding relationship is determined, memory allocation and memory merging and return can be performed according to the coroutine's request for memory application and memory release.
[0044] In addition, it should be pointed out that in order to avoid the situation where the target coroutine is executed from a coroutine scheduling thread other than the coroutine scheduling thread to which it belongs, and the target coroutine accesses the memory allocation area bound to the other coroutine scheduling thread, resulting in concurrent access, the execution subject can configure to prohibit cross-thread scheduling of the coroutine, or implement optimistic locking, for example, version number mechanism, CAS algorithm, etc., to solve the problem of cross-thread access to the memory allocation area.
[0045] Furthermore, in some optional methods, the above steps 201, 202, and 203 of the present application can be executed via the callback interface provided by the hook through the library functions in the specified coroutine library, such as malloc, free, calloc, realloc, memalign, valloc, etc. This method can realize the "transparent" replacement of the memory allocator, that is, the replacement of the memory allocator without the application's perception, improve the memory management performance, and at the same time, on the basis of the high-performance concurrent scheduling capability provided by the coroutine library, further improve the overall concurrent throughput capability.
[0046] For a coroutine library without a callback interface, such as libco, libgo, etc., an override of the library functions related to the coroutine-based memory management operation can be added, and a coroutine-based memory manager can be implemented based on steps 201, 202, and 203.
[0047] In some optional embodiments, the method further includes: in response to obtaining a first request for memory from a first coroutine, searching for a target memory space corresponding to the first request in a thread local cache area corresponding to a memory allocation area bound to a coroutine group to which the first coroutine is located; in response to a successful search, allocating the target memory space; in response to a failed search, searching for and allocating the target memory space corresponding to the first request in a memory allocation area bound to the coroutine group to which the first coroutine is located.
[0048] In this implementation, when the application calls the library function free to release memory space, the released memory is not directly released back to the memory allocation area, but is placed in the thread local cache area corresponding to the memory allocation area. In response to obtaining the first request for memory from the first coroutine, the execution subject can first search for the target memory space corresponding to the first request in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the first coroutine is located, that is, the memory space suitable for the size of this request. If the search is successful, the target memory space is allocated in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the first coroutine is located; if the search fails, the target memory space corresponding to the first request is further searched in the memory allocation area bound to the coroutine group where the first coroutine is located and allocated.
[0049] This implementation method achieves memory allocation for the request for memory from the coroutine by, in response to obtaining a first request for memory from the first coroutine, searching for the target memory space corresponding to the first request in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the first coroutine is located; in response to a successful search, allocating the target memory space; in response to a failed search, searching for the target memory space corresponding to the first request in the memory allocation area bound to the coroutine group where the first coroutine is located and allocating it.
[0050] In some optional embodiments, the method further includes: in response to obtaining a second request from the second coroutine to release memory, merging the free memory space in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the second coroutine is located to obtain a first memory space; in response to determining that the first memory space meets a first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located, returning the first memory space to the memory allocation area and merging the free memory space of the memory allocation area to obtain a second memory space; in response to determining that the second memory space meets a second preset condition for returning it to the operating system, returning the second memory space to the operating system.
[0051] In this implementation, in response to obtaining the second request from the second coroutine to release memory, the execution subject may first merge the free memory space in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the second coroutine is located, obtain the first memory space, and determine whether the first memory space meets the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located; in response to determining that the first memory space meets the first preset condition for returning the memory allocation area, return the first memory space to the memory allocation area, and merge the free memory space of the memory allocation area to obtain the second memory space; in response to determining that the second memory space meets the second preset condition for returning to the operating system, return the second memory space to the operating system, and in response to not meeting the condition, terminate the return operation.
[0052] Here, the first preset condition and the second preset condition may be the same or different. The first preset condition and the second preset condition may be set based on experience and actual needs. For example, the size of continuous memory with continuous addresses and aligned start addresses is greater than or equal to a preset size threshold, the residence time in the storage area (cache area or memory allocation area) is greater than or equal to a preset time threshold, etc. This application does not limit this.
[0053] The implementation method obtains a first memory space by, in response to obtaining a second request from the second coroutine to release memory, merging the free memory space in the thread local cache area corresponding to the memory allocation area bound to the coroutine group to which the second coroutine is located; in response to determining that the first memory space meets a first preset condition for returning the memory allocation area bound to the coroutine group to which the second coroutine is located, returning the first memory space to the memory allocation area and merging the free memory space in the memory allocation area to obtain a second memory space; in response to determining that the second memory space meets a second preset condition for returning it to the operating system, returning the second memory space to the operating system, thereby realizing the return of memory space for the coroutine's request to release memory.
[0054] In some optional embodiments, the method further includes: in response to determining that the first memory space does not meet the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located or the second memory space does not meet the second preset condition for returning it to the operating system, ending the operation.
[0055] In this implementation, the execution entity can determine whether the first memory space meets the first preset condition for returning the memory allocation area bound to the coroutine group to which the second coroutine is located, and terminate the operation in response to the first memory space not meeting the first preset condition or the second memory space not meeting the second preset condition for returning to the operating system.
[0056] This implementation method returns the memory space requested by the coroutine to release memory by terminating the operation in response to determining that the first memory space does not meet the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located, or the second memory space does not meet the second preset condition for returning it to the operating system.
[0057] In some optional ways, the configuration prohibits cross-thread scheduling of coroutines.
[0058] In this implementation, in order to avoid the situation where the target coroutine is executed from a coroutine scheduling thread other than the coroutine scheduling thread to which it belongs, and the target coroutine accesses the memory allocation area bound to the other coroutine scheduling thread, resulting in concurrent access, the execution subject can be configured to prohibit cross-thread scheduling of coroutines.
[0059] This implementation method can effectively avoid concurrent access and ensure memory allocation / recycling performance by configuring to prohibit cross-thread scheduling of coroutines.
[0060] Continue to see Figure 3 , Figure 3 It is an architecture diagram of an application scenario of the coroutine-based memory management method according to this embodiment.
[0061] exist Figure 3In the application scenario, the execution subject can start at least one coroutine scheduling thread; configure at least one memory allocation area for each coroutine scheduling thread, such as arena0, arena1, arena2...arenaN, wherein each memory allocation area in the at least one memory allocation area corresponds to a thread local cache area; start the coroutine, and when submitting the coroutine to the coroutine scheduling thread in at least one coroutine scheduling thread, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread submitted by the coroutine according to a preset load balancing algorithm, wherein the coroutines bound to the same memory allocation area constitute a coroutine group, for example, coroutine group 0 is bound to arena0, and tcache0 corresponds to (belongs to) arena0; coroutine group 1 is bound to arena1, and tcache1 corresponds to (belongs to) arena1...coroutine group N is bound to arenaN, and tcacheN corresponds to (belongs to) arenaN.
[0062] Figure 4a A process 400 of another embodiment of a memory management method based on coroutines that can be applied to the present application is shown. In this embodiment, the memory management method based on coroutines includes the following steps:
[0063] Step 401, start at least one coroutine scheduling thread.
[0064] In this embodiment, the implementation details and technical effects of step 401 can refer to the description of step 201 and will not be repeated here.
[0065] Step 402: configure at least one memory allocation area for each coroutine scheduling thread.
[0066] In this embodiment, the implementation details and technical effects of step 402 can refer to the description of step 202 and will not be repeated here.
[0067] Step 403, start the coroutine, and when submitting the coroutine to a coroutine scheduling thread in at least one coroutine scheduling thread, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0068] In this embodiment, the implementation details and technical effects of step 403 may refer to the description of step 203 and will not be repeated here.
[0069] Step 404, under the condition that cross-thread scheduling of coroutines is allowed in the configuration, before performing a memory allocation operation based on a coroutine memory request, or before performing a memory merge and return operation based on a coroutine memory release request, mark the atomic variable.
[0070] In this embodiment, in order to avoid the target coroutine being executed from other coroutine scheduling threads other than the coroutine scheduling thread to which it belongs, and the target coroutine accesses the memory allocation area bound by the other coroutine scheduling thread, resulting in concurrent access, under the condition that the cross-thread scheduling of coroutines is allowed, the execution subject can configure an atomic variable based on the CAS (Compare And Swap) operation. Before performing a memory allocation operation based on a coroutine application memory request, or before performing a memory merge and return operation based on a coroutine release memory request, the atomic variable is marked.
[0071] Among them, CAS is a hardware synchronization primitive provided by a processor (CPU) that supports concurrency. The CAS operation contains three operands, namely the memory location (V), the expected original value (A), and the new value (B), written as CAS (V, A, B). If the value of the memory location matches the expected original value, the processor will automatically update the location value to the new value. Otherwise, the processor does not do anything.
[0072] Step 405, in response to determining that the memory allocation operation is completed or the memory merge and return operation is completed, the atomic variable is cleared.
[0073] In this embodiment, the execution subject clears the atomic variable in response to determining that the memory allocation operation is completed or the memory merge and return operation is completed.
[0074] Specifically, Figure 4b As shown, the execution subject can configure atomic variables based on the CAS operation, and before executing the memory allocation operation, mark the atomic variables, that is, set the atomic variables. The memory allocation operation may include: searching for the memory space corresponding to the first request in the thread local cache area, that is, the memory space suitable for the size of this request. If the search is successful, the target memory space corresponding to the first request is directly allocated; if the search fails, the target memory space corresponding to the first request is further searched in the memory allocation area. If the search is successful, the target memory space corresponding to the first request is directly allocated.
[0075] In response to the memory allocation operation completing, the atomic variable is cleared to zero.
[0076] Another example Figure 4cAs shown, the execution subject can configure atomic variables based on CAS operations, mark atomic variables before performing memory merging and returning operations, and the memory merging and returning operations may include: merging the free memory space in the thread local cache area to obtain a first memory space, and judging whether the first memory space meets the first preset condition for returning the memory allocation area; in response to the first memory space not meeting the first preset condition, the merging and returning operation is terminated. In response to the first memory space meeting the first preset condition for returning the memory allocation area, the first memory space is returned to the memory allocation area, and the free memory space of the memory allocation area is merged to obtain a second memory space; in response to the second memory space meeting the second preset condition for returning the operating system, the second memory space is returned to the operating system; in response to not meeting, the operation is terminated.
[0077] In response to determining that the memory merge and return operation is complete, the atomic variable is cleared.
[0078] Furthermore, if marking an atomic variable fails and the number of failures is greater than or equal to the preset threshold, the busy loop is no longer performed, but the coroutine is suspended to delay, swapped out once and swapped in again, and then the CAS operation is retried. This is "yield & resume", that is, swapping out first, and then continuing the subsequent operation when it is swapped in again.
[0079] As can be seen from Figure 4, Figure 2 Compared with the corresponding embodiments, the process 400 of the coroutine-based memory management method in this embodiment reflects that under the condition that the configuration allows cross-thread scheduling of coroutines, before executing the memory allocation operation based on the coroutine application memory request, or before executing the memory merging and returning operation based on the coroutine release memory request, the atomic variable is marked; in response to determining that the memory allocation operation is completed or the memory merging and returning operation is completed, the atomic variable is cleared (that is, load balancing is supported, such as when a large number of coroutines on certain coroutine scheduling threads end and cause uneven load, the coroutines on other coroutine scheduling threads with heavier loads can be scheduled to scheduling threads with relatively lighter loads) while avoiding concurrent access.
[0080] Further references Figure 5 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a memory management device based on a coroutine, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0081] like Figure 5 As shown, the coroutine-based memory management device 500 of this embodiment includes: a startup module 501, a configuration module 502 and a binding module 503.
[0082] Among them, the starting module 501 can be configured to start at least one coroutine scheduling thread.
[0083] The configuration module 502 may be configured to configure at least one memory allocation area for each coroutine scheduling thread.
[0084] The binding module 503 can be configured to start the coroutine and, when submitting the coroutine to a coroutine scheduling thread in at least one coroutine scheduling thread, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm.
[0085] In some optional embodiments of this embodiment, the device also includes a marking module, which is configured to mark the atomic variable before performing a memory allocation operation based on a coroutine memory request, or before performing a memory merge and return operation based on a coroutine memory release request, under the condition that the configuration allows cross-thread scheduling of coroutines; in response to determining that the memory allocation is completed or the memory merge and return are completed, the atomic variable is cleared.
[0086] In some optional embodiments of this embodiment, the device also includes an allocation module, which is configured to, in response to obtaining a first request for memory from a first coroutine, search for a target memory space corresponding to the first request in a thread local cache area corresponding to a memory allocation area bound to a coroutine group where the first coroutine is located; in response to a successful search, allocate the target memory space; in response to a failed search, search for and allocate the target memory space corresponding to the first request in a memory allocation area bound to the coroutine group where the first coroutine is located.
[0087] In some optional embodiments of this embodiment, the device also includes a release module, which is configured to, in response to obtaining a second request from the second coroutine to release memory, merge the free memory space in the thread local cache area corresponding to the memory allocation area bound to the coroutine group where the second coroutine is located, to obtain a first memory space; in response to determining that the first memory space meets the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located, return the first memory space to the memory allocation area, and merge the free memory space of the memory allocation area to obtain a second memory space; in response to determining that the second memory space meets the second preset condition for returning it to the operating system, return the second memory space to the operating system.
[0088] In some optional embodiments of this embodiment, the device also includes a return module, which is configured to terminate the operation in response to determining that the first memory space does not meet the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located, or the second memory space does not meet the second preset condition for returning it to the operating system.
[0089] In some optional aspects of this embodiment, the device further includes: a configuration module configured to prohibit cross-thread scheduling of coroutines.
[0090] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0091] like Figure 6 , is a block diagram of an electronic device according to a coroutine-based memory management method according to an embodiment of the present application.
[0092] 600 is a block diagram of an electronic device according to a memory management method based on coroutines according to an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.
[0093] like Figure 6 As shown, the electronic device includes: one or more processors 601, a memory 602, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are interconnected using different buses and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 601 is taken as an example.
[0094] The memory 602 is a non-transient computer-readable storage medium provided in the present application. The memory stores instructions executable by at least one processor to enable the at least one processor to execute the memory management method based on coroutines provided in the present application. The non-transient computer-readable storage medium of the present application stores computer instructions, which are used to enable a computer to execute the memory management method based on coroutines provided in the present application.
[0095] The memory 602 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as program instructions / modules corresponding to the memory management method based on coroutines in the embodiment of the present application (for example, the attached Figure 5 The processor 601 executes various functional applications and data processing of the server by running the non-transient software programs, instructions and modules stored in the memory 602, that is, the memory management method based on coroutine in the above method embodiment is implemented.
[0096] The memory 602 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required by at least one function; the data storage area may store data created by the use of an electronic device based on coroutine memory management, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory 602 may optionally include a memory remotely arranged relative to the processor 601, and these remote memories may be connected to the electronic device based on coroutine memory management via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0097] The electronic device of the memory management method based on coroutine may further include: an input device 603 and an output device 604. The processor 601, the memory 602, the input device 603 and the output device 604 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0098] The input device 603 can receive input digital or character information, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a track ball, a joystick, and other input devices. The output device 604 may include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.
[0099] Various implementations of the systems and techniques described herein can be realized in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0100] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for programmable processors and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or means (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0102] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0103] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
[0104] According to the technical solution of the embodiment of the present application, the memory occupancy rate is reduced while the system concurrency performance is effectively improved.
[0105] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this application can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this application can be achieved, and this document is not limited here.
[0106] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this application should be included in the protection scope of this application.
Claims
1. A memory management method based on coroutine, the method comprising: Start at least one coroutine scheduling thread; Configure at least one memory allocation area for each coroutine scheduling thread, wherein each memory allocation area in the at least one memory allocation area corresponds to a thread local cache area; A coroutine is started, and when the coroutine is submitted to a coroutine scheduling thread of the at least one coroutine scheduling thread, the coroutine is bound to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm, and after the binding relationship is determined, according to the coroutine's request to release memory, memory is merged and returned based on the memory allocation area bound to the coroutine group to which the coroutine is located, wherein coroutines bound to the same memory allocation area constitute a coroutine group.
2. The method according to claim 1, further comprising: Under the condition that cross-thread scheduling of coroutines is allowed, atomic variables are marked before performing memory allocation operations based on coroutine memory requests, or before performing memory merging and returning operations based on coroutine memory release requests; In response to determining that the memory allocation operation is completed or the memory merge and return operation is completed, the atomic variable is cleared.
3. The method according to claim 1, further comprising: In response to obtaining a first request from a first coroutine for applying for memory, searching for a target memory space corresponding to the first request in a thread local cache area corresponding to a memory allocation area bound to a coroutine group to which the first coroutine is located; In response to the search being successful, allocating the target memory space; In response to a search failure, a target memory space corresponding to the first request is searched for and allocated in a memory allocation area bound to the coroutine group where the first coroutine is located.
4. The method according to claim 1, further comprising: In response to obtaining a second request from the second coroutine to release memory, merging free memory spaces in a thread local cache area corresponding to a memory allocation area bound to a coroutine group to which the second coroutine is located to obtain a first memory space; In response to determining that the first memory space meets a first preset condition for returning the memory allocation area bound to the coroutine group in which the second coroutine is located, returning the first memory space to the memory allocation area, and merging the free memory space of the memory allocation area to obtain a second memory space; In response to determining that the second memory space meets a second preset condition for returning the second memory space to the operating system, the second memory space is returned to the operating system.
5. The method according to claim 4, further comprising: In response to determining that the first memory space does not meet the first preset condition for returning the memory allocation area bound to the coroutine group where the second coroutine is located or the second memory space does not meet the second preset condition for returning it to the operating system, the operation is terminated.
6. The method according to any one of claims 3 to 5, further comprising: Configuration prohibits cross-thread scheduling of coroutines.
7. A memory management device based on coroutine, the device comprising: A startup module, configured to start at least one coroutine scheduling thread; A configuration module is configured to configure at least one memory allocation area for each coroutine scheduling thread, wherein each memory allocation area in the at least one memory allocation area corresponds to a thread local cache area; A binding module is configured to start a coroutine and, when submitting the coroutine to one of the at least one coroutine scheduling threads, bind the coroutine to a memory allocation area corresponding to the coroutine scheduling thread to which the coroutine is submitted according to a preset load balancing algorithm, and after the binding relationship is determined, merge and return the memory based on the memory allocation area bound to the coroutine group to which the coroutine is located according to the coroutine's request to release memory, wherein the coroutines bound to the same memory allocation area constitute a coroutine group.
8. The device according to claim 7, further comprising: A marking module is configured to mark atomic variables before performing a memory allocation operation based on a memory request of a coroutine, or before performing a memory merging and returning operation based on a memory release request of a coroutine, under the condition that cross-thread scheduling of coroutines is allowed; In response to determining that the memory allocation operation is completed or the memory merge and return operation is completed, the atomic variable is cleared.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.