Memory copying method and system, electronic equipment and storage medium
By building a memory copy service within the operating system and utilizing technologies such as fine-grained copy tasks, synchronization queues, and hardware resource acceleration, the performance limitations of existing memory copy solutions are addressed, resulting in faster and more efficient memory copying.
Patent Information
- Application Number
- CN202511158752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-17
AI Technical Summary
Existing memory copying solutions cannot meet the high demands of real-world copying, especially in hardware optimization and zero-copy methods, which suffer from performance deficiencies, page alignment challenges, and security issues.
By building a memory copy service within the operating system, leveraging the operating system service architecture, and employing technologies such as fine-grained copy tasks, synchronization queues, hardware resource acceleration, parallel processing, and dependency tracking, the memory copy process is optimized, supporting both asynchronous and synchronous copy modes and fully utilizing system hardware capabilities.
It accelerates memory copy performance, reduces time overhead, increases copy speed, meets copy needs in multiple scenarios, avoids unnecessary copy operations, and improves overall copy efficiency.
Smart Images

Figure CN120803736A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, in particular to a memory copy method and system, an electronic device and a storage medium. BACKGROUND
[0002] Memory copy refers to a process of copying data from one address region to another address region in memory. A memory copy usually needs to specify the source memory region (start address and length) and the target memory region (start address and length) of the copied data. The two memory regions can exist in interleaving. Memory copy is usually completed based on virtual memory in modern computers, but can also be completed directly on physical memory. This operation usually involves the movement of a large amount of data, which has a direct impact on data processing and transmission performance. However, the existing memory copy schemes cannot meet the actual copy requirements of high requirements. SUMMARY
[0003] The technical problem to be solved by the present disclosure is to overcome the defects in the prior art that cannot meet the actual copy requirements of high requirements, and the purpose is to provide a memory copy method, system, electronic device and storage medium.
[0004] The present disclosure solves the above technical problems by the following technical solutions:
[0005] In a first aspect, the present disclosure provides a memory copy method, comprising:
[0006] An operating system service for memory copy constructed in an operating system is adopted to receive a copy request initiated by a preset object, and memory copy processing is performed according to the copy request.
[0007] Optionally, the step of performing memory copy processing according to the copy request comprises:
[0008] The target copy task corresponding to the copy request is divided into a plurality of segment copy tasks by a preset granularity, and memory copy processing is performed on different segments.
[0009] Optionally, the memory copy method further comprises:
[0010] The copy progress of the segment copy task is marked based on a descriptor specified by a preset object; wherein each segment copy task corresponds to a descriptor.
[0011] Optionally, when the descriptor of any segment copy task represents that the copy is completed, the copy data corresponding to the segment copy task is provided for use by the preset object.
[0012] Optionally, the operating system service is provided with a synchronization queue.
[0013] When the descriptor of the sub-copy task of the target segment represents an unfinished copy, the memory copy method further comprises:
[0014] The synchronization queue is used to separately submit the sub-copy task of the target segment, so as to preferentially execute the asynchronous copy operation of the sub-copy task of the target segment.
[0015] Optionally, the operating system service has the attribute of being able to use all hardware resources of the system.
[0016] The memory copy processing step according to the copy request further comprises:
[0017] Based on each sub-copy task, a matching target hardware resource is determined, and the target hardware resource is used to accelerate the copy processing of the corresponding sub-copy task.
[0018] Optionally, the step of using the target hardware resource to accelerate the copy processing of the corresponding sub-copy task comprises:
[0019] A SIMD (Single Instruction Multiple Data) instruction set is used to accelerate the movement processing of the sub-copy task.
[0020] Optionally, the memory copy method further comprises:
[0021] A central processing unit (CPU) and a direct memory access (DMA) are used to perform parallel processing on different sub-copy tasks.
[0022] Optionally, the step of using a central processing unit (CPU) and a direct memory access (DMA) to perform parallel processing on different sub-copy tasks comprises:
[0023] The central processing unit (CPU) is used to process the sub-copy task with a copy data volume less than a first preset value, and the direct memory access (DMA) is used to process the sub-copy task with a copy data volume greater than or equal to the first preset value.
[0024] The time length difference between the task completion time length of the central processing unit (CPU) and the task completion time length of the direct memory access (DMA) is less than a second preset value.
[0025] Optionally, the step of performing the memory copy processing according to the copy request comprises:
[0026] In response to the copy request corresponding to a plurality of copy tasks, performing a merging processing on the plurality of copy tasks to obtain a merged task;
[0027] performing the memory copy processing on the merged task, or performing the copy processing on the merged task first and then performing the copy processing on each of the copy tasks before merging.
[0028] Optionally, the step of performing the memory copy processing according to the copy request comprises:
[0029] In response to the copy request corresponding to a plurality of copy tasks, extracting a first dependency relationship between initial to-be-copied data in each of the copy tasks;
[0030] Based on the first dependency relationship, filtering out final to-be-copied data in each of the copy tasks, and performing the memory copy processing based on the final to-be-copied data;
[0031] Wherein, there is no intersection between the final to-be-copied data corresponding to different copy tasks.
[0032] Optionally, before the step of performing the memory copy processing according to the copy request, the memory copy method further comprises:
[0033] Obtaining a first barrier task submitted by a kernel after a system subsidence; wherein, the first barrier task is used to record first copy indication information of a user copy queue;
[0034] Obtaining a second barrier task submitted by the kernel before returning to a user mode; wherein, the second barrier task is used to record second copy indication information of a kernel queue;
[0035] Based on the first barrier task and the second barrier task, indicating a second dependency relationship corresponding between the user copy queue and the kernel queue;
[0036] The step of performing the memory copy processing according to the copy request comprises:
[0037] In response to the copy request corresponding to a plurality of copy tasks, based on the second dependency relationship, obtaining a copy execution order of each of the copy tasks in the user copy queue and the kernel queue, and performing an asynchronous copy operation on the plurality of copy tasks.
[0038] Optionally, a copy queue is provided in the operating system service;
[0039] The step of performing the memory copy processing according to the copy request comprises:
[0040] adding a target copy task in the copy request initiated by the preset object into the copy queue for memory copy processing;
[0041] Alternatively, the operating system service interacts with the preset object by calling a system call.
[0042] Optionally, the memory copy method further comprises:
[0043] In response to the presence of at least two preset objects, the copy request of the preset object with the shortest length of service copy data or other preset strategies is selected by a scheduling unit for management.
[0044] Optionally, the operating system service supports asynchronous copy and synchronous copy.
[0045] In a second aspect of the present disclosure, a memory copy system is provided, comprising a processor, which performs the following operations:
[0046] An operating system service for memory copy constructed in an operating system is used to receive a copy request initiated by a preset object, and perform memory copy processing according to the copy request.
[0047] In a third aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, wherein the processor implements the memory copy method of the first aspect when executing the computer program.
[0048] In a fourth aspect of the present disclosure, a computer readable storage medium is provided, which stores a computer program, wherein the computer program is executed by a processor to implement the memory copy method of the first aspect.
[0049] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program is executed by a processor to implement the memory copy method of the first aspect.
[0050] On the basis of common sense in the art, the above-mentioned preferred conditions can be combined arbitrarily, i.e., to obtain each preferred example of the present disclosure.
[0051] The positive progress effect of the present disclosure is that:
[0052] In the present disclosure, the memory copy service is provided through the operating system service architecture, which can support full utilization of various hardware capabilities of the system to accelerate the memory copy performance, support utilization of the global view of the system service to optimize the copy, etc., so as to achieve the effects of accelerating the memory copy rate of the preset object such as an application program, ensuring the copy processing efficiency, and reducing the overall memory copy time overhead. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 Flow chart of the memory copy method of Embodiment 1 of the present disclosure;
[0054] Figure 2 Interaction diagram between the preset object and the operating system service in Embodiment 2 of the present disclosure;
[0055] Figure 3 First flow chart of the memory copy method of Embodiment 2 of the present disclosure;
[0056] Figure 4 Second flow chart of the memory copy method of Embodiment 2 of the present disclosure;
[0057] Figure 5 Schematic diagram of the multi-hardware hybrid copy of Embodiment 2 of the present disclosure;
[0058] Figure 6 Schematic diagram of the memory copy method of Embodiment 2 of the present disclosure not involving copy absorption;
[0059] Figure 7 Schematic diagram of the memory copy method of Embodiment 2 of the present disclosure involving copy absorption;
[0060] Figure 8 Schematic diagram of the different queue dependency tracking of Embodiment 2 of the present disclosure;
[0061] Figure 9 Schematic diagram of the task execution order corresponding to the different queue dependency tracking in Embodiment 2 of the present disclosure; Figure 7
[0062] Figure 10 Module schematic diagram of the memory copy system of Embodiment 3 of the present disclosure;
[0063] Figure 11 Structure schematic diagram of the electronic device of Embodiment 5 of the present disclosure. DETAILED DESCRIPTION
[0064] The present disclosure will be further described below by way of examples, but the present disclosure is not limited in the scope of the examples.
[0065] The prefix words such as "first", "second" in the embodiments of the present disclosure are only used to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of ordinal words and other prefix words in the embodiments of the present disclosure to distinguish the described objects does not constitute a limitation on the described objects, and the description of the described objects should be referred to the description of the context in the claims or embodiments, and should not constitute an unnecessary limitation because of the use of such prefix words. In addition, in the description of the embodiments, unless otherwise stated, the meaning of "a plurality of" is two or more than two.
[0066] Memory copy refers to the process of copying data from one address region to another address region in memory. A memory copy usually needs to specify the source memory region (start address and length) and the target memory region (start address and length) of the copied data. The two memory regions can exist staggered. Memory copy is usually completed based on virtual memory (Virtual Memory) in modern computers, but can also be completed directly on physical memory (Physical Memory). This operation usually involves the movement of a large amount of data, which has a direct impact on data processing and transmission performance.
[0067] There are mainly two schemes for current memory copy: (1) optimizing synchronous copy based on library by combining hardware capabilities, which has difficulty in fully utilizing all hardware optimization, and the memory copy performance still cannot meet the use requirements of many scenes; (2) zero-copy technology, which belongs to a zero-copy method based on remapping or shared memory, which has many problems such as page alignment challenge, security challenge, single copy challenge and the like in actual application.
[0068] Based on this, the present disclosure proposes a memory copy service provided through an operating system service architecture to provide performance acceleration in memory copy, specifically:
[0069] Embodiment 1
[0070] As shown in Figure 1 , the memory copy method of the present embodiment includes:
[0071] S101, an operating system service for memory copy constructed in an operating system is adopted to receive a copy request initiated by a preset object; wherein the operating system service is specially used to provide a memory copy service.
[0072] Specifically, the preset object includes but is not limited to an application program (such as a database and the like), other system kernel services (such as a network protocol stack, a file system, a memory manager and the like).
[0073] S102, memory copy processing is performed according to the copy request.
[0074] In the present solution, the memory copy service is provided through the operating system service architecture, which can support full utilization of various hardware capabilities of the system to accelerate the memory copy performance, support optimization of the copy using the global view of the system service, etc., so as to accelerate the memory copy rate of the preset object such as an application program, ensure the copy processing efficiency, and reduce the overall memory copy time overhead, etc.
[0075] Embodiment 2
[0076] The memory copy method of the present embodiment is a further improvement of Embodiment 1, specifically:
[0077] In an implementable solution, the operating system service supports both asynchronous copy mode and synchronous copy mode.
[0078] In the present solution, after the memory copy is made asynchronous in the asynchronous copy mode, the preset object can still use the memory copy primitive similar to the synchronous copy to complete the memory copy, and the time of the memory copy can be removed from the critical path, the calculation and the copy can be parallelized, the time delay overhead caused by the copy can be effectively avoided, the response time delay can be reduced, and the throughput can be improved.
[0079] In an implementable solution, the operating system service is provided with a copy queue (QCopy), a synchronization queue (QSync), etc., and the operating system service (which can be referred to as Copier) interacts with the preset object through the above-mentioned queues.
[0080] The queue structure of each queue can be set or adjusted according to actual conditions; for example, as shown in Figure 2 The queue structure CopyQueue of the copy queue QCopy corresponds to the copy task CopyTask, which includes Source (source address), Destination (target address), Length (length), Type (type), Func (function), Descriptor (descriptor), Granularity (granularity), etc.; the queue structure SyncQueue of the synchronization queue QSync corresponds to the synchronization task SyncTask, which includes the target address (or page) Destination and the length Length. Generally, the task processing priority (High priority) in the synchronization queue QSync is higher than the task processing priority (Normal priority) in the copy queue QCopy.
[0081] The operating system service can interact with the preset object by calling the system call, and can also interact with the preset object through the queue;
[0082] It should be noted that in the way of interaction through the queue, no privilege level context switching is needed, and the interaction can be efficiently completed.
[0083] The step S102 comprises:
[0084] The target copy task in the copy request initiated by the preset object is added to the copy queue for memory copy processing.
[0085] Specifically, the operating system service adds the target copy task in the copy request initiated by the preset object to the copy queue QCopy to start the memory copy operation.
[0086] The target copy task provides detailed information of the source memory and the target memory, which includes but is not limited to address information, page identification information, and copy length information.
[0087] In the scheme, the copy queue QCopy set by the operating system service realizes timely and reliable response to the copy request pair initiated by the preset object, and guarantees the feasibility and reliability of the entire memory copy process.
[0088] In an implementable scheme, as shown in Figure 3 The step S102 comprises:
[0089] S1021, the target copy task corresponding to the copy request is divided into a plurality of segment copy tasks by using a preset granularity, and the memory copy processing is performed on different segments.
[0090] The copy request further comprises a requirement for the preset granularity of segmenting the copy data. The granularity determines the size of the copy data corresponding to each segment copy task. The preset granularity can be a fixed value determined based on actual experience, or can be adjusted according to actual needs.
[0091] In the scheme, the segment-based granularity copy scheme is introduced in the operating system service. Specifically, the operating system service divides the target copy task in the copy request submitted by the preset object into a plurality of segments to process the copy data of each segment copy task one by one, thereby relieving the pressure of data copy. Before use, it is not necessary to complete the copy of the entire data, but only to ensure the completion of the copy of the current required data to meet the availability, so as to balance the memory copy efficiency and more reasonably meet the data copy demand of the actual scene.
[0092] In an implementable scheme, the memory copy method further comprises:
[0093] copy progress of the sub-copy task is marked by the descriptor.
[0094] As shown in Figure 2 the descriptor (Descriptor) of each sub-copy task is included in the copy request, i.e., a bitmap (bitmap) is used to track the copy status of each sub-copy task through the descriptor, and the progress of the completed copy is marked by the bitmap; the preset object confirms the copy progress of each sub-copy task by checking the descriptor; and the operation service system updates the related bitmap of the descriptor after each sub-copy task is completed. Figure 1 bitmap 0 indicates that the copy is in progress (copy in progress), and bitmap 1 indicates that the copy is completed (copy finished).
[0095] In this scheme, the copy progress of each task is monitored in real time by recording the copy progress of each sub-copy task in real time through the descriptor, and the entire memory copy process is intuitive and orderly.
[0096] In an implementable scheme, when the descriptor of any sub-copy task indicates that the copy is completed, the copy data corresponding to the sub-copy task is provided to the preset object for use.
[0097] In this scheme, through timely updating of the descriptor, the preset object can know the copy status of each sub-copy task in time, and the corresponding copy data can be used by the preset object after the completion of the copy of any sub-copy task, so that the data can be used while copying, further meeting the copy demand of the memory copy scene.
[0098] In an implementable scheme, the copy queue QCopy can process multiple target copy tasks corresponding to multiple copy requests, and the multiple target copy tasks can be initiated by the same preset object or by different preset objects respectively.
[0099] For the case of target copy tasks initiated by the same preset object, for example, an application can submit a copy request corresponding to a copy task: copy data from A to B (A→B), and then submit another copy request corresponding to a subsequent copy task: copy data from B to C (B→C).
[0100] In this scheme, the operating system service can continuously process different multiple copy tasks without introducing other complex processes, thereby ensuring the feasibility and efficiency of multiple copy task processing.
[0101] In an implementable solution, the memory copying method further comprises:
[0102] The sub-copy task of the target segment is submitted separately by using the synchronization queue, so as to preferentially execute the asynchronous copying operation of the sub-copy task of the target segment.
[0103] In the solution, when the preset object needs to use the copying data corresponding to the sub-copy task of a segment, but it is determined based on the descriptor that the corresponding sub-copy task has not been completed, that is, the copying data of the segment is unavailable, at this time, the sub-copy task of the segment can be submitted separately by using the auxiliary synchronization task of the synchronization queue QSync. The operating system service always gives priority to the task in the synchronization queue QSync, so as to improve the priority of the segment and the task dependent thereon. Thus, the operating system service will preferentially process the sub-copy task of the segment, so as to relieve the head blocking and reduce the program delay, thereby guaranteeing the response processing efficiency of the sub-copy task of the target segment and further optimizing the acceleration performance of the memory copying.
[0104] The queue structure of the synchronization queue QSync is relatively simple compared with the queue structure of the copying queue QCopy, and can only include a target address (or page) and a length. Of course, the queue structure of the synchronization queue QSync can also be adjusted according to actual needs.
[0105] In an implementable solution, the operating system service has the attribute of being able to use all hardware resources of the system.
[0106] The hardware resources include, but are not limited to, AVX (Advanced Vector Extensions), CPU, DMA, etc.
[0107] As shown in FIG. 1, step S102 further comprises: Figure 4
[0108] S1022, based on each sub-copy task, determining a matching target hardware resource, and using the target hardware resource to perform accelerated copying processing on the corresponding sub-copy task.
[0109] In the solution, based on the attribute that the operating system service can use all hardware resources of the system, for any sub-copy task, one or more matching hardware resources can be automatically used to perform accelerated processing, thereby realizing the performance acceleration of the overall memory copying.
[0110] In an implementable solution, the step of using the target hardware resource to perform accelerated copying processing on the corresponding sub-copy task in step S1022 comprises:
[0111] The SIMD instruction set is used to accelerate the moving processing of the sub-copy task.
[0112] In the scheme, the SIMD instruction set is used to accelerate the moving processing of the sub-copy task, so as to effectively improve the copy rate of each sub-copy task and accelerate the overall memory copy process.
[0113] In an implementable scheme, the memory copy method further includes:
[0114] The central processing unit (CPU) and direct memory access (DMA) are used to perform parallel processing on different sub-copy tasks.
[0115] Specifically, since the copy based on DMA is limited to continuous PA (Physical Address), the copy task is generally divided into multiple sub-tasks; for example, as shown in Figure 5 in the worst case where all pages are not continuous, the task is divided into 6 DMA tasks (i.e., corresponding to each sub-copy task), Task1, Task 2, Task3, Task4, Task5, Task6; for small copy data volume, the overhead of submitting DMA tasks and retrieving the results (330 cycles) exceeds the time required for CPU-based copy, so DMA is an inefficient choice; performance analysis of DMA copy of different sizes shows that for copy data with a copy data volume ≥ 1.4KB, DMA becomes a data copy method worth using; based on this, a hybrid copy strategy is proposed, which uses CPU functions to process small copy data volume copy tasks and lets DMA handle larger copy data volume copy tasks; for example, Figure 5 in which Task1, Task3, Task5, and Task6 are processed by CPU, and Task 2 and Task 4 are processed by DMA.
[0116] Figure 5 in which the arrow points to the page boundary, Src corresponds to the source operand, which is the starting position of data transfer, and its content is transferred to the target position; Dst (Destination) corresponds to the destination operand, which is the target position of data transfer and is used to receive the data of the source operand.
[0117] In the scheme, parallel copying is performed by the central processing unit (CPU) and direct memory access (DMA) to improve the throughput of the operating system service, thereby further improving the processing efficiency of memory copy and reducing the overall memory copy time overhead.
[0118] In one feasible solution, the steps of using a central processing unit (CPU) and direct memory access (DMA) to process different sub-copy tasks in parallel include:
[0119] The central processing unit (CPU) is used to process the sub-copy task whose copy data amount is less than the first preset value; at the same time, the direct memory access (DMA) is used to process the sub-copy task whose copy data amount is greater than or equal to the first preset value;
[0120] The difference between the task completion time of the central processing unit (CPU) and the task completion time of the direct memory access (DMA) is smaller than a second preset value.
[0121] In this solution, a scheduler is introduced into the operating system copier service to assign different sub-copy tasks to various hardware units, then bundle and send these sub-copy tasks to the DMA and AVX hardware to complete the copy tasks in parallel. The operating system copier service selects larger sub-copy tasks from a larger task (or multiple consecutive smaller tasks) for DMA processing and assigns smaller sub-copy tasks to the CPU. This ensures that the CPU and DMA tasks complete in similar times, minimizing latency and further accelerating memory copy performance.
[0122] Among them, the first preset value and the second preset value can be determined or adjusted based on factors such as hardware performance and actual scenario requirements.
[0123] In one feasible solution, the step of performing memory copy processing according to the copy request includes:
[0124] In response to a plurality of copy tasks corresponding to the copy request, the plurality of copy tasks are merged to obtain a merged task;
[0125] Perform memory copy processing on the merged task; or, first perform copy processing on the merged task, and then perform copy processing on each copy task before the merge.
[0126] For example, when Redis processes a SET, it first copies the value from the kernel (K) to the input I / O buffer (I) via recv(), and then needs to copy the value to the DB (D) after parsing the copy request; copy absorption eliminates the unnecessary copying of unused data to the input buffer (K→I) by design.
[0127] In this solution, the operating system service Copier enables the copy absorption function (i.e., task merging) based on the global view. This function can effectively simplify the processing flow of multiple copy tasks, avoid unnecessary copies, and achieve the effect of optimized copying.
[0128] In an implementable solution, the step of performing memory copy processing according to the copy request comprises:
[0129] In response to the copy request corresponding to a plurality of copy tasks, a first dependency relationship between initial copy data in each copy task is extracted;
[0130] The operating system service Copier uses the descriptor to track the dependency relationship between the copy tasks corresponding to different copy requests. The dependency relationship reflects the dependency relationship of the data (determined by address interval overlap), specifically:
[0131] If the descriptor (i.e. bitmap) corresponding to a part of the memory of a copy task is marked, it means that the part of the memory of the copy task may be modified, i.e. the current copy task depends on the part of the memory whose descriptor is marked.
[0132] On the contrary, if the descriptor (i.e. bitmap) corresponding to a part of the memory of a copy task is not marked, it means that there is no possibility of modification of the part of the memory in the copy task, i.e. the current copy task does not depend on the part of the memory, but more on the corresponding part of the copy task corresponding to the copy request of the previous layer.
[0133] Based on the first dependency relationship, the final copy data in each copy task is filtered out, and memory copy processing is performed based on the final copy data;
[0134] Wherein, there is no intersection between the final copy data corresponding to different copy tasks.
[0135] For example, during the copy processing, the operating system service Copier first copies the copy data in B to C (using the bitmap of the descriptor), and then copies the remaining part in A to C, thereby realizing only necessary copy.
[0136] Specifically, as shown in Figure 6 Before introducing the function of layered absorption (No Absorb), the copy data in A is copied to B in sequence, and then the data obtained by combining the copy data in A and the copy data in B is copied to C;
[0137] As shown in Figure 7 After introducing the function of layered absorption (Layered Copy Absorb), the descriptor is used to determine the copy data D in B that is different from the copy data in A and copy it to C first, and then copy the copy data in A except the copy data D corresponding to the position to C, thereby ensuring the accuracy of data copy while reducing unnecessary data copy.
[0138] In the scheme, the operating system service Copier uses descriptors to track dependencies and copy expected data from the latest "layer", effectively avoiding unnecessary copy processing operations, ensuring the correctness and efficiency of the copy processing.
[0139] In an implementable scheme, before the step of performing memory copy processing according to the copy request, the memory copy method further comprises:
[0140] obtaining a first barrier task submitted by the kernel after the system sinks; wherein the first barrier task is used to record first copy indication information of the user copy queue;
[0141] obtaining a second barrier task submitted by the kernel before returning to the user mode; wherein the second barrier task is used to record second copy indication information of the kernel queue;
[0142] based on the first barrier task and the second barrier task, indicating a second dependency relationship between the user copy queue and the kernel queue;
[0143] the step of performing memory copy processing according to the copy request, comprising:
[0144] in response to a plurality of copy tasks corresponding to the copy request, based on the second dependency relationship, obtaining the copy execution order of each copy task in the user copy queue and the kernel queue, and performing asynchronous copy operation on the plurality of copy tasks.
[0145] Specifically, absorbing cross-privilege copy (cross-privilege copy refers to cross-kernel and user state copy, from kernel state memory copy to user state memory, or from user state copy to kernel state memory) needs to track the dependency relationship between multiple copy queues (the dependency relationship corresponds to the dependency relationship of task submission order / sequence); wherein, in the case of maintaining 2 copy queues for each process (for security considerations), a kernel queue (K Queue, i.e. the copy queue QCopy in the kernel state) is used for kernel submission tasks (these tasks always contain kernel addresses), and a user copy queue (UQueue, i.e. the copy queue QCopy in the user state) is used for application programs; in order to achieve correctness and excellent absorption performance, the dependency relationship of the copy tasks submitted to the two queues needs to be considered.
[0146] wherein, for the dependency relationship between the multiple copy queues to be tracked, the submission order of the tasks in different queues needs to be determined, and the determination of the submission order is specifically implemented by the following process:
[0147] system sink (kernel state memory copy to user state memory, for example, system call) and return event (user state memory copy to kernel state memory) are used as indicators to track the dependency relationship; for the two copy queues U Queue and KQueue, as followsFigure 8 As shown, specifically, a barrier task is introduced: each time after the sink, a barrier task is submitted before the first copy task of the kernel queue (K Queue) is submitted, which records the current position of the user copy queue (U Queue) (and the number of reuse from the beginning), and the barrier task is marked as Barrier start (BS); wherein the current position is the position directly read from the queue when the task is submitted, that is, the index of the queue head; since the ring queue involves queue reuse, the queue state needs to be uniquely marked in combination with the reuse number and the queue head index;
[0148] Similarly, before returning to the user mode, the kernel submits a barrier task, and the barrier task submitted before returning is marked as Barrier end (BE).
[0149] In this way, the operating system service Copier can explicitly know the dependency relationship between the two queues (U Queue and K Queue) in the cross-privilege copy process: as Figure 9 As shown, after the tasks U1 and U2 in the user copy queue U Queue are executed, the tasks (K1-K4) in the kernel queue (K Queue) are executed synchronously, and the tasks U3 and U4 in the user copy queue U Queue are executed; then the tasks U5 and U6 in the user copy queue U Queue are executed, that is, the tasks (K1-K4) in the kernel queue (K Queue) are processed after the task U2 in the user copy queue U Queue and before the task U5.
[0150] In this way, the operating system service Copier can explicitly know the submission order of different tasks in the two queues, so as to execute the corresponding copy tasks in sequence based on the submission order, and ensure the correctness of the copy operation in the cross-privilege copy process. Of course, for the case of other multiple queues, the implementation process is similar, which will not be described here.
[0151] In an implementable scheme, the memory copy method further includes:
[0152] In response to the existence of at least two preset objects, the copy request of the preset object with the shortest copy data length or other preset strategy is selected for management by the scheduling unit.
[0153] Specifically, the operating system service Copier is provided with a scheduler, and the operating system service Copier maintains a total copy length for each preset object, and selects the preset object with the shortest copy length for service each time the scheduling is performed, so as to ensure the fairness of the memory copy process; wherein the administrator can adjust the copy slice (analogous to the time slice) of the operating system service Copier, and set the maximum copy length each time the scheduling is performed.
[0154] The operating system service copier scheduler works in each cgroup (control group) and schedules between cgroups to maintain isolation. Users can achieve performance isolation between processes by setting limits.
[0155] The following uses accelerated asynchronous copy as an example to specifically illustrate the implementation principle of the memory copy method of this embodiment:
[0156] (1) During the operating system startup process, the operating system service Copier for memory copying is started;
[0157] Among them, the operating system service Copier provides a memory copy service interface for preset objects (applications, other system kernel services, etc., and the following takes the application as an example);
[0158] (2) The operating system service Copier interacts with the application through the operating system call (Syscall) or the shared memory queue (SharedQueue). The application can call the system call or write the copy task corresponding to the copy request into the shared memory queue;
[0159] After receiving the copy task initiated by the application (including the starting address and length of the copy source memory, as well as the starting address and length of the destination address, etc.), the operating system service Copier starts processing the copy request;
[0160] (3) The operating system service Copier marks the progress of the copy by marking the bitmap provided by the application. Each bit indicates the completion status of the copy of a small segment of memory (0 means not yet completed, 1 means completed);
[0161] (4) In asynchronous copy mode, the application can continue to perform subsequent operations after submitting the copy task to the operating system service Copier, that is, asynchronous copy does not block the subsequent execution of the application;
[0162] In synchronous copy mode, the application will block and wait for the request to complete before continuing with subsequent operations;
[0163] (5) Each time the application accesses data related to the asynchronous copy, it needs to check the corresponding bit in the bitmap to determine whether the copy is complete. If the copy is complete, the application can continue to access the data. Otherwise, the application can choose to wait until the copy is complete.
[0164] (6) When involving multiple applications, the operating system service Copier can determine multiple scheduling strategies based on the length of copying completed by each application in a period of time, or the service time of each application, etc., to provide a scheduler mechanism to ensure the fairness of memory copying operations;
[0165] (7) When involving multiple applications, the operating system service Copier can determine multiple quota strategies according to the quota configured by the application, using the length of copying completed by each application in a period of time, or the service time of each application, etc., to achieve performance isolation between multiple applications;
[0166] (8) The operating system service Copier can use regular memory access instructions (load / store), hardware acceleration units (such as Intel I / OAT, etc. supporting memory copying DMA extension), single instruction multiple data (SIMD) instructions (such as Intel AVX series instructions, Intel AMX series instructions, etc.) to accelerate data copying, that is, multiple instructions and devices can be used to complete memory copying;
[0167] (9) The operating system service Copier can optimize memory copying using its global perspective, and perform scheduling and merging optimization on submitted copying tasks;
[0168] For example, if there are copying tasks from A and B and copying tasks from B to C in the copying queue; the operating system service Copier can merge them into copying from A to C according to other information (such as user semantics, etc.), or preferentially execute copying from A to C, to ensure the optimization of copying process and improve the acceleration effect of memory copying.
[0169] Based on the above implementation process, the scheme of the embodiment can achieve the following effects:
[0170] Compared with the existing memory copying mechanism based on library functions, in the operating system service, acceleration including SIMD instruction extension and DMA engine can be used at the same time, so that the hardware capability can be fully utilized, and the shortcoming that only part of the hardware can be accelerated in the existing scheme can be avoided.
[0171] Compared with other zero-copy methods using shared memory, the optimization of memory copying operations below 4KB (a physical page size) can be supported, the size of memory copying that can be optimized is reduced to the level of hundreds of bytes, and the time overhead is further optimized;
[0172] Compared with other hardware optimization-based zero-copy methods, the cost of modifying hardware is avoided, and the modification of the application program is simplified;
[0173] Compared with the existing memory copy mechanism based on the library function, the global perspective of the system service can be fully utilized, multiple copy requests can be considered as a whole, and unnecessary copy operations can be optimized, thereby further optimizing the acceleration effect of memory copy.
[0174] Compared with the existing memory copy mechanism based on the library function, the global perspective of the system service can be fully utilized, multiple copy requests can be considered as a whole, and unnecessary copy operations can be optimized, thereby further optimizing the acceleration effect of memory copy.
[0175] Embodiment 3
[0176] As shown in Figure 10 The memory copy system of the embodiment includes a processor 1, which performs the following operations:
[0177] An operating system service for memory copy constructed in the operating system is adopted to receive a copy request initiated by a preset object, and memory copy processing is performed according to the copy request.
[0178] The operating system service is specially used to provide memory copy service.
[0179] Specifically, the preset object includes but is not limited to an application program (such as a database application), other system kernel services (such as a network protocol stack, a file system, a memory manager, etc.).
[0180] In this scheme, the memory copy service is provided through the operating system service architecture, which can support full utilization of various hardware capabilities of the system to accelerate the memory copy performance, support optimization of copy by using the global perspective of the system service, etc., so as to accelerate the memory copy rate of the preset object such as the application program, ensure the copy processing efficiency, and reduce the overall memory copy time overhead.
[0181] Embodiment 4
[0182] The memory copy system of the embodiment is a further improvement of embodiment 3, and specifically:
[0183] In an implementable scheme, the processor 1 is further configured to divide the target copy task corresponding to the copy request into a plurality of segment copy tasks according to a preset granularity, and perform memory copy processing on different segments.
[0184] In an implementable scheme, the processor 1 is further configured to mark the copy progress of the segment copy task based on a descriptor specified by the preset object; wherein each segment copy task corresponds to a descriptor.
[0185] In an implementable scheme, when the descriptor of any segment copy task represents that the copy is completed, the copy data corresponding to the segment copy task is provided for use by the preset object.
[0186] In an implementation, the operating system service has a synchronization queue.
[0187] When the descriptor of the sub-copy task of the target segment represents that the copy is not completed, the processor 1 is configured to separately submit the sub-copy task of the target segment to the synchronization queue to perform the asynchronous copy operation of the sub-copy task of the target segment preferentially.
[0188] In an implementation, the operating system service has a property of using all hardware resources corresponding to the system.
[0189] The processor 1 is configured to determine a target hardware resource matched with each sub-copy task, and perform accelerated copy processing on the corresponding sub-copy task by using the target hardware resource.
[0190] In an implementation, the processor 1 performs accelerated movement processing on the sub-copy task by using a SIMD instruction set.
[0191] In an implementation, the processor 1 performs parallel processing on different sub-copy tasks by using a central processing unit (CPU) and a direct memory access (DMA).
[0192] In an implementation, the processor 1 processes a sub-copy task with a copy data volume less than a first preset value by using a central processing unit (CPU), and processes a sub-copy task with a copy data volume greater than or equal to the first preset value by using a direct memory access (DMA).
[0193] A time difference between a task completion time of the central processing unit (CPU) and a task completion time of the direct memory access (DMA) is less than a second preset value.
[0194] In an implementation, the processor 1 is configured to, in response to a plurality of copy tasks corresponding to a copy request, perform merging processing on the plurality of copy tasks to obtain a merged task.
[0195] The merged task is subjected to memory copy processing, or the merged task is subjected to copy processing, and each copy task before the merging is subjected to copy processing.
[0196] In an implementation, the processor 1 is configured to, in response to a plurality of copy tasks corresponding to a copy request, extract a first dependency relationship between initial to-be-copied data in each copy task.
[0197] The final to-be-copied data in each copy task is filtered based on the first dependency relationship, and memory copy processing is performed based on the final to-be-copied data.
[0198] The final to-be-copied data corresponding to different copy tasks have no intersection.
[0199] In an implementable solution, the first barrier task submitted by the kernel after the acquisition system sinks is acquired; wherein the first barrier task is used to record the first copy indication information of the user copy queue;
[0200] The second barrier task submitted by the kernel before returning to the user mode is acquired; wherein the second barrier task is used to record the second copy indication information of the kernel queue;
[0201] Based on the first barrier task and the second barrier task, the second corresponding dependency relationship between the user copy queue and the kernel queue is indicated;
[0202] The processor 1 is also used to, in response to the copy request corresponding to a plurality of copy tasks, based on the second dependency relationship, obtain the copy execution order of each copy task in the user copy queue and the kernel queue, and perform asynchronous copy operation on the plurality of copy tasks.
[0203] In an implementable solution, the copy queue is provided in the operating system service;
[0204] The processor 1 is also used to add the target copy task in the copy request initiated by the preset object to the copy queue for memory copy processing;
[0205] Alternatively, the operating system service interacts with the preset object in the form of calling system call.
[0206] In an implementable solution, the processor 1 is also used to, in response to the existence of at least two preset objects, manage through the scheduling unit, select the copy request of the preset object with the shortest length of service copy data or other preset strategy.
[0207] In an implementable solution, the operating system service supports asynchronous copy and synchronous copy.
[0208] For the system embodiment, since it basically corresponds to the method embodiment, the related parts are described in the part of the method embodiment. The system embodiment described above is only schematic, and the units described as separate components can or can not be physically separated, and the components of the unit can or can not be physical units, that is, they can be located in one place, or also distributed on multiple network units. According to actual needs, part or all of the modules can be selected to achieve the purpose of the present disclosure.
[0209] Embodiment 5
[0210] Figure 11This is a structural diagram of an electronic device showing an example embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor. When the processor executes the computer program, it implements the memory copy method described in any of the above embodiments. Figure 11 The electronic device 110 shown is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0211] like Figure 11 As shown, the electronic device 110 may be a general-purpose computing device, such as a server device. Components of the electronic device 110 may include, but are not limited to, the at least one processor 111, the at least one memory 112, and a bus 113 connecting different system components (including the memory 112 and the processor 111).
[0212] The bus 113 includes a data bus, an address bus, and a control bus.
[0213] The memory 112 may include a volatile memory, such as a random access memory (RAM) 1121 and / or a cache memory 1122 , and may further include a read-only memory (ROM) 1123 .
[0214] The memory 112 may also include a program tool 1125 (or utility) having a set (at least one) of program modules 1124, such program modules 1124 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0215] The processor 111 executes various functional applications and data processing by running the computer program stored in the memory 112, such as the memory copy method provided in any of the above embodiments.
[0216] The electronic device 110 can also communicate with one or more external devices 114 (e.g., a keyboard, pointing device, etc.). Such communication can be performed via an input / output (I / O) interface 115. Furthermore, the electronic device 110 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 116. As shown, the network adapter 116 communicates with other modules of the electronic device 110 via a bus 113. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 110, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0217] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into units / modules embodied by multiple units / modules.
[0218] Embodiment 6
[0219] The embodiments of the present disclosure further provide a computer readable storage medium, which has stored thereon a computer program. The program, when executed by a processor, implements the memory copy method provided by any of the above embodiments.
[0220] More specifically, the readable storage medium can include, but is not limited to, a portable disc, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0221] Embodiment 7
[0222] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The computer program, when executed by a processor, implements the memory copy method according to any of the above embodiments.
[0223] The program code for carrying out the computer program product of the present disclosure can be written in any combination of one or more programming languages, and can be executed entirely on the user device, partially on the user device, as a standalone software package, partially on the user device and partially on a remote device, or entirely on a remote device.
[0224] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an illustration, and the protection scope of the present disclosure is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, and such changes and modifications all fall within the protection scope of the present disclosure.
Claims
1. A memory copy method, characterized in that: The memory copy method includes: An operating system service for memory copy constructed in the operating system is used to receive a copy request initiated by a preset object, and perform memory copy processing according to the copy request.
2. The memory copy method according to claim 1, wherein: The step of performing memory copy processing according to the copy request includes: The target copy task corresponding to the copy request is divided into several segment copy tasks using a preset fine granularity, and memory copy processing is performed on the different segments.
3. The memory copy method according to claim 2, wherein: The memory copy method further includes: The copy progress of the sub-copy task is marked based on a descriptor specified by a preset object; wherein each sub-copy task corresponds to one descriptor.
4. The memory copy method according to claim 3, wherein: When the descriptor of any segment of the sub-copy task indicates that copying is completed, the copy data corresponding to the sub-copy task is provided to the preset object for use.
5. The memory copy method according to claim 3, wherein: The operating system service is provided with a synchronization queue; When the descriptor of the sub-copy task of the target segment indicates that the copy is not completed, the memory copy method further includes: The synchronous queue is used to submit the sub-copy task of the target segment separately, so as to preferentially execute the asynchronous copy operation of the sub-copy task of the target segment.
6. The memory copy method according to claim 2, wherein: The operating system service has the property of being able to use all hardware resources corresponding to the system; The step of performing memory copy processing according to the copy request further includes: Based on each of the sub-copy tasks, a matching target hardware resource is determined, and the target hardware resource is used to perform accelerated copy processing on the corresponding sub-copy task.
7. The memory copy method according to claim 6, wherein: The step of using the target hardware resources to accelerate the copy processing of the corresponding sub-copy task includes: The SIMD instruction set is used to accelerate the movement processing of the sub-copy task.
8. The memory copy method according to claim 6, wherein: The memory copy method further includes: A central processing unit (CPU) and direct memory access (DMA) are used to process the different sub-copy tasks in parallel.
9. The memory copy method according to claim 8, wherein: The step of using a central processing unit (CPU) and direct memory access (DMA) to process different sub-copy tasks in parallel includes: The central processing unit (CPU) is used to process the sub-copy task whose copy data amount is less than a first preset value; and at the same time, the direct memory access (DMA) is used to process the sub-copy task whose copy data amount is greater than or equal to the first preset value. The difference between the task completion time of the central processing unit (CPU) and the task completion time of the direct memory access (DMA) is smaller than a second preset value.
10. The memory copy method according to claim 1, wherein: The step of performing memory copy processing according to the copy request includes: In response to the copy request corresponding to multiple copy tasks, merging the multiple copy tasks to obtain a merged task; Perform memory copy processing on the merged task; or, first perform copy processing on the merged task, and then perform copy processing on each of the copy tasks before the merge.
11. The memory copy method according to claim 1, wherein: The step of performing memory copy processing according to the copy request includes: In response to the copy request corresponding to a plurality of copy tasks, extracting a first dependency relationship between initial to-be-copied data in each of the copy tasks; Filtering out final data to be copied in each of the copy tasks based on the first dependency relationship, and performing memory copy processing based on the final data to be copied; The final data to be copied corresponding to different copy tasks have no intersection.
12. The memory copy method according to claim 1, wherein: Before the step of performing memory copy processing according to the copy request, the memory copy method further includes: Obtaining a first barrier task submitted by the kernel after the system crashes; wherein the first barrier task is used to record first copy indication information of the user copy queue; Obtaining a second barrier task submitted by the kernel before returning to user mode; wherein the second barrier task is used to record second copy indication information of the kernel queue; Indicating a corresponding second dependency relationship between the user copy queue and the kernel queue based on the first barrier task and the second barrier task; The step of performing memory copy processing according to the copy request includes: In response to the copy request corresponding to multiple copy tasks, the copy execution order of each copy task in the user copy queue and the kernel queue is obtained based on the second dependency relationship, and asynchronous copy operations are performed on the multiple copy tasks.
13. The memory copy method according to claim 1, wherein: The operating system service is provided with a copy queue; The step of performing memory copy processing according to the copy request includes: Adding the target copy task in the copy request initiated by the preset object to the copy queue to perform memory copy processing; Alternatively, the operating system service interacts with the preset object by calling a system call.
14. The memory copy method according to claim 1, wherein: The memory copy method further includes: In response to the existence of at least two preset objects, the scheduling unit manages and selects the copy request or other preset strategy of the preset object with the shortest length of service copy data.
15. The memory copy method according to any one of claims 1 to 14, characterized in that: The operating system service supports asynchronous copy and synchronous copy.
16. A memory copy system, characterized in that: The memory copy system includes a processor, and the processor performs the following operations: An operating system service for memory copy constructed in the operating system is used to receive a copy request initiated by a preset object, and perform memory copy processing according to the copy request.
17. An electronic device comprising a memory, a processor, and a computer program stored in the memory and configured to run on the processor, wherein: When the processor executes the computer program, the memory copy method according to any one of claims 1 to 15 is implemented.
18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the memory copy method according to any one of claims 1 to 15 is implemented.
19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the memory copy method according to any one of claims 1 to 15 is implemented.