Data processing method, device and system based on paravirtualized device
By aggregating multiple initial data in the paravirtualized device and sending a DMA request, the problem of DMA performance degradation caused by frequent interaction between the paravirtualized device and the host is solved, and the effect of improving DMA performance is achieved.
Patent Information
- Application Number
- CN202210153414.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In the virtualization implementation of virtio device that combines soft and hard software, the interaction between the paravirtualized device and the host frequently leads to a degradation of DMA performance.
By acquiring a plurality of initial data stored in the completion queue of the paravirtualized device, determining a plurality of first data that meets the preset conditions, performing an aggregation operation to generate a first aggregation result, and sending a direct memory access request carrying the first aggregation result to the memory of the host.
Reduces the number of operations generated by updating the used ring, avoids the backpressure of the PCIe interface on the device side, and improves DMA performance.
Smart Images

Figure CN114637574B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtualization technology, and in particular to a data processing method, device and system based on a semi-virtualized device. Background Art
[0002] At present, in the implementation of virtio device virtualization that combines software and hardware, the host and the device are connected through PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard). According to the virtio device specification, when the device receives data, it needs to submit it to the CPU (Central Processing Unit) through multiple steps, and each step is to write the host memory through DMA (Direct Memory Access). However, when each step is too frequent, the CPU PCIe subsystem will have a bottleneck, the device side will see PCIe interface back pressure, and DMA performance will decrease.
[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, apparatus and system based on a para-virtualized device, so as to at least solve the technical problem in the related art that the para-virtualized device frequently interacts with the host, resulting in a decrease in DMA performance.
[0005] According to one aspect of an embodiment of the present application, a data processing method based on a para-virtualized device is provided, including: obtaining multiple initial data stored in a completion queue of the para-virtualized device, wherein the multiple initial data are used to characterize descriptive information of original data that has been processed by the para-virtualized device but has not been submitted to the host; determining multiple first data that meet preset conditions in the multiple initial data; performing an aggregation operation on the multiple first data to generate a first aggregation result; and sending a direct memory access request carrying the first aggregation result to the host's memory.
[0006] According to another aspect of an embodiment of the present application, a data processing device based on a para-virtualized device is also provided, including: a data acquisition module, used to acquire multiple initial data stored in a completion queue of the para-virtualized device, wherein the multiple initial data are used to characterize the descriptive information of the original data that has been processed by the para-virtualized device but has not been submitted to the host; a data determination module, used to determine multiple first data that meet preset conditions in the multiple initial data; an aggregation module, used to perform aggregation operations on the multiple first data to generate a first aggregation result; and a sending module, used to send a direct memory access request carrying the first aggregation result to the host's memory.
[0007] According to another aspect of an embodiment of the present application, a data processing system based on a para-virtualized device is also provided, including: a host, including: a memory and a completion queue; a para-virtualized device, connected to the host, for obtaining multiple initial data stored in the completion queue, wherein the multiple initial data are used to characterize the descriptive information of the original data that has been processed by the para-virtualized device but has not been submitted to the host; determining multiple first data that meet preset conditions in the multiple initial data; performing an aggregation operation on the multiple first data to generate a first aggregation result; and sending a direct memory access request carrying the first aggregation result to the host's memory.
[0008] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, the computer-readable storage medium including a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned data processing method based on the semi-virtualized device.
[0009] According to another aspect of an embodiment of the present application, a computer terminal is further provided, comprising: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the above-mentioned data processing method based on a semi-virtualized device is executed when the program is run.
[0010] In an embodiment of the present application, when it is necessary to update the used ring, multiple initial data stored in the used can be obtained, and then multiple first data that meet the preset conditions can be screened out from the multiple initial data, and the multiple first data are aggregated to generate a first aggregation result, and a DMA request carrying the first aggregation result is sent to the memory, so as to achieve the purpose of updating multiple queue items of the used ring at one time. It is easy to notice that by initiating a DMA request through an aggregation operation, there is no need to initiate a DMA request for each queue item, thereby achieving the technical effect of reducing the number of operations generated by updating the used ring, avoiding the device side from seeing the back pressure of the PCIe interface, and improving the DMA performance, thereby solving the technical problem in the related technology that the semi-virtualized device and the host frequently interact with each other, resulting in a decrease in DMA performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0012] Figure 1 It is a schematic diagram of updating a used ring according to the prior art;
[0013] Figure 2It is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method based on a paravirtualized device according to an embodiment of the present application;
[0014] Figure 3 is a flow chart of a data processing method based on a paravirtualized device according to an embodiment of the present application;
[0015] Figure 4 It is a schematic diagram of an optional software-hardware combined virtio device virtualization implementation architecture according to an embodiment of the present application;
[0016] Figure 5 is a schematic diagram of an optional update of a used ring according to an embodiment of the present application;
[0017] Figure 6 is a schematic diagram of an optional queue scheduling according to an embodiment of the present application;
[0018] Figure 7 is a schematic diagram of a data processing device based on a paravirtualized device according to an embodiment of the present application;
[0019] Figure 8 is a schematic diagram of a data processing system based on a paravirtualized device according to an embodiment of the present application;
[0020] Fig. 9 It is a structural block diagram of a computer terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following explanations:
[0024] virtio: virtio is an I / O paravirtualization solution, a set of general I / O device virtualization programs, and an abstraction of a set of general I / O devices in the paravirtualized hypervisor. Devices that use the virtio protocol are called virtio devices.
[0025] Data buffer: stores the data received by the device.
[0026] Used ring: It is the completion queue of the virtio device. When the device completes a request sent by the driver, the hardware notifies the driver that the command has been completed by submitting the used ring. The used ring points to the structure of the data buffer and only contains the description information of the data (address, length, etc.).
[0027] Used ring index: is the pointer to the used ring.
[0028] Currently, after the current device receives data, it needs to go through the following three steps to submit it to the CPU: write data buffer; write used ring; write used ring index. For example, take the submission of used ring as an example. Figure 1As shown, usedring contains 8 queue items, the submitted queue items are shown with solid boxes, and the unsubmitted queue items are shown with hollow boxes. The existing steps of writing used ring are as follows: submit queue item 0 to the CPU, and update the used ring index to 1, indicating that it has not been submitted since queue 1; submit queue 1 to the CPU, and update the used ring index to 2, indicating that it has not been submitted since queue 2; submit queue 2 to the CPU, and update the used ring index to 3, indicating that it has not been submitted since queue 3; submit queue 3 to the CPU, and update the used ring index to 4, indicating that it has not been submitted since queue 4.
[0029] Therefore, when multiple queue items in the used ring need to be updated, or the used ring index needs to be updated multiple times, multiple DMA requests need to be initiated, resulting in a large number of operations and reduced DMA performance.
[0030] In order to solve the above problems, the present application provides an aggregate submission solution, which reduces the number of updates and improves DMA performance by aggregating multiple DMA requests that need to be initiated into one.
[0031] Example 1
[0032] According to an embodiment of the present application, a data processing method based on a paravirtualized device is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0033] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 2 The hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method based on a semi-virtualized device is shown. Figure 2 As shown, the computer terminal 20 (or mobile device) may include one or more (202a, 202b, ..., 202n are used to illustrate) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 204 for storing data, and a transmission device 206 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 2The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 2 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0034] It should be noted that the one or more processors 202 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuits may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 20 (or mobile device). The data processing circuits may be used as a processor to control (e.g., the selection of a variable resistor terminal path connected to an interface).
[0035] The memory 204 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the data processing method based on the semi-virtualized device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 204, that is, realizing the above-mentioned data processing method based on the semi-virtualized device. The memory 204 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 204 may further include a memory remotely arranged relative to the processor 202, and these remote memories may be connected to the computer terminal 20 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0036] The transmission device 206 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 20. In one example, the transmission device 206 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 206 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0037] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 20 (or mobile device).
[0038] It should be noted that, in some optional embodiments, the above Figure 2The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 2 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).
[0039] Under the above operating environment, this application provides Figure 3 The data processing method based on the paravirtualized device is shown. Figure 3 1 is a flow chart of a data processing method based on a paravirtualized device according to an embodiment of the present application. Figure 3 As shown, the method comprises the following steps:
[0040] Step S302 : acquiring a plurality of initial data stored in a completion queue of the para-virtualized device, wherein the plurality of initial data are used to represent description information of original data that has been processed by the para-virtualized device but has not yet been submitted to the host.
[0041] The semi-virtualized device in the above steps can be installed on the host, for example, it can be a semi-virtualized network card, or it can be a semi-virtualized memory, but it is not limited to this. The initial data in the above steps can be a data item stored in the used ring that is not submitted to the CPU, and each data item stores description information of the corresponding original data, and the description information may include the address of the data buffer storing the original data, the length of the original data, etc., but it is not limited to this. For different types of semi-virtualized devices, the type of original data is different. For example, for a semi-virtualized network card, the original data can be an original message.
[0042] For example, Figure 4 The hardware-software combined virtio device virtualization implementation architecture shown is used as an example for explanation. After the device receives the original data, it first executes the write buffer step to write the original data into different buffers and waits for the host to process the original data. After the virtio device completes the processing, that is, after writing the original data into different buffers, the corresponding description information can be stored in the used ring. For example, the original data stored in buffer 0 to buffer 2 have been written into the corresponding buffers, and the corresponding description information can be stored in the used ring, corresponding to queue items 0 to 2 respectively. Then the write used ring step is executed, and the corresponding description information in queue items 0 to 2 can be used as the initial data.
[0043] Step S304: determining a plurality of first data satisfying a preset condition among the plurality of initial data.
[0044] The preset condition in the above steps may be an aggregation condition set in advance according to actual needs. For example, the condition may be to directly aggregate all unsubmitted queue items in the used ring; the condition may also be to aggregate unsubmitted queue items corresponding to the same network card transmission queue, but is not limited thereto.
[0045] For example, Figure 4 The software and hardware combined virtio device virtualization implementation architecture is used as an example for explanation. It is assumed that the preset condition is to submit all unsubmitted queue items in the used ring. Therefore, queue items 0 to 2 can be used as the first data.
[0046] Step S306: performing an aggregation operation on the plurality of first data to generate a first aggregation result.
[0047] In an optional embodiment, the above-mentioned aggregation operation may be to concatenate the description information corresponding to the multiple first data, for example, to concatenate the addresses and sum up the lengths, so as to obtain the above-mentioned first aggregation result, but is not limited to this, and other aggregation operations may also be used.
[0048] Step S308: Send a direct memory access request carrying the first aggregation result to the memory of the host.
[0049] In an optional embodiment, for the first aggregation result, in order to update the memory corresponding to the used ring at one time, the first aggregation result can be encapsulated according to the DMA protocol to obtain a DMA request, and the DMA request is sent to the host's memory to complete the purpose of updating the used ring.
[0050] For example, Figure 5 Taking the used ring shown as an example, the 8 queue items contained in the used ring have not been submitted to the memory. At this time, queue item 0 has been submitted, and queue items 1 to 7 can be used as initial data. Queue items 1 to 3 can be screened out as the first data through preset conditions, and the three queue items can be aggregated to generate a DMA request to update items 1 to 3 of the used ring at one time, and the boxes corresponding to queue items 1 to 3 are changed to solid boxes, indicating that queue items 1 to 3 have been submitted.
[0051] According to the solution provided by the above-mentioned embodiment of the present application, when it is necessary to update the used ring, multiple initial data stored in the used can be obtained, and then multiple first data that meet the preset conditions can be screened out from the multiple initial data, and the multiple first data are aggregated to generate a first aggregation result, and a DMA request carrying the first aggregation result is sent to the memory, so as to achieve the purpose of updating multiple queue items of the used ring at one time. It is easy to notice that by initiating a DMA request by means of an aggregation operation, there is no need to initiate a DMA request for each queue item, thereby achieving the technical effect of reducing the number of operations generated by updating the used ring, avoiding the device side from seeing the back pressure of the PCIe interface, and improving the DMA performance, thereby solving the technical problem in the related technology that the semi-virtualized device and the host frequently interact with each other, resulting in a decrease in DMA performance.
[0052] In the above-mentioned embodiment of the present application, determining multiple first data satisfying preset conditions among multiple initial data includes at least one of the following: acquiring data stored in a target cache to obtain multiple first data, wherein the multiple initial data are cached to the target cache in sequence; when a preset timing time is reached, determining the multiple initial data to be multiple first data; when the number of the multiple initial data is greater than or equal to a preset number, determining the multiple initial data to be multiple first data.
[0053] The target cache may be a pre-set adaptive cache, and when a bottleneck occurs in PCIe performance, data is accumulated in the adaptive cache. In an optional embodiment, queue items in the used ring are sequentially stored in the adaptive cache before being submitted to the CPU, so all data items stored in the adaptive cache may be used as the first data.
[0054] The above-mentioned preset timing time can be a time of a preset timer, which can be set according to actual needs. In an optional embodiment, before the preset timing time arrives, there is no need to perform any processing on the queue items stored in the used ring, and after the preset timing time arrives, all unsubmitted queue items stored in the used ring can be used as the first data.
[0055] The above-mentioned preset number may be a preset maximum aggregation number, which may be set according to actual needs. In an optional embodiment, before the number of unsubmitted queue items stored in the used ring reaches the preset number, there is no need to perform any processing on the queue items stored in the used ring. After the number of unsubmitted queue items reaches the preset number, all unsubmitted queue items stored in the used ring may be used as the first data.
[0056] It should be noted that the above three conditions can be used alone or in any combination, for example, combining the target cache with a preset time, combining the target cache with a preset number, combining the preset time and the preset number, and combining the target cache, the preset time and the preset number. The specific combination can be determined according to actual needs, and this application does not make specific limitations on this.
[0057] For example, Figure 4 Taking the software and hardware combined virtio device virtualization implementation architecture as an example, the preset conditions may include adaptive cache, timer and maximum aggregation number. The above three methods can be used in combination to determine the first data and perform aggregation operation on the first data.
[0058] In the above embodiment of the present application, an aggregation operation is performed on multiple first data to generate a first aggregation result, including: determining a first transmission queue corresponding to each first data, wherein the first transmission queue is used to transmit the original data; obtaining the first data corresponding to the same first transmission queue to obtain a first aggregation result.
[0059] Since the virtio device supports multiple queues, the first transmission queue mentioned above may be a queue queue supported by the virtio device. The virtio device may use multiple queues to transmit original data to the host, and each queue corresponds to a used ring.
[0060] In an optional embodiment, for each queue item, the queue to which the corresponding original data is sent can be determined. Because before the aggregation operation, queue items corresponding to different queues are often intertwined and cannot be directly aggregated, the queue items corresponding to the same queue can be arranged together by scheduling by queue, and then the queue items corresponding to the same queue can be aggregated to obtain the first aggregation result.
[0061] For example, Figure 4 The hardware and software combined virtio device virtualization implementation architecture is used as an example to illustrate Figure 6 As shown in the figure, boxes filled with different patterns represent queue items corresponding to different queues, and different numbers represent the sequence numbers in the queue. Before scheduling, queue items corresponding to different queues are intertwined. After scheduling, queue items corresponding to the same queue are adjacent. Therefore, queue items corresponding to the same queue can be aggregated and submitted to the memory at one time.
[0062] In the above-mentioned embodiment of the present application, obtaining the first data corresponding to the same first transmission queue and obtaining the first aggregation result includes: determining the first transmission queue corresponding to the first first data among multiple first data to obtain the target transmission queue; obtaining the first data corresponding to the target transmission queue among multiple first data to obtain the first aggregation result.
[0063] In an optional embodiment, since queue items with earlier storage positions in the used ring indicate earlier processing times by the semi-virtualized device, in order to reduce the waiting time of the device sending the original data, the queue corresponding to the queue item with the earliest storage position in the first data can be used as the target queue, and all first data corresponding to the target queue can be aggregated to obtain a first aggregation result, while the first data corresponding to other queues need to wait until the first data corresponding to the target queue is submitted before being processed.
[0064] In the above embodiment of the present application, the method also includes: determining multiple second data among multiple initial data based on the initial queue identifier corresponding to the completion queue, wherein the initial queue identifier is used to represent the identification information of the data that has been submitted, and the multiple second data are used to represent the data currently submitted to the host; performing an aggregation operation on the multiple second data to obtain a second aggregation result; and updating the initial queue identifier based on the second aggregation result.
[0065] The above initial queue identifier may be a used ring index, pointing to the first uncommitted queue item in the used ring. The above update may be to update the value of the used ring index.
[0066] In an optional embodiment, after the used ring is updated, the used ring index needs to be updated. Since multiple queue items in the used ring are updated, the used ring index needs to be updated multiple times. In order to avoid multiple operations caused by multiple updates of the used ring index, the latest submitted queue item can be determined as the second data, and multiple DMA requests can be merged into one, that is, an aggregation operation is performed on the second data to obtain a second aggregation result, and the used ring index is updated once.
[0067] For example, Figure 4Taking the hardware-software combined virtio device virtualization implementation architecture as an example, after executing the write used ring step, the write used ring index step can be executed to aggregate queue items 0 to 2, and the used ring index is updated based on the second aggregation result, and the value is updated to 3.
[0068] For example, Figure 5 Take the used ring shown as an example. None of the 8 queue items in the used ring have been submitted to the memory. At this time, queue item 0 has been submitted. At this time, the value of the used ring index is 1, that is, the initial queue identifier is 1. In addition, the 1st to 3rd items of the used ring are updated at one time. Therefore, the value of the used ring index can be directly changed to 4.
[0069] It should be noted that the queue structures of the completion queues of different versions of virtio devices are different. For new versions of virtio devices, if the used ring index does not need to be updated, there is no need to perform the above steps.
[0070] In the above embodiment of the present application, performing an aggregation operation on multiple second data to obtain a second aggregation result includes: determining a second transmission queue corresponding to each second data; obtaining the second data corresponding to the same second transmission queue to obtain a second aggregation result.
[0071] The second transmission queue mentioned above may also be a queue queue supported by the virtio device. The virtio device may use multiple queues to transmit original data to the host, and each queue corresponds to a used ring.
[0072] In an optional embodiment, similar to the aggregation operation of queue items in the used ring, queue items corresponding to the same queue can be arranged together by queue scheduling, and then queue items corresponding to the same queue can be aggregated to obtain a second aggregation result.
[0073] For example, Figure 4 The hardware and software combined virtio device virtualization implementation architecture is used as an example to illustrate Figure 6 As shown, boxes filled with different patterns represent queue items corresponding to different queues. Before scheduling, queue items corresponding to different queues are intertwined. After scheduling, queue items corresponding to the same queue are adjacent. Therefore, queue items corresponding to the same queue can be aggregated and the used ring index can be updated at one time.
[0074] It should be noted that the queue scheduling method can be performed only once, that is, if the queue scheduling method has been used to perform aggregation operations during the process of updating the used ring, there is no need to use the queue scheduling method to perform aggregation operations when updating the used ring index; if the queue scheduling method has not been used to perform aggregation operations during the process of updating the used ring, the queue scheduling method is used to perform aggregation operations when updating the used ring index.
[0075] In the above embodiment of the present application, the method also includes: acquiring multiple original data; determining a third transmission queue corresponding to each original data; sorting the multiple original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; and writing the sorted data into the data buffer of the host in sequence.
[0076] The third transmission queue mentioned above may also be a queue queue supported by the virtio device. The virtio device may use multiple queues to transmit original data to the host.
[0077] In an optional embodiment, since before the aggregation operation, the original data corresponding to different queues are often intertwined and cannot be directly aggregated, the original data corresponding to the same queue can be arranged together by queue scheduling, and then written into the data buffer in sequence according to the sorted data.
[0078] For example, Figure 4 The hardware and software combined virtio device virtualization implementation architecture is used as an example to illustrate Figure 6 As shown, boxes filled with different patterns represent queue items corresponding to different queues. Before scheduling, data corresponding to different queues are intertwined. After scheduling, data corresponding to the same queue are adjacent. Therefore, the data can be stored in the data buffer in sequence, thereby ensuring that the queue items stored in the used ring belong to the same queue and are adjacent.
[0079] It should be noted that the queue scheduling method can be performed only once, that is, if the original data is cached and scheduled by queue scheduling before writing the data buffer, there is no need to perform aggregation operations by queue scheduling when subsequently updating the used ring and the used ring index; if the original data is not cached and scheduled by queue scheduling before writing the data buffer, then if the aggregation operation is performed by queue scheduling during the process of updating the used ring, there is no need to perform aggregation operations by queue scheduling when updating the used ring index; if the aggregation operation is not performed by queue scheduling during the process of updating the used ring, the aggregation operation is performed by queue scheduling when updating the used ring index.
[0080] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0081] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0082] Example 2
[0083] According to an embodiment of the present application, a data processing device based on a paravirtualized device for implementing the above-mentioned data processing method based on a paravirtualized device is also provided. Figure 7 As shown, the device 700 includes: a data acquisition module 702 , a data determination module 704 , an aggregation module 706 and a sending module 708 .
[0084] Among them, the data acquisition module 702 is used to obtain multiple initial data stored in the completion queue of the semi-virtualized device, wherein the multiple initial data are used to represent the description information of the original data that has been processed by the semi-virtualized device but has not been submitted to the host; the data determination module 704 is used to determine multiple first data that meet preset conditions in the multiple initial data; the aggregation module 706 is used to perform aggregation operations on the multiple first data to generate a first aggregation result; the sending module 708 is used to send a direct memory access request carrying the first aggregation result to the host's memory.
[0085] It should be noted that the above data acquisition module 702, data determination module 704, aggregation module 706 and sending module 708 correspond to steps S302 to S308 in Example 1, and the four modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the above Example 1. It should be noted that the above modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0086] In the above embodiments of the present application, the data determination module includes at least one of the following: a data acquisition unit, a first data determination unit, and a second data determination unit.
[0087] Among them, the data acquisition unit is used to acquire the data stored in the target cache to obtain multiple first data, wherein the multiple initial data are cached in the target cache in sequence; the first determination unit is used to determine that the multiple initial data are multiple first data when a preset timing time arrives; the second determination unit is used to determine that the multiple initial data are multiple first data when the number of the multiple initial data is greater than or equal to a preset number.
[0088] In the above embodiment of the present application, the aggregation module includes: a queue determination unit and a result acquisition unit.
[0089] The queue determination unit is used to determine the first transmission queue corresponding to each first data, wherein the first transmission queue is used to transmit the original data; the result acquisition unit is used to acquire the first data corresponding to the same first transmission queue to obtain a first aggregation result.
[0090] In the above embodiment of the present application, the result acquisition unit is also used to determine the first transmission queue corresponding to the first first data among multiple first data, obtain the target transmission queue, and obtain the first data corresponding to the target transmission queue among multiple first data to obtain the first aggregation result.
[0091] In the above embodiment of the present application, the device also includes: an update module.
[0092] Among them, the data determination module is also used to determine multiple second data among multiple initial data based on the initial queue identifier corresponding to the completion queue, wherein the initial queue identifier is used to represent the identification information of the data that has been submitted, and the multiple second data are used to represent the data currently submitted to the host; the aggregation module is also used to perform aggregation operations on the multiple second data to obtain a second aggregation result; the update module is used to update the initial queue identifier based on the second aggregation result.
[0093] In the above embodiment of the present application, the aggregation module includes: a queue determination unit and a result acquisition unit.
[0094] The queue determination unit is used to determine the second transmission queue corresponding to each second data; the result acquisition unit is used to acquire the second data corresponding to the same second transmission queue to obtain a second aggregation result.
[0095] In the above embodiment of the present application, the device also includes: a data acquisition module, a queue determination module, a sorting module and a writing module.
[0096] Among them, the data acquisition module is used to acquire multiple original data; the queue determination module is used to determine the third transmission queue corresponding to each original data; the sorting module is used to sort the multiple original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; the writing module is used to write the sorted data into the data buffer of the host in sequence.
[0097] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0098] Example 3
[0099] According to an embodiment of the present application, a data processing system based on a para-virtualized device for implementing the above-mentioned data processing method based on a para-virtualized device is also provided. Figure 8 As shown, the system includes:
[0100] The host 82 includes a memory and a completion queue.
[0101] The paravirtualized device 84 is connected to the host and is used to obtain multiple initial data stored in the completion queue, wherein the multiple initial data are used to represent the description information of the original data that has been processed by the paravirtualized device but has not been submitted to the host; determine multiple first data that meet preset conditions in the multiple initial data; perform aggregation operations on the multiple first data to generate a first aggregation result; and send a direct memory access request carrying the first aggregation result to the host's memory.
[0102] In the above-mentioned embodiment of the present application, the semi-virtualized device is also used to perform at least one of the following steps: obtaining data stored in the target cache to obtain multiple first data, wherein the multiple initial data are cached to the target cache in sequence; when a preset timing time is reached, determining that the multiple initial data are multiple first data; when the number of the multiple initial data is greater than or equal to a preset number, determining that the multiple initial data are multiple first data.
[0103] In the above embodiment of the present application, the semi-virtualized device is also used to determine the first transmission queue corresponding to each first data, and obtain the first data corresponding to the same first transmission queue to obtain a first aggregation result, wherein the first transmission queue is used to transmit the original data.
[0104] In the above embodiment of the present application, the semi-virtualized device is also used to determine the first transmission queue corresponding to the first first data among multiple first data, obtain the target transmission queue, and obtain the first data corresponding to the target transmission queue among multiple first data to obtain the first aggregation result.
[0105] In the above embodiment of the present application, the semi-virtualized device is also used to determine multiple second data among multiple initial data based on the initial queue identifier corresponding to the completion queue, wherein the initial queue identifier is used to represent the identification information of the data that has been submitted, and the multiple second data are used to represent the data currently submitted to the host; perform aggregation operations on the multiple second data to obtain a second aggregation result; and update the initial queue identifier based on the second aggregation result.
[0106] In the above embodiment of the present application, the semi-virtualized device is also used to determine the second transmission queue corresponding to each second data; the result acquisition unit is used to obtain the second data corresponding to the same second transmission queue to obtain the second aggregation result.
[0107] In the above embodiment of the present application, the host also includes: a data buffer; the semi-virtualized device is also used to obtain multiple original data; determine the third transmission queue corresponding to each original data; sort the multiple original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; and write the sorted data into the data buffer in sequence.
[0108] It should be noted that the preferred implementation scheme involved in the above embodiments of the present application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0109] Example 4
[0110] The embodiment of the present application may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal may also be replaced by a terminal device such as a mobile terminal.
[0111] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
[0112] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the data processing method based on the para-virtualized device: obtaining multiple initial data stored in the completion queue of the para-virtualized device, wherein the multiple initial data are used to characterize the description information of the original data that has been processed by the para-virtualized device but has not been submitted to the host; determining multiple first data that meet preset conditions in the multiple initial data; performing aggregation operations on the multiple first data to generate a first aggregation result; and sending a direct memory access request carrying the first aggregation result to the host's memory.
[0113] Optionally, Fig. 9 is a structural block diagram of a computer terminal according to an embodiment of the present application. Fig. 9 As shown, the computer terminal A may include: one or more (only one is shown in the figure) processors 902 , and a memory 904 .
[0114] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device based on the semi-virtualized device in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned data processing method based on the semi-virtualized device. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories can be connected to the terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0115] The processor can call the information and application stored in the memory through the transmission device to execute the following steps: obtain multiple initial data stored in the completion queue of the para-virtualized device, wherein the multiple initial data are used to represent the description information of the original data that has been processed by the para-virtualized device but has not been submitted to the host; determine multiple first data that meet preset conditions in the multiple initial data; perform aggregation operations on the multiple first data to generate a first aggregation result; and send a direct memory access request carrying the first aggregation result to the host's memory.
[0116] Optionally, the processor may also execute program code of the following steps: obtaining data stored in the target cache to obtain multiple first data, wherein multiple initial data are cached in the target cache in sequence; and / or, when a preset timing time is reached, determining that the multiple initial data are multiple first data; and / or, when the number of the multiple initial data is greater than or equal to a preset number, determining that the multiple initial data are multiple first data.
[0117] Optionally, the processor may also execute program code of the following steps: determining a first transmission queue corresponding to each first data, wherein the first transmission queue is used to transmit original data; obtaining first data corresponding to the same first transmission queue to obtain a first aggregation result.
[0118] Optionally, the processor may also execute program code of the following steps: determining a first transmission queue corresponding to a first first data among multiple first data to obtain a target transmission queue; obtaining first data corresponding to the target transmission queue among multiple first data to obtain a first aggregation result.
[0119] Optionally, the processor may also execute the program code of the following steps: based on the initial queue identifier corresponding to the completion queue, determine multiple second data among multiple initial data, wherein the initial queue identifier is used to represent the identification information of the data that has been submitted, and the multiple second data are used to represent the data currently submitted to the host; perform aggregation operations on the multiple second data to obtain a second aggregation result; and update the initial queue identifier based on the second aggregation result.
[0120] Optionally, the processor may also execute program code of the following steps: determining a second transmission queue corresponding to each second data; acquiring second data corresponding to the same second transmission queue to obtain a second aggregation result.
[0121] Optionally, the processor may also execute program code of the following steps: obtaining multiple original data; determining a third transmission queue corresponding to each original data; sorting the multiple original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; and writing the sorted data into the data buffer of the host in sequence.
[0122] By adopting the embodiment of the present application, a data processing solution based on a paravirtualized device is provided. A DMA request is initiated once by means of an aggregate operation, without initiating a DMA request for each queue item, thereby achieving the technical effect of reducing the number of operations generated by updating the used ring, avoiding the device side from seeing the back pressure of the PCIe interface, and improving the DMA performance, thereby solving the technical problem in the related technology that the paravirtualized device and the host frequently interact, resulting in a decrease in DMA performance.
[0123] It can be understood by those skilled in the art that Fig. 9 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig. 9 It does not limit the structure of the above electronic device. For example, the computer terminal A may also include Fig. 9 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig. 9 Different configurations are shown.
[0124] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0125] Example 5
[0126] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the data processing method based on the semi-virtualized device provided by the above embodiment.
[0127] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0128] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: obtaining multiple initial data stored in a completion queue of a para-virtualized device, wherein the multiple initial data are used to characterize descriptive information of original data that has been processed by the para-virtualized device but has not been submitted to the host; determining multiple first data that meet preset conditions in the multiple initial data; performing an aggregation operation on the multiple first data to generate a first aggregation result; and sending a direct memory access request carrying the first aggregation result to the host's memory.
[0129] Optionally, the storage medium is also configured to store program codes for executing the following steps: obtaining data stored in a target cache to obtain multiple first data, wherein multiple initial data are sequentially cached in the target cache; and / or, when a preset timing time is reached, determining that the multiple initial data are multiple first data; and / or, when the number of the multiple initial data is greater than or equal to a preset number, determining that the multiple initial data are multiple first data.
[0130] Optionally, the storage medium is also configured to store program code for executing the following steps: determining a first transmission queue corresponding to each first data, wherein the first transmission queue is used to transmit original data; obtaining first data corresponding to the same first transmission queue to obtain a first aggregation result.
[0131] Optionally, the storage medium is also configured to store program code for executing the following steps: determining a first transmission queue corresponding to a first first data among multiple first data to obtain a target transmission queue; obtaining first data corresponding to the target transmission queue among multiple first data to obtain a first aggregation result.
[0132] Optionally, the storage medium is also configured to store program codes for executing the following steps: determining multiple second data among multiple initial data based on an initial queue identifier corresponding to a completion queue, wherein the initial queue identifier is used to represent identification information of data that has been submitted, and the multiple second data are used to represent data currently submitted to the host; performing aggregation operations on the multiple second data to obtain a second aggregation result; and updating the initial queue identifier based on the second aggregation result.
[0133] Optionally, the storage medium is further configured to store program codes for executing the following steps: determining a second transmission queue corresponding to each second data; acquiring second data corresponding to the same second transmission queue to obtain a second aggregation result.
[0134] Optionally, the storage medium is also configured to store program codes for executing the following steps: obtaining multiple original data; determining a third transmission queue corresponding to each original data; sorting the multiple original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; and writing the sorted data into the data buffer of the host in sequence.
[0135] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0136] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0137] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0138] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0140] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0141] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A data processing method based on a paravirtualized device, characterized in that: include: Acquire a plurality of initial data stored in the completion queue of the paravirtualized device, wherein the plurality of initial data are used to represent description information of original data that has been processed by the paravirtualized device but has not been submitted to the host, and the description information includes an address and a length of the data; Determining a plurality of first data satisfying a preset condition among the plurality of initial data; Performing an aggregation operation on the multiple first data to generate a first aggregation result, wherein the aggregation operation is used to represent concatenating the description information corresponding to the multiple first data, and the concatenation includes at least one of the following: concatenating addresses in the description information corresponding to the multiple first data, and summing the lengths in the description information corresponding to the multiple first data; Sending a direct memory access request carrying the first aggregation result to the memory of the host.
2. The method according to claim 1, characterized in that Determining a plurality of first data satisfying a preset condition among the plurality of initial data comprises at least one of the following: Acquire the data stored in the target cache to obtain the plurality of first data, wherein the plurality of initial data are sequentially cached in the target cache; When the preset timing time arrives, determining the multiple initial data to be the multiple first data; When the number of the plurality of initial data is greater than or equal to a preset number, the plurality of initial data is determined to be the plurality of first data.
3. The method according to claim 1, characterized in that Performing the aggregation operation on the plurality of first data to generate a first aggregation result includes: Determine a first transmission queue corresponding to each first data, wherein the first transmission queue is used to transmit the original data; The first data corresponding to the same first transmission queue is obtained, and the aggregation operation is performed on the first data to obtain the first aggregation result.
4. The method according to claim 3, characterized in that Acquiring first data corresponding to the same first transmission queue, and performing the aggregation operation on the first data to obtain the first aggregation result includes: Determine a first transmission queue corresponding to a first first data among the plurality of first data, and obtain a target transmission queue; The first data corresponding to the target transmission queue among the multiple first data is obtained, and the aggregation operation is performed on the first data to obtain the first aggregation result.
5. The method according to claim 1, characterized in that The method further comprises: Based on the initial queue identifier corresponding to the completion queue, determine a plurality of second data among the plurality of initial data, wherein the initial queue identifier is used to represent identification information of data that has been submitted, and the plurality of second data is used to represent data currently submitted to the host; Performing the aggregation operation on the plurality of second data to obtain a second aggregation result; The initial queue identifier is updated based on the second aggregation result.
6. The method according to claim 5, characterized in that Performing the aggregation operation on the plurality of second data to obtain a second aggregation result includes: Determine a second transmission queue corresponding to each second data; The second data corresponding to the same second transmission queue is obtained, and the aggregation operation is performed on the second data to obtain the second aggregation result.
7. The method according to claim 1, characterized in that The method further comprises: Get multiple raw data; Determine a third transmission queue corresponding to each original data; sorting the plurality of original data according to the third transmission queue to obtain sorted data, wherein the original data corresponding to the same third transmission queue are adjacent; The sorted data are written into the data buffer of the host in sequence.
8. A data processing device based on a paravirtualized device, characterized in that: include: A data acquisition module, used to acquire a plurality of initial data stored in the completion queue of the paravirtualized device, wherein the plurality of initial data are used to represent the description information of the original data that has been processed by the paravirtualized device but has not been submitted to the host, and the description information includes the address and length of the data; A data determination module, used to determine a plurality of first data satisfying a preset condition among the plurality of initial data; an aggregation module, configured to perform an aggregation operation on the plurality of first data to generate a first aggregation result, wherein the aggregation operation is used to represent concatenating the description information corresponding to the plurality of first data, and the concatenation includes at least one of the following: concatenating addresses in the description information corresponding to the plurality of first data, and summing the lengths in the description information corresponding to the plurality of first data; A sending module is used to send a direct memory access request carrying the first aggregation result to the memory of the host.
9. A data processing system based on a paravirtualized device, characterized in that: include: Host, including: memory and completion queue; The paravirtualized device is connected to the host and is used to obtain multiple initial data stored in the completion queue, wherein the multiple initial data are used to represent the description information of the original data that has been processed by the paravirtualized device but has not been submitted to the host, and the description information includes the address and length of the data; determine multiple first data that meet preset conditions in the multiple initial data; perform an aggregation operation on the multiple first data to generate a first aggregation result, wherein the aggregation operation is used to represent splicing the description information corresponding to the multiple first data, and the splicing includes at least one of the following: splicing the addresses in the description information corresponding to the multiple first data, and summing the lengths in the description information corresponding to the multiple first data; and sending a direct memory access request carrying the first aggregation result to the memory of the host.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the data processing method based on a semi-virtualized device according to any one of claims 1 to 7.
11. A computer terminal, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program, when running, executes the data processing method based on a semi-virtualized device as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method for optimizing network throughput in virtualization environment of embedded network
CN104615495A
Cited By
Data processing method, apparatus and system based on para-virtualization device
WO2023155698A1