Data transmission method and device and storage medium
By dividing the data into multiple data blocks and setting preset conditions and fence information in the descriptor, the problem of inefficient GPU processing when CE transmits big data is solved, pipeline processing is realized and the performance of the replication engine is improved.
Patent Information
- Application Number
- CN202510637405.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-16
AI Technical Summary
When the replication engine (CE) needs to transmit big data, the GPU of the target device needs to wait for the CE to pass all the data before starting processing, resulting in inefficient data processing.
By dividing the target data into multiple data blocks and setting preset conditions and fence information in the descriptor, the target data block is allowed to be processed or forwarded immediately after writing to the destination address, avoiding waiting for all data blocks to be written.
Pipeline processing between CE and hardware units or between multiple CEs is realized, which greatly improves the efficiency of data processing, and completes the transmission of multiple data blocks through one descriptor, improving the performance of the replication engine.
Smart Images

Figure CN120162286A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and in particular, to a data transmission method, apparatus, and storage medium. Background Art
[0002] In a computer system, a copy engine (CE) is a key component for processing data copy operations. As a direct memory access (DMA) controller in a graphics processing unit (GPU), through the CE, data can be copied from one device or the memory of the current device to the memory of the current device or other devices, and the data is transmitted to the GPU of the target device for subsequent processing. However, in the current technical solutions, when the CE needs to transmit a very large block of data, the GPU of the target device needs to wait for the CE to transfer all the data before it can start processing, resulting in low data processing efficiency. Summary of the Invention
[0003] In view of this, the present disclosure provides a data transmission method, apparatus, and storage medium.
[0004] According to an aspect of the present disclosure, a data transmission method is provided. The method can be used for a copy engine of a graphics processing unit (GPU), and the method includes:
[0005] In response to a target data block in the target data satisfying a preset condition, based on a descriptor corresponding to the target data, writing the target data block from a source address to a destination address, where the target data is composed of multiple data blocks, and the preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data;
[0006] In response to the target data block having been written from the source address to the destination address, writing a first preset value to a position associated with the target data block in the first fence information of the descriptor, where the first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit, and the first fence is a storage space of a preset size.
[0007] In a possible implementation, in response to a second preset value being written to a position associated with the target data block in the second fence information of the descriptor, it is determined that the target data block satisfies the preset condition, and the second fence is a storage space of a preset size.
[0008] In a possible implementation, the descriptor further includes one or more of the following parameters: source address information, destination address information, the size of the target data, and the size of each data block in the target data.
[0009] In a possible implementation, being forwarded to the target hardware unit includes being forwarded to the target hardware unit from the hardware unit where the destination address is located via other hardware units except the hardware unit where the destination address is located, or being directly forwarded to the target hardware unit from the hardware unit where the destination address is located.
[0010] In a possible implementation, the method further includes:
[0011] In response to the position associated with the target data block in the first fence information in the descriptor being written with a first preset value, writing the first preset value to the position associated with the target data block in the first fence corresponding to the destination address.
[0012] In a possible implementation, the method further includes:
[0013] In response to the target data having been processed by the hardware unit where the destination address is located or having been forwarded to other hardware units, writing a second preset value to the position associated with the next data block of the target data block in the second fence information of the descriptor, where each data block in the target data is written to the same destination address.
[0014] In a possible implementation, in response to the position associated with the next data block of the target data block in the second fence corresponding to the destination address being written with the second preset value, it is determined that the target data block has been processed by the hardware unit where the destination address is located or has been forwarded to other hardware units.
[0015] In a possible implementation, the size of each data block in the target data is related to the time for the hardware unit where the destination address is located to process the data block, so that the time for transmitting each data block matches the time for the hardware unit where the destination address is located to process each data block.
[0016] In a possible implementation, the hardware unit includes a graphics processing unit (GPU).
[0017] According to another aspect of the present disclosure, a data transmission device is provided. The device can be used in the copy engine of a graphics processing unit (GPU), and the device includes:
[0018] A first writing module, configured to, in response to a target data block in the target data meeting a preset condition, write the target data block from a source address to a destination address based on a descriptor corresponding to the target data, where the target data is composed of multiple data blocks, and the preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data;
[0019] A second writing module, configured to, in response to a target data block being written from a source address to a destination address, write a first preset value to a position associated with the target data block in the first fencing information of a descriptor, where the first preset value indicates that the target data block can be used to be processed by a hardware unit where the destination address is located or be forwarded to a target hardware unit, and the first fence is a storage space with a preset size.
[0020] In a possible implementation, in response to a second preset value being written to a position associated with the target data block in the second fencing information of the descriptor, it is determined that the target data block meets a preset condition, and the second fence is a storage space with a preset size.
[0021] In a possible implementation, the descriptor further includes one or more of the following parameters: source address information, destination address information, size of target data, and size of each data block in the target data.
[0022] In a possible implementation, being forwarded to the target hardware unit includes being forwarded from the hardware unit where the destination address is located to the target hardware unit via other hardware units except the hardware unit where the destination address is located, or being directly forwarded from the hardware unit where the destination address is located to the target hardware unit.
[0023] In a possible implementation, the apparatus further includes:
[0024] A third writing module, configured to, in response to a first preset value being written to a position associated with the target data block in the first fencing information in the descriptor, write the first preset value to a position associated with the target data block in the first fence corresponding to the destination address.
[0025] In a possible implementation, the apparatus further includes:
[0026] A fourth writing module, configured to, in response to the target data being processed by the hardware unit where the destination address is located or being forwarded to other hardware units, write a second preset value to a position associated with the next data block of the target data block in the second fencing information of the descriptor, where each data block in the target data is written to the same destination address.
[0027] In a possible implementation, in response to a second preset value being written to a position associated with the next data block of the target data block in the second fence corresponding to the destination address, it is determined that the target data block has been processed by the hardware unit where the destination address is located or has been forwarded to other hardware units.
[0028] In a possible implementation, the size of each data block in the target data is related to the time for the hardware unit where the destination address is located to process the data block, so that the time for transmitting each data block matches the time for the hardware unit where the destination address is located to process each data block.
[0029] In one possible implementation, the hardware unit includes a Graphics Processing Unit (GPU).
[0030] According to another aspect of the present disclosure, there is provided a data transmission device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.
[0031] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0032] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0033] According to an embodiment of the present application, by responding that a target data block in target data meets a preset condition, based on a descriptor corresponding to the target data, the target data block is written from a source address to a destination address. The preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data. In response to the target data block having been written from the source address to the destination address, a first preset value is written to a position associated with the target data block in the first fence information of the descriptor. The first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to a target hardware unit. The first fence is a storage space of a preset size, which enables the target data block to perform subsequent operations without waiting for all data blocks in the target data to be written to the destination address. After the target data block is written to the destination address, the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to a target hardware unit, so that a pipeline can be formed between the CE and the hardware unit or between multiple CEs, greatly improving the data processing efficiency. Moreover, each time the target data is transmitted, only one descriptor is required to complete the transmission of multiple data blocks, improving the performance of the CE.
[0034] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.
[0036] Figure 1 A schematic diagram showing an application scenario of an embodiment of the present application.
[0037] Figure 2 The flowchart showing the data transmission method according to an embodiment of the present application.
[0038] Figure 3 The schematic diagram showing the structure of a descriptor according to an embodiment of the present application.
[0039] Figure 4 The schematic diagram showing the first fence and the second fence according to an embodiment of the present application.
[0040] Figure 5 The structural diagram showing the data transmission device according to an embodiment of the present application.
[0041] Figure 6 It is a block diagram of a device 1900 for data transmission shown according to an exemplary embodiment. Detailed implementation manners
[0042] The following will describe in detail various exemplary embodiments, features and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0043] The special term "exemplary" herein means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.
[0044] In addition, for better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0045] In a computer system, a copy engine (CE) is a key component for processing data copy operations. As a direct memory access (DMA) controller in a graphics processing unit (GPU), through the CE, data can be copied from one device or the memory of the current device to the memory of the current device or other devices, and the data is transmitted to the GPU of the target device for subsequent processing. However, in the current technical solutions, when the CE needs to transmit a very large block of data, the GPU of the target device needs to wait for the CE to transfer all the data before starting to process, resulting in low data processing efficiency.
[0046] In view of this, the present application provides a data transmission method, apparatus, and storage medium. The method of the embodiments of the present application can be used for the CE on the GPU, and the target data can be divided into multiple data blocks to implement fragmented transmission of the data. In the present application, in response to the target data block in the target data satisfying a preset condition, based on the descriptor corresponding to the target data, the target data block is written from the source address to the destination address. The preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data. By responding that the target data block has been written from the source address to the destination address, a first preset value is written to the position associated with the target data block in the first fence information of the descriptor. The first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit. The first fence is a storage space of a preset size, which enables the target data block to perform subsequent operations without waiting for all data blocks in the target data to be written to the destination address. After the target data block is written to the destination address, it can be used to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit, so that a pipeline can be formed between the CE and the hardware unit or between multiple CEs, greatly improving the data processing efficiency. Moreover, each time the target data is transmitted, only one descriptor is required to complete the transmission of multiple data blocks, improving the performance of the CE.
[0047] Figure 1 Schematic diagram showing an application scenario of an embodiment of the present application. As Figure 1 shown, the present application can be applied to the scenario where the CE moves data from the memory of device 1 to the memory of device 2 and then moves the data from the memory of device 2 to the video memory of device 2. Among them, the CE can be on device 2. By moving the data from the memory of device 1 to the video memory of device 2, the hardware unit (such as the GPU) on device 2 can process the data.
[0048] In this application scenario, the data is moved twice. The method of the embodiments of the present application can be used to fragment the data into multiple data blocks. The multiple data blocks are first sequentially moved from the memory of device 1 to the memory of device 2. Among them, whenever a data block is moved from the memory of device 1 to the memory of device 2, based on the method of the embodiments of the present application, it is not necessary to wait for other data blocks to complete the move, and the data block can be continuously moved from the memory of device 2 to the video memory of device 2. Whenever a data block is moved from the memory of device 2 to the video memory of device 2, based on the method of the embodiments of the present application, it is not necessary to wait for other data blocks to complete the move, and the GPU on device 2 can process the data block. Thus, a pipeline process is formed for each data block in the data, greatly improving the efficiency of processing the data.
[0049] It should be noted that in the above application scenarios, two data transfers are used as an example, and it is also possible to perform only one data transfer (such as only transferring data from the memory of device 2 to the video memory of device 2), or perform more than two data transfers. This application does not limit this.
[0050] The above-mentioned device 1 and device 2 can be terminal devices or servers. Among them, the terminal device can be any one or more of a mobile phone, a foldable electronic device, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), and a vehicle-mounted device. The specific type of the terminal device in the embodiments of this application is not specially limited and can have wired or wireless communication functions.
[0051] The server can be located locally or in the cloud, and can be a physical device or a virtual device, such as a virtual machine, a container, etc., and has a wireless communication function. Among them, the wireless communication function can be set in the chip (system) or other components or assemblies of the server. The wireless communication function can be realized, for example, through mobile communication technologies such as 2G / 3G / 4G / 5G, as well as Wi-Fi, Bluetooth, frequency modulation (FM), data transmission radio, satellite communication, etc. It can also communicate through a wired connection to achieve interaction with other devices.
[0052] The data transmission method of the embodiments of this application will be introduced below.
[0053] Figure 2 The flowchart showing the data transmission method according to an embodiment of this application. This method can be used in the copy engine CE of the GPU, such as Figure 2 As shown, this method can include:
[0054] Step S201, in response to a target data block in the target data meeting a preset condition, write the target data block from the source address to the destination address based on the descriptor corresponding to the target data.
[0055] Among them, the target data is composed of multiple data blocks, and the target data can be divided into multiple data blocks of the same size. The preset condition can include that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data.
[0056] Among them, the descriptor corresponding to the target data can be filled in by a driver, and is used to move the target data from the source address (which can be referred to as SrcAddress) to the destination address (which can be referred to as DstAddress). The source address and the destination address can be the memory address and the video memory address on the same device respectively. For example, the source address is the memory address of the above-mentioned device 2, and the destination address is the video memory address of the above-mentioned device 2; the source address and the destination address can also be the memory address / video memory address on different devices respectively. For example, the source address is the memory address of the above-mentioned device 1, and the destination address is the memory address of the above-mentioned device 2.
[0057] Figure 3 A schematic diagram showing the structure of a descriptor according to an embodiment of the present application. As Figure 3 shown, the parameters of the descriptor may include: first fence (which can be referred to as SignalFence) information and second fence (which can be referred to as WaitFence) information. The descriptor may also include one or more of the following parameters: source address information, destination address information, the size of the target data (which can be referred to as DataSize), and the size of each data block in the target data (which can be referred to as SegmentSize).
[0058] Among them, the first fence information and the second fence information can respectively represent two arrays, and the length of the array can be equal to the total number of data blocks in the target data (i.e., (DataSize + SegmentSize - 1) / SegmentSize). Each position in the array can be associated with each data block in the target data respectively. The second fence information can be used to indicate whether the data block can start to be written from the source address to the destination address, and the first fence information can be used to indicate whether the data block has been written from the source address to the destination address. When a specific position in the array corresponding to the second fence information is written with a second preset value (for example, 1), it can indicate that the data block corresponding to this position can start to be transmitted, that is, it can start to be written from the source address to the destination address; when a specific position in the array corresponding to the first fence information is written with a first preset value (the first preset value can be the same as the second preset value, for example, 1), it can indicate that the data block corresponding to this position can be used for processing by the hardware unit where the destination address is located.
[0059] It can be determined that the target data block meets the preset condition in response to a specific position in the second fence information of the descriptor associated with the target data block being written with the second preset value. The driver can write the second preset value into the descriptor, and the second preset value indicates that the target data block can start to be written from the source address to the destination address.
[0060] Among them, the second fence information can be associated with a second fence, which can be set in the memory or video memory where the destination address is located, and is a storage space of a preset size. The size of the storage space can be associated with the target data and can be a storage space in the video memory or memory. Refer to Figure 4 , which shows a schematic diagram of a first fence and a second fence according to an embodiment of the present application. As Figure 4 shown, for the second fence, reference can be made to the second fence 1 (WaitFence1) and the second fence 0 (WaitFence0) in the figure, which are respectively associated with the data block 1 (data1) and the data block 0 (data0) representing data blocks.
[0061] The second fence may include storage spaces respectively associated with each data block. In response to the position associated with the target data block in the second fence information in the descriptor being written with a second preset value, the second preset value can be written to the position associated with the target data block in the second fence corresponding to the destination address. Thus, the value of the second fence in the actual storage space can be synchronized with the second fence information in the descriptor.
[0062] Step S202: In response to the target data block having been written from the source address to the destination address, write a first preset value to the position associated with the target data block in the first fence information of the descriptor.
[0063] The position associated with the target data block in the first fence information may be the position associated with the target data block in the array corresponding to the first fence information. The first preset value may indicate that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit.
[0064] The first fence information can be associated with a first fence, which can be set in the memory or video memory where the destination address is located, and is a storage space of a preset size. The size of the storage space can be associated with the target data and can be a storage space in the video memory or memory. For an example of the first fence, reference can be made to Figure 4 the first fence 1 (SignalFence1) and the first fence 0 (SignalFence0) therein, which are respectively associated with the data block 1 (data1) and the data block 0 (data0).
[0065] In response to the position associated with the target data block in the first fence information in the descriptor being written with the first preset value, the first preset value can be written to the position associated with the target data block in the first fence corresponding to the destination address. Thus, the value of the first fence in the actual storage space can be synchronized with the first fence information in the descriptor.
[0066] The above-mentioned hardware unit may include a Graphics Processing Unit (GPU). The hardware unit where the destination address is located may be the GPU in the device where the destination address is located. After it obtains the position associated with the target data block in the first fence and writes the first preset value, it can access the target data block to process the target data block. Such processing may be, for example, performing image rendering, image post-processing, or computing tasks for deep learning inference, etc.
[0067] Among them, the size of each data block in the target data may be related to the time for the hardware unit where the destination address is located to process the data block, so that the time for transmitting each data block matches the time for the hardware unit where the destination address is located to process each data block. For example, through the preset data block size, the time for transmitting each data block can be made as close as possible to the time for the hardware unit where the destination address is located to process each data block, so that the transmission of data blocks and the processing of data blocks by the hardware unit can form a pipeline.
[0068] After writing the target data block to the destination address, the hardware unit where the destination address is located may not process the target data block, but forward the target data block to a target hardware unit other than the hardware unit where the destination address is located, so that the target hardware unit processes the target data block. At this time, the first preset value may indicate that the target data block can be used for forwarding to the target hardware unit. Being forwarded to the target hardware unit may include forwarding from the hardware unit where the destination address is located to the target hardware unit via other hardware units except the hardware unit where the destination address is located, or directly forwarding from the hardware unit where the destination address is located to the target hardware unit.
[0069] At this time, the device driver corresponding to the hardware unit where the destination address is located may fill in the descriptor, and data transmission is performed according to the method in the above steps S201 - S202. When the SignalFence corresponding to the destination address is written with the first preset value, the WaitFence corresponding to the address of the forwarded hardware unit can be written with the second preset value, so that the target data block can be written into the memory / video memory of the forwarded hardware unit.
[0070] The above forwarding may be one or more times, that is, the target data block can be processed by the target hardware unit after one forwarding, or after multiple forwardings among multiple hardware units, it can be processed by the target hardware unit written by the last forwarding. The process of transmitting data across hardware units (i.e., across chips) can be implemented by using chip interconnection technologies such as the Peripheral Component Interconnect Express (PCIe) or Chip to Chip (C2C). Thus, by using the first fence and the second fence in the present application, it is beneficial to achieve pipeline management among the CEs of multiple chips and improve the processing efficiency.
[0071] According to an embodiment of the present application, by responding that a target data block in target data meets a preset condition, based on a descriptor corresponding to the target data, the target data block is written from a source address to a destination address. The preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data. In response to the target data block having been written from the source address to the destination address, a first preset value is written to a position associated with the target data block in the first fence information of the descriptor. The first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to a target hardware unit. The first fence is a storage space of a preset size, which enables the target data block to perform subsequent operations without waiting for all data blocks in the target data to be written to the destination address, and makes the target data block available to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit, so that a pipeline can be formed between the CE and the hardware unit or between multiple CEs, greatly improving the data processing efficiency. Moreover, each time the target data is transmitted, only one descriptor is required to complete the transmission of multiple data blocks, improving the performance of the CE.
[0072] The method may further include:
[0073] In response to the target data having been processed by the hardware unit where the destination address is located or having been forwarded to other hardware units, a second preset value is written to a position associated with the next data block of the target data block in the second fence information of the descriptor.
[0074] In response to the second preset value being written to the position associated with the next data block of the target data block in the second fence information of the descriptor, the second preset value may be written to the position associated with the next data block of the target data block in the second fence corresponding to the destination address.
[0075] When the second preset value is written to the position associated with the next data block of the target data block in the second fence corresponding to the destination address, it can be determined that the target data block has been processed by the hardware unit where the destination address is located or has been forwarded to other hardware units. Before the next data block of the target data block is written from the source address to the destination address, the second preset value has been written to the positions associated with all the data blocks before this data block in the second fence information of the descriptor, and the first preset value has been written to the positions associated with all the data blocks before this data block in the first fence information of the descriptor. Thus, the orderly transmission of each data block in the target data can be ensured.
[0076] Among them, each data block in the target data can be written to the same destination address. Thus, while saving storage space, pipeline control can be achieved using the second fence information, improving the processing efficiency.
[0077] Figure 5The structural diagram of a data transmission device according to an embodiment of the present application is shown. This device can be used in the copy engine of a graphics processing unit (GPU), such as Figure 5 As shown, the device includes:
[0078] A first writing module 501, configured to, in response to a target data block in target data satisfying a preset condition, write the target data block from a source address to a destination address based on a descriptor corresponding to the target data. The target data is composed of multiple data blocks, and the preset condition includes that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data;
[0079] A second writing module 502, configured to, in response to the target data block having been written from the source address to the destination address, write a first preset value to a position associated with the target data block in the first fence information of the descriptor. The first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or be forwarded to a target hardware unit. The first fence is a storage space with a preset size.
[0080] In a possible implementation manner, in response to a second preset value being written to a position associated with the target data block in the second fence information of the descriptor, it is determined that the target data block satisfies the preset condition. The second fence is a storage space with a preset size.
[0081] In a possible implementation manner, the descriptor further includes one or more of the following parameters: source address information, destination address information, the size of the target data, and the size of each data block in the target data.
[0082] In a possible implementation manner, being forwarded to the target hardware unit includes being forwarded from the hardware unit where the destination address is located to the target hardware unit via other hardware units except the hardware unit where the destination address is located, or being directly forwarded from the hardware unit where the destination address is located to the target hardware unit.
[0083] In a possible implementation manner, the device further includes:
[0084] A third writing module, configured to, in response to a first preset value being written to a position associated with the target data block in the first fence information in the descriptor, write the first preset value to a position associated with the target data block in the first fence corresponding to the destination address.
[0085] In a possible implementation manner, the device further includes:
[0086] A fourth writing module, configured to, in response to the target data having been processed by the hardware unit where the destination address is located or having been forwarded to other hardware units, write a second preset value to a position associated with the next data block of the target data block in the second fence information of the descriptor, where each data block in the target data is written to the same destination address.
[0087] In a possible implementation, in response to a second preset value being written to a position associated with the next data block of the target data block in the second fence corresponding to the destination address, it is determined that the target data block has been processed by the hardware unit where the destination address is located or has been forwarded to other hardware units.
[0088] In a possible implementation, the size of each data block in the target data is related to the time for the hardware unit where the destination address is located to process the data block, so that the time for transmitting each data block matches the time for the hardware unit where the destination address is located to process each data block.
[0089] In a possible implementation, the hardware unit includes a Graphics Processing Unit (GPU).
[0090] According to the embodiments of the present application, by responding to the target data block in the target data meeting the preset conditions, based on the descriptor corresponding to the target data, the target data block is written from the source address to the destination address. The preset conditions include that the previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data. In response to the target data block having been written from the source address to the destination address, a first preset value is written to the position associated with the target data block in the first fence information of the descriptor. The first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or be forwarded to the target hardware unit. The first fence is a storage space with a preset size, which enables the target data block to perform subsequent operations without waiting for all the data blocks in the target data to be written to the destination address after the target data block is written to the destination address, and enables the target data block to be used to be processed by the hardware unit where the destination address is located or be forwarded to the target hardware unit, so that a pipeline can be formed between the CE and the hardware unit or between multiple CEs, greatly improving the data processing efficiency, and only one descriptor is required to complete the transmission of multiple data blocks each time the target data is transmitted, improving the performance of the CE.
[0091] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0092] The embodiments of the present disclosure further provide a data transmission device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0093] The embodiments of the present disclosure further provide a non - volatile computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0094] Embodiments of the present disclosure also provide a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0095] Figure 6 FIG. 1900 is a block diagram of an apparatus 1900 for data transmission shown according to an exemplary embodiment. For example, the apparatus 1900 may be provided as a server or a terminal device. Referring to Figure 6 , the apparatus 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0096] The apparatus 1900 may also include a power component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0097] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions. The above computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method.
[0098] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0099] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0100] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0101] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0102] Aspects of the present disclosure are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.
[0103] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0104] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0105] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0106] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A data transmission method, characterized in that: The method is used for a copy engine of a graphics processing unit (GPU), and the method comprises: In response to a target data block in the target data satisfying a preset condition, based on a descriptor corresponding to the target data, writing the target data block from a source address to a destination address, the target data being composed of a plurality of data blocks, the preset condition including that a previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data; In response to the target data block having been written from the source address to the destination address, a first preset value is written to a position associated with the target data block in the first fence information of the descriptor, wherein the first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or to be forwarded to the target hardware unit, and the first fence is a storage space of a preset size.
2. The method according to claim 1, characterized in that In response to a second preset value being written into a position associated with the target data block in the second fence information of the descriptor, it is determined that the target data block meets the preset condition, and the second fence is a storage space of a preset size.
3. The method according to claim 2, characterized in that The descriptor further includes one or more parameters of the following: source address information, destination address information, the size of the target data, and the size of each data block in the target data.
4. The method according to claim 1, characterized in that The forwarding to the target hardware unit includes forwarding from the hardware unit where the destination address is located to the target hardware unit via other hardware units except the hardware unit where the destination address is located, or forwarding from the hardware unit where the destination address is located directly to the target hardware unit.
5. The method according to claim 1 or 2, characterized in that: The method further comprises: In response to a first preset value being written into a position associated with the target data block in the first fence information in the descriptor, the first preset value is written into a position associated with the target data block in the first fence corresponding to the destination address.
6. The method according to claim 2, characterized in that The method further comprises: In response to the target data having been processed by the hardware unit where the destination address is located or having been forwarded to other hardware units, a second preset value is written to a position associated with the next data block of the target data block in the second fence information of the descriptor, wherein each data block in the target data is written to the same destination address.
7. The method according to claim 6, characterized in that In response to a second preset value being written into a position associated with a next data block of the target data block in a second fence corresponding to the destination address, it is determined that the target data block has been processed by the hardware unit where the destination address is located or has been forwarded to other hardware units.
8. The method according to claim 1, characterized in that The size of each data block in the target data is related to the time taken by the hardware unit at the destination address to process the data block, so that the time taken to transmit each data block matches the time taken by the hardware unit at the destination address to process each data block.
9. The method according to claim 1, characterized in that: The hardware unit includes a graphics processor GPU.
10. A data transmission device, characterized in that: The device is used for a copy engine of a graphics processor GPU, and the device comprises: A first writing module, configured to write the target data block from a source address to a destination address based on a descriptor corresponding to the target data in response to a target data block in the target data satisfying a preset condition, wherein the target data is composed of a plurality of data blocks, and the preset condition includes that a previous data block of the target data block has been written to the destination address or the target data block is the first data block in the target data; A second writing module is used to write a first preset value to a position associated with the target data block in the first fence information of the descriptor in response to the target data block having been written from the source address to the destination address, wherein the first preset value indicates that the target data block can be used to be processed by the hardware unit where the destination address is located or forwarded to the target hardware unit, and the first fence is a storage space of a preset size.
11. A data transmission device, comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.
12. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product, comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Data move engine to move a block of data
CN105408874A
Method, device, and system for executing writing task
CN105528371A
Data transmission method and device
CN118331903A
Data writing control method and system, terminal and medium
CN119311229A
Using a tree-based data structure to map logical addresses to physical addresses on a storage device
US20170123665A1