Data transmission method and device, storage medium and electronic equipment
By storing the target data to be sent in the target memory of the computing platform, and sending notification messages to the receiver based on the Ethernet communication link, and using the direct memory access communication link for data transmission, the problem of high delay and high computing power consumption of cross-chip system transmission based on Ethernet is solved, and efficient and real-time data exchange is achieved.
Patent Information
- Application Number
- CN202412000508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
The cross-chip system-on-chip data transmission is high and the computing power consumption is large, which cannot meet the needs of scenarios such as high latency requirements such as autonomous driving.
The target data to be sent is stored in the target direct memory access memory of the computing platform, and a notification message is sent to the receiver based on the Ethernet communication link, instructing it to obtain data from the target direct memory access memory, and use the direct memory access communication link for data transmission.
It improves the efficiency of data transmission, reduces the load of the CPU, optimizes resource utilization, enhances the robustness of communication and system stability, and meets the real-time data exchange requirements in high-demand scenarios such as autonomous driving.
Smart Images

Figure CN120045507A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a data transmission method, apparatus, storage medium, and electronic device. Background Art
[0002] The Peripheral Component Interconnect Express (PCIE) is mainly used to connect motherboard peripheral devices, and has characteristics such as high bandwidth, low latency, and low overhead. The Data Distribution Service (DDS) is a highly reliable and efficient distributed communication technology that adopts a publish / subscribe model and provides real-time, efficient, and flexible data communication services. Due to the increasing demand for chip computing power in fields such as autonomous driving, the System on a Chip (SOC) alone can no longer meet the actual computing power requirements. Computing devices based on multiple SOCs can significantly improve computing performance, but cross-SOC communication has the defects of high latency and low efficiency. Since DDS mainly transmits data based on Ethernet, although it can achieve reliable and efficient communication services within a single device, for the scenario of large data transmission across SOCs, there are many drawbacks and it cannot meet the requirements of scenarios with high latency requirements such as autonomous driving and robotics. Specifically, it can be divided into the following three aspects:
[0003] 1. The latency of cross-SOC data transmission based on Ethernet is relatively high, up to dozens of milliseconds or even hundreds of milliseconds, making it difficult to meet the service scenarios with low latency requirements; 2. When transmitting large data based on Ethernet, the CPU is required to parse the data. As the data volume increases, the CPU computing power consumption also increases; 3. The large data transmission across SOCs is often image-type data. The processing and use of this type of data are usually at the Graphics Processing Unit (GPU) end. Using DDS for transmission will cause the transmission link (GPU-CPU-DDS-CPU-GPU) to be too long, resulting in waste of resources. Considering the above three points, the current DDS method has high latency and high computing power consumption in large data cross-SOC communication.
[0004] In view of the problem of low transmission efficiency of cross-system-on-chip data transmission based on Ethernet in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of the present application provide a data transmission method, apparatus, storage medium, and electronic device to at least solve the problem of low transmission efficiency of cross-system-on-chip data transmission based on Ethernet.
[0006] According to an embodiment of the present application, a data transmission method is provided, which is applied to a first system-on-chip among multiple system-on-chips in a computing platform. The multiple system-on-chips further include a second system-on-chip, and the method includes: when target data is to be sent to the second system-on-chip, storing the target data in a target direct memory access memory in the computing platform, where the target direct memory access memory is a direct memory access memory corresponding to the first system-on-chip and the second system-on-chip; and sending a notification message to the second system-on-chip based on an Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0007] In an exemplary embodiment, the storing the target data in the target direct memory access memory in the computing platform includes: storing first data to be sent by a central processing unit in the first system-on-chip in the target direct memory access memory in the computing platform; and / or storing second data to be sent by a graphics processing unit in the first system-on-chip in the target direct memory access memory in the computing platform, where the data in the target direct memory access memory is to be obtained by a central processing unit and / or an image processing unit in the second system-on-chip, and the target data includes the first data and / or the second data.
[0008] In an exemplary embodiment, the storing the target data in the target direct memory access memory in the computing platform includes: serializing the target data into data of a target type using a sequence number interface corresponding to the data type of the target data, where the target type of data is allowed to be stored in the direct memory access memory; and writing the target type of data into the target direct memory access memory.
[0009] In an exemplary embodiment, after sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, the method further includes: monitoring the Ethernet interface of the first system-on-chip to determine whether a target identifier sent by the second system-on-chip through the Ethernet communication link is obtained, where the target identifier is an identifier generated before the first system-on-chip stores the target data into the target direct memory access memory in the computing platform, the notification message carries the target identifier, and the second system-on-chip sends the target identifier to the first system-on-chip through the Ethernet communication link after obtaining the target data from the target direct memory access memory; in the case of obtaining the target identifier sent by the second system-on-chip through the Ethernet communication link within a predetermined time, determining that the direct memory access communication link between the first system-on-chip and the second system-on-chip is normal, where the direct memory access communication link is a link through which the first system-on-chip and the second system-on-chip communicate through the target direct memory access memory; in the case of not obtaining the target identifier sent by the second system-on-chip through the Ethernet communication link within the predetermined time, determining that the direct memory access communication link between the first system-on-chip and the second system-on-chip is abnormal, switching the data transmission link between the first system-on-chip and the second system-on-chip to the Ethernet communication link, and sending the target data to the second system-on-chip based on the Ethernet communication link; continuously monitoring the link state of the direct memory access communication link, and in the case where the link state of the direct memory access communication link becomes a normal state, sending a recovery message to the second system-on-chip based on the Ethernet communication link, where the recovery message is used to indicate to the second system-on-chip that the data transmission link between the first system-on-chip and the second system-on-chip is switched back to the direct memory access communication link.
[0010] In an exemplary embodiment, the computing platform further includes a target central processing unit. Before storing the target data into the target direct memory access memory in the computing platform, the method further includes: determining the target direct memory access memory designated by the target central processing unit; where the target direct memory access memory is a direct memory access memory among the M direct memory access memories of the computing platform, the data volume allowed to be stored in which is greater than or equal to the data volume of the target data, and the difference between the data volume allowed to be stored and the target data volume is less than or equal to a preset threshold. The M direct memory access memories have a one-to-one correspondence with M processes in the computing platform. The M processes include a target process, and the target process is used to instruct the first system-on-chip to send the target data to the second system-on-chip. M is an integer greater than or equal to 2.
[0011] In an exemplary embodiment, the computing platform further includes a target central processing unit. After sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, the method further includes: determining whether the second system-on-chip successfully obtains the target data from the target direct memory access memory, and determining whether there is data to be sent to the second system-on-chip; in a case where it is determined that the second system-on-chip successfully obtains the target data from the target direct memory access memory and there is no data to be sent to the second system-on-chip, sending indication information to the target central processing unit, where the indication information is used to instruct the target central processing unit to release the target direct memory access memory and mark the state of the target direct memory access memory as an idle state.
[0012] In an exemplary embodiment, storing the target data in the target direct memory access memory in the computing platform includes: in a case where the target data includes N sub-data, dividing the target direct memory access memory into N sub-direct memory access memories according to the data amounts of the N sub-data, where N is a positive number greater than or equal to 2; storing the N sub-data correspondingly in the N sub-direct memory access memories.
[0013] According to another embodiment of the present application, there is also provided a data transmission device, which is applied to a first system-on-chip among a plurality of system-on-chips in a computing platform, and the plurality of system-on-chips further includes a second system-on-chip; including: a storage module, configured to store the target data in the target direct memory access memory in the computing platform in a case where the target data is to be sent to the second system-on-chip, where the target direct memory access memory is a direct memory access memory corresponding to the first system-on-chip and the second system-on-chip; a sending module, configured to send a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0014] According to still another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0015] According to still another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0016] According to another embodiment of the present application, there is also provided a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0017] With the present application, when a first system-on-chip in a plurality of system-on-chips in a computing platform is to send target data to a second system-on-chip, the target data is stored in the target direct memory access memory of the computing platform, and a notification message is sent to the second system-on-chip based on an Ethernet communication link, thereby instructing the second system-on-chip to obtain the target data from the target direct memory access memory. Since the data transmission is performed using the direct memory access communication link between the first system-on-chip and the second system-on-chip, the data transmission efficiency is improved, and the problem of low transmission efficiency of cross-system-on-chip data transmission based on Ethernet is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a hardware structure block diagram of a server device for a data transmission method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a data transmission method according to an embodiment of the present application;
[0021] Figure 3 is a schematic architecture diagram of a data processing module according to an embodiment of the present application;
[0022] Figure 4 is a schematic architecture diagram of a memory sharing module according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of a memory partition according to an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of a status monitoring according to an embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a data transmission effect according to an embodiment of the present application;
[0026] Figure 8 is a block diagram of a structure of a data transmission device according to an embodiment of the present application;
[0027] Figure 9It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0028] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.
[0030] The embodiment of the data transmission method provided in the embodiment of the present application can be executed in a server device or a similar computing device. Taking the operation on the server device as an example, Figure 1 It is a hardware structure block diagram of a server device of a data transmission method according to an embodiment of the present application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor (abbreviated as MPU) or a field-programmable gate array (abbreviated as FPGA)) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a configuration different from Figure 1 shown in the figure.
[0031] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the data transmission method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the server device through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0032] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of a server device. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0033] To solve the above problems, a data transmission method is provided in this embodiment, which is applied to the first system-on-chip among multiple system-on-chips in a computing platform, and the multiple system-on-chips further include a second system-on-chip; as Figure 2 shown, the method includes the following steps S202 - S204:
[0034] Step S202: When the target data is to be sent to the second system-on-chip, store the target data in the target direct memory access memory in the computing platform, where the target direct memory access memory is the direct memory access memory corresponding to the first system-on-chip and the second system-on-chip;
[0035] That is to say, the first SOC (i.e., the first system-on-chip) and the second SOC (i.e., the second system-on-chip) can realize cross-SOC data transmission through the direct memory access communication link corresponding to the corresponding target DMA memory (i.e., the target direct memory access memory).
[0036] It should be noted that the Direct Memory Access (DMA) technology can directly transfer data from the memory of one computer to another computer without the intervention of the Central Processing Unit (CPU).
[0037] It should be noted that since the DMA operation is directly carried out between the memory and the device without the parsing and scheduling of the CPU, the load of the CPU is reduced, the resource allocation of the computing platform is optimized, and the data transmission efficiency is improved.
[0038] Step S204: Send a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0039] It should be noted that when the target data has been stored in the target direct memory access memory, the first SOC needs to notify the second SOC of the data availability. In the embodiment of the present application, through the publish-subscribe mechanism of DDS and using the established Ethernet communication link, a notification message is sent to the second SOC, so that the second SOC can read the data located in the target direct memory access memory when subscribing to the notification message, without an additional query or request process. The real-time communication feature of DDS ensures the low latency and high reliability of this operation. The publish-subscribe mode of DDS allows for broadcast-style data notifications among multiple SOCs, making the communication process more flexible, easy to expand and maintain. This model is particularly suitable for application scenarios such as autonomous driving that require real-time data sharing and processing.
[0040] It should be noted that the first SOC publishes a notification message based on the Ethernet communication link, enabling the second SOC to quickly locate the data position and directly read it through DMA, avoiding CPU resource consumption and network latency, thereby achieving efficient and real-time data exchange in cross-SOC large data transmission. At the same time, this combination method also enhances the robustness of the communication architecture, providing a more stable data transmission guarantee for applications such as autonomous driving and robotics.
[0041] In the above steps S202 - S204, when the first system-on-chip in the multiple system-on-chips in the computing platform is to send target data to the second system-on-chip, the target data is stored in the target direct memory access memory of the computing platform, and a notification message is sent to the second system-on-chip based on the Ethernet communication link, thereby instructing the second system-on-chip to obtain the target data from the target direct memory access memory. Since the direct memory access communication link between the first system-on-chip and the second system-on-chip is used for data transmission, the data transmission efficiency is improved, and the problem of low transmission efficiency of cross-system-on-chip data transmission based on Ethernet is solved.
[0042] In an exemplary embodiment, storing the target data in the target direct memory access memory of the computing platform is implemented by including the following step S11 and / or step S12:
[0043] Step S11: Store the first data to be sent by the central processing unit in the first system-on-chip in the target direct memory access memory of the computing platform;
[0044] It should be noted that as Figure 4As shown, the first System on Chip (SOC) writes the data to be sent from the Central Processing Unit (CPU) in the first SOC into the DMA memory, which can achieve more efficient and low-latency data transmission to the CPU and / or Graphics Processing Unit (GPU) in the second SOC, avoiding the extra copy of data between the CPU and the network interface or the GPU, significantly reducing the latency of data processing, and improving the overall communication efficiency of the system.
[0045] Step S12: Store the second data to be sent from the graphics processing unit in the first system on chip in the target direct memory access memory in the computing platform, where the data in the target direct memory access memory is to be acquired by the central processing unit and / or the image processor in the second system on chip, and the target data includes: the first data and / or the second data.
[0046] It should be noted that for a large amount of data generated by the GPU, such as images, point clouds, or intermediate results of deep learning models, the first SOC stores the data to be sent from the GPU in the target direct memory access memory, which can achieve more efficient and low-latency data transmission to the CPU and / or GPU in the second SOC. Directly using the DMA memory for transmission avoids the performance loss caused by frequent copying of data between the GPU and the CPU, and realizes high-performance exchange of GPU data.
[0047] It should be noted that in this embodiment, the DMA memory uses the Unified Memory (UM) architecture. Optionally, as Figure 4 shown, the unified memory architecture in a single device realizes memory sharing between the CPU and the GPU. For the DMA memory sharing result of two devices in this solution, due to its unified memory characteristics, memory sharing between the CPUs and GPUs on the two devices can be realized. Two devices, Host A (with SOC1, that is, the first system on chip) and Host B (with SOC2, that is, the second system on chip), share memory through DMA, and Host A and Host B respectively implement the UM unified memory in this shared memory area, that is, the GPU and CPU in Host A can directly achieve zero-copy without going through the memory copy process, and the same conclusion applies to Host B. Based on the above results, Host A and Host B can both achieve shared memory between CPUs, shared memory between GPUs, and memory sharing between the CPU of Host A and the GPU of Host B, realizing zero-copy of CPU / GPU memory between the two devices.
[0048] It should be noted that through the above steps, the CPU and / or GPU in the second SOC can read the stored data of the CPU and / or GPU in the first SOC, enabling zero-copy data transfer between different SOCs, greatly reducing data transfer latency, lowering the CPU load, optimizing resource utilization, and improving communication robustness and system stability. This has significant technical effects and application value for achieving efficient and real-time cross-SOC data exchange, especially in high-demand scenarios such as autonomous driving.
[0049] In an exemplary embodiment, storing the target data in the target direct memory access memory in the computing platform can be achieved through the following steps S21 - S22:
[0050] Step S21: Serialize the target data into data of a target type using a sequence number interface corresponding to the data type of the target data, where the target type of data is allowed to be stored in the direct memory access memory.
[0051] Optionally, as Figure 3 shown, the application layer contains data such as images and point clouds. The target data first enters the DMA domain, where it is serialized in the DMA domain and then written into the memory set by the DMA.
[0052] It should be noted that the data types for big data communication across SOCs are often data types such as images and point clouds. Since the data type required in the DMA memory is unit_8, the above data types cannot be directly written into the DMA memory. For different data types, a unified interface for serialization and deserialization needs to be provided to achieve writing data into the DMA memory and reading data from the DMA memory.
[0053] Optionally, the embodiment of the present application provides a data serialization interface for the data type of Robot Operating System 2 (ROS2). This interface realizes the conversion of the ROS2 data type to the unit_8 type, enabling direct storage of image (Image) and point cloud (PointCloud) data types in ROS2 into the memory. In addition to the serialization interface, data restoration can also be performed, that is, deserializing the data in the unit_8 format into the image (Image) and point cloud (PointCloud) data types.
[0054] It should be noted that the serialization interface ensures that all types of data can be converted into a unified format supported by DMA memory, which simplifies the data processing process and enables different data types to be efficiently transmitted under the same communication architecture. By using the serialization interface specifically for ROS2 data types (such as Image and PointCloud), the data conversion process becomes efficient and fast, reducing the data processing time and improving the overall communication efficiency. In addition, the design of the serialization interface takes into account possible future-added data types, which means that new data formats can be easily integrated into the existing communication architecture, enhancing the system's compatibility and scalability.
[0055] Step S22: Write the data of the target type into the target direct memory access memory.
[0056] It should be noted that steps S21 and S22, through the serialization interface and the DMA mechanism, achieve efficient and low-latency data transmission across SOCs. This not only improves the communication efficiency, reduces the CPU burden, but also enhances the system's compatibility, scalability, and communication reliability. In the autonomous driving system, it can significantly reduce the data transmission latency, improve the real-time performance and stability of the system, and is an effective way to achieve high-performance data communication.
[0057] In an exemplary embodiment, after sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, the method further includes the following steps S31-S34:
[0058] Step S31: Monitor the Ethernet interface of the first system-on-chip to determine whether a target identifier sent by the second system-on-chip through the Ethernet communication link is obtained, where the target identifier is the identifier generated before the first system-on-chip stores the target data into the target direct memory access memory in the computing platform, the notification message carries the target identifier, and the second system-on-chip sends the target identifier to the first system-on-chip through the Ethernet communication link after obtaining the target data from the target direct memory access memory;
[0059] Step S32: When the target identifier sent by the second system-on-chip through the Ethernet communication link is obtained within a predetermined time, determine that the direct memory access communication link between the first system-on-chip and the second system-on-chip is normal, where the direct memory access communication link is the link through which the first system-on-chip and the second system-on-chip communicate via the target direct memory access memory;
[0060] Optionally, during data transmission, the first SOC stores the target data in the target direct memory access memory and sends a notification message with the target identifier to the second SOC via the Ethernet communication link. After successfully obtaining the data from the DMA memory, the second SOC returns the identifier to the first SOC via the Ethernet link. The first SOC checks whether the target identifier returned by the second SOC is received within a predetermined time by monitoring the Ethernet interface to confirm whether the DMA communication link is normal.
[0061] It should be noted that by monitoring the return of the identifier, the abnormality of the DMA communication link can be detected within a short time, the fault tolerance mechanism can be quickly started, and the impact of communication failures on the system can be reduced.
[0062] Step S33: In the case where the target identifier sent by the second system-on-chip via the Ethernet communication link is not obtained within the predetermined time, it is determined that the direct memory access communication link between the first system-on-chip and the second system-on-chip is abnormal, the data transmission link between the first system-on-chip and the second system-on-chip is switched to the Ethernet communication link, and the target data is sent to the second system-on-chip based on the Ethernet communication link;
[0063] Optionally, in the case where it is determined that the DMA communication link between the first SOC and the second SOC is abnormal, the first SOC switches the data transmission link (i.e., the DMA communication link) between the first SOC and the second SOC to the Ethernet communication link, and the first SOC sends the target data to the second SOC via the Ethernet communication link while continuing to monitor the status of the DMA link, thereby ensuring the continuity and reliability of cross-SOC data transmission.
[0064] It should be noted that by quickly switching the data transmission link to the Ethernet communication link, even when the DMA communication link fails, the continuous transmission of data can be ensured, and the system's ability to cope with sudden failures is improved.
[0065] Step S34: Continuously monitor the link status of the direct memory access communication link, and in the case where the link status of the direct memory access communication link becomes normal, send a recovery message to the second system-on-chip based on the Ethernet communication link, where the recovery message is used to instruct the second system-on-chip that the data transmission link between the first system-on-chip and the second system-on-chip is switched back to the direct memory access communication link.
[0066] Optionally, after switching to the Ethernet communication link, the first SOC continuously monitors the status of the DMA communication link. Once the DMA communication link returns to normal, the first SOC sends a recovery message to the second SOC based on the Ethernet communication link, instructing the second SOC to switch the data transmission link (i.e., the Ethernet communication link) back to the DMA communication link to make full use of the low latency and high efficiency characteristics of DMA.
[0067] It should be noted that the recovery of DMA communication can optimize resource utilization and reduce the resource consumption of Ethernet communication. Especially for large - volume data transmission, it can significantly improve the overall communication efficiency of the system. The above steps S31 to S34 ensure the continuity, reliability, and efficiency of data transmission across SOCs through the processes of monitoring, confirmation, fault switching, and status recovery, effectively handle the situation of communication link anomalies, improve the robustness and real - time performance of the system, and have significant technical effects and application values for scenarios with extremely high requirements for real - time performance and communication reliability such as autonomous driving.
[0068] Optionally, for the possible PCIE failures in the strategy of using DMA combined with DDS for cross - SOC large - data communication, which may lead to the inability to use DMA for data transmission, the embodiments of this application design corresponding fault - tolerance mechanisms to ensure that cross - SOC data transmission can work properly in the case where DMA cannot be used due to PCIE failures. At the same time, a monitoring mechanism is designed. As Figure 6 shown, before the node A in the first SOC sends the target data through the DMA link, it generates an identifier (a boolean type indicating whether the DMA data writing is completed). The target data is transmitted to the node B in the second SOC through the DMA link. After successfully receiving the target data, the node B returns the identifier to the node A through the Ethernet link as a reception confirmation signal. When the node A sends the identifier, a timer is started, and a preset time (for example, 10 milliseconds) is set. This preset time is set according to the actual communication latency and performance requirements of the system to ensure that normal network fluctuations can be tolerated. The node A listens to the Ethernet interface during the running of the timer to detect whether the node B returns the corresponding identifier. If the node A successfully receives the identifier returned by the node B within the preset time, it is confirmed that the identifier matches the current data. If the match is successful, it is determined that the current DMA link is working properly, and data transmission continues using the DMA mode. If the node A fails to receive any tag within T time, it immediately determines that the DMA link has failed and starts the fault - tolerance mechanism to switch to the Ethernet communication mode;
[0069] When the DMA link fails and the fault tolerance mechanism is started, Node A and Node B communicate via Ethernet. Node A publishes corresponding topic messages and Node B receives them. The data transmission method switches from the DMA method to the Ethernet method (i.e., from the DMA communication link to the Ethernet communication link). In this mode, data is transmitted normally via Ethernet, and the timer function of Node A continues to work, continuously detecting the status of the DMA link;
[0070] After switching to the Ethernet communication mode, Node A will continuously monitor the status of the DMA link. If it is detected that the DMA link returns to normal, Node A will restart the DMA communication mode and stop the Ethernet communication. After switching back to the DMA mode, Node A re - establishes the communication context to ensure the efficiency of subsequent data transmission. Node A notifies Node B to resume the DMA receiving mode to synchronize the communication path.
[0071] In an exemplary embodiment, the computing platform further includes a target central processing unit. Before storing the target data into the target direct memory access memory in the computing platform, the method further includes the following steps: determining the target direct memory access memory designated by the target central processing unit; wherein, the target direct memory access memory is a direct memory access memory among the M direct memory access memories of the computing platform, where the amount of data allowed to be stored is greater than or equal to the amount of the target data, and the difference between the amount of data allowed to be stored and the amount of the target data is less than or equal to a preset threshold. The M direct memory access memories have a one - to - one correspondence with the M processes in the computing platform. The M processes include a target process, and the target process is used to instruct the first system - on - chip to send the target data to the second system - on - chip, and M is an integer greater than or equal to 2.
[0072] Optionally, the target central processing unit of the computing platform selects a target direct memory access memory that is large enough and exceeds the transmission volume requirement from the multiple DMA memories that have been partitioned according to the estimated data transmission volume of the target process. This step ensures that the target data will not encounter a memory shortage problem during transmission and avoids efficiency losses caused by over - allocating memory.
[0073] Optionally, the target central processing unit of the computing platform analyzes the requirements of the application program, configuration files, or historical data to predict the number of processes (M) in the processes of the computing platform and the data transmission volume of each process, and then partitions the DMA memory in the computer platform in advance.
[0074] It should be noted that by predicting the number of processes and the amount of data transmission, the computing platform can reasonably plan the DMA memory resources during the initialization phase, avoiding inefficiencies or resource waste caused by insufficient or excessive memory allocation. Moreover, pre-allocating the DMA memory area avoids the additional latency brought by dynamically allocating memory during data transmission, thereby improving the speed and real-time performance of data transmission. In addition, predicting the amount of data transmission helps determine the size of the DMA memory area, thereby reducing the waiting time during data processing and transmission and enhancing the efficiency of cross-SOC communication.
[0075] Optionally, the computing platform will divide the DMA memory area of the computing platform according to the number of processes and the amount of data transmission of each process.
[0076] Optionally, for M process tasks, it is necessary to effectively allocate and manage memory between different processes. As Figure 5 shown, memory allocation is mainly divided into two aspects. One is system memory, and the other is DMA memory. Memory needs to be allocated according to the actual application scenario. For DMA memory, since the number of processes for large data transmission is controllable and the amount of data transmitted by the processes is prior data, a fixed partition method is used for memory partitioning. In the fixed partition scheme, according to the number of processes M and the size of the transmitted data m (m is a positive number), the memory corresponding to DMA is divided into n fixed-size areas (that is, the DMA memory area is divided according to the number of processes. The i-th area in the DMA memory area corresponds to the i-th process among multiple processes, and i is a positive integer less than or equal to M), and the memory of each area needs to be greater than or equal to m to ensure normal data transmission. According to this scheme, the memory partitioning of different processes is completed before the process starts, and the start and end addresses of the memory corresponding to each process are determined. Therefore, before the program starts, the memory addresses for data serialization and deserialization of different processes are different. In this way, memory address conflicts are effectively avoided, different processes can be effectively isolated, and the utilization efficiency of memory can be effectively guaranteed.
[0077] It should be noted that the fixed partition DMA memory management strategy improves the efficiency and speed of memory allocation, reduces the memory management overhead during data transmission, enables each process to have an independent DMA memory area, avoids data conflicts or losses that may be caused by multiple processes accessing the same memory area simultaneously, and through the pre-allocation and management of DMA memory, reduces the computational load of the CPU in memory management, enabling the CPU to focus on more core computing tasks and improving the overall computing performance of the computing platform. The fixed DMA memory area allocation strategy enhances the robustness of cross-SOC communication, ensuring the transmission quality of data and the stability of the communication link even when facing sudden data transmission requirements in some areas.
[0078] In an exemplary embodiment, after sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, the method further includes the following steps S41 - S42:
[0079] Step S41: Determine whether the second system-on-chip successfully obtains the target data from the target direct memory access memory, and determine whether there is data to be sent to the second system-on-chip;
[0080] Step S42: In the case where it is determined that the second system-on-chip successfully obtains the target data from the target direct memory access memory and there is no data to be sent to the second system-on-chip, send indication information to the target central processing unit, where the indication information is used to instruct the target central processing unit to release the target direct memory access memory and mark the state of the target direct memory access memory as an idle state.
[0081] Optionally, when the target process completes data transmission, the target central processing unit immediately reclaims the corresponding DMA memory area and marks it as an idle state. Thereafter, if a new process is started, the reclaimed memory area will be preferentially used.
[0082] Optionally, when all DMA memory areas of the computing platform are occupied and there is no idle memory area, the target central processing unit will start a memory allocation failure handling mechanism, trigger an alarm, and execute an emergency handling program to ensure the security and stability of the system.
[0083] In an exemplary embodiment, storing the target data in the target direct memory access memory in the computing platform is implemented through the following steps S51 - S52:
[0084] Step S51: When the target data includes N sub-data, divide the target direct memory access memory into N sub-direct memory access memories according to the data volume of the N sub-data, where N is a positive number greater than or equal to 2;
[0085] Step S52: Store the N sub-data correspondingly in the N sub-direct memory access memories.
[0086] Optionally, the N sub-data correspond to different processes, that is, N processes in the first SOC will respectively instruct the transmission of the N sub-data.
[0087] Optionally, the target DMA memory area is divided into different regions. Assume that region 1 is the DMA memory area shared by the first SOC and the second SOC. Then, according to the N sub-data processed by the first SOC and the second SOC, region 1 is further divided into N sub-memory areas (i.e., N sub-direct memory access memories). Among them, the first sub-memory area in the N memory areas is used to store the first sub-data among the N sub-data to be transmitted from the first SOC to the second SOC.
[0088] In an exemplary embodiment, the target central processing unit may also obtain the data communication requirements between the multiple SOCs, and optimize the sizes and address ranges of the non-prefetch memory and the prefetch memory in the computing platform based on the data communication requirements.
[0089] Optionally, before starting data transmission, it is necessary to comprehensively master the data communication requirements between each SOC, including but not limited to data type, data volume, communication frequency, real-time requirement, and any specific application requirements. For example, in an autonomous driving system, it may be necessary to frequently and in real-time transmit a large amount of image data from a camera SOC to a processing unit SOC, or transmit point cloud data from a lidar SOC to a data analysis SOC. Then, according to the obtained communication requirements, the sizes and address ranges of the non-prefetch memory (Non-Temporary Memory, used to store data that will not be pre-read by the CPU) and the prefetch memory (Temporary Memory, used to store data that may be pre-read by the CPU) in the computing platform are optimized. The optimization strategies may include but are not limited to adjusting the size of the memory area, dynamically allocating the address range of the memory area, and adjusting the configuration of the non-prefetch memory and the prefetch memory according to the communication requirement priority. For example, for large data transmission, the size of the non-prefetch memory can be increased to reduce the direct intervention of the CPU and improve the efficiency of DMA.
[0090] Optionally, the Board Support Package (BSP) is a key component of an embedded device, providing underlying software support for the hardware platform. Correctly configuring the BSP enables the operating system and applications to properly use the hardware resources. To improve data transmission efficiency and reduce transmission latency, it is necessary to reasonably configure the memory mapping strategy of PCIE. For the large data communication requirements across SOCs, optimizing the memory mapping range can effectively improve the overall system performance. By adjusting the PCIE memory mapping parameters in the device tree, the embodiments of this application optimize the sizes and address ranges of non-prefetch memory and prefetch memory, achieving the optimization of non-prefetch memory mapping and prefetch memory mapping. By modifying the mapping configuration of the PCIE controller in the driver program, this embodiment realizes the dynamic adjustment of the starting address of the resource. Through the above methods, the sizes of non-prefetch memory and prefetch memory are increased, and the starting address of the mapping is dynamically adjusted, which can better support large data transmission across SOCs and reduce the latency or data loss caused by insufficient memory during the transmission process.
[0091] In an exemplary embodiment, the method further includes the following steps S61 - S62:
[0092] Step S61: Generate a reference check code based on the target data through a check algorithm;
[0093] Optionally, the check algorithm includes but is not limited to: hash algorithm, cyclic redundancy check algorithm.
[0094] It should be noted that the reference check code reflects the integrity and consistency status of the target data before being written into the target DMA memory, providing a benchmark for subsequent integrity verification. By generating the reference check code, it is possible to effectively detect whether any changes have occurred to the target data during the transmission process, including data corruption and malicious tampering, thus ensuring the accuracy and security of the data during cross-SOC transmission. This is particularly crucial for applications such as autonomous driving and robotics that have strict requirements for data reliability.
[0095] Step S62: During the process of storing the target data into the target direct memory access memory in the computing platform, store the reference check code into the target direct memory access memory, where the notification message is further used to instruct the second SOC to obtain the target data from the target direct memory access memory, generate a target check code based on the obtained data through a check algorithm, and confirm that the content of the target data is correct and complete when it is confirmed that the target check code is consistent with the reference check code.
[0096] Optionally, after receiving the notification message, the second SOC reads the target data from the target direct memory access memory and generates a target check code according to the same verification algorithm. Subsequently, the generated target check code is compared with the reference check code carried in the notification message. If the two are consistent, it indicates that the data has not been tampered with or damaged during transmission, and the data content remains intact and error-free.
[0097] It should be noted that by comparing the target check code with the reference check code, the integrity and accuracy verification of cross-SOC data transmission are realized, further enhancing the security of the entire communication system. Even in a complex network environment, the original state of the data can be ensured, which is crucial for applications involving public safety such as autonomous driving systems and can effectively avoid potential safety hazards caused by data errors.
[0098] Obviously, the above-described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. To better understand the above method, the following will describe the above process in conjunction with embodiments, but it is not used to limit the technical solutions of the embodiments of the present invention:
[0099] The embodiments of this application mainly include six modules: a system tuning module (used to optimize the size and address range of non-prefetch memory and prefetch memory in the computing platform), a data scheduling module (used to send cross-SOC data), a data processing module (used to perform serialization and deserialization processing on data), a shared memory module (used to implement the unified memory architecture of the GPU and CPU), a memory management module (used to manage the memory in the computing platform), and a fault tolerance mechanism module (used to monitor and switch communication links). The above six modules have gone through a series of steps such as operating system layer DMA settings, DDS and DMA cooperative scheduling, data serialization and deserialization, GPU and CPU memory sharing, DMA memory management, and fault tolerance mechanisms, and have completed the overall solution from the underlying system to the upper-layer application.
[0100] The following will further describe the embodiments of the present invention in detail in combination with the big data transmission scenario in autonomous driving:
[0101] Robot Operating System (ROS2) is a development framework and operating system for robot-related applications, with characteristics such as real-time performance, distribution, and security. This system can be applied to autonomous driving and can be used as an intermediate interface between the operating system of autonomous driving and specific applications. The in-vehicle computing platform is an important hardware carrier for running autonomous driving applications. The in-vehicle computing platform used in this embodiment is a heterogeneous distributed computing platform, and its computing units are 4*Orin + X86 CPU. The communication methods between Orin (a highly integrated SOC) include Ethernet and PCIE DMA. The specific device information is shown in Table 1 below:
[0102] Table 1
[0103]
[0104] Based on a heterogeneous distributed computing platform, a big data communication method using DDS combined with DMA for cross-SOC can achieve efficient and low-latency transmission of ROS2 messages across SOCs, thus achieving goals such as cross-device transmission of camera data, reducing the end-to-end latency of the autonomous driving service, and making the autonomous driving system more stable and secure. To achieve the above goals, this example is mainly divided into the following steps:
[0105] I. Cross-SOC communication initialization:
[0106] In the system initialization stage, each Orin computing unit in the vehicle-mounted computing platform needs to establish a stable communication connection. In this embodiment, two communication methods, Ethernet and PCIE DMA, are adopted to support diverse communication requirements.
[0107] (1) DDS configuration:
[0108] First, configure the DDS environment for each Orin computing unit. In this embodiment, Cyclonedds is used as the communication middleware, and Quality of Service (QoS) parameters, including reliability, persistence, and history depth, are set to meet the requirements for real-time performance and reliability in autonomous driving applications. As shown in the following code, the basic DDS communication configuration is performed by setting the domain identification (domain id) and the minimum socket receive buffer size:
[0109] <?xml version="1.0" encoding="UTF-8"?>
[0110] <CycloneDDS xmlns="https: / / cdds.io / config" xmlns:xsi="http: / / www.w3.org / 2001 / XMLSchema-instance" xsi:schemaLocation="https: / / cdds.io / config;
[0111] https: / / raw.githubusercontent.com / eclipse-cyclonedds / cyclonedds / master / etc / cyclonedds.xsd">;
[0112] <Domain id="any">;
[0113] <internal> ;
[0114] <minimumsocketreceivebuffersize>10MB< / minimumsocketreceivebuffersize> ;
[0115] < / internal> ;
[0116] ;
[0117] ;
[0118] (2) DMA channel establishment:
[0119] Use PCIE DMA to establish a high-speed data channel between different SOCs to ensure low-latency transmission of large data. The DMA channel establishment process is mainly divided into: PCIE hardware connection, startup configuration, and startup communication test. First, connect two Orins using a PCIE adapter card; when the Orin starts, the device side performs the following configurations to complete the DMA channel configuration:
[0120] cd / sys / kernel / config / pci_ep / ;
[0121] mkdir functions / pci_epf_nv_test / func1;
[0122] echo 0x10de > functions / pci_epf_nv_test / func1 / vendorid;
[0123] echo 0x0001 > functions / pci_epf_nv_test / func1 / deviceid;
[0124] ln -s functions / pci_epf_nv_test / func1 controllers / 141a0000.pcie_ep / ;
[0125] echo 1 > controllers / 141a0000.pcie_ep / start;
[0126] At the receiving end (end point, abbreviated as EP) and the sending end (root point, abbreviated as RP), perform a communication test to ensure normal DMA communication:
[0127] # Input at the ep end;
[0128] busybox devmem 0x4307b8000 32 0x98765432;
[0129] # Input at the rp end, and the output at the rp end is 0x98765432;
[0130] busybox devmem 0x3a300000;
[0131] # Input at the rp end;
[0132] busybox devmem 0x3a300000 32 0x12345678;
[0133] # Input at the ep end, and the output at the ep end is 0x12345678;
[0134] busybox devmem 0x4307b8000; 123456789;
[0136] II. Data transmission process:
[0137] (1) Data serialization processing:
[0138] For big data communication across SOCs, the data types are often image, point cloud and other data types, and the above data types cannot be directly written into memory. Similar to DDS serialization, for different data types, it is necessary to provide a unified interface for serialization and deserialization to implement writing data to memory and reading data from memory. This embodiment provides a data serialization interface for ROS2 data types, which realizes the conversion of ROS2 data types to unit_8 type, and realizes the direct writing of image (Image) and point cloud (PointCloud) data types in ROS2 to memory.
[0139] (2) Writing data to DMA memory:
[0140] The serialized data is written to the specified DMA memory area through the PCIE DMA channel. During the writing process, the system will update the boolean type identifier in the DDS domain in real time to identify the completion status of data writing.
[0141] (3) DDS message publishing:
[0142] Once the data writing to the DMA memory area is completed, DDS broadcasts this identifier to other Orin computing units through Ethernet to notify the receiving end that the data is ready to be read.
[0143] (4) Data reading and deserialization:
[0144] When the Orin computing unit at the receiving end receives the DDS message, it immediately starts the process of reading DMA memory data. The read data stream will be converted back to the ROS2 message format through the deserialization interface and handed over to the application layer for processing.
[0145] III. Memory Management and Optimization:
[0146] For multiple process tasks, it is necessary to effectively allocate and manage memory between different processes. Memory allocation is mainly divided into two aspects, one is system memory and the other is DMA memory, and memory needs to be allocated according to the actual application scenario. In order to effectively manage and allocate memory between multiple computing units, this embodiment implements DMA and system memory partitioning.
[0147] (1) Memory Partition Management:
[0148] According to the task load situation in the autonomous driving system, the DMA memory is divided into several fixed-size partitions, and the memory allocation strategy is dynamically adjusted according to different data transfer tasks. The size of each partition is larger than the possible data block to ensure the integrity of data transfer.
[0149] (2) Memory Allocation at Process Startup:
[0150] Whenever a new process starts and requires large data transfer, the system allocates a fixed partition for this process. Since the partition size is determined during the system initialization phase, the allocation process is simple and efficient, reducing the latency caused by memory allocation.
[0151] (3) Memory Recycling and Reallocation:
[0152] When the process completes data transfer, the system immediately reclaims the corresponding DMA memory partition and marks it as idle. After that, if a new process starts, the system will preferentially use the reclaimed memory partition.
[0153] (4) Handling of Memory Allocation Failure:
[0154] In extremely rare cases, if all partitions are occupied and there is no free partition, the system will start the memory allocation failure handling mechanism, trigger an alarm and execute an emergency handling program to ensure the safety and stability of the system.
[0155] IV. Memory Sharing:
[0156] Due to its unified memory feature, it is possible to achieve CPU and GPU memory sharing on two devices. Based on this feature, combined with DDS+DMA proposed in the data scheduling module, the following functions are realized: The GPU of Host A generates data and writes it into the shared memory; the CPU of Host A publishes a flag service to Host B through DDS. The CPU of Host B receives the identifier and directly accesses the shared memory through the DMA channel; the GPU of Host B can directly read data from the shared memory without copying the data.
[0157] V. Fault Tolerance Mechanism:
[0158] (1) Link Failure Judgment:
[0159] Before node A sends a data packet through the DMA link, it generates an identifier. The data is transmitted through the DMA link to node B. After node B successfully receives the data, it returns the identifier to node A through the Ethernet link as a reception confirmation signal. At the same time as node A sends the identifier, a timer is started, and a predetermined waiting time T (for example, 10 milliseconds) is set. This waiting time T is set according to the actual communication delay and performance requirements of the system to ensure that normal network fluctuations can be tolerated. Node A listens to the Ethernet interface during the operation of the timer to detect whether node B returns the corresponding identifier. If node A successfully receives the identifier returned by node B within time T, it is confirmed that the identifier matches the current data. If the match is successful, it is determined that the current DMA link is working properly, and data transmission continues using the DMA mode. If node A does not receive any label within time T, it immediately determines that the DMA link has failed, starts the fault tolerance mechanism, and switches to the Ethernet communication mode.
[0160] (2) Data Link Switching:
[0161] When the DMA link fails and the fault tolerance mechanism is started, node A and node B transmit through the Ethernet. Node A publishes the corresponding topic message and node B receives it. The data transmission method is switched from the DMA method to the Ethernet method. In this mode, data is normally transmitted through the Ethernet, and the timer function of node A still continues to work, continuously detecting the status of the DMA link.
[0162] (3) Link Monitoring Mechanism:
[0163] After switching to the Ethernet communication mode, node A will continuously monitor the status of the DMA link. If it is detected that the DMA link returns to normal, node A will restart the DMA communication mode and stop the Ethernet communication. After switching back to the DMA mode, node A re-establishes the communication context to ensure the efficiency of subsequent data transmission. Node A notifies node B to resume the DMA reception mode to achieve synchronization of the communication path.
[0164] VI. Performance Optimization and Evaluation:
[0165] To verify the effectiveness of this embodiment, the big data transmission performance between two Orins was evaluated, as Figure 7 shown below:
[0166] Compared with data communication through Cyclonedds (Ethernet), when transmitting big data between devices, the DDS+DMA method has a more obvious effect. When transmitting messages larger than 500Kb, the performance of the DMA method begins to exceed that of DDS. When the message size is 100Mb, the DMA transmission time is only one-tenth of that of DDS. This fully demonstrates that the data communication model based on DDS+DMA has higher communication efficiency and lower latency than Cyclonedds when transmitting big data across SOCs.
[0167] It should be noted that the embodiments of this application design a big data communication method and device based on DDS and DMA. On the one hand, this method utilizes the shared memory feature of DMA to achieve big data communication across SOCs, and at the same time uses DDS to establish a message communication publish-subscribe mechanism to complete the problem of efficient big data communication across SOCs; meanwhile, a fault tolerance mechanism is designed to ensure that big data communication can still be carried out in the event of a DMA failure. This effectively improves the big data communication efficiency and communication cost across SOCs, significantly reduces the CPU resource utilization rate, and can efficiently implement GPU-CPU, GPU-GPU, and CPU-CPU communication across SOCs, achieving zero-copy of data across SOCs, significantly reducing communication latency, reducing system resource consumption, and effectively enhancing the stability and real-time performance of the system; at the same time, the fault tolerance mechanism can carry out data communication in the event of DMA failure, effectively enhancing the robustness of the system's data communication, and has high practical value for complex autonomous driving-related applications.
[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0169] In this embodiment, a data transmission device is further provided, which is applied to the first system-on-chip among multiple system-on-chips in a computing platform, and the multiple system-on-chips further include a second system-on-chip; for implementing the above embodiments and preferred implementation manners, those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the modules described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0170] Figure 8 is a structural block diagram of a data transmission device according to an embodiment of the present application. The device includes:
[0171] A storage module 802, configured to store the target data in the target direct memory access memory in the computing platform when the target data is to be sent to the second system-on-chip, where the target direct memory access memory is the direct memory access memory corresponding to the first system-on-chip and the second system-on-chip;
[0172] A sending module 804, configured to send a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0173] In the above device, when the first system-on-chip among multiple system-on-chips in the computing platform is to send target data to the second system-on-chip, the target data is stored in the target direct memory access memory of the computing platform, and a notification message is sent to the second system-on-chip based on the Ethernet communication link, so as to instruct the second system-on-chip to obtain the target data from the target direct memory access memory. Since the direct memory access communication link between the first system-on-chip and the second system-on-chip is used for data transmission, the data transmission efficiency is improved, and the problem of low transmission efficiency of cross-system-on-chip data transmission based on Ethernet is solved.
[0174] In an exemplary embodiment, the storage module 802 is further configured to store the first data to be sent by the central processing unit in the first system-on-chip in the target direct memory access memory in the computing platform; and / or store the second data to be sent by the graphics processing unit in the first system-on-chip in the target direct memory access memory in the computing platform, where the data in the target direct memory access memory is to be obtained by the central processing unit and / or the image processor in the second system-on-chip, and the target data includes: the first data and / or the second data.
[0175] In an exemplary embodiment, the storage module 802 is further configured to serialize the target data into data of a target type by using a sequence number interface corresponding to the data type of the target data, where the target type of data is allowed to be stored in the direct memory access memory; and write the data of the target type into the target direct memory access memory.
[0176] In an exemplary embodiment, the above device further includes: a monitoring module, configured to, after sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, monitor the Ethernet interface of the first system-on-chip to determine whether a target identifier sent by the second system-on-chip through the Ethernet communication link is obtained, where the target identifier is an identifier generated before the first system-on-chip stores the target data in the target direct memory access memory in the computing platform, the notification message carries the target identifier, and the second system-on-chip sends the target identifier to the first system-on-chip through the Ethernet communication link after obtaining the target data from the target direct memory access memory; determine that the direct memory access communication link between the first system-on-chip and the second system-on-chip is normal when the target identifier sent by the second system-on-chip through the Ethernet communication link is obtained within a predetermined time, where the direct memory access communication link is a link through which the first system-on-chip and the second system-on-chip communicate through the target direct memory access memory; determine that the direct memory access communication link between the first system-on-chip and the second system-on-chip is abnormal when the target identifier sent by the second system-on-chip through the Ethernet communication link is not obtained within the predetermined time, switch the data transmission link between the first system-on-chip and the second system-on-chip to the Ethernet communication link, and send the target data to the second system-on-chip based on the Ethernet communication link; continuously monitor the link state of the direct memory access communication link, and when the link state of the direct memory access communication link becomes normal, send a recovery message to the second system-on-chip based on the Ethernet communication link, where the recovery message is used to instruct the second system-on-chip that the data transmission link between the first system-on-chip and the second system-on-chip is switched back to the direct memory access communication link.
[0177] In an exemplary embodiment, the computing platform further includes a target central processing unit, and the above device further includes: a first determination module, configured to determine the target direct memory access memory designated by the target central processing unit before storing the target data into the target direct memory access memory in the computing platform; wherein, the target direct memory access memory is a direct memory access memory among the M direct memory access memories in the computing platform where the amount of data allowed to be stored is greater than or equal to the amount of the target data, and the difference between the amount of data allowed to be stored and the amount of the target data is less than or equal to a preset threshold, and the M direct memory access memories have a one-to-one correspondence with M processes in the computing platform, the M processes include a target process, and the target process is used to instruct the first system-on-chip to send the target data to the second system-on-chip, and M is an integer greater than or equal to 2.
[0178] In an exemplary embodiment, the computing platform further includes a target central processing unit, and the above device further includes: a second determination module, configured to determine whether the second system-on-chip successfully obtains the target data from the target direct memory access memory and determine whether there is data to be sent to the second system-on-chip after sending a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip; in the case of determining that the second system-on-chip successfully obtains the target data from the target direct memory access memory and there is no data to be sent to the second system-on-chip, send indication information to the target central processing unit, where the indication information is used to instruct the target central processing unit to release the target direct memory access memory and mark the state of the target direct memory access memory as an idle state.
[0179] In an exemplary embodiment, the storage module 802 is further configured to, when the target data includes N sub-data, divide the target direct memory access memory into N sub-direct memory access memories according to the amounts of the N sub-data, where N is a positive number greater than or equal to 2; and store the N sub-data correspondingly in the N sub-direct memory access memories.
[0180] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0181] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0182] Optionally, in this embodiment, the above computer program may be configured to perform the following steps by a computer program:
[0183] S1. When the target data is to be sent to the second system-on-chip, store the target data in the target direct memory access memory in the computing platform, where the target direct memory access memory is the direct memory access memory corresponding to the first system-on-chip and the second system-on-chip;
[0184] S2. Send a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0185] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0186] The embodiment of the present application also provides an electronic device, as Figure 9 shown. The electronic device includes a memory 902 and a processor 904. A computer program is stored in the memory 902, and the processor 904 is configured to perform the steps in any of the above method embodiments by the computer program.
[0187] Optionally, in this embodiment, the above processor 904 may be configured to perform the following steps by a computer program:
[0188] S1. When the target data is to be sent to the second system-on-chip, store the target data in the target direct memory access memory in the computing platform, where the target direct memory access memory is the direct memory access memory corresponding to the first system-on-chip and the second system-on-chip;
[0189] S2. Send a notification message to the second system-on-chip based on the Ethernet communication link between the first system-on-chip and the second system-on-chip, where the notification message is used to instruct the second system-on-chip to obtain the target data from the target direct memory access memory.
[0190] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0191] Optionally, those of ordinary skill in the art can understand that Figure 9 the structure shown is only schematic, Figure 9 and does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components (such as a network interface, etc.) than those shown in Figure 9 , or have a different configuration from that shown in Figure 9 .
[0192] Among them, the memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the data transmission method and the data transmission device in the embodiments of the present application. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, implements the above-mentioned data transmission method. The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 902 may further include a memory remotely disposed relative to the processor 904, and these remote memories can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 902 can specifically but not limitedly be used to store information such as system configuration files. As an example, as Figure 9 shown, the above-mentioned memory 902 may include but is not limited to the storage module 802 and the sending module 804 in the above-mentioned data transmission device. In addition, it may also include but is not limited to other module units in the above-mentioned data transmission device, which will not be elaborated in this example.
[0193] Optionally, the above-mentioned transmission device 906 is used to receive or send data via a network. Specific examples of the above network may include wired networks and wireless networks. In one instance, the transmission device 906 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one instance, the transmission device 906 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly.
[0194] In addition, the above-mentioned electronic device further includes: a display 908; and a connection bus 910 for connecting each module component in the above-mentioned electronic device.
[0195] The embodiments of the present application also provide a computer program product, and the above computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0196] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.
[0197] An embodiment of the present application also provides a computer program, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0198] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0199] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A data transmission method, characterized in that: A first system on chip in a plurality of systems on chips in a computing platform, wherein the plurality of systems on chips also includes a second system on chip; include: In a case where target data is to be sent to the second system on chip, storing the target data in a target direct memory access memory in the computing platform, wherein the target direct memory access memory is a direct memory access memory corresponding to the first system on chip and the second system on chip; A notification message is sent to the second system on chip based on an Ethernet communication link between the first system on chip and the second system on chip, wherein the notification message is used to instruct the second system on chip to obtain the target data from the target direct memory access memory.
2. The method according to claim 1, characterized in that The storing the target data into a target direct memory access memory in the computing platform comprises: storing first data to be sent by the central processor in the first system-on-chip into the target direct memory access memory in the computing platform; and / or storing the second data to be sent by the graphics processor in the first system-on-chip into the target direct memory access memory in the computing platform, The data in the target direct memory access memory is to be acquired by the central processing unit and / or the image processor in the second system on chip, and the target data includes: the first data and / or the second data.
3. The method according to claim 1, characterized in that The storing the target data into a target direct memory access memory in the computing platform comprises: Serializing the target data into data of a target type using a serial number interface corresponding to the data type of the target data, wherein the direct memory access memory allows the data of the target type to be stored; Writing data of the target type to the target direct memory access memory.
4. The method according to claim 1, characterized in that: After sending a notification message to the second system on chip based on the Ethernet communication link between the first system on chip and the second system on chip, the method further includes: monitoring the Ethernet interface of the first system-on-chip to determine whether a target identifier sent by the second system-on-chip through the Ethernet communication link is obtained, wherein the target identifier is an identifier generated by the first system-on-chip before storing the target data in a target direct memory access memory in the computing platform, the notification message carries the target identifier, and the second system-on-chip sends the target identifier to the first system-on-chip through the Ethernet communication link after obtaining the target data from the target direct memory access memory; In the case where a target identifier sent by the second system on chip through the Ethernet communication link is acquired within a predetermined time, determining that a direct memory access communication link between the first system on chip and the second system on chip is normal, wherein the direct memory access communication link is a link for the first system on chip and the second system on chip to communicate through the target direct memory access memory; If the target identifier sent by the second system on chip through the Ethernet communication link is not obtained within the predetermined time, determining that the direct memory access communication link between the first system on chip and the second system on chip is abnormal, switching the data transmission link between the first system on chip and the second system on chip to the Ethernet communication link, and sending the target data to the second system on chip based on the Ethernet communication link; Continuously monitor the link status of the direct memory access communication link, and when the link status of the direct memory access communication link becomes normal, send a recovery message to the second system on chip based on the Ethernet communication link, wherein the recovery message is used to instruct the second system on chip that the data transmission link between the first system on chip and the second system on chip is switched back to the direct memory access communication link.
5. The method according to claim 1, characterized in that The computing platform also includes a target central processing unit, Before storing the target data in the target direct memory access memory in the computing platform, the method further includes: Determining the target direct memory access memory specified by the target central processing unit; Among them, the target direct memory access memory is a direct memory access memory among the M direct memory access memories of the computing platform, and the amount of data allowed to be stored is greater than or equal to the amount of data of the target data, and the difference between the amount of data allowed to be stored and the target amount of data is less than or equal to a preset threshold. The M direct memory access memories have a one-to-one correspondence with the M processes in the computing platform, and the M processes include a target process. The target process is used to instruct the first system on chip to send the target data to the second system on chip, and M is an integer greater than or equal to 2.
6. The method according to claim 1, characterized in that The computing platform also includes a target central processing unit, After sending a notification message to the second system on chip based on the Ethernet communication link between the first system on chip and the second system on chip, the method further includes: determining whether the second system on chip successfully obtains the target data from the target direct memory access memory, and determining whether there is data to be sent to the second system on chip; When it is determined that the second system on chip successfully obtains the target data from the target direct memory access memory and there is no data to be sent to the second system on chip, an indication message is sent to the target central processor, wherein the indication message is used to instruct the target central processor to release the target direct memory access memory and mark the state of the target direct memory access memory as an idle state.
7. The method according to claim 1, characterized in that Storing the target data in a target direct memory access memory in the computing platform includes: In the case where the target data includes N sub-data, dividing the target direct memory access memory into N sub-direct memory access memories according to the data amount of the N sub-data, where N is a positive number greater than or equal to 2; The N sub-data are correspondingly stored in the N sub-direct memory access memories.
8. A data transmission device, characterized in that: A first system on chip in a plurality of systems on chips in a computing platform, wherein the plurality of systems on chips also includes a second system on chip; include: a storage module, configured to store the target data in a target direct memory access memory in the computing platform when the target data is to be sent to the second system on chip, wherein the target direct memory access memory is a direct memory access memory corresponding to the first system on chip and the second system on chip; A sending module, used for sending a notification message to the second system on chip based on the Ethernet communication link between the first system on chip and the second system on chip, wherein the notification message is used to instruct the second system on chip to obtain the target data from the target direct memory access memory.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
RDMA-based low-delay data transmission method and related device
CN113886294A
Data shuffling method, device and equipment, computer readable storage medium and product
CN116932521A
Communication link detection method and device
CN117240748A
Data transmission method and device of controller, controller, storage medium and vehicle
CN118282931A
Direct memory access device and method, electronic equipment and readable storage medium
CN119046187A
Cited By
Memory management system and method, compiling method, equipment, medium and product
CN121255475A
Memory management system, method, compiling method, device, medium and product
CN121255475B