Vehicle-mounted communication system based on distributed architecture and vehicle
By configuring the root port and end node in the on-board communication system, and using PCIe switches to transmit data in the shared storage space, the blocking problem of large-scale data throughput under Ethernet communication is solved, and efficient data transmission and communication is achieved.
Patent Information
- Application Number
- CN202510855104.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In the case of large-scale data throughput under the communication interconnection of Ethernet, blocking is prone to occur, resulting in a decrease in communication efficiency.
Using a distributed architecture-based on-vehicle communication system, multiple edge computing devices are configured as root ports and end nodes, and PCIe switches are used to realize high-speed connection between root ports and end nodes, and a shared storage space is set so that the root ports and end nodes can be directly transmitted in the shared storage space through PCIe switches.
It greatly improves data transmission efficiency, reduces transmission delay, solves the blocking problem caused by large-scale data throughput under Ethernet communication interconnection, and improves communication efficiency.
Smart Images

Figure CN120358213A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and particularly to an in-vehicle communication system and a vehicle based on a distributed architecture. Background Art
[0002] With the development of technologies, the design concept of the autonomous driving framework has changed from modular to an integrated perception and planning one, which has become the mainstream in the industry. The vehicle-end computing platform is more restricted by many practical factors such as power supply, heat dissipation, safety, and deployment cost. The in-vehicle computing platform usually uses multiple system-on-chips to form high computing power to support the efficient and stable inference of the upper-layer framework.
[0003] In related technologies, the device interconnection between multiple system-on-chips is achieved through Ethernet, which is prone to blocking in the case of a large volume of data throughput, resulting in a decline in communication efficiency. Summary of the Invention
[0004] This application provides an in-vehicle communication system and a vehicle based on a distributed architecture to at least solve the problem in related technologies that blocking is likely to occur in the case of a large volume of data throughput under Ethernet communication interconnection, resulting in a decline in communication efficiency.
[0005] This application provides an in-vehicle communication system based on a distributed architecture, including: Multiple edge computing devices; wherein, at least one edge computing device is configured as a root port, and at least one edge computing device is configured as an end node; A PCIe switch, PCIe-connected to the root port and the end node respectively; Wherein, the root port includes a shared storage space, and the root port and the end node perform data transmission in the shared storage space through the PCIe switch.
[0006] This application also provides a vehicle, including: the above-mentioned in-vehicle communication system based on a distributed architecture.
[0007] Through this application, by respectively configuring multiple edge computing devices as root ports and end nodes, and using a PCIe switch to achieve a high-speed connection between the root port and the end node, the high bandwidth and low latency characteristics of PCIe can be fully utilized, greatly improving the data transmission efficiency, thereby ensuring the smoothness of the in-vehicle computing platform when processing a large amount of data. The root port is provided with a shared storage space, enabling the root port and the end node to directly perform data transmission in the shared storage space through the PCIe switch, further reducing the intermediate links of data transmission and lowering the transmission latency. Therefore, the problem that blocking is likely to occur in the case of a large volume of data throughput under Ethernet communication interconnection, resulting in a decline in communication efficiency, can be solved, achieving the technical effect of improving communication efficiency. Description of the Drawings
[0008] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0009] Figure 1 is a schematic diagram of the architecture of a vehicle-mounted communication system based on a distributed architecture provided by an embodiment of the present application; Figure 2 is a schematic diagram of the mapping process of a physical address provided by an embodiment of the present application; Figure 3 is a schematic diagram of the data transmission process between a publishing node and a subscribing node provided by an embodiment of the present application; Figure 4 is a schematic diagram of the data transmission process of a graphics processor provided by an embodiment of the present application; Figure 5 is a schematic diagram of the structure of the communication logic provided by an embodiment of the present application. Detailed implementation manners
[0010] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0011] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0012] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0013] In combination with the specific application environment architecture or specific hardware architecture on which the vehicle-mounted communication system based on a distributed architecture depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0014] Embodiments of the present application provide a vehicle-mounted communication system based on a distributed architecture, and the system is described in detail in combination with the architecture of the vehicle-mounted communication system based on the distributed architecture.
[0015] Specifically, Figure 1 FIG. is a schematic diagram of the architecture of a vehicle-mounted communication system based on an embodiment of the present application.
[0016] As Figure 1 shown, the vehicle-mounted communication system based on the distributed architecture includes a plurality of edge computing devices and a PCIe (Peripheral Component Interconnect Express) switch.
[0017] The edge computing device is a device with powerful computing capabilities. By setting the edge computing device in the vehicle-mounted communication system, the computing task can be transferred from the central cloud or data center to a location closer to the data source, and data processing can be performed near the data source to reduce data transmission pressure and latency, and improve the intelligence and automation level of the vehicle. In some embodiments, the edge computing device adopted may be a system-on-chip (SoC), such as the Nvidia Drive Orin device, and each edge computing device may include an independent CPU (Central Processing Unit) and GPU (Graphics Processing Unit).
[0018] In a vehicle-mounted scenario, different edge computing devices can cooperate with each other. For example, some edge computing devices are responsible for processing image data collected by cameras, and some edge computing devices are responsible for processing sensor data such as radars, to achieve distributed computing, thereby improving the overall computing power level and jointly providing support for autonomous driving decisions.
[0019] As Figure 1 shown, the plurality of edge computing devices may include an edge computing device 101, an edge computing device 102, an edge computing device 103, and an edge computing device 104. It should be noted that Figure 1 the number of edge computing devices is for illustrative purposes, and the number of edge computing devices may be 2, 3, 5, or other numbers, and the embodiments of the present application do not limit this.
[0020] The Edge Computing Devices 101, 102, 103, and 104 communicate with each other via PCIe connections. Among them, PCIe is a high-speed serial computer expansion bus standard, which is widely used in the fields of computers and servers. PCIe has the characteristics of high bandwidth and low latency, and can support the rapid transmission of a large amount of data. Compared with traditional Ethernet, PCIe provides higher data transmission efficiency at the hardware level, especially when dealing with high-density data streams. The PCIe connection utilizes the high bandwidth and low latency characteristics of the PCIe bus, enabling the Edge Computing Devices to transmit data quickly and efficiently.
[0021] In a vehicle communication system based on a distributed architecture, multiple Edge Computing Devices are divided into a Root Port (RootPort System, RP) and End Nodes (EndPoint, EP). For example, in Figure 1 , Edge Computing Device 101 is the Root Port, and Edge Computing Devices 102, 103, and 104 are End Nodes.
[0022] The Root Port is the core node of the entire communication architecture, responsible for managing and coordinating the data transmission of the entire system. The End Node is the terminal device for data transmission, responsible for executing specific computing tasks or data processing. By dividing the Edge Computing Devices into a Root Port and End Nodes, the data transmission can be organized more efficiently, enabling the orderly flow of data between different devices.
[0023] Since PCIe direct connection can only achieve point-to-point connection, in order to achieve high-speed connection between the Root Port and End Nodes, the embodiment of this application uses a PCIe switch. A PCIe switch is a device specifically used to manage PCIe connections, including multiple PCIe interfaces. The Root Port and End Nodes can be connected to the PCIe switch through the PCIe interfaces respectively. The Root Port and End Nodes can perform high-speed data transmission through the PCIe interfaces. The core function of the PCIe switch is to achieve fast data forwarding and routing, enabling data to be efficiently transmitted on different PCIe links. By using the PCIe switch, flexible interconnection between multiple Edge Computing Devices is achieved, and a stable and reliable distributed communication network is built.
[0024] In the embodiments of the present application, the root port includes a shared storage space. The shared storage space is a memory space of a certain size in the root port, which allows the root port and the end node to directly perform data read and write operations in this space, thereby realizing the DMA (Direct Memory Access) interconnection between the root port and the end node. Among them, DMA is a technology that allows certain hardware subsystems to directly access the system memory without the control of the CPU, which can significantly improve the data transmission efficiency and reduce the burden on the CPU.
[0025] In the embodiments of the present application, the root port and the end node can perform data transmission in the shared storage space through a PCIe switch. Specifically, the end node can write the processed data into the shared storage space through DMA, and the root port can read the data from the shared storage space through DMA for further processing or forwarding to other end nodes. This way of directly exchanging data in the shared storage space makes full use of the low-latency characteristics of PCIe, enabling the root port and the end node to bypass the CPU access when reading and writing in the shared memory space, and realizing the zero-copy direct access of point-to-point data between devices. This method further reduces the intermediate links of data transmission and reduces the transmission latency.
[0026] Through the present application, by respectively configuring multiple edge computing devices as the root port and the end node and using a PCIe switch to achieve a high-speed connection between the root port and the end node, the high-bandwidth and low-latency characteristics of PCIe can be fully utilized, greatly improving the data transmission efficiency, thereby ensuring the smoothness of the in-vehicle computing platform when processing a large amount of data. The root port is provided with a shared storage space, enabling the root port and the end node to directly perform data transmission in the shared storage space through a PCIe switch, further reducing the intermediate links of data transmission and reducing the transmission latency. Therefore, the problem of easy blockage and resulting in a decline in communication efficiency in the case of a large amount of data throughput under Ethernet communication interconnection can be solved, achieving the technical effect of improving communication efficiency.
[0027] In some embodiments, before performing data transmission in the shared storage space through a PCIe switch, it includes: The end node maps the physical address of the end node to the shared storage space of the root port through the base address register, so as to facilitate the end node to determine the physical address of the end node through the shared storage space.
[0028] In this embodiment, the base address register (BAR) is an important component in the PCIe architecture, which is used to store the physical address range information of the device and allows devices to access each other's memory spaces through mapping.
[0029] In this embodiment, the end node configures the base address register during the initialization phase, and defines the mapping position of the memory space of the end node in the shared storage space of the root port through the base address register. Specifically, as Figure 2 shown, when the end node starts up or accesses the system, it sends a configuration request to the root port through the PCIe bus, carrying the information of its own base address register. After receiving this information, the root port allocates a corresponding address space for the end node in the shared storage space and records the address mapping relationship. In this way, the physical address of the end node is mapped to the shared storage space of the root port.
[0030] After the address mapping is completed, the end node can determine its own physical address through the shared storage space. When the end node needs to perform data transmission, it only needs to access the shared storage space according to the mapped address without caring about the specific location of the physical address.
[0031] In this embodiment, the end node maps its own physical address to the shared storage space of the root port through the base address register, enabling the end node to accurately access the shared storage space and improving the accuracy of data transmission. The root port can also optimize the data transmission path and improve the communication performance by uniformly managing the address mapping of the shared storage space.
[0032] In some embodiments, data transmission is performed in the shared storage space through a PCIe switch, including: The data writing parties in the root port and the end node serialize the data into binary to obtain serialized data; compress the serialized data and write the compressed serialized data into the shared storage space.
[0033] In this embodiment, the data writing party is the party in the root port and the end node that writes data into the shared storage space.
[0034] Serialization is a process of converting a data structure or object state into a storable or transmittable format, and binary serialization is a process of converting a data structure or object state into a binary format. For example, in ROS2 (Robot Operating System 2), the data writing party can call the serialize_message() method in the C++ client library (rclcpp) to convert a ROS2 format Topic (topic) message into a vector<uint8_t> type. For example, a Topic that has been registered with a message format in ROS2 can be converted into binary serialized data t_binary, such as 0110001011110101.
[0035] However, since the underlying BAR0 (Base Address Register 0) transfers data in units of page memory in 32-bit hexadecimal format, directly transferring binary serialized data may result in too low semantic density of the message carrier. To improve the transmission efficiency, the serialized data can be losslessly compressed.
[0036] Lossless compression is a compression method that can fully restore the original data. It reduces the volume of data by reducing redundant information in the data. In one example, after lossless compression, the effective length of the data can be compressed to 1 / 8 of the original, and the semantic density of the compressed data t_hex is 8 times that of the uncompressed t_binary. The compression operation not only reduces the bandwidth requirement for data transmission but also improves the speed and efficiency of data transmission.
[0037] After serialization and compression are completed, the data writer can write the processed data into the register address of the root port by means of memory copy, thereby implementing DMA transmission.
[0038] In this embodiment, complex data can be converted into a binary sequence through serialization operations, which facilitates data transmission and processing between different devices and improves communication compatibility; by compressing the serialized data, not only the bandwidth requirement for data transmission is reduced, but also the semantic density of the message carrier is increased.
[0039] In some embodiments, data is transmitted in the shared storage space through a PCIe switch, including: The data readers in the root port and the end node read the stored data in the shared storage space and restore the read data through deserialization.
[0040] In this embodiment, the data reader is the party in the root port and the end node that reads data from the shared storage space.
[0041] In this embodiment, the data reader can read data from the shared storage space through DMA read mode. Taking the data reader as the end node as an example, the data stored in the shared storage space can be read from the register address of the end node by means of memory copy.
[0042] Since the data stored by the data writer is compressed data, for example, the compressed data t_hex is in hexadecimal format. After the data reader reads the hexadecimal format data, it needs to deserialize these data back into binary format data. For example, bitwise operations and encoding rules can be used for deserialization to convert the t_hex data back into the binary sequence t_binary. Then, the serialize_message() method in the C++ client library (rclcpp) of ROS2 is called to convert the binary sequence t_binary into the ROS2 format. In this way, the ROS2 application design based on PCIE communication for different edge computing technologies is realized.
[0043] In this embodiment, the integrity and semantic consistency of data during transmission are improved through deserialization operations, enabling the data reader to correctly understand and process the received data and enhancing the data transmission efficiency between different devices.
[0044] In some embodiments, a vehicle-mounted communication system based on a distributed architecture may further include: An Ethernet switch 130, which is respectively connected to the root port and the end nodes via Ethernet; A message middleware, which is used to transmit messages between the root port and the end nodes and between different end nodes through the Ethernet switch 130.
[0045] In this embodiment, in order to further enhance the communication capabilities of the vehicle-mounted communication system based on a distributed architecture, an Ethernet switch 130 may also be set in the vehicle-mounted communication system based on a distributed architecture.
[0046] The Ethernet switch 130 is a network device that can realize Ethernet connections between multiple devices. As a widely used network technology, Ethernet has good compatibility and scalability and can support various types of devices and communication protocols.
[0047] As Figure 1 shown, in this embodiment, the Ethernet switch 130 is respectively connected to the root port and the end nodes via Ethernet, enabling data transmission between the root port and the end nodes via Ethernet.
[0048] In this embodiment, in addition to the hardware connection provided by the Ethernet switch, a message middleware is introduced. The message middleware is a software component used to transmit messages between different devices. In some embodiments, DDS (Data Distribution Service) can be used as the message middleware. DDS is a high-performance middleware protocol that can meet application scenarios with extremely high requirements for real-time performance and reliability, such as autonomous driving.
[0049] In this embodiment, the role of the message middleware is to transmit messages between the root port and the end nodes, as well as between different end nodes, through an Ethernet switch. The message middleware adopts a publish-subscribe or point-to-point communication mode, allowing devices to interact in a loosely coupled manner. For example, an end node can act as a message publishing node and publish specific types of data to the message middleware; the root port or other end nodes can act as subscription nodes and subscribe to the message types they are interested in. The message middleware is responsible for receiving, storing, and forwarding these messages.
[0050] In this embodiment, the high-speed data transmission channel provided by the PCIe switch is suitable for scenarios of large-volume and low-latency data transmission, such as the real-time transmission and processing of sensor data; while the Ethernet switch and the message middleware provide flexibility and scalability for communication, and are suitable for scenarios such as control instruction transmission and configuration information exchange, reducing the problem of resource waste caused by occupying PCIe.
[0051] In this embodiment, by combining the distributed communication architectures of PCIe and Ethernet, the advantages of both technologies are fully utilized. The high bandwidth and low latency characteristics of PCIe improve the transmission efficiency of critical data, meeting the strict real-time requirements of the autonomous driving system; while the intelligent coordination of Ethernet and the message middleware enables the communication process to also adapt to different types of data transmission requirements, support diverse application scenarios, improve the scalability and maintainability of the system, and reduce the development and deployment costs.
[0052] In some embodiments, message transmission through an Ethernet switch includes: Using a topic as the carrier of message transmission and encapsulating data tags in the topic for message transmission; the data tags characterize the attributes of the data.
[0053] In the traditional DDS communication model, a topic usually directly encapsulates the encoded raw data, such as image data in the sensor_msgs::msg::Image format or point cloud data in the sensor_msgs::msg::PointCloud2 format. However, with the development of autonomous driving technology, the number and accuracy of sensors carried by vehicles have been continuously improved, and the amount of data generated has increased explosively. The characteristic that DDS only supports Ethernet in device-to-device communication makes it prone to congestion due to the actual bandwidth of Ethernet when undertaking large-data transmission tasks, resulting in communication delays or loss of important data frames.
[0054] Taking an entry-level autonomous commercial vehicle as an example, when equipped with 6 cameras with 200W pixels and 1 lidar with 128 lines and collecting data at a rate of 30Hz, the data generation rate is as high as 26.4 Gb / s, far exceeding the effective transmission capacity of general in-vehicle 10 Gigabit Ethernet (10 Gb / s), which easily leads to message congestion and data frame loss.
[0055] In this embodiment, the traditional communication model is optimized. DDS is retained, and the architecture and publish-subscribe mode of Ethernet + DDS are still adopted. However, instead of directly encapsulating the original data in the topic, data tags are encapsulated. The data tag contains key attribute information of the data, such as data encryption type, the starting address of memory storage, the length of memory occupation, etc.
[0056] In one example, the data tag can be defined as: { string name; Message name string type; Message type long start; Starting pointer address of the message long size; Message length } In this way, the amount of information carried by the topic is greatly reduced, thus significantly reducing the burden of Ethernet transmission. For example, for a topic that originally needed to transmit a large amount of image or point cloud data, now only a short message containing the data tag needs to be transmitted, greatly improving the efficiency of message transmission.
[0057] The original data itself is written into the shared storage space through the PCIe + DMA method. When the subscription node receives the topic containing the data tag, it can directly read the corresponding original data from the shared storage space according to the memory storage starting address and length information in the data tag. Therefore, the volume of the topic will be very small, far less than the average bandwidth of Ethernet, and the storage of the original data makes full use of the high-speed data channel provided by the PCIe switch and the advantages of the shared storage space, improving the communication efficiency.
[0058] In some embodiments, data transmission in the shared storage space through the PCIe switch includes: The publishing node writes the target data into the shared storage space. In the case of successful writing, it encapsulates the data tag of the target data in the target topic and transmits the target topic to the subscription node through the message middleware; Among them, the publishing node is the data writing party among the root port and the end nodes.
[0059] In this embodiment, the data transmission process may include data writing and message sending steps.
[0060] The data transmission process is as follows Figure 3 As shown, the publishing node writes the target data into the shared storage space through the PCIe switch according to the mapping relationship determined by the base address register. After the data is successfully written into the shared storage space, the publishing node generates a write success signal and encapsulates the write success signal in the target topic. The target topic includes data tags, such as key information like the memory start address, data length, encryption type, etc. of the data in the shared storage space. The publishing node transmits the target topic to the subscribing node through DDS using the Ethernet switch.
[0061] In this embodiment, by writing the data into the shared storage space and using the PCIe switch for transmission, the high bandwidth and low latency characteristics of PCIe are fully utilized, enabling the data to be quickly transmitted between devices. The encapsulation of data tags and the transmission mechanism of DDS topics can transmit key information without directly transmitting a large amount of raw data, reducing the occupancy of the Ethernet bandwidth and improving the efficiency of data transmission.
[0062] In some embodiments, data transmission in the shared storage space through the PCIe switch includes: The subscribing node receives the target topic transmitted by the message middleware and reads the target data from the shared storage space according to the data tags in the target topic; Among them, the subscribing node is the data reading party among the root port and the end node.
[0063] In this embodiment, as Figure 3 shown, after the subscribing node receives the target topic containing data tags, it reads the target data from the shared storage space according to the information in the data tags.
[0064] Since the data tags contain accurate memory address and data length information, the subscribing node can directly access the shared storage space quickly through the PCIe switch to read the required data.
[0065] In this embodiment, the high bandwidth characteristic of the PCIe switch enables the data to be quickly written and read from the shared storage space, meeting the strict real-time requirements of the autonomous driving system. The encapsulation of data tags and the transmission mechanism of DDS topics can transmit key information without directly transmitting a large amount of raw data, reducing the occupancy of the Ethernet bandwidth and improving the efficiency of data transmission.
[0066] In some embodiments, data transmission in the shared storage space through the PCIe switch includes: In the case of data transmission related to the graphics processor task, the root port and the end node use the graphics processor direct communication technology to perform data transmission in the shared storage space through the PCIe switch.
[0067] Graphics processing units (GPUs) were originally designed specifically for graphics rendering. With their powerful parallel computing capabilities, they are now widely used in fields such as deep learning and computer vision. In vehicle scenarios, GPU tasks can include various computing tasks ranging from basic image rendering to complex model inference. For example, autonomous vehicles need to process image data collected by multiple cameras in real time, and perform image recognition, object detection, and scene segmentation through GPUs. These all belong to GPU tasks. Model inference tasks, on the other hand, utilize trained deep learning models to analyze and predict sensor data, and determine the positions and motion states of targets such as pedestrians, vehicles, and obstacles in the vehicle's surrounding environment, thereby providing a basis for autonomous driving decisions.
[0068] Data related to GPU tasks is characterized by large data volumes and high real-time requirements. Taking autonomous driving as an example, multiple high-definition cameras installed in the vehicle generate a large amount of image data per second, and sensors such as lidar also generate a vast amount of point cloud data. When processing these data for GPU tasks, they need to be quickly transmitted and processed so that the vehicle can respond in a timely manner.
[0069] GPU Direct is a technology that allows direct data transfer between GPUs or between GPUs and other devices. Data transfer no longer depends on the central processing unit for data forwarding, but instead directly establishes a data transfer channel between devices through a high-speed PCIe bus.
[0070] When performing data transfer related to GPU tasks, the root port and end nodes can use GPU Direct to transfer data in the shared storage space. Specifically, when the end node obtains data related to GPU tasks, it directly writes the data into a pre-allocated area in the shared storage space through a PCIe switch without going through CPU processing. The root port or other nodes that need to use this data can directly read the data from the shared storage space through GPU Direct, and utilize the high bandwidth and low latency characteristics of PCIe during the transfer process to achieve fast data acquisition.
[0071] In this embodiment, GPU Direct reduces the intermediate links in data transfer, reduces the steps of multiple copies of data between the host memory and GPU memory, thereby reducing latency and improving data transfer efficiency.
[0072] In this embodiment, the GPU task includes a model inference task. Transmitting data related to the model inference task through GPU Direct can complete data transfer and processing faster, thereby improving the real-time performance and reliability of the autonomous driving system.
[0073] In some embodiments, the direct communication technology of the graphics processor is adopted to transfer data in the shared storage space through a PCIe switch, including: The data writer in the root port and the end node writes data into the shared storage space through the graphics processor.
[0074] In this embodiment, when the upper-layer application initiates a data processing request related to the graphics processor task, the type and requirements of the task can be analyzed. If it is determined that the graphics processor needs to be used for processing, such as the inference output of a deep learning model, the image processing result, or the feature extraction result of sensor data, the corresponding graphics processor core and the shared storage space area can be allocated. For example, when processing the image recognition task in autonomous driving, a specific CUDA (Compute Unified Device Architecture) core of the graphics processor can be allocated for the camera data for processing, and a continuous memory area can be reserved in the shared storage space to store the intermediate data and the final result during the processing.
[0075] The graphics processors in the root port and the end node can establish a direct communication channel through the PCIe switch. After the communication channel is established, the graphics processor of the data writer starts to write data into the shared storage space.
[0076] In this embodiment, by combining the direct communication technology of the graphics processor and PCIe during the data transfer process, the graphics processors can quickly write data to each other, improving the communication efficiency of the system and providing strong support for the complex computing tasks in the autonomous driving framework.
[0077] In some embodiments, the direct communication technology of the graphics processor is adopted to transfer data in the shared storage space through a PCIe switch, including: The data reader in the root port and the end node reads data from the shared storage space through the graphics processor.
[0078] In this embodiment, when the upper-layer application needs to access the data in the shared storage space, a read instruction will be sent to the graphics processor of the data reader. The read can include data tag information, such as the start address, length, and format of the data in the shared storage space, etc. The graphics processor of the data reader identifies the location and characteristics of the required data by parsing these tags, and transfers the data from the shared storage space to the local memory.
[0079] In an example, the data transfer process of the data writer in the root port and the end node through the graphics processor is as Figure 4 shown.
[0080] In this example, taking the edge computing device as a system-on-chip, the system-on-chip may include a central processing unit and a graphics processing unit.
[0081] In the traditional solution, taking the topic as a carrier, the data is transmitted and relayed to the graphics processing unit through the central processing unit.
[0082] In the embodiment of the present application, the original data is no longer encapsulated in the topic, but the data tag is encapsulated, and the message middleware is still used to transmit the topic. According to the data tag in the topic, the graphics processing unit can directly read the data from the shared storage space through the graphics processing unit direct communication technology combined with PCIe.
[0083] As Figure 4 shown, the graphics processing unit executes complex model inference tasks. Models N-1, N, N+1, etc. can represent different deep learning models for performing different tasks. The graphics processing unit can communicate with the model through the data tag, perform inference and output the result.
[0084] In this embodiment, by combining the graphics processing unit direct communication technology and PCIe during the data transmission process, the graphics processing units can quickly read data from each other, reducing the waiting time for data transmission and improving the data processing efficiency.
[0085] In some embodiments, the shared storage space includes a central processing unit storage space and a graphics processing unit storage space; The central processing unit storage space is used to store data related to the central processing unit tasks; The graphics processing unit storage space is used to store data related to the graphics processing unit tasks.
[0086] In this embodiment, the shared storage space can be a multi-functional storage area, including a central processing unit storage space and a graphics processing unit storage space.
[0087] The central processing unit storage space is used to store data related to the central processing unit tasks. The central processing unit tasks are tasks processed by the central processing unit of the edge computing device, including tasks with strong logic and high sequentiality, such as system control, sensor data preprocessing, path planning, etc. The data related to the central processing unit tasks can include control instructions, sensor sampling frequency configuration, etc.
[0088] In this embodiment, when the root port or end node is the data writer, the corresponding storage space can be automatically selected according to the data type. For data related to central processing unit (CPU) tasks, the data writer can write the data into the CPU storage space through a PCIe switch. For data related to graphics processing unit (GPU) tasks, the GPU can directly write the data into the GPU storage space by using the GPU direct communication technology.
[0089] In this embodiment, by dividing the shared storage space into a CPU storage space and a GPU storage space, different types of data can be separated and stored, reducing the problem of conflicts generated during the data transmission process between the CPU and the GPU, and improving the resource utilization efficiency.
[0090] In some embodiments, the in-vehicle communication system based on a distributed architecture may further include: a message scheduling layer, a data transmission layer, and a model inference layer; the message scheduling layer and the data transmission layer transmit data related to CPU tasks, and the model inference layer transmits data related to the GPU; Among them, the message scheduling layer performs message transmission by combining Ethernet and a message middleware; the data processing layer performs data transmission by combining PCIe and direct memory access; the model inference layer performs data transmission by combining PCIe and the GPU direct communication technology.
[0091] In this embodiment, as Figure 5 shown, the in-vehicle communication system based on a distributed architecture may further include: a message scheduling layer, a data transmission layer, and a model inference layer, and each layer has specific functions and transmission mechanisms to adapt to different types of processor tasks and data transmission requirements.
[0092] The message scheduling layer is mainly responsible for processing data related to CPU tasks and can perform message transmission by combining Ethernet and a message middleware. The role of the message scheduling layer is to manage and schedule the message flow in the system. By using the wide compatibility of Ethernet and the efficient transmission mechanism of the message middleware, reliable transmission of CPU task data is achieved. In the message scheduling layer, data is usually encapsulated and transmitted in the form of messages with topics as carriers, and the messages contain data tags.
[0093] The data transmission layer performs data transmission by combining PCIe and direct memory access technology. The data transmission layer focuses on achieving fast and efficient data transmission, especially for CPU tasks that require a large amount of data transfer, such as the original state data collected by sensors. PCIe provides a high-bandwidth and low-latency hardware channel, and the direct memory access technology allows data to be directly transmitted between the memory and the device without the intervention of the CPU, thereby reducing the burden on the CPU and improving the data transmission efficiency.
[0094] The model inference layer is specifically designed to process data related to GPU tasks. The model inference layer performs data transmission by combining PCIe and direct communication technology between the GPU and the system memory. The main responsibility of the model inference layer is to support the GPU in performing efficient model inference tasks, such as the execution of deep learning models. Through PCIe and direct communication technology between the GPU and the system memory, data can be directly transmitted between the GPU and the shared storage space without passing through the central processing unit (CPU) or the host memory, thus achieving high-speed data transmission and processing.
[0095] In this embodiment, through this hierarchical distributed architecture, optimized support for different types of processor tasks is achieved. Specifically, the message scheduling layer improves the efficiency of message scheduling by combining Ethernet and message middleware; the data transmission layer realizes fast and efficient data transmission by combining PCIe and direct memory access technology, reducing the burden on the CPU; the model inference layer realizes high-speed data transmission and processing of GPU task data by combining PCIe and direct communication technology between the GPU and the system memory, providing support for complex computing tasks in the autonomous driving system.
[0096] An embodiment of the present application also provides a vehicle, including the in-vehicle communication system based on the distributed architecture described above.
[0097] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0098] The above has provided a detailed introduction to an in-vehicle communication system and a vehicle based on a distributed architecture according to the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A vehicle-mounted communication system based on a distributed architecture, characterized in that, Comprising: Multiple edge computing devices; wherein, at least one edge computing device is configured as a root port, and at least one edge computing device is configured as an end node; A PCIe switch, PCIe-connected to the root port and the end node respectively; Wherein, the root port includes a shared storage space, and the root port and the end node perform data transmission in the shared storage space through the PCIe switch.
2. The system according to claim 1, characterized in that, Before performing data transmission in the shared storage space through the PCIe switch, it includes: The end node maps the physical address of the end node to the shared storage space of the root port through a base address register, so that the end node can determine the physical address of the end node through the shared storage space.
3. The system according to claim 1, wherein The data transmission in the shared storage space through the PCIe switch includes: The data writing party among the root port and the end node performs binary serialization on the data to obtain serialized data; compresses the serialized data, and writes the compressed serialized data into the shared storage space.
4. The system according to claim 3, characterized in that, The data transmission in the shared storage space through the PCIe switch includes: The data reading party among the root port and the end node reads the stored data in the shared storage space, and restores the read data through deserialization.
5. The system according to claim 1, wherein The system further includes: An Ethernet switch, Ethernet-connected to the root port and the end node respectively; A message middleware, used for message transmission between the root port and the end node and between different end nodes through the Ethernet switch.
6. The system according to claim 5, characterized in that The message transmission through the Ethernet switch includes: Using a topic as the carrier of message transmission, and performing message transmission by encapsulating data tags in the topic; the data tags characterize the attributes of the data.
7. The system according to claim 6, wherein The data tags include at least one of data encryption type, memory storage start address, and memory occupancy length.
8. The system according to claim 5, wherein The data transmission in the shared storage space through the PCIe switch includes: The publishing node writes the target data into the shared storage space, and in the case of successful writing, encapsulates the data tag of the target data in the target topic, and transmits the target topic to the subscribing node through the message middleware; Wherein, the publishing node is the data writing party among the root port and the end node.
9. The system according to claim 8, wherein The data transmission in the shared storage space through the PCIe switch includes: The subscribing node receives the target topic transmitted by the message middleware, and reads the target data from the shared storage space according to the data tag in the target topic; Wherein, the subscribing node is the data reading party among the root port and the end node.
10. The system according to claim 1, wherein The data transmission in the shared storage space through the PCIe switch includes: In the case of performing data transmission related to a graphics processor task, the root port and the end node use a graphics processor direct communication technology to perform data transmission in the shared storage space through the PCIe switch.
11. The system according to claim 10, wherein The data transmission in the shared storage space by using the direct communication technology of the graphics processor through the PCIe switch includes: The data writing parties in the root port and the end nodes write data into the shared storage space through the graphics processor.
12. The system according to claim 11, characterized in that, The data transmission in the shared storage space by using the direct communication technology of the graphics processor through the PCIe switch includes: The data reading parties in the root port and the end nodes read data from the shared storage space through the graphics processor.
13. The system according to claim 11, wherein The graphics processor tasks include model inference tasks.
14. The system according to claim 1, wherein The shared storage space includes a central processing unit storage space and a graphics processor storage space; The central processing unit storage space is used to store data related to central processing unit tasks; The graphics processor storage space is used to store data related to graphics processor tasks.
15. A vehicle, characterized in that, It includes the vehicle-mounted communication system based on the distributed architecture according to any one of claims 1-14.
Citation Information
Patent Citations
PCIe-based vehicle resource sharing method and device, equipment, medium and vehicle
CN117851298A
Information processing method and system
CN117978828A
Vehicle-mounted communication system and vehicle
CN118101684A
GPU cross-host communication interconnection system based on PCIe NTB
CN119149480A
Computing device with ethernet connectivity for virtual machines on several systems on a chip
EP3992791A1
Cited By
Vehicle-mounted communication method and system
CN121056486A
Data interaction method, platform, device, medium and program product
CN121262546A
Data interaction method, platform, device, medium and program product
CN121262546B
Data network interaction system for rail transit data transmission
CN121567706A