In-vehicle communication system and vehicle based on distributed architecture

By adopting a distributed architecture and PCIe switches in the in-vehicle communication system, edge computing devices are configured as root ports and end nodes to achieve high-speed data transmission, solving the problem of large-scale data throughput congestion under Ethernet communication and improving communication efficiency.

CN120358213BActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510855104.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-12
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When Ethernet communication is interconnected and large amounts of data are being processed, congestion is likely to occur, resulting in decreased communication efficiency.

Method used

A distributed architecture is adopted, multiple edge computing devices are configured as root ports and end nodes, and PCIe switches are used to achieve high-speed connections between root ports and end nodes. Data is transmitted through shared storage space, reducing intermediate links and delays.

Benefits of technology

It improves data transmission efficiency, ensures the smoothness of the on-board computing platform when processing large amounts of data, solves the problem of large-scale data throughput congestion under Ethernet communication interconnection, and improves communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358213B_ABST
    Figure CN120358213B_ABST
Patent Text Reader

Abstract

The present application discloses an on-vehicle communication system and vehicle based on a distributed architecture, which relates to the field of communication technology, including configuring multiple edge computing devices as root ports and end nodes respectively, and using PCIe switches to achieve high-speed connections between the root ports and end nodes, which can fully utilize the high bandwidth and low latency characteristics of PCIe, greatly improve data transmission efficiency, thereby ensuring the smoothness of the on-vehicle computing platform when processing large amounts of data. The root port sets a shared storage space, so that the root port and the end node can directly transmit data in the shared storage space through the PCIe switch, further reducing the intermediate links of data transmission and reducing transmission delay. Therefore, it can solve the problem that blocking is easy to occur in the case of large-scale data throughput under Ethernet communication interconnection, resulting in reduced communication efficiency, and achieve the technical effect of improving communication efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to an in-vehicle communication system and a vehicle based on a distributed architecture. Background Art

[0002] With the advancement of technology, autonomous driving frameworks have shifted from modularity to integrated perception and planning, becoming the industry's mainstream design approach. Vehicle-side computing platforms are increasingly constrained by practical factors such as power supply, heat dissipation, security, and deployment costs. In-vehicle computing platforms typically utilize multiple system-on-chips to generate high computing power to support efficient and stable reasoning in the upper-layer framework.

[0003] In related technologies, the interconnection between multiple system-level chips is achieved through Ethernet, which is prone to congestion in the case of large-scale data throughput, resulting in reduced communication efficiency. Summary of the Invention

[0004] The present application provides an in-vehicle communication system and vehicle based on a distributed architecture, so as to at least solve the problem in the related art that congestion is prone to occur when large amounts of data are throughputed under Ethernet communication interconnection, resulting in reduced communication efficiency.

[0005] This application provides a distributed architecture-based vehicle communication system, including:

[0006] A plurality of edge computing devices; wherein at least one edge computing device is configured as a root port and at least one edge computing device is configured as an end node;

[0007] A PCIe switch is connected to the root port and the end node PCIe respectively;

[0008] The root port includes a shared storage space, and the root port and the end node perform data transmission in the shared storage space through the PCIe switch.

[0009] The present application also provides a vehicle, comprising: the above-mentioned vehicle-mounted communication system based on distributed architecture.

[0010] Through this application, by configuring multiple edge computing devices as root ports and end nodes respectively, and using PCIe switches to achieve high-speed connections between root ports and end nodes, it is possible to fully utilize the high bandwidth and low latency characteristics of PCIe, greatly improving data transmission efficiency, thereby ensuring the smoothness of the on-board computing platform when processing large amounts of data. The root port sets a shared storage space, so that the root port and the end node can directly transmit data in the shared storage space through the PCIe switch, further reducing the intermediate links in data transmission and reducing transmission delay. Therefore, it can solve the problem of easy congestion in the case of large-scale data throughput under Ethernet communication interconnection, resulting in reduced communication efficiency, and achieve the technical effect of improving communication efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 Schematic diagram of the architecture of a distributed-based in-vehicle communication system provided in an embodiment of the present application;

[0013] Figure 2 This is a schematic diagram of the physical address mapping process provided by an embodiment of the present application;

[0014] Figure 3 This is a schematic diagram of the data transmission process between the publishing node and the subscribing node provided in an embodiment of the present application;

[0015] Figure 4 Schematic diagram of the data transmission process of the graphics processor provided in an embodiment of the present application;

[0016] Figure 5 It is a structural diagram of the communication logic provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0019] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0020] In conjunction with the specific application environment architecture or specific hardware architecture that the in-vehicle communication system based on the distributed architecture relies on, the specific application environment architecture or specific hardware architecture is described here.

[0021] The embodiments of the present application provide a vehicle-mounted communication system based on a distributed architecture, and the system is described in detail in conjunction with the architecture of the vehicle-mounted communication system based on the distributed architecture.

[0022] Specifically, Figure 1 This is a schematic diagram of the architecture of an in-vehicle communication system based on a distributed architecture provided according to an embodiment of the present application.

[0023] like Figure 1 As shown in the figure, the in-vehicle communication system based on a distributed architecture includes multiple edge computing devices and PCIe (Peripheral Component Interconnect Express) switches.

[0024] Edge computing devices are devices with powerful computing capabilities. Installing edge computing devices in vehicle communication systems can shift computing tasks from a central cloud or data center to a location closer to the data source, allowing data processing to be performed closer to the data source. This reduces data transmission pressure and latency, and enhances the vehicle's intelligence and automation. In some embodiments, the edge computing device used can be a system-on-chip (SoC), such as the Nvidia Drive Orin device. Each edge computing device can include an independent CPU (Central Processing Unit) and GPU (Graphics Processing Unit).

[0025] In vehicle-mounted scenarios, different edge computing devices can work together in a division of labor. For example, some edge computing devices are responsible for processing image data collected by cameras, while others are responsible for processing data from sensors such as radars, enabling distributed computing. This improves the overall computing power level and jointly provides support for autonomous driving decisions.

[0026] like Figure 1 As shown, the plurality of edge computing devices may include edge computing device 101, edge computing device 102, edge computing device 103 and edge computing device 104. It should be noted that, Figure 1 The number of edge computing devices is for illustrative purposes only. The number of edge computing devices can be 2, 3, 5 or any other number, which is not limited in the embodiments of the present application.

[0027] Edge computing devices 101, 102, 103, and 104 communicate with each other via PCIe connections. PCIe is a high-speed serial computer expansion bus standard widely used in computers and servers. PCIe boasts high bandwidth and low latency, enabling rapid transmission of large amounts of data. Compared to traditional Ethernet, PCIe offers higher data transmission efficiency at the hardware level, particularly when handling high-density data streams. PCIe connections leverage the high bandwidth and low latency of the PCIe bus, enabling fast and efficient data transmission between edge computing devices.

[0028] In a distributed architecture-based in-vehicle communication system, multiple edge computing devices are divided into root ports (RP) and end nodes (EP). Figure 1 In the example, edge computing device 101 is a root port, and edge computing device 102, edge computing device 103, and edge computing device 104 are end nodes.

[0029] The root port is the core node of the entire communication architecture, responsible for managing and coordinating data transmission across the entire system. The end node is the terminal device for data transmission, responsible for performing specific computing tasks or data processing. By dividing edge computing devices into root ports and end nodes, data transmission can be organized more efficiently, ensuring an orderly flow of data between different devices.

[0030] Since PCIe direct connection can only achieve point-to-point connection, in order to achieve high-speed connection between the root port and the end node, the embodiment of the present application adopts a PCIe switch. A PCIe switch is a device specifically used to manage PCIe connections, including multiple PCIe interfaces. The root port and the end node can be connected to the PCIe switch respectively through the PCIe interface. The root port and the end node can perform high-speed data transmission through the PCIe interface. The core function of the PCIe switch is to achieve fast forwarding and routing of data, so that data can be efficiently transmitted on different PCIe links. By using a PCIe switch, flexible interconnection between multiple edge computing devices is achieved, and a stable and reliable distributed communication network is built.

[0031] In the embodiments of the present application, the root port includes a shared memory space. This shared memory space is a memory space of a certain size within the root port that allows the root port and end nodes to directly read and write data within this space, thereby enabling DMA (Direct Memory Access) interconnection between the root port and end nodes. DMA is a technology that allows certain hardware subsystems to directly access system memory without CPU control, significantly improving data transfer efficiency and reducing CPU burden.

[0032] In an embodiment of the present application, the root port and the end node can perform data transmission in a shared storage space through a PCIe switch. Specifically, the end node can write the processed data into the shared storage space through DMA, and the root port can read the data from the shared storage space through DMA for further processing or forwarding to other end nodes. This method of directly exchanging data in the shared storage space fully utilizes the low latency characteristics of PCIe, so that the reading and writing between the root port and the end node in the shared memory space can bypass the CPU access, and realize point-to-point zero-copy direct access to data between devices. This method further reduces the intermediate links in data transmission and reduces transmission delay.

[0033] This application, by configuring multiple edge computing devices as root ports and end nodes respectively and using PCIe switches to achieve high-speed connections between the root ports and end nodes, can fully utilize the high bandwidth and low latency characteristics of PCIe, greatly improving data transmission efficiency, thereby ensuring the smoothness of the on-board computing platform when processing large amounts of data. The root port sets up a shared storage space, allowing the root port and end nodes to directly transmit data in the shared storage space through the PCIe switch, further reducing the intermediate links in data transmission and reducing transmission latency. Therefore, it can solve the problem of congestion that is prone to occur in the case of large-scale data throughput under Ethernet communication interconnection, resulting in reduced communication efficiency, and achieve the technical effect of improving communication efficiency.

[0034] In some embodiments, before data is transmitted in the shared storage space through the PCIe switch, the following steps are included:

[0035] The end node maps the physical address of the end node to the shared memory space of the root port through the base address register, so that the end node can determine the physical address of the end node through the shared memory space.

[0036] In this embodiment, the base address register (BAR) is an important component in the PCIe architecture, which is used to store the physical address range information of the device and allow devices to access each other's memory space through mapping.

[0037] In this embodiment, the end node configures the base address register during the initialization phase, and defines the mapping location of the end node's memory space in the root port shared memory space through the base address register. Figure 2 As shown in the figure, when an end node boots up or connects to the system, it sends a configuration request to the root port via the PCIe bus, carrying its base address register information. After receiving this information, the root port allocates a corresponding address space for the end node in the shared memory space and records the address mapping. In this way, the end node's physical address is mapped to the root port's shared memory space.

[0038] After address mapping is complete, the end node can determine its own physical address through the shared storage space. When the end node needs to transmit data, it only needs to access the shared storage space according to the mapped address, without having to worry about the specific location of the physical address.

[0039] In this embodiment, end nodes use base address registers to map their physical addresses to the root port's shared memory space, enabling accurate access to the shared memory space and improving data transmission accuracy. The root port also optimizes data transmission paths and improves communication performance by centrally managing address mappings to the shared memory space.

[0040] In some embodiments, transmitting data in a shared storage space through a PCIe switch includes:

[0041] The data writers in the root port and the end node perform binary serialization on the data to obtain serialized data; compress the serialized data, and write the compressed serialized data into the shared storage space.

[0042] In this embodiment, the data writer is the one of the root port and the end node that writes data into the shared storage space.

[0043] Serialization is the process of converting a data structure or object state into a storable or transferable format, while binary serialization is the process of converting a data structure or object state into a binary format. For example, in ROS2 (Robot Operating System 2), the data writer can call the serialize_message() method in the C++ client library (rclcpp) to convert a Topic message in ROS2 format into a vector<uint8_t> Type. For example, a Topic that has registered a message format in ROS2 can be converted into binary serialized data t_binary, such as 0110001011110101.

[0044] However, since the underlying BAR0 (base address register 0) is transmitted in 32-bit hexadecimal format in units of page memory, directly transmitting binary serialized data may result in a low semantic density of the message carrier. To improve transmission efficiency, lossless compression can be performed on the serialized data.

[0045] Lossless compression is a compression method that can completely restore the original data. It reduces data size by eliminating redundant information. In one example, after lossless compression, the effective length of data can be compressed to 1 / 8 of its original length, and the semantic density of the compressed data t_hex is 8 times that of the uncompressed t_binary. Compression not only reduces data bandwidth requirements but also improves data transmission speed and efficiency.

[0046] After serialization and compression are completed, the data writer can write the processed data to the register address of the root port through memory copy, thereby realizing DMA transmission.

[0047] In this embodiment, complex data can be converted into a binary sequence through serialization operations to facilitate data transmission and processing between different devices, thereby improving communication compatibility; by compressing the serialized data, not only the bandwidth requirement for data transmission is reduced, but also the semantic density of the message carrier is increased.

[0048] In some embodiments, transmitting data in a shared storage space through a PCIe switch includes:

[0049] The data readers in the root port and the end node read the data stored in the shared storage space and restore the read data through deserialization.

[0050] In this embodiment, the data reader is the one of the root port and the end node that reads data from the shared storage space.

[0051] In this embodiment, the data reader can read data from the shared memory space by DMA reading. Taking the data reader as an end node as an example, the data stored in the shared memory space can be read from the register address of the end node by memory copying.

[0052] Because the data writer stores compressed data, such as compressed data t_hex in hexadecimal format, the data reader needs to deserialize the hexadecimal data back into binary format. For example, this can be done by using bitwise operations and encoding rules to convert the t_hex data back into the binary sequence t_binary. The serialize_message() method in the ROS2 C++ client library (rclcpp) can then be called to convert the binary sequence t_binary into the ROS2 format. This enables ROS2 application designs based on PCIE communication for different edge computing technologies.

[0053] In this embodiment, the integrity and semantic consistency of data during transmission are improved through the deserialization operation, so that the data reader can correctly understand and process the received data, thereby improving the data transmission efficiency between different devices.

[0054] In some embodiments, the in-vehicle communication system based on a distributed architecture may further include:

[0055] Ethernet switch 130, connected to the root port and the end node Ethernet;

[0056] The message middleware is used to transmit messages between the root port and the end node, and between different end nodes through the Ethernet switch 130 .

[0057] In this embodiment, in order to further enhance the communication capability of the in-vehicle communication system based on a distributed architecture, an Ethernet switch 130 may be further provided in the in-vehicle communication system based on a distributed architecture.

[0058] The Ethernet switch 130 is a network device that can realize Ethernet connection between multiple devices. As a widely used network technology, Ethernet has good compatibility and scalability and can support various types of devices and communication protocols.

[0059] like Figure 1 As shown, in this embodiment, the Ethernet switch 130 establishes Ethernet connections with the root port and the end node respectively, so that data can be transmitted between the root port and the end node via Ethernet.

[0060] In this embodiment, in addition to the hardware connectivity provided by the Ethernet switch, message middleware is also introduced. Message middleware is a software component used to transmit messages between different devices. In some embodiments, DDS (Data Distribution Service) can be used as message middleware. DDS is a high-performance middleware protocol that meets the extremely high real-time and reliability requirements of application scenarios such as autonomous driving.

[0061] In this embodiment, the message middleware is responsible for transmitting messages between root ports and end nodes, as well as between different end nodes, through Ethernet switches. The message middleware uses a publish-subscribe or point-to-point communication model, allowing for loosely coupled interaction between devices. For example, an end node can act as a message publisher, publishing specific types of data to the message middleware; a root port or other end node can act as a subscriber, subscribing to messages of interest. The message middleware is responsible for receiving, storing, and forwarding these messages.

[0062] In this embodiment, the high-speed data transmission channel provided by the PCIe switch is suitable for large-scale, low-latency data transmission scenarios, such as real-time transmission and processing of sensor data; while the Ethernet switch and message middleware provide flexibility and scalability for communication, which is suitable for scenarios such as control instruction transmission and configuration information exchange, reducing the problem of resource waste caused by PCIe occupation.

[0063] This embodiment leverages the advantages of both PCIe and Ethernet technologies through a distributed communication architecture. PCIe's high bandwidth and low latency improve the transmission efficiency of critical data, meeting the stringent real-time requirements of autonomous driving systems. The intelligent coordination of Ethernet and message-based middleware enables the communication process to adapt to different data transmission requirements, supporting diverse application scenarios, improving system scalability and maintainability, and reducing development and deployment costs.

[0064] In some embodiments, transmitting a message through an Ethernet switch includes:

[0065] The topic is used as the carrier of message transmission, and the message transmission is carried out by encapsulating data tags in the topic; the data tags represent the attributes of the data.

[0066] In the traditional DDS communication model, topics typically directly encapsulate encoded raw data, such as image data in the sensor_msgs::msg::Image format or point cloud data in the sensor_msgs::msg::PointCloud2 format. However, with the advancement of autonomous driving technology, the number and accuracy of vehicle-mounted sensors are continuously improving, and the amount of data generated is exploding. DDS only supports Ethernet for inter-device communication, making it prone to congestion when handling large data transmission tasks due to the actual Ethernet bandwidth. This can cause communication delays or loss of important data frames.

[0067] Taking an entry-level autonomous commercial vehicle as an example, it is equipped with six 2-megapixel cameras and one 128-line lidar. When collecting data at a rate of 30Hz, the data generation rate is as high as 26.4Gb / s, far exceeding the effective transmission capacity of general-purpose automotive 10Gb / s Ethernet, which can easily lead to message congestion and data frame loss.

[0068] This embodiment optimizes the traditional communication model, retaining DDS and still using the Ethernet + DDS architecture and publish-subscribe model. However, topics no longer directly encapsulate raw data, but instead encapsulate data tags. Data tags contain key data attributes, such as data encryption type, memory storage address, and memory usage.

[0069] In an example, a data label can be defined as:

[0070] {

[0071] string name; message name

[0072] string type; message type

[0073] long start; message first pointer address

[0074] long size; message length

[0075] }

[0076] This significantly reduces the amount of information carried by topics, significantly reducing the burden on Ethernet transmission. For example, topics that originally required the transmission of large amounts of image or point cloud data now only require the transmission of short messages containing data tags, greatly improving message transmission efficiency.

[0077] The raw data itself is written to shared storage via PCIe+DMA. When a subscribing node receives a topic containing a data tag, it can read the corresponding raw data directly from shared storage based on the memory storage address and length information in the data tag. As a result, the topic size is very small, far less than the average Ethernet bandwidth. The storage of raw data fully utilizes the high-speed data channels and shared storage space provided by the PCIe switch, improving communication efficiency.

[0078] In some embodiments, transmitting data in a shared storage space through a PCIe switch includes:

[0079] The publishing node writes the target data into the shared storage space. If the write is successful, the target data's data tag is encapsulated in the target topic and the target topic is transmitted to the subscribing node through the message middleware.

[0080] The publishing node is the data writer in the root port and end node.

[0081] In this embodiment, the data transmission process may include data writing and message sending steps.

[0082] The data transmission process is as follows Figure 3 As shown in the figure, the publishing node writes the target data to the shared storage space via the PCIe switch, based on the mapping relationship determined by the base address register. After the data is successfully written to the shared storage space, the publishing node generates a write success signal and encapsulates the write success signal in the target topic. The target topic includes key information such as the data's first memory address in the shared storage space, data length, and encryption type. The publishing node transmits the target topic to the subscribing node via DDS and an Ethernet switch.

[0083] In this embodiment, by writing data into a shared storage space and transmitting it through a PCIe switch, the high bandwidth and low latency characteristics of PCIe are fully utilized, so that data can be quickly transmitted between devices. The encapsulation of data tags and the transmission mechanism of DDS topics can deliver key information without directly transmitting a large amount of original data, thereby reducing the Ethernet network bandwidth usage and improving the efficiency of data transmission.

[0084] In some embodiments, transmitting data in a shared storage space through a PCIe switch includes:

[0085] The subscription node receives the target topic transmitted by the message middleware and reads the target data from the shared storage space according to the data tag in the target topic;

[0086] The subscribing nodes are the data readers in the root ports and end nodes.

[0087] In this embodiment, if Figure 3 As shown, after the subscription node receives the target topic containing the data tag, it reads the target data from the shared storage space according to the information in the data tag.

[0088] Since the data tag contains accurate memory address and data length information, the subscribing node can quickly access the shared storage space directly through the PCIe switch and read the required data.

[0089] In this embodiment, the high-bandwidth characteristics of the PCIe switch enable data to be quickly written to and read from shared storage space, meeting the strict real-time requirements of the autonomous driving system. The encapsulation of data tags and the transmission mechanism of DDS topics can transmit key information without directly transmitting large amounts of raw data, reducing the use of Ethernet network bandwidth and improving the efficiency of data transmission.

[0090] In some embodiments, transmitting data in a shared storage space through a PCIe switch includes:

[0091] In the case of data transmission related to GPU tasks, the root port and the end node use GPU direct communication technology to transmit data in the shared storage space through the PCIe switch.

[0092] Originally designed for graphics rendering, graphics processors (GPUs) are now widely used in fields such as deep learning and computer vision due to their powerful parallel computing capabilities. In automotive scenarios, GPU tasks can range from basic image rendering to complex model reasoning. For example, autonomous vehicles need to process image data collected by multiple cameras in real time, using the GPU for image recognition, object detection, and scene segmentation. These tasks fall under the category of GPUs. Model reasoning, on the other hand, utilizes trained deep learning models to analyze and predict sensor data, determining the position and motion of pedestrians, vehicles, obstacles, and other objects in the vehicle's surroundings, thereby providing a basis for autonomous driving decisions.

[0093] The data associated with GPU tasks is characterized by high volume and high real-time requirements. For example, in autonomous driving, the vehicle's multiple high-definition cameras generate a massive amount of image data every second, and sensors like LiDAR also generate massive amounts of point cloud data. This data must be transmitted and processed quickly when GPU tasks are being processed to enable the vehicle to respond promptly.

[0094] GPU Direct (GPU Direct) allows direct data transfer between GPUs, or between GPUs and other devices. Instead of relying on a central processing unit (CPU) for data transfer, data transfer is established directly between devices via the high-speed PCIe bus.

[0095] When transferring data related to GPU tasks, the root port and end nodes can use GPU Direct Communication technology to transfer data in shared memory. Specifically, after an end node obtains data related to a GPU task, it writes the data directly to a pre-allocated area in shared memory via the PCIe switch, without CPU processing. The root port or other nodes that need to access this data can read the data directly from shared memory using GPU Direct Communication technology, leveraging PCIe's high bandwidth and low latency to achieve rapid data retrieval.

[0096] In this embodiment, the graphics processor direct communication technology reduces the intermediate links in data transmission and reduces the steps of multiple data copying between the host memory and the GPU memory, thereby reducing latency and improving data transmission efficiency.

[0097] In this embodiment, the graphics processor tasks include model reasoning tasks. Transmitting data related to the model reasoning tasks through graphics processor direct communication technology can complete data transmission and processing more quickly, thereby improving the real-time performance and reliability of the autonomous driving system.

[0098] In some embodiments, GPU direct communication technology is used to transmit data in a shared storage space through a PCIe switch, including:

[0099] Data writers in the root port and end node write data to the shared storage space through the graphics processor.

[0100] In this embodiment, when an upper-layer application initiates a data processing request related to a GPU task, the type and requirements of the task can be analyzed. If it is determined that the GPU is required for processing, such as the inference output of a deep learning model, image processing results, or feature extraction results from sensor data, the corresponding GPU core and shared memory space area can be allocated. For example, when processing image recognition tasks in autonomous driving, a specific GPU CUDA (Compute Unified Device Architecture) core can be assigned to process camera data, and a contiguous memory area can be reserved in the shared memory space to store intermediate data and final results during the processing process.

[0101] The root port and the GPU in the end node can establish a direct communication channel through the PCIe switch. Once the communication channel is established, the GPU writing the data begins writing the data to the shared storage space.

[0102] In this embodiment, by combining graphics processor direct communication technology and PCIe during the data transmission process, data can be written quickly between graphics processors, improving the communication efficiency of the system and providing strong support for complex computing tasks in the autonomous driving framework.

[0103] In some embodiments, GPU direct communication technology is used to transmit data in a shared storage space through a PCIe switch, including:

[0104] Data readers in the root port and end node read data from the shared memory space through the graphics processor.

[0105] In this embodiment, when an upper-layer application needs to access data in the shared storage space, it sends a read instruction to the GPU on the data reader. The read instruction may include data tag information, such as the data's starting address, length, and format in the shared storage space. The GPU on the data reader parses these tags to identify the location and characteristics of the required data and transfers the data from the shared storage space to local memory.

[0106] In one example, the data writing process between the root port and the end node is as follows: Figure 4 shown.

[0107] In this example, taking the edge computing device as a system-on-chip as an example, the system-on-chip may include a central processing unit and a graphics processing unit.

[0108] In traditional solutions, data is transmitted and transferred from the central processing unit to the graphics processing unit using topics as carriers.

[0109] In this embodiment of the present application, topics no longer encapsulate raw data, but instead encapsulate data tags, and still use message middleware to transmit topics. Based on the data tags in the topics, the GPU can directly read data from the shared storage space through GPU direct communication technology combined with PCIe.

[0110] like Figure 4 As shown, the GPU performs complex model inference tasks. Model N-1, Model N, Model N+1, and so on can represent different deep learning models, used to perform different tasks. The GPU can communicate with the model through data labels, perform inference, and output results.

[0111] In this embodiment, by combining the graphics processor direct communication technology and PCIe during the data transmission process, data can be read quickly between graphics processors, which reduces the waiting time for data transmission and improves data processing efficiency.

[0112] In some embodiments, the shared memory space includes a CPU memory space and a GPU memory space;

[0113] CPU storage space, used to store data related to CPU tasks;

[0114] GPU memory space, used to store data related to GPU tasks.

[0115] In this embodiment, the shared memory space may be a multifunctional memory area including a CPU memory space and a GPU memory space.

[0116] The CPU storage space is used to store data related to CPU tasks. CPU tasks are processed by the edge computing device's CPU. These tasks include highly logical and sequential tasks such as system control, sensor data preprocessing, and path planning. Data related to CPU tasks can include control instructions, sensor sampling frequency configuration, and other data.

[0117] In this embodiment, when a root port or end node is used as a data writer, the corresponding storage space can be automatically selected based on the data type. For data related to CPU tasks, the data writer can write the data to the CPU storage space via a PCIe switch. For data related to GPU tasks, the GPU can directly write the data to the GPU storage space using GPU direct communication technology.

[0118] In this embodiment, by dividing the shared storage space into a CPU storage space and a GPU storage space, different types of data can be stored separately, reducing conflicts between the CPU and GPU during data transmission and improving resource utilization efficiency.

[0119] In some embodiments, the in-vehicle communication system based on a distributed architecture may further include: a message scheduling layer, a data transmission layer, and a model reasoning layer; the message scheduling layer and the data transmission layer transmit data related to central processing unit tasks, and the model reasoning layer transmits data related to a graphics processing unit;

[0120] Among them, the message scheduling layer transmits messages by combining Ethernet and message middleware; the data processing layer transmits data by combining PCIe and direct memory access; and the model inference layer transmits data by combining PCIe and graphics processor direct communication technology.

[0121] In this embodiment, Figure 5 As shown, the in-vehicle communication system based on a distributed architecture can also include: a message scheduling layer, a data transmission layer, and a model reasoning layer. Each layer has specific functions and transmission mechanisms to adapt to different types of processor tasks and data transmission requirements.

[0122] The message scheduling layer primarily handles data related to CPU tasks. It utilizes a combination of Ethernet and message-based middleware for message transmission. The message scheduling layer manages and schedules message flows within the system. Leveraging the broad compatibility of Ethernet and the efficient transmission mechanisms of message-based middleware, it ensures reliable transmission of CPU task data. Within the message scheduling layer, data is typically encapsulated and transmitted in the form of messages, using topics as carriers. Messages contain data tags.

[0123] The data transfer layer combines PCIe and direct memory access technologies for data transmission. This layer focuses on fast and efficient data transfer, particularly for CPU-intensive tasks that require large amounts of data movement, such as raw state data collected by sensors. PCIe provides a high-bandwidth, low-latency hardware channel, while direct memory access technology allows data to be transferred directly between memory and devices without CPU intervention, reducing CPU burden and improving data transfer efficiency.

[0124] The model inference layer specifically handles data related to GPU tasks. This layer utilizes a combination of PCIe and direct communication with the GPU for data transmission. Its primary responsibility is to support efficient model inference tasks, such as executing deep learning models, on the GPU. Through direct communication with the GPU, data can be transferred directly between the GPU and shared storage, bypassing the CPU or host memory, achieving high-speed data transmission and processing.

[0125] In this embodiment, this layered, distributed architecture enables optimized support for different types of processor tasks. Specifically, the message scheduling layer improves message scheduling efficiency by combining Ethernet and message-based middleware. The data transmission layer achieves fast and efficient data transmission, reducing the burden on the central processing unit (CPU), by combining PCIe and direct memory access technology. The model inference layer achieves high-speed transmission and processing of GPU task data by combining PCIe and GPU direct communication technology, providing support for complex computing tasks in autonomous driving systems.

[0126] An embodiment of the present application also provides a vehicle, comprising the above-mentioned in-vehicle communication system based on a distributed architecture.

[0127] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] The above is a detailed introduction to the in-vehicle communication system and vehicle based on a distributed architecture provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A vehicle communication system based on a distributed architecture, characterized in that: include: A plurality of edge computing devices; wherein at least one edge computing device is configured as a root port and at least one edge computing device is configured as an end node; A PCIe switch is connected to the root port and the end node PCIe respectively; Wherein, the root port includes a shared storage space, and the root port and the end node perform data transmission in the shared storage space through the PCIe switch; including: the publishing node writes the target data into the shared storage space, and if the writing is successful, encapsulates the data tag of the target data in the target topic, and transmits the target topic to the subscription node through the message middleware; the subscription node receives the target topic transmitted by the message middleware, and reads the target data from the shared storage space according to the data tag in the target topic; wherein, the publishing node is the data writer in the root port and the end node, and the subscription node is the data reader in the root port and the end node; the message middleware transmits messages between the publishing node and the subscription node through the Ethernet switch; The shared storage space includes a CPU storage space and a GPU storage space; the CPU storage space is used to store data related to CPU tasks; the GPU storage space is used to store data related to GPU tasks; when the root port or the end node acts as a data writer, the corresponding storage space is selected according to the data type; The system includes a message scheduling layer, a data transmission layer and a model reasoning layer; The message scheduling layer is used to combine Ethernet and the message middleware to transmit messages; The data transmission layer is used to combine PCIe and direct memory access technology to perform data transmission; The model inference layer is used to process data related to graphics processor tasks and combine PCIe and graphics processor direct communication technology for data transmission.

2. The system according to claim 1, wherein: Before data is transmitted in the shared storage space through the PCIe switch, the method includes: The end node maps the physical address of the end node to the shared memory space of the root port through the base address register, so that the end node can determine the physical address of the end node through the shared memory space.

3. The system according to claim 1, wherein: The transmitting data in the shared storage space through the PCIe switch includes: The data writers in the root port and the end node perform binary serialization on the data to obtain serialized data; compress the serialized data, and write the compressed serialized data into the shared storage space.

4. The system according to claim 3, characterized in that The transmitting data in the shared storage space through the PCIe switch includes: The data readers in the root port and the end node read the data stored in the shared storage space, and restore the read data through deserialization.

5. The system according to claim 1, wherein: The system further comprises: an Ethernet switch, connected to the root port and the end node Ethernet respectively; Message middleware is used to transmit messages between the root port and the end node, and between different end nodes through the Ethernet switch.

6. The system according to claim 5, characterized in that The message transmission through the Ethernet switch includes: The topic is used as a carrier for message transmission, and the message transmission is performed by encapsulating a data tag in the topic; the data tag represents the attribute of the data.

7. The system according to claim 6, characterized in that The data tag includes at least one of a data encryption type, a memory storage first address, and a memory occupied length.

8. The system according to claim 1, wherein: The transmitting data in the shared storage space through the PCIe switch includes: In the case of data transmission related to a graphics processor task, the root port and the end node use a graphics processor direct communication technology to perform data transmission in the shared storage space through the PCIe switch.

9. The system according to claim 8, characterized in that The adopting of the graphics processor direct communication technology to transmit data in the shared storage space through the PCIe switch includes: The data writers in the root port and the end node write data into the shared storage space through a graphics processor.

10. The system according to claim 9, characterized in that The adopting of the graphics processor direct communication technology to transmit data in the shared storage space through the PCIe switch includes: The data readers in the root port and the end node read data from the shared storage space through a graphics processor.

11. The system according to claim 9, wherein: The graphics processor tasks include model inference tasks.

12. A vehicle, characterized in that: It includes the vehicle communication system based on the distributed architecture as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Information processing method and system

    CN117978828A