Unmanned driving software development method applicable to multi-chip and multi-domain controllers

By adopting atomic operator-based software development method, the complexity of autonomous driving software development is reduced, and the problems of large workload, long cycle and high cost under traditional development methods are solved, and efficient and convenient software development and reuse are achieved.

CN114443134BActive Publication Date: 2025-06-27HANGZHOU HONG JING DRIVE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111655839.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-06-27
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Traditional autonomous driving software development methods lead to large workload, long cycles, high costs, and large differences in hardware platforms of different models, resulting in difficulties in software reuse and transplantation.

Method used

Using a software algorithm architecture based on reusable atomic operators, the algorithms required by external devices are converted into atomic operator data streams, combined into a complete software process, and the processes are divided on multiple chips for adaptation, a code framework is generated, and unit testing and system-level verification are performed.

Benefits of technology

It greatly reduces the cost and cycle of software development, improves development efficiency and code reuse rate, reduces testing workload, and facilitates software transplantation between different models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114443134B_ABST
    Figure CN114443134B_ABST
Patent Text Reader

Abstract

The present invention discloses a driverless software development method applicable to a multi-chip multi-domain controller, comprising the following steps: (1) determining a software algorithm architecture based on reusable atomic operators; (2) respectively converting the algorithms required for each external device into atomic operator data streams; (3) combining the atomic operator data streams of each external device into a complete software process, dividing multiple atomic operators in the software process into multiple processes, and each of the processes includes at least one atomic operator; (4) adapting the computing power required by all processes to the computing power of multiple chips; (5) generating a code framework for each of the processes; (6) performing software development and unit testing on each atomic operator; (7) verifying each process; (8) performing system-level verification. In the present invention, the atomic operators can be reused, which is convenient for transplantation between different vehicle models and improves the development efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for developing driverless software applicable to a multi-chip multi-domain controller, belonging to the technical field of autonomous driving. Background Art

[0002] The complexity of high-level driverless software is very high. Generally, SOA (Service-Oriented Architecture) is adopted for software development on SoC (System on Chip), and it can be based on operating systems such as Linux and QNX. In SOA, each functional module will be combined into some independently running processes, and communication can be carried out between processes.

[0003] There are often multiple chips inside the domain controller for high-level driverless, including MCU, AI (Artificial Intelligence) chips, and chips with high logical computing power, etc. Each process needs to be allocated to different chips for operation.

[0004] Such as Figure 1 shown, it is a typical hardware architecture of a multi-chip multi-domain controller, which has two domain controllers, a primary and a backup. Among them, the primary domain controller has 3 SoCs and 1 MCU (Microcontroller Unit), and 2 of the SoCs include AI units with deep learning inference functions. The backup domain controller has 1 SoC and 1 MCU, mainly used for functional safety redundancy, and performs operations such as safe parking when the primary domain controller fails.

[0005] For automobile OEMs, different vehicle models have different functional requirements. Since it is necessary to minimize hardware costs, a unified hardware platform will not be used. Instead, some chips with just enough computing power are selected according to specific functional requirements to form a new domain controller. In this way, different vehicle models use different hardware, different chip computing powers, and different communication mechanisms between chips. However, at the software level, for software such as perception, positioning, fusion, and planning, it is hoped that it can be reused as much as possible to save software development costs and shorten the development cycle.

[0006] Such as Figure 2 shown, it is a typical traditional software development process based on SoC. Figure 2 The modules in correspond to processes, such as the map positioning module, multi-sensor fusion module, etc. The traditional software development process mainly has the following problems:

[0007] 1. There are significant differences in the number and performance of SoCs in the domain controllers used in the hardware platforms of different vehicle models. Some large modules may not be able to be carried by any SoC, so the modules need to be split. In this case, the code needs to be massively refactored.

[0008] 2. Among different vehicle models, there are differences in the functional requirements of modules. If one tries to reuse the software of corresponding modules of other vehicle models, a large number of modifications are required, and then a new software branch needs to be maintained. As the number of vehicle models increases, more and more engineers will be needed to maintain each software branch.

[0009] 3. The developers of each module need to be familiar with the overall software architecture and the associations with other modules, and select the communication methods between modules, which requires relatively high capabilities of the developers.

[0010] In short, if developed according to the traditional software development method, the software for each vehicle model needs to be developed from scratch, with high costs and long cycles. Summary of the Invention

[0011] The object of the present invention is to solve the technical problems of large workload, long cycle and high cost in the traditional development method of driverless software pointed out in the background technology.

[0012] To achieve the above object of the invention, the present invention provides a driverless software development method applicable to multi-chip multi-domain controllers, including the following steps: (1) determining a software algorithm architecture based on reusable atomic operators; (2) respectively converting the algorithms required to be executed by each external device into atomic operator data streams; (3) combining the atomic operator data streams of each external device into a complete software process, and dividing the multiple atomic operators in the software process into multiple processes, each of the processes includes at least one atomic operator; (4) adapting the computing power required by all processes to the computing power of multiple chips; (5) generating a code framework for each of the processes; (6) performing software development and unit testing on each atomic operator; (7) verifying each process; (8) performing system-level verification. The code framework is equivalent to a complete code directory structure and code files, and the functions of the atomic operators in the code files are empty inside, and the subsequent software development needs to complete these atomic operator codes. System-level verification is to execute all processes simultaneously and verify whether the verification results meet the expectations.

[0013] Further, in step (1), the atomic operator is a basic operation unit suitable for implementing at least one algorithm step.

[0014] Further, in step (2), the external devices at least include lidar, millimeter-wave radar, camera, positioning terminal and inertial sensor.

[0015] Further, the atomic operators at least include lidar data access, point cloud pre - processing, point cloud clustering, obstacle recognition, obstacle fusion, obstacle tracking, lidar perception data publishing, lidar deep learning inference, post - processing of deep learning, millimeter - wave radar data access, millimeter - wave radar perception data publishing, camera data access, image pre - processing, image distortion removal, camera deep learning inference, lane line recognition, camera perception data publishing, time synchronization, multi - sensor matching, fused data publishing, positioning data access, inertial sensor data access, rough positioning, Kalman filtering, positioning publishing, high - definition map reading, high - definition map publishing, map and visual positioning fusion, map and visual positioning fusion, and map and visual positioning fusion result publishing.

[0016] Further, in step (2), the computing power of the chip includes AI computing power and logic computing power.

[0017] Further, in step (3), processes are automatically divided based on atomic aggregation. The atomic aggregation is to combine multiple atomic operators with high coupling degree into the same process. The coupling degree is measured from three aspects: 1. The communication bandwidth between two atomic operators; 2. The complexity of the API between two atomic operators (i.e., the number of variables in the API); 3. Manually set logical associations. Among them, the third aspect will be automatically recorded after manual adjustment at the end. Atomic operators assigned to one process will increase the coupling degree, and the coupling degree between atomic operators split into two processes before and after will decrease; the coupling degree between atomic operators = normalization constant * communication bandwidth * API complexity+manual coupling degree adjustment; when judging the high or low coupling degree, a default coupling degree threshold will be used, and this threshold will be adjusted according to the actual coupling degree threshold of the confirmed architecture each time. For example, after manual process splitting, it means that the previous coupling degree threshold was too low and will be automatically increased.

[0018] Further, during the process of dividing processes in step (3), when the computing power required by a process exceeds the computing power range of the chip, the process is split into at least two processes, and the split multiple processes are adapted to different chips until the computing power required by the process is fully adapted.

[0019] Further, during the process of dividing processes in step (3), when there are multiple parts with low logical coupling degree in a process, each part is split into an independent process.

[0020] Further, in step (6), the atomic operators that have completed unit testing are automatically imported into the atomic operator library for reuse.

[0021] Further, in the data stream of the atomic operators in step (2), program codes for communication between atomic operators are automatically generated, and combined program codes for multiple aggregated atomic operators are automatically generated.

[0022] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0023] 1. Greatly reduce the software development cost of porting platform software to different hardware platforms. The code volume for porting existing functions can be reduced by more than 80%, and unit tests for existing atomic operators can be omitted, and the process testing workload can also be reduced by more than 50%. For software that traditionally takes 6 months to develop, it is expected to only take 2 - 3 months using this method, which can shorten the development cycle by more than half.

[0024] 2. Since atomic operators can be reused, during the software testing process of different vehicle models, it is equivalent to having completely tested and verified all the atomic operators used therein, thus ensuring the stability of these algorithms, reducing a large amount of additional testing workload, and improving development efficiency.

[0025] 3. Facilitate transplantation between different vehicle models, saving R & D costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a schematic block diagram of a multi - chip multi - domain controller in the prior art;

[0027] Figure 2 is a software development flow chart in the prior art for driverless;

[0028] Figure 3 is a flow chart of an embodiment of the present invention;

[0029] Figure 4 is a schematic diagram of atomic aggregation in an embodiment of the method of the present invention;

[0030] Figure 5 is a software algorithm architecture diagram of an embodiment of the present invention;

[0031] Figure 6 is a hardware architecture diagram of an embodiment of the present invention;

[0032] Figure 7 is a hardware architecture setting interface of an embodiment of the present invention;

[0033] Figure 8 is a software architecture setting interface of an embodiment of the present invention;

[0034] Figure 9 is an atomic aggregation setting interface of an embodiment of the present invention;

[0035] Figure 10 is a preliminary process allocation diagram in an embodiment of the present invention;

[0036] Figure 11It is the process allocation diagram after automatic adjustment in an embodiment of the present invention;

[0037] Figure 12 It is the process allocation diagram after manual fine-tuning in an embodiment of the present invention;

[0038] Figure 13 It is the interface for importing atomic operators in an embodiment of the present invention. Detailed implementation manners

[0039] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0040] As Figure 3 shown, an embodiment of the software development method for driverless vehicles applicable to multi-chip multi-domain controllers of the present invention includes the following steps: (1) determining a software algorithm architecture based on reusable atomic operators; (2) respectively converting the algorithms required for each external device into atomic operator data streams; (3) combining the atomic operator data streams of each external device into a complete software process, dividing multiple atomic operators in the software process into multiple processes, and each of the processes includes at least one atomic operator; (4) adapting the computing power required by all processes to the computing power of multiple chips; (5) generating code frameworks for each of the processes; (6) performing software development and unit testing on each atomic operator; (7) verifying each process; (8) performing system-level verification.

[0041] Figure 3 The process shown adopts an algorithm mode based on atomic operators compared with the traditional V-shaped development mode, and the core lies in the automatic aggregation mechanism of atoms. It should be noted that for algorithms based on AI chips, such as deep learning, there is generally also a concept of operators, which is similar to the concept of mathematical operators, such as convolution. The operators of AI are only used to build networks and generally cannot be reused between multiple platforms, and software still needs to be implemented by writing program codes. However, the atomic operators in the embodiments of the present invention are a kind of reusable basic algorithm operation units, not limited to mathematical operators, and it represents an intermediate step of a complete algorithm. For example, each step including picture format conversion, scaling, cropping, distortion removal, etc. in an image processing algorithm can belong to an atomic operator for image preprocessing. A complete software is established based on atomic operators, and the main purpose is to facilitate code reuse between different vehicle models and multiple platforms. The definition of atomic operators can be modified according to actual situations. For example, if the picture format conversion consumes a high amount of computing power or becomes a complex logic, it can be made into an independent atomic operator.

[0042] Atomize the algorithms involving all external devices. For example, visual perception can be atomized into multiple atomic operators such as sensor access, visual pre-processing, deep learning obstacle recognition, deep learning lane line recognition, obstacle post-processing, lane line post-processing, and result publishing.

[0043] During the software architecture design phase, it is necessary to refine the data flow level of the atomic operator and consider reusing operators in the existing atomic operator library to reduce the development workload. Compared with the traditional process, there is an additional system adaptation phase, which mainly inputs the hardware architecture and task arrangement, and then the software automatically completes the atomic aggregation process and the generation of the code framework. After that, development and unit testing are all oriented to atomic operators, and the verified atoms will enter the atomic operator library.

[0044] like Figure 4 As shown in the figure, in the automatic atomic aggregation method, first, according to the atomic architecture diagram, the dependency relationship between different atoms is calculated. If there is data transmission between two atoms, an edge is drawn in the middle. The larger the amount of data transmission between two atoms, the greater the coupling, the greater the weight of the edge, and the shorter the distance between them. Conversely, the distance is longer. The dependency relationship here refers to whether there is data communication between two atomic operators.

[0045] In this way, we formed a two-dimensional graph of the distribution of atoms. Then, according to the hardware architecture diagram, we learned that these algorithms need to be deployed on several chips and the computing power of each chip. Then the processes (services) are divided according to two principles: first, according to the coupling degree of atoms, the ones with high coupling degree are allocated together as much as possible; second, if too many atoms are coupled, the overall computing power of the process is too high and exceeds the chip carrying range, and the process will be split. It is worth noting that if we regard each atom as a process, although it simplifies the allocation mechanism, it will lead to the problem that the communication cost is too high when the highly coupled atoms are not in a process, so it is necessary to make a trade-off between combination and division. For example, the interface between the two atoms of camera access and visual pre-processing is the camera raw data, and the data volume and bandwidth requirements are very large, so the coupling degree is very high, and memory can be effectively shared within a process to improve efficiency. Finally, based on the result of this automatic allocation, appropriate manual adjustments can be made, and the communication code between atoms can be automatically generated. In this way, the codes of all atoms are unified, and the combination code of atoms can be automatically generated according to different projects, which ensures the convenience of project transplantation and saves R&D costs. The following takes the development of a specific unmanned driving software as an example to illustrate the process of the method of the present invention.

[0046] Figure 5 and Figure 6They are the software and hardware architectures respectively, which serve as the input sources. In the hardware architecture, the AI SoC has AI computing power in TOPS, and all SoCs have logical computing power in DMIPS. Squares represent chips, and ellipses represent external devices. In the software architecture, ellipses represent the decomposed atomic operators, and squares represent the external associated devices of the SoC. Each operator needs to estimate the AI computing power and logical computing power. The edge between two operators will include an execution cycle and a bandwidth, and the cycle is event, which means it is triggered by an upstream event.

[0047] Step 1: Input the hardware architecture and refer to the interface in Figure 7 . You can drag elements from the left to the main view and input the hardware architecture. Confirm the computing power of each chip, as well as the type and bandwidth of the channels between chips. MCU and Device are both associated devices for the SoC, and their hardware paths are also noted in this diagram.

[0048] Step 2: Input the software architecture and refer to the interface in Figure 8 . Drag components from the newly added elements to the main view, or directly import the atomic operators in the existing library through the atomic operator list. Fill in the computing power in the node properties on the right side. For those with AI computing, the TOPS computing power needs to be filled in additionally. Then select the trigger mode. For tasks triggered by a fixed cycle, the trigger frequency needs to be written. Event trigger means that the calculation is triggered according to the previous input data, and the cycle does not need to be filled in. Finally, confirm the channels between atoms or between atoms and external devices. Synchronous sending means that the data is directly sent after the calculation is completed, while asynchronous sending requires setting the sending cycle and using an independent thread for sending. The header file of the interface needs to be added to the interface structure, and then the interface size (in bytes) needs to be filled in to calculate the bandwidth requirement.

[0049] Step 3: Execute the automatic process allocation and refer to the interface in Figure 9 . After clicking the automatic aggregation button, the automatically aggregated processes will be generated in the main view, and the computing power requirements after statistics for each process can be seen. Based on the software architecture in Figure 5 and the hardware architecture in Figure 6 , the automatic aggregation process is described as follows. First, make a preliminary process allocation:

[0050] a) Set the communication bandwidth between any two atoms in the diagram as the weight of the edge between the two atoms;

[0051] b) For the atoms at both ends of the edge with a large weight, try to ensure that they are in the same process, at least in the same chip, to improve the data transmission efficiency;

[0052] c) For the edges triggered by events (the event part in the diagram), considering the delay of message passing, try to ensure that the atomic operators at both ends are in the same process and can be directly serially called to avoid explicit message passing.

[0053] According to these principles, the Figure 10 allocation result is obtained. Figure 10 In, multiple atomic operators with the same connection line between adjacent atomic operators (referring to the same symbols at the ends of the connection lines, such as hollow arrows, solid arrows, hollow circles, solid circles, etc.) form a process. Note that the same atomic operator can only belong to one process, that is, two processes cannot share the same atomic operator. The processes of lidar, millimeter-wave radar, and camera all terminate at the data publishing atomic operator, and the time synchronization atomic operator below it belongs to another process. Figure 11 , Figure 12 is the same.

[0054] Then, the allocated processes are adapted to different chips, and there are several principles:

[0055] a) Processes that require TOPS computing power need to be placed on the AI SoC;

[0056] b) The external device link must be corresponding to the hardware architecture;

[0057] c) The computing power of each chip needs to meet the limit.

[0058] However, calculated according to this condition, a total of 60KDMIPS of logical computing power is required on the AI SoC, which exceeds the 28KDMIPS of the AI SoC. Therefore, at the stage of deploying to the chip, the process needs to be split according to the limit conditions, so as to automatically adjust to obtain the Figure 11 new process allocation result shown. Figure 11 In, multiple atomic operators with the same connection line between adjacent atomic operators form a process.

[0059] According to this adjustment, the processes related to lidar deep learning and the processes related to the camera are both allocated on the AI SoC for execution (sharing 21.5KDMIPS + 25TOPS), while other processes are allocated on non-AI SoC for execution (sharing 74.5KDMIPS), meeting the requirements.

[0060] Fourth step, perform manual fine-tuning considering the logical characteristics between atomic operators. Figure 9 In the interface of, if it is considered that the process arrangement is illogical (automatic aggregation does not consider the logical meaning between atomic operators), the atomic operator can be adjusted to different processes by manual dragging. At the same time, the software will also learn the logical coupling degree between atomic operators. When performing automatic aggregation again, it will tend to give priority to placing the atomic operators that have been manually set in one process in one process.

[0061] For example Figure 12In this case, since the logical coupling degree or correlation between the two atomic operators of map publishing and map visual positioning fusion is low, they can be split into two processes. Figure 12 In this case, multiple atomic operators that use the same connection line between adjacent atomic operators form a process.

[0062] Step 5: Generate the framework code, including the existing atomic operators. Here, a C++ class will be generated for each process, responsible for the management of process scheduling and member variables, and then a C++ class will be generated for each atomic operator, including some basic frameworks of the class (constructor and destructor functions, function declarations). Then, different communication codes will be added between different atomic operators according to the allocated processes, chips, etc.:

[0063] a) For the operators within the same process, directly call two functions serially in the process scheduling function, and place the intermediate interface parameters in the member variables of the process class;

[0064] b) For the operators of different processes on the same chip, use shared memory for communication, and this part can be implemented by the middleware of the unmanned driving system (such as AP AUTOSAR, etc.);

[0065] c) For the operators on different chips, if the bandwidth is relatively small, Socket can be used for communication, and if the bandwidth is relatively large, it may be necessary to adapt to proprietary interfaces such as PCIE.

[0066] Step 6: Continue development and verification based on the code framework, and this part of the work can be carried out based on common C++ development tools.

[0067] Step 7: For the verified atomic operators, they can be imported into the atomic operator library for reuse. As Figure 13 shown, only the source code and dependent libraries need to be added during import. The imported atomic operators will appear in the Figure 8 atomic operator list in the lower left corner.

[0068] Compared with the traditional software development method, this method has the following advantages:

[0069] 1. Since the atomic operator code is small enough, it will not be split according to the project needs, and the input and output of the atom are fixed. Therefore, it can basically be reused between different projects without maintaining redundant branches;

[0070] 2. Since the atomic operators have strong reusability, and the reused atomic operators themselves have been verified in a large number of previous projects, the software quality can be guaranteed, and a large amount of unit testing and module testing work can also be reduced;

[0071] 3. Developers of atomic operators only need to know the input, output, and functions of atomic operators, and do not have to have a complete understanding of the overall software architecture. The communication code between atomic operators is also automatically generated. Therefore, there is no need for particularly experienced engineers to engage in development work.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A software development method for driverless applications applicable to a multi-chip multi-domain controller, characterized in that, It includes the following steps: (1) Determine the software algorithm architecture based on reusable atomic operators; (2) Respectively convert the algorithms required to be executed by each external device into atomic operator data streams, where the external devices at least include lidar, millimeter-wave radar, camera, positioning terminal, and inertial sensor; (3) Combine the atomic operator data streams of each external device into a complete software process, and divide multiple atomic operators in the software process into multiple processes, where each of the processes at least includes one atomic operator; Automatically divide processes based on atomic aggregation. The atomic aggregation is to combine multiple atomic operators with high coupling degree into the same process. During the process of dividing processes, when the computing power required by a process exceeds the computing power range of the chip, then split the process into at least two processes, and adapt the split multiple processes to different chips until the computing power required by the process is fully adapted. When there are multiple parts with low logical coupling degree in a process, split each part into an independent process, where: coupling degree between atomic operators = normalization constant * communication bandwidth between two atomic operators * number of variables in the APIs between two atomic operators + manual coupling degree adjustment; (4) Adapt the computing power required by all processes to the computing power of multiple chips, where the computing power of the chip includes AI computing power and logical computing power; (5) Generate the code framework for each of the processes; (6) Perform software development and unit testing on each atomic operator; (7) Verify each process; (8) Conduct system-level verification; The atomic operators at least include lidar data access, point cloud preprocessing, point cloud clustering, obstacle recognition, obstacle fusion, obstacle tracking, lidar perception data publishing, lidar deep learning inference, post deep learning processing, millimeter-wave radar data access, millimeter-wave radar perception data publishing, camera data access, image preprocessing, image distortion correction, camera deep learning inference, lane line recognition, camera perception data publishing, time synchronization, multi-sensor matching, fused data publishing, positioning data access, inertial sensor data access, rough positioning, Kalman filtering, positioning publishing, high-precision map reading, high-precision map publishing, map and visual positioning fusion, map and visual positioning fusion result publishing.

2. The driverless software development method applicable to a multi-chip multi-domain controller according to claim 1, characterized in that The atomic operator in step (1) is a basic operation unit suitable for implementing at least one algorithm step.

3. The driverless software development method applicable to a multi-chip multi-domain controller according to claim 1, characterized in that, The atomic operators that complete unit testing in step (6) are automatically imported into the atomic operator library for reuse.

4. The driverless software development method applicable to a multi-chip multi-domain controller according to claim 1, characterized in that In the atomic operator data stream in step (2), program code for communication between atomic operators is automatically generated, and combined program code for multiple aggregated atomic operators is automatically generated.

Citation Information

Patent Citations

  • Method of automatic generation of internet-of-things data process flow based on RFID (radio frequency identification)

    CN103197961A

  • Code generation method and device

    CN110297632A