Unmanned driving system dynamic control method and device, electronic equipment and storage medium

By employing a heterogeneous computing platform with dual system-on-a-chip in the autonomous driving system, dynamic deployment and load balancing of tasks between the two computing units are achieved, solving the problems of low computing power utilization and high response latency caused by static task deployment, and improving the system's computing power utilization, real-time response capability and robustness.

CN120994340APending Publication Date: 2025-11-21SANY INTELLIGENT MINING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511115034.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The current autonomous driving system uses a static task deployment method, which lacks dynamic adjustment based on real-time task load and equipment failure status. This results in low computing power utilization, high system response latency, and difficulty in meeting the requirements of high reliability and high real-time performance for autonomous driving.

Method used

The heterogeneous computing platform, which adopts dual system-on-a-chip, achieves dynamic deployment and load balancing of tasks to be processed between two computing units through device type identification, task priority evaluation, computing load monitoring and system health status feedback mechanisms. It also utilizes a fault diagnosis and monitoring module to perform task migration and fault switching in abnormal situations.

Benefits of technology

It improves computing power utilization, real-time response capability and system robustness, ensures high reliability and security of the system in case of failure, and supports dynamic task allocation and resource optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994340A_ABST
    Figure CN120994340A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned driving system dynamic control method and device, electronic equipment and a storage medium, and relates to the technical field of unmanned driving, and the method comprises the steps: calling an equipment type recognition interface, obtaining an equipment type, and loading a configuration file; establishing a task priority table, a task computing power demand table and an equipment load table; acquiring priorities and computing power requirements of the plurality of to-be-processed tasks, and distributing the plurality of to-be-processed tasks to two computing units for processing; and under the condition that the operation data of one calculation unit is abnormal, at least one part of the to-be-processed tasks needing to be processed by the abnormal calculation unit is transferred to the other calculation unit. According to the technical scheme, through mechanisms such as equipment type identification, task priority evaluation and operation data monitoring, dynamic deployment and load balancing of the to-be-processed task between the two computing units are realized, and the computing power utilization rate, the real-time response capability and the system robustness of the automatic driving system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned driving, in particular to an unmanned system dynamic control method and device, electronic equipment and storage medium. BACKGROUND

[0002] In related technologies, the task deployment mode of the computing platform based on the unmanned system is mostly static division, and lacks the ability to dynamically adjust according to real-time task load and device fault condition, resulting in low computing power utilization and high system response delay, which is difficult to meet the high reliability and high real-time automatic driving requirements. SUMMARY

[0003] In order to solve or improve the technical problems of low computing power utilization and high system response delay of the traditional task deployment mode, one purpose of the present application is to provide an unmanned system dynamic control method.

[0004] Another purpose of the present application is to provide an unmanned system dynamic control device.

[0005] Another purpose of the present application is to provide an electronic equipment.

[0006] Another purpose of the present application is to provide a readable storage medium.

[0007] To achieve the above purpose, the first aspect of the present application provides an unmanned system dynamic control method applied to a computing platform of a dual system level chip, the computing platform comprising two computing units based on system level chips, a configuration management module, a task scheduling engine module and a fault diagnosis monitoring module.

[0008] The unmanned system dynamic control method comprises: calling a device type identification interface through the configuration management module, obtaining device types corresponding to the two computing units, and loading configuration files corresponding to the device types; based on the configuration files, initializing a task scheduling strategy through the task scheduling engine module, establishing a task priority table, a task computing power demand table and a device load table; according to the task priority table, the task computing power demand table and the device load table, obtaining priorities and computing power demands of a plurality of to-be-processed tasks, and distributing the plurality of to-be-processed tasks to the two computing units for processing; collecting running data of the two computing units through the fault diagnosis monitoring module; in the case that the running data of one of the computing units is abnormal, at least part of the plurality of to-be-processed tasks that need to be processed by the abnormal computing unit is transferred to the other computing unit.

[0009] The application aims to provide a dynamic control method for an unmanned system, based on the priority of tasks and the computing power requirement, a plurality of to-be-processed tasks are processed by two computing units respectively. The operation data is monitored by the fault diagnosis monitoring module, and in the case that one of the computing units is abnormal, at least part of the plurality of to-be-processed tasks to be processed is transferred to the other computing unit. Through the mechanisms of device type identification, task priority evaluation and operation data monitoring, the dynamic deployment and load balancing of the to-be-processed tasks between the two computing units are realized. This design is beneficial to improve the computing power utilization rate, real-time response capability and system robustness of the automatic driving system.

[0010] In some technical solutions, according to the task priority table, the task computing power requirement table and the device load table, the priority and computing power requirement of the plurality of to-be-processed tasks are obtained, and the plurality of to-be-processed tasks are distributed to the two computing units for processing, comprising: according to the task priority table, the task computing power requirement table and the device load table, the priority and computing power requirement of the plurality of to-be-processed tasks are determined; according to the priority and computing power requirement of the to-be-processed tasks, the plurality of to-be-processed tasks are divided into emergency tasks, high-priority tasks and low-priority tasks; based on the task category, the computing power requirement and the current load state of the device load table, the plurality of to-be-processed tasks are distributed to the two computing units for processing.

[0011] In this technical solution, the computing platform intelligently schedules and switches the to-be-processed tasks between the two computing units, and realizes the efficient operation and resource optimization of the unmanned system through dynamic task distribution and computing power sharing.

[0012] In some technical solutions, optionally, the emergency task includes one or a combination of the following: perception fusion task and emergency braking task; the high-priority task includes path planning task; and the low-priority task includes log recording task.

[0013] In this technical solution, by finely dividing the to-be-processed tasks, the computing power requirement of each type of task is determined, so that the system can dynamically adjust the computing power allocation according to the environment, which is beneficial to optimize resource allocation and avoid overload or resource idling of a computing unit.

[0014] In some technical solutions, optionally, the configuration file includes device computing power capability information, task mapping relationship information, heartbeat period information and fault code definition information.

[0015] In this technical solution, the device computing power capability information, the task mapping relationship information, the heartbeat period information and the fault code definition information jointly constitute the "operation rule library" of the system, ensuring that the two computing units can dynamically allocate tasks, monitor the health status and handle faults according to the designed logic, which is the core support for realizing computing power optimization and high reliability.

[0016] In some embodiments, the operation data includes load state data, heartbeat state data, process running state data, and fault code data.

[0017] In this embodiment, the refined collection and analysis of operation data is the core support for the dynamic deployment, fault self-recovery, and safe operation of the dual-system level chip. Through data-driven decision optimization, the reliability, efficiency, and maintainability of the system are significantly improved.

[0018] In some embodiments, the dynamic control method of the unmanned system further includes: collecting operation data of the two computing units through the fault diagnosis monitoring module; in the case that the operation data of one of the computing units is abnormal, at least part of a plurality of to-be-processed tasks that need to be processed by the abnormal computing unit is transferred to the other computing unit, then fault information is reported to the gateway, and the abnormal computing unit is diagnosed and recovered according to the fault information.

[0019] In this embodiment, after a fault occurs, part of the tasks are transferred, the fault information is reported and the fault is diagnosed and recovered. Through the whole process design of "accurate monitoring-intelligent transfer-closed loop recovery", the reliability and safety of the unmanned system are significantly improved.

[0020] In some embodiments, the computing power requirement of the to-be-processed task includes CPU utilization and GPU utilization.

[0021] In this embodiment, by defining the computing power requirements of CPU utilization and GPU utilization in detail, a "quantifiable and adaptable" basis can be provided for task scheduling and resource allocation of the unmanned system, the resource allocation accuracy is improved, the waste or overload of computing power is avoided, the cache and video memory utilization is optimized, and the computing efficiency is improved.

[0022] The second aspect of the present application provides a dynamic control device of an unmanned system, comprising: a device type identification unit configured to call a device type identification interface through a configuration management module to obtain device types corresponding to two computing units; a configuration file loading unit configured to load configuration files corresponding to the device types; a task scheduling strategy initialization unit configured to initialize a task scheduling strategy through a task scheduling engine module based on the configuration files to establish a task priority table, a task computing power demand table, and a device load table; a task processing unit configured to obtain priorities and computing power demands of a plurality of to-be-processed tasks according to the task priority table, the task computing power demand table, and the device load table, and distribute the plurality of to-be-processed tasks to the two computing units for processing; a running data acquisition unit configured to acquire running data of the two computing units through a fault diagnosis monitoring module; and a task dynamic deployment unit configured to, in a case where the running data of one of the computing units is abnormal, transfer at least part of a plurality of to-be-processed tasks that need to be processed by the abnormal computing unit to the other computing unit.

[0023] The present application aims to provide a dynamic control device of an unmanned system, which processes a plurality of to-be-processed tasks through two computing units based on priorities and computing power demands of the tasks. The running data is monitored through a fault diagnosis monitoring module, and at least part of a plurality of to-be-processed tasks that need to be processed is transferred to the other computing unit in a case where one of the computing units is abnormal. Through mechanisms such as device type identification, task priority evaluation, and running data monitoring, dynamic deployment and load balancing of to-be-processed tasks between the two computing units are achieved. This design is conducive to improving the computing power utilization rate, real-time response capability, and system robustness of the autonomous driving system.

[0024] The third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores programs or instructions executable on the processor, and the processor implements the steps of the dynamic control method of the unmanned system in any of the above technical solutions when executing the programs or instructions. The electronic device has the beneficial effects of any of the above technical solutions, which are not repeated here.

[0025] The fourth aspect of the present application provides a readable storage medium, which stores programs or instructions, and the programs or instructions implement the steps of the dynamic control method of the unmanned system in any of the above technical solutions when executed by a processor. The readable storage medium has the beneficial effects of any of the above technical solutions, which are not repeated here.

[0026] Additional aspects and advantages of the technical solutions of the present application will become apparent from the following description section, or will be understood by those skilled in the art through practice of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1A structural block diagram of a computing platform according to one embodiment of the present application is shown;

[0028] Figure 2 A structural block diagram of a computing platform according to another embodiment of the present application is shown;

[0029] Figure 3 A structural block diagram of a computing platform according to another embodiment of the present application is shown;

[0030] Figure 4 A flow chart of a dynamic control method of an unmanned system according to one embodiment of the present application is shown;

[0031] Figure 5 A flow chart of a dynamic control method of an unmanned system according to another embodiment of the present application is shown;

[0032] Figure 6 A flow chart of a dynamic control method of an unmanned system according to another embodiment of the present application is shown;

[0033] Figure 7 A structural block diagram of a dynamic control device of an unmanned system according to one embodiment of the present application is shown;

[0034] Figure 8 A structural block diagram of an electronic device according to one embodiment of the present application is shown.

[0035] wherein, Figures 1 to 8 The correspondence between the reference signs and the component names in the accompanying drawings is as follows:

[0036] 100: computing platform; 110: computing unit; 111: system on chip; 120: configuration management module; 130: task scheduling engine module; 140: fault diagnosis monitoring module; 150: communication middleware module; 160: OTA deployment module; 300: dynamic control device of unmanned system; 310: device type identification unit; 320: configuration file loading unit; 330: task scheduling strategy initialization unit; 340: task processing unit; 350: running data acquisition unit; 360: task dynamic deployment unit; 400: electronic device; 410: memory; 420: processor. DETAILED DESCRIPTION

[0037] In order to enable a clearer understanding of the above-mentioned purposes, features and advantages of the embodiments of the present application, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0038] Many specific details are set forth in the following description in order to provide a thorough understanding of the application. However, embodiments of the application can be practiced without resorting to the details specifically set forth in the following description, as those skilled in the art will understand that the scope of the application is measured by the claims.

[0039] With the development of automatic driving technology, the computing power demand of vehicle-mounted computing platforms is increasing. As a relatively mainstream automatic driving chip platform at present, NVIDIA Orin (a product name of a system and chip) has high integration and strong computing power. However, in actual application, the vehicle-mounted computing platform usually only uses a single system-level chip. A single system-level chip is difficult to meet the computing power demand of complex tasks such as multi-sensor fusion, path planning, behavior prediction and the like for high-order automatic driving, and there is a single-point failure risk.

[0040] The present application aims to provide an unmanned system dynamic control method, device, electronic equipment and storage medium, the computing platform includes two computing units based on system-level chips, and a plurality of to-be-processed tasks are processed by the two computing units respectively. The running data is monitored by the fault diagnosis monitoring module, and at least part of the plurality of to-be-processed tasks that need to be processed is transferred to the other computing unit in the case that one of the computing units is abnormal. Through the mechanisms of device type identification, task priority evaluation and running data monitoring, dynamic deployment and load balancing of to-be-processed tasks between the two computing units are realized. This design is beneficial to improve the computing power utilization rate, real-time response capability and system robustness of the automatic driving system.

[0041] In the related art, although two system-level chips are used, the task deployment method is mostly static division, and lacks the ability to dynamically adjust according to real-time task load, system health state and environmental changes, resulting in low computing power utilization rate, high system response delay, and difficulty in meeting the high reliability and high real-time automatic driving demand.

[0042] In the technical solution of the present application, the computing platform is a heterogeneous computing platform with two system-level chips, and through the mechanisms of device type identification, task priority evaluation, computing power load monitoring and system health state feedback, dynamic deployment and load balancing of to-be-processed tasks between the two computing units are realized, which is beneficial to improve the computing power utilization rate, real-time response capability and system robustness of the automatic driving system.

[0043] It should be emphasized that the computing platform intelligently schedules and switches the to-be-processed tasks between the two computing units, and realizes efficient operation and resource optimization of the unmanned system through dynamic task allocation and computing power sharing.

[0044] The following description will be made with reference to the accompanying drawings. Figures 1 to 8The application discloses an unmanned system dynamic control method, device, electronic equipment and storage medium.

[0045] In an embodiment of the application, the unmanned system dynamic control method is applied to a computing platform 100 of a dual-system chip. Figure 1 and Figure 2 As shown in the figure, the computing platform 100 includes two computing units 110 based on system chips 111, a configuration management module 120, a task scheduling engine module 130 and a fault diagnosis monitoring module 140.

[0046] The computing platform 100 is a heterogeneous computing platform adopting a dual-Orin (a kind of system chip 111) architecture, and includes two computing units 110 based on system chips 111. Through mechanisms such as device type identification, task priority evaluation, computing power load monitoring and system health state feedback, dynamic deployment and load balancing of tasks to be processed between the two computing units 110 are realized, which is beneficial to improving the computing power utilization rate, real-time response capability and system robustness of the automatic driving system.

[0047] Optionally, the two computing units 110 are respectively denoted as Orin-A and Orin-B. Figure 3 In the figure, “Orin-A” represents the first computing unit 110, and “Orin-B” represents the second computing unit 110. “Camera x 7” represents seven visual sensors, such as a camera or a camera, etc. GMSL (Gigabit Multimedia Serial Link) represents a gigabit multimedia serial link. “Lidar x 4” represents four laser radars. ENet (Ethernet) represents Ethernet. “GPS + IMU” represents a GPS (Global Positioning System) positioning device and an IMU (Inertial Measurement Unit) device. CAN (Controller Area Network) represents a controller area network. “Radar x 3” represents three millimeter wave radars.

[0048] Among them, the Orin-A is connected with the visual sensor (such as “Camera”). The Orin-B is connected with other types of sensors (such as “Lidar”, “GPS”, “IMU” and “Radar”). The connection here is “electrical connection” or “communication connection”.

[0049] ACU (Autonomous Control Unit) represents an autonomous control unit. Switch represents a switch. Orin-A is connected with one of the switches, and Orin-B is connected with the other switch. MCU (Microcontroller Unit) represents a microcontroller unit. TC397 represents an automotive-grade microcontroller. The switch is connected with the microcontroller unit.

[0050] T-BOX (Telematics BOX) represents a telematics box. The telematics box is communicatively connected with the autonomous control unit.

[0051] The configuration management module 120 can obtain the current device type through the device type identification interface, start the process management service, load the corresponding configuration file, and perform system service.

[0052] It should be noted that the OS (Operating System) as a basic support layer provides a running environment for the application process and system service (such as process management, resource management, fault diagnosis, etc.) of the upper layer, is responsible for scheduling CPU, memory, storage, network and other hardware resources, coordinates the execution of various tasks, and is a basic software platform for the dual Orin computing unit 110 to realize dynamic deployment, task scheduling and fault processing.

[0053] Optionally, in the configuration management module 120, the system framework and service layer include basic services, system services, a system framework and a basic library. The system services include process management services. The process management services include resource management, permission management, device state management, life cycle management and configuration management. The system framework includes a resource management framework. The resource management framework includes CPU (Central Processing Unit) resources, network resources, memory resources and storage resources. The permission framework includes permission configuration and permission contribution. The basic library can be boost (Boost C++ Libraries, an open source cross-platform function library collection).

[0054] The task scheduling engine module 130 is used for dynamically determining a task deployment target according to task priority, computing power demand, current load state and the like.

[0055] Optionally, in the task scheduling engine module 130, the system framework and service layer includes basic services, system services, system framework, and basic libraries. Among them, the system services include process management services. The process management services include resource management, permission management, device state management, life cycle management, and configuration management. The system framework includes a resource management framework. The resource management framework includes CPU resources, network resources, memory resources, and storage resources. The permission framework includes permission configuration and permission contribution. The basic library can be boost.

[0056] The fault diagnosis monitoring module 140 is used to monitor the health status of the dual Orin device (two computing units 110) in real time, including heartbeat, process state, and system resource usage.

[0057] Optionally, in the fault diagnosis monitoring module 140, the system framework and service layer includes basic services, system services, system framework, and basic libraries. Among them, the system services include fault diagnosis services. The fault diagnosis services include fault receiver, fault diagnosis processing, fault storage, fault reporting, and fault management services. The system framework includes a fault diagnosis framework. The fault diagnosis framework includes fault protocol format, fault transmission framework, and fault reporting framework.

[0058] Optionally, in the fault diagnosis monitoring module 140, the system services include state management services. The state management services include device state, sensor state, resource state, business state, process state, and state management.

[0059] In some embodiments, optionally, as shown in Figure 2 The computing platform 100 further includes a communication middleware module 150. The communication middleware module 150 supports DDS (Data Distribution Service)-based cross-device task state synchronization and fault reporting.

[0060] Optionally, in the communication middleware module 150, the system framework and service layer includes basic services, system services, system framework, and basic libraries. The system framework includes a communication software bus framework. The communication software bus framework includes SOA (Service-Oriented Architecture) and DOA (Data-Oriented Architecture).

[0061] Among them, SOA includes IDL (Interface Definition Language), Event, Field, Method, service discovery and Transport. Transport continues to encapsulate on the basis of DOA. Service discovery is distributed. "Event", "Field" and "Method" represent access patterns.

[0062] In addition, DOA includes IDL, PUB (Publish), SUB (Subscribe), service discovery and Transport. Service discovery supports distribution.

[0063] In some embodiments, as shown in Figure 2 The computing platform 100 also includes an OTA deployment module 160. Among them, "OTA" is Over-The-Air, which means over-the-air download. The OTA deployment module 160 supports unified packaging and differential deployment of dual-device configuration files.

[0064] As shown in Figure 4 The unmanned system dynamic control method includes:

[0065] S202, call the device type identification interface through the configuration management module, obtain the device types corresponding to the two computing units, and load the configuration files corresponding to the device types.

[0066] When the unmanned system starts, the device type identification interface is called to obtain the current device type (such as Orin-A and Orin-B) to distinguish two different working environments.

[0067] According to the device type, the corresponding configuration file is loaded. For example: config_orin_a.yaml (a file format) or config_orin_b.yaml (a file format).

[0068] Optionally, the configuration file includes device computing power capability information, task mapping relationship information, heartbeat period information and fault code definition information.

[0069] S204, based on the configuration file, initialize the task scheduling strategy through the task scheduling engine module, establish the task priority table, the task computing power demand table and the device load table.

[0070] The purpose of this step is to initialize the task scheduling engine, establish the task priority table, the task computing power demand table and the device load table.

[0071] S206, according to the task priority table, the task computing power demand table and the device load table, the priority and the computing power demand of the plurality of to-be-processed tasks are acquired, and the plurality of to-be-processed tasks are distributed to the two computing units for processing.

[0072] The task priority includes an emergency task (such as perception fusion, emergency braking), a high-priority task (such as path planning) and a low-priority task (such as log recording).

[0073] Each to-be-processed task is marked, and the marking content includes the computing power demand, the real-time requirement and the fault tolerance level.

[0074] According to the marking content, the plurality of to-be-processed tasks are distributed to the two computing units for processing, so as to optimize the resource configuration.

[0075] S208, the running data of the two computing units are collected by the fault diagnosis monitoring module; in the case that the running data of one of the computing units is abnormal, at least part of the plurality of to-be-processed tasks that need to be processed by the abnormal computing unit are transferred to the other computing unit.

[0076] The running data includes load state data, heartbeat state data, process running state data and fault code data.

[0077] If it is detected that a certain device (a certain computing unit) is abnormal (such as heartbeat loss, process crash), a fault switching mechanism is triggered.

[0078] If the load of a certain device (a certain computing unit) is too high or a fault occurs, part of the tasks are migrated to another device or another computing unit. This dynamic control mode supports a hot switching mechanism and ensures that the system operation is not affected during the task migration process.

[0079] The application aims to provide a dynamic control method for an unmanned driving system, based on the priority and the computing power demand of tasks, a plurality of to-be-processed tasks are processed by two computing units. Through the fault diagnosis monitoring module, the running data are monitored, and in the case that one of the computing units is abnormal, at least part of the plurality of to-be-processed tasks that need to be processed are transferred to the other computing unit. Through the device type identification, the task priority evaluation and the running data monitoring mechanism, the dynamic deployment and the load balancing of the to-be-processed tasks between the two computing units are realized. This design mode is beneficial to improving the computing power utilization rate, the real-time response capability and the system robustness of the automatic driving system.

[0080] In some embodiments, optionally, in the case that one of the computing units is abnormal, the fault diagnosis monitoring module notifies the task scheduling engine, and at least part of the to-be-processed tasks on this computing unit are migrated to the other computing unit through the task scheduling engine.

[0081] Report fault information to the gateway through the communication middleware module to trigger the fault diagnosis and recovery process. Differentiate device types through fault codes to facilitate fault location and processing.

[0082] In some embodiments, optionally, the configuration files of the two computing units are packaged uniformly through the OTA deployment module; the configuration is selectively loaded according to the device type during updating; the remote update task mapping relationship, computing power strategy and fault handling mechanism are supported.

[0083] The unmanned system dynamic control method of the present application has the following advantages:

[0084] 1. Improved computing power utilization: through the task dynamic deployment mechanism, the computing power of the double Orin device (two computing units) is fully utilized;

[0085] 2. Enhanced system reliability: support task migration and high availability switching when the device fails;

[0086] 3. Improved system real-time performance: dynamically schedule tasks according to task priority and real-time requirements;

[0087] 4. Flexible and scalable: support OTA remote update deployment strategy and configuration;

[0088] 5. Strong compatibility: support differential configuration of double devices, adapt to different vehicle models and platforms.

[0089] In some embodiments, optionally, as shown in Figure 5 S206 (determining the priority and computing power requirement of the plurality of to-be-processed tasks according to the task priority table, the task computing power requirement table and the device load table) includes:

[0090] S2062, according to the task priority table, the task computing power requirement table and the device load table, determine the priority and computing power requirement of the plurality of to-be-processed tasks.

[0091] Optionally, the basic priority coefficient of each to-be-processed task is extracted from the task priority table, and the priority of the to-be-processed task is corrected in combination with the task correlation degree (such as the dependency relationship between perception tasks and decision tasks) to generate the final priority weight.

[0092] It should be noted that the basic priority coefficient is divided into 1 to 10, wherein 10 represents the highest priority.

[0093] Optionally, the number of CPU (Central Processing Unit) cores, GPU (Graphics Processing Unit) computing power, and memory occupation of each to-be-processed task are parsed from the task computing power requirement table to form a standardized computing power requirement matrix.

[0094] Optionally, current load data (such as CPU utilization, GPU utilization, remaining memory, temperature, etc.) of the two computing units are collected in real time from the device load table, and a load trend (such as an average load growth rate within 5 seconds) is calculated through a sliding window algorithm.

[0095] S2064, according to the priority and computing power requirement of the to-be-processed task, the plurality of to-be-processed tasks are divided into emergency tasks, high-priority tasks, and low-priority tasks.

[0096] The priority of the emergency task is the highest, and the computing power requirement fluctuation is small. The computing power requirement of the high-priority task is medium and can be dynamically adjusted. The computing power requirement of the low-priority task is low and can be interrupted.

[0097] The emergency task includes but is not limited to a perception fusion task and an emergency braking task. The high-priority task includes but is not limited to a path planning task. The low-priority task includes but is not limited to a log recording task.

[0098] S2066, based on the task category, computing power requirement, and current load state of the device load table, the plurality of to-be-processed tasks are assigned to the two computing units for processing.

[0099] Optionally, the device load table is refreshed every first time length, and if the load rate of one of the computing units exceeds a first load threshold, a task migration mechanism is triggered, and low-priority tasks are preferentially moved. If the load rate of this computing unit still exceeds the first load threshold, part of the high-priority tasks (such as non-real-time path planning) are migrated to the other computing unit, and the migration process synchronizes the task state through the DDS protocol to ensure data consistency.

[0100] Optionally, the first time length is 80 ms to 120 ms.

[0101] In one specific embodiment, the first time length is 80 ms.

[0102] In one specific embodiment, the first time length is 90 ms.

[0103] In one specific embodiment, the first time length is 100 ms.

[0104] In one specific embodiment, the first time length is 110 ms.

[0105] In one specific embodiment, the first time length is 120 ms.

[0106] Optionally, the first load threshold is 70% to 90%.

[0107] In one specific embodiment, the first load threshold is 70%.

[0108] In one specific embodiment, the first load threshold is 80%.

[0109] In one specific embodiment, the first load threshold is 90%.

[0110] The load balancing-based allocation logic controls the average load difference between the two computing units within 15%, and the GPU utilization rate is increased from 60% of the traditional static allocation to 80%, avoiding overloading of a certain computing unit or idling of resources.

[0111] The computing platform intelligently schedules and fails over the tasks to be processed between the two computing units, and realizes efficient operation and resource optimization of the unmanned system through dynamic task allocation and computing power sharing.

[0112] In some embodiments, optionally, the emergency task includes one or a combination of the following: a perception fusion task and an emergency braking task.

[0113] The perception fusion task includes fusing data of multiple sensors (such as “Lidar”, “Camera” and “Radar”) (fusing point cloud data and image data), obstacle dynamic identification, trajectory prediction and surrounding environment perception.

[0114] The emergency braking task includes obstacle collision risk assessment (calculating collision time based on relative speed and distance), brake command generation (combining road surface friction coefficient and current vehicle speed), and actuator (brake system) response confirmation.

[0115] The high-priority task includes a path planning task.

[0116] The path planning task includes short-term path planning, long-term path planning and lane change decision.

[0117] The low-priority task includes a log recording task.

[0118] Timestamped recording of sensor raw data (such as radar point cloud segments and visual sensor image frames), system running status (CPU / GPU load, task execution time consumption), and abnormal events (such as temporary sensor failure and computing power fluctuation).

[0119] By finely dividing the tasks to be processed, the computing power requirements of various tasks are determined, so that the system can dynamically adjust the computing power allocation according to the environment, which is beneficial to optimize resource allocation and avoid overloading or idling of a certain computing unit.

[0120] In some embodiments, optionally, the configuration file includes device computing power capability information, task mapping relationship information, heartbeat cycle information, and fault code definition information.

[0121] The device computing power capability information refers to the information set of the hardware performance boundary and running limit of the two computing units. The device computing power capability information includes chip specifications (such as CPU core number, GPU architecture, and computing power upper limit), resource threshold (such as CPU / GPU safe utilization rate and memory occupation warning line), and dynamic adjustment rules (such as computing power frequency reduction strategy when the temperature is too high).

[0122] The task mapping relationship information refers to the information defining the binding rules between unmanned driving tasks and computing units and hardware resources. The task mapping relationship information includes task default deployment target (such as emergency task preferentially assigned to Orin-A), task and resource forced association (such as perception fusion task must occupy GPU resource), and data dependency between tasks (such as path planning needs to wait for perception fusion result).

[0123] The heartbeat cycle information refers to the information defining the system health monitoring timing rules. The heartbeat cycle information includes computing unit heartbeat frequency (such as computing unit sending heartbeat every 10 ms), sensor heartbeat frequency (such as sensor sending heartbeat every 50 ms), timeout threshold (such as 3 consecutive heartbeat loss determining as fault), and heartbeat data format (carrying real-time load, temperature, and other information).

[0124] The fault code definition information refers to the information classifying, coding, and binding processing strategies for possible system faults. The fault code definition information includes fault classification coding (such as hardware fault, software fault, emergency level, and warning level), fault reason description, and corresponding processing strategy (such as emergency fault triggering task migration, and warning fault only recording log).

[0125] The device computing power capability information, task mapping relationship information, heartbeat cycle information, and fault code definition information jointly constitute the "operation rule library" of the system, ensuring that the two computing units can dynamically allocate tasks according to the design logic, monitor the health status, and handle faults, which is the core support for realizing computing power optimization and high reliability.

[0126] In some embodiments, optionally, the running data includes load state data, heartbeat state data, process running state data, and fault code data.

[0127] The load state data refers to quantitative data describing real-time occupation of system hardware resources (such as computing, storage, network) and environment-related states.

[0128] The load state data includes one or a combination of the following: CPU utilization, GPU computing power occupation, video memory usage, chip core temperature, real-time power consumption, power supply voltage stability, hard disk read / write speed, cache hit rate, and storage capacity occupation ratio.

[0129] The heartbeat state data refers to monitoring data generated by periodically sending a "survival signal" (heartbeat packet) by a device, module or component, for determining whether the target is online and whether the communication is normal.

[0130] The heartbeat state data includes one or a combination of the following: heartbeat packet sending interval, actual receiving delay, receiving success rate, computing unit heartbeat data, and sensor heartbeat data.

[0131] The process running state data refers to monitoring data of real-time behavior and resource occupation of each task process (such as sensing, planning, and control tasks) in the system, for feeding back whether the tasks are normally executed.

[0132] The process running state data includes one or a combination of the following: start-up time, running duration, process-level CPU / GPU occupation, memory allocation, single task time consumption (such as sensing fusion time consumption less than or equal to 50 ms), and time consumption fluctuation range (the smaller the standard deviation, the more stable).

[0133] The fault code data refers to structured coding and associated information generated according to preset rules when the system has an abnormality, for accurately identifying fault types, locations and contexts.

[0134] The fault code data includes one or a combination of the following: fault code, fault level (emergency, warning, and prompt), and fault occurrence timestamp.

[0135] Fine collection and analysis of running data are core supports for the dual-system-level chip to achieve "dynamic deployment, fault self-recovery, and safe operation", and through data-driven decision optimization, the reliability, efficiency and maintainability of the system are significantly improved.

[0136] Optionally, the fault diagnosis and monitoring module includes a health state monitoring module. The health state monitoring module continuously collects current loads (CPU / GPU / memory) of the two computing units, heartbeat states (software dog and hardware dog), process running states, and fault code reports.

[0137] It should be noted that in the heartbeat state, "soft dog" and "hard dog" refer to software watchdog and hardware watchdog, which are two mechanisms for monitoring the running state of the system. The soft dog (software watchdog) realizes the watchdog function through the software timer, and the time is also essentially dependent on the hardware timer on the hardware peripheral. The hard dog (hardware watchdog) realizes the watchdog function through the mechanism of the hardware itself, and is essentially based on the timer principle.

[0138] In some embodiments, optionally, as shown in S208 (collecting running data of two computing units by a fault diagnosis monitoring module; in the case that the running data of one of the computing units is abnormal, transferring at least part of a plurality of to-be-processed tasks that need to be processed by the computing unit with abnormal running data to another computing unit), the dynamic control method of the unmanned system further comprises: Figure 6

[0139] S210, reporting fault information to a gateway, and performing fault diagnosis and recovery on the computing unit with abnormal running data according to the fault information.

[0140] The fault information includes basic information, context data and processing trajectory. The basic information includes the number corresponding to the computing unit with abnormal running data, the abnormal time stamp and the abnormal level.

[0141] The context data includes the load curve (CPU / GPU utilization rate change) of the last 3 seconds before the abnormality, the heartbeat loss record and the associated process state (such as sensing whether the process crashes).

[0142] The processing trajectory includes the list of tasks that have been transferred, the current load of the target computing unit and the transfer time consumption.

[0143] After the fault occurs, part of the tasks are transferred, the fault information is reported, and the fault diagnosis and recovery are performed. This dynamic control method significantly improves the reliability and safety of the unmanned system through the whole process design of "accurate monitoring-intelligent transfer-closed loop recovery".

[0144] In some embodiments, optionally, the computing power requirement of the to-be-processed task includes CPU utilization rate and GPU utilization rate.

[0145] By defining the computing power requirement of CPU utilization rate and GPU utilization rate in detail, the "quantifiable and adaptable" basis can be provided for the task scheduling and resource allocation of the unmanned system, the resource allocation accuracy is improved, the waste or overload of computing power is avoided, the cache and video memory utilization is optimized, and the computing efficiency is improved.

[0146] In an embodiment of the present application, as shown in Figure 7 ​As shown, the unmanned system dynamic control device 300 includes a device type identification unit 310, a configuration file loading unit 320, a task scheduling strategy initialization unit 330, a task processing unit 340, a running data acquisition unit 350, and a task dynamic deployment unit 360.

[0147] The device type identification unit 310 is configured to call a device type identification interface through the configuration management module 120 to obtain the device types corresponding to the two computing units 110.

[0148] When the unmanned system is started, the device type identification interface is called to obtain the current device types (such as Orin-A and Orin-B) to distinguish two different working environments.

[0149] The configuration file loading unit 320 is configured to load the configuration file corresponding to the device type.

[0150] The corresponding configuration file is loaded according to the device type. For example, config_orin_a.yaml (a file format) or config_orin_b.yaml (a file format).

[0151] Optionally, the configuration file includes device computing power capability information, task mapping relationship information, heartbeat period information, and fault code definition information.

[0152] The task scheduling strategy initialization unit 330 is configured to initialize the task scheduling strategy based on the configuration file through the task scheduling engine module 130 to establish a task priority table, a task computing power demand table, and a device load table.

[0153] The task processing unit 340 is configured to obtain the priorities and computing power demands of a plurality of to-be-processed tasks according to the task priority table, the task computing power demand table, and the device load table, and distribute the plurality of to-be-processed tasks to the two computing units 110 for processing.

[0154] The task priority includes an emergency task (such as perception fusion and emergency braking), a high-priority task (such as path planning), and a low-priority task (such as log recording).

[0155] Each to-be-processed task is marked, and the marking content includes computing power demand, real-time requirement, and fault tolerance level.

[0156] According to the marking content, the plurality of to-be-processed tasks are distributed to the two computing units 110 for processing to optimize resource allocation.

[0157] The running data acquisition unit 350 is configured to acquire the running data of the two computing units 110 through the fault diagnosis monitoring module 140.

[0158] The running data includes load state data, heartbeat state data, process running state data and fault code data.

[0159] The task dynamic deployment unit 360 is configured to, in the case where the running data of one of the computing units 110 is abnormal, transfer at least part of the multiple to-be-processed tasks that need to be processed by the abnormal computing unit 110 to another computing unit 110.

[0160] If it is detected that a certain device (a certain computing unit 110) is abnormal (for example, heartbeat loss or process crash), a fault switching mechanism is triggered.

[0161] If the load of a certain device (a certain computing unit 110) is too high or a fault occurs, part of the tasks are migrated to another device (another computing unit 110). This dynamic control mode supports a hot switching mechanism and ensures that the system operation is not affected during the task migration process.

[0162] The present application aims to provide a dynamic control device 300 of an unmanned system, which processes multiple to-be-processed tasks through two computing units 110 based on the priority of the tasks and the computing power requirement. The running data is monitored by a fault diagnosis monitoring module 140, and in the case where one of the computing units 110 is abnormal, at least part of the multiple to-be-processed tasks that need to be processed are transferred to another computing unit 110. Through the mechanisms of device type identification, task priority evaluation and running data monitoring, the dynamic deployment and load balancing of the to-be-processed tasks between the two computing units 110 are realized. This design is beneficial to improving the computing power utilization rate, real-time response capability and system robustness of the automatic driving system.

[0163] In an embodiment of the present application, as shown in Figure 8 The electronic device 400 includes a memory 410 and a processor 420. The memory 410 stores programs or instructions that can be run on the processor 420. When the processor 420 executes the programs or instructions, the steps of the dynamic control method of the unmanned system in any of the above embodiments are implemented. The electronic device 400 has the beneficial effects of any of the above embodiments, which are not repeated here.

[0164] In an embodiment of the present application, a readable storage medium stores programs or instructions, which are executed by a processor to implement the steps of the dynamic control method of the unmanned system in any of the above embodiments. The readable storage medium has the beneficial effects of any of the above embodiments, which are not repeated here.

[0165] In the present application, the terms "first", "second", "third" are only used for descriptive purpose, and should not be understood as indicating or implying relative importance. The term "multiple" refers to two or more, unless otherwise explicitly limited. The terms "mount", "connect", "connection", "fix", and the like should be interpreted broadly, for example, "connection" can be fixed connection, or detachable connection, or integral connection; "connection" can be direct connection, or indirect connection through intermediate medium. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0166] In the description of the present application, it should be understood that the terms "upper", "lower", "left", "right", "front", "back", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the device or unit referred to must have a particular direction, be constructed and operated in a particular orientation, therefore, should not be understood as a limitation on the present application.

[0167] In the description of the present application, the terms "one embodiment", "some embodiments", "a specific embodiment", and the like, mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0168] The above is only the preferred embodiment of the present application, and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A dynamic control method for unmanned systems, characterized in that, The application discloses a computing platform applied to a dual-system level chip, and the computing platform comprises two system level chip-based computing units, a configuration management module, a task scheduling engine module and a fault diagnosis monitoring module. The dynamic control method of the unmanned system comprises: The configuration management module calls a device type identification interface to obtain device types corresponding to the two computing units and loads configuration files corresponding to the device types; Based on the configuration files, the task scheduling engine module initializes a task scheduling strategy, establishes a task priority table, a task computing power demand table and a device load table; According to the task priority table, the task computing power demand table and the device load table, the priorities and computing power demands of a plurality of to-be-processed tasks are obtained, and the plurality of to-be-processed tasks are distributed to the two computing units for processing; The fault diagnosis monitoring module collects running data of the two computing units; in the case that the running data of one of the computing units is abnormal, at least part of the to-be-processed tasks that need to be processed by the computing unit with the abnormal running data is transferred to the other computing unit.

2. The dynamic control method of an unmanned system according to claim 1, wherein, The method comprises: According to the task priority table, the task computing power demand table and the device load table, the priorities and computing power demands of a plurality of to-be-processed tasks are obtained, and the plurality of to-be-processed tasks are distributed to the two computing units for processing; According to the task priority table, the task computing power demand table and the device load table, the priorities and computing power demands of a plurality of to-be-processed tasks are obtained, and the plurality of to-be-processed tasks are distributed to the two computing units for processing; The emergency task comprises one or a combination of the following: a perception fusion task and an emergency braking task; 3. The dynamic control method of an unmanned system according to claim 2, wherein, The high-priority task comprises a path planning task; The low-priority task comprises a log recording task. The configuration file comprises device computing power capability information, task mapping relationship information, heartbeat cycle information and fault code definition information.

4. The dynamic control method of an unmanned system according to any one of claims 1 to 3, wherein, The running data comprises load state data, heartbeat state data, process running state data and fault code data.

5. The dynamic control method of an unmanned system according to any one of claims 1 to 3, wherein, Further comprising:

6. The dynamic control method of an unmanned system according to any one of claims 1 to 3, wherein, The fault diagnosis monitoring module collects running data of the two computing units; In the case that the running data of one of the computing units is abnormal, at least part of the to-be-processed tasks that need to be processed by the computing unit with the abnormal running data is transferred to the other computing unit, then fault information is reported to a gateway, and fault diagnosis and recovery are performed on the computing unit with the abnormal running data according to the fault information. The computing power demand of the to-be-processed task comprises CPU utilization and GPU utilization.

7. The dynamic control method of an unmanned system according to any one of claims 1 to 3, wherein, The device type identification unit (310) is configured to call a device type identification interface through the configuration management module (120) to obtain device types corresponding to the two computing units (110); 8. An unmanned system dynamic control apparatus, comprising: ​ ​ A configuration file loading unit (320) is configured to load a configuration file corresponding to the device type; A task scheduling strategy initialization unit (330) is configured to initialize a task scheduling strategy by a task scheduling engine module (130) based on the configuration file, and establish a task priority table, a task computing power requirement table, and a device load table; A task processing unit (340) is configured to obtain priorities and computing power requirements of a plurality of to-be-processed tasks according to the task priority table, the task computing power requirement table, and the device load table, and distribute the plurality of to-be-processed tasks to the two computing units (110) for processing; An operation data acquisition unit (350) is configured to acquire operation data of the two computing units (110) by a fault diagnosis monitoring module (140); A task dynamic deployment unit (360) is configured to, in a case where operation data of one of the computing units (110) is abnormal, transfer at least part of a plurality of to-be-processed tasks that need to be processed by the computing unit (110) with the abnormal operation data to the other computing unit (110).

9. An electronic device, comprising: Comprise: A memory (410) and a processor (420), wherein the memory (410) stores programs or instructions executable on the processor (420), and the processor (420) implements the steps of the dynamic control method of the unmanned system according to any one of claims 1 to 7 when executing the programs or the instructions.

10. A readable storage medium, characterized by, The readable storage medium stores programs or instructions, and the programs or the instructions are executed by the processor to implement the steps of the dynamic control method of the unmanned system according to any one of claims 1 to 7.