A multi-model adaptive scheduling method for automatic driving
By employing a multi-model adaptive scheduling method and utilizing discrete optimization coding and a dynamic pruning evaluation framework, the high precision and low latency issues of autonomous driving systems in complex scenarios are addressed. This achieves efficient and dynamic model scheduling, thereby improving system performance and safety.
Patent Information
- Application Number
- CN202511047623.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing autonomous driving systems struggle to simultaneously meet the requirements of high precision and low latency in complex scenarios. Static scheduling lacks flexibility, while dynamic scheduling ignores scenario characteristics, leading to unstable performance. Existing methods also struggle to achieve real-time model switching.
A multi-model adaptive scheduling method is adopted, which transforms the network selection problem into a 0-1 planning problem through discrete optimization coding technology. Combined with scene density discretization algorithm and dynamic pruning evaluation framework, the optimal network combination is selected quickly, thus optimizing computational complexity and real-time performance.
It significantly improves the performance of autonomous driving systems in diverse scenarios, achieving efficient and dynamic model scheduling, avoiding the performance bottlenecks of static scheduling and the blindness of simple dynamic scheduling, and ensuring safe driving in complex environments.
Smart Images

Figure CN120540105B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of autonomous driving, specifically to a multi-model adaptive scheduling method for autonomous driving. Background Technology
[0002] A complete autonomous driving system typically comprises three main modules: perception, decision-making, and control. These three modules work together to achieve autonomous navigation in complex environments. The perception module relies on data fusion from multiple sensors (such as cameras and LiDAR) and combines it with deep learning models (such as CNNs for object detection and RNNs for trajectory prediction) to perform real-time identification of environmental elements (roads, vehicles, pedestrians, etc.). The decision-making module performs path planning and behavioral decisions based on multimodal perception results. The control module maps decision commands into vehicle dynamics control signals. However, with the exponential increase in scenario complexity, significant contradictions exist between existing modules: on the one hand, a single model struggles to simultaneously meet the requirements of real-time performance, robustness, and generalization in dynamic scenarios; on the other hand, the computational constraints of in-vehicle embedded platforms further exacerbate the challenges of multi-model collaboration. Therefore, it is urgent to construct an intelligent model scheduling mechanism to achieve precise allocation of computing resources through dynamic model combination selection and execution order optimization. This mechanism, through dynamic model combination selection, achieves a balance between accuracy and latency under the computational limitations of in-vehicle embedded platforms, ensuring that the system maintains high reliability and low-latency response during complex scenario transitions.
[0003] Traditional autonomous driving model scheduling methods mainly fall into two categories: static scheduling and rule-driven dynamic scheduling. Static scheduling methods allocate computational resources by pre-setting fixed model combinations. The PointPillars model proposed by Alex H. Lang et al. for complex urban scenarios, which fuses LiDAR point clouds with 3D object detection, achieved over 90% detection accuracy in complex urban environments. However, its fixed model configuration leads to significant resource waste in sparse target scenarios. Wang et al. also pointed out that when vehicles enter tunnels from highways, the pre-set fixed high-precision model experiences a sharp increase in visual detection miss rate to 25% due to LiDAR failure, causing a surge in decision latency. The static nature of such methods poses extremely high safety risks. Rule-driven dynamic scheduling methods attempt to achieve model switching through conditional triggers, but their modeling depth and real-time performance remain bottlenecks. Cevher et al.'s cross-modal fusion scheme only fuses LiDAR and visual features with fixed weights, without designing a dynamic weight allocation mechanism. This causes the low-latency advantage of the visual model in complex scenarios to be offset by redundant LiDAR computation, resulting in an overall system energy efficiency ratio of only 65% of the theoretical peak. The scheduling strategy proposed by Shi et al., based on GPU load thresholds, triggers model degradation by setting 85% GPU utilization. However, it fails to consider the dynamic characteristics of the target, resulting in severe model switching lag during scene changes. These methods attempt to trigger model switching through hardware metrics, but the lack of deep analysis of scene semantics leads to decision-making lag, making it difficult to achieve real-time model switching under complex scene changes.
[0004] In summary, existing model scheduling methods for autonomous driving systems have significant shortcomings when adapting to complex scenarios. Static scheduling lacks flexibility and cannot cope with dynamic changes in target density; simple dynamic scheduling rules ignore scenario characteristics, leading to unstable performance; while improved methods have enhanced adaptability to some extent, they still suffer from low real-time performance, high switching overhead, or excessive computational complexity. These limitations make it difficult for existing systems to simultaneously meet the requirements of high accuracy and low latency in diverse autonomous driving scenarios, especially in complex traffic environments with drastic changes in target density. Therefore, there is an urgent need for a method that can dynamically schedule multiple models based on scenario target density to optimize the performance of autonomous driving systems. Summary of the Invention
[0005] To address the shortcomings of current research, this invention provides a multi-model adaptive scheduling method for autonomous driving, aiming to overcome the deficiencies of existing technologies in dynamic model scheduling, such as insufficient scene adaptability and difficulty in co-optimizing accuracy and latency. By utilizing discrete optimization coding techniques for network configuration, the multi-component network selection problem is transformed into a discrete optimization problem, significantly reducing decision complexity and ensuring the system can quickly select the optimal network combination in real-time operation. A scene density discretization algorithm based on target perception is employed to achieve real-time density perception in complex scenes and efficiently calculate scene complexity. A quantitative prediction model of network performance and scene density is constructed to accurately fit the relationship between accuracy and latency and density and network schemes. A network selection optimization framework with dynamic pruning and parallel evaluation is constructed, modeling the network selection problem as a single-objective optimization 0-1 programming problem. The solution space is systematically explored and pruned to quickly find the optimal network combination, significantly reducing computational complexity, thereby achieving efficient and dynamic model scheduling in real-time autonomous driving scenarios.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A multi-model adaptive scheduling method for autonomous driving includes:
[0008] S1 abstracts the autonomous driving system into four components: feature extraction module, trajectory tracking module, semantic segmentation module, and planning and decision-making module.
[0009] S2 provides multiple network options for feature extraction, trajectory tracking, and semantic segmentation, forming various different network configuration schemes. Each network configuration scheme is identified in the form of a triple.
[0010] S3 performs discretization encoding on all network configuration schemes;
[0011] S4, Define the optimization objective;
[0012] S5 collects density data samples from multiple scenarios to build diverse, high-quality datasets;
[0013] S6 quantifies the density of scene targets and classifies target density levels;
[0014] S7 utilizes diverse high-quality datasets and groups these diverse high-quality datasets according to target density levels;
[0015] S8 was used to test the model accuracy under scene data of different densities and different network configurations.
[0016] S9, establish the scene target density, network configuration scheme, and accuracy density-network-accuracy fitting function;
[0017] S10, testing latency under different target densities and network configuration schemes in different scenarios;
[0018] S11, Establish the scene target density, network configuration scheme, and latency density-network-latency fitting function;
[0019] S12, the optimal network selection scheme is modeled based on the density-network-accuracy fitting function and the density-network-time delay fitting function, and established as a discrete combinatorial optimization problem;
[0020] S13. For discrete combinatorial optimization problems, an improved branch and bound method is used to solve them and output the identifier of the optimal solution.
[0021] S14: Based on the optimal solution identifier, call the corresponding feature extraction network, trajectory tracking network, and semantic segmentation network from the pre-loaded network library.
[0022] Furthermore, in S1:
[0023] The feature extraction module is used to extract key feature information from the raw sensor data;
[0024] The trajectory tracking module is used to identify and track dynamic targets in a scene using the results of feature extraction, and to predict their future trajectory.
[0025] The semantic segmentation module is used to classify scene pixels;
[0026] The planning and decision-making module is used to integrate the results of trajectory tracking and semantic segmentation to formulate vehicle driving strategies.
[0027] Further, S2 includes:
[0028] It provides multiple backbone networks from ResNet34, ResNet50, ResNet101, and ResNet152 for feature extraction;
[0029] Two networks are provided for trajectory tracking: DenseTrackNet and SparseTrackNet;
[0030] Two networks are provided for semantic segmentation: UNet and ENet;
[0031] These networks can be combined independently to form a variety of different network configuration schemes, each of which is identified by a triple.
[0032] S3 includes:
[0033] One-hot encoding quantifies the network configuration scheme, ensuring that the network selection for each component is unique and mutually exclusive; feature extraction uses a 4-bit vector, trajectory tracking uses a 2-bit vector, semantic segmentation uses a 2-bit vector, and finally concatenates them into an 8-dimensional 0-1 vector.
[0034] Further, S4 includes:
[0035] We construct optimization indices using weighted interpolation:
[0036] ,
[0037] Where H is the evaluation index used to evaluate the balance between accuracy and latency, ACC is the accuracy of system operation, Latency is the latency of system operation, and λ is the weighting coefficient used to adjust the priority of accuracy and latency. When the system is running, the H value of each network configuration scheme will be calculated based on the real-time target density using the fitted accuracy and latency functions, and the network configuration scheme with the largest H value will be selected as the optimal configuration.
[0038] S5 includes:
[0039] In the simulation environment Carla, different target density scenarios are actively constructed, and data samples are generated by adjusting the target density of vehicles and pedestrians. Real scene data with similar density to the Carla constructed sample are inserted into different target density scenarios. Finally, data samples with different target densities will be cleaned to remove erroneous samples.
[0040] Further, S6 includes:
[0041] The pre-trained lightweight target detection model YOLOV8 is used to analyze the images input from the sensor, detect and count the number and categories of targets in the scene, assign different weights W to different target categories, and divide the scene into multiple target density levels based on the number of targets.
[0042] Further, S7 includes:
[0043] The multi-scene density data samples collected by S5 are used, and the data of different target densities are grouped according to the multiple target density levels divided by S6.
[0044] S8 includes: for each set of data samples, executing multiple network configuration schemes defined in S2, testing the comprehensive accuracy of each network configuration scheme, defining it as the accuracy rate of planning decisions, and finally obtaining multiple accuracy results.
[0045] Further, S9 includes:
[0046] The scene target density is treated as a discrete value, and an 8-dimensional 0-1 vector defined by S3 is sampled to represent each network configuration scheme. The input data of the density-network-accuracy fitting function is all triples [density, network encoding, accuracy]. The function is fitted by a regression model to obtain an accuracy function that can predict the system accuracy based on a given target density and network configuration scheme.
[0047] Further, S11 includes:
[0048] The scene target density is treated as a discrete value, and an 8-dimensional 0-1 vector defined by S3 is sampled to represent each network configuration scheme. The input data of the density-network-delay fitting function is all triples [density, network encoding, delay], which are fitted by a regression model.
[0049] Further, S12 includes:
[0050] By maximizing the optimization objective defined in S4 Choose the optimal network configuration and model the task as a constrained discrete combinatorial optimization problem:
[0051] ,
[0052] Where max represents taking the maximum value, f acc f represents the precision function of the target density D and the network configuration scheme O. Latency The time delay function represents the target density D and the network configuration scheme O;
[0053] Constraints: The value of O is a 0-1 code representing the selection results of eight network configuration schemes, where 0 represents not selecting the scheme and 1 represents selecting the scheme. The sum of the first to fourth selection results is 1, the sum of the fifth to sixth selection results is 1, and the sum of the seventh to eighth selection results is 1.
[0054] ,
[0055] The discrete combinatorial optimization problem was ultimately modeled as a 0-1 integer programming problem, a single-objective combinatorial optimization problem with a finite feasible solution space.
[0056] Further, S13 includes:
[0057] For this discrete combinatorial optimization, an improved branch and bound method is used to solve the problem. All network schemes are constructed as a tree-like decision space, with the root node representing the unselected state and each branch corresponding to a network configuration scheme. By pre-calculating the accuracy-delay benchmark values of various network configurations under the target density, a priority search queue is generated in descending order of the estimated objective function values. During the traversal, the upper and lower bounds of each node are dynamically calculated.
[0058] Compared with the prior art, the beneficial effects of the present invention are:
[0059] 1) Through innovative dynamic scheduling mechanisms and optimized solution strategies, the performance of autonomous driving systems in diverse scenarios has been significantly improved. Discrete optimization coding techniques are used to abstract the multi-network selection problem into 0-1 integer programming, using binary encoding (0-1 vectors) to represent the activation state of model combinations. A scene density discretization algorithm based on target perception is used to divide scene density into 10 levels, quickly locating the complexity level of the current scene. Experimental data is used to fit accuracy (mAP), latency (ms), scene density (D), and model combinations. A quantitative prediction model of multi-network performance and scene density is established to quantify the performance-resource trade-offs of different model combinations in specific density scenarios. A network selection optimization framework with dynamic pruning and parallel evaluation is designed, transforming the final network decision into a single-objective 0-1 programming problem, with maximizing the performance-resource utility ratio as the objective function.
[0060] 2) To address the real-time adaptation of autonomous driving systems in complex scenarios, considering the system's real-time requirements, a real-time scene density estimation method needs to be established. This involves using the lightweight object detection network YOLOv8 to extract the number, category, and speed information of targets in the current scene in real time, and assigning different weights to different target categories using a weighting coefficient table. Based on the weight, quantity, category, and velocity information, these are input into the scene density function to obtain continuous scene density values. The continuous values are then divided into discrete values according to the scene density value partitioning table, ultimately achieving real-time density estimation for different scenes.
[0061] 3) For network scheme selection in a 0-1 programming problem with a limited solution space, considering the algorithm's scalability, instead of exhaustive sampling to calculate 16 combinations, an improved branch-and-bound algorithm is used to find the optimal combination. The Pareto front solution set for all network combinations is pre-calculated. Dynamic weight pruning is used to eliminate candidate schemes with objective function values lower than the current optimum. The remaining candidate networks are prioritized based on scene density levels. A double-buffered parallel evaluation strategy is used to simultaneously calculate the real-time performance indicators of each scheme, and finally, the comprehensive optimal solution is selected. This achieves the selection of the best network combination scheme adapted to the current scene.
[0062] 4) This invention can dynamically adjust the model configuration according to scenario requirements. Through dynamic network adaptation and optimized solution strategies, it avoids the performance bottleneck caused by static scheduling, while reducing the blindness of simple dynamic scheduling, significantly improving the overall performance of the autonomous driving system. Whether in high-density or low-density scenarios, the system can achieve an optimized balance between accuracy and latency, providing strong support for safe driving in complex environments. Attached Figure Description
[0063] Figure 1 This is a flowchart of the present invention.
[0064] Figure 2This is a network structure diagram of the present invention. Detailed Implementation
[0065] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0066] A multi-model adaptive scheduling method for autonomous driving, such as Figure 1 As shown, the main steps include the following:
[0067] Step 1: Abstract the autonomous driving system into the following four components:
[0068] Feature extraction module: Extracts key feature information from raw sensor data to support subsequent modules.
[0069] Trajectory tracking module: Utilizes the results of feature extraction to identify and track dynamic targets in the scene and predict their future trajectories.
[0070] Semantic segmentation module: Classifies scene pixels and identifies elements such as roads, vehicles, and pedestrians.
[0071] Planning and Decision Module: Integrates the results of trajectory tracking and semantic segmentation to formulate vehicle driving strategies.
[0072] Step 2: Based on the abstracted autonomous driving components, configure multiple network options for feature extraction, trajectory tracking, and semantic segmentation. Feature extraction provides four backbone networks: ResNet34, ResNet50, ResNet101, and ResNet152. Trajectory tracking provides two networks: DenseTrackNet and SparseTrackNet. Semantic segmentation provides two networks: UNet and ENet. The network structures of each module are as follows: Figure 2 As shown. By combining these networks independently, a total of 4×2×2 = 16 different network configuration schemes are formed. Each scheme is identified in the form of a triple, such as [ResNet34, DenseTrackNet, ENet].
[0073] Step 3: To facilitate automatic selection and configuration of the optimal solution by the system, these 16 solutions will be quantized and encoded using one-hot encoding to ensure that the network configuration solution for each component is unique and mutually exclusive. Feature extraction uses a 4-bit vector, trajectory tracking uses a 2-bit vector, and semantic segmentation uses a 2-bit vector, which are then concatenated into an 8-dimensional 0-1 vector. For example, if the triplet is identified as [ResNet34, DenseTrackNet, ENet], then the feature extraction module vector is represented as [1,0,0,0], the trajectory tracking module vector is represented as [1,0], and the semantic segmentation module vector is represented as [0,1]. These three vectors are then concatenated into an 8-dimensional 0-1 vector of [1,0,0,0,1,0,0,1].
[0074] The network scheme encoding method is as follows:
[0075] In the feature extraction networks, ResNet34 is [1,0,0,0], ResNet50 is [0,1,0,0], ResNet101 is [0,0,1,0], and ResNet152 is [0,0,0,1]. In the trajectory tracking networks, DenseTrackNet is [1,0] and SparseTrackNet is [0,1]. In the semantic segmentation networks, UNet is [1,0] and ENet is [0,1]. The vector representation of the network selection scheme triple [ResNet50, SparseTrackNet, UNet] is [0,1,0,0,0,1,1,0].
[0076] Step 4: To select the optimal solution from multiple network combination schemes, an optimization objective needs to be defined to find the best balance between accuracy and latency. Specifically, a weighted interpolation method is used to construct the optimization index:
[0077] (1)
[0078] Here, H is the evaluation metric used to assess the balance between accuracy and latency, ACC is the accuracy of system operation, Latency is the latency of system operation, and λ is a weighting coefficient used to adjust the priority of accuracy and latency. During system operation, based on the real-time target density, the H value of each network configuration scheme is calculated using the fitted accuracy and latency functions, and the network configuration scheme with the largest H value is selected as the optimal configuration.
[0079] Step 5: Construct diverse, high-quality datasets to support the fitting data required for subsequent modeling. In the Carla simulation environment, scenarios with different target densities are actively constructed, and samples are generated by adjusting targets such as vehicles and pedestrians. Considering the differences between simulation data and real-world scenarios, real-world scene data of the corresponding density will be inserted into different density scenarios. Finally, erroneous samples will be removed through data cleaning.
[0080] Step 6: Quantize scene target density to support network selection and performance optimization based on scene complexity. The pre-trained lightweight target detection model YOLOv8 is used to analyze the images input from the sensor, detecting and counting the number and categories of targets in the scene. Different weights W are assigned to different target categories, and the scene is divided into 10 density levels based on the number of targets. Specifically: the number of targets in each category is multiplied by the target's weight, then divided by the image area occupied by the identified targets in the current image, and finally multiplied by 1 plus the logarithm of the sum of the target motion velocities. The sum of the calculation results for each target category yields the target density of the current scene. The scene density quantization formula is:
[0081] (2)
[0082] Where C is the total number of target categories, and N is the total number of target categories. i W represents the number of targets of type i. i Let A be the weight coefficient for the i-th type of target, and A be the effective perception area of the scene (unit: m). 2 V represents the average velocity of the target (unit: m / s). The specific weighting coefficient table and density partitioning scheme are as follows:
[0083] Weighting coefficient (W) i ): Pedestrians (1.5), bicycles (1.2), cars (1.0), buses (1.8), motorcycles (1.3), special vehicles (2.0). For density grading, an improved exponential segmentation method is used to map continuous density values S to 10 levels: Level 1 (S≤0.1), Level 2 (0.1<S≤0.5), Level 3 (0.5<S≤1.2), Level 4 (1.2<S≤2.5), Level 5 (2.5<S≤5), Level 6 (5<S≤10), Level 7 (10<S≤20), Level 8 (20<S≤40), Level 9 (40<S≤80), Level 10 (S>80).
[0084] Step 7: Utilize the multi-scene density data samples collected in Step 5, and group the data according to the 10 target density levels defined in Step 6.
[0085] Step 8: For each data sample, execute the 16 network configuration schemes defined in Step 2, and test the overall accuracy of each network configuration scheme, defined as the accuracy of the planning decision (the degree to which the final path plan matches the actual ground conditions). Ultimately, 16 × 10 = 160 accuracy results will be obtained.
[0086] Step 9: Establish the density-network-accuracy fitting function for scene target density, network configuration scheme, and accuracy, specifically including:
[0087] After obtaining 160 accuracy test results, the relationship between accuracy and scene target density and network configuration scheme is established by fitting an accuracy function. The target density is treated as a discrete value (levels 1-10), and each network scheme is represented by an 8-dimensional 0-1 vector defined in step 3. The input data consists of 160 sets of triples [density, network encoding, accuracy]. A regression model is used for fitting, and the density-network-accuracy fitting function is:
[0088] (3)
[0089] Where β0, β1, and β2 are fitting coefficients, D is the scene density, O is the one-hot encoding network selection scheme, and α i O represents the fitting coefficient. i For the one-hot encoding of the i-th network selection scheme, δ i γ i β0 + β1D + β2D are the fitting coefficients, and ε is the fitting error. 2 The relationship between accuracy and density is described, α i O i The relationship between accuracy and network selection is described, (γ) i D+δ i D 2 )O i The joint effect of network density on accuracy is described. The model captures the nonlinear effects of density and scheme on accuracy. To ensure model accuracy and generalization ability, 5-fold cross-validation is used to evaluate the fit, and mean squared error (MSE) is calculated to quantify the prediction error. The final accuracy function can predict the system accuracy based on a given target density and network scheme.
[0090] Step 10: In autonomous driving systems, accuracy is crucial for system planning. However, in practical applications, latency must also be considered to ensure real-time performance. Therefore, a model relating latency to scene density and network scheme is needed. Similar to Step 7, for data samples with different scene densities, the 16 network combination schemes defined in Step 2 are executed, and their overall accuracy is tested, defined as the latency for planning decisions. Ultimately, 16 × 10 = 160 latency results will be obtained.
[0091] Step 11: Treat the scene target density as a discrete value and sample an 8-dimensional 0-1 vector defined by S3 to represent each network configuration scheme. The input data for the density-network-delay fitting function is all triples [density, network encoding, delay], which are then fitted using a regression model. Specifically: The density-network-delay fitting function consists of two stages. In the first stage, a fitting function is constructed between the system delay and the target density, and the density parameters θ0 and θ1 are fitted by minimizing the function result. After fixing the density parameters θ0 and θ1, in the second stage, a fitting function is constructed between the system delay, the target density, and the network scheme, and the parameter η corresponding to the i-th network scheme is fitted by minimizing the function result. i κ i and φ i Finally, 10-fold cross-validation was used to evaluate the fitting effect, and the root mean square error was calculated to quantify the error, ultimately yielding a delay function that can predict system delay based on a given target density and network configuration.
[0092] After obtaining 160 latency test results, similar to step 9, the input data consists of 160 sets of triples [density, network encoding, latency]. However, the fitting function is slightly different because latency increases exponentially due to the network scheme, while accuracy is slightly smoother. The specific density-network-latency fitting function is as follows:
[0093] (4)
[0094] Where L represents latency, exp represents the exponential function, D is the scene density, and O i A network configuration scheme for one-hot encoding, θ0, θ1, η i κ i and φ i To fit the parameters, a staged fitting strategy is adopted to ensure stability. First, the parameters are fitted using all data. and :
[0095] (5)
[0096] min represents taking the minimum value; after fixing the fitting parameter θ, the subnetwork fits η. i κ i and φ i :
[0097] (6)
[0098] To ensure the model's prediction accuracy and generalization ability, 10-fold cross-validation was used to evaluate the fitting effect, and the root mean square error (RMSE) was calculated to quantify the error. The validation results show that the error is controlled within an acceptable range, and the final time delay function can predict the system time delay based on the given target density and network scheme.
[0099] Step 12: After obtaining the density-network-accuracy fitting function and the density-network-time delay fitting function based on steps 9 and 11, maximize the optimization objective. Choose the optimal network configuration. Model this task as a constrained discrete combinatorial optimization problem:
[0100] (7)
[0101] Where max represents taking the maximum value, f acc f represents the precision function of the target density D and the network configuration scheme O. Latency The time delay function represents the target density D and the network configuration scheme O.
[0102] The constraint is that the value of O ranges from 0 to 1, representing the 0-1 encoding of the eight network configuration scheme selection results. 0 indicates that the scheme is not selected, and 1 indicates that the scheme is selected. The sum of the first to fourth selection results is 1, the sum of the fifth to sixth selection results is 1, and the sum of the seventh to eighth selection results is 1. The expression is as follows:
[0103] (8)
[0104] The discrete combinatorial optimization problem was ultimately modeled as a 0-1 integer programming problem, a single-objective combinatorial optimization problem with a finite feasible solution space.
[0105] Step 13: For this discrete combinatorial optimization, the branch and bound method is used for solution. The 16 network schemes are constructed as a tree-like decision space, with the root node representing the unselected state and each branch corresponding to a network activation decision. A priority search queue is generated by pre-calculating the accuracy-delay baseline value of each network at the target density, sorted in descending order of the estimated objective function value. During the traversal, the upper and lower bounds of each node are dynamically calculated. Specifically, the H value of the currently traversed network is summed with the maximum possible H value of the untraversed network (H value is the evaluation index defined in S4), and the maximum value is taken to obtain the estimated upper bound. The upper bound estimation function is:
[0106] (9)
[0107] Where UB is the upper bound of the prediction, Current H is the H value of the selected network, and Potential H is the H value of the remaining networks. The Potential H value is quickly estimated by taking the maximum possible value among the remaining networks for accuracy and the minimum value among the remaining networks for delay. The upper bound is based on the measured value of the current path and the estimate of the maximum theoretical gain of the remaining networks, while the lower bound is the determined value of the actual selected networks. When the upper bound of a branch is lower than the current global optimum, pruning is triggered to eliminate invalid search paths. Finally, by shrinking the solution space layer by layer, the optimal solution that maximizes H is obtained.
[0108] Step 14: Based on the optimal solution identifier output in Step 13, the system calls the corresponding feature extraction network, trajectory tracking network, and semantic segmentation network from the pre-loaded network library.
[0109] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solutions of the present invention, as long as they do not depart from the spirit and scope of the technical solutions of the present invention, should be covered within the scope of the claims of the present invention.
[0110] This invention provides a multi-model adaptive scheduling method for autonomous driving, comprising the following: 1) a discrete optimization coding technique for network configuration, which transforms network selection into a discrete optimization problem through binary coding, significantly reducing decision complexity; 2) a scene density discretization algorithm based on target perception, which uses the perception results of target detection of the surrounding environment to map continuous scene density to discrete scene density; 3) a quantitative prediction model for network performance and scene density, establishing a functional relationship between accuracy and latency and density and network schemes; 4) a network selection optimization framework with dynamic pruning and parallel evaluation, which uses dynamic weight pruning to eliminate candidate schemes with objective function values lower than the current optimal one, prioritizes the remaining candidate networks based on scene density level, uses a double-buffered parallel evaluation strategy to synchronously calculate the real-time performance indicators of each scheme, and finally selects the comprehensive optimal solution to obtain the optimal network selection scheme.
[0111] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-model adaptive scheduling method for autonomous driving, characterized in that, include: S1 abstracts the autonomous driving system into four components: feature extraction module, trajectory tracking module, semantic segmentation module, and planning and decision-making module. S2 provides multiple network options for feature extraction, trajectory tracking, and semantic segmentation, forming various different network configuration schemes. Each network configuration scheme is identified in the form of a triple. S3 performs discretization encoding on all network configuration schemes; S4 defines the optimization objective, including: We construct optimization indices using weighted interpolation: , Where H is the evaluation index used to evaluate the balance between accuracy and latency, ACC is the accuracy of system operation, Latency is the latency of system operation, and λ is the weighting coefficient used to adjust the priority of accuracy and latency. When the system is running, the H value of each network configuration scheme will be calculated based on the real-time target density using the fitted accuracy and latency functions, and the network configuration scheme with the largest H value will be selected as the optimal configuration. S5 collects density data samples from multiple scenarios to build diverse, high-quality datasets; S6 quantifies the density of scene targets and classifies target density levels; S7 utilizes diverse high-quality datasets and groups these diverse high-quality datasets according to target density levels; S8 was used to test the model accuracy under scene data of different densities and different network configurations. S9, establish the scene target density, network configuration scheme, and accuracy density-network-accuracy fitting function; S10, testing latency under different target densities and network configuration schemes in different scenarios; S11, Establish the scene target density, network configuration scheme, and latency density-network-latency fitting function; S12, the optimal network selection scheme is modeled based on the density-network-accuracy fitting function and the density-network-time delay fitting function, and established as a discrete combinatorial optimization problem; S13, For discrete combinatorial optimization problems, an improved branch and bound method is used to solve them, and the optimal solution identifier is output, including: The upper bound is obtained by summing the H values of the currently traversed network and the maximum possible H value of the untraversed network, and taking the maximum value. The upper bound estimation function is: , Where UB is the upper bound of the prediction, Current H is the H value of the selected network, and Potential H is the H value of the remaining network. The Potential H value is quickly estimated by taking the maximum possible value of the remaining network for accuracy and the minimum value of the remaining network for delay. The upper bound is estimated based on the measured value of the current path and the maximum theoretical gain of the remaining network, while the lower bound is the determined value of the actual selected network. When the upper bound of a branch is lower than the current global optimal solution, pruning is triggered to eliminate invalid search paths. Finally, by shrinking the solution space layer by layer, the optimal solution that maximizes H is obtained. S14: Based on the optimal solution identifier, call the corresponding feature extraction network, trajectory tracking network, and semantic segmentation network from the pre-loaded network library.
2. The multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, In S1: The feature extraction module is used to extract key feature information from the raw sensor data; The trajectory tracking module is used to identify and track dynamic targets in a scene using the results of feature extraction, and to predict their future trajectory. The semantic segmentation module is used to classify scene pixels; The planning and decision-making module is used to integrate the results of trajectory tracking and semantic segmentation to formulate vehicle driving strategies.
3. The multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, S2 includes: It provides multiple backbone networks from ResNet34, ResNet50, ResNet101, and ResNet152 for feature extraction; Two networks are provided for trajectory tracking: DenseTrackNet and SparseTrackNet; Two networks are provided for semantic segmentation: UNet and ENet; These networks can be combined independently to form a variety of different network configuration schemes, each of which is identified by a triple. S3 includes: One-hot encoding quantifies the network configuration scheme, ensuring that the network selection for each component is unique and mutually exclusive; feature extraction uses a 4-bit vector, trajectory tracking uses a 2-bit vector, semantic segmentation uses a 2-bit vector, and finally concatenates them into an 8-dimensional 0-1 vector.
4. The multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, S5 includes: In the simulation environment Carla, different target density scenarios are actively constructed, and data samples are generated by adjusting the target density of vehicles and pedestrians. Real scene data with similar density to the Carla constructed sample are inserted into different target density scenarios. Finally, data samples with different target densities will be cleaned to remove erroneous samples.
5. The multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, S6 includes: The pre-trained lightweight target detection model YOLOV8 is used to analyze the images input from the sensor, detect and count the number and categories of targets in the scene, assign different weights W to different target categories, and divide the scene into multiple target density levels based on the number of targets.
6. The multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, S7 includes: The multi-scene density data samples collected by S5 are used, and the data of different target densities are grouped according to the multiple target density levels divided by S6. S8 includes: for each set of data samples, executing multiple network configuration schemes defined in S2, testing the comprehensive accuracy of each network configuration scheme, defining it as the accuracy rate of planning decisions, and finally obtaining multiple accuracy results.
7. The multi-model adaptive scheduling method for autonomous driving according to claim 3, characterized in that, S9 includes: The scene target density is treated as a discrete value, and an 8-dimensional 0-1 vector defined by S3 is sampled to represent each network configuration scheme. The input data of the density-network-accuracy fitting function is all triples [density, network encoding, accuracy]. The function is fitted by a regression model to obtain an accuracy function that can predict the system accuracy based on a given target density and network configuration scheme.
8. The multi-model adaptive scheduling method for autonomous driving according to claim 3, characterized in that, S11 includes: The scene target density is treated as a discrete value, and an 8-dimensional 0-1 vector defined by S3 is sampled to represent each network configuration scheme. The input data of the density-network-delay fitting function is all triples [density, network encoding, delay], which are fitted by a regression model.
9. A multi-model adaptive scheduling method for autonomous driving according to claim 4, characterized in that, S12 includes: By maximizing the optimization objective defined in S4 Choose the optimal network configuration and model the task as a constrained discrete combinatorial optimization problem: , Where max represents taking the maximum value, f acc f represents the precision function of the target density D and the network configuration scheme O. Latency The time delay function represents the target density D and the network configuration scheme O; Constraints: The value of O is a 0-1 code representing the selection results of eight network configuration schemes, where 0 represents not selecting the scheme and 1 represents selecting the scheme. The sum of the first to fourth selection results is 1, the sum of the fifth to sixth selection results is 1, and the sum of the seventh to eighth selection results is 1. , The discrete combinatorial optimization problem was ultimately modeled as a 0-1 integer programming problem, a single-objective combinatorial optimization problem with a finite feasible solution space.
10. A multi-model adaptive scheduling method for autonomous driving according to claim 1, characterized in that, S13 includes: For this discrete combinatorial optimization, an improved branch and bound method is used to solve the problem. All network schemes are constructed as a tree-like decision space, with the root node representing the unselected state and each branch corresponding to a network configuration scheme. By pre-calculating the accuracy-delay benchmark values of various network configurations under the target density, a priority search queue is generated in descending order of the estimated objective function values. During the traversal, the upper and lower bounds of each node are dynamically calculated.
Citation Information
Patent Citations
Multi-scene adaptive pedestrian trajectory prediction method based on attention mechanism
CN119579642A
Neural network driven vehicle adaptive cruise control method
CN120096565A