A Vehicle-Mountain Spatiotemporal Joint Decision-Making Method and System Based on a Distillation-Based Lightweight Large Model
By constructing lightweight models and 4D raster models for path planning and decision-making, the bottleneck of intelligent driving model deployment and the problem of excessive latency have been solved, realizing an in-vehicle decision-making system with low computing power requirements and high safety.
Patent Information
- Application Number
- CN202511223582.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing large-scale intelligent driving models suffer from bottlenecks in model deployment, excessively high decision-making latency, and insufficient safety redundancy, failing to meet the computing power requirements of automotive-grade chips and the real-time requirements for emergency obstacle avoidance.
A vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model is adopted. The method involves constructing a lightweight model for sensor data processing, combining it with a 4D raster model for path planning and spatiotemporal joint decision-making, and using a safety kernel to verify the vehicle body control parameters.
It reduces vehicle computing power requirements, decreases decision latency, improves security and real-time decision-making, and provides dual-core security verification functionality.
Smart Images

Figure CN120742695B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent driving technology, specifically to an onboard spatiotemporal joint decision-making method and system based on a distillation-based lightweight large model. Background Technology
[0002] The existing large-scale intelligent driving models have the following pain points: (1) There is a bottleneck in model deployment. Traditional large-scale models (>10B parameters) require 200+ TOPS of computing power, which exceeds the carrying capacity of automotive-grade chips (Orin-X only has 254 TOPS); (2) The decision-making delay is too high. Traditional planning algorithms are executed step by step (perception → prediction → planning), with a cumulative delay of >300ms, which cannot meet the emergency obstacle avoidance requirements (<100ms); (3) Insufficient safety redundancy. Single-chip solutions lack ASIL-D level real-time monitoring, and hardware failures can lead to system crashes. Summary of the Invention
[0003] In view of this, the purpose of this invention is to provide a vehicle-mounted spatiotemporal joint decision-making method and system based on a distillation-based lightweight large model, so as to solve the above-mentioned technical problems.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] The present invention provides a vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model, comprising the following steps:
[0006] Acquire sensor fusion data of the current scene and the vehicle's current location;
[0007] The sensor fusion data is input into a pre-built lightweight model to obtain inference results for one or more targets. The lightweight model is obtained by distilling a local student model using a cloud-based teacher model. The inference results include target labels, target cue information, target location information, and motion vector information.
[0008] A 4D raster model of the current scene is constructed based on the reasoning results of one or more targets. In the 4D raster model, path planning is performed based on the reasoning results and the current position of the vehicle to obtain multiple traffic routes. The 4D raster model includes a three-dimensional spatial dimension and a time dimension.
[0009] A spatiotemporal joint decision is made on multiple routes to obtain the optimal path; the vehicle control parameters corresponding to the optimal path are sent to the safety kernel and compared with the preset vehicle control threshold to obtain the comparison result; and the vehicle control decision is executed based on the comparison result.
[0010] In one embodiment of this application, the method for constructing the lightweight model includes:
[0011] Obtain teacher models from the cloud;
[0012] A loss function that incorporates time-series execution is constructed and used in conjunction with the student model deployed on the vehicle by the teacher model for guided training until the training completion condition is met, resulting in a lightweight model. The training completion condition includes either the total loss no longer changing or the number of training iterations reaching a preset threshold. The mathematical expression of the loss function is:
[0013]
[0014] In the formula, For the total loss, For time-series loss weighting coefficients, express Divergence loss, This represents the feature map of the student model at time t. This represents the feature map of the teacher model at time t.
[0015] In one embodiment of this application, constructing a 4D rasterized model of the current scene based on the inference results of one or more targets includes:
[0016] Obtain the spatial and temporal parameters of the scene, wherein the temporal parameters are determined by the target prompt information;
[0017] Based on the spatial parameters, a three-dimensional model is obtained; the three-dimensional model is rasterized to obtain multiple spatial grids; and the time parameters are sliced to obtain multiple time layers.
[0018] The vehicle's current position is mapped to a corresponding spatial grid; and based on the target's motion vector information and position information, the target is mapped to spatial grids corresponding to multiple time layers.
[0019] The spatial grids corresponding to the multiple time layers where the target is located are marked as risk grids, resulting in a 4D rasterized model.
[0020] In one embodiment of this application, path planning is performed in the 4D rasterized model based on the inference result and the current position of the vehicle to obtain multiple travel routes, including:
[0021] With the objectives of traffic efficiency, traffic safety, and traffic comfort respectively, a path search is performed starting from the current grid where the vehicle is currently located to obtain the next node grid. Priority rating ;
[0022] Among them, priority scoring based on traffic efficiency. The mathematical expression is:
[0023]
[0024] Priority scoring for traffic safety purposes The mathematical expression is:
[0025]
[0026] Priority rating based on travel comfort The mathematical expression is:
[0027]
[0028] In the formula, As the weight of the first safety item, As the weight of the second safety item, As the weight of the third safety item, For adjustment coefficients, The minimum distance to the nearest obstacle of the vehicle. The weight of the first comfort item, As the weight of the second comfort item, As the third comfort item, The current rate of change of vehicle acceleration. To move from the current grid to the next node grid The corresponding maximum rate of change of acceleration, As the weight of the first efficiency term, As the weight of the second efficiency term, As the weight of the third efficiency term, For vehicle speed, For speed limits; among them, , , ;
[0029] The next node grid with the highest priority score The target grid is used as the current grid, and the path search is performed starting from the current grid where the vehicle is currently located, until multiple routes are obtained through the current scene.
[0030] In one embodiment of this application, a spatiotemporal joint decision is made on multiple travel routes to obtain the optimal path, including:
[0031] Statistics on the total travel time of multiple routes Acceleration through each grid Maximum curvature The minimum distance to the nearest obstacle when passing through each grid. ,in, Indicates the grid number;
[0032] Calculate the variance of multiple accelerations for multiple traffic routes. Multiple minimum distances average ;
[0033] Overall travel time for multiple routes Variance of acceleration Maximum curvature ,average value Normalization was performed separately to obtain the normalized travel time. Normalized variance Normalized maximum curvature and normalized mean ;
[0034] Based on the normalized travel time The normalized variance The normalized maximum curvature and the normalized average value Calculate the overall evaluation value for each route. The comprehensive evaluation value The mathematical expression is:
[0035]
[0036] In the formula, As the first weight, As the second weight, It is the third weight;
[0037] It will meet the time constraints and the comprehensive evaluation value The highest-traffic route is taken as the optimal path, where the time constraint is determined by the time dimension of the 4D rasterized model.
[0038] In one embodiment of this application, the vehicle body control parameters corresponding to the optimal path include acceleration and torque.
[0039] In one embodiment of this application, performing a vehicle body control decision based on the comparison result includes:
[0040] When the vehicle control parameters corresponding to the optimal path do not exceed the safety kernel and the preset vehicle control threshold, vehicle control is performed based on the optimal path; when the vehicle control parameters corresponding to the optimal path exceed the safety kernel and the preset vehicle control threshold, a safe parking protocol is initiated.
[0041] This application also provides an in-vehicle spatiotemporal joint decision-making system based on a distillation-based lightweight large model, including:
[0042] The acquisition module is used to acquire sensor fusion data of the current scene and the current position of the vehicle;
[0043] The inference module is used to input the sensor fusion data into a pre-built lightweight model to obtain inference results for one or more targets. The lightweight model is obtained by distillation training of a local student model by a cloud-based teacher model. The inference results include target labels, target prompt information, target location information, and motion vector information.
[0044] The planning module is used to construct a 4D raster model of the current scene based on the reasoning results of one or more targets. In the 4D raster model, path planning is performed based on the reasoning results and the current position of the vehicle to obtain multiple traffic routes. The 4D raster model includes a three-dimensional spatial dimension and a time dimension.
[0045] The decision module is used to perform spatiotemporal joint decision-making on multiple routes to obtain the optimal path; and send the vehicle control parameters corresponding to the optimal path to the safety kernel for comparison with the preset vehicle control threshold to obtain the comparison result; and execute the vehicle control decision based on the comparison result.
[0046] This application also provides an electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method described above.
[0047] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.
[0048] The beneficial effects of this invention are as follows: The in-vehicle spatiotemporal joint decision-making method and system based on a distilled lightweight large model reduces the computational power required by the vehicle by compressing the large in-vehicle model through knowledge distillation. Utilizing the lightweight large model, this application performs scene recognition and target segmentation on multi-sensor fusion data, and combines this with a 4D rasterized scene for route search, obtaining multiple possible routes. Then, based on spatiotemporal joint decision-making, a decision is made among these multiple routes to obtain the optimal route. Furthermore, the vehicle control parameters resulting from the optimal route are compared with the parameter thresholds within the safety kernel, providing a dual-core safety verification function. This application has advantages such as low computational power requirement, low decision latency, and high security. Attached Figure Description
[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0050] Figure 1This is a flowchart illustrating an embodiment of the in-vehicle spatiotemporal joint decision-making method based on a distillation lightweight large model.
[0051] Figure 2 This is a hardware structure diagram of one embodiment of this application;
[0052] Figure 3 This is a system architecture diagram according to one embodiment of this application;
[0053] Figure 4 This is a structural diagram of an in-vehicle spatiotemporal joint decision-making system based on a distillation lightweight large model, as shown in one embodiment of this application. Detailed Implementation
[0054] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0055] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the layers related to the present invention and are not drawn according to the actual number, shape and size ratio of the layers in the actual implementation. In the actual implementation, the form and number of each layer can be arbitrarily changed, and the layer layout may also be more complex.
[0056] Numerous details are explored in the following description to provide a more thorough explanation of embodiments of the invention; however, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details.
[0057] Figure 1 This is a flowchart illustrating an embodiment of the in-vehicle spatiotemporal joint decision-making method based on a distillation lightweight large model, as shown in this application. Figure 1 As shown: The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model in this embodiment may include steps S110 to S140:
[0058] S110, acquire sensor fusion data of the current scene and the current position of the vehicle;
[0059] In this application, the sensor fusion data includes images captured by a camera and echo data collected by radar. The data is fused using a data-level fusion method to serve as input for a lightweight model.
[0060] The vehicle's current location is location data.
[0061] S120, the sensor fusion data is input into a pre-built lightweight model to obtain inference results for one or more targets, wherein the lightweight model is obtained by distillation training of a local student model by a cloud teacher model, and the inference results include target labels, target prompt information, target position information and motion vector information;
[0062] In this application, the cloud model is pre-deployed in the vehicle using distillation learning, specifically including:
[0063] (1) Obtain the teacher model in the cloud; use the original UniAD model (4B parameters) as the teacher network.
[0064] (2) Construct a loss function that incorporates time-series execution, and combine it with the student model deployed on the vehicle by the teacher model for guided training until the training completion condition is met, thus obtaining a lightweight model. The training completion condition includes either the total loss no longer changing or the number of training iterations reaching a preset threshold. The mathematical expression of the loss function is:
[0065]
[0066] In the formula, For the total loss, For time-series loss weighting coefficients, express Divergence loss, This represents the feature map of the student model at time t. This represents the feature map of the teacher model at time t.
[0067] in, The value of the divergence loss is determined based on the difference in probability distribution of the single-frame output of the teacher-student model. Introducing the divergence loss can constrain the similarity of spatial features and ensure the accuracy of static scene perception.
[0068] The value of is determined by grid search, with an optimal value of 0.7. Its function is to balance the spatial / temporal loss. After testing, A value greater than 0.7 will cause overfitting of the motion trajectory. A value less than 0.7 will reduce temporal consistency.
[0069] Conv-Transformer hybrid features from the original UniAD model to provide a standard for spatiotemporal joint knowledge;
[0070] It comes from the output of a lightweight CNN and is a compressed representation that mimics the characteristics of a teacher.
[0071] In this application, the compressed and lightweight model can output multi-target detection boxes, relative position data with respect to the vehicle, and motion vector information based on sensor fusion data. Target position information can be obtained by performing coordinate transformation between the relative position data with respect to the vehicle and the vehicle's current position.
[0072] In addition, for identifiable targets such as signs and traffic lights, the device can also extract directional information.
[0073] S130, construct a 4D raster model of the current scene based on the reasoning results of one or more targets, perform path planning in the 4D raster model based on the reasoning results and the current position of the vehicle to obtain multiple traffic routes, wherein the 4D raster model includes a three-dimensional spatial dimension and a time dimension;
[0074] This application discretizes the 3D spatial and temporal dimensions into a 4D grid (resolution 0.1m×0.1m×0.1m×0.2s), and then performs path planning and decision-making within the 4D grid.
[0075] Methods for constructing 4D rasterized models include:
[0076] S1301, Obtain the spatial and temporal parameters of the scene;
[0077] For example, in the current scenario of a traffic light intersection, in addition to identifying multiple dynamic targets (pedestrians and vehicles), the traffic light is also identified. For example, in terms of spatial parameters, the intersection size is 40m × 30m; in terms of temporal parameters, the green light has 5.0s remaining.
[0078] S1302, perform three-dimensional modeling based on the spatial parameters to obtain a three-dimensional model; rasterize the three-dimensional model to obtain multiple spatial grids; and slice the time parameters to obtain multiple time layers;
[0079] This application can be based on an existing 3D map or can use real-time rendering for modeling; no limitation is imposed here. The resolution of the rasterized 3D model is 0.1m × 0.1m, generating a 400 × 300 grid when the intersection size is 40m × 30m. Time slices: 0.2s / grid, which can generate 25 time slices when there are 5.0s remaining on the green light.
[0080] S1303, map the current position of the vehicle to the corresponding spatial grid; and map the target to spatial grids corresponding to multiple time layers based on the target's motion vector information and the target's position information;
[0081] S1304 marks the spatial grids corresponding to multiple time layers where the target is located as risk grids, thus obtaining a 4D rasterized model.
[0082] Finally, the target's running information and location information are used to perform location mapping and time mapping. For example, pedestrian speed: 1.5m / s, risk zone marker: OBJ2 covers the grid [201:205, 80:85] at t+0.8s.
[0083] In addition, the vehicle's position and motion information need to be mapped into the 3D model. For example: current vehicle speed: 30km / h (8.3m / s), initial pose: grid (50,10)@t=0s. Current vehicle maximum steering angle: ±35°, motion constraint curvature radius ≥6m.
[0084] After constructing the 4D raster model, this application performs path planning based on the ideas of path node expansion and reward calculation. The specific process is as follows:
[0085] S1311, with the objectives of traffic efficiency, traffic safety, and traffic comfort respectively, performs a path search starting from the current grid where the vehicle is currently located, to obtain the next node grid of the current grid. Priority rating ;
[0086] In this application, starting from the current grid, the next node grid in the next time layer is found. The next node grid can be any grid within a circle with the current vehicle position as the center and a radius of R. The radius R is determined by the current vehicle speed. The greater the speed, the larger the radius R. However, it is necessary to ensure that the angle between the line connecting the next node grid and the current grid and the current direction of the vehicle is less than a set value, so as to ensure that the planned route moves along the navigation route and avoid large curvature turns that affect comfort or even cause dangerous situations.
[0087] In this application, routes are planned based on three objectives: efficiency priority, safety priority, and comfort priority.
[0088] Among them, priority scoring based on traffic efficiency. The mathematical expression is:
[0089]
[0090] Priority scoring for traffic safety purposes The mathematical expression is:
[0091]
[0092] Priority rating based on travel comfort The mathematical expression is:
[0093]
[0094] In the formula, As the weight of the first safety item, As the weight of the second safety item, As the weight of the third safety item, For adjustment coefficients, The minimum distance to the nearest obstacle of the vehicle. The weight of the first comfort item, As the weight of the second comfort item, As the third comfort item, The current rate of change of vehicle acceleration. To move from the current grid to the next node grid The corresponding maximum rate of change of acceleration, As the weight of the first efficiency term, As the weight of the second efficiency term, As the weight of the third efficiency term, For vehicle speed, For speed limits; among them, , , ;
[0095] In the above formula, when The reward index decays at distances less than 2 meters, thus ensuring a safe distance. Other comfort and efficiency items are evaluated using the rate of change of acceleration and the percentage of speed limit compliance, respectively.
[0096] Exemplarily, in one embodiment of this application, , When efficiency is the priority, the weight is increased to 0.3. , When safety is the priority, the weight is increased to 0.8. , When comfort is the priority, the weight is increased to 0.3.
[0097] S1312, select the next node mesh with the highest priority score. The target grid is used as the current grid, and the path search is performed starting from the current grid where the vehicle is currently located, until multiple routes are obtained through the current scene.
[0098] In this application, a target grid is found at each evaluation angle and at each time level (the risk grid is different at each time level). After the above process is repeated cyclically, the travel routes under the three angles of efficiency priority, safety priority, and comfort priority can be found respectively.
[0099] S140, perform spatiotemporal joint decision-making on multiple routes to obtain the optimal path; send the vehicle control parameters corresponding to the optimal path to the safety kernel and compare them with the preset vehicle control threshold to obtain the comparison result; and execute the vehicle control decision based on the comparison result.
[0100] Finally, in order to find the optimal route, this application also conducts a comprehensive evaluation to obtain the optimal path, the specific process of which is as follows:
[0101] S1401, Calculate the total travel time for multiple routes. Acceleration along each grid Maximum curvature The minimum distance to the nearest obstacle when passing through each grid. ,in, Indicates the grid number;
[0102] Among them, the total travel time The number of time layers traversed to reach the target location determines the passage efficiency evaluation item. Acceleration of each grid cell. Maximum curvature This is calculated by the vehicle's infotainment chip and is used for evaluating driving comfort. Minimum distance. This is a safety evaluation item.
[0103] S1402, Calculate the variance of multiple accelerations for multiple traffic routes. Multiple minimum distances average ;
[0104] S1403, overall travel time for multiple routes. Variance of acceleration Maximum curvature ,average value Normalization was performed separately to obtain the normalized travel time. Normalized variance Normalized maximum curvature and normalized mean ;
[0105] S1404, based on the normalized travel time The normalized variance The normalized maximum curvature and the normalized average value Calculate the overall evaluation value for each route. To facilitate comprehensive evaluation using standardized metrics, this application employs the min-max normalization method to normalize the same type of evaluation items for different routes, and then performs weighted summation to obtain the following mathematical expression:
[0106]
[0107] In the formula, As the first weight, As the second weight, It is the third weight;
[0108] For example, the order is safety first, comfort second, and efficiency lowest, with the following weights: , , .
[0109] For example, the table below shows a comparison of the overall evaluation of the three routes:
[0110] Table 1. Comprehensive Evaluation of Traffic Routes
[0111]
[0112] S1405 will meet the time constraint and the comprehensive evaluation value The highest-traffic route is taken as the optimal path, where the time constraint is determined by the time dimension of the 4D rasterized model.
[0113] As shown in the table above, Option 2 (safety first) has the highest score, but it will cause the timeout and prevent passage through the traffic lights. Therefore, Option 3 (comfort first) is selected as the final passage option.
[0114] Furthermore, the acceleration and torque of the vehicle body must not be too large when making decisions to ensure safety. To ensure safe passage, vehicle body control decisions also need to be made based on the aforementioned comparison results, including:
[0115] When the vehicle control parameters corresponding to the optimal path do not exceed the safety kernel and the preset vehicle control threshold, vehicle control is performed based on the optimal path; when the vehicle control parameters corresponding to the optimal path exceed the safety kernel and the preset vehicle control threshold, a safe parking protocol is initiated.
[0116] An exemplary process includes:
[0117] Step 1: Orin (intelligent driving chip) outputs decision instructions to TC397 (security core).
[0118] Step 2: TC397 safety check to confirm whether the torque value exceeds the limit (<300Nm).
[0119] Step 3: If the torque exceeds the limit, activate the safety stop protocol; otherwise, execute it.
[0120] Figure 2 This is a hardware structure diagram of one embodiment of this application, such as... Figure 2 As shown, based on the hardware structure, this application first uses the GMSL2 interface to input the sensor fusion data from the camera / radar into the Orin intelligent driving chip, which then inputs the data into the lightweight model for inference and decision-making. The decision-making circuitry is input to the TC397 security core for verification. If the verification passes, further execution is performed. If the TC397 security core is unavailable, the data is sent to a cloud-based diagnostic platform for verification.
[0121] Figure 3 This is a system architecture diagram of one embodiment of this application, such as... Figure 3 As shown, this application compresses the original 5B parameter model to 500M by distilling a lightweight model, reducing computational power dependence. Simultaneously, a spatiotemporal joint decision-making algorithm is employed to improve response speed and decision rationality. Finally, a dual-core lockstep verification is implemented using a safety redundancy controller to ensure security.
[0122] This invention discloses a vehicle-mounted spatiotemporal joint decision-making method based on a distilled lightweight large model. This application reduces the computational power required by the vehicle by compressing the large vehicle model through knowledge distillation. Utilizing the lightweight large model, this application performs scene recognition and target segmentation on multi-sensor fusion data, and combines this with a 4D rasterized scene for route search, obtaining multiple possible routes. Then, based on spatiotemporal joint decision-making, a decision is made among these multiple routes to obtain the optimal route. Furthermore, the vehicle control parameters resulting from the optimal route are compared with the parameter thresholds within the safety kernel, providing a dual-core safety verification function. This application has advantages such as low computational power requirement, low decision latency, and high security.
[0123] like Figure 4 As shown, this application also provides an in-vehicle spatiotemporal joint decision-making system based on a distillation-based lightweight large model, including:
[0124] The acquisition module is used to acquire sensor fusion data of the current scene and the current position of the vehicle;
[0125] The inference module is used to input the sensor fusion data into a pre-built lightweight model to obtain inference results for one or more targets. The lightweight model is obtained by distillation training of a local student model by a cloud-based teacher model. The inference results include target labels, target prompt information, target location information, and motion vector information.
[0126] The planning module is used to construct a 4D raster model of the current scene based on the reasoning results of one or more targets. In the 4D raster model, path planning is performed based on the reasoning results and the current position of the vehicle to obtain multiple traffic routes. The 4D raster model includes a three-dimensional spatial dimension and a time dimension.
[0127] The decision module is used to perform spatiotemporal joint decision-making on multiple routes to obtain the optimal path; and send the vehicle control parameters corresponding to the optimal path to the safety kernel for comparison with the preset vehicle control threshold to obtain the comparison result; and execute the vehicle control decision based on the comparison result.
[0128] This invention relates to a vehicle-mounted spatiotemporal joint decision-making system based on a distilled lightweight large-scale model. This application reduces the computational power required by the vehicle by compressing the large-scale vehicle model through knowledge distillation. Utilizing the lightweight large-scale model, this application performs scene recognition and target segmentation on multi-sensor fusion data, and combines this with a 4D rasterized scene for route search, obtaining multiple possible routes. Then, based on spatiotemporal joint decision-making, a decision is made among these multiple routes to obtain the optimal route. Furthermore, the vehicle control parameters resulting from the optimal route are compared with the parameter thresholds within the safety kernel, providing a dual-core safety verification function. This application has advantages such as low computational power requirement, low decision latency, and high security.
[0129] This embodiment also provides an electronic device, including: a processor and a memory;
[0130] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory so that the terminal performs any of the methods in this embodiment.
[0131] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0132] The electronic device provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic device performs the various steps of the above method.
[0133] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0134] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0135] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. The embodiments of the invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims.
[0136] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model, characterized in that, Including the following steps: Acquire sensor fusion data of the current scene and the vehicle's current location; The sensor fusion data is input into a pre-built lightweight model to obtain inference results for one or more targets. The lightweight model is obtained by distilling a local student model using a cloud-based teacher model. The inference results include target labels, target cue information, target location information, and motion vector information. A 4D rasterized model of the current scene is constructed based on the inference results of one or more targets. In this 4D rasterized model, path planning is performed based on the inference results and the vehicle's current position to obtain multiple travel routes. The 4D rasterized model includes a three-dimensional spatial dimension and a temporal dimension. The construction of the 4D rasterized model based on the inference results of one or more targets includes: obtaining spatial and temporal parameters of the scene, wherein the temporal parameters are determined by target prompt information; performing three-dimensional modeling based on the spatial parameters to obtain a three-dimensional model; rasterizing the three-dimensional model to obtain multiple spatial grids; slicing the temporal parameters to obtain multiple temporal layers; mapping the vehicle's current position to the corresponding spatial grid; mapping the target to the spatial grids corresponding to the multiple temporal layers based on the target's motion vector information and the target's position information; and marking the spatial grids corresponding to the multiple temporal layers where the target is located as risk grids to obtain the 4D rasterized model. A spatiotemporal joint decision is made on multiple routes to obtain the optimal path; the vehicle control parameters corresponding to the optimal path are sent to the safety kernel and compared with the preset vehicle control threshold to obtain the comparison result; and the vehicle control decision is executed based on the comparison result.
2. The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model according to claim 1, characterized in that, The method for constructing the lightweight model includes: Obtain teacher models from the cloud; A loss function that incorporates time-series execution is constructed and used in conjunction with the student model deployed on the vehicle by the teacher model for guided training until the training completion condition is met, resulting in a lightweight model. The training completion condition includes either the total loss no longer changing or the number of training iterations reaching a preset threshold. The mathematical expression of the loss function is: In the formula, For the total loss, For time-series loss weighting coefficients, express Divergence loss, This represents the feature map of the student model at time t. This represents the feature map of the teacher model at time t.
3. The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model according to claim 1, characterized in that, Based on the inference results and the vehicle's current position, path planning is performed in the 4D rasterized model to obtain multiple travel routes, including: With the objectives of traffic efficiency, traffic safety, and traffic comfort respectively, a path search is performed starting from the current grid where the vehicle is currently located to obtain the next node grid. Priority rating ; Among them, priority scoring based on traffic efficiency. The mathematical expression is: Priority scoring for traffic safety purposes The mathematical expression is: Priority rating based on travel comfort The mathematical expression is: In the formula, As the weight of the first safety item, As the weight of the second safety item, As the weight of the third safety item, For adjustment coefficients, The minimum distance to the nearest obstacle of the vehicle. The first comfort item weight, As the weight of the second comfort item, As the third comfort item, The current rate of change of vehicle acceleration. To move from the current grid to the next node grid The corresponding maximum rate of change of acceleration, As the weight of the first efficiency term, As the weight of the second efficiency term, As the weight of the third efficiency term, For vehicle speed, For speed limits; among them, , , ; The next node grid with the highest priority score The target grid is used as the current grid, and the path search is performed starting from the current grid where the vehicle is currently located, until multiple routes are obtained through the current scene.
4. The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model according to claim 1, characterized in that, A spatiotemporal joint decision-making process is performed on multiple travel routes to obtain the optimal path, including: Statistics on the total travel time of multiple routes Acceleration along each grid Maximum curvature The minimum distance to the nearest obstacle when passing through each grid. ,in, Indicates the grid number; Calculate the variance of multiple accelerations for multiple traffic routes. Multiple minimum distances average ; Overall travel time for multiple routes Variance of acceleration Maximum curvature ,average value Normalization was performed separately to obtain the normalized travel time. Normalized variance Normalized maximum curvature and normalized mean ; Based on the normalized travel time The normalized variance The normalized maximum curvature and the normalized average value Calculate the overall evaluation value for each route. The comprehensive evaluation value The mathematical expression is: In the formula, As the first weight, As the second weight, It is the third weight; It will meet the time constraints and the comprehensive evaluation value The highest-traffic route is taken as the optimal path, where the time constraint is determined by the time dimension of the 4D rasterized model.
5. The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model according to claim 1, characterized in that, The vehicle control parameters corresponding to the optimal path include acceleration and torque.
6. The vehicle-mounted spatiotemporal joint decision-making method based on a distillation-based lightweight large model according to claim 1, characterized in that, Based on the comparison results, a vehicle body control decision is made, including: When the vehicle control parameters corresponding to the optimal path do not exceed the safety kernel and the preset vehicle control threshold, vehicle control is performed based on the optimal path; when the vehicle control parameters corresponding to the optimal path exceed the safety kernel and the preset vehicle control threshold, a safe parking protocol is initiated.
7. A vehicle-mounted spatiotemporal joint decision-making system based on a lightweight distillation model, characterized in that: include: The acquisition module is used to acquire sensor fusion data of the current scene and the current position of the vehicle; The inference module is used to input the sensor fusion data into a pre-built lightweight model to obtain inference results for one or more targets. The lightweight model is obtained by distillation training of a local student model by a cloud-based teacher model. The inference results include target labels, target prompt information, target location information, and motion vector information. The planning module is used to construct a 4D rasterized model of the current scene based on the inference results of one or more targets. In the 4D rasterized model, path planning is performed based on the inference results and the current position of the vehicle to obtain multiple traffic routes. The 4D rasterized model includes a three-dimensional spatial dimension and a temporal dimension. Constructing the 4D rasterized model of the current scene based on the inference results of one or more targets includes: obtaining spatial and temporal parameters of the scene, wherein the temporal parameters are determined by target prompt information; performing three-dimensional modeling based on the spatial parameters to obtain a three-dimensional model; rasterizing the three-dimensional model to obtain multiple spatial grids; slicing the temporal parameters to obtain multiple temporal layers; mapping the current position of the vehicle to the corresponding spatial grid; mapping the target to the spatial grids corresponding to the multiple temporal layers based on the target's motion vector information and the target's position information; and marking the spatial grids corresponding to the multiple temporal layers where the target is located as risk grids to obtain the 4D rasterized model. The decision module is used to perform spatiotemporal joint decision-making on multiple routes to obtain the optimal path; and send the vehicle control parameters corresponding to the optimal path to the safety kernel for comparison with the preset vehicle control threshold to obtain the comparison result; and execute the vehicle control decision based on the comparison result.
8. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Automobile auxiliary driving monocular distance measurement method and device based on lightweight occupancy prediction network
CN119399595A
Visual information-based four-dimensional occupation grid panoramic sensing method and device
CN119810367A