Intelligent driving system data enhancement method based on spatio-temporal joint hierarchical prior knowledge

CN122414271BActive Publication Date: 2026-09-04UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610899860.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-04
Estimated Expiration
2046-06-22

AI Technical Summary

Technical Problem

如果不能有效利用这些时空联合分层的先验知识,数据增强的效果将大打折扣,无法满足智能驾驶系统对数据多样性和准确性的严格要求

Benefits of technology

本发明有效解决了智能驾驶系统传统机器学习算法中先验知识利用不充分、分层不明确的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414271B_ABST
    Figure CN122414271B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent driving and discloses a smart driving system data enhancement method based on spatiotemporal joint layered prior knowledge, which comprises the following steps: separating prior knowledge on scientific knowledge and task categories; generating synthetic perception data containing diversified driving scenes through a diffusion model integrated with a spatiotemporal attention mechanism; guiding a time series generation model to predict vehicle behavior data based on algebraic equations and differential equations of kinematic models and dynamic models; generating control data according to the logical rules of vehicle control strategies; and realizing end-to-end training of a model in a dynamic driving scene based on original perception data, synthetic perception data, predicted vehicle behavior data and control data. The application integrates layered prior knowledge into the training process adaptively, alleviates the problem of insufficient training data and effectively enhances the performance and generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent driving technology, specifically to a data augmentation method for intelligent driving systems based on spatiotemporal joint hierarchical prior knowledge. Background Technology

[0002] With the rapid development of intelligent driving technology, its safety and reliability have become critical issues. Intelligent driving systems rely on large amounts of data for training and decision-making, and the quality and diversity of this data have a significant impact on system performance. However, in practical applications, obtaining sufficient and high-quality data faces numerous challenges.

[0003] On the one hand, real-world road scenarios are complex and varied, encompassing diverse weather conditions, lighting conditions, and traffic situations. Collecting comprehensive data is not only costly but also extremely difficult in practice. On the other hand, existing data collection methods often fail to cover all possible scenarios, resulting in data limitations. This may cause intelligent driving systems to misjudge or fail to make accurate decisions in rare or special situations.

[0004] Traditional data augmentation methods can expand the amount of data to some extent, but they have significant shortcomings for the field of autonomous driving. Most of these methods simply perform geometric transformations and color jitter on images or data, without fully considering the spatiotemporal characteristics of autonomous driving data and the prior knowledge it contains. Data in autonomous driving has strong spatiotemporal correlations; the vehicle's state at different times and the changes in the surrounding environment are continuous, and certain prior rules and patterns exist in specific scenarios. If these spatiotemporally and hierarchically layered prior knowledge cannot be effectively utilized, the effect of data augmentation will be greatly reduced, failing to meet the stringent requirements of autonomous driving systems for data diversity and accuracy.

[0005] Furthermore, traditional machine learning algorithms for intelligent driving systems suffer from insufficient utilization of prior knowledge and unclear hierarchical structures. Therefore, researching data augmentation methods for intelligent driving systems based on spatiotemporally joint hierarchical prior knowledge has significant practical implications and application value. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides a data augmentation method for intelligent driving systems based on spatiotemporal joint hierarchical prior knowledge.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: A data augmentation method for intelligent driving systems based on spatiotemporal joint hierarchical prior knowledge includes: By using a hierarchical prior knowledge representation method, prior knowledge is separated into scientific knowledge and task categories. Prior knowledge includes algebraic equations, differential equations, physical simulation results, space invariance, and logical rules. Synthetic perception data containing diverse driving scenarios is generated through a diffusion model that integrates spatiotemporal attention mechanisms under physical simulation results and spatial invariance constraints; a time-series generation model is guided to predict vehicle behavior data based on algebraic and differential equations of kinematic and dynamic models; and control data is generated according to the logical rules of vehicle control strategies. Based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data, end-to-end training of driving models in dynamic driving scenarios is achieved.

[0008] In one embodiment, the diffusion model, which uses physical simulation results and an integrated spatiotemporal attention mechanism under spatial invariance constraints, generates synthetic perception data containing diverse driving scenarios, specifically including: By integrating a diffusion model with a spatiotemporal attention mechanism, combining physical simulation results for radar data and the principle of spatial invariance for all perception data, synthetic perception data containing diverse driving scenarios is generated based on the original perception data. The perception data includes images, radar data, and videos containing driving scenarios.

[0009] In one embodiment, the diffusion model integrating a spatiotemporal attention mechanism, combined with physical simulation results for radar data and the principle of spatial invariance for all perception data, generates synthetic perception data containing diverse driving scenarios based on the original perception data, specifically including: The physical simulation results are simulated standard radar data. Based on physical characteristics including radar electromagnetic wave propagation, scattering attenuation, and multipath clutter, standard radar data under different weather conditions, different obstructions, and different reflection characteristics are simulated and generated. The standard radar data is used as the physical constraint in the process of reverse denoising of the diffusion model to generate radar data. The principle of spatial invariance refers to adding spatial constraints during the generation of all perception data in the diffusion model by utilizing the properties of rigid targets, including vehicles, pedestrians, and road facilities, that are invariant in translation, rotation, and topology in driving scenarios. Noise is gradually added to the original perceived data through the forward diffusion process of the diffusion model until the original perceived data becomes completely random noise. By introducing a spatiotemporal attention mechanism in the inverse denoising process of the diffusion model, different levels of attention are given to the perceived data at different spatial locations and time points to capture the correlation between features at different spatiotemporal locations and generate synthetic perceived data containing diverse driving scenarios. The spatiotemporal attention mechanism includes a spatial attention mechanism and a temporal attention mechanism. The spatial attention mechanism is used to focus on important areas in the spatial dimension, and the temporal attention mechanism is used to focus on dynamic changes in the temporal dimension and adaptively adjust the diffusion model's attention to different time points. The perceived data includes images, radar data, and videos containing driving scenarios.

[0010] In one embodiment, when noise is progressively added to the original sensing data, such as images, radar data, or videos, through a forward diffusion process using a diffusion model, the resulting noise is completely random. for: ; ; ; in, For the noisy image, noisy radar data, or noisy video obtained by adding standard Gaussian noise in step T, The noise distribution is a multivariate Gaussian noise distribution with a mean of 0 and a covariance matrix of identity matrix I. Here, represents the cumulative noise scaling factor, indicating the cumulative noise impact from step 1 to step T. For the first The noise scaling factor of the step. Let be the noise factor at step T, and t represent the index at step t.

[0011] In one embodiment, the algebraic and differential equations based on kinematic and dynamic models guide the time-series generation model to predict vehicle behavior data, specifically including: Based on the kinematic model, algebraic equations are established to describe the relationship between the vehicle's position, velocity, and acceleration. Based on the dynamic model, differential equations are established to describe the relationship between the vehicle's forces and motion state. The vehicle's position, velocity, and acceleration determined by the kinematic and dynamic models are organized in the form of a time series and used as prior knowledge for the time series generation model. The time series generation model is trained to learn the changing patterns of vehicle behavior and is used to predict vehicle behavior data.

[0012] In one embodiment, the vehicle behavior data includes the vehicle's trajectory, instantaneous speed, acceleration, heading angle change, driving curvature, relative distance, and relative speed under driving conditions such as straight driving, lane changing, turning, following, yielding, idling, starting and stopping, and acceleration and deceleration.

[0013] In one embodiment, the logical rules of the vehicle control strategy include: speed control rules, steering control rules, and braking control rules.

[0014] In one embodiment, generating control data based on the logical rules of the vehicle control strategy specifically includes: Speed ​​control rules include cruise control logic and adaptive cruise control logic; Cruise control logic: After setting the target speed, the engine throttle opening is automatically adjusted by comparing the actual vehicle speed with the target speed; when the actual vehicle speed is lower than the target speed, the vehicle accelerates; when the actual vehicle speed is higher than the target speed, the vehicle decelerates. Adaptive cruise control logic: If the distance to the vehicle in front is detected to be greater than the safe distance threshold and the current vehicle speed is lower than the set speed, the vehicle will accelerate; if the distance to the vehicle in front is less than the safe distance threshold, the vehicle will decelerate. Steering control rules include lane keeping logic and lane changing logic; The lane-keeping logic includes: if the vehicle deviates to the left from the lane centerline by a distance of x, then a right turn angle is generated. Control commands, It is directly proportional to x; if the vehicle deviates to the right from the lane centerline by a distance of x, then a leftward turning angle is generated. Control commands, It is directly proportional to x; The lane-changing logic includes: when the driver triggers a lane-changing signal or the vehicle's automatic driving system determines that the lane-changing conditions are met, generating steering control data based on the position of the target lane, and turning the steering wheel towards the target lane at a set angular velocity. The steering wheel's rotation angle and duration are determined based on the current relative position and speed of the vehicle and the target lane. The braking control rules include: if braking is required, the braking intensity is gradually increased according to the relationship between vehicle speed and distance, based on the kinematic model, so that the vehicle can stop at the target position.

[0015] In one embodiment, the end-to-end training of the driving model in dynamic driving scenarios based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data specifically includes: By utilizing raw perception data, synthetic perception data, predicted vehicle behavior data, and control data, a diverse range of simulated driving scenarios covering urban roads, highways, intersections, different weather conditions, and varying traffic densities are constructed. Raw perception data serves as the real-world benchmark, supervising the driving model's basic perception feature learning. Synthetic perception data serves as supplementary samples, enhancing the driving model's generalization ability. Predicted vehicle behavior data serves as dynamic spatiotemporal priors, guiding the driving model to learn traffic interaction evolution patterns. Control data serves as supervised ground truth, constraining the accuracy of the driving model's decision-making and control outputs. A multi-task learning framework is built to simultaneously optimize perception, planning, decision-making, and control tasks. By fusing multi-source features using a shared network structure and loss function, end-to-end joint training of the driving model is achieved.

[0016] In one embodiment, the loss function for: ; in, It is the uncertainty estimate for the i-th task. It is the uncertainty weight of the i-th task. It is a regularization term. Let be the loss function for the i-th task.

[0017] Compared with the prior art, the beneficial technical effects of the present invention are: This invention effectively solves the problems of insufficient utilization of prior knowledge and unclear hierarchical structure in traditional machine learning algorithms for intelligent driving systems.

[0018] This invention generates high-quality training data based on prior knowledge of scientific knowledge and task categories, which can increase the training sample size of intelligent driving system models.

[0019] This invention integrates hierarchical prior knowledge and adapts it to the training process of intelligent driving systems, alleviating the problem of insufficient training data and significantly improving the performance and generalization ability of the model.

[0020] Other beneficial effects of the present invention will be explained in detail through the introduction of specific technical features and technical solutions in specific embodiments. Those skilled in the art should be able to understand the beneficial technical effects brought about by these technical features and technical solutions through the introduction of these technical features and technical solutions. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the method flow in an embodiment of the present invention.

[0022] Figure 2 This is a data augmentation technology roadmap in an embodiment of the present invention. Detailed Implementation

[0023] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.

[0024] like Figure 1 As shown, a data augmentation method for intelligent driving systems based on spatiotemporal joint hierarchical prior knowledge in this invention includes the following steps: S1 utilizes a hierarchical prior knowledge representation method to separate prior knowledge in terms of scientific knowledge and task categories. Prior knowledge includes algebraic equations, differential equations, physical simulation results, space invariance, and logical rules. S2 generates synthetic perception data containing diverse driving scenarios through a diffusion model that integrates spatiotemporal attention mechanisms under physical simulation results and spatial invariance constraints. S3, based on algebraic and differential equations of kinematic and dynamic models, guides time-series generation models to predict vehicle behavior data; S4 generates control data based on the logical rules of the vehicle control strategy; S5 enables end-to-end training of driving models in dynamic driving scenarios based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data.

[0025] In one embodiment, step S1, which utilizes a hierarchical prior knowledge representation method to separate prior knowledge in terms of scientific knowledge and task categories, specifically includes: From the perspective of scientific knowledge, fine-grained prior knowledge separation is carried out from algebraic equations, differential equations, simulation results, space invariance, logical rules, knowledge graphs, probabilistic relationships, and human feedback; in terms of task categories, multi-dimensional and highly discriminative prior knowledge separation is carried out from the environmental understanding and perception, motion planning, behavioral decision-making, and control of the intelligent driving system.

[0026] Then, high-quality training data is generated based on prior knowledge of scientific knowledge and task categories, increasing the training sample size of the model.

[0027] In one embodiment, step S2 generates synthetic perception data containing diverse driving scenarios using a diffusion model based on physical simulation results and an integrated spatiotemporal attention mechanism under spatial invariance constraints. Specifically, this includes: By integrating a diffusion model with a spatiotemporal attention mechanism, high-quality synthesis is performed on the original perception data to generate images, radar data, and videos containing diverse driving scenarios as synthetic perception data. Combining physical simulation results and the principle of spatial invariance, a multimodal diffusion model is used to enhance the diversity and accuracy of the perception data. Furthermore, by simulating radar data under different weather conditions, different occlusion and reflection characteristics, the coverage of the training samples is expanded.

[0028] The diffusion model integrating spatiotemporal attention mechanism includes: generating data samples by gradually adding and removing noise. In the forward diffusion process, the model gradually adds noise to the original perceived data, causing it to gradually lose its distinguishing characteristics until the data becomes completely random noise. In the reverse denoising process, a spatiotemporal attention mechanism is introduced to give different attention to images at different spatial locations and time points in order to capture the correlation between features at different spatiotemporal locations.

[0029] The spatiotemporal attention mechanism consists of spatial attention and temporal attention. Spatial attention focuses on important regions in the spatial dimension, weighting features at each location to emphasize key areas and de-emphasize irrelevant parts. Temporal attention focuses on dynamic changes in the temporal dimension, enhancing the accuracy of sequence prediction by learning temporal dependencies. It can adaptively adjust the model's attention to different time points based on factors such as the similarity of frame images.

[0030] The expression for the attention mechanism is: ; Where P is the node feature matrix, , and Both are linear projections of P. More specifically, V can be viewed as a vector representing a single input feature, while Q and K are feature vectors for calculating the weights W(Q, K) in the attention mechanism, obtained from the input features.

[0031] For the output of the i-th node, The formula is: ; in, Let be the output value of the i-th node, and N be the number of output nodes, i.e., the number of categories. (This is achieved through...) The function can convert the output values ​​of multi-class classification into a probability distribution in the range [0, 1] with a sum of 1.

[0032] In the attention mechanism, (Q, K, V) is V multiplied by the corresponding weight according to the degree of attention. That is, the similarity between the current query and all keys is calculated, and this similarity value is passed through the Softmax layer to obtain a set of weights. The value under the attention mechanism is obtained by summing the product of this set of weights and the corresponding value.

[0033] Then, combining the simulation results of physics and the principle of spatial invariance, a multimodal diffusion model is used to enhance the diversity and accuracy of radar data. Furthermore, by simulating radar data under different weather conditions, with varying degrees of obstruction and reflection characteristics, the coverage of the training samples is expanded. The specific operations are as follows: The sensed data (images, radar data, or video) undergoes forward diffusion processing, gradually adding standard Gaussian noise until the sensed data evolves into pure noise. The state data at each step of adding standard Gaussian noise is recorded; where the pure noise... The formula for generating it is as follows: ; ; ; in, For the noisy image, noisy radar data, or noisy video obtained by adding standard Gaussian noise in step T, The noise distribution is a multivariate Gaussian noise distribution with a mean of 0 and a covariance matrix of identity matrix I. Here, represents the cumulative noise scaling factor, indicating the cumulative noise impact from step 1 to step T. For the first The noise scaling factor of the step. Let be the noise factor at step T, and t represent the index at step t.

[0034] In one embodiment, the algebraic and differential equations based on the kinematic and dynamic models in step S3 guide the time-series generation model to predict vehicle behavior data, specifically including: A kinematic model is used to describe the changes in a vehicle's position, velocity, and acceleration over time. When a vehicle is traveling in a straight line, the change in its position x with time t can be represented by kinematic equations. To describe, among which This is the initial position. It is the initial velocity. It's acceleration.

[0035] Based on a dynamic model, various forces acting on the vehicle, such as engine driving force, air resistance, and friction, are considered. Newton's second law is used to establish the relationship between the vehicle's motion and the forces acting on it. Among these, air resistance is a crucial factor during the vehicle's movement. The calculation formula is as follows: ; in, It is the air drag coefficient. A is air density, and A is the vehicle's frontal area. It refers to the vehicle's speed.

[0036] Then, a Long Short-Term Memory (LSTM) network was selected as the time-series generation model. The key motion feature data such as vehicle position, speed, and acceleration determined by the kinematic and dynamic models were organized in the form of a time series and used as the prior knowledge of the LSTM network to train the time-series generation model, learn the changing patterns of vehicle behavior, and use it to predict future vehicle behavior.

[0037] In one embodiment, step S4, which generates control data based on the logical rules of the vehicle control strategy, specifically includes: The logical rules of the vehicle control strategy mainly include: speed control rules, steering control rules, and braking control rules.

[0038] The speed control rules include cruise control logic and adaptive cruise control logic.

[0039] Cruise control logic: After setting the target speed, the control system automatically adjusts the engine throttle opening by comparing the actual vehicle speed with the target speed. When the actual vehicle speed is lower than the target speed, the throttle opening is increased to increase engine speed and driving force, thus accelerating the vehicle; conversely, the throttle opening is decreased to reduce engine speed and driving force, thus decelerating the vehicle.

[0040] Adaptive cruise control logic: If the distance to the vehicle in front is detected to be greater than the safe distance threshold and the current vehicle speed is lower than the set speed, increase power output to accelerate; if the distance to the vehicle in front is less than the safe distance threshold, brake or decelerate.

[0041] Steering control rules include lane keeping logic and lane changing logic.

[0042] Lane keeping logic: If the vehicle deviates left from the lane centerline by a distance of x, then a right turn angle is generated. Control commands, It is directly proportional to x; conversely, a deviation to the right generates a command to turn left.

[0043] Lane change logic: When the driver triggers a lane change signal or the vehicle's automatic driving system determines that the lane change conditions are met, it will generate steering control data based on the position of the target lane and turn the steering wheel toward the target lane at a certain angular velocity. The turning angle and duration are determined based on factors such as the relative position of the vehicle and the target lane and the vehicle speed. The higher the vehicle speed, the smaller the steering angular velocity, to ensure a smooth lane change.

[0044] Braking control rules include emergency braking logic and conventional braking logic. When braking is required, the braking intensity is gradually increased according to the relationship between vehicle speed and distance, following a reasonable kinematic model, so that the vehicle can stop just at the target position, ensuring that the vehicle braking process is safe, smooth, and efficient.

[0045] In one embodiment, step S5, based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data, enables end-to-end training of the driving model in dynamic driving scenarios, specifically including: By utilizing collected and augmented data, diverse simulated driving scenarios are created for training. These scenarios include urban roads, highways, intersections, different weather conditions, and different traffic densities. A multi-task learning framework is designed to enable the model to simultaneously optimize perception, planning, decision-making, and control tasks during the learning process. End-to-end training of the driving model is achieved by sharing the network structure and loss function.

[0046] The driving model used in this embodiment is an end-to-end Transformer autonomous driving decision model. This driving decision model is based on a multi-head self-attention mechanism to build an overall encoding-decoding architecture. It does not require manual design and screening of scene features and rule features. It can directly use the fused data sequence consisting of raw perception data, synthetic perception data, vehicle behavior prediction data and control data as input to the driving decision model. The encoder completes the global correlation modeling of multi-dimensional scene information, and the decoder directly outputs continuous driving decision control commands such as vehicle steering, acceleration and braking. It completes the end-to-end training of the entire link from raw scene data to vehicle driving control commands. It can accurately adapt to complex dynamic driving scenarios such as rain, snow, congestion, lane changing and overtaking, and greatly improve the model's decision accuracy and generalization ability in real road conditions.

[0047] The main steps in constructing a shared network structure are as follows: Defining the shared layer: In multi-task learning, the shared layer is a layer that provides input to multiple tasks simultaneously. Through the shared layer, the model can leverage the correlation between tasks, improve training efficiency and generalization performance, and learn common features of perception, planning, decision-making, and control tasks, including spatiotemporal features of vehicle motion, road structure features, etc.

[0048] Design task-specific output layers: Each task has its own output layer, which is responsible for transforming the input of the shared layers into task-specific outputs. These task-specific output layers address the unique features of perception, planning, decision-making, and control tasks respectively. A specific task layer is added after the shared layers to transform the shared features into specific outputs for each task.

[0049] Parameter sharing: In multi-task learning, parameter sharing refers to multiple tasks sharing a portion of model parameters. This approach can reduce model complexity and overfitting, and improve the model's generalization performance. Parameter sharing can be achieved through hard parameter sharing (all tasks share the same hidden layers) or soft parameter sharing (model parameters for different tasks remain similar within a certain range, but are not completely identical).

[0050] In machine learning, the loss function is a function that measures the difference between the model's predictions and the true values, and it is an indispensable part of optimization algorithms. The goal of the loss function is to improve the model's predictive performance by minimizing its value. More specifically, the loss function provides a way to quantify model error, allowing machine learning algorithms to use this information to adjust model parameters. Multi-task learning involves training multiple related tasks simultaneously, each with its own loss function. Therefore, a comprehensive mechanism is needed to coordinate these different loss functions.

[0051] In multi-task learning, uncertainty-weighted loss is a method that dynamically adjusts the weights of loss functions for different tasks, aiming to enable the model to learn more effectively from multiple related tasks. This method is particularly suitable for situations where the magnitudes of loss values ​​may differ significantly between different tasks, preventing one task from dominating the entire learning process due to a large loss value.

[0052] In a setup containing multiple tasks, the loss function for each task can be expressed as: , where i represents the task number. The total loss function of uncertainty-weighted loss. It can be written as: ; in, It is the uncertainty estimate of the i-th task, usually called the variance. It is the uncertainty weight, which determines the proportion of each task's loss in the total loss. It is a regularization term used to penalize the increase of uncertainty in the estimate, preventing the model from growing infinitely. To minimize the first term.

[0053] Reference Figure 2As shown, this invention utilizes a hierarchical generation model based on task categories to categorize the intelligent driving system into perception priors, planning priors, decision priors, and control priors. The perception prior refers to generating synthetic perception data encompassing diverse driving scenarios through a diffusion model that integrates spatiotemporal attention mechanisms under physical simulation results and spatial invariance constraints. The planning and decision priors refer to guiding the temporal generation model to predict vehicle behavior data based on algebraic and differential equations from kinematic and dynamic models. The control prior refers to generating control data according to the logical rules of the vehicle control strategy. Furthermore, fine-grained prior knowledge separation is performed on each of these four stages from a scientific knowledge perspective. Finally, this hierarchical prior knowledge is used to expand and optimize the dataset, achieving data augmentation for each stage of the intelligent driving system, thereby improving the model's performance and generalization ability.

[0054] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0055] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0056] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0057] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.

[0058] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A data augmentation method for intelligent driving systems based on spatiotemporal joint hierarchical prior knowledge, characterized in that, include: Using a hierarchical prior knowledge representation method, prior knowledge is separated based on scientific knowledge and task categories. From the perspective of scientific knowledge, prior knowledge is separated from algebraic equations, differential equations, physical simulation results, space invariance, logical rules, knowledge graphs, probabilistic relationships, and human feedback. In terms of task categories, prior knowledge is separated from the environmental understanding and perception, motion planning, behavioral decision-making and control aspects of the intelligent driving system. A diffusion model integrating spatiotemporal attention mechanism under physical simulation results and spatial invariance constraints generates synthetic perception data containing diverse driving scenarios; algebraic and differential equations based on kinematic and dynamic models guide the temporal generation model to predict vehicle behavior data. Based on the logical rules of the vehicle control strategy, control data is generated; Based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data, end-to-end training of driving models in dynamic driving scenarios is achieved.

2. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 1, characterized in that, The diffusion model, which integrates spatiotemporal attention mechanisms under physical simulation results and spatial invariance constraints, generates synthetic perception data containing diverse driving scenarios, specifically including: By integrating a diffusion model with a spatiotemporal attention mechanism, combining physical simulation results for radar data and the principle of spatial invariance for all perception data, synthetic perception data containing diverse driving scenarios is generated based on the original perception data. The perception data includes images, radar data, and videos containing driving scenarios.

3. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 2, characterized in that, The diffusion model, which integrates a spatiotemporal attention mechanism, combines physical simulation results for radar data with the principle of spatial invariance for all perception data to generate synthetic perception data containing diverse driving scenarios based on the original perception data. Specifically, this includes: The physical simulation results are simulated standard radar data. Based on physical characteristics including radar electromagnetic wave propagation, scattering attenuation, and multipath clutter, standard radar data under different weather conditions, different obstructions, and different reflection characteristics are simulated and generated. The standard radar data is used as the physical constraint in the process of reverse denoising of the diffusion model to generate radar data. The principle of spatial invariance refers to adding spatial constraints during the generation of all perception data in the diffusion model by utilizing the properties of rigid targets, including vehicles, pedestrians, and road facilities, that are invariant in translation, rotation, and topology in driving scenarios. Noise is gradually added to the original perceived data through the forward diffusion process of the diffusion model until the original perceived data becomes completely random noise. By introducing a spatiotemporal attention mechanism in the inverse denoising process of the diffusion model, different levels of attention are given to the perceived data at different spatial locations and time points to capture the correlation between features at different spatiotemporal locations and generate synthetic perceived data containing diverse driving scenarios. The spatiotemporal attention mechanism includes a spatial attention mechanism and a temporal attention mechanism. The spatial attention mechanism is used to focus on important areas in the spatial dimension, and the temporal attention mechanism is used to focus on dynamic changes in the temporal dimension and adaptively adjust the diffusion model's attention to different time points. The perceived data includes images, radar data, and videos containing driving scenarios.

4. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 3, characterized in that, When noise is gradually added to the original sensor data, such as images, radar data, or videos, through the forward diffusion process of a diffusion model, the resulting completely random noise is... for: ; ; ; in, For the noisy image, noisy radar data, or noisy video obtained by adding standard Gaussian noise in step T, The noise distribution is a multivariate Gaussian noise distribution with a mean of 0 and a covariance matrix of identity matrix I. Here, represents the cumulative noise scaling factor, indicating the cumulative noise impact from step 1 to step T. For the first The noise scaling factor of the step. Let be the noise factor at step T, and t represent the index at step t.

5. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 1, characterized in that, The algebraic and differential equations based on kinematic and dynamic models guide the time-series generation model to predict vehicle behavior data, specifically including: Based on the kinematic model, algebraic equations are established to describe the relationship between the vehicle's position, velocity, and acceleration. Based on the dynamic model, differential equations are established to describe the relationship between the vehicle's forces and motion state. The vehicle's position, velocity, and acceleration determined by the kinematic and dynamic models are organized in the form of a time series and used as prior knowledge for the time series generation model. The time series generation model is trained to learn the changing patterns of vehicle behavior and is used to predict vehicle behavior data.

6. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 1, characterized in that, The vehicle behavior data includes the vehicle's trajectory, instantaneous speed, acceleration, heading angle change, driving curvature, relative distance, and relative speed under driving conditions such as straight driving, lane changing, turning, following, yielding, idling, starting and stopping, and acceleration and deceleration.

7. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 1, characterized in that, The logical rules of the vehicle control strategy include: speed control rules, steering control rules, and braking control rules.

8. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 7, characterized in that, The generation of control data based on the logical rules of the vehicle control strategy specifically includes: Speed ​​control rules include cruise control logic and adaptive cruise control logic; Cruise control logic: After setting the target speed, the engine throttle opening is automatically adjusted by comparing the actual vehicle speed with the target speed; when the actual vehicle speed is lower than the target speed, the vehicle accelerates; when the actual vehicle speed is higher than the target speed, the vehicle decelerates. Adaptive cruise control logic: If the distance to the vehicle in front is detected to be greater than the safe distance threshold and the current vehicle speed is lower than the set speed, the vehicle will accelerate; if the distance to the vehicle in front is less than the safe distance threshold, the vehicle will decelerate. Steering control rules include lane keeping logic and lane changing logic; The lane-keeping logic includes: if the vehicle deviates to the left from the lane centerline by a distance of x, then a right turn angle is generated. Control commands, It is directly proportional to x; if the vehicle deviates to the right from the lane centerline by a distance of x, then a leftward turning angle is generated. Control commands, It is directly proportional to x; The lane-changing logic includes: when the driver triggers a lane-changing signal or the vehicle's automatic driving system determines that the lane-changing conditions are met, generating steering control data based on the position of the target lane, and turning the steering wheel towards the target lane at a set angular velocity. The steering wheel's rotation angle and duration are determined based on the current relative position and speed of the vehicle and the target lane. The braking control rules include: if braking is required, the braking intensity is gradually increased according to the relationship between vehicle speed and distance, based on the kinematic model, so that the vehicle can stop at the target position.

9. The data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 1, characterized in that, The method of training the driving model end-to-end in dynamic driving scenarios based on raw perception data, synthetic perception data, predicted vehicle behavior data, and control data specifically includes: By utilizing raw perception data, synthetic perception data, predicted vehicle behavior data, and control data, a diverse range of simulated driving scenarios covering urban roads, highways, intersections, different weather conditions, and varying traffic densities are constructed. Raw perception data serves as a real-world benchmark, supervising the driving model's basic perception feature learning. Synthetic perception data serves as supplementary samples, enhancing the driving model's generalization ability. Predicted vehicle behavior data serves as a dynamic spatiotemporal prior, guiding the driving model to learn traffic interaction evolution patterns. Control data serves as supervised ground truth, constraining the accuracy of the driving model's decision-making and control outputs. A multi-task learning framework is built to simultaneously optimize perception, planning, decision-making, and control tasks. Multi-source features are fused using a shared network structure and loss function to achieve end-to-end joint training of the driving model. The driving model is an end-to-end Transformer autonomous driving decision-making model.

10. A data augmentation method for an intelligent driving system based on spatiotemporal joint hierarchical prior knowledge according to claim 9, characterized in that, The loss function for: ; in, It is the uncertainty estimate for the i-th task. It is the uncertainty weight of the i-th task. It is a regularization term. Let be the loss function for the i-th task.

Citation Information

Patent Citations

  • Automatic driving test scene simulation generalization generation method and system based on knowledge distillation

    CN121744955A

  • Automatic driving behavior regulation and control method and system based on perceptual reliability prior

    CN122166145A