Automatic driving obstacle intention prediction and avoidance method and system
By fusing the static and dynamic features of image and point cloud data to predict obstacle intentions and establish an avoidance strategy model with dual optimization objectives, the problems of inaccurate obstacle identification and unreasonable avoidance strategies are solved, achieving efficient and safe avoidance of the autonomous driving system.
Patent Information
- Application Number
- CN202511132004.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-13
AI Technical Summary
In existing autonomous driving technologies, obstacle recognition is inaccurate, dynamic behavior prediction is insufficient, and avoidance strategies are unreasonable, resulting in low driving safety and road traffic efficiency.
By acquiring image data and three-dimensional point cloud data in real time, the static and dynamic features of obstacles are extracted respectively, and the attention mechanism is used for weighted fusion to form a fused feature vector. Multiple independently trained deep learning models are input to predict the intention. Based on the predicted intention, the associated constraints between the vehicle operation status and input control are set, and an avoidance strategy model with path length and time cost as the dual optimization objectives is established. The input control is continuously updated until the optimization target value is iterated to the minimum, and the final execution instruction is generated.
It achieves high-precision prediction of obstacle intentions and flexible optimization of avoidance strategies, improving driving safety and road traffic efficiency, reducing the false triggering rate of emergency avoidance, and improving operation smoothness and traffic flow integration.
Smart Images

Figure CN120635865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method and system for predicting and avoiding obstacle intentions in autonomous driving. Background Art
[0002] In autonomous driving technology, accurate obstacle identification and efficient obstacle avoidance are key to ensuring driving safety and improving road efficiency. Existing obstacle identification methods primarily rely on, but are not limited to, data collected by sensors such as radar, lidar, and cameras, as well as image recognition technologies based on this data. While these technologies can achieve obstacle detection and location to a certain extent, they are often limited to analyzing data from a single sensor or focusing solely on static recognition of the obstacle's appearance.
[0003] For example, radar sensors can accurately measure the distance and relative speed of obstacles, but they have limitations in identifying the specific type, shape, and texture of obstacles. Camera sensors can capture rich visual information, including the color, shape, and texture of obstacles, but their performance can be significantly affected in complex environments (such as at night or in rainy and foggy weather). Furthermore, while existing image recognition technologies can identify the appearance of obstacles, they often lack a deep understanding of their dynamic behavior, making it difficult to accurately predict their future movement trajectory and intentions. Summary of the Invention
[0004] The technical problem to be solved by the embodiments of the present invention is to provide a method and system for predicting and avoiding obstacle intentions in autonomous driving, so as to solve the problem that the existing technology cannot accurately identify obstacles in an autonomous driving environment and the avoidance strategy is unreasonable.
[0005] The present invention discloses a method for predicting and avoiding obstacle intentions in an autonomous driving vehicle, comprising: Real-time acquisition of image data and 3D point cloud data of obstacles ahead during autonomous driving; extracting static features of the obstacle from the image data, and extracting dynamic features of the obstacle from the three-dimensional point cloud data; The static features and the dynamic features are weighted respectively by using an attention mechanism, and the weighted static features and the dynamic features are fused to obtain a fused feature vector; Inputting the fused feature vectors into multiple independently trained deep learning models respectively, and based on the model differences of the multiple deep learning models, weightedly fusing the output results of the multiple deep learning models to obtain the predicted intention of the obstacle; Setting associated constraints between the vehicle's operating state and input control based on the predicted intention, and establishing an avoidance strategy model with path length and time cost as dual optimization objectives; Continuously updating the input control in the prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within the associated constraints, and outputting a time series of the input control to constitute an avoidance strategy; The real-time operating status of the vehicle during the automatic driving process is obtained, and the final vehicle execution instruction is generated by combining the avoidance strategy and the real-time operating status.
[0006] Optionally, the acquiring of image data and three-dimensional point cloud data of a front obstacle during autonomous driving includes: Real-time collection of image data and three-dimensional point cloud data of obstacles ahead during autonomous driving, wherein the image data includes appearance information of the obstacles and the three-dimensional point cloud data includes shape and position information of the obstacles; The acquired image data is preprocessed, including first removing noise from the image data using wavelet transform, and then enhancing the contrast of the image data using histogram equalization. The function expression of the image data preprocessing is:
[0007] Where, represents the preprocessed image data, represents the image data before preprocessing, represents the composite transformation function for noise removal and contrast enhancement, Represents residual noise.
[0008] Optionally, extracting the static features of the obstacle from the image data includes: An appearance flow network constructed using multiple convolutional layers is used to extract static features of the obstacle from the preprocessed image data. The static features include the shape, texture, and color histogram of the obstacle. The function expression for extracting static features by the appearance flow network is:
[0009] Where, represents the feature map output by the appearance flow network, represents a nonlinear activation function, represents the weight matrix of the convolution kernel, represents the image data input to the appearance flow network, Represents the convolution bias term.
[0010] Optionally, extracting the dynamic features of the obstacle from the three-dimensional point cloud data includes: A motion flow network constructed using a gated loop is used to extract the dynamic features of the obstacle from the three-dimensional point cloud data. The dynamic features include the speed, direction, and acceleration of the obstacle. The function expression for extracting the dynamic features of the motion flow network is:
[0011]
[0012]
[0013] Where, Represents the reset gate, Reset gate weights. Represents the reset gate bias term, represents the update gate, represents the update gate weight, represents the update gate bias term, represents the candidate hidden state, represents the candidate hidden state weight, represents the candidate hidden state bias term, represents a nonlinear activation function, represents the hidden state at the previous time step, Represents the three-dimensional point cloud data input at the current time step, represents the hyperbolic tangent activation function, Represents element-wise multiplication.
[0014] Optionally, the autonomous driving obstacle intention prediction and avoidance method further includes a method for fusing the static features and the dynamic features, including: Define the first weight matrix and the first bias term of the attention mechanism, use the defined first weight matrix to map the static features to the calculation space of the attention weight, and obtain the attention weight of the static features through normalization function calculation. The calculation function expression of the static feature attention weight is:
[0015] Define the second weight matrix and the second bias term of the attention mechanism, use the defined second weight matrix to map the dynamic feature to the calculation space of the attention weight, and obtain the attention weight of the dynamic feature through normalization function calculation. The calculation function expression of the dynamic feature attention weight is:
[0016] Where, represents the attention weight of static features, represents the first weight matrix, Represents static features, represents the first bias term, represents the attention weight of dynamic features, represents the second weight matrix, Represents dynamic features, represents the second bias term, represents the normalization function; The weighted static features and the dynamic features are concatenated to obtain the fused feature vector. The function expression for concatenating the static features and the dynamic features is:
[0017] Where, represents the concatenated fusion feature vector, Represents element-wise multiplication.
[0018] Optionally, the autonomous driving obstacle intention prediction and avoidance method further includes a method of using multiple deep learning models to predict the fused feature vector to obtain the predicted intention, including: Obtain the fused feature vectors of multiple known intention obstacles during historical autonomous driving processes and use them as training data; Using multiple deep learning models constructed by convolutional neural networks, fully connected layers, and nonlinear activation functions, and independently training the multiple deep learning models using different initialization parameters and training data; Inputting the fused feature vectors of the obstacles obtained in real time into the multiple trained deep learning models respectively, and using the random forest algorithm to calculate the weight of the output results of each deep learning model; The output result of each deep learning model is multiplied by the corresponding weight, and the weighted output results of all the deep learning models are added and fused to obtain the final predicted intention. The function expression of the weighted fusion of multiple deep learning models is:
[0019] Where, Indicates the final prediction intention, represents the output result of the i-th deep learning model, represents the weight of the i-th deep learning model.
[0020] Optionally, setting associated constraints between the vehicle operating state and input control according to the predicted intention, and establishing an avoidance strategy model with path length and time cost as dual optimization objectives, includes: The association between the operating state and the input control during the vehicle's automatic driving process is set according to the vehicle's dynamic characteristics and road constraints. The functional expression for the association between the operating state and the input control is:
[0021] Where, represents the vehicle operating status at time t, represents the vehicle input control at time t, A system dynamics model representing the vehicle; According to the predicted intention of the obstacle, the constraints of the operating state and the input control during the vehicle automatic driving process are set. The functional expressions of the operating state and the input control constraints are:
[0022] Where, represents a set of constraints, represents the obstacle prediction intention at time t; An avoidance strategy model in the prediction time domain is established with path length and time cost as dual optimization objectives. The function expression of the avoidance strategy model is:
[0023] Where, represents the dual optimization objective value output by the avoidance strategy model, and represents the weight coefficient, and T represents the prediction time domain.
[0024] Optionally, acquiring the real-time operating status of the vehicle during the automatic driving process and generating a final vehicle execution instruction in combination with the avoidance strategy and the real-time operating status includes: Determining the desired operating state of the vehicle based on the avoidance strategy, and acquiring the current actual operating state of the vehicle during the autonomous driving process; The target acceleration of the vehicle control execution is calculated based on the expected operating state and the actual operating state. The calculation function expression of the target acceleration is:
[0025] Where, represents the target acceleration for vehicle control execution, represents the expected speed, represents the expected acceleration, Indicates the current actual speed of the vehicle. Indicates the current actual acceleration of the vehicle. is the proportional gain coefficient, is the differential gain coefficient; Output corresponding vehicle control instructions based on the calculated target acceleration.
[0026] The present invention also discloses a prediction and avoidance system, which adopts the above-mentioned automatic driving obstacle intention prediction and avoidance method, and the system includes: The data acquisition module is used to obtain real-time image data and 3D point cloud data of obstacles ahead during autonomous driving; a feature extraction module, configured to extract static features of the obstacle from the image data and dynamic features of the obstacle from the three-dimensional point cloud data; a feature fusion module, configured to weight the static features and the dynamic features respectively using an attention mechanism, and fuse the weighted static features and the dynamic features to obtain a fused feature vector; an intention prediction module, configured to input the fused feature vector into a plurality of independently trained deep learning models, and based on model differences among the plurality of deep learning models, weightedly fuse the output results of the plurality of deep learning models to obtain a predicted intention of the obstacle; a model building module for setting associated constraints between vehicle operating states and input controls based on the predicted intent, and establishing an avoidance strategy model with path length and time cost as dual optimization objectives; an avoidance strategy generation module, configured to continuously update the input control in a prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within associated constraints, and output a time series of the input control to constitute an avoidance strategy; The control instruction generation module is used to obtain the real-time operating status of the vehicle during the automatic driving process, and generate the final vehicle execution instruction in combination with the avoidance strategy and the real-time operating status.
[0027] The present invention also discloses a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned automatic driving obstacle intention prediction and avoidance method is implemented.
[0028] Compared with the prior art, the method and system for predicting and avoiding obstacle intentions in autonomous driving provided by the embodiments of the present invention have the following beneficial effects: By acquiring image data and three-dimensional point cloud data in real time, static features and dynamic features are extracted respectively. These features are then weighted and fused using an attention mechanism to form a fused feature vector. The fused feature vector is then fed into multiple independently trained deep learning models, and the output of the model difference weighted fusion is used to obtain the predicted intent. Based on the predicted intent, constraints associated with the operating state and input control are set. An avoidance strategy model is established with path length and time cost as dual optimization objectives. Input control is continuously updated within the prediction time domain until the optimized target value is iteratively minimized. A time series of avoidance strategies is output, and the final execution instructions are generated based on the real-time operating state. This solves the problems of existing technologies in complex environments, such as inaccurate obstacle recognition, insufficient intention prediction, and unreasonable avoidance strategies. It achieves more accurate intention prediction, more flexible constraint and optimization objective setting, and more efficient iterative control, significantly improving driving safety and road traffic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments, in which: Figure 1 A schematic block diagram of the overall steps of the method for predicting and avoiding obstacles in an autonomous driving vehicle according to an embodiment of the present invention; Figure 2 A schematic block diagram of the intention prediction and avoidance process of the autonomous driving obstacle intention prediction and avoidance method provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. Now, in conjunction with the accompanying drawings, the preferred embodiments of the present invention will be described in detail.
[0031] The present invention discloses a method for predicting and avoiding obstacles in an autonomous driving vehicle. Figure 1 and Figure 2 Shown, including: S1, real-time acquisition of image data and 3D point cloud data of obstacles ahead during autonomous driving; S2, extracting static features of obstacles from image data and dynamic features of obstacles from 3D point cloud data; S3. Use the attention mechanism to weight the static features and dynamic features respectively, and fuse the weighted static features and dynamic features to obtain a fused feature vector; S4. Input the fused feature vectors into multiple independently trained deep learning models, and based on the model differences of the multiple deep learning models, perform weighted fusion of the output results of the multiple deep learning models to obtain the predicted intention of the obstacle; S5. Set the associated constraints between the vehicle's operating state and input control based on the predicted intention, and establish an avoidance strategy model with path length and time cost as dual optimization objectives; S6. Continuously updating the input control in the prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within the associated constraints, and the time series of the output and input control constitutes the avoidance strategy; S7. Obtain the real-time operating status of the vehicle during the autonomous driving process, and generate the final vehicle execution instruction based on the avoidance strategy and the real-time operating status.
[0032] By implementing the aforementioned embodiment of the obstacle intention prediction and avoidance method for autonomous driving, real-time image data and three-dimensional point cloud data are acquired and processed in parallel, enabling simultaneous and in-depth extraction of both static and dynamic features of obstacles, overcoming the limitations of single-data-source analysis. An attention mechanism is then employed to weight the two types of features separately, dynamically adjusting the weights based on feature importance. These features are then fused to form a high-dimensional fused feature vector with complementary information, significantly improving feature representation. The fused feature vectors are then fed into multiple independently trained deep learning models for multimodal analysis. Leveraging the diverse outputs generated by differences in model structure, the weighted fusion mechanism integrates the predictive strengths of multiple models, effectively overcoming the bias of individual models and achieving highly accurate predictions of the obstacle's future motion state and behavioral motivations. Based on this prediction, constraints are set for the vehicle's operating state and input controls. A mathematical optimization model with dual optimization objectives of path length and time cost is constructed. The input controls are continuously updated within the prediction time domain until the dual optimization objectives converge to a minimum while satisfying all associated constraints. Finally, an avoidance strategy consisting of an optimized input control sequence is output. The final execution instructions are thus dynamically generated in real time in combination with the actual operating status of the vehicle, thereby solving the problem of misjudgment of obstacle intentions and rigid avoidance strategies caused by the limitations of static feature recognition and the lack of dynamic behavior prediction in the existing technology. That is, the embodiment of the present invention uses a multi-level feature fusion mechanism, a multi-model collaborative prediction mechanism, and a dynamic optimization solution mechanism to achieve full-dimensional and accurate perception of obstacle intentions in complex environments, probabilistic modeling and prediction of motion states, and real-time adaptive optimization of avoidance decisions, so that the autonomous driving system has the ability to proactively deal with emergencies, while ensuring the optimality and timeliness of the driving trajectory. This can not only significantly reduce the false triggering rate of emergency avoidance, but also improve the smoothness of continuous avoidance operations and the degree of traffic flow integration, thereby breaking through the technical bottleneck of traditional systems based on static features and preset rules, and significantly optimizing road traffic efficiency while improving driving safety redundancy.
[0033] Furthermore, image data and 3D point cloud data of obstacles ahead during autonomous driving are obtained, including: Real-time collection of image data and 3D point cloud data of obstacles ahead during autonomous driving. The image data contains information about the appearance of the obstacles, while the 3D point cloud data contains information about the shape and location of the obstacles. The acquired image data is preprocessed, including using wavelet transform to remove noise from the image data, and then using histogram equalization to enhance the contrast of the image data. The function expression of image data preprocessing is:
[0034] Where, represents the preprocessed image data, represents the image data before preprocessing, represents the composite transformation function for noise removal and contrast enhancement, Represents residual noise.
[0035] By implementing the aforementioned embodiment of the autonomous driving obstacle intention prediction and avoidance method, the preprocessed image data significantly improves noise resistance and texture clarity while preserving the integrity of appearance information. Noise removal suppresses invalid pixel perturbations caused by environmental interference, contrast enhancement enhances target edges and detailed features, and the inclusion of residual noise prevents feature loss due to oversmoothing, enabling the 3D point cloud data to be transmitted in its original state, fully maintaining the spatial accuracy and real-time performance of its shape and position information. The high-quality preprocessed image data provides a low-noise, high-resolution input foundation for static feature extraction, overcoming feature distortion caused by single data quality degradation in complex scenarios such as rain, fog, and nighttime. It provides high-fidelity data support for the construction of fused feature vectors, thereby enhancing the robustness of the autonomous driving system's perception of the essential attributes (appearance, shape, and position) of obstacles in harsh environments.
[0036] As mentioned above, while the autonomous vehicle is driving, it preferably uses sensors such as high-precision cameras (resolution of 4096x3072 pixels, frame rate of 60fps) and LiDAR (scanning frequency of 20Hz, point cloud density of 2 million points per second) to collect real-time image data and 3D point cloud data of obstacles ahead. The image data details the appearance of the obstacles, such as posture (standing, walking, running, etc.), clothing color (red, blue, green, etc.), and size (height, width). The 3D point cloud data accurately depicts the shape of the obstacles (such as the outline of a human body or the shape of a suitcase) and position information (distance and azimuth relative to the vehicle). To improve the accuracy of subsequent processing, the collected data is preprocessed: wavelet transform (Daubechies 4 wavelet basis, 5-layer decomposition) is used to remove high-frequency noise from the image data, and histogram equalization (contrast stretching to the range of 0-255) is used to enhance image contrast.
[0037] Furthermore, static features of obstacles are extracted from the image data, including: The appearance flow network constructed by multiple convolutional layers is used to extract the static features of obstacles from the preprocessed image data. The static features include the shape, texture, and color histogram of the obstacles. The function expression of the appearance flow network to extract static features is:
[0038] Where, represents the feature map output by the appearance flow network, represents a nonlinear activation function, represents the weight matrix of the convolution kernel, represents the image data input to the appearance flow network, Represents the convolution bias term.
[0039] Through the implementation of the aforementioned autonomous driving obstacle intention prediction and avoidance method embodiment, an appearance flow network is constructed to perform deep spatial feature learning on the input image data using a convolution kernel weight matrix. This network, combined with a nonlinear activation function, enhances the high-order nonlinear modeling of obstacle geometry and surface properties. The introduction of a convolution bias term effectively adapts to the feature distribution shift of image data under varying brightness conditions. This allows the network's multi-layered convolutional structure to extract microscopic texture details while preserving the integrity of macroscopic shape contours. The output feature map exhibits both local feature resolution and global structural consistency, significantly enhancing the robustness of the representation of the essential attributes of obstacles in complex scenarios such as rain, fog, and low illumination at night, providing high-quality static feature input that is resistant to environmental interference for intention prediction.
[0040] As mentioned above, the appearance flow network preferably uses a deep convolutional neural network, specifically the ResNet-101 model (Deep Residual Network with 101 layers), for feature extraction. ResNet-101 contains 101 convolutional layers, with the final fully connected layer removed to output feature maps. These feature maps contain static features such as the shape, texture, and color histogram of the obstacle. The nonlinear activation function used is the ReLU (Rectified Linear Unit) nonlinear activation function.
[0041] Furthermore, the dynamic features of obstacles are extracted from the 3D point cloud data, including: The motion flow network constructed by gated loop is used to extract the dynamic features of obstacles from the 3D point cloud data. The dynamic features include the speed, direction, and acceleration of the obstacles. The function expression of the motion flow network to extract dynamic features is:
[0042]
[0043]
[0044] Where, Represents the reset gate, Reset gate weights. Represents the reset gate bias term, represents the update gate, represents the update gate weight, represents the update gate bias term, represents the candidate hidden state, represents the candidate hidden state weight, represents the candidate hidden state bias term, represents a nonlinear activation function, represents the hidden state at the previous time step, Represents the three-dimensional point cloud data input at the current time step, represents the hyperbolic tangent activation function, Represents element-wise multiplication.
[0045] By implementing the above-mentioned embodiment of the automatic driving obstacle intention prediction and avoidance method, the dynamic features (including speed, direction, acceleration) of the obstacle are extracted from the three-dimensional point cloud data using the motion flow network constructed by the gated loop, and the dynamic feature extraction process is precisely defined by the function expression: Reset the gate By weight and bias Calculate the hidden state of the previous time step With the current input A nonlinear combination of update gates By weight and bias Calculate the combination of the two, the candidate hidden state Based on the reset gate Hidden state for the previous time step The weighted result and the current input By weight and bias terms The hyperbolic tangent activation output is calculated, and the dual-gating mechanism of the reset gate and the update gate is used to adaptively filter the effective information of the obstacle's dynamic features and suppress redundant interference. The spatiotemporal correlation modeling of real-time 3D point cloud data using candidate hidden states accurately captures the continuous evolution of the obstacle's motion characteristics. This enables the dynamic feature extraction process to deeply integrate the obstacle's temporal behavior pattern and instantaneous state changes, significantly enhancing the robustness of the representation of the dynamic evolution of speed, direction, and acceleration in complex motion trajectories (such as accelerated lane changes and sharp turns), providing a high-precision dynamic feature foundation with strong temporal dependence for intention prediction.
[0046] As mentioned above, the motion flow network preferably uses a gated recurrent unit (GRU) to convert the point cloud data collected by the lidar into time series data, which is then input into the motion flow network to extract the dynamic characteristics of obstacles. The motion flow network can capture the speed, direction, acceleration, and other motion states of obstacles, as well as their changing trends over time.
[0047] Furthermore, the autonomous driving obstacle intention prediction and avoidance method also includes a method of fusing static features and dynamic features, including: Define the first weight matrix and the first bias term of the attention mechanism, use the defined first weight matrix to map the static features to the calculation space of the attention weight, and obtain the attention weight of the static features through the normalization function calculation. The calculation function expression of the static feature attention weight is:
[0048] Define the second weight matrix and the second bias term of the attention mechanism, use the defined second weight matrix to map the dynamic features to the calculation space of the attention weight, and obtain the attention weight of the dynamic features through the normalization function calculation. The calculation function expression of the dynamic feature attention weight is:
[0049] Where, represents the attention weight of static features, represents the first weight matrix, Represents static features, represents the first bias term, represents the attention weight of dynamic features, represents the second weight matrix, Represents dynamic features, represents the second bias term, represents the normalization function; The weighted static features and dynamic features are spliced together to obtain a fused feature vector. The function expression for splicing static features and dynamic features is:
[0050] Where, represents the concatenated fusion feature vector, Represents element-wise multiplication.
[0051] Through the implementation of the aforementioned autonomous driving obstacle intention prediction and avoidance method embodiment, a dual-parameter space is constructed using separate first and second weight matrices. Static and dynamic features are independently nonlinearly mapped during the attention weight calculation process. Attention weights that strictly match the feature importance are then generated using a normalization function. Subsequently, the weighted static and dynamic features are subjected to a refined fusion process with dimension alignment using element-wise multiplication. This mechanism overcomes the feature representation confusion caused by weight sharing in traditional feature concatenation, significantly enhancing the fused feature vector's ability to collaboratively represent the multimodal attributes of obstacles in complex scenarios. Specifically, when environmental interference weakens the reliability of static features, the first weight matrix automatically reduces the weight assigned to fuzzy textures, while the second weight matrix simultaneously enhances the contribution of dynamic features. This maintains the information completeness of the fused feature vector even in complex environments, providing a robust input foundation for adaptive environmental changes in subsequent obstacle intention prediction.
[0052] As described above, preferably, the normalization function adopts the Softmax function (Soft Maximum Function).
[0053] To further illustrate, let's take an example to handle the obstacle detection task in front of an autonomous vehicle: The appearance flow network is responsible for extracting static features of obstacles from a single image frame, such as shape, color, texture, etc. These features are represented as a vector with a dimension of 128 The motion flow network is responsible for extracting the dynamic features of obstacles from continuous video frames, such as speed, acceleration, direction change, etc. These features are also represented as a vector with a dimension of 128 .
[0054] In order to fuse these two feature vectors into a fused feature vector, the attention mechanism is introduced. First, define two weight matrices and , their dimensions are all 1x128, which are used to map appearance features and motion features into the calculation space of attention weights. At the same time, two bias terms are also defined and , their dimensions are all 1 and are used to adjust the calculation of attention weights.
[0055] Next, the softmax function is used to calculate the static features and dynamic features The attention weight and ; Assuming that the calculation = [0.6, 0.2, 0.1, ..., 0.1] (only some elements are shown, it is actually a 128-dimensional vector), = [0.5, 0.3, 0.1, ..., 0.1] (only some elements are shown).
[0056] Attention weight and Indicates that during the fusion process, static features and dynamic features The importance of each element.
[0057] Then, the weighted features are concatenated to obtain a fused feature vector with a dimension of 256, which contains the fusion information of appearance features and motion features, providing richer feature representation for subsequent obstacle classification and detection tasks.
[0058] Furthermore, the autonomous driving obstacle intention prediction and avoidance method also includes a method for using multiple deep learning models to predict the fused feature vector to obtain the predicted intention, including: Obtain the fused feature vectors of multiple known intention obstacles during historical autonomous driving processes and use them as training data; Multiple deep learning models, each constructed from convolutional neural networks, fully connected layers, and nonlinear activation functions, were trained independently using different initialization parameters and training data. The fusion feature vectors of obstacles acquired in real time are input into multiple trained deep learning models respectively, and the random forest algorithm is used to calculate the weight of the output results of each deep learning model; The output of each deep learning model is multiplied by the corresponding weight, and the weighted output results of all deep learning models are added and fused to obtain the final prediction intent. The function expression of weighted fusion of multiple deep learning models is:
[0059] Where, Indicates the final prediction intention, represents the output result of the i-th deep learning model, represents the weight of the i-th deep learning model.
[0060] Through the implementation of the aforementioned autonomous driving obstacle intention prediction and avoidance method embodiment, a random forest algorithm dynamically calculates the weights of each deep learning model's output. Each independently trained deep learning model exhibits inherent heterogeneity due to differences in initialization parameters and training data, resulting in significant diversity in the prediction results of different deep learning models for the same fused feature vector. Based on historical training data of fused feature vectors of known obstacle intentions during autonomous driving, the random forest algorithm adaptively assigns weights to each deep learning model through feature importance analysis, ensuring that the decision contribution of highly adaptable deep learning models is automatically strengthened when the obstacle's motion pattern suddenly changes. Finally, through a weighted fusion mechanism that multiplies the weights by the outputs of the corresponding deep learning models and then adds them together, this overcomes the generalization bottleneck and scenario-dependency limitations of a single deep learning model. This results in predictions that combine the advantages of multi-model redundancy with the dynamic calibration capabilities of the random forest algorithm, significantly improving the robustness of obstacle identification and prediction error tolerance in complex interactive scenarios, such as lane changes with multiple obstacles and deceleration with ambiguous intentions.
[0061] As mentioned above, preferably, the nonlinear activation function of the deep learning model adopts the softmax activation function (SoftMaximum Activation Function). To improve the accuracy of classification, an ensemble learning method such as random forest or gradient boosting tree is used to fuse the output results of multiple deep learning models.
[0062] Further examples: Consider an autonomous driving scenario with 10 obstacles, each with 256-dimensional comprehensive features. Three deep learning models (e.g., Model A, Model B, and Model C) are trained to predict the intentions of these obstacles. These models have the same network structure but use different initialization parameters and training data.
[0063] During the training process, cross-validation can be used to evaluate the performance of each model and select the optimal model parameters. Assume that the accuracy of Model A, Model B, and Model C on the validation set are 90%, 85%, and 88%, respectively.
[0064] Then, a random forest is used as an ensemble learning method to calculate the weight of each model. Random forest improves classification accuracy by constructing multiple decision trees and combining their prediction results. In this embodiment of the present invention, 50 decision trees can be constructed to calculate the weight of each model.
[0065] Through the calculation of random forest, the weights of model A, model B and model C are 0.4, 0.3 and 0.3 respectively. These weights reflect the importance of each model in the final prediction.
[0066] Finally, the output of each model is multiplied by its corresponding weight, and the weighted outputs are summed to obtain the final predicted intent. This predicted intent represents the probability distribution of each obstacle's intent and can be used for decision-making and control of the autonomous driving system.
[0067] Furthermore, the associated constraints of the vehicle's operating state and input control are set according to the predicted intention, and an avoidance strategy model is established with path length and time cost as the dual optimization objectives, including: The relationship between the vehicle's operating state and input control during the vehicle's autonomous driving process is set based on the vehicle's dynamic characteristics and road constraints. The functional expression for the relationship between the operating state and input control is:
[0068] Where, represents the vehicle operating status at time t, represents the vehicle input control at time t, A system dynamics model representing the vehicle; According to the predicted intention of the obstacle, the constraints of the vehicle's operating state and input control during the automatic driving process are set. The functional expressions of the operating state and input control constraints are:
[0069] Where, represents a set of constraints, represents the obstacle prediction intention at time t; An avoidance strategy model in the prediction time domain is established with path length and time cost as dual optimization objectives. The function expression of the avoidance strategy model is:
[0070] Where, represents the dual optimization objective value output by the avoidance strategy model, and represents the weight coefficient, and T represents the prediction time domain.
[0071] By implementing the above-mentioned automatic driving obstacle intention prediction and avoidance method embodiment, the system dynamics model is used to accurately embed the vehicle's own physical characteristics (such as steering inertia, drive response) and road geometric constraints (such as curvature boundaries and slope limits), so that the operation state update process strictly follows the actual mechanical dynamics laws. At the same time, the time-varying obstacle prediction intention is used as the adaptive input of the constraint function, forcing the avoidance strategy to respond to the evolution of external threats in real time within the prediction time domain, and through the dual optimization objective function, the rolling joint optimization of path length and time cost within the prediction time domain is achieved. Among them, Indicates the vehicle operating status at time t, such as vehicle position, speed, acceleration, etc.; Represents the vehicle input control at time t, such as accelerator, brake, steering, etc.; represents the system dynamics model of the vehicle, taking into account the vehicle's dynamic characteristics and road constraints; g is a set of constraints used to ensure that the avoidance strategy does not violate traffic rules, avoid collisions, keep lanes, etc. These constraints may include speed limits, steering angle limits, avoidance of collisions with other vehicles or obstacles, etc.; weight coefficient and It is used to balance the priorities between different optimization objectives. Its coordinated adjustment breaks through the barriers of efficiency and safety in traditional single-objective optimization. Ultimately, under the dual closed-loop effects of state transfer constraints and intention-driven constraints, the continuously updated input control The sequence systematically converges to a dynamic trajectory solution that minimizes the dual optimization objective value under the premise of satisfying the vehicle's own motion boundary and obstacle avoidance safety threshold. That is, by iteratively solving the dual optimization objective problem, that is, under the premise of satisfying the constraints, find the input control sequence that minimizes the avoidance strategy model J. , thus obtaining a series of optimal input controls, which will constitute the avoidance strategy to guide the vehicle's automatic avoidance behavior.
[0072] Based on the above content, in detail: input control in the prediction time domain Perform closed-loop rolling optimization to make the dual optimization objective values of the avoidance strategy model in the system dynamics model and intention constraint set Continuous iteration and update under the dual boundary constraints. Each iteration is based on the current vehicle operation status Resolve the constrained optimization problem and adaptively adjust the input control by gradient descent method The spatial and temporal distribution of the dual optimization objective value converges to the minimum value in the current prediction time domain. This mechanism breaks through the static limitations of traditional open-loop decision-making, making the time series of the final output input control strictly comply with the mechanical constraints of vehicle steering inertia and road curvature, the safety avoidance boundary of the dynamic coupling obstacle prediction intention, and the path length and time cost in the weight coefficient. and The global optimal balance under precise ratio is achieved, and then an avoidance strategy that is physically feasible, responsive to intentions, and has optimal energy consumption is generated at the output end.
[0073] Furthermore, the real-time operating status of the vehicle during the autonomous driving process is obtained, and the final vehicle execution instructions are generated by combining the avoidance strategy and the real-time operating status, including: Determine the desired operating state of the vehicle based on the avoidance strategy, and collect the actual operating state of the vehicle during the autonomous driving process; The target acceleration of vehicle control execution is calculated based on the expected operating state and the actual operating state. The calculation function expression of the target acceleration is:
[0074] Where, represents the target acceleration for vehicle control execution, represents the expected speed, represents the expected acceleration, Indicates the current actual speed of the vehicle. Indicates the current actual acceleration of the vehicle. is the proportional gain coefficient, is the differential gain coefficient; Output corresponding vehicle control instructions based on the calculated target acceleration.
[0075] By implementing the above-mentioned automatic driving obstacle intention prediction and avoidance method embodiment, the proportional gain coefficient is used. Expected speed and actual speed The difference is dynamically scaled and adjusted, and the differential gain coefficient is used Expected acceleration With actual acceleration The deviation is corrected and compensated in real time, forming a dual closed-loop negative feedback mechanism of speed and acceleration. This feedforward-feedback composite control strategy that integrates the actual operating state can break through the limitations of traditional open-loop control that only relies on the expected operating state, making the target acceleration The calculation result can immediately respond to the transient error between the vehicle's motion state and the avoidance strategy requirements, that is, when the actual acceleration Delayed acceleration due to sudden road adhesion or mechanical delay When the differential term A compensatory increment will be generated to force the closed loop to converge; when the actual speed Deviation from expected speed When the proportional term Automatically generate restorative adjustments. The synergistic effect of the two ensures that the resulting vehicle execution instructions have both dynamic correction capabilities and motion continuity assurance. While meeting the safety constraints of the avoidance strategy, it completely solves the acceleration jerks, speed oscillations, and response delays caused by the lack of state feedback in traditional methods, significantly improving the smoothness and adaptability of autonomous driving control.
[0076] As mentioned above, the avoidance strategy is transmitted to the control module of the autonomous driving system, which controls the vehicle's throttle, brake, steering and other actuators to achieve automatic avoidance of the vehicle. When executing the avoidance strategy, the control module will The system calculates corresponding control instructions, such as adjusting the throttle opening, braking force, or steering angle. These control instructions achieve precise control of the vehicle by controlling the vehicle's throttle, brakes, steering, and other actuators, allowing the autonomous vehicle to accurately track the desired acceleration and speed for safe and stable avoidance maneuvers.
[0077] The present invention also discloses a prediction and avoidance system, which adopts the above-mentioned automatic driving obstacle intention prediction and avoidance method, and the system includes: The data acquisition module is used to obtain real-time image data and 3D point cloud data of obstacles ahead during autonomous driving; Feature extraction module, used to extract static features of obstacles from image data and dynamic features of obstacles from 3D point cloud data; The feature fusion module is used to weight the static features and dynamic features separately using the attention mechanism, and fuse the weighted static features and dynamic features to obtain a fused feature vector; The intention prediction module is used to input the fused feature vectors into multiple independently trained deep learning models. Based on the model differences of multiple deep learning models, the output results of multiple deep learning models are weighted and fused to obtain the predicted intention of the obstacle; The model building module is used to set the associated constraints of the vehicle's operating state and input control according to the predicted intention, and to establish an avoidance strategy model with path length and time cost as the dual optimization objectives; The avoidance strategy generation module is used to continuously update the input control in the prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within the associated constraints. The time series of the output input control constitutes the avoidance strategy. The control instruction generation module is used to obtain the real-time operating status of the vehicle during autonomous driving, and generate the final vehicle execution instructions based on the avoidance strategy and real-time operating status.
[0078] The present invention also discloses a computer-readable storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, the above-mentioned automatic driving obstacle intention prediction and avoidance method is implemented.
[0079] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, characterized in that the processor implements the above-mentioned automatic driving obstacle intention prediction and avoidance method when executing the computer program.
[0080] The present invention is described based on flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to specific embodiments. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0081] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0083] It should be understood that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Those skilled in the art may modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents; and all these modifications and replacements should fall within the scope of protection of the present invention.
Claims
1. A method for predicting and avoiding obstacles in an autonomous driving vehicle, characterized in that: include: Real-time acquisition of image data and 3D point cloud data of obstacles ahead during autonomous driving; extracting static features of the obstacle from the image data, and extracting dynamic features of the obstacle from the three-dimensional point cloud data; The static features and the dynamic features are weighted respectively by using an attention mechanism, and the weighted static features and the dynamic features are fused to obtain a fused feature vector; Inputting the fused feature vectors into multiple independently trained deep learning models respectively, and based on the model differences of the multiple deep learning models, weightedly fusing the output results of the multiple deep learning models to obtain the predicted intention of the obstacle; Setting associated constraints between the vehicle's operating state and input control based on the predicted intention, and establishing an avoidance strategy model with path length and time cost as dual optimization objectives; Continuously updating the input control in the prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within the associated constraints, and outputting a time series of the input control to constitute an avoidance strategy; The real-time operating status of the vehicle during the automatic driving process is obtained, and the final vehicle execution instruction is generated by combining the avoidance strategy and the real-time operating status.
2. The method for predicting and avoiding obstacles in autonomous driving according to claim 1, wherein: The obtaining of image data and three-dimensional point cloud data of obstacles ahead during the autonomous driving process includes: Real-time collection of image data and three-dimensional point cloud data of obstacles ahead during autonomous driving, wherein the image data includes appearance information of the obstacles and the three-dimensional point cloud data includes shape and position information of the obstacles; The acquired image data is preprocessed, including first removing noise from the image data using wavelet transform, and then enhancing the contrast of the image data using histogram equalization. The function expression of the image data preprocessing is: Where, represents the preprocessed image data, represents the image data before preprocessing, represents the composite transformation function for noise removal and contrast enhancement, Represents residual noise.
3. The method for predicting and avoiding obstacles in autonomous driving according to claim 2, wherein: The extracting the static features of the obstacle from the image data includes: An appearance flow network constructed using multiple convolutional layers is used to extract static features of the obstacle from the preprocessed image data. The static features include the shape, texture, and color histogram of the obstacle. The function expression for extracting static features by the appearance flow network is: Where, represents the feature map output by the appearance flow network, represents a nonlinear activation function, represents the weight matrix of the convolution kernel, represents the image data input to the appearance flow network, Represents the convolution bias term.
4. The method for predicting and avoiding obstacles in autonomous driving according to claim 1, wherein: The extracting the dynamic features of the obstacle from the three-dimensional point cloud data includes: A motion flow network constructed using a gated loop is used to extract the dynamic features of the obstacle from the three-dimensional point cloud data. The dynamic features include the speed, direction, and acceleration of the obstacle. The function expression for extracting the dynamic features of the motion flow network is: Where, Represents the reset gate, Reset gate weights. Represents the reset gate bias term, represents the update gate, represents the update gate weight, represents the update gate bias term, represents the candidate hidden state, represents the candidate hidden state weight, represents the candidate hidden state bias term, represents a nonlinear activation function, represents the hidden state at the previous time step, Represents the three-dimensional point cloud data input at the current time step, represents the hyperbolic tangent activation function, Represents element-wise multiplication.
5. The automatic driving obstacle intention prediction and avoidance method according to claim 1, characterized in that: The autonomous driving obstacle intention prediction and avoidance method also includes a method for fusing the static features and the dynamic features, including: Define the first weight matrix and the first bias term of the attention mechanism, use the defined first weight matrix to map the static features to the calculation space of the attention weight, and obtain the attention weight of the static features through normalization function calculation. The calculation function expression of the static feature attention weight is: Define the second weight matrix and the second bias term of the attention mechanism, use the defined second weight matrix to map the dynamic feature to the calculation space of the attention weight, and obtain the attention weight of the dynamic feature through normalization function calculation. The calculation function expression of the dynamic feature attention weight is: Where, represents the attention weight of static features, represents the first weight matrix, Represents static features, represents the first bias term, represents the attention weight of dynamic features, represents the second weight matrix, Represents dynamic features, represents the second bias term, represents the normalization function; The weighted static features and the dynamic features are concatenated to obtain the fused feature vector. The function expression for concatenating the static features and the dynamic features is: Where, represents the concatenated fusion feature vector, Represents element-wise multiplication.
6. The method for predicting and avoiding obstacles in an autonomous driving process according to claim 1, wherein: The autonomous driving obstacle intention prediction and avoidance method also includes a method for using multiple deep learning models to predict the fused feature vector to obtain the predicted intention, including: Obtain the fused feature vectors of multiple known intention obstacles during historical autonomous driving processes and use them as training data; Using multiple deep learning models constructed by convolutional neural networks, fully connected layers, and nonlinear activation functions, and independently training the multiple deep learning models using different initialization parameters and training data; Inputting the fused feature vectors of the obstacles obtained in real time into the multiple trained deep learning models respectively, and using the random forest algorithm to calculate the weight of the output results of each deep learning model; The output result of each deep learning model is multiplied by the corresponding weight, and the weighted output results of all the deep learning models are added and fused to obtain the final predicted intention. The function expression of the weighted fusion of multiple deep learning models is: Where, Indicates the final prediction intention, represents the output of the i-th deep learning model, represents the weight of the i-th deep learning model.
7. The method for predicting and avoiding obstacles in an autonomous driving process according to claim 1, wherein: The method of setting associated constraints between the vehicle operating state and input control according to the predicted intention and establishing an avoidance strategy model with path length and time cost as dual optimization objectives includes: The association between the operating state and the input control during the vehicle's automatic driving process is set according to the vehicle's dynamic characteristics and road constraints. The functional expression for the association between the operating state and the input control is: Where, represents the vehicle operating status at time t, represents the vehicle input control at time t, A system dynamics model representing the vehicle; According to the predicted intention of the obstacle, the constraints of the operating state and the input control during the vehicle automatic driving process are set. The functional expressions of the operating state and the input control constraints are: Where, represents a set of constraints, represents the obstacle prediction intention at time t; An avoidance strategy model in the prediction time domain is established with path length and time cost as dual optimization objectives. The function expression of the avoidance strategy model is: Where, represents the dual optimization objective value output by the avoidance strategy model, and represents the weight coefficient, and T represents the prediction time domain.
8. The automatic driving obstacle intention prediction and avoidance method according to claim 1, characterized in that: The obtaining of the real-time operating status of the vehicle during the automatic driving process and generating a final vehicle execution instruction in combination with the avoidance strategy and the real-time operating status include: Determine the desired operating state of the vehicle according to the avoidance strategy, and acquire the current actual operating state of the vehicle during the autonomous driving process; The target acceleration of the vehicle control execution is calculated based on the expected operating state and the actual operating state. The calculation function expression of the target acceleration is: Where, represents the target acceleration for vehicle control execution, represents the expected speed, represents the expected acceleration, Indicates the current actual speed of the vehicle. Indicates the current actual acceleration of the vehicle. is the proportional gain coefficient, is the differential gain coefficient; Output corresponding vehicle control instructions based on the calculated target acceleration.
9. A prediction and avoidance system, using the automatic driving obstacle intention prediction and avoidance method according to any one of claims 1 to 8, characterized in that: The system comprises: The data acquisition module is used to obtain real-time image data and 3D point cloud data of obstacles ahead during autonomous driving; a feature extraction module, configured to extract static features of the obstacle from the image data and dynamic features of the obstacle from the three-dimensional point cloud data; a feature fusion module, configured to weight the static features and the dynamic features respectively using an attention mechanism, and fuse the weighted static features and the dynamic features to obtain a fused feature vector; an intention prediction module, configured to input the fused feature vector into a plurality of independently trained deep learning models, and based on model differences among the plurality of deep learning models, weightedly fuse the output results of the plurality of deep learning models to obtain a predicted intention of the obstacle; a model building module for setting associated constraints between vehicle operating states and input controls based on the predicted intent, and establishing an avoidance strategy model with path length and time cost as dual optimization objectives; an avoidance strategy generation module, configured to continuously update the input control in a prediction time domain until the dual optimization objective value of the avoidance strategy model is iteratively minimized within associated constraints, and output a time series of the input control to constitute an avoidance strategy; The control instruction generation module is used to obtain the real-time operating status of the vehicle during the automatic driving process, and generate the final vehicle execution instruction in combination with the avoidance strategy and the real-time operating status.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for predicting and avoiding obstacle intentions in an autonomous driving manner as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Automatic driving obstacle recognition method and system based on deep learning
CN119580226A
Power grid safety production detection method and device, electronic equipment and storage medium
CN119886702A
Image abnormal state classification method and device, computer equipment and storage medium
CN119942206A
Autonomous navigation of a continuum robot
WO2025019373A1
Cited By
Automatic driving control method and system for mining truck
CN120922179A
Traffic incident detection system based on Leiyu fusion
CN120954195A
Real vehicle testing and human-like trajectory tracking method based on subjective control stability evaluation
CN121256962A
Photoelectric hybrid analog-to-digital conversion method and system
CN121283419A