An automatic driving control method and system for a mining truck

By using multimodal data fusion and temporal modeling techniques, combined with ground mechanical parameters, vehicle control commands are generated, which solves the problem of insufficient perception and adaptability in autonomous driving in mining areas, improves the vehicle's passability and safety on soft roads, and reduces operational risks.

CN120922179BActive Publication Date: 2026-01-13ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511467958.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-01-13
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing autonomous driving solutions cannot effectively perceive the mechanical properties of the ground in mining areas, causing vehicles to easily slip and get stuck on soft and rugged surfaces. Furthermore, they cannot adapt to the frequent changes in the mining environment and specific operational processes, resulting in low overall operational efficiency.

Method used

By employing multimodal data fusion and temporal modeling techniques, combined with ground mechanical parameters, lateral and longitudinal control commands for the vehicle are generated. These commands are then corrected using a terrain passability model to generate the final wheel torque and steering control values.

Benefits of technology

It improves vehicle passability and safety on unstructured roads, enables more accurate intelligent decision-making, reduces operational risks and rescue costs, and expands the operational range of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120922179B_ABST
    Figure CN120922179B_ABST
Patent Text Reader

Abstract

The application provides a mine truck automatic driving control method and system, and belongs to the technical field of automatic driving control. The method comprises the following steps: extracting feature vectors of each mode data after pre-processing of acquired multi-mode data; connecting the feature vectors of each mode data to form a fusion feature vector; inputting the fusion feature vectors of continuous multiple historical moments into a time sequence model for time sequence modeling; decoding the output of the time sequence model by a decoder to generate lateral and longitudinal control instructions of a vehicle; estimating ground mechanics parameters of a current driving area in real time to construct a terrain passability model; correcting the lateral control instructions and the longitudinal control instructions based on the terrain passability model to generate final wheel torque and steering control amount. Based on the method, a mine truck automatic driving control system is also provided. The application improves the passability and safety of an automatic driving vehicle on an unstructured road surface, and realizes more accurate and more scene-adaptive intelligent decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous driving control technology, and specifically relates to an autonomous driving control method and system for mining trucks. Background Technology

[0002] Autonomous driving technology in open-pit mines is of great significance for ensuring safe production and improving operational efficiency. However, the mining scenario still faces many challenges, including increased resource extraction depth, harsh working environments, high levels of danger, rising labor costs, and increasingly stringent regulations and environmental requirements in recent years.

[0003] Currently, most existing autonomous driving solutions are designed for structured roads, and their core technologies typically rely on high-precision maps, localization, and rule-based control logic. When these solutions are directly applied to mining areas, several limitations become apparent: First, traditional solutions lack awareness and consideration of ground mechanical properties, failing to adjust vehicle control strategies in real time based on parameters such as road surface bearing capacity and shear characteristics. This leads to risks such as vehicle slippage and getting stuck on soft, rough surfaces, resulting in poor maneuverability. Second, the mining environment changes drastically, and roads change frequently as mining progresses. Pre-built high-precision maps cannot be created, making solutions based on pre-set maps insufficiently adaptable. Third, most solutions focus on obstacle avoidance and path tracking, failing to deeply integrate with specific mining operations (such as precise parking in unloading areas and avoiding heavy equipment), resulting in low overall operational efficiency and an inability to cope with challenges posed by glare or low-light conditions. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes an autonomous driving control method and system for mining trucks. This significantly improves the vehicle's passability and safety on unstructured road surfaces, and enables more precise and scenario-appropriate intelligent decision-making.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] An automatic driving control method for mining trucks includes the following steps:

[0007] Collect multimodal data, which includes at least image data, lidar point clouds, and vehicle status information;

[0008] After preprocessing the multimodal data, feature vectors of each modality are extracted; the feature vectors of each modality are then concatenated to form a fused feature vector.

[0009] The fused feature vectors from multiple consecutive historical moments are input into a time-series model for time-series modeling. The output of the time-series model is decoded by a decoder to generate lateral and longitudinal control commands for the vehicle. The ground mechanical parameters of the current driving area are estimated in real time to construct a terrain passability model.

[0010] Based on the terrain passability model, the lateral and longitudinal control commands are modified to generate the final wheel torque and steering control values.

[0011] The present invention also proposes an automatic driving control system for mining trucks, comprising: a data acquisition module, a feature extraction module, a decoding module, and a control module;

[0012] The acquisition module is used to acquire multimodal data, which includes at least image data, lidar point clouds, and vehicle status information;

[0013] The feature extraction module is used to extract the feature vectors of each modality data after preprocessing the multimodal data; and to connect the feature vectors of each modality data to form a fused feature vector.

[0014] The decoding module is used to input the fused feature vectors from multiple consecutive historical moments into the time series model for time series modeling, decode the output of the time series model through the decoder to generate lateral control commands and longitudinal control commands for the vehicle; and estimate the ground mechanical parameters of the current driving area in real time to construct a terrain passability model.

[0015] The control module is used to modify the lateral control commands and longitudinal control commands based on the terrain passability model, and generate the final wheel torque and steering control values.

[0016] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects:

[0017] This invention proposes an autonomous driving control method and system for mining trucks, belonging to the field of autonomous driving control technology. The method includes the following steps: collecting multimodal data, which includes at least image data, LiDAR point clouds, and vehicle status information; preprocessing the multimodal data and extracting feature vectors for each modality; concatenating the feature vectors of each modality to form a fused feature vector; inputting the fused feature vector from multiple consecutive historical moments into a time-series model for time-series modeling; decoding the output of the time-series model to generate lateral and longitudinal control commands for the vehicle; estimating the ground mechanical parameters of the current driving area in real time to construct a terrain passability model; and correcting the lateral and longitudinal control commands based on the terrain passability model to generate the final wheel torque and steering control values. Based on this autonomous driving control method for mining trucks, an autonomous driving control system for mining trucks is also proposed. This invention can address the difficulty of scene recognition under strong light and low light conditions. This invention no longer blindly executes preset steering and throttle commands, but can "sense" the hardness and slipperiness of the ground. For example, when the system detects that it is on a soft, muddy road surface, it uses a model to predict the probability of passing through different paths, automatically selects the path with the highest passability, and adjusts torque distribution and slip ratio thresholds in advance to prevent the wheels from spinning and getting stuck. This greatly reduces operational risks, reduces rescue costs, and expands the range of operating conditions for autonomous driving systems.

[0018] This invention deeply integrates multimodal fusion, temporal modeling, and ground mechanics control to form a hierarchical perception-decision-control system. Multimodal fusion ensures the redundancy and robustness of environmental perception; temporal modeling enables the system to understand dynamic changes in the scene (such as slope trends); and finally, the ground mechanics model gives physical meaning to this perceived information and outputs control quantities that conform to the vehicle's dynamic characteristics. This makes the operation of autonomous vehicles not only safe but also smooth and energy-efficient. Attached Figure Description

[0019] Figure 1 This is a flowchart of an automatic driving control method for mining trucks proposed in Embodiment 1 of the present invention;

[0020] Figure 2 This is a schematic diagram of the sensor arrangement on the mining truck proposed in Embodiment 1 of the present invention;

[0021] Figure 3 This is an architecture diagram of the timing modeling output of the lateral and longitudinal control commands of the vehicle proposed in Embodiment 1 of the present invention;

[0022] Figure 4 The architecture diagram of the Long Short-Term Memory (LSTM) network proposed in Embodiment 1 of this invention is shown.

[0023] Figure 5 This is a schematic diagram of the structure of a single LSTM unit in the Long Short-Term Memory (LSTM) network proposed in Embodiment 1 of the present invention;

[0024] Figure 6 This is an architecture diagram of the Transformer decoder proposed in Embodiment 1 of the present invention;

[0025] Figure 7 This is an architecture diagram of the timing model GRU proposed in Embodiment 1 of the present invention;

[0026] Figure 8 This is a schematic diagram of the GRU unit proposed in Embodiment 1 of the present invention;

[0027] Figure 9 This is a framework diagram for capturing the feature relationships of different subspaces using multi-head attention, as proposed in Embodiment 1 of the present invention.

[0028] Figure 10 This is a framework diagram of scene analysis and driving intention analysis proposed in Embodiment 1 of the present invention;

[0029] Figure 11 This is a schematic diagram of an automatic driving control system for mining trucks proposed in Embodiment 2 of the present invention. Detailed Implementation

[0030] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.

[0031] Example 1

[0032] Embodiment 1 of the present invention proposes an automatic driving control method for mining trucks, which is used to solve the key technical problems faced by automatic driving of mining trucks in the prior art.

[0033] Figure 1 This is a flowchart of an automatic driving control method for mining trucks proposed in Embodiment 1 of the present invention;

[0034] In step S100, the process begins.

[0035] In step S110, multimodal data is collected, which includes at least image data, lidar point cloud, and vehicle status information.

[0036] Figure 2 This is a schematic diagram of the sensor arrangement on a mining truck according to Embodiment 1 of the present invention; a front camera and a rear camera are installed on the mining truck to collect image data from two directions; a lidar (LiDAR) is deployed on the top of the vehicle to collect lidar point cloud data.

[0037] Figure 3 This is an architecture diagram of the lateral and longitudinal control commands of the vehicle output by the time-series modeling proposed in Embodiment 1 of the present invention; a large language model is used to add semantic labels to the dataset, and the decision-making reasons and driving intentions are explained based on the information of each image and the decision-making information of human experts.

[0038] The preprocessing of multimodal data includes: scaling and normalizing pixel values ​​of images; denoising, random sampling, and coordinate standardization of LiDAR point clouds; and standardization of vehicle state information to eliminate dimensions.

[0039] In step S120, after preprocessing the multimodal data, feature vectors of each modality are extracted; the feature vectors of each modality are then concatenated to form a fused feature vector.

[0040] In this application, when extracting feature vectors of each modality data, a convolutional neural network is used to extract image feature vectors of the preprocessed image. Specifically, the preprocessed image is input into the convolutional neural network, and after convolution, batch normalization and ReLU activation, feature extraction is performed through residuals. An adaptive average pooling layer is introduced to downsample the feature map to a fixed size, and finally, a fully connected layer is used to compress and output a visual feature vector of a specified dimension.

[0041] In terms of visual feature extraction: In addition to ResNet-18, other backbone networks such as ResNet-34, EfficientNet, MobileNet or Vision Transformer can be used as alternatives, as long as they can effectively extract image features.

[0042] In image preprocessing: In addition to compressing the image to a resolution of 320×180, it can also be compressed to other resolutions, such as 224×224.

[0043] This application uses a one-dimensional convolutional network or a point cloud processing network to extract the feature vectors of LiDAR point clouds when extracting feature vectors from each modality of data. Specifically, the preprocessed LiDAR point cloud data is convolved into multiple layers to map the four-dimensional input (x, y, z, reflection intensity) to a feature space of a specified dimension. Each convolutional layer is followed by batch normalization and a GELU activation function. Finally, a global max pooling layer generates point cloud feature vectors of the specified dimension, effectively capturing the spatial distribution features of the point cloud. This invention uses channel-level feature extraction instead of spatial voxelization, which preserves the original geometric information and avoids the huge computational cost of 3D convolution. For point cloud processing, in addition to the 1D convolutional structure, PointNet++, Point Transformer, or other point cloud processing networks can be used as alternatives, as long as they can effectively process LiDAR point cloud data.

[0044] In this application, when extracting feature vectors from each modality of data, a fully connected layer is used to encode vehicle state feature vectors of vehicle state information. Specifically, the preprocessed vehicle state information is mapped to the hidden space through convolution, and a vehicle state feature vector of a specified dimension is output. This step can perceive the vehicle's motion state and use speed as one of the decision parameters. Figure 10 This is a framework diagram of scene analysis and driving intention analysis proposed in Embodiment 1 of the present invention; wherein, The input image is used. Regarding interpretability enhancement: In addition to Qwen-VL, other large-scale visual language models such as LLaVA and BLIP-2 can be used as alternatives, as long as they can effectively provide semantic explanations for driving decisions.

[0045] The feature vectors of each modality are concatenated end-to-end to form a longer vector:

[0046] V_fused=Concat(V_vision,V_lidar,V_state);

[0047] For example: V_vision (512 dimensions) + V_lidar (512 dimensions) + V_state (128 dimensions) = V_fused (1152 dimensions).

[0048] In step S130, the fused feature vectors from multiple consecutive historical moments are input into the time series model for time series modeling, and the output of the time series model is decoded by the decoder to generate the vehicle's lateral control command and longitudinal control command.

[0049] The multimodal fusion feature vectors from multiple consecutive historical moments are organized into a time-series sequence in chronological order;

[0050] The time series sequence is input into a time series model for time series modeling to capture long-term dependencies and dynamic change features in the sequence. The time series model includes a Long Short-Term Memory (LSTM) network, a gated recurrent unit (GRU), or a Transformer decoder.

[0051] Figure 4 This is an architecture diagram of the Long Short-Term Memory (LSTM) network proposed in Embodiment 1 of the present invention; wherein, For the current moment The input multimodal fusion feature vector; for The input multimodal fusion feature vector; for The input multimodal fusion feature vector; for The input multimodal fusion feature vector.

[0052] Figure 5 This is a schematic diagram of the structure of a single LSTM unit in the Long Short-Term Memory (LSTM) network proposed in Embodiment 1 of the present invention; in this application Figure 4 The publicly disclosed Long Short-Term Memory (LSTM) network includes four... Figure 5 The LSTM cell shown contains a forget gate. Input gate Output gate ;

[0053] ;

[0054] ;

[0055] ;

[0056] ;

[0057] ;

[0058] .

[0059] in, This is the multimodal fusion feature vector input at the current time. The current state vector is hidden, containing historical context information about the vehicle's environment and state from the past to the present. In a cellular state, information is stored for a long time; Candidate cell states are used to create new candidate values ​​that can be added to cell states; Use the Sigmoid activation function; It is the hyperbolic tangent activation function; This is the forget gate bias vector; The input gate bias vector; This is the candidate cell state bias vector; This is the output gate bias vector; This is the forget gate weight matrix; The input gate weight matrix; This is the candidate cell state weight matrix; This is the output gate weight matrix.

[0060] Figure 6 This is an architecture diagram of the Transformer decoder proposed in Embodiment 1 of the present invention; This is the multimodal fusion feature vector input at the current time. for The input multimodal fusion feature vector.

[0061] Figure 7 This is an architecture diagram of the timing model GRU proposed in Embodiment 1 of the present invention; Figure 7 middle, For the current moment The input multimodal fusion feature vector; for The input multimodal fusion feature vector; for The input multimodal fusion feature vector; for The input multimodal fusion feature vector.

[0062] Figure 8 This is a schematic diagram of the GRU unit proposed in Embodiment 1 of the present invention. The GRU unit includes an update gate. Reset the door ;

[0063] ;

[0064] ;

[0065] ;

[0066] ;

[0067] in, This is the multimodal fusion feature vector input at the current time. Output the state vector for the current moment, which includes both long-term and short-term information; This is the candidate hidden state, used to create new candidate values; Use the Sigmoid activation function; It is the hyperbolic tangent activation function; To update the weight matrix of the gate; To reset the weight matrix of the gate; Let be the weight matrix of the candidate hidden states.

[0068] This application extracts the temporal feature representation output by the temporal model to obtain a feature vector containing historical context information; the temporal feature representation is input into a decoder composed of fully connected layers for dimensionality reduction and decoding processing, and outputs a two-dimensional control vector; the two-dimensional control vector is parsed into the vehicle's lateral control command and longitudinal control command respectively, where the first dimension corresponds to the lateral control command, representing the steering control quantity of the steering wheel, and the second dimension corresponds to the longitudinal control command, representing the combined control quantity of the accelerator and brake.

[0069] The multimodal fusion feature vectors from multiple consecutive historical moments (e.g., 4 frames) are organized into a sequence in chronological order. The fusion feature vector at each moment (e.g., 2048 dimensions) contains the visual, point cloud, and vehicle state information of the current moment. This chronological sequence is represented as: [x(t-3), x(t-2), x(t-1), x(t)], where x(t) is the multimodal fusion feature vector input at the current moment.

[0070] The prepared time series is input into a time series model (such as LSTM, GRU or Transformer) for modeling.

[0071] LSTM / GRU processing: The recurrent neural network processes the feature vector at each time step sequentially, selectively retaining or discarding information through gating mechanisms to capture long-term dependencies in the sequence. The hidden state is propagated between time steps, and the final output is the hidden state of the last time step as a temporal feature representation.

[0072] Transformer processing: Figure 9 This is a framework diagram of multi-head attention for capturing feature relationships in different subspaces, as proposed in Embodiment 1 of the present invention. A self-attention mechanism is used to calculate the relevance weights of each position in the sequence to all other positions, processing the entire sequence in parallel. Multi-head attention (e.g., 4 heads) captures feature relationships in different subspaces, which are then processed by a feedforward neural network to output an enhanced feature representation for each time step. Figure 9 middle, This represents the multimodal fusion feature vector input at four time points. It is the set of the entire input feature sequence. This represents the enhanced feature vectors at four time points. This represents the set of the entire output feature sequence.

[0073] For LSTM / GRU: the output is the final hidden state (e.g., 1024-dimensional); for Transformer: the output of the last time step can be selected, or temporal average pooling can be used to obtain a comprehensive feature vector (e.g., 1024-dimensional). This feature vector contains historical context information and can reflect the trend of vehicle operation and the dynamic changes of the scene.

[0074] The feature vector output from the time-series model is input into the decoder (usually a fully connected neural network) for decoding: The first fully connected layer maps the 1024-dimensional feature vector to an intermediate dimension (e.g., 512-dimensional), using the ReLU activation function to introduce a non-linear transformation. The second fully connected layer further compresses the intermediate-dimensional features to a 2-dimensional output, corresponding to the horizontal and vertical control variables.

[0075] The 2D vectors output by the decoder correspond to: Lateral control commands: representing the steering wheel angle or steering angular velocity; positive values ​​indicate right turn; negative values ​​indicate left turn; the absolute value indicates the steering amplitude. Longitudinal control commands: representing the combined control quantity of the throttle and brake; positive values ​​indicate acceleration commands (throttle opening); negative values ​​indicate deceleration commands (brake intensity); the absolute value indicates the control intensity.

[0076] The longitudinal control command is based on a preset allocation strategy, which converts the scalar command into independent throttle opening signal and brake pressure signal; the allocation strategy includes, but is not limited to: a state machine switching strategy based on threshold and hysteresis zone, or an optimized allocation strategy based on the vehicle inverse dynamics model.

[0077] The generated raw control commands are post-processed: a low-pass filter is applied to eliminate command abrupt changes and ensure control smoothness; the command range is constrained according to the vehicle's physical limitations to prevent exceeding the actuator's capabilities; a dead zone is set near the zero value to avoid frequent minor movements of the actuator.

[0078] In step S140, the ground mechanical parameters of the current driving area are estimated in real time to construct a terrain passability model; this is used to map the data from perception sensors (cameras, lidar) and vehicle body sensors (wheel speed, torque) into parameters describing the mechanical properties of the ground, and finally calculate whether the vehicle can pass safely and how to pass.

[0079] Ground mechanical parameters, including soil bearing strength and shear properties, are estimated in real time by fusing wheel speed and slip ratio data from vehicle status information, as well as the identification results of ground material from images and LiDAR.

[0080] Visual data consists of images of the road ahead captured by a camera; point cloud data consists of 3D environmental point clouds provided by LiDAR.

[0081] Vehicle status data includes: wheel speed and slip ratio, drive torque, and vehicle attitude.

[0082] The speed and slip ratio of each wheel are obtained through the ABS wheel speed sensor. From the formula Calculation, where It is the angular velocity of the wheel. It is the wheel radius. The vehicle speed is estimated by GPS / IMU. Slip ratio is a key indicator reflecting the interaction between the wheels and the ground.

[0083] Drive torque is obtained from the CAN bus, which displays the drive torque output by the engine or motor.

[0084] Vehicle attitude, provided by the IMU: vehicle pitch angle (reflecting the gradient) and roll angle.

[0085] The process of semantic recognition of ground materials is as follows: Visual data is input into a pre-trained image classification CNN (such as ResNet). This network is trained to identify common mining area ground types, such as "compacted hard soil," "loose gravel," "mud," "rock," and "asphalt." It outputs a material category label L_terrain and its confidence probability P. Different materials have typical ranges of mechanical properties (e.g., concrete has a high adhesion coefficient, while mud has a low adhesion coefficient).

[0086] The process of terrain geometric feature extraction includes analyzing the preprocessed and segmented lidar ground point cloud to output key geometric features. Specifically, these include roughness, slope, and step height.

[0087] The process of back-deriving parameters based on vehicle dynamics response includes: combining the current vehicle's driving torque T with the measured slip ratio. As input, it is substituted into a simplified vehicle dynamics model for reverse engineering.

[0088] The terrain passability model is constructed and evaluated by inputting semantic information L_terrain, geometric features (roughness, slope), and mechanical parameters θ_terrain into a fusion module (such as a small fully connected neural network). The fusion module learns the correlations between different information sources. For example, when visually identified as "muddy," it tends to trust the low c and φ values ​​derived from the dynamic model; while the importance of geometric features decreases when vehicles are traveling on smooth, hard rock. The output is a set of revised and enhanced, more reliable integrated ground mechanical parameters θ_fused.

[0089] The process of passability prediction is as follows: input the comprehensive parameter θ_fused and the current vehicle state (weight, attitude) into the vehicle dynamics model to perform forward simulation prediction.

[0090] The predictions include: maximum adhesion, driving resistance, and predicted slip ratio curve; the output passability indicators include: passability probability, optimal slip ratio, and torque limit.

[0091] In step S150, the lateral control command and longitudinal control command are modified based on the terrain passability model to generate the final wheel torque and steering control amount.

[0092] The longitudinal control correction process involves adjusting the desired throttle or brake opening based on the predicted driving resistance. Torque limiting and optimal slip ratio are used as core parameters and delegated to the underlying controller (such as a slip ratio control system), so that it no longer pursues maximum acceleration but rather prioritizes "avoiding getting stuck" and "high efficiency." For example, on muddy roads, the system actively limits torque output, keeping the slip ratio near the optimal slip ratio.

[0093] The process of lateral control correction involves providing the path planner with a passability cost map. Path planning no longer considers only geometric feasibility (whether it is possible to bypass obstacles), but prioritizes paths with high passability and low driving resistance. Based on the predicted maximum lateral force, the vehicle speed during turns is limited to prevent sideslip.

[0094] In step S160, a large visual language model is used to semantically interpret the driving decision-making process, and the weight ratio of lateral control loss and longitudinal control loss in the training loss function is adjusted based on the interpretation results to optimize model training.

[0095] Multimodal scene information during autonomous driving system decision-making is input into a large visual language model to generate a semantic explanation of the current driving decision. The semantic explanation includes scene description, decision reason, driving intention, and risk identification.

[0096] The semantic interpretation is compared and evaluated with the interpretation annotated by human experts. Based on the comparison and evaluation results, the weight allocation of different control objectives in the model training loss function is dynamically adjusted, including adjusting the weights of the horizontal control loss, the vertical control loss, and the safety constraint loss according to the needs of the scenario.

[0097] The training samples are reweighted based on the quality and importance of the semantic interpretation, and the semantic interpretation is used as a reward signal for reinforcement learning to encourage the model to learn interpretable driving strategies that conform to human understanding.

[0098] Based on semantic interpretation, we identify systemic defects and erroneous decision-making patterns in autonomous driving models, establish a semantic interpretation knowledge base, and accumulate high-quality interpretations in different scenarios to guide the iterative optimization and training of subsequent models.

[0099] In step S170, the process ends.

[0100] The autonomous driving control method for mining trucks proposed in Embodiment 1 of this invention no longer blindly executes preset steering and throttle commands, but can instead "sensor" the softness and slipperiness of the ground. For example, when the system detects that it is currently on a soft, muddy road surface, it predicts the probability of passing through different paths through a model, automatically selects the path with the highest passability, and adjusts the torque distribution and slip rate threshold in advance to avoid wheel spin and getting stuck. This greatly reduces operational risks, reduces rescue costs, and expands the operating conditions that the autonomous driving system can operate under.

[0101] The autonomous driving control method for mining trucks proposed in Embodiment 1 of this invention deeply integrates multimodal fusion, temporal modeling, and ground mechanics control to form a hierarchical perception-decision-control system. Multimodal fusion ensures the redundancy and robustness of environmental perception; temporal modeling enables the system to understand the dynamic changes in the scene (such as the trend of slope change); finally, the ground mechanics model gives physical meaning to this perceived information and outputs control quantities that conform to the vehicle's dynamic characteristics. This makes the vehicle's operation not only safe but also smooth and energy-efficient.

[0102] Example 2

[0103] Based on the automatic driving control method for mining trucks proposed in Embodiment 1 of the present invention, Embodiment 2 of the present invention also proposes an automatic driving control system for mining trucks. Figure 11 This is a schematic diagram of an automatic driving control system for mining trucks according to Embodiment 2 of the present invention. The system includes: a data acquisition module, a feature extraction module, a decoding module, and a control module.

[0104] The acquisition module is used to acquire multimodal data, which includes at least image data, lidar point clouds, and vehicle status information;

[0105] The feature extraction module is used to extract the feature vectors of each modality after preprocessing the multimodal data; and to connect the feature vectors of each modality to form a fused feature vector.

[0106] The decoding module is used to input the fused feature vectors from multiple consecutive historical moments into the time series model for time series modeling, and decode the output of the time series model through the decoder to generate lateral control commands and longitudinal control commands for the vehicle; and to estimate the ground mechanical parameters of the current driving area in real time to construct a terrain passability model.

[0107] The control module is used to modify the lateral control commands and longitudinal control commands based on the terrain passability model, and generate the final wheel torque and steering control values.

[0108] The system also includes an optimization module. This module utilizes a large-scale visual language model to semantically interpret the driving decision-making process and adjusts the weight ratio of lateral control loss and longitudinal control loss in the training loss function based on the interpretation results to optimize model training. The specific execution process is as follows: Multimodal scene information during autonomous driving system decision-making is input into the large-scale visual language model to generate a semantic interpretation of the current driving decision. This semantic interpretation includes scene description, decision reasons, driving intention, and risk identification. The semantic interpretation is compared and evaluated with interpretations annotated by human experts. Based on the comparison and evaluation results, the weight allocation of different control objectives in the model training loss function is dynamically adjusted, including adjusting the weights of lateral control loss, longitudinal control loss, and safety constraint loss according to scene requirements. Training samples are reweighted based on the quality and importance of the semantic interpretation, using the semantic interpretation as a reward signal for reinforcement learning to encourage the model to learn interpretable driving strategies that conform to human understanding. Based on the semantic interpretation, systematic defects and erroneous decision-making patterns of the autonomous driving model are identified, a semantic interpretation knowledge base is established, and high-quality interpretations in different scenarios are accumulated to guide subsequent iterative optimization training of the model.

[0109] During the data acquisition process, the data acquisition module also collects human expert decision-making information and removes invalid data from the dataset. A large language model is used to add semantic labels to the dataset, interpreting the reasons for decisions and driving intentions based on the information from each image and the human expert decision-making information.

[0110] The preprocessing of multimodal data during the feature extraction module includes: scaling and normalizing pixel values ​​of the image; denoising, random sampling, and coordinate standardization of the LiDAR point cloud; and standardizing vehicle state information to eliminate dimensions.

[0111] When extracting feature vectors for each modality of data, a convolutional neural network is used to extract image feature vectors from the preprocessed image. Specifically, the preprocessed image is input into the convolutional neural network, and after convolution, batch normalization and ReLU activation, feature extraction is performed through residuals. An adaptive average pooling layer is introduced to downsample the feature map to a fixed size, and finally, a fully connected layer is used to compress and output a visual feature vector of a specified dimension.

[0112] When extracting feature vectors from each modality of data, a one-dimensional convolutional network or a point cloud processing network is used to extract the point cloud feature vectors of the LiDAR point cloud. Specifically, the preprocessed LiDAR point cloud data is convolved into multiple layers to map the four-dimensional input LiDAR point cloud data to a feature space of a specified dimension. Each convolutional layer is followed by batch normalization and the GELU activation function. Finally, a global max pooling layer generates the point cloud feature vector of the specified dimension.

[0113] When extracting feature vectors from each modality of data, a fully connected layer is used to encode the vehicle state feature vector of the vehicle state information; specifically, the preprocessed vehicle state information is mapped to the hidden space through convolution, and the vehicle state feature vector of the specified dimension is output.

[0114] In the decoding module, the fused feature vectors from multiple consecutive historical moments are input into a temporal model for temporal modeling. The output of the temporal model is decoded by a decoder to generate lateral and longitudinal control commands for the vehicle. Specifically, the multimodal fused feature vectors from multiple consecutive historical moments are organized into a temporal sequence in chronological order. The temporal sequence is input into the temporal model for temporal modeling to capture long-term dependencies and dynamic change features in the sequence. The temporal model includes a Long Short-Term Memory (LSTM) network, a Gated Recurrent Unit (GRU), or a Transformer decoder. The temporal feature representation output by the temporal model is extracted to obtain a feature vector containing historical context information. The temporal feature representation is input into a decoder composed of fully connected layers for dimensionality reduction and decoding, outputting a two-dimensional control vector. The two-dimensional control vector is parsed into lateral and longitudinal control commands for the vehicle, where the first dimension corresponds to the lateral control command, representing the steering control amount of the steering wheel, and the second dimension corresponds to the longitudinal control command, representing the combined control amount of the accelerator and brake.

[0115] The longitudinal control command is based on a preset allocation strategy, which converts the scalar command into independent throttle opening signal and brake pressure signal; the allocation strategy includes, but is not limited to: a state machine switching strategy based on threshold and hysteresis zone, or an optimized allocation strategy based on the vehicle inverse dynamics model.

[0116] Ground mechanical parameters include soil bearing strength and shear properties, which are estimated in real time by fusing wheel speed and slip ratio data from the vehicle status information, as well as the ground material identification results from the image and lidar.

[0117] In the control module, the longitudinal control correction process involves adjusting the desired throttle or brake opening based on the predicted driving resistance. Torque limiting and optimal slip ratio are used as core parameters and delegated to the underlying controller (such as a slip ratio control system), so that it no longer pursues maximum acceleration but rather prioritizes "avoiding getting stuck" and "high efficiency." For example, on muddy roads, the system actively limits torque output, keeping the slip ratio near the optimal slip ratio.

[0118] The process of lateral control correction involves providing the path planner with a passability cost map. Path planning no longer considers only geometric feasibility (whether it is possible to bypass obstacles), but prioritizes paths with high passability and low driving resistance. Based on the predicted maximum lateral force, the vehicle speed during turns is limited to prevent sideslip.

[0119] The autonomous driving control system for mining trucks proposed in Embodiment 2 of this invention no longer blindly executes preset steering and throttle commands, but can instead "sensor" the softness and slipperiness of the ground. For example, when the system detects that it is currently on a soft, muddy road surface, it predicts the probability of passing through different paths through a model, automatically selects the path with the highest passability, and adjusts the torque distribution and slip rate threshold in advance to prevent the wheels from spinning and getting stuck. This greatly reduces operational risks, reduces rescue costs, and expands the operating conditions that the autonomous driving system can operate under.

[0120] The autonomous driving control system for mining trucks proposed in Embodiment 2 of this invention deeply integrates multimodal fusion, temporal modeling, and ground mechanics control to form a hierarchical perception-decision-control system. Multimodal fusion ensures the redundancy and robustness of environmental perception; temporal modeling enables the system to understand dynamic changes in the scene (such as slope change trends); finally, the ground mechanics model gives physical meaning to this perceived information and outputs control quantities that conform to the vehicle's dynamic characteristics. This makes vehicle handling not only safe but also smooth and energy-efficient.

[0121] The description of the relevant parts of the mining truck automatic driving control system provided in Embodiment 2 of this application can be found in the detailed description of the corresponding parts of the mining truck automatic driving control method provided in Embodiment 1 of this application, and will not be repeated here.

[0122] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.

[0123] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for automatically driving a mine truck, characterized by The method comprises the following steps: Collecting multi-modal data, the multi-modal data at least comprising image data, laser radar point cloud and vehicle state information; Extracting feature vectors of each modality data after pre-processing the multi-modal data respectively; connecting the feature vectors of each modality data to form a fusion feature vector; Inputting the fusion feature vectors of continuous multiple historical time points into a time sequence model for time sequence modeling, decoding the output of the time sequence model by a decoder to generate lateral control instructions and longitudinal control instructions of the vehicle; and estimating ground mechanics parameters of a current driving area in real time to construct a terrain passability model; Inputting the fusion feature vectors of continuous multiple historical time points into a time sequence model for time sequence modeling, decoding the output of the time sequence model by a decoder to generate lateral control instructions and longitudinal control instructions of the vehicle, specifically: organizing the multi-modal fusion feature vectors of continuous multiple historical time points into a time sequence in chronological order; inputting the time sequence into the time sequence model for time sequence modeling to capture long-term dependence and dynamic change characteristics in the sequence, the time sequence model comprising a long short-term memory network (LSTM), a gated recurrent unit (GRU) or a Transformer decoder; extracting time sequence feature representations output by the time sequence model to obtain feature vectors containing historical context information; inputting the time sequence feature representations into a decoder composed of full connection layers for dimension reduction and decoding processing to output a two-dimensional control vector; and parsing the two-dimensional control vector into lateral control instructions and longitudinal control instructions of the vehicle respectively, wherein the first dimension corresponds to the lateral control instructions, representing a steering control amount of a steering wheel, and the second dimension corresponds to the longitudinal control instructions, representing a comprehensive control amount of the throttle and brake; The longitudinal control instructions are converted into independent throttle opening degree signals and brake pressure signals based on a preset distribution strategy; The distribution strategy comprises a state machine switching strategy based on a threshold and a hysteresis zone, or an optimal distribution strategy based on a vehicle inverse dynamics model; The lateral control instructions and the longitudinal control instructions are corrected based on the terrain passability model to generate final wheel torque and steering control amount.

2. The automatic driving control method of a mine truck according to claim 1, characterized in that, The method further comprises: Performing semantic interpretation on the driving decision process by using a large visual language model; and adjusting the weight proportion of the lateral control loss and the longitudinal control loss in the training loss function based on the interpretation result to optimize the model training.

3. The method of claim 1, wherein, The pre-processing process of the multi-modal data comprises: scaling and pixel value normalization of the image, noise reduction, random sampling and coordinate standardization of the laser radar point cloud, and standardization of the vehicle state information to eliminate dimension; and adding semantic labels of interpretation decision reasons and driving intentions to the image frames in the training data set by using a large visual language model.

4. The method of claim 1, wherein, When extracting the feature vectors of each modality data, a convolutional neural network is used to extract the image feature vectors of the pre-processed image, specifically: inputting the pre-processed picture into the convolutional neural network, performing convolution, batch normalization and ReLU activation, and then performing feature extraction through residual; introducing an adaptive average pooling layer to downsample the feature map to a fixed size, and finally compressing the visual feature vector of a specified dimension through a full connection layer.

5. The method of claim 1, wherein, The one-dimensional convolution network or the point cloud processing network is used to extract the point cloud feature vector of the laser radar point cloud when extracting the feature vectors of the modal data, specifically: The preprocessed laser radar point cloud data is multi-layer convoluted, and the four-dimensional input laser radar point cloud data is mapped to a feature space of a specified dimension; a batch normalization and a GELU activation function are connected after each convolution layer; and a global maximum pooling layer generates a point cloud feature vector of a specified dimension.

6. The method of claim 1, wherein, The vehicle state feature vector of the vehicle state information is encoded by using a fully connected layer when extracting the feature vectors of the modal data, specifically: the preprocessed vehicle state information is mapped to a hidden space by convolution, and a vehicle state feature vector of a specified dimension is output.

7. The method of claim 2, wherein, A large visual language model is used to perform semantic interpretation on the driving decision-making process, and the weight ratio of the lateral control loss and the longitudinal control loss in the training loss function is adjusted based on the interpretation result to optimize the model training, specifically: The multi-modal scene information during the decision-making of the autonomous driving system is input into the large visual language model to generate a semantic interpretation of the current driving decision, which includes scene description, decision reason, driving intention and risk identification; The semantic interpretation is compared and evaluated with the interpretation annotated by human experts, and the weight distribution of different control targets in the model training loss function is dynamically adjusted based on the comparison and evaluation result, including adjusting the lateral control loss weight, the longitudinal control loss weight and the safety constraint loss weight according to the scene demand; The training samples are reweighted according to the quality and importance of the semantic interpretation, the semantic interpretation is used as a reward signal for reinforcement learning, and the model is encouraged to learn an interpretable driving strategy consistent with human understanding; Based on the semantic interpretation, the systematic defects and erroneous decision-making patterns of the autonomous driving model are identified, a semantic interpretation knowledge base is established, and high-quality interpretations in different scenes are accumulated to guide the iterative optimization training of subsequent models.

8. The method of claim 1, wherein, The ground mechanics parameters include soil bearing strength and shear characteristics, which are estimated in real time by fusing the wheel speed and slip rate data in the vehicle state information and the identification results of the ground material by the image and laser radar.

9. A mining truck autonomous driving control system for performing a mining truck autonomous driving control method according to any one of claims 1 to 8, characterized in that, It comprises: a collection module, a feature extraction module, a decoding module and a control module; The collection module is used to collect multi-modal data, and the multi-modal data at least includes image data, laser radar point cloud and vehicle state information; The feature extraction module is used to extract feature vectors of each modal data after preprocessing the multi-modal data; the feature vectors of each modal data are connected to form a fusion feature vector; The decoding module is used to input the fusion feature vectors of continuous multiple historical time points into a time series model for time series modeling, decode the output of the time series model by a decoder, generate lateral control instructions and longitudinal control instructions of the vehicle, and estimate ground mechanics parameters of the current driving area in real time to construct a terrain passability model; The control module is used to correct the lateral control instructions and the longitudinal control instructions based on the terrain passability model to generate the final wheel torque and steering control amount.

Citation Information

Patent Citations

  • Off-road vehicle path tracking stability control method used in complex terrain

    CN118192344A

  • Vehicle control method and device, vehicle and computer readable storage medium

    CN118419067A