Vehicle driving intention recognition method and system

By combining multi-view visual information fusion and causal reasoning technology with progressive group convolution and hypergraph neural network, the problems of insufficient robustness and accuracy of driver intention recognition in existing technologies are solved, and more accurate driving intention recognition is achieved.

CN120635655AActive Publication Date: 2025-09-12XI AN JIAOTONG UNIV

Patent Information

Application Number
CN202511126984.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-09-12
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing driving intention recognition methods fail to comprehensively consider the correlation between driver behavior and the external environment, and ignore the high-order synergy between multiple factors such as the driver's facial state, body movements, and vehicle driving environment, resulting in insufficient robustness and accuracy of intention recognition.

Method used

By collecting and fusing multi-perspective visual information, combining cross-perspective consistency contrastive learning and causal reasoning technology, a progressive group convolutional neural network (PGCNN) is used to extract features, a hypergraph structure is established to model causal hyperedge relationships, hypergraph convolution is used for intention reasoning, and a hypergraph neural network guided by prior knowledge (CI-HGCN) is used to identify driver intentions.

Benefits of technology

It achieves more accurate and comprehensive understanding and prediction of driving intentions, and improves the accuracy and robustness of driving intention recognition, especially the recognition accuracy in scenarios such as lane changing, turning, going straight, and deceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635655A_ABST
    Figure CN120635655A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle driving intention recognition method and system, and relates to the technical field of intelligent driving, and the method comprises the steps: collecting the front image and side image data of a to-be-detected driver and the forward image data of a vehicle; performing feature extraction on the multi-view image data through a hierarchical grouping strategy of a progressive grouping convolutional neural network; optimizing the features of the extracted multi-view-angle image data by adopting cross-view-angle consistency learning; respectively calculating state information corresponding to each feature; inputting each piece of state information into a causal reasoning hypergraph neural network based on priori knowledge guidance, extracting node features of each piece of state information, performing hypergraph construction based on priori knowledge, updating the node features through a hypergraph convolutional layer, and outputting a driving intention category; according to the method, the driving intention of the driver can be more fully captured and understood by fusing the visual information inside and outside the vehicle cockpit and high-order correlation causal reasoning, and the driving intention recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent driving technology, and in particular to a vehicle driving intention recognition method and system. Background Art

[0002] With the rapid development of intelligent driving technology, human-machine co-driving has become the mainstream model of intelligent driving in the foreseeable future. In this scenario, accurately identifying the driver's driving intention is crucial to ensuring driving safety and improving the driving experience.

[0003] Existing driving intention recognition methods often rely on vehicle-based signal sensors (such as steering wheel angle, accelerator pedal position, and brake signals) to estimate driving intention by analyzing the vehicle's historical driving status. However, these methods primarily focus on vehicle status and ignore key behavioral factors such as driver behavior and status, resulting in a less accurate and in-depth understanding of driving intention.

[0004] In recent years, some studies have introduced cameras facing the driver's front to assess fatigue by monitoring their facial features, providing supplementary information about driving intent. While this method can reflect the driver's state to a certain extent, it only provides simple fatigue information and struggles to accurately infer and deeply understand complex driving intent. This results in low robustness and accuracy in driver intent recognition in scenarios such as lane changes, steering, driving straight, and deceleration. Summary of the Invention

[0005] In view of the fact that existing technologies are unable to comprehensively consider the correlation between driver behavior and the external environment, and also ignore the high-order synergy between multiple factors such as the driver's facial state, body movements, and vehicle driving environment, resulting in insufficient robustness and accuracy in intention recognition, the present invention proposes a vehicle driving intention recognition method and system, which achieves accurate recognition of the driver's intention through the collection and fusion of multi-perspective visual information, combined with cross-perspective consistency comparison learning and causal reasoning technology, thereby solving the problems existing in the existing technology.

[0006] A method for identifying vehicle driving intention comprises the following steps: Collect the front image, side image and forward image data of the driver to be tested and the vehicle he is driving; Feature extraction is performed on the frontal and side images of the driver and the forward image data of the vehicle being driven; the extracted facial and full-body image features of the driver and the image features of the environment in front of the vehicle are aligned and optimized; and the corresponding state information is calculated based on the optimized image features; the state information includes the driver's head posture, line of sight, fatigue state, body movements, and the environment outside the vehicle being driven; Map each state information into a feature vector, and establish high-order causal hyperedge relationships between each state information based on prior knowledge. Build a hypergraph structure using feature vectors as nodes and high-order causal hyperedge relationships between each state information as edges. Based on the hypergraph structure, reason about the evolution of the feature vector corresponding to each state information and update the nodes. After the update, concatenate and fuse the features corresponding to all nodes to generate a driving scene representation vector. The driving intention category is identified based on the driving scene representation vector.

[0007] Furthermore, a hierarchical grouping strategy of the progressive group convolutional neural network (PGCNN) is used to extract features from the front image, side image, and forward image data of the driver to be tested, including the following steps: The input front and side images of the driver and the forward image data of the vehicle are subjected to progressive group convolution calculation; wherein, in the first layer, the input image data is subjected to convolution operation to obtain the feature map of the image data ; in the layer, the output feature map of the previous layer Divided along the channel dimension Group, for each group of feature maps Perform convolution operations independently, expressed as: ; in, For the The convolution kernel of the group; The feature map of each group output Splice and get the feature map ; Use ReLU activation function and batch normalization to transform the feature map Process and extract the feature vector of each image data , which includes the frontal image feature vector , silhouette feature vector and the forward image feature vector .

[0008] Furthermore, cross-view consistency contrast learning is used to align and optimize the extracted driver's facial image features, full-body image features, and vehicle front environment image features, specifically including the following steps: According to the feature vectors of the frontal image, side image and forward image Construct positive and negative samples; in a training batch, the feature pairs of the same scene are considered as positive sample pairs, which are close to each other in the feature space; feature pairs of different scenes and are regarded as negative sample pairs, which are far away from each other in the feature space; Calculate eigenvectors Consistency loss function for: ; ; in, is the cosine similarity, is the temperature coefficient, and is the index number of the sample in a data batch during training, represents the number of samples in the batch, represents a positive sample pair; represents a negative sample pair; By optimizing the target The PGCNN weight parameters are optimized with back-propagation calculation so that the feature vectors of the same scene are close to each other in the feature space, and the optimized image features are obtained.

[0009] Furthermore, the step of calculating the state information corresponding to each optimized image feature includes the following steps: The optimized frontal image feature vector , silhouette feature vector and the forward image feature vector Input them into the fully connected neural network FCNN respectively and calculate their corresponding states: ; in, Indicates the status of head posture, gaze direction, fatigue, body movements and the external environment: Indicates head posture, Indicates the direction of sight, Indicates fatigue state, Indicates body movements, Indicates the environment outside the vehicle.

[0010] Furthermore, the method of mapping each state information into a feature vector, establishing a high-order causal hyperedge relationship between each state information based on prior knowledge, and constructing a hypergraph structure using the feature vector as a node and the high-order causal hyperedge relationship between each state information as an edge specifically includes the following steps: Head posture , sight direction , fatigue state , body movements 、External environment Input into the causal reasoning-based hypergraph neural network CI-HGCN to extract the node features of each state information; Based on prior knowledge of driving behavior, we define high-order causal hyperedge relationships between each state information, including: hyperedges connecting head posture, gaze direction, and body movements to model lane change or steering intentions; hyperedges connecting fatigue state, body movements, and the external environment to model braking intentions; and hyperedges connecting gaze direction, body movements, and the external environment to model acceleration or deceleration intentions. According to the node characteristics and high-order causal hyperedge relationships, the hypergraph structure is constructed using the hyperedge indicator matrix, which is expressed as: .

[0011] Furthermore, the method of inferring the evolution of the feature vector corresponding to each state information based on the hypergraph structure and updating the nodes; concatenating and fusing the features corresponding to all updated nodes to generate a driving scene representation vector; specifically includes the following steps: Based on the hypergraph structure, the node features are updated by hypergraph convolutional reasoning based on the evolution of node features. The update process is expressed as: ; in: is a node The updated features of is the hyperedge weight, is the super-edge degree, Indicates the index number of the node, is the total number of nodes, is the node degree, is the weight matrix of the hypergraph convolutional layer, is the ReLU activation function, represents the index number of the hyperedge, represents the number of hyperedges; The updated features of all nodes are concatenated and fused to generate a driving scene representation vector.

[0012] Furthermore, it also includes supervising the learning optimization of PGCNN by establishing an intention estimation loss function, and driving the update of PGCNN parameters by stimulating the output of the PGCNN model to be consistent with the actual situation; the intention estimation loss function is expressed as: ; in, is the true intention label, is the predicted probability, is the correction coefficient of the regularization term in the loss function, is the L1 norm of the weight parameter, Indicates the sample index number, represents the number of samples used for training, Indicates the category index number, Indicates the number of categories contained in the training data.

[0013] Furthermore, the acquisition of the front image and side image of the driver to be tested and the forward image data of the vehicle he is driving specifically includes the following steps: Visual sensors are used to collect the driver's front image, side image and vehicle forward image respectively; the front image data includes the driver's facial image, head posture and line of sight direction image; the side image includes the driver's upper body movement image; the vehicle forward image includes video images of road conditions, traffic signals and obstacle distribution.

[0014] Furthermore, after collecting the front image and side image of the driver to be tested and the forward image data of the vehicle he is driving, the front image, side image and forward image data of the driver to be tested are preprocessed, and the process specifically includes the following steps: Gaussian filtering is used to denoise the image data; Perform contrast stretching on the denoised image data; Align the timestamps of each processed image data; The image data after timestamp alignment is resampled by bilinear interpolation to obtain the processed image data.

[0015] The present invention also includes a vehicle driving intention recognition system, comprising: An acquisition module is used to acquire the front image and side image of the driver to be tested and the forward image data of the vehicle he is driving; The state calculation module is used to extract features from the front and side images of the driver and the forward image data of the vehicle being driven; align and optimize the extracted facial and full-body image features of the driver and the image features of the environment in front of the vehicle; and calculate the corresponding state information based on the optimized image features; the state information includes the driver's head posture, line of sight, fatigue state, body movements, and the environment outside the vehicle being driven; The hypergraph structure construction module is used to map each state information into a feature vector and establish high-order causal hyperedge relationships between each state information based on prior knowledge. The hypergraph structure is constructed using feature vectors as nodes and high-order causal hyperedge relationships between each state information as edges. The hypergraph structure is used to infer the evolution of the feature vector corresponding to each state information and update the nodes. The features corresponding to all updated nodes are spliced ​​and fused to generate a driving scene representation vector. The recognition module is used to identify the driving intention category based on the driving scene representation vector.

[0016] The present invention provides a method for identifying vehicle driving intention, which has the following beneficial effects: The present invention proposes a driving intention recognition method based on multi-perspective visual information fusion perception based on the driver's facial information, the driver's body movement information, and the vehicle's forward environmental information, achieving a more accurate and comprehensive understanding and prediction of driving intentions; by aligning and optimizing the visual features of multiple perspectives, establishing associations for cross-perspective feature extraction, improving the multi-perspective visual information fusion effect, and improving the accuracy of driving intention recognition. At the same time, a causal reasoning hypergraph neural network guided by prior knowledge is proposed, which improves the rationality and accuracy of intention reasoning through high-order associations of various states guided by human cognition. This method can more fully capture, understand, and infer the driver's driving intentions through visual information inside and outside the vehicle cockpit. The accuracy of driver intention recognition in scenarios such as lane changing, turning, going straight, and deceleration is significantly improved compared to existing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A schematic diagram of the structure of a vehicle driving intention recognition system in an embodiment of the present invention; Figure 2 This is a schematic diagram of the operation flow of the vehicle driving intention recognition system in an embodiment of the present invention; Figure 3 This is a flowchart of a multi-perspective visual fusion perception driving intention recognition method in an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0019] This paper proposes a vehicle driving intention recognition method. By collecting and fusing multi-perspective visual information, combined with cross-perspective consistency comparative learning and causal reasoning technology, it achieves an accurate and in-depth understanding of the driver's intention, providing reliable support for intelligent driving decision-making.

[0020] like Figure 3 As shown, the method specifically includes the following steps: S1. Data Collection and Annotation. Build a basic dataset for model training. Deploy multi-view sensors in a real-world driving environment to capture driver frontal images, side images, and vehicle forward images.

[0021] (1) Front vision sensor: installed above the steering wheel, using a 1080p high-resolution camera, equipped with infrared fill light, with a frame rate of 120 frames per second, to collect the driver's facial image, covering head posture, line of sight direction and other information.

[0022] (2) Side vision sensor: installed on the side of the driver's seat, using a wide-angle camera to collect the driver's upper body movements, such as holding the steering wheel, shifting gears, etc.

[0023] (3) Forward vision sensor: installed on the front of the vehicle, it uses a monocular camera combined with a lidar to collect road conditions (such as lane lines), traffic signals (such as traffic light status) and obstacle distribution (such as the distance to the vehicle ahead).

[0024] The collection process covers a variety of scenes such as urban roads and highways, as well as different conditions such as sunny days, rainy days, daytime and nighttime. The duration of each set of image data is no less than 10 seconds, and the frame rate is set to 120 frames per second. Professionals manually annotate the images, record driving intentions, such as "turn left" or "slow down", and annotate related states, such as line of sight and body movements. This forms a data set containing 5,000 sets of samples, 80% of which are used for training, 10% for verification, and 10% for testing. The collected images include front ,side , forward , the marking intention is recorded as , for example, “turn left”, the dataset size is 5000 groups.

[0025] S2. Data preprocessing: (1) The original image is preprocessed, including denoising (using Gaussian filtering, standard deviation σ =1.5), image enhancement (contrast stretching, enhancement factor 1.2) and time synchronization (ensuring the alignment of timestamps of the three-view images with an error of <10ms).

[0026] (2) The original image is resampled to a size of 224×224 suitable for the network through bilinear interpolation.

[0027] (3) The annotation data is stored in JSON format, including the image path, intent label (such as "turn left") and state label (such as line of sight direction).

[0028] S3. Feature Extraction: A progressive grouped convolutional neural network (PGCNN) is used to extract key features from multi-view images for subsequent analysis. Frontal, side, and frontal images are fed into the PGCNN, which converts them into feature vectors that allow the computer to understand the image content.

[0029] PGCNN adopts a hierarchical grouping strategy. In the shallow layers, more groups are used, for example, 8 groups, focusing on low-order sparse features such as edges and textures to capture local details. In the deep layers, the number of groups is gradually reduced, for example, 1 group, to extract high-order dense features such as overall shape and scene semantics, achieving feature representation from local to global: (1) PGCNN network structure: PGCNN contains 5 convolutional layers, each layer uses different group convolutions, and the group density changes with the depth. The specific structure is as follows: Layer 1 (input layer): number of packets G = 1, convolution kernel size 3×3, number of output channels 32, stride 2, padding 1. The input is the preprocessed image (resolution 224×224×3), and the output feature map size is 112×112×32.

[0030] Layer 2: Number of groups G =8, convolution kernel size 3×3, number of output channels 64, stride 1, padding 1, output feature map size 56×56×64.

[0031] Layer 3: Number of Groups G =4, convolution kernel size 3×3, number of output channels 128, stride 2, padding 1, output feature map size 28×28×128.

[0032] Layer 4: Number of Groups G =2, convolution kernel size 3×3, number of output channels 256, stride 2, padding 1, output feature map size 14×14×256.

[0033] Layer 5: Number of Groups G =1, convolution kernel size 3×3, number of output channels 256, stride 1, padding 1, output feature map size 7×7×512.

[0034] Global average pooling: Pool the 5th layer feature map into a 1×1×512 feature vector, expressed as ,in Indicates the viewing angle (front, side, or forward).

[0035] (2) Calculation process: Input: front, side, and forward images (resolution normalized to 224×224×3).

[0036] Grouped convolution calculation: In order to establish efficient feature calculation from shallow to deep layers, the network uses different numbers of groups from shallow to deep layers. The number of groups in shallow layers is large, and the number of groups gradually decreases as the layers go deeper. Layer, input feature map Divided into groups, each group performs convolution operation independently, the formula is: ; in, For the The convolution kernel of the group outputs the feature map Splice to .

[0037] Activation and normalization: In each layer, after convolution calculation, ReLU activation function and batch normalization are applied to process the feature map to obtain the activated and normalized feature map.

[0038] Output: Each view image generates a 512-dimensional feature vector , a total of three eigenvectors (positive ,side , forward ).

[0039] S4. Cross-view consistency contrastive learning: Through contrastive learning, we ensure that the features extracted from different viewpoints in the same scene can match each other while staying away from the features of other scenes, thereby improving the fusion effect of multi-view information.

[0040] In a training batch, the front, side, and forward features of the same scene are considered positive pairs, requiring them to be as close as possible in feature space. Features from different scenes are considered negative pairs, requiring them to be as far apart as possible. This process is achieved through a contrastive loss function, training the model to enhance feature alignment and differentiation, ensuring that the three perspectives can collaboratively express driving intent.

[0041] The calculation process includes: Input: Feature vector extracted by PGCNN , corresponding to the front, side, and forward perspectives of the same scene.

[0042] Positive and negative sample pair construction: In a training batch (batch size ), the feature pairs of the same scene are regarded as positive sample pairs, which are required to be close in the feature space; feature pairs of different scenes (such as , ) are regarded as negative sample pairs and are required to be far away from each other in the feature space.

[0043] Consistency loss function: In order to avoid feature association errors caused by factors such as background noise when extracting features from multi-view images, a consistency loss function is used to constrain the joint learning optimization process. The calculation formula of the consistency loss function is: ; ; in, is the cosine similarity, is the temperature coefficient, and is the index number of the sample in a data batch during training, represents the number of samples in the batch, represents a positive sample pair; represents a negative sample pair.

[0044] Feature alignment optimization: By optimizing the target The weight parameters of the neural network are optimized with back-propagation calculations to make the feature vectors of the same scene close in the feature space and enhance cross-view consistency.

[0045] S5. Behavior, State, and Environment Identification: Extracted feature vectors are used to estimate the driver's key states, providing a basis for intent recognition. A fully connected neural network (FCNN) processes these features to calculate information such as head pose, gaze direction, fatigue status, body movements, and the external environment. Head pose reflects head rotation angle, gaze direction indicates eye focus, fatigue status determines alertness or fatigue, body movements identify hand movements, and the external environment assesses whether the road is clear. These state results are compared with ground truth annotations, and the model is optimized to improve estimation accuracy.

[0046] (1) FCNN network structure: FCNN contains two fully connected layers: Layer 1: input dimension 512, output dimension 256.

[0047] Layer 2: Input dimension is 256, and output dimension is determined by label dimension.

[0048] (2) Calculation process: enter: Enter positive features , Input side features , Input forward features .

[0049] State estimation: Each FCNN processes its input features independently and calculates the corresponding state. The formula is: ; in, Indicates the status of each aspect: (Head posture, determined by output), (The direction of sight, given by output), (Fatigue state, caused by output), (Body movements, by output), (External environment, by output).

[0050] During the model training phase, based on the model’s predicted output and the labeled true value, the loss function of the FCNN output prediction and true value corresponding to each feature vector is calculated, and a weighted sum is performed to obtain the total loss: ; in: are the angular error losses between the head pose and gaze direction predicted by the model and the true value; are the cross entropy losses between the fatigue state, limb movement, and external environment predicted by the model and the true value; 、 are the weights corresponding to their respective states; In the inference phase, the head pose predicted by the model is , sight direction , fatigue state , body movements 、External environment as output variable.

[0051] S6. Intention Reasoning: By integrating all state information, the present invention implements driver intention reasoning through a causal inference hypergraph convolutional network (CI-HGCN). The CI-HGCN analyzes the interactions between states such as head posture and gaze direction. For example, turning the head and looking to the left may indicate an intention to turn left. The network treats these states as nodes, constructs a relationship network, and after multiple layers of calculation, outputs the driving intention, such as "turn left" or "slow down." This is then compared and optimized with the actual intention. Traditional driving intention reasoning methods are based solely on point-to-point relationship modeling and cannot effectively capture the joint causal influences between multiple state variables. They also lack domain knowledge guidance, making causal structure learning susceptible to noise and inferential stability. Furthermore, the limited expressive power of node relationships makes it difficult to explain the multi-factor linkage mechanisms underlying complex driving behavior. To address these issues, CI-HGCN incorporates prior knowledge of human cognition, constructs a hypergraph based on the prior relationships between various factors, and then performs reasoning calculations through hypergraph convolution to accurately estimate the driver's intention. The specific implementation method is as follows: (1) Driver status node initialization: CI-HGCN treats states (head pose, gaze direction, fatigue status, body movements, and external environment) as nodes in a graph, with edges between nodes representing causal relationships. The network consists of three graph convolutional layers, specifically modeled as follows:

[0052] Node initialization: each state Mapped to 32-dimensional node features , through the fully connected layer: ; in, is a node The initial feature representation of 、 are learnable parameters, It is an activation function that ensures nonlinear feature mapping. By uniformly mapping node features, it ensures that different state information can be associated and modeled in the same feature space, improving the expression consistency of multi-source information fusion.

[0053] (2) Constructing high-order hypergraph structures based on prior knowledge: Traditional graph modeling can only represent the relationship between two nodes, and cannot express the real situation where multiple states jointly affect a driving intention. To this end, the present invention is based on the knowledge of the driving behavior field, based on the initialization node obtained in the previous step, and based on the prior knowledge of the driving behavior field to define the high-order causal hyperedge relationship between each state information, which includes: connecting the head posture, line of sight direction, and body movement to model the intention to change lanes or turn, connecting the fatigue state, body movement, and the external environment to model the intention to brake, and connecting the line of sight direction, body movement, and the external environment to model the intention to accelerate or decelerate: hyperedge :connect , used to model lane change or turning intention; hyperedge :connect , used to model the braking intention caused by fatigue driving; hyperedge :connect , used to model intentions such as acceleration and deceleration.

[0054] According to the node characteristics and high-order causal hyperedge relationships, the hypergraph structure is constructed using the hyperedge indicator matrix, which is expressed as: .

[0055] Through hypergraph modeling, we can explicitly capture the complex joint causal relationships between multiple variables, make up for the limitation of traditional causal graphs that can only describe binary relationships, and improve the realism and accuracy of causal modeling.

[0056] (3) Hypergraph convolutional reasoning node feature update: Based on the hypergraph structure, the node features are updated by hypergraph convolutional reasoning based on the evolution of node features. The update process is expressed as: ; in: is a node Updated features, is the hyperedge weight, is the super-edge degree, Indicates the index number of the node, is the total number of nodes, is the node degree, is the weight matrix of the hypergraph convolutional layer, is the ReLU activation function, represents the index number of the hyperedge, Represents the number of hyperedges.

[0057] This hypergraph convolution enables the joint aggregation of node features across multiple nodes within a hyperedge, enabling nodes to obtain information from high-order causal relationships during reasoning. Through hypergraph convolution, information can be jointly propagated across multiple state nodes, capturing complex causal links and improving the accuracy of intent reasoning and the physical plausibility of reasoning paths.

[0058] (4) Driving intention output: After completing several layers of hypergraph convolution, the updated features of all nodes are fused to form a unified driving scene representation. The fusion method is feature splicing: ; Then, the driving intention classification prediction is performed through the fully connected layer and the Softmax function, and the intention category (such as "turn left", "turn right", "go straight", "slow down") is output: ; in and are the weight and bias parameters of the fully connected layer respectively; The above calculation uniformly encodes the state features after causal reasoning into an expression of driving intention, completing the closed loop from state recognition to intention reasoning.

[0059] Loss function: To accurately understand the driver's intention, we establish an intention estimation loss function to supervise the learning optimization of the overall model. This function drives the update of model parameters by encouraging the model to output intention results that are consistent with the actual situation. This loss function combines cross-entropy loss and L1 regularization loss and is calculated as follows: ; in, is the true intention label (such as "go straight", "turn left", "turn right", "change lane left", "change lane right"), is the predicted probability, is the correction coefficient of the regularization term in the loss function, is the L1 norm of the weight parameter, Indicates the sample index number, represents the number of samples used for training, Indicates the category index number, Indicates the number of categories contained in the training data. The accuracy of intent estimation is ensured through cross-entropy loss, and regularized sparsity constraints are used to prevent the hypergraph structure from becoming complex, maintaining the generalization performance and interpretability of the inference network.

[0060] This paper proposes a driving intention recognition method based on multi-view visual information fusion perception based on the driver's facial information, driver body movements, and the vehicle's forward environment. This method achieves more accurate and comprehensive understanding and prediction of driving intention. By aligning and optimizing visual features from multiple perspectives and establishing cross-view feature extraction correlations, this method enhances the multi-view visual information fusion effect and improves the accuracy of driving intention recognition. A visual feature extraction network based on progressive group convolution is proposed. This network uses sparse group convolutions in the shallow layers of the convolutional neural network and sparser group convolutions in the deep layers. This allows the network to improve the quality of visual information extraction while maintaining computational efficiency, thereby enhancing the accuracy of driving intention recognition. A causal reasoning hypergraph neural network guided by prior knowledge is also proposed. This method improves the rationality and accuracy of intention inference by leveraging high-order correlations between various states guided by human cognition. This method can more fully capture, understand, and infer the driver's driving intention using visual information from both inside and outside the vehicle cockpit. It significantly improves the accuracy of driver intention recognition in scenarios such as lane changing, turning, going straight, and deceleration compared to existing methods.

[0061] The method based on hypergraph modeling and hypergraph convolutional reasoning proposed in this invention solves the problems of insufficient causal modeling and lack of multi-variable joint reasoning in existing driving intention reasoning, significantly improves the accuracy of driving intention recognition and the interpretability of the reasoning process, and shows higher robustness and stability in complex driving scenarios (such as lane changing, turning, straight driving, and deceleration).

[0062] Example: In an intelligent driving vehicle, a frontal vision sensor is mounted above the steering wheel, a side vision sensor is mounted to the side of the driver's seat, and a forward vision sensor is mounted at the front of the vehicle. The graphics computing unit uses a high-performance graphics processor, such as the HUAWEI MDC610 or NVIDIA Orin X, to run a trained driver intention recognition model. The system collects multi-view image data in real time and outputs driving intention using the proposed method.

[0063] To train the model, we collected approximately 5,000 sets of image data representing driving scenarios. Each set included frontal, side, and forward images, with intents labeled as "turn left" or "turn right." We used PGCNN to extract features, and optimized the model using a combination of consistency and classification loss functions. After training, the model's performance was verified on a validation set. Once it met performance requirements, it was ready for deployment.

[0064] The present invention proposes a method and system for recognizing driving intentions through multi-perspective visual information fusion perception based on driver facial information, driver body movement information, and vehicle forward environmental information, achieving more accurate and comprehensive understanding and prediction of driving intentions. It also proposes a multi-perspective fusion calculation method based on cross-perspective consistency contrastive learning. By performing consistency constraint-driven contrastive learning on visual features from multiple perspectives, cross-perspective feature extraction associations are established, improving the multi-perspective visual information fusion effect and the accuracy of driving intention recognition. A visual feature extraction network based on progressive group convolution employs sparse group convolution in the shallow layers of the convolutional neural network and sparse group convolution in the deep layers. This allows the visual feature extraction network to improve the quality of visual information extraction while maintaining computational efficiency, thereby enhancing the accuracy of driving intention recognition. In summary, the present invention can more fully capture and understand the driver's driving intentions through visual information inside and outside the vehicle cockpit, significantly improving the accuracy of driver intention recognition in scenarios such as lane changing, turning, going straight, and deceleration compared to existing methods.

[0065] Based on the same inventive concept, the present invention proposes a vehicle driving intention recognition system, such as Figure 1 Shown, including: Acquisition module: used to collect the driver's front image, side image and vehicle forward image data; specifically includes: 1) Driver front vision sensor: This sensor is installed in the cockpit, usually above the steering wheel or on the instrument panel, and is used to capture the driver's facial image in real time.

[0066] The sensor uses a high-resolution camera with 1080p or higher resolution and is equipped with infrared fill light to accommodate nighttime or low-light conditions. By analyzing facial images, the system extracts key information such as the driver's head posture, gaze direction, and fatigue level. Head posture includes parameters such as pitch and yaw angles, gaze direction reflects eye focus, and fatigue level is assessed using indicators such as eyelid closure frequency and yawning frequency. The frontal vision sensor generates a frontal image data stream for subsequent feature extraction and state estimation.

[0067] 2) Driver side vision sensor: This sensor is installed on the side of the cockpit, usually on the side of the seat or inside the door, to capture the driver's body movements.

[0068] The sensor uses a wide-angle camera that covers the driver's upper body, recording hand movements and body posture. Hand movements include gripping the steering wheel and operating the gearshift, while body postures include leaning forward and sideways. These movements reflect the driver's intentions, such as turning or changing lanes, as well as their attention allocation. The lateral vision sensor generates a side-view data stream that reflects the driver's dynamic behavior.

[0069] 3) Vehicle forward vision sensor: This sensor is installed in front of the vehicle, usually located on the front of the vehicle or on the top of the windshield, and is used to perceive the vehicle's external environment information.

[0070] The sensor can use a monocular camera or a binocular stereo vision system, and can be combined with LiDAR to enhance perception capabilities. By capturing information such as road conditions, traffic signals, and obstacle distribution, the system can understand the context of the vehicle's movement. Road conditions include lane position, traffic signals include traffic light status, and obstacle distribution includes data such as the distance to the vehicle ahead. This environmental information provides a critical basis for understanding driver intent. The forward-facing vision sensor generates a forward-facing image data stream that describes the external scene in which the vehicle is traveling.

[0071] The state calculation module is used to extract features from the front image, side image and forward image data of the driver to be tested; align and optimize the extracted features; and calculate the state information corresponding to each feature based on the optimized features.

[0072] The recognition module extracts node features for each state. Based on these node features, it establishes high-order causal hyperedge relationships between each state based on prior knowledge, thereby constructing a hypergraph structure. It uses hypergraph convolution to infer the evolution of each state feature, updates the node features, and then concatenates and fuses all updated node features to generate a driving scene representation vector. It then identifies the driving intention category based on the driving scene representation vector. This module, the core processing unit of the system, utilizes a high-performance graphics processor, such as the HUAWEI MDC610 or similar hardware, and is responsible for implementing the driver intention recognition model. The graphics computing unit receives data streams from frontal, lateral, and forward vision sensors and processes multi-view images in real time. This processing includes feature extraction, cross-view fusion, state estimation, and intention inference, ultimately outputting a driving intention result, such as "turn left" or "slow down." Furthermore, this module communicates with the vehicle control system via the CAN bus or Ethernet, providing information support for assisted driving decisions, such as triggering automatic braking or steering prompts. The graphics computing unit generates a driving intention classification result and associated confidence level.

[0073] The overall operation process of the system is as follows Figure 2As shown: The front, side, and forward vision sensors simultaneously capture image data at a fixed frame rate of 30 frames per second, forming a multi-view visual data stream. The graphics computing unit preprocesses the raw images, ensuring input data quality and consistency through operations such as denoising, image enhancement, and temporal synchronization. This preprocessed data undergoes feature extraction, cross-view fusion, and intention inference modules to calculate driving intent. The graphics computing unit transmits the intent results to the vehicle control system and records intermediate features during the inference process, such as gaze direction and fatigue status, to support subsequent optimization or manual verification.

[0074] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A vehicle driving intention recognition method, characterized in that: The following steps are involved: Collect the front image, side image and forward image data of the driver to be tested and the vehicle he is driving; Feature extraction is performed on the frontal and side images of the driver and the forward image data of the vehicle being driven; the extracted facial and full-body image features of the driver and the image features of the environment in front of the vehicle are aligned and optimized; and the corresponding state information is calculated based on the optimized image features; the state information includes the driver's head posture, line of sight, fatigue state, body movements, and the environment outside the vehicle being driven; Map each state information into a feature vector, and establish high-order causal hyperedge relationships between each state information based on prior knowledge; A hypergraph structure is constructed using feature vectors as nodes and high-order causal hyperedge relationships between each state information as edges. The evolution of the feature vector corresponding to each state information is inferred based on the hypergraph structure, and the nodes are updated. The features corresponding to all updated nodes are concatenated and fused to generate a driving scene representation vector. The driving intention category is identified based on the driving scene representation vector.

2. A vehicle driving intention recognition method according to claim 1, characterized in that: The hierarchical grouping strategy of the progressive group convolutional neural network (PGCNN) is used to extract features from the frontal image, side image, and forward image data of the driver and his vehicle, respectively. The specific steps include: The input front and side images of the driver and the forward image data of the vehicle are subjected to progressive group convolution calculation; wherein, in the first layer, the input image data is subjected to convolution operation to obtain the feature map of the image data ; in the layer, the output feature map of the previous layer Divided along the channel dimension Group, for each group of feature maps Perform convolution operations independently, expressed as: ; in, For the The convolution kernel of the group; The feature map of each group output Splice and get the feature map ; Use ReLU activation function and batch normalization to transform the feature map Process and extract the feature vector of each image data , which includes the frontal image feature vector , silhouette feature vector and the forward image feature vector .

3. A vehicle driving intention recognition method according to claim 2, characterized in that: Cross-view consistency contrast learning is used to align and optimize the extracted features of the driver's facial image, full-body image, and vehicle front environment image. The specific steps include: According to the feature vectors of the frontal image, side image and forward image Construct positive and negative samples; in a training batch, the feature pairs of the same scene are considered as positive sample pairs, which are close to each other in the feature space; feature pairs of different scenes and are regarded as negative sample pairs, which are far away from each other in the feature space; Calculate eigenvectors Consistency loss function for: ; ; in, is the cosine similarity, is the temperature coefficient, and is the index number of the sample in a data batch during training, represents the number of samples in the batch, represents a positive sample pair; represents a negative sample pair; By optimizing the target The PGCNN weight parameters are optimized with back-propagation calculation so that the feature vectors of the same scene are close to each other in the feature space, and the optimized image features are obtained.

4. A vehicle driving intention recognition method according to claim 3, characterized in that: Calculating the state information corresponding to each optimized image feature includes the following steps: The optimized frontal image feature vector , silhouette feature vector and the forward image feature vector Input them into the fully connected neural network FCNN respectively and calculate their corresponding states: ; in, Indicates the status of head posture, gaze direction, fatigue, body movements and the external environment: Indicates head posture, Indicates the direction of sight, Indicates fatigue state, Indicates body movements, Indicates the environment outside the vehicle.

5. The method for identifying vehicle driving intention according to claim 4, characterized in that: The mapping of each state information into a feature vector and establishing a high-order causal hyperedge relationship between each state information based on prior knowledge; The hypergraph structure is constructed by taking the feature vectors as nodes and the high-order causal hyperedge relationships between each state information as edges. The specific steps include: Head posture , sight direction , fatigue state , body movements 、External environment Input into the causal reasoning-based hypergraph neural network CI-HGCN to extract the node features of each state information; Based on prior knowledge of driving behavior, we define high-order causal hyperedge relationships between each state information, including: hyperedges connecting head posture, gaze direction, and body movements to model lane change or steering intentions; hyperedges connecting fatigue state, body movements, and the external environment to model braking intentions; and hyperedges connecting gaze direction, body movements, and the external environment to model acceleration or deceleration intentions. According to the node characteristics and high-order causal hyperedge relationships, the hypergraph structure is constructed using the hyperedge indicator matrix, which is expressed as: 。 6. A vehicle driving intention recognition method according to claim 5, characterized in that: The method of inferring the evolution of the feature vector corresponding to each state information based on the hypergraph structure and updating the nodes; concatenating and fusing the features corresponding to all updated nodes to generate a driving scene representation vector; specifically includes the following steps: Based on the hypergraph structure, the node features are updated by hypergraph convolutional reasoning based on the evolution of node features. The update process is expressed as: ; in: is a node The updated features of is the hyperedge weight, is the super-edge degree, Indicates the index number of the node, is the total number of nodes, is the node degree, is the weight matrix of the hypergraph convolutional layer, is the ReLU activation function, represents the index number of the hyperedge, represents the number of hyperedges; The updated features of all nodes are concatenated and fused to generate a driving scene representation vector.

7. The method for identifying vehicle driving intention according to claim 2, wherein: It also includes supervising the learning optimization of PGCNN by establishing an intention estimation loss function, and driving the update of PGCNN parameters by stimulating the output of the PGCNN model to be consistent with the actual situation; the intention estimation loss function is expressed as: ; in, is the true intention label, is the predicted probability, is the correction coefficient of the regularization term in the loss function, is the L1 norm of the weight parameter, Indicates the sample index number, represents the number of samples used for training, Indicates the category index number, Indicates the number of categories contained in the training data.

8. The vehicle driving intention recognition method according to claim 1, characterized in that: The collecting of the front image and side image of the driver to be tested and the forward image data of the vehicle he is driving specifically includes the following steps: Visual sensors are used to collect the driver's front image, side image and vehicle forward image respectively; the front image data includes the driver's facial image, head posture and line of sight direction image; the side image includes the driver's upper body movement image; the vehicle forward image includes video images of road conditions, traffic signals and obstacle distribution.

9. The method for identifying vehicle driving intention according to claim 1, wherein: The method further includes preprocessing the front image, side image and forward image data of the driver and the vehicle after collecting the front image and side image of the driver and the vehicle, wherein the preprocessing process specifically includes the following steps: Gaussian filtering is used to denoise the image data; Perform contrast stretching on the denoised image data; Align the timestamps of each processed image data; The image data after timestamp alignment is resampled by bilinear interpolation to obtain the processed image data.

10. A vehicle driving intention recognition system, characterized in that: include: An acquisition module is used to acquire the front image and side image of the driver to be tested and the forward image data of the vehicle he is driving; The state calculation module is used to extract features from the front and side images of the driver and the forward image data of the vehicle being driven; align and optimize the extracted facial and full-body image features of the driver and the image features of the environment in front of the vehicle; and calculate the corresponding state information based on the optimized image features; the state information includes the driver's head posture, line of sight, fatigue state, body movements, and the environment outside the vehicle being driven; A hypergraph structure building module is used to map each state information into a feature vector and establish high-order causal hyperedge relationships between each state information based on prior knowledge; A hypergraph structure is constructed using feature vectors as nodes and high-order causal hyperedge relationships between each state information as edges. The evolution of the feature vector corresponding to each state information is inferred based on the hypergraph structure, and the nodes are updated. The features corresponding to all updated nodes are concatenated and fused to generate a driving scene representation vector. The recognition module is used to identify the driving intention category based on the driving scene representation vector.

Citation Information

Patent Citations

  • Driver expressway lane changing intention prediction method, system and device

    CN110427850A

  • Driver driving behavior identification method and system based on cyclic graph convolutional network

    CN114078243A

  • Driver intention recognition method

    CN117485348A

  • Method for identifying steering intention of driver by fusing images inside and outside vehicle

    CN118823439A

  • Driver behavior recognition method based on geometric space fusion feature deep attention network

    CN119919916A

Cited By

  • Motion motion state determination method and device, storage medium and program product

    CN122020434A