Driver behavior recognition method
By combining hyperspectral imaging and attention mechanism models, the problem of insufficient global contextual information modeling in driving behavior recognition is solved, and high-precision driving behavior recognition and early warning are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN QIYANG SPECIAL EQUIP TECH ENG CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-06-02
AI Technical Summary
Existing driving behavior recognition methods struggle to effectively model the global contextual information of driver behavior, especially the spatial semantic relationship between hands and face, resulting in insufficient resolution accuracy when processing hyperspectral images.
The system uses hyperspectral imaging technology to acquire driver images, extracts a three-dimensional cube after dimensionality reduction through principal component analysis, and uses an attention mechanism model for inference. It also combines a Gaussian modulation attention module and a complementary Transformer module for global semantic modeling to achieve high-precision discrimination of driving behavior.
It improves the accuracy and robustness of driving behavior recognition, effectively identifying dangerous behaviors and activating the vehicle's early warning mechanism in complex scenarios, thus avoiding misjudgments.
Smart Images

Figure CN122135343A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle control, and more specifically, to a method for driver behavior recognition. Background Technology
[0002] With the development of intelligent driving and advanced driver assistance systems (ADAS), real-time and accurate identification of driver behavior has become crucial for improving driving safety. Dangerous behaviors such as talking on the phone, smoking, and closing one's eyes due to fatigue can easily lead to distraction and significantly increase the risk of traffic accidents.
[0003] Existing driving behavior recognition methods typically extract features from video or image data and then classify them. However, when dealing with high-dimensional visual data (such as hyperspectral images or spatiotemporal cubes), relying solely on local feature extractors (such as convolutional networks) makes it difficult to effectively model global contextual information related to behavior, such as the spatial semantic association between hands and faces.
[0004] Therefore, there is an urgent need for a driving behavior recognition method that can take into account both local detail discrimination and global semantic modeling capabilities, in order to improve the parsing accuracy of high-dimensional driving scene data and thus achieve more reliable behavior discrimination and early warning. Summary of the Invention
[0005] In view of the above problems, this application proposes a driver behavior recognition method that can solve the above problems.
[0006] This application provides a driver behavior recognition method applied to a vehicle. The method includes: acquiring a hyperspectral image containing the driver; extracting multiple three-dimensional cubes from the dimensionality-reduced hyperspectral data, wherein the label of each three-dimensional cube is determined by the category label of its central pixel; inputting the multiple three-dimensional cubes into an attention mechanism model for inference to obtain the semantic category of each three-dimensional cube; determining whether the driver is currently in a preset dangerous behavior state based on the semantic category, and activating an onboard warning mechanism when the dangerous behavior state meets a preset triggering condition.
[0007] Therefore, an attention mechanism model is used to reason about the 3D cubes extracted from hyperspectral data. The attention mechanism model has the ability to model global contextual information (e.g., hand-face spatial association), thus overcoming the shortcomings of traditional local feature extractors (e.g., convolutional networks) in capturing global semantics of behavior. At the same time, dangerous state judgment and warning are performed by combining the semantic category of each 3D cube, achieving high-precision and reliable behavior discrimination of high-dimensional driving data, effectively taking into account both global semantic modeling and practical safety application needs. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0009] Figure 1 A flowchart illustrating a driver behavior recognition method provided in an embodiment of this application is shown.
[0010] Figure 2 A schematic diagram of the structure of a driver behavior recognition device provided in an embodiment of this application is shown.
[0011] Figure 3 A schematic diagram of the structure of a vehicle provided in an embodiment of this application is shown.
[0012] Figure 4 This illustration shows a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of this application. Detailed Implementation
[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0014] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0015] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0016] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0018] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0019] This application provides a driver behavior recognition method for use in vehicles. The method includes: acquiring a hyperspectral image containing the driver; extracting multiple three-dimensional cubes from the dimensionality-reduced hyperspectral data, wherein the label of each three-dimensional cube is determined by the category label of its central pixel; inputting the multiple three-dimensional cubes into an attention mechanism model for inference to obtain the semantic category of each three-dimensional cube; determining whether the driver is currently in a preset dangerous behavior state based on the semantic category, and activating an onboard warning mechanism when the dangerous behavior state meets preset triggering conditions.
[0020] Therefore, an attention mechanism model is used to reason about the 3D cubes extracted from hyperspectral data. The attention mechanism model has the ability to model global contextual information (e.g., hand-face spatial association), thus overcoming the shortcomings of traditional local feature extractors (e.g., convolutional networks) in capturing global semantics of behavior. At the same time, dangerous state judgment and warning are performed by combining the semantic category of each 3D cube, achieving high-precision and reliable behavior discrimination of high-dimensional driving data, effectively taking into account both global semantic modeling and practical safety application needs.
[0021] Please see Figure 1 , Figure 1 This illustration shows a flowchart of a driver behavior recognition method provided in an embodiment of this application, which can be applied to vehicles. Figure 1 As shown, the method may include steps 110 to 140.
[0022] In step 110, a hyperspectral image containing the driver is acquired.
[0023] In existing technologies, driving behavior analysis mostly relies on visible light (RGB) cameras. In scenarios with drastic changes in lighting (e.g., strong backlight, nighttime), or when the driver is wearing sunglasses or their face is partially obscured, problems such as blurred features and failure to make judgments may occur.
[0024] Hyperspectral imaging technology can simultaneously acquire the spatial information of a target and the spectral response of dozens to hundreds of consecutive narrow bands. It is highly sensitive to material composition and physiological state (e.g., blood oxygenation and skin reflectivity), and can effectively penetrate ordinary optical interference, providing more robust behavioral discrimination criteria.
[0025] In embodiments of this application, a hyperspectral imaging device is installed inside the vehicle to acquire hyperspectral images containing the driver. In one specific embodiment, the hyperspectral imaging device may include a hyperspectral camera, a synchronization control unit, and an onboard computing platform.
[0026] The hyperspectral camera is mounted inside the vehicle's windshield or above the dashboard, facing the driver's seat area, to continuously acquire hyperspectral data cubes of the driver's head and upper body. A synchronization control unit triggers the hyperspectral camera to capture images at a preset frame rate (e.g., 5–30 Hz). The onboard computing platform communicates with the hyperspectral camera to perform principal component analysis, 3D cube extraction, attention mechanism model inference, and behavioral warning output.
[0027] Hyperspectral cameras can be designed using push-broom, snapshot, or filter rotation architectures, with a preferred spectral coverage of 400–1000 nm (covering visible and near-infrared light) to balance environmental adaptability and sensitivity to physiological signals.
[0028] In one specific implementation, a snapshot hyperspectral camera (e.g., IMEC's Snapshot Hyperspectral Imager or Viavi Nano-Hyperspec) is used. Hyperspectral cameras are based on filter arrays or computational imaging principles and can simultaneously acquire spatial and spectral information within a single frame exposure. They do not require mechanical scanning, are highly vibration resistant, and are suitable for automotive environments.
[0029] In one specific implementation: the hyperspectral camera uses a spectral range of 450–950nm (covering visible and near-infrared light), uses 32–64 spectral bands (configurable), has a spatial resolution of 640×480 pixels, uses a frame rate of 15–30FPS, and has a module size of less than 5cm³, which can be embedded in the rearview mirror base or above the dashboard.
[0030] The hyperspectral camera connects to the vehicle domain controller (e.g., NVIDIA Jetson Orin NX or Qualcomm Snapdragon Ride™ platform) via a USB 3.0 or GMSL interface. The vehicle domain controller simultaneously runs perception algorithms and vehicle control logic. The hyperspectral camera is mounted in the central area inside the windshield, with a field of view (FOV) covering the driver's head and hand operating area (approximately 60° horizontal × 40° vertical).
[0031] Specifically, a hyperspectral image can be denoted as... .in, It can be hyperspectral data, For hyperspectral image tag set, This represents the spatial size of a hyperspectral image, while Indicates the number of bands in a hyperspectral image. This represents the maximum number of category labels. Although hyperspectral images carry a wealth of useful spectral information, they contain many redundant bands, which increases computational complexity and can potentially reduce classification performance.
[0032] Therefore, in some implementations, after the step of "acquiring a hyperspectral image containing the driver", the method further includes: performing principal component analysis on the hyperspectral image to reduce its dimensionality, thereby obtaining the reduced hyperspectral data.
[0033] This reduces dimensionality and removes redundant information, thereby reducing bandwidth and computational load. After preprocessing using principal component analysis, the number of bands in the hyperspectral image is reduced from... Reduce to The output becomes .
[0034] In one specific implementation, the hyperspectral image (e.g., 640×480×64) is first subjected to principal component analysis (PCA) dimensionality reduction at the edges, retaining the first 10–15 principal components (cumulative variance contribution rate >95%), thus compressing the number of bands to [a specific value]. =10 15. Significantly reduces the input dimensions of subsequent steps (e.g., Conv3D).
[0035] In step 120, multiple three-dimensional cubes are extracted from the dimensionality-reduced hyperspectral data, and the label of each three-dimensional cube is determined by the category label of its center pixel.
[0036] For the dimensionality-reduced hyperspectral data Perform 3D cube extraction and generate Adjacent 3D cubes .in, This represents the spatial dimensions of each cube. It is particularly important to note that all three-dimensional cubes... center pixel ( , ) is defined as its label, namely a three-dimensional cube. The label is determined by the center pixel of its corresponding position, where, , .
[0037] In the embodiments of this application, not all extracted 3D cubes are used in the attention mechanism model used in subsequent steps. This application removes samples labeled "background" and retains only non-background data with clear behavioral semantics (i.e., labels). (Samples). During the training of the attention mechanism model, non-background data is divided into training and test sets to ensure that the attention mechanism model learns discriminative features directly related to driving behavior.
[0038] In step 130, multiple 3D cubes are input into the attention mechanism model for inference to obtain the semantic category of each 3D cube.
[0039] In the embodiments of this application, the overall structure of the attention mechanism model mainly consists of three parts: a convolution-based feature extraction module, a Gaussian modulation attention module, and a complementary Transformer module. Through the collaborative work of the convolution-based feature extraction module, the Gaussian modulation attention module, and the complementary Transformer module, which are respectively responsible for local feature extraction, channel adaptive weighting, and global semantic modeling, high-precision recognition of driving behavior is achieved.
[0040] In one specific implementation, the convolution-based feature extraction module sequentially includes a three-dimensional convolutional layer (Conv3D), a spectral compression layer (using a 1×1×b convolutional kernel), and a two-dimensional convolutional layer (Conv2D).
[0041] In one specific implementation, the Gaussian modulation attention module includes a channel weight generation unit, a Gaussian modulation unit, and a residual connection unit. The channel weight generation unit performs channel statistics and transformations on the input spatial feature map; the Gaussian modulation unit applies Gaussian distribution modulation to the generated channel weights; and the residual connection unit adds the modulated weighted features to the original input to output the final feature map.
[0042] In one specific implementation, the complementary Transformer module includes a token embedding layer, a complementary multi-head self-attention mechanism (CMHSA) layer, a feedforward network (MLP) layer, and a layer normalization unit. The token embedding layer flattens the input feature map and concatenates learnable classification tokens and positional codes to form an input sequence; the CMHSA layer consists of a standard multi-head self-attention branch and parallel convolutional branches, and their outputs are fused; the MLP layer consists of two fully connected layers and an intermediate activation function; residual connections and layer normalization are used between the sub-layers, and the entire module can be stacked in multiple layers.
[0043] The semantic category of each three-dimensional cube can be a high-level semantic label of the actions performed by the driver within the corresponding spatiotemporal segment, and its value comes from a predefined set of driving behavior categories. In the embodiments of this application, the semantic categories include, but are not limited to: normal driving, making a phone call, smoking, drinking water, closing eyes, yawning, and taking both hands off the steering wheel.
[0044] Multiple 3D cubes are input into an attention mechanism model for end-to-end inference to achieve fine-grained semantic discrimination of driving behavior. Specifically, in some implementations, the step "inputting multiple 3D cubes into an attention mechanism model for inference to obtain the semantic category of each 3D cube" may include the following steps: (1) Through the convolution-based feature extraction module, three-dimensional convolution, spectral compression and two-dimensional convolution are performed on multiple three-dimensional cubes in sequence to obtain the spatial feature map corresponding to each three-dimensional cube; (2) By using the Gaussian modulation attention module, channel weighting and residual connection operations are performed based on the spatial feature map to output the final feature map that takes into account both the main and secondary features; (3) Through the complementary Transformer module, global semantic modeling and behavior classification are performed on the final feature map to obtain the semantic category of each three-dimensional cube.
[0045] Considering that hyperspectral images (HSI) can provide spectral resolution far exceeding that of traditional RGB images in driving scenarios, containing rich spectral dimensions and spatial structure information, this application first utilizes three-dimensional convolution (Conv3D) to extract features from the input hyperspectral data cube to capture the joint correlation between spectral bands and spatial locations; simultaneously, two-dimensional convolution (Conv2D) is introduced to perform refined spatial feature modeling on a single spatial plane (e.g., a specific band or a fused feature map). Specifically, in some embodiments, the step "by sequentially performing three-dimensional convolution, spectral compression, and two-dimensional convolution on multiple three-dimensional cubes using a convolution-based feature extraction module to obtain the spatial feature map corresponding to each three-dimensional cube" may include the following steps: (1) By using three-dimensional convolution, the spectral-spatial joint features in each three-dimensional cube are extracted to obtain the spectral-spatial joint features corresponding to each three-dimensional cube; (2) By spectral compression, the spectral-spatial joint features are processed to obtain the processed spectral-spatial joint features; (3) By using two-dimensional convolution, spatial feature modeling is performed on the processed spectral-spatial joint features to obtain the spatial feature map corresponding to each three-dimensional cube.
[0046] By employing a dual-path feature extraction strategy involving both 3D and 2D convolution, the attention mechanism model can both preserve the "spectral-spatial" coupling characteristics of hyperspectral data and enhance the spatial detail representation of key regions, providing a highly discriminative feature foundation for subsequent attention mechanisms and behavior discrimination. Specifically: Using 3D convolution (Conv3D) technology on each 3D cube Extracting spectral-spatial joint features. This feature extraction method effectively captures the interaction between spectral information and spatial structure, thus providing more accurate behavioral analysis results. The spectral-spatial joint features corresponding to each 3D cube are represented as follows: in, For the first Layer Zhang's feature map is located at ( , , The activation values of neurons in the three-dimensional cube (which comprehensively reflect the spatial activity of the cube) are: , ) and spectrum ( (joint information across dimensions) For activation function, , and Representing the first The height, width, and depth of the 3D convolutional kernel. For the first A feature cube at position ( , , The weight parameter at position ) For the first Layer Zhang's feature map is located at ( , , The neuron activation value, This is a bias term.
[0047] Next, spectral compression techniques are used to process the extracted spectral-spatial joint features. By applying a 1×1×n convolution kernel (i.e., dimensionality reduction in the spectral dimension), the high-dimensional spectral information is fused into a single-channel representation, thereby reducing redundant information and enhancing computational efficiency.
[0048] Simultaneously, two-dimensional convolution (Conv2D) is introduced to further refine the spectral compression features, enhancing the spatial detail representation of key regions and laying a solid foundation for subsequent attention mechanisms. The two-dimensional convolution (Conv2D) consists of Conv2D layers, BN layers, and ReLU layers, and its output feature size remains [missing information]. .
[0049] In one specific implementation, spatial feature modeling is performed on each 3D cube using 2D convolution to obtain a spatial feature map corresponding to each 3D cube, which can be represented as: in, For the first Layer Zhang's feature map is located at... , Spatial feature map, For activation function, and These represent the height and width of the two-dimensional convolution kernel, respectively. For the first Zhang's feature map is in position , The weight parameters at that location, For the first Layer Zhang's feature map is located at... , Spatial feature map, This is a bias term.
[0050] Thus, this application has achieved efficient utilization of hyperspectral data through effective preprocessing and feature extraction strategies for the original HSI data, from the representation of the original HSI data to PCA dimensionality reduction, and finally through Conv3D and Conv2D feature extraction. This significantly improves the accuracy and robustness of driver behavior recognition.
[0051] Furthermore, in the embodiments of this application, two types of feature regions are defined: primary features correspond to discriminative regions, and secondary features represent important but easily overlooked regions. Primary features can be discriminative regions with high response intensity in the spatial-spectral domain and high correlation with the target behavior category. For example, in the "making a phone call" behavior, the spectral reflectance characteristics of the hand near the ear and changes in facial orientation typically exhibit strong activation, constituting the main basis for the model's decision. Secondary features are regions with weaker responses, easily ignored by attention mechanisms or subsequent classifiers, but still have supplementary value for behavior discrimination. For example, subtle signals such as slight wrist movements of the driver, spectral differences in sleeve material, and slight eyelid closure, while not dominating the classification result, can effectively improve the robustness of the model in complex scenes (e.g., occlusion, low light).
[0052] Therefore, it is evident that primary features are crucial for improving discriminative ability, while secondary features also contribute to obtaining better classification results. Thus, to further improve the performance of the Gaussian Modulation Attention Block (GMA), particularly by enhancing the role of these secondary features, this application uses the Gaussian Modulation Attention Block (GMA) in the attention mechanism model to redistribute feature distribution to enhance secondary features in the channels, while retaining the original primary features.
[0053] In some implementations, the step "using a Gaussian modulation attention module to perform channel weighting and residual connection operations based on the spatial feature map, and outputting a final feature map that takes into account both primary and secondary features" may include the following steps: (1) Perform average pooling on the spatial feature map to obtain global context features; (2) Perform linear transformation and nonlinear activation on the global context features to generate intermediate features containing channel dependencies; (3) Based on the intermediate features and the Gaussian modulation function, the weights of each channel are redistributed through the Gaussian distribution to enhance the secondary features and obtain the modulation weight map; (4) Add the modulation weight map and the spatial feature map pixel by pixel to output the final feature map.
[0054] Gaussian modulation attention module receives spatial feature maps as input ( (representing the number of channels) is subjected to average pooling. This reduces computational complexity and extracts global context information to obtain global context features. Next, the global context features... A series of linear transformation layers and activation function layers are fed in. This allows for the introduction of nonlinear factors and the learning of intermediate features that include channel dependence. Subsequently, intermediate features Feeded into the Gaussian modulation function ( In this process, the Gaussian modulation function is used to adjust the weight distribution among the channels, giving special emphasis to those channels that carry secondary features but contribute to the final classification, thereby generating a modulated weight map. Finally, the modulation weight map The corresponding weights are applied channel by channel to the spatial feature map. And by adding them element by element, they are combined with the spatial feature map. By combining these methods, we can ensure that secondary features are enhanced while retaining the original important features, resulting in the final output feature map. .
[0055] In one specific implementation, a Gaussian modulation attention module performs channel weighting and residual connection operations based on the spatial feature map, outputting a final feature map that takes into account both primary and secondary features, which can be represented as: in, For average pooling function, Includes linear transformation and activation function layers. It is a Gaussian modulation function. This is a channel-by-channel weighted operation.
[0056] Therefore, in the Gaussian modulation attention module, the Gaussian modulation function is used to redistribute the channel responses of intermediate features to achieve explicit enhancement of secondary features. Specifically, in some implementations, the step "redistributing the channel weights according to the intermediate features and the Gaussian modulation function through a Gaussian distribution to enhance secondary features and obtain a modulation weight map" may include the following steps: (1) Map the intermediate features to a Gaussian distribution to generate a modulation weight map; the intermediate features contain N channel response values, where N is a positive integer; (2) Calculate the arithmetic mean of the response values of N channels as the mean of the Gaussian distribution; (3) Calculate the root mean square deviation of the response value of each channel relative to the arithmetic mean, and use it as the standard deviation of the Gaussian distribution; (4) Substitute each channel response value into the Gaussian probability density function parameterized by the mean and standard deviation to calculate the corresponding weight element in the modulation weighting diagram.
[0057] Assume the intermediate features input to the Gaussian modulation function are , Indicates the total number of channels. For the first The scalar response values of each channel. The Gaussian modulation function first calculates the intermediate features. The statistical properties (i.e., the mean and standard deviation of the Gaussian distribution) will determine the response of each channel. Mapped to mean Centered, standard deviation On a Gaussian probability density function of scale , the corresponding weight elements in the modulation weight map are obtained.
[0058] By modulating the weight map to reweight the spatial feature map at the channel level, the attention mechanism model expands its focus from solely on high-response primary features to include secondary features with moderate responses. Due to the bell-shaped nature of the Gaussian distribution, channels with response values close to the mean of all channels (typically corresponding to auxiliary cues in behavior discrimination, such as finger gestures or sleeve reflections) receive higher weights. This enriches the semantic hierarchy of the feature representation without losing dominant information, significantly improving the accuracy and robustness of driving behavior classification.
[0059] In one specific implementation, the arithmetic mean of the response values of N channels is calculated as the mean of a Gaussian distribution, which can be expressed as: in, The mean of a Gaussian distribution is given. The total number of channels. For the first The scalar response value of each channel.
[0060] In one specific implementation, the root mean square deviation of each channel response value relative to the arithmetic mean is calculated, which serves as the standard deviation of the Gaussian distribution and can be expressed as: in, Let be the standard deviation of the Gaussian distribution. The total number of channels. For the first The scalar response value of each channel. This is the arithmetic mean of the responses from all channels.
[0061] By modulating the weight map, strong response channels (primary features) far from the mean receive lower weights; medium and weak response channels (secondary features) close to the mean are relatively enhanced because they are near the peak of the Gaussian distribution; the overall weight distribution is bell-shaped, avoiding extreme suppression or amplification and maintaining feature diversity. This effectively alleviates the bias problem of "stronger features getting stronger and weaker features getting weaker" in traditional Gaussian modulation attention modules, enabling the model to consider both dominant regions and easily overlooked but semantically valuable secondary regions.
[0062] Furthermore, although the convolution-based feature extraction modules (including 3D and 2D convolutions) in attention mechanism models can effectively capture the local spectral-spatial joint features of driver behavior, their inherent local receptive field limits their ability to model global contextual information (e.g., coordinated hand and face movements, cross-regional behavioral associations).
[0063] In contrast, the complementary Transformer module, through its self-attention mechanism, can directly establish dependencies between any two locations, thereby effectively capturing long-range dependencies between remote features and deep semantics. This is crucial for understanding complex driving behaviors (e.g., "looking at a phone while steering").
[0064] Given the complementarity between the convolution-based feature extraction module and the complementary Transformer module in feature learning—the convolution-based feature extraction module excels in local details and translation invariance, while the complementary Transformer module excels in global semantic and structural modeling—the attention mechanism model adopted in this application also includes a complementary Transformer module to fuse and enhance the multi-scale feature representations output from the aforementioned feature extraction stage. Specifically: In the embodiments of this application, the complementary Transformer module mainly includes positional encoding and a complementary multi-head self-attention submodule. The complementary multi-head self-attention submodule is the core of the complementary Transformer module, used to explicitly model long-range semantic associations across regions and channels while preserving local discriminative cues, ultimately improving the accuracy and robustness of behavior classification. In some embodiments, the step "using the complementary Transformer module to perform global semantic modeling and behavior classification on the final feature map to obtain the semantic category of each 3D cube" may include the following steps: (1) Flatten the final feature map in terms of spatial dimensions to obtain a two-dimensional feature matrix composed of multiple spatial location feature vectors; (2) By using a linear mapping layer, the two-dimensional feature matrix is projected onto a preset latent space dimension to generate a content token sequence; (3) Add a category token at the beginning of the content token sequence to form a complete token sequence containing both the category token and the content token; (4) Add the position code to each token in the complete token sequence to obtain the input sequence adapted to the complementary Transformer module; (5) Input the input sequence into the complementary Transformer module, and after processing by multi-head self-attention mechanism, layer normalization and multilayer perceptron, extract the output representation of the classification token, and determine the semantic category of the corresponding three-dimensional cube based on the output representation of the classification token.
[0065] Assume the final feature map input to the complementary Transformer module is , For the number of channels, The spatial dimension is used to adapt to the serialization input requirements of the Transformer. First, the spatial dimension of the final feature map is flattened, resulting in a shape of... The two-dimensional feature matrix is then projected onto a predefined latent space dimension through a learnable linear mapping layer. Generate a sequence of input representations (i.e., content token sequence).
[0066] Subsequently, a learnable classification token is appended to the beginning of the content token sequence. , forming a collection A complete sequence of tokens; to preserve the original spatial location information, positional encoding (PE) is introduced and added to the input token sequence. The token sequence includes a learnable classification token (denoted as ). )and A content token obtained by feature projection Its combination form is: in, For classification tokens (a learnable vector used for final behavior category prediction), and These are content tokens, For position encoding.
[0067] The core of the complementary Transformer module is the multi-head self-attention mechanism, which dynamically establishes dependencies between any two positions by calculating the correlation between the query matrix, key matrix, and value matrix, thereby capturing long-range semantic associations in driving behavior (e.g., the collaborative pattern of hand movements and head posture). Specifically, in some implementations, the method further includes the following steps: (1) Generate a query matrix, a key matrix, and a value matrix from the input sequence through a linear mapping; (2) The query matrix, key matrix and value matrix are processed separately to obtain multiple sub-attention outputs; (3) Concatenate multiple sub-attention outputs along the channel dimension and perform a linear transformation through the weight parameters to obtain the self-attention branch output; (4) Reshape the value matrix into a spatial tensor form and input it into a convolution function containing convolutional layers and batch normalization layers to extract local features and obtain the output of the convolutional branch; (5) The output of the self-attention branch and the output of the convolution branch are added and fused element by element to obtain the enhanced feature representation, so as to determine the semantic category of the corresponding three-dimensional cube based on the enhanced feature representation.
[0068] First, the input sequences are used to generate query matrices through learnable linear mappings. Key matrix Sum matrix Subsequently, standard scaled dot product self-attention computation is performed on each attention head to obtain multiple sub-attention outputs. In one specific implementation, each sub-attention output can be represented as: in, For child attention output, For querying the matrix, The key matrix, Let be the dimension of the key matrix. It is a value matrix.
[0069] After calculating all sub-attention outputs, the sub-attention outputs are concatenated along the channel dimension and transformed using a learnable linear transformation matrix. Mapping to the target dimension yields the output representation of the self-attention branch (i.e., the self-attention branch output).
[0070] Crucially, this application further introduces a convolution branch: the value matrix... Remodeled into a space tensor form (e.g., The input is fed into a convolutional function that contains convolutional layers and batch normalization (BN) layers. Local feature extraction is performed to enhance the sensitivity to high-frequency information such as edges and textures, and the output of the convolution branch is obtained.
[0071] Finally, the outputs of the self-attention branch and the convolution branch are added and fused element by element to obtain the final output of the complementary multi-head self-attention mechanism (i.e., the enhanced feature representation).
[0072] In one specific implementation, the enhanced feature representation can be expressed as: in, For the enhanced feature representation, , and These are the sub-attention outputs from the first layer to the h-th layer. It is a linear transformation matrix. It is a value matrix.
[0073] By fusing the outputs of the self-attention branch with the outputs of the convolutional branch, the attention mechanism model can simultaneously capture low-frequency semantic information (modeled by self-attention) and high-frequency spatial details (modeled by convolution), thereby gaining a more comprehensive understanding of the visual patterns of driving behavior.
[0074] The enhanced feature representation CMHSA is then fed into a Layer Normalization (LayerNorm) layer and a Multilayer Perceptron (MLP) layer to complete the processing flow of a single-layer encoder. Specifically, after obtaining the enhanced feature representation CMHSA, it is represented as a... A sequence of tokens ,in For category tokens, the rest ( () is a space content token.
[0075] The enhanced feature representation CMHSA first undergoes a first-layer normalization (LayerNorm) to stabilize the training process and accelerate convergence; then it is input into a multilayer perceptron (MLP), which consists of two fully connected layers with GELU or ReLU activation functions in between to introduce nonlinear transformation capabilities; finally, it undergoes a second-layer normalization process and is residually concatenated with the original CMHSA output to form a complete single-layer encoder output, which can be represented as: in, For the output of a single-layer encoder, For layer normalization function, It is a multilayer perceptron. This is the output sequence of CMHSA.
[0076] In practical implementations, the encoder structure described above can be stacked in multiple layers (e.g., 6 or 12 layers), progressively deepening the semantic expressive power of the features. After processing through all encoding layers, the final enhanced token sequence representation is obtained. .
[0077] Crucially, this application utilizes only the classification tokens in the sequence. Behavior discrimination is performed. During the multi-layer self-attention and convolution fusion process, the token has automatically aggregated the global context information of the entire 3D cube, including cross-regional pose correlations and local detail cues.
[0078] Specifically, The input is fed into a classification head, which contains a learnable fully connected layer (FClayer) with a weight matrix. Map features to preset values. A space of behavioral categories ( Given the total number of driving behavior categories (e.g., normal driving, talking on the phone, smoking, drinking water, fatigue, etc.), output the logits vector: in, This is the output logits vector. This is the weight matrix of the classification heads. For the classification token in the final encoder output, This is the bias vector for the classification head.
[0079] Finally, the semantic category of the 3D cube is determined by taking the category index corresponding to the maximum response value in the logits vector, which can be represented as: in, For the predicted semantic category index, The function that takes the index corresponding to the maximum value. For the first The logits value for each category.
[0080] Ultimately, this enables accurate semantic understanding and classification of driving behavior represented by each three-dimensional cube.
[0081] In step 140, based on the semantic category, it is determined whether the driver is currently in a preset dangerous behavior state, and when the dangerous behavior state meets the preset triggering conditions, the vehicle warning mechanism is activated.
[0082] In the embodiments of this application, a list of dangerous behavior categories is pre-established. For example, the list of dangerous behavior categories includes: "making a phone call", "closing eyes", "smoking", "drinking water", "taking both hands off the steering wheel", etc.; if the semantic category of the current three-dimensional cube belongs to the list of dangerous behavior categories, it is determined that the driver is currently in a dangerous behavior state.
[0083] Furthermore, to avoid frequent false alarms caused by momentary misjudgments, this application introduces multi-dimensional triggering conditions to control the early warning mechanism. In some implementations, the preset triggering conditions include, but are not limited to: duration conditions, confidence level conditions, and behavior combination conditions.
[0084] For example, the number of consecutive frames of the same dangerous behavior state reaches or exceeds a preset threshold (e.g., 5 consecutive frames are classified as "eyes closed", corresponding to approximately 0.5 seconds, assuming a frame rate of 10fps). For example, after Softmax normalization, the predicted probability of the dangerous category of the logits output by the classification head is not lower than a preset threshold (e.g., P("making a phone call") ≥ 0.9). For example, multiple related dangerous behaviors occur successively in a short period of time (e.g., first detecting "hand close to face", then detecting "holding an object", then jointly judging it as a high-risk event of "making a phone call").
[0085] The vehicle activates its in-vehicle warning mechanism when any one of the following triggering conditions—duration condition, confidence condition, and behavior combination condition—is met. In some implementations, the in-vehicle warning mechanism includes, but is not limited to: issuing voice prompts (e.g., "Please focus on driving"), triggering dashboard warning lights or head-up display (HUD) alarms, controlling the seat vibration module to generate vibration alerts, and sending alarm logs and corresponding video clips to a remote monitoring platform.
[0086] Thus, this application not only achieves fine-grained semantic understanding of driving behavior, but also combines temporal context and confidence assessment to effectively improve the accuracy and practicality of the warning system and avoid false alarms caused by misjudgment of a single frame.
[0087] Please see Figure 2 , Figure 2 This illustration shows a schematic diagram of a driver behavior recognition device according to an embodiment of this application, applied to a vehicle. The driver behavior recognition device 200 includes: a data acquisition module 210, an extraction module 220, an acquisition module 230, and an execution module 240. Specifically: Acquisition module 210 is used to acquire hyperspectral images containing the driver; Extraction module 220 is used to extract multiple three-dimensional cubes from the dimensionality-reduced hyperspectral data, and the label of each three-dimensional cube is determined by the category label of its center pixel; Module 230 is obtained, which is used to input multiple 3D cubes into the attention mechanism model for inference and obtain the semantic category of each 3D cube; The execution module 240 is used to determine whether the driver is currently in a preset dangerous behavior state based on the semantic category, and to activate the vehicle warning mechanism when the dangerous behavior state meets the preset triggering conditions.
[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0089] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0090] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0091] Please see Figure 3 , Figure 3 The diagram illustrates the structure of a vehicle according to an embodiment of this application. The vehicle 300 in this application may include one or more of the following components: a processor 310, a memory 320, and one or more application programs. The one or more application programs may be stored in the memory 320 and configured to be executed by one or more processors 310. The one or more programs are configured to perform the driver behavior recognition method as described in the foregoing method embodiments.
[0092] Processor 310 may include one or more processing cores. Processor 310 connects to various parts within the vehicle 300 using various interfaces and lines, and performs various functions and processes data of the vehicle 300 by running or executing instructions, programs, code sets, or instruction sets stored in memory 320, and by calling data stored in memory 320. Optionally, processor 310 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 310 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 310 and may be implemented separately using a communication chip.
[0093] The memory 320 may include random access memory (RAM) or read-only memory (ROM). The memory 320 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 320 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the vehicle 300 during use.
[0094] Please see Figure 4 , Figure 4 The diagram shows a computer-readable storage medium 400 provided in an embodiment of this application. The computer-readable storage medium 400 stores program code, which can be called by a processor to execute the driver behavior recognition method described in the above method embodiment.
[0095] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 for performing the steps according to the method embodiments of this application. This program code can be read from or written to one or more computer program devices. The program code 410 for performing the steps according to the method embodiments of this application may be compressed, for example, in a suitable form.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for driver behavior recognition, characterized in that, Applied to vehicles, the method includes: Acquire hyperspectral images containing the driver; Multiple three-dimensional cubes are extracted from the dimensionality-reduced hyperspectral data, and the label of each three-dimensional cube is determined by the category label of its central pixel; The multiple 3D cubes are input into an attention mechanism model for inference to obtain the semantic category of each 3D cube; Based on the semantic category, it is determined whether the driver is currently in a preset dangerous behavior state, and when the dangerous behavior state meets the preset triggering conditions, the vehicle warning mechanism is activated.
2. The driver behavior recognition method according to claim 1, characterized in that, The attention mechanism model includes a convolution-based feature extraction module, a Gaussian modulation attention module, and a complementary Transformer module. The steps involve inputting the multiple 3D cubes into the attention mechanism model for inference to obtain the semantic category of each 3D cube, including: The convolution-based feature extraction module sequentially performs 3D convolution, spectral compression and 2D convolution on the multiple 3D cubes to obtain spatial feature maps corresponding to each 3D cube. The Gaussian modulation attention module performs channel weighting and residual connection operations based on the spatial feature map to output a final feature map that takes into account both primary and secondary features. The complementary Transformer module performs global semantic modeling and behavior classification on the final feature map to obtain the semantic category of each three-dimensional cube.
3. The driver behavior recognition method according to claim 2, characterized in that, The steps involve using the convolution-based feature extraction module to sequentially perform 3D convolution, spectral compression, and 2D convolution on the multiple 3D cubes to obtain spatial feature maps corresponding to each 3D cube, including: By using 3D convolution, the spectral-spatial joint features in each of the 3D cubes are extracted to obtain the spectral-spatial joint features corresponding to each of the 3D cubes. The spectral-spatial joint features are processed by spectral compression to obtain the processed spectral-spatial joint features. Spatial feature modeling is performed on the processed spectral-spatial joint features through two-dimensional convolution to obtain spatial feature maps corresponding to each of the three-dimensional cubes.
4. The driver behavior recognition method according to claim 2, characterized in that, The spectral-spatial joint features corresponding to each of the three-dimensional cubes are represented as follows: in, For activation function, , and Representing the first The height, width, and depth of the 3D convolutional kernel. For the first A feature cube at position ( , , The weight parameter at position ) For the first Layer Zhang's feature map is located at ( , , The neuron activation value, For bias terms; And / or, the spatial feature map corresponding to each of the three-dimensional cubes is represented as follows: in, For activation function, and These represent the height and width of the two-dimensional convolution kernel, respectively. For the first Zhang's feature map is in position , The weight parameters at that location, For the first Layer Zhang's feature image is located at... , Spatial feature map, This is a bias term.
5. The driver behavior recognition method according to claim 2, characterized in that, The steps involve using the Gaussian modulation attention module to perform channel weighting and residual connection operations based on the spatial feature map, outputting a final feature map that takes into account both primary and secondary features, including: The spatial feature map is subjected to average pooling to obtain global context features; The global context features are subjected to linear transformation and nonlinear activation to generate intermediate features containing channel dependencies; Based on the intermediate features and the Gaussian modulation function, the weights of each channel are redistributed through a Gaussian distribution to enhance the secondary features, resulting in a modulation weight map. The modulation weight map and the spatial feature map are added pixel by pixel to output the final feature map.
6. The driver behavior recognition method according to claim 5, characterized in that, The step involves redistributing the channel weights based on the intermediate features and the Gaussian modulation function using a Gaussian distribution to enhance the secondary features, thereby obtaining a modulation weight map. This includes: The intermediate features are mapped to a Gaussian distribution to generate a modulation weight map; the intermediate features contain N channel response values, where N is a positive integer. Calculate the arithmetic mean of the N channel response values, and use it as the mean of the Gaussian distribution; Calculate the root mean square deviation of each channel response value relative to the arithmetic mean, and use it as the standard deviation of the Gaussian distribution; Substitute each channel response value into a Gaussian probability density function parameterized by the mean and standard deviation to calculate the corresponding weight element in the modulation weighting graph.
7. The driver behavior recognition method according to claim 2, characterized in that, The steps involve using the complementary Transformer module to perform global semantic modeling and behavior classification on the final feature map to obtain the semantic category of each 3D cube, including: The final feature map is flattened in space to obtain a two-dimensional feature matrix composed of multiple spatial location feature vectors; The two-dimensional feature matrix is projected onto a preset latent space dimension through a linear mapping layer to generate a content token sequence; A category token is appended to the beginning of the content token sequence to form a complete token sequence containing both a category token and a content token. The position code is added to each token in the complete token sequence to obtain an input sequence adapted to the complementary Transformer module; The input sequence is input to the complementary Transformer module, and after processing by a multi-head self-attention mechanism, layer normalization and multilayer perceptron, the output representation of the classification token is extracted, and the semantic category of the corresponding 3D cube is determined based on the output representation of the classification token.
8. The driver behavior recognition method according to claim 7, characterized in that, The method further includes: The input sequence is used to generate a query matrix, a key matrix, and a value matrix through a linear mapping. The query matrix, the key matrix, and the value matrix are processed separately to obtain multiple sub-attention outputs; The multiple sub-attention outputs are concatenated along the channel dimension and linearly transformed using weight parameters to obtain the self-attention branch output. The value matrix is reshaped into a spatial tensor form and input into a convolution function containing convolutional layers and batch normalization layers for local feature extraction to obtain the convolutional branch output. The output of the self-attention branch and the output of the convolution branch are added and fused element by element to obtain an enhanced feature representation, and the semantic category of the corresponding three-dimensional cube is determined based on the enhanced feature representation.
9. The driver behavior recognition method according to claim 8, characterized in that, The output of each sub-attention is represented as: in, For child attention output, For the query matrix, The key matrix, Let be the dimension of the key matrix. Let be the value matrix.
10. The driver behavior recognition method according to claim 1, characterized in that, After acquiring a hyperspectral image containing the driver, the method further includes: Principal component analysis is performed on the hyperspectral image to reduce its dimensionality, resulting in the dimensionality-reduced hyperspectral data.