Driving distraction identification method and device, equipment and storage medium
By introducing a driving distraction recognition model with deep separable convolution and multi-spectral attention mechanism, the traditional method has solved the shortage of timeliness and accuracy, and achieved efficient identification of driver slight behavior changes, significantly improving the accuracy and real-timeness of the recognition.
Patent Information
- Application Number
- CN202510459797.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Traditional driving distraction recognition solutions are difficult to balance the timeliness and accuracy of recognition, and are difficult to capture the driver's complex behavioral characteristics, especially minor driving behavior changes.
The driving distraction recognition model is adopted with a deep separable convolution and multi-spectral attention mechanism. The feature extraction is performed through the depth separable convolution layer, and the spectrum division and attention fusion are performed in combination with the multi-spectral attention layer. Finally, the input pooling layer and feature flattening layer are processed to generate the driving behavior feature vector to be identified.
It significantly improves the performance of distracted driving detection, especially a good balance between efficient real-time and high-precision detection, and can effectively capture small but crucial feature changes, greatly improving the accuracy of the model.
Smart Images

Figure CN119992522A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision and deep learning technology, and in particular to a method, device, equipment and storage medium for identifying driver distraction. Background Art
[0002] With the rapid development of intelligent transportation systems and autonomous driving technology, driving safety has become one of the focuses of social attention, and distracted driving behaviors, such as drivers using mobile phones, eating, talking to passengers, etc., significantly increase the risk of traffic accidents. Therefore, developing efficient and accurate methods for identifying distracted driving is of great practical significance.
[0003] In related technologies, traditional driver distraction recognition solutions, due to considerations of the operation and resource usage of vehicle-mounted equipment, usually perform feature extraction based on feature extraction algorithms such as Histogram of Oriented Gradients (HOG), Scale-Invariant Feature Transform (SIFT), and shallow neural network models, and then classify and identify driving distraction behaviors based on classifiers (such as support vector machines, decision trees, etc.), which can achieve the recognition of driving distraction to a certain extent.
[0004] However, the feature extraction method of traditional driving distraction recognition scheme has high computational complexity when processing high-dimensional data, which makes it difficult to meet the needs of real-time recognition. In addition, due to the limitations of the feature extraction method, it is difficult to fully capture the subtle characteristics of the driver's distracted behavior, resulting in low accuracy of driving distraction recognition. Summary of the invention
[0005] The present application provides a method, device, equipment and storage medium for identifying distracted driving, which are used to solve the technical problem that traditional distracted driving identification solutions are difficult to balance the timeliness and accuracy of identification.
[0006] The first aspect of the present application provides a method for identifying distracted driving, comprising: inputting original driver image data into a preset distracted driving identification model, performing preprocessing through an input layer, and obtaining candidate driver image data; The candidate driver image data is subjected to feature extraction through a depth-wise separable convolutional layer to obtain an initial driving behavior feature map; The initial driving behavior feature map is input into the multi-spectral attention layer for spectrum division, and the feature maps of each spectrum range are subjected to attention fusion operation to obtain a fused feature map; The fused feature map is input into the pooling layer and the feature flattening layer for processing to obtain the feature vector of the driving behavior to be identified; The feature vector of the driving behavior to be identified is input into the classifier to perform driving distraction behavior identification and obtain the driving distraction identification result.
[0007] A second aspect of the present application provides a driving distraction recognition device, comprising: a preprocessing unit, configured to input original driver image data into a preset driving distraction recognition model, perform preprocessing through an input layer, and obtain candidate driver image data; A feature extraction unit, used to extract features from the candidate driver image data through a depth-separable convolutional layer to obtain an initial driving behavior feature map; The frequency division attention unit is used to input the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and perform attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; A feature processing unit, used to input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; The recognition unit is used to input the feature vector of the driving behavior to be recognized into the classifier to recognize the distracted driving behavior and obtain the distracted driving recognition result.
[0008] A third aspect of the present application provides a driving distraction identification device, comprising: a memory and at least one processor, wherein instructions are stored in the memory; and at least one processor calls the instructions in the memory so that the driving distraction identification device executes the above-mentioned driving distraction identification method.
[0009] A fourth aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned method for identifying distracted driving.
[0010] In the technical solution provided by the present application, by introducing a driving distraction recognition model with deep separable convolution and multi-spectral attention mechanism, the performance of distracted driving detection is significantly improved, especially a good balance is achieved between efficient real-time and high-precision detection, and the traditional driving distraction recognition solution relies on manually designed local feature information such as edges and corners, but is insufficient in capturing the complex behavioral characteristics of drivers, especially small changes in driving behavior, and shallow neural networks are difficult to deeply understand the deep semantic information in image data. The present application significantly reduces the computational complexity and parameter amount through deep separable convolution, while effectively retaining the feature extraction capability, thereby meeting the dual requirements of real-time and accuracy. The multi-spectral attention mechanism can adaptively weight according to the different spectral features of the input image data, thereby enhancing the model's sensitivity to key driving behavior features, overcoming the defect of traditional methods in insufficient capture of instantaneous behavior, thereby effectively capturing these small but crucial feature changes, and greatly improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a schematic diagram of a first embodiment of a method for identifying distracted driving in this application; Figure 2 This is a schematic diagram of a second embodiment of the method for identifying distracted driving in this application; Figure 3 A schematic diagram of a network architecture of a distracted driving recognition model for this application; Figure 4 A schematic diagram of a third embodiment of the method for identifying distracted driving in this application; Figure 5 A schematic diagram of an embodiment of a driving distraction identification device in the present application; Figure 6 It is a schematic diagram of another embodiment of the driving distraction identification device in the present application; Figure 7 This is a schematic diagram of an embodiment of a distracted driving identification device in the present application. DETAILED DESCRIPTION
[0012] The present application provides a method, device, equipment and storage medium for identifying distracted driving, which are used to ensure the accuracy and real-time performance of distracted driving identification.
[0013] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0014] It should be noted that the image data, videos and the privacy information that may be involved used in this application are implemented with the authorization of the identified user and / or in compliance with relevant legal provisions.
[0015] For ease of understanding, the specific process of this application is described below. Figure 1 , an embodiment of the method for identifying distracted driving in the present application includes: 101. Input the original driver image data into a preset driving distraction recognition model, perform preprocessing through the input layer, and obtain candidate driver image data.
[0016] It is understandable that the execution subject of this application can be a distracted driving identification device, or various driving terminals (such as vehicles) or driving assistance systems, intelligent driving systems, automatic driving systems, etc., and in appropriate circumstances, various road monitoring terminals and systems, which are not limited here. This application is explained by taking a vehicle as the execution subject.
[0017] In this embodiment, the original driver image data can come from relevant videos collected by a vehicle-mounted camera or other channels, and the image data records the driver's status. For example, a vehicle-mounted camera installed above the windshield, near the rearview mirror, or below the dashboard captures the driver's real-time image data to ensure that the driver's facial expressions, head posture, eye movements, hands and other body parts can be clearly captured.
[0018] It can be understood that the distracted driving recognition model of the present embodiment can be applied to the distracted driving behavior recognition of a single image or the distracted driving behavior recognition of the entire video, and the present embodiment does not specifically limit this.
[0019] It should be noted that for the recognition of distracted driving behavior in a single image, only one frame of original driver image data can be input, or a preset number of consecutive frames of original driver image data can be input to improve the single-frame recognition accuracy. For example, 10 frames of images with motion changes are input, and the 5th frame is recognized. The model extracts the timing change information in these 10 frames, and combines the timing change information to adjust the attention weight to determine whether there is distracted driving behavior in the 5th frame.
[0020] It is understandable that the network architecture provided in this embodiment can be used for both the recognition of distracted driving behavior applied to videos and single image data. For the scenario of whether there is distracted driving behavior in the entire video, the architecture provided in this embodiment can use multi-frame image data to obtain temporal changes by differentiating or fusing the information of continuous frames, that is, by extracting the inter-frame changes of each feature map to obtain inter-frame change features, and further through the accuracy of the allocation of each attention weight in the multi-spectral attention layer. In order to extract inter-frame change features, it is possible to further consider combining optical flow, recurrent neural network (RNN), long short-term memory network (LSTM) or Transformer architecture to improve the accuracy of temporal change information extraction. For details, it can be executed with reference to conventional solutions, and no further elaboration is given here.
[0021] For the scenario of whether there is distracted driving behavior in the entire video, the evaluation result can be a comprehensive evaluation result of the recognition results of each frame of the image, for example, the proportion of the number of image frames identified as distracted driving behavior in the total number of frames, the duration, frequency, etc., to improve the accuracy of the final evaluation; it can also include the total number of frames of all images identified as distracted driving behavior, the frame number, and the behavior type (that is, the classification label) to mark whether there is distracted driving behavior in the video, and / or all image frames with distracted driving behavior, without specific limitation.
[0022] To facilitate understanding, the following explanation will be given using a single-image distracted driving behavior recognition scenario as an example, that is, the original driver image data is input, one frame of the image is evaluated, and the remaining images are used to identify temporal change information to optimize the distribution of attention fusion operations.
[0023] The driving distraction recognition model of this embodiment introduces deep separable convolution and multi-spectral attention mechanism to provide an efficient and accurate distracted driving detection solution, which is suitable for real-time monitoring and detection of driver distraction behavior.
[0024] In this embodiment, the network architecture of the driving distraction recognition model includes an input layer (or image data processing layer), a depthwise separable convolution layer (or feature extraction layer), a multi-spectral attention layer, a pooling layer, a feature flattening layer and a classifier, which are connected in sequence.
[0025] In a feasible implementation, the original driver image data is adjusted to the standard required by the subsequent depthwise separable convolutional layer processing through the input layer to be compatible with changes in various situations, so as to ensure that the driving distraction recognition model can be stably inferred. Exemplarily, the preprocessing of the input layer includes but is not limited to cropping, rotation, flipping and normalization techniques.
[0026] It should be noted that the driving distraction recognition model can be a pre-trained model, or the model parameters can be automatically iterated and optimized step by step during use to improve the recognition accuracy. This embodiment does not limit the training process of the model.
[0027] 102. The candidate driver image data is subjected to feature extraction through a depthwise separable convolutional layer to obtain an initial driving behavior feature map.
[0028] In this embodiment, the depthwise separable convolution layer (DSC) mainly includes two steps: depthwise convolution and pointwise convolution. The depthwise convolution (or spatial convolution) performs convolution operations on each input channel separately without mixing the information between channels, while the pointwise convolution performs linear combination in the channel dimension to generate a new feature map, that is, the initial driving behavior feature map. Compared with the standard convolution, the depthwise separable convolution significantly reduces the amount of calculation and the number of parameters while maintaining good performance.
[0029] In this embodiment, the depth separable convolution layer may include at least one depth sub-volume module, each of which includes two steps: depth convolution and point-by-point convolution. This embodiment does not limit the specific number of depth sub-volume modules. For two or more depth sub-volume modules, they are connected by module cascade, that is, the output of the previous depth sub-volume module is the input of the next depth sub-volume module, until the last depth sub-volume module outputs the initial driving behavior feature map. Furthermore, the spatial dimensions of each depth sub-volume module can be set to increase in sequence to gradually improve the accuracy of feature extraction.
[0030] It is understandable that as the number of deep volume modules increases, the accuracy of feature extraction increases, and the consumption of computing resources also increases. In practical applications, an appropriate number of deep volume modules can be selected to extract the initial driving behavior features.
[0031] Optionally, the depthwise separable convolutional layer includes multiple depth sub-volume modules, which input the candidate driver image data into each depth sub-volume module in sequence, perform spatial convolution operation on the input feature map through deep convolution, and integrate cross-channel information through point-by-point convolution and then input it into the next depth sub-volume module, until the last deep sub-volume module outputs the initial driving behavior feature map.
[0032] Optionally, the depthwise separable convolution layer includes a depthwise convolution module. The depthwise convolution uses a 3×3 convolution kernel and performs convolution within each input channel. The formula is expressed as:
[0033] in, Represents the feature map output after the deep convolution operation, is the depth convolution kernel, Represents the convolution operation; Represents the input feature map; Point-by-point convolution uses a 1×1 convolution kernel for channel integration, and the formula is expressed as:
[0034] in, Represents the feature map output after point-by-point convolution operation, It is a deep convolution kernel, which is used to integrate the channels of the feature map after deep convolution.
[0035] That is, when the depth-separable convolutional layer includes a depth-wise convolutional module, its formula can be expressed as:
[0036] in, Represents the initial driving behavior feature map output by the depthwise separable convolutional layer.
[0037] In this embodiment, the structure adopted by the depthwise separable convolutional layer significantly reduces the computational complexity and number of parameters of the model while maintaining the effect of feature extraction. The computational complexity is reduced from Down to ,in is the size of the convolution kernel, that is, the spatial dimension of the convolution kernel, is the number of input channels, is the number of output channels.
[0038] In this embodiment, the initial driving behavior feature map is used to indicate the map of shallow features, or relatively low-level features, extracted after processing by the depth-separable convolution layer, such as directional edges, color changes, and local textures. These features are crucial for distinguishing the background from key parts (such as hands and faces). For example, edge detection helps to identify hand contours, facial features (such as eyes and mouths), and the boundaries of objects in the car, while color changes and local textures can help detect the contrast differences between hands and steering wheels, mobile phones, or other objects. Through these low-level features, the model can preliminarily determine whether the driver's hands are on the steering wheel, or whether there are obvious handheld objects, such as mobile phones or beverage bottles.
[0039] It should be noted that the introduction of inter-frame variation features can further improve the accuracy of distracted driving behavior recognition. For example, it has reference value for the allocation of attention weights in the subsequent multi-spectral attention mechanism. This embodiment does not limit the specific number of the indicated initial driving behavior feature graphs. Usually, each frame image corresponds to an initial driving behavior feature graph to calculate the temporal variation information between consecutive frames, that is, the inter-frame variation features, to better guide the accurate recognition of instantaneous behavior features.
[0040] 103. The initial driving behavior feature map is input into the multi-spectral attention layer for spectrum division, and the feature maps of each spectrum range are subjected to attention fusion operation to obtain a fused feature map.
[0041] In this embodiment, the Multi-Spectral Attention (MSA) layer mainly includes three steps: spectrum division, attention weight calculation and feature weighted fusion. The initial driving behavior feature map is divided into multiple spectrum ranges, and the feature map of each spectrum range is assigned attention weights based on the attention mechanism and weighted fused to obtain a fused feature map. This embodiment uses the multi-spectral attention mechanism to make up for the shortcomings of insufficient feature information in the existing channel attention mechanism, introduces more frequency components to make full use of feature information, and dynamically adjusts the weights of different spectrum features, thereby enhancing the model's sensitivity and recognition ability to key features and improving the accuracy of distracted driving detection.
[0042] It is understandable that the multi-spectral attention mechanism can adaptively weight according to the different spectral features of the input image data, enhancing the model's sensitivity to key driving behavior features. The characteristics of distracted driving are often manifested as subtle changes in the driver's behavior, such as momentary deviations of the eyes, slight turns of the head, or involuntary movements of the hands. These subtle changes are difficult to capture in traditional methods. However, after adopting the multi-spectral attention mechanism, the network can adaptively weight features in different spectral ranges, assigning larger weights to spectral features that are more associated with distracted driving behavior, and assigning smaller weights to spectral features that are less associated with distracted driving behavior, thereby effectively capturing these subtle but crucial feature changes and greatly improving the accuracy of the model.
[0043] In this embodiment, the multi-spectral attention layer can perform spectrum division using spectrum conversion functions such as discrete cosine transform function, discrete Fourier transform function, wavelet transform function, etc., without specific limitation.
[0044] After obtaining the feature graphs of each spectrum range through the above-mentioned spectrum conversion function, the importance of the feature graphs of each spectrum range can be determined through the preset attention mechanism, and then the effects of the feature graphs of different spectrum ranges can be determined, so as to learn to allocate the attention weight to the spectrum features that are more relevant to the recognition of distracted driving behavior. By extracting the features of the image through this network, the semantic features of the image can be obtained, and the tiny feature position information of the image space can be retained. The global context view can selectively aggregate the context, so that the key tiny features can be more compact and aggregated, and the subsequent classification accuracy can be improved with a smaller number of parameters.
[0045] The multi-spectral attention layer of this embodiment processes the initial driving behavior feature map output by the DSC. Compared with processing the ordinary feature map output by the conventional convolutional neural network (CNN), it can make fuller use of the technical advantages of DSC. DSC first effectively screens and optimizes the features, and then sends them to the MSA for global feature enhancement, thereby reducing the amount of calculation of the MSA and improving the effect of attention enhancement. This method is more effective than using MSA alone, and can more accurately capture key areas in the driving scene (such as gestures, facial movements, etc.). At the same time, it is more lightweight than the MSA solution that directly acts on CNN features, so that better recognition effects can still be obtained when computing resources are limited. By using more spectral variable information provided in the feature map of each spectral range to more accurately allocate attention weights, the attention to the associated features of driving distraction behavior in the fusion feature map can be effectively improved.
[0046] Compared with the traditional CNN structure, the distracted driving recognition model of this embodiment has been optimized in many aspects of the network architecture. The traditional CNN has high computational complexity. This architecture has made some optimizations in the combination of computational efficiency, feature extraction capability and attention mechanism, which can improve recognition accuracy while ensuring computational efficiency, and is particularly suitable for fine-grained classification tasks such as distracted driving detection.
[0047] In this embodiment, the fused feature map is used to indicate the intermediate features after further processing by the multi-spectral attention layer. The DSC at this stage is combined with the MSA mechanism to extract more complex structural features and behavior patterns. For example, the model can detect dynamic changes in the position of the hand to determine whether the hand has left the steering wheel or is operating the mobile phone. In addition, facial posture information is also effectively extracted at this level, such as the driver's line of sight direction, head angle, etc., to help determine whether the driver is paying attention to the road. At the same time, the multi-spectral attention mechanism (MSA) enhances the focus on key areas, allowing the model to more accurately distinguish different object interaction modes. For example, if the hand features frequently overlap with a highlighted area (possibly a mobile phone), it can be inferred that the driver is using a mobile phone.
[0048] 104. Input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain the feature vector of the driving behavior to be identified.
[0049] This embodiment downsamples the feature map processed by the attention mechanism through the pooling layer to effectively reduce the spatial dimension of the feature map, while retaining important spatial information and reducing the computational complexity of subsequent processing. The feature flattening layer is used to convert the downsampled feature map into a one-dimensional feature vector to ensure that the feature vector of the driving behavior to be identified is adapted to the subsequent classifier, which is convenient for subsequent classification and detection tasks.
[0050] In this embodiment, the driving behavior feature vector to be identified is at the back end of the network, and the pooling layer and MSA further extract high-level features. The pooling layer may include global max pooling (GMP), so that the model has the ability to understand the driver's overall behavior pattern. At this time, the model can not only recognize the local information of the driver's hands and face, but also infer the complete driving behavior pattern based on these features. For example, combined with hand, face, and body posture information, the model can accurately determine whether the driver is performing typical distracted driving behaviors, such as making phone calls, eating, or lowering his head to operate the central control screen. At the same time, MSA can also help the model better understand the global context and reduce environmental interference, such as changes in lighting in the car and interference from background objects (such as passengers in the co-pilot), thereby improving the robustness of recognition.
[0051] 105. Input the feature vector of the driving behavior to be identified into the classifier to perform driving distraction behavior identification, and obtain a driving distraction identification result.
[0052] The classifier of this embodiment is responsible for real-time detection and classification of distracted driving behavior and outputting the final detection result, which may include: forward propagating the flattened feature vector through the classifier to obtain the category probability distribution, converting the output into a probability distribution through the Softmax function, and selecting the category with the highest probability as the final prediction result.
[0053] It can be understood that the distracted driving recognition model of this embodiment covers three levels in terms of feature extraction: low-level features (such as edges, colors, and textures) in the initial driving behavior feature map, intermediate features (such as gestures, faces, and object interactions) in the fused feature map, and high-level features (global behavior patterns) in the feature vector of the driving behavior to be identified, so that it can accurately capture information related to distracted driving behavior. Among them, low-level features help distinguish between the background and key parts, intermediate features establish the driver's behavior pattern, and high-level features ultimately complete the judgment of the driver's behavior.
[0054] In this embodiment, the driving analysis recognition result is used to indicate whether the driver in the image has distracted driving behavior. The driving behavior labels (or categories) in this embodiment may include but are not limited to: safe driving, texting using right hand, talking on the phone using right hand, texting using left hand, talking on the phone using left hand, operating the radio, drinking water, reaching behind, doing hair and makeup, and talking to passengers.
[0055] Optionally, when the driving distraction recognition model recognizes that the driver is distracted from driving, it can further remind the driver through various prompting methods or enable corresponding assisted driving solutions, such as controlling the vehicle speed to reduce, to ensure driving safety, without specific restrictions.
[0056] It can be understood that the driving distraction recognition model provided in this embodiment can be applied to the recognition of real-time behavior or non-real-time behavior, such as inputting a violation video for image extraction and then identifying whether there is driving distraction recognition.
[0057] In this embodiment, by introducing a driving distraction recognition model with deep separable convolution and multi-spectral attention mechanism, the performance of distracted driving detection is significantly improved, especially a good balance is achieved between efficient real-time and high-precision detection, and the traditional driving distraction recognition scheme relies on manually designed local feature information such as edges and corners, but is insufficient in capturing the complex behavior characteristics of drivers, especially small changes in driving behavior. Traditional shallow neural networks are difficult to deeply understand the deep semantic information in image data. The present application significantly reduces the computational complexity and parameter amount through deep separable convolution, while effectively retaining the feature extraction capability, thereby meeting the dual requirements of real-time and precision. The multi-spectral attention mechanism can adaptively weight according to the different spectral features of the input image data, thereby enhancing the model's sensitivity to key driving behavior features, overcoming the defect of traditional methods in insufficiently capturing instantaneous behavior and small behaviors, thereby effectively capturing these small but crucial feature changes, greatly improving the accuracy of the model.
[0058] See also Figure 2 and Figure 3Another embodiment of the method for identifying distracted driving in the present application includes: 201. Input the original driver image data into a preset driving distraction recognition model, perform preprocessing through the input layer, and obtain candidate driver image data.
[0059] Step 201 can be performed with reference to step 101 and will not be described in detail here.
[0060] 202. Perform feature extraction on the candidate driver image data through a depthwise separable convolutional layer to obtain an initial driving behavior feature map.
[0061] Reference Figure 3 Schematic diagram of the network architecture, in this embodiment, the depth separable convolution layer is explained by taking two cascaded depth sub-volume modules as an example, and each depth sub-volume module includes depth convolution and point-by-point convolution. That is, the depth separable convolution layer includes a first depth sub-volume module and a second depth sub-volume module, and the spatial dimension corresponding to the convolution kernel in the first depth sub-volume module is smaller than the spatial dimension corresponding to the convolution kernel in the second depth sub-volume module.
[0062] Exemplarily, the candidate driver image data is input into the first deep convolution module, and a spatial convolution operation is performed on the input feature map through deep convolution, and the cross-channel information is integrated through point-by-point convolution, and then input into the second deep convolution module; the second deep convolution module performs a spatial convolution operation on the input feature map through deep convolution, and the cross-channel information is integrated through point-by-point convolution to obtain an initial driving behavior feature map.
[0063] Optionally, the first depth convolution module can adopt a 1*1 convolution kernel, and the second depth convolution module can adopt a 3*3 convolution kernel.
[0064] It is understandable that this embodiment uses a two-layer cascaded DSC structure in terms of deep separable convolution (DSC), rather than a traditional single-layer DSC, in which the first layer of DSC performs channel transformation through 1×1 convolution to reduce computational redundancy and improve expression ability, while the second layer of DSC performs local feature extraction through 3×3 convolution to improve the network's spatial perception ability. This two-layer DSC structure has stronger nonlinear expression ability than a single-layer DSC, while still maintaining a low amount of computation. Compared with the standard CNN, it greatly reduces the number of parameters and computational complexity, while enhancing the feature extraction capability, and is particularly suitable for tasks such as distracted driving detection that require high fine-grained features.
[0065] The network architecture of this embodiment is optimized on the object of MSA, using the first DSC to reduce the amount of calculation, and using the second DSC to improve the feature extraction capability, so that it achieves a better balance between calculation efficiency and attention enhancement effect. Compared with the ordinary single-layer DSC+MSA structure, the use of double-layer DSC improves the local feature extraction capability, and combines with MSA for global feature enhancement, which can improve recognition accuracy while ensuring calculation efficiency, and is particularly suitable for fine-grained classification tasks such as distracted driving detection.
[0066] 203. Divide the spectrum according to the corresponding spectrum range through each spectrum module, extract the global information of each spectrum feature graph, and obtain the global feature graph corresponding to each spectrum range.
[0067] The multi-spectral attention layer of this embodiment may include multiple spectrum modules, attention mechanism modules and fusion modules. Each spectrum module independently processes the features of different spectrum ranges, calculates the attention weights of the global feature maps corresponding to each spectrum range based on the attention mechanism, and dynamically adjusts the weights of each spectrum feature in the final feature fusion process.
[0068] It should be understood that one spectrum module corresponds to one spectrum range, and each spectrum module outputs a spectrum feature graph of its corresponding spectrum range. The number of spectrum feature graphs is equal to the number of spectrum modules, and the spectrum range corresponding to each spectrum module can be set according to actual conditions. This embodiment divides the spectrum range so that the model can more finely capture the driver's behavioral characteristics in different spectrum ranges, especially those small but critical behavioral changes, thereby significantly improving the accuracy and robustness of distracted driving detection.
[0069] Specifically, the initial driving behavior feature map is divided according to a number of preset spectrum ranges to obtain multiple spectrum feature maps; and the global information of each spectrum feature map is extracted through global average pooling and global maximum pooling operations.
[0070] For example, assuming there are D spectrum modules, each spectrum feature graph can be expressed as , Each spectrum module may include a first 2D convolution, a ReLU function, a second 2D convolution and a Sigmoid function, wherein the first 2D convolution and the second 2D convolution may use a 1*1 convolution kernel.
[0071] Optionally, taking the cosine discrete transformation function as an example, the processing process of each spectrum module is explained: the above-mentioned initial driving behavior characteristic map is divided according to several preset spectrum ranges to obtain multiple spectrum characteristic maps, including: the initial driving behavior characteristic map is changed according to the preset spectrum range through the cosine discrete transformation function, the coordinates in the spatial domain are mapped with the coordinates in the frequency domain to obtain the corresponding frequency information, and the high-frequency area in the spectrum characteristic map generally stores feature information that is strongly correlated with the identification of distracted driving behavior.
[0072] In a feasible implementation, the initial driving behavior feature map is evenly split into multiple parts along the channel dimension, a frequency index is assigned to each part, and the frequency feature is obtained by DCT transformation according to the assigned frequency index. The larger the index, the higher the frequency and the faster the change. The high-frequency information can be selected to extract micro-behavior features and instantaneous behavior features, such as instantaneous deviation of the eyes, micro-rotation of the head, or involuntary movements of the hands.
[0073] The above global average pooling formula is as follows:
[0074] in, Used to represent the height of the feature map, Used to indicate the width of the feature map; Indicates The characteristic graph height of the spectrum, Indicates The characteristic graph width of the spectrum.
[0075] Exemplarily, the global maximum pooling formula is as follows:
[0076] 204. Calculate the attention weight of the global feature map corresponding to each spectrum range through the attention mechanism module.
[0077] Optionally, the pooled features are passed through a fully connected layer and an activation function to generate the attention weights for each spectral submodule. The formula is as follows:
[0078] in, Used to indicate the The attention weight of each spectrum; represents the Sigmoid activation function, Indicates a fully connected operation. The attention weight of each spectral feature graph in this embodiment reflects the importance of the spectral feature graph, so as to pay more attention to the features associated with distracted driving behavior during fusion.
[0079] Optionally, the attention mechanism module can further combine the temporal change information to allocate the attention weights of the global feature maps corresponding to each spectral range, wherein the temporal change information is used to indicate the feature changes between consecutive frames. This embodiment uses more spectral variable information provided in the feature map of each spectral range to improve the recognition accuracy of minor behaviors. Further combining the inter-frame change features can more accurately allocate attention weights. The resulting fused feature map can take into account the recognition of both instantaneous behaviors and minor behaviors, which is beneficial to the later classification of driving behaviors to improve the overall classification accuracy.
[0080] Exemplarily, the extraction of the above-mentioned time-series change information can be to extract the inter-frame change features of the initial driving behavior feature map corresponding to each frame image, that is, to identify the changes of the same features between consecutive frames. In actual applications, the driver's instantaneous and unconscious movements and behaviors will cause the features recognized as hands to change in position between the above-mentioned consecutive frames.
[0081] It should be understood that in the above scenario, the depth-separable convolution layer of step 202 outputs the initial driving behavior feature map of continuous frames. It is possible to choose to set a temporal change recognition layer between the depth-separable convolution layer and the multi-spectral attention layer, or set a temporal change recognition module in parallel with multiple spectrum modules in the multi-spectral attention layer, wherein the temporal change recognition layer or the temporal change recognition module can identify the temporal change information of each region in the continuous frames, that is, used to extract inter-frame change features.
[0082] In a feasible implementation, taking the example of setting a temporal change recognition layer between the depthwise separable convolution layer and the multi-spectral attention layer, the initial driving behavior feature map of continuous frames is extracted with the temporal change recognition layer to obtain the inter-frame change features, and the temporal change information is obtained; the temporal change information and the initial driving behavior feature map are input into the multi-spectral attention layer; the initial driving behavior feature map is divided according to the corresponding spectrum range through each spectrum module, and the global information of each spectrum feature map is extracted to obtain the global feature map corresponding to each spectrum range; the attention weight of the global feature map corresponding to each spectrum range is allocated in combination with the temporal change information through the attention mechanism module; each spectrum feature map is multiplied by the corresponding attention weight and then fused through the fusion module to obtain a fused feature map.
[0083] In a feasible implementation, taking a multi-spectral attention layer setting and a timing change recognition module in parallel with multiple spectrum modules as an example, the attention mechanism module is respectively connected to the timing change recognition module and each spectrum module. Specifically, the initial driving behavior feature map of continuous frames is input into the timing change recognition module to extract the inter-frame change features, and the timing change information is obtained and input into the attention mechanism module; The initial driving behavior feature map to be identified is input into each spectrum module for division according to the corresponding spectrum range, and the global information of each spectrum feature map is extracted to obtain the global feature map corresponding to each spectrum range, and then input into the attention mechanism module; The attention mechanism module is used to allocate attention weights to the global feature maps corresponding to each spectral range in combination with the temporal change information. The fusion module multiplies each spectral feature map with the corresponding attention weight and then fuses them to obtain a fused feature map.
[0084] It should be understood that if only a single image needs to be identified whether distracted driving behavior exists, then the input to each spectrum module is only the initial driving behavior feature map of a single frame to be identified; if the entire video or multiple images need to be identified whether distracted driving behavior exists, then the input to each spectrum module is the initial driving behavior feature map of each frame to be identified.
[0085] The above-mentioned continuous frames are used to indicate images of a continuous preset number of frames. For example, the initial driving behavior feature graph corresponding to the 1st frame to the 8th frame obtains the change of the temporal characteristics of the continuous frames by differentiating or fusing the information of the continuous frames.
[0086] To facilitate understanding, an example is provided: if the features of a local area (such as the driver's eyes, hands, or head) change significantly in consecutive frames, this usually means that there is dynamic behavior in the area (such as rapid eye movement, hands leaving the steering wheel, or involuntary hand movements). In this embodiment, the multi-spectral attention mechanism can use this temporal change information to assign higher attention weights to these dynamic areas during the feature fusion process, so that the model pays more attention to these important areas that may be related to distracted driving. On the other hand, by analyzing the information of different spectra, multi-spectral attention can identify which areas show high-frequency changes (which may correspond to rapid movements) or low-frequency changes (which may correspond to continuous behaviors) in the time domain, and then adjust the attention distribution of the corresponding areas. This adaptive adjustment on the spectrum enables the model to more accurately capture subtle changes in driver behavior.
[0087] 205. Each spectral feature map is multiplied by the corresponding attention weight through the fusion module and then fused to obtain a fused feature map.
[0088] Each spectral feature map is multiplied by the corresponding attention weight and then fused. The formula is as follows:
[0089] This embodiment can enhance the model's ability to focus on key features and improve the accuracy of distracted driving detection by dynamically adjusting the weights of different spectral features.
[0090] 206. Input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified.
[0091] Perform the maximum pooling operation on the fused feature map, and the formula is expressed as:
[0092] in, Represents the pooling feature map, which is used to represent the output after the maximum pooling operation. is the spatial position index of the fused feature map; represents the fused feature map, which is used to indicate the output after the fusion of each spectrum feature map; g represents the offset of the pooling core in the vertical direction, and q represents the offset of the pooling core in the horizontal direction; the pooling feature map is flattened into a one-dimensional vector to obtain the feature vector of the driving behavior to be identified.
[0093] For example, using a pooling kernel of 28×28 and a stride of 1, the value range of g, q is [0, 27].
[0094] The feature map Flattened to a one-dimensional vector , its formula is expressed as:
[0095] Assuming that the size of the feature map after pooling is C×28×28, the length of the flattened feature vector is C×28×28.
[0096] In this embodiment, downsampling through the maximum pooling operation can effectively reduce the spatial dimension of the feature map, while retaining important spatial information, reducing the amount of computation for subsequent processing, and the downsampled feature map is flattened and converted into a one-dimensional feature vector to meet the input requirements of the subsequent classifier.
[0097] 207. Input the feature vector of the driving behavior to be identified into a classifier to perform driving distraction behavior identification, and obtain a driving distraction identification result.
[0098] In this embodiment, the classifier architecture may include a first batch of normalization layers, a first fully connected layer, a Dropout layer, a second batch of normalization layers, an ELU activation function, and a second fully connected layer.
[0099] Specifically, the driving behavior feature vector to be identified is standardized through the first normalization layer, and mapped to the target dimension through the first fully connected layer to obtain the preferred driving behavior feature vector; the preferred driving behavior feature vector is randomly set to zero with a preset probability through the Dropout layer, standardized through the second normalization layer, activated through the ELU activation function, and then input into the second fully connected layer for mapping to obtain the driving distraction recognition result.
[0100] In this embodiment, the flattened feature vector is forward propagated through the classifier to obtain the category probability distribution, the output is converted into a probability distribution through the Softmax function, and the category with the highest probability is selected as the final prediction result.
[0101] The above batch normalization layer can speed up the training process, stabilize the model, and reduce internal covariate shift.
[0102]
[0103] in, Represents input; Expressed as the mean of the batch; Expressed as the variance of the batch, and is a trainable parameter, Small constant to prevent division by zero.
[0104] The first fully connected layer (Fully Connected Layer) above uses linear transformation to map the input feature vector to a lower dimensional space. The formula is as follows:
[0105] in, is the weight matrix, is the bias vector, Represents the feature vector input to the fully connected layer. Exemplarily, the standardized feature vector is mapped to a 512-dimensional space.
[0106] The above Dropout layer randomly sets the input unit to zero with a certain probability during training to prevent overfitting. The formula is as follows:
[0107] Where x represents input, and p represents the discard probability. For example, the input unit is randomly set to zero with a probability of 50%, that is, p=0.5, to prevent overfitting and improve the generalization ability of the model.
[0108] The above ELU activation function (Exponential Linear Unit) can enhance nonlinear expression and is defined as follows:
[0109] Among them, x represents the input; α is an adjustable parameter, usually set to 1.0; ELU can enhance the nonlinear expression ability of the model and accelerate training.
[0110] The second fully connected layer maps the feature vector to the final multiple category outputs, and its formula is:
[0111] in, represents the first fully connected layer; To represent the second fully connected layer, represents the first batch of normalization layers; To represent the second batch of normalized layers, the weight matrix and bias parameters of the second fully connected layer can be set differently from those of the first fully connected layer. The second fully connected layer maps the feature vector to a final target number J of category outputs, such as 10 categories or other numbers, without specific limitation.
[0112] The Softmax function converts the output into a probability distribution for easy classification. Its formula is:
[0113] in, is the probability of the predicted category; is the category index; is the probability of the mth category currently being paid attention to, that is, the one-hot encoding of the true label; It represents the probability of predicting the class with index j out of the total number of classes J.
[0114] In order to further improve the performance of classification and detection, the following mechanisms can also be introduced into the classifier: such as residual connections, attention mechanisms, multi-layer perceptrons (MLP) and normalization strategies.
[0115] The above residual connections add residual connections between fully connected layers to alleviate the gradient vanishing problem and facilitate the training of deep networks.
[0116]
[0117] The above-mentioned attention mechanism embeds the self-attention mechanism in the classifier, so that the model can dynamically focus on different parts of the feature vector during the classification process, thereby improving the accuracy of classification. Its formula is:
[0118] Among them, Q, K, and V are query, key, and value matrices respectively. The dimension of the key.
[0119] The above-mentioned Multi-Layer Perceptron (MLP) adds multiple fully connected layers and nonlinear activation functions to the classifier to enhance the expressive power of the model.
[0120]
[0121] The above normalization strategy further stabilizes the training process and prevents overfitting by introducing a strategy combining batch normalization and Dropout.
[0122] Through the above design, the classifier not only has the basic batch normalization, full connection, Dropout and activation functions, but also can significantly improve the classification performance and stability of the model through residual connection, attention mechanism and multi-layer perceptron structures.
[0123] In this embodiment, by introducing a driving distraction recognition model with deep separable convolution and multi-spectral attention mechanism, the performance of distracted driving detection is significantly improved, especially a good balance is achieved between efficient real-time and high-precision detection, and the traditional driving distraction recognition solution relies on manually designed local feature information such as edges and corners, which is insufficient in capturing the complex behavioral characteristics of drivers, especially the slight changes in driving behavior, and the problem that shallow neural networks are difficult to deeply understand the deep semantic information in image data. In this embodiment, a two-layer cascaded DSC structure is used in the deep separable convolution. The first layer of DSC module reduces computational redundancy and improves expression ability through channel transformation, and the second layer of DSC performs local feature extraction to improve the spatial perception ability of the network. The deep separable convolution layer of this embodiment has stronger nonlinear expression ability while still maintaining a low amount of computation, which is particularly suitable for tasks such as distracted driving detection that require high fine-grained features. By optimizing the object of MSA, dividing the spectrum of the initial driving behavior feature map, and combining global average pooling, global maximum pooling and attention mechanism to weightedly fuse the spectrum feature maps, the model can more finely capture the driver's behavioral characteristics in different spectrum ranges, especially those small but critical behavioral changes, thereby significantly improving the accuracy and robustness of distracted driving detection. Furthermore, the MSA layer of this embodiment can further combine the time series change information to guide the attention allocation of multi-spectral feature maps, so that the model can pay more attention to instantaneous behavior and small behavior. Finally, the classifier significantly improves the classification performance and stability of the model through batch normalization, full connection, Dropout and activation functions and other optimization mechanisms.
[0124] See also Figure 4 The third embodiment of the method for identifying distracted driving in the present application describes the training process of the distracted driving identification model of the present application, including: 401. Build training set and validation set.
[0125] Specifically, the driver's behavior video data or image data is collected through a vehicle-mounted camera, the sample video data or sample image data is preprocessed, and the training set and the verification set are divided according to a preset ratio.
[0126] It can be understood that the present embodiment can collect sample video data or sample image data of driver behavior through the vehicle-mounted camera to cover videos of different time periods (day, night), weather conditions (sunny, rainy), road scenes (urban roads, highways) and individual differences of drivers (age, gender, accessories worn, etc.); and pre-mark corresponding labels, including normal driving (holding the steering wheel with both hands), distracted behavior (using mobile phones, eating, turning heads to talk, etc.).
[0127] The above preprocessing may include but is not limited to extracting frames from video data to obtain corresponding images, and performing image adjustment, random cropping and padding, random horizontal flipping, and standardization on each image, and then randomly dividing the data in an 8:2 ratio to ensure that the two data sets are balanced in terms of illumination, viewing angle, and distraction behavior category distribution. The validation set may also contain additional extreme scene data (such as strong glare and severe occlusion) to evaluate the robustness of the model.
[0128] In a feasible implementation, the sample image data is preprocessed, including: adjusting each sample image to a preset first target pixel, performing random cropping and padding on the adjusted image, and further performing random horizontal flipping and normalization to obtain sample image data to enhance the generalization ability and robustness of the model.
[0129] For example, the input image data may be scaled to 256×256 pixels. This process uses a bilinear interpolation method, which can smoothly adjust the image size and retain more detail information.
[0130] Exemplarily, the random cropping and padding is performed on a 256×256 image by randomly cropping it to 224×224 pixels, and a 4-pixel padding is applied during the cropping process to increase the diversity of the image data.
[0131] Exemplarily, the random horizontal flipping flips the image horizontally with a probability of 50% to simulate different perspective changes of the driver.
[0132] Exemplarily, the above-mentioned standardization process normalizes the pixel values of the image data to the range of [-1, 1], and eliminates the influence of illumination changes by converting the pixel values to the same scale range; each channel is standardized separately (R, G, B channels) to avoid deviations between different channels and enhance the adaptability of the model to color changes under different illumination environments. The formula is expressed as:
[0133] in, Indicates the input The image is standardized. The input image includes R, G, and B channels. The mean of the R, G, and B channels, Indicates the standard deviation of each of the R, G, and B channels.
[0134] For example, you can set , , whose values correspond to the three channels of RGB respectively. In this embodiment, the influence of illumination change is eliminated by converting the RGB channels of the pixel into the same scale range.
[0135] For example, in order to adapt to changes in lighting conditions, this embodiment uses brightness changes and normalization processing to simulate different lighting environments such as daytime, nighttime, cloudy days, and strong light. This helps the model extract effective features under complex lighting conditions and improves its robustness.
[0136] Exemplarily, in order to adapt to changes in perspectives of different cameras, this embodiment introduces operations such as random cropping, random padding, and horizontal flipping to simulate changes in perspective caused by factors such as camera angle offset or vehicle shaking during driving, thereby improving the model's adaptability in multi-angle scenarios.
[0137] For example, considering the various interference factors that may exist during driving, the complexity of the driving environment, such as the adjustment of the seat position in the car, the occlusion of the sun visor or rearview mirror, and the driver wearing glasses, hats, masks, etc., all of which will interfere with the image features. This embodiment can effectively increase the robustness of the model to these changes through data enhancement strategies.
[0138] In this embodiment, the preprocessing step is intended to enhance the generalization and robustness of the model, and brightness changes and standardization processing help the model to perform stable reasoning under different lighting conditions (such as daytime, nighttime, cloudy days, strong light, etc.). Random cropping, padding, and horizontal flipping simulate different perspective changes to help the model adapt to the possible camera angle changes during driving. The above steps ensure that the model can adapt to different lighting, perspectives, and driving environments. The preprocessing taken by the model input layer is to input unified and standardized candidate image data. The actual detection environment of the model is based on the real original image data. Its preprocessing, such as adjusting the image data to a preset size and other standardized processing procedures.
[0139] 402. Train the initial model using the training set, and perform back propagation optimization on the model parameters using a preset loss function to obtain a candidate model.
[0140] Specifically, the training set is input into the initial model for driving distraction recognition to obtain a predicted label; the loss value between the predicted label and the true label is calculated by a preset cross entropy loss function; the model parameters are updated by a preset optimization strategy and loss value to obtain a candidate model. This embodiment adjusts the model parameters through repeated iterative training to gradually reduce the cross entropy loss function, thereby continuously improving the accuracy of detection and the stability of the system.
[0141] Exemplarily, the above loss function may adopt the formula of cross-entropy loss function:
[0142] Among them, L is the cross entropy loss function; J is the total number of categories; is the probability of the true label, that is, the probability of the category with the actual classification index o, which can be represented by one-hot encoding, is the predicted probability, that is, the probability of predicting the category with classification index o.
[0143] Exemplarily, the back propagation may use a gradient descent method, such as Stochastic Gradient Descent (SGD), to update model parameters. Specifically, the gradient of the loss function with respect to the model parameters is calculated by the chain rule, and the model parameters are adjusted according to the gradient information and the learning rate.
[0144] Exemplarily, this embodiment may further introduce optimization strategies, such as momentum optimization and weight decay strategies. Specifically, by accumulating historical gradient information, momentum optimization can reduce the update step size in the steep direction and increase the update step size in the gentle direction. By limiting the size of the model weights, overfitting is prevented, thereby improving the generalization ability of the model. By penalizing larger weight values, weight decay can prompt the model to choose simpler solutions and avoid over-reliance on certain features. In high-dimensional space, weight decay helps maintain the numerical stability of model parameters and avoid numerical problems caused by excessive weight values.
[0145] Exemplarily, this embodiment may also introduce a cosine annealing learning rate scheduling strategy (Cosine Annealing) to dynamically adjust the learning rate, accelerate the convergence speed and improve the model performance, and the learning rate gradually decreases as the training progresses.
[0146] This embodiment further ensures the training efficiency and convergence performance of the model by adjusting the learning rate through back propagation based on the cross entropy loss function and a preset optimization strategy. While maintaining high efficiency and real-time performance, this embodiment achieves more accurate and stable recognition of distracted driving behavior.
[0147] 403. The candidate model is evaluated through the validation set, and the candidate model corresponding to the best model parameters is used as the driving distraction recognition model.
[0148] In this embodiment, the validation set can be used to calculate the accuracy, precision, recall rate, F1 score and other indicators of the candidate model for the classification task, and the candidate model corresponding to the best model parameters after the evaluation is confirmed as the driving distraction recognition model.
[0149] 404. Inputting the original driver image data into a preset driving distraction recognition model, preprocessing through an input layer, and obtaining candidate driver image data; 405. Extracting features from the candidate driver image data through a depthwise separable convolutional layer to obtain an initial driving behavior feature map; 406. Input the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and perform an attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; 407. Input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; 408. Input the feature vector of the driving behavior to be identified into a classifier to perform driving distraction behavior identification, and obtain a driving distraction identification result.
[0150] Steps 404-408 may be performed with reference to the above steps 101-105 and will not be described in detail here.
[0151] In this embodiment, the sample image data in the training set and the validation set are preprocessed to ensure that the model can adapt to different lighting, viewing angles and driving environments. The loss function is used for back propagation, and the efficiency of the training process is improved based on the optimization strategy. The distracted driving recognition model that has passed the evaluation is used for distracted driving recognition, which significantly improves the performance of distracted driving detection and achieves a good balance between efficient real-time and high-precision detection. This embodiment significantly reduces the computational complexity and parameter quantity through deep separable convolution, while effectively retaining the feature extraction capability, thereby meeting the dual requirements of real-time and precision. The multi-spectral attention mechanism can be adaptively weighted according to the different spectral features of the input image data, enhancing the model's sensitivity to key driving behavior features, overcoming the defect of traditional methods in insufficiently capturing instantaneous behaviors, thereby effectively capturing these tiny but crucial feature changes, and greatly improving the accuracy of the model.
[0152] The above describes the method for identifying distracted driving in this application. The following describes the device for identifying distracted driving in this application. Figure 5 , an embodiment of the driving distraction identification device in the present application includes: A preprocessing unit 501 is used to input the original driver image data into a preset driving distraction recognition model, and perform preprocessing through an input layer to obtain candidate driver image data; A feature extraction unit 502 is used to extract features from the candidate driver image data through a depthwise separable convolutional layer to obtain an initial driving behavior feature map; The frequency division attention unit 503 is used to input the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and perform attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; A feature processing unit 504 is used to input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; The identification unit 505 is used to input the feature vector of the driving behavior to be identified into the classifier to perform driving distraction behavior identification and obtain a driving distraction identification result.
[0153] In this embodiment, by introducing a driving distraction recognition model with deep separable convolution and multi-spectral attention mechanism, the performance of distracted driving detection is significantly improved, especially a good balance is achieved between efficient real-time and high-precision detection, and the traditional driving distraction recognition scheme relies on manually designed local feature information such as edges and corners, but is insufficient in capturing the complex behavior characteristics of drivers, especially small changes in driving behavior, and shallow neural networks are difficult to deeply understand the deep semantic information in image data. The present application significantly reduces the computational complexity and parameter amount through deep separable convolution, while effectively retaining the feature extraction capability, thereby meeting the dual requirements of real-time and precision. The multi-spectral attention mechanism can adaptively weight according to the different spectral features of the input image data, thereby enhancing the model's sensitivity to key driving behavior features, overcoming the defect of traditional methods in insufficient capture of instantaneous behavior, thereby effectively capturing these small but crucial feature changes, and greatly improving the accuracy of the model.
[0154] See also Figure 6 Another embodiment of the driving distraction identification device in the present application includes: A preprocessing unit 501 is used to input the original driver image data into a preset driving distraction recognition model, and perform preprocessing through an input layer to obtain candidate driver image data; A feature extraction unit 502 is used to extract features from the candidate driver image data through a depthwise separable convolutional layer to obtain an initial driving behavior feature map; The frequency division attention unit 503 is used to input the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and perform attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; A feature processing unit 504 is used to input the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; The identification unit 505 is used to input the feature vector of the driving behavior to be identified into the classifier to perform driving distraction behavior identification and obtain a driving distraction identification result.
[0155] Optionally, the depth separable convolution layer includes a first depth convolution module and a second depth convolution module, the spatial dimension corresponding to the convolution kernel in the first depth convolution module is smaller than the spatial dimension corresponding to the convolution kernel in the second depth convolution module, and the feature extraction unit 502 is specifically used to: input the candidate driver image data into the first depth convolution module, perform a spatial convolution operation on the input feature map through depth convolution, and integrate the cross-channel information through point-by-point convolution and then input it into the second depth convolution module; The second deep convolution module performs spatial convolution on the input feature map through deep convolution, and integrates cross-channel information through point-by-point convolution to obtain the initial driving behavior feature map.
[0156] Optionally, the multi-spectral attention layer includes multiple spectrum modules, attention mechanism modules and fusion modules; the frequency division attention unit 503 includes: The frequency division subunit 5031 is used to divide the corresponding spectrum range through each spectrum module, and extract the global information of each spectrum feature map to obtain the global feature map corresponding to each spectrum range; An attention calculation subunit 5032 is used to calculate the attention weight of the global feature map corresponding to each spectrum range through an attention mechanism module; The attention fusion subunit 5033 is used to fuse each spectrum feature map by multiplying the corresponding attention weight through a fusion module to obtain a fused feature map.
[0157] Optionally, the multi-spectral attention layer further includes a temporal change recognition module, and the frequency division attention unit 503 further includes: The time series change subunit 5034 is used to extract the inter-frame change features of the initial driving behavior feature graph of the continuous frames through the time series change recognition layer, obtain the time series change information, and input it into the attention mechanism module.
[0158] Optionally, the attention calculation subunit 5032 is also used for: the attention mechanism module allocates attention weights of the global feature maps corresponding to each spectral range in combination with the temporal change information, wherein the temporal change information is used to indicate feature changes between consecutive frames.
[0159] Optionally, the feature processing unit 504 is specifically used to perform a maximum pooling operation on the fused feature map to obtain a pooled feature map; Flatten the pooled feature map into a one-dimensional vector to obtain the feature vector of the driving behavior to be identified.
[0160] Optionally, the classifier includes a first batch of normalization layers, a first fully connected layer, a Dropout layer, a second batch of normalization layers, an ELU activation function, and a second fully connected layer connected in sequence; the recognition unit 505 is specifically used to standardize the driving behavior feature vector to be identified through the first batch of normalization layers, and map it to the target dimension through the first fully connected layer to obtain the preferred driving behavior feature vector; The preferred driving behavior feature vector is randomly set to zero with a preset probability through the Dropout layer, standardized through the second batch normalization layer, activated through the ELU activation function, and then input into the second fully connected layer for mapping to obtain the driving distraction recognition result.
[0161] Optionally, the driving distraction recognition device further includes: A construction unit 506, used to construct a training set and a validation set; The training unit 507 is used to train the initial model through the training set, and perform back propagation optimization of the model parameters through a preset loss function to obtain a candidate model; The evaluation unit 508 is used to evaluate the candidate model through the verification set, and use the candidate model corresponding to the best model parameters as the driving distraction recognition model.
[0162] Optionally, the evaluation unit 508 is specifically used to input the training set into the initial model for driving distraction recognition to obtain a predicted label; calculate the loss value between the predicted label and the true label through a preset cross entropy loss function; update the model parameters through a preset optimization strategy and loss value to obtain a candidate model.
[0163] In this embodiment, the sample image data in the training set and the validation set are enriched by preprocessing to ensure that the model can adapt to different lighting, viewing angles and driving environments. The loss function is used for back propagation, and the efficiency of the training process is improved based on the optimization strategy. The distracted driving recognition model that has passed the evaluation is used for distracted driving recognition, which significantly improves the performance of distracted driving detection. The distracted driving recognition model that introduces deep separable convolution and multi-spectral attention mechanism significantly improves the performance of distracted driving detection, especially a good balance between efficient real-time and high-precision detection. The traditional distracted driving recognition solution relies on manually designed local feature information such as edges and corners, and is insufficient in capturing the complex behavioral characteristics of drivers, especially small changes in driving behavior. The shallow neural network is difficult to deeply understand the deep semantic information in the image data. In this embodiment, a two-layer cascaded DSC structure is used in the deep separable convolution. The first layer of the DSC module reduces computational redundancy and improves expression ability through channel transformation. The second layer of the DSC extracts local features to improve the spatial perception ability of the network. The deep separable convolutional layer of this embodiment has stronger nonlinear expression ability while still maintaining a low amount of computation, and is particularly suitable for distracted driving detection, a task that requires high fine-grained features. By optimizing the object of MSA, by performing spectral division on the initial driving behavior feature map, and combining global average pooling, global maximum pooling and attention mechanism to weightedly fuse each spectral feature map, the model can more finely capture the driver's behavioral characteristics in different spectral ranges, especially those small but critical behavioral changes, thereby significantly improving the accuracy and robustness of distracted driving detection. Furthermore, the MSA layer of this embodiment can further combine the time series change information to guide the attention allocation of multi-spectral feature maps, so that the model can pay more attention to the attention to instantaneous behavior and small behavior. Finally, the classifier significantly improves the classification performance and stability of the model through batch normalization, full connection, Dropout and activation functions and other optimization mechanisms.
[0164] above Figure 5 and Figure 6 The driving distraction identification device in the present application is described in detail from the perspective of modular functional entities, and the driving distraction identification device in the present application is described in detail from the perspective of hardware processing.
[0165] See also Figure 7 As shown, the driving distraction identification device includes a processor 700 and a memory 701. The memory 701 stores machine executable instructions that can be executed by the processor 700. The processor 700 executes the machine executable instructions to implement the above-mentioned driving distraction identification method.
[0166] further, Figure 7The illustrated driving distraction identification device further includes a bus 702 and a communication interface 703 , and the processor 700 , the communication interface 703 and the memory 701 are connected via the bus 702 .
[0167] The memory 701 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 703 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 702 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0168] The processor 700 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 700. The above processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The disclosed methods, steps and logic block diagrams in the embodiments of the present disclosure can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present disclosure can be directly embodied as a hardware decoding processor to be executed, or a combination of hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 701 , and the processor 700 reads the information in the memory 701 and completes the method steps of the above-mentioned embodiment in combination with its hardware.
[0169] The present application also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are executed on a computer, the computer executes the steps of the above-mentioned driving distraction identification method.
[0170] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0172] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying distracted driving, characterized in that: The method for identifying distracted driving includes: The original driver image data is input into a preset driving distraction recognition model, and preprocessed through the input layer to obtain candidate driver image data; Extracting features from the candidate driver image data through a depthwise separable convolutional layer to obtain an initial driving behavior feature map; Inputting the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and performing an attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; Inputting the fused feature map into the pooling layer and the feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; The driving behavior feature vector to be identified is input into a classifier to perform driving distraction behavior identification, and obtain a driving distraction identification result.
2. The method for identifying distracted driving according to claim 1, characterized in that: The depth-separable convolution layer includes a first depth convolution module and a second depth convolution module, wherein the spatial dimension corresponding to the convolution kernel in the first depth convolution module is smaller than the spatial dimension corresponding to the convolution kernel in the second depth convolution module; The step of extracting features from the candidate driver image data through a depth-separable convolutional layer to obtain an initial driving behavior feature map includes: Input the candidate driver image data into the first deep convolution module to perform spatial convolution operation on the input feature map through deep convolution, and integrate the cross-channel information through point-by-point convolution and then input it into the second deep convolution module; The second deep convolution module performs spatial convolution on the input feature map through deep convolution, and integrates cross-channel information through point-by-point convolution to obtain an initial driving behavior feature map.
3. The method for identifying distracted driving according to claim 1, characterized in that: The multi-spectral attention layer includes multiple spectrum modules, attention mechanism modules and fusion modules; The initial driving behavior feature map is input into the multi-spectral attention layer for spectrum division, and the feature map of each spectrum range is subjected to attention fusion operation to obtain a fused feature map, including: The initial driving behavior characteristic graph is divided according to the corresponding spectrum range by each spectrum module, and the global information of each spectrum characteristic graph is extracted to obtain the global characteristic graph corresponding to each spectrum range; Calculate the attention weight of the global feature map corresponding to each spectrum range through the attention mechanism module; Each of the spectral feature maps is multiplied by the corresponding attention weights through the fusion module and then fused to obtain a fused feature map.
4. The method for identifying distracted driving according to claim 3, characterized in that: The attention mechanism module further includes: The attention weights of the global feature maps corresponding to each spectral range are allocated in combination with the temporal variation information, wherein the temporal variation information is used to indicate feature changes between consecutive frames.
5. The method for identifying distracted driving according to claim 3 or 4, characterized in that: The multi-spectral attention layer also includes a temporal change recognition module; the output of the depth-separable convolution layer is an initial driving behavior feature map of continuous frames, and one frame of candidate driver image corresponds to one frame of initial driving behavior feature map; The method for identifying distracted driving also includes: The initial driving behavior feature map of continuous frames is passed through the temporal change recognition layer to extract the inter-frame change features, obtain the temporal change information, and input it into the attention mechanism module.
6. The method for identifying distracted driving according to claim 1, characterized in that: The classifier includes a first batch of normalized layers, a first fully connected layer, a Dropout layer, a second batch of normalized layers, an ELU activation function, and a second fully connected layer connected in sequence; The step of inputting the to-be-identified driving behavior feature vector into a classifier to perform driving distraction behavior identification to obtain a driving distraction identification result includes: The driving behavior feature vector to be identified is normalized by the first normalization layer, and mapped to the target dimension by the first fully connected layer to obtain a preferred driving behavior feature vector; The preferred driving behavior feature vector is randomly set to zero with a preset probability through the Dropout layer, standardized through the second batch normalization layer, activated through the ELU activation function, and then input into the second fully connected layer for mapping to obtain a driving distraction recognition result.
7. The method for identifying distracted driving according to claim 1, characterized in that: Before the original driver image data is input into the preset driving distraction recognition model and preprocessed through the input layer to obtain the candidate driver image data, the method further includes: Construct training and validation sets; The initial model is trained using the training set, and model parameters are optimized by back-propagation using a preset loss function to obtain a candidate model; The candidate model is evaluated through a validation set, and the candidate model corresponding to the best model parameters is used as a driving distraction recognition model.
8. A driving distraction identification device, characterized in that: The driving distraction identification device comprises: A preprocessing unit, used for inputting the original driver image data into a preset driving distraction recognition model, performing preprocessing through an input layer, and obtaining candidate driver image data; A feature extraction unit, used to extract features from the candidate driver image data through a depth-separable convolutional layer to obtain an initial driving behavior feature map; A frequency division attention unit, used for inputting the initial driving behavior feature map into the multi-spectral attention layer for spectrum division, and performing an attention fusion operation on the feature maps of each spectrum range to obtain a fused feature map; A feature processing unit, used for inputting the fused feature map into a pooling layer and a feature flattening layer for processing to obtain a feature vector of the driving behavior to be identified; The identification unit is used to input the driving behavior feature vector to be identified into a classifier to perform driving distraction behavior identification to obtain a driving distraction identification result.
9. A driving distraction identification device, characterized in that: The driving distraction identification device comprises: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instructions in the memory to enable the driving distraction identification device to execute the driving distraction identification method according to any one of claims 1 to 7.
10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are read and executed, the method for identifying distracted driving as described in any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Face image tampering passive detection method and device, terminal equipment and storage medium
CN115272240A
Forgery detection method based on feature enhancement and spectral analysis
CN115829909A
Intelligent cabin-oriented driver distraction behavior identification method and device
CN116883975A
Pedestrian crossing intention recognition method based on multi-source information fusion
CN117173663A
Automobile driver driving behavior identification method and device
CN119296085A
Cited By
Driver abnormal behavior identification method and system based on continuous learning
CN122090425A
A driver abnormal behavior recognition method and system based on continuous learning
CN122090425B