A lane-changing behavior monitoring method integrating in-vehicle driving video and GPS speed information

By integrating in-vehicle driving video with GPS speed information, combining visual feature coding and frequency domain self-attention module, lane-changing behavior monitoring is realized without coordinate system calibration, solving the problems of difficult equipment calibration, short battery life and environmental impact in the prior art, and it has strong robustness and continuous monitoring capabilities.

CN114120295BActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111441802.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-05-13
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

When monitoring the driver's lane change behavior, the prior art has problems such as difficulty in calibration of the coordinate system between the equipment and the vehicle, short battery life of the equipment, and environmental factors affecting the detection accuracy.

Method used

Using a method of integrating in-car driving video and GPS speed information, the steering wheel area image is collected through the in-car camera, combined with visual feature encoding and frequency domain self-attention module, to detect lane change behavior.

Benefits of technology

It realizes lane-changing behavior monitoring without coordinate system calibration and does not rely on wearable devices, and has strong robustness and continuous monitoring capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120295B_ABST
    Figure CN114120295B_ABST
Patent Text Reader

Abstract

The present invention discloses a lane-changing behavior monitoring method that integrates in-vehicle driving video and GPS speed information. The method obtains a steering wheel area image sequence collected by an in-vehicle camera, inputs it into a visual feature encoding module for feature extraction, and obtains visual features; and obtains GPS speed information of each frame in the image sequence, fuses it with the visual features, and obtains time-domain fusion features; finally, the time-domain fusion features are converted into frequency-domain fusion features, and lane-changing behavior detection is performed through a frequency-domain self-attention module, and the lane-changing behavior recognition result is output. The present invention does not rely on the precise calibration of the monitoring device and the vehicle coordinate system, nor does it require the driver to wear a wearable device, and has strong robustness to changes in the driver, vehicle, driving speed, and environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of automobile driving assistance technology, and in particular, relates to a lane change behavior monitoring method that integrates in-vehicle driving video and GPS speed information. Background Art

[0002] Frequent lane changes due to overtaking are considered to be one of the important factors causing traffic accidents. Monitoring drivers' lane changing behavior and providing real-time warnings can help improve traffic safety.

[0003] There are many technical solutions for driving monitoring in the prior art. For example, by reading the vehicle status from the vehicle bus, driving behavior is monitored. However, this solution requires the installation of a dedicated vehicle reader in the car. There are also technical solutions that use smartphones as a substitute method to read the IMU sensor of the mobile phone to measure the vehicle posture and monitor driving behavior, but this technical solution requires precise calibration between the coordinate systems of the mobile phone and the vehicle. There are also technical solutions that install the mobile phone on the steering wheel to measure its rotational acceleration to detect lane changes, but the disadvantage of this method is that the rotating mobile phone is difficult to charge to support continuous monitoring. There are also technical solutions that transplant the monitoring algorithm to the wearable device worn by the driver, but the limited battery of the wearable device is difficult to support continuous monitoring. At the same time, once the driver takes the hand wearing the device off the steering wheel, these methods will not be able to detect lane changes normally. There are also technical solutions that try to detect by using dashcam videos, but the quality of dashcam videos will be significantly affected by many environmental factors, such as blurred lane markings, bad weather, etc., resulting in deterioration of detection accuracy. Summary of the invention

[0004] The purpose of this application is to provide a lane-changing behavior monitoring method that integrates in-vehicle driving video and GPS speed information, extracts basic visual features that describe steering wheel rotation, and combines GPS speed data to infer the resulting lane-changing behavior.

[0005] In order to achieve the above purpose, the technical solution of this application is as follows:

[0006] A lane-changing behavior monitoring method integrating in-vehicle driving video and GPS speed information, comprising:

[0007] Obtain a steering wheel area image sequence captured by the in-vehicle camera, input it into the visual feature encoding module for feature extraction, and obtain visual features;

[0008] Obtain the GPS speed information of each frame in the image sequence, fuse it with the visual features, and obtain the time domain fusion features;

[0009] The time domain fusion features are converted into frequency domain fusion features, and the lane changing behavior is detected through the frequency domain self-attention module, and the lane changing behavior recognition results are output.

[0010] Furthermore, the visual feature encoding module adopts a mobile visual convolutional network MobileNetV3, and replaces the down-sampled inverse residual module Inverted Residual in the mobile visual convolutional network MobileNetV3 with an inverted pooling module Inverted Pool module, and the inverted pooling module Inverted Pool module includes:

[0011] The first point-wise convolutional layer, the maximum pooling layer, the multi-scale CBAM module and the second point-wise convolutional layer.

[0012] Furthermore, the multi-scale CBAM module includes a multi-scale channel attention module and a multi-scale spatial attention module, and the multi-scale CBAM module performs the following operations:

[0013] Input the first feature map before the maximum pooling layer and the second feature map after the maximum pooling layer into the multi-scale channel attention module to obtain a third feature map;

[0014] Multiplying the second feature map by the third feature map to obtain a fourth feature map;

[0015] Combine the fourth feature map with the first feature Figure 1 And send it to the multi-scale spatial attention module to obtain the fifth feature map;

[0016] Multiply the fourth feature map by the fifth feature map to obtain the output feature map.

[0017] Furthermore, the multi-scale channel attention module performs the following operations:

[0018] The first feature map and the second feature map are subjected to maximum and mean pooling respectively, connected through a multi-layer perceptron with shared weights, and finally channel attention is obtained through Sigmoid activation.

[0019] Furthermore, the multi-scale spatial attention module performs the following operations:

[0020] The link where the first feature map is located is convolved with a convolution with a stride of 2 to adjust the size of the output feature map to match the size of the convolution result with a stride of 1 of the fourth feature map, and the convolution results of the two feature maps are connected, and then convolved once and activated by Sigmoid to obtain spatial attention.

[0021] Furthermore, the lane-changing behavior detection is performed by the frequency-domain self-attention module, and the lane-changing behavior recognition result is output, including:

[0022] For each column of the frequency domain fusion feature, an MLP is used to extract the key frequency points, thereby obtaining a compressed fusion feature containing the key frequency points;

[0023] For each row of the frequency domain fusion feature, an MLP is used to extract the key feature points, thereby obtaining a compressed fusion feature containing the key feature points;

[0024] The compressed fusion features containing key frequency points and the compressed fusion features containing key feature points are concatenated, and then three vectors Q, K and V are obtained through three different MLPs, input into the self-attention module, and finally pass through another layer of MLP to obtain the classification result of whether the current image sequence contains lane changes.

[0025] The lane change behavior monitoring method proposed in this application, which integrates in-vehicle driving video and GPS speed information, has the advantages of being simple and reliable compared with existing methods. The technical solution of this application does not rely on the precise calibration of the monitoring equipment and the vehicle coordinate system, nor does it require the driver to wear a wearable device, and is highly robust to changes in the driver, vehicle, driving speed and environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flow chart of the lane change behavior monitoring method for this application;

[0027] Figure 2 This is a schematic diagram of the overall network structure of this application;

[0028] Figure 3 It is a schematic diagram of the structure of the visual feature encoding module;

[0029] Figure 4 It is a schematic diagram of the multi-scale CBAM module structure;

[0030] Figure 5 Schematic diagram of the multi-scale channel attention module structure;

[0031] Figure 6 Schematic diagram of the multi-scale spatial attention module structure;

[0032] Figure 7 Schematic diagram of the frequency domain self-attention module structure. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0034] This application monitors the driver's lane-changing behavior by using an in-car camera to monitor the rotation of the steering wheel. As a vision-based method, this application does not require coordinate system calibration, nor does it require the driver to wear any monitoring equipment. Existing visual human behavior recognition methods are mainly used to detect large movements. During the lane change process, the steering wheel rotation driven by the driver's arm movement may be very slight, and existing methods are difficult to meet application requirements.

[0035] In one embodiment, the present application provides a lane change behavior monitoring method that integrates in-vehicle driving video and GPS speed information, such as Figure 1 As shown, including:

[0036] Step S1, obtaining a steering wheel area image sequence captured by an in-vehicle camera, and inputting the image sequence into a visual feature encoding module for feature extraction to obtain visual features.

[0037] The present application sets up an in-car camera in the car to collect video images containing the steering wheel area, and performs frame sampling at fixed intervals to reduce the amount of input data. Then the steering wheel detection is performed on the first frame, and each frame is cropped with the obtained bounding box to only contain the steering wheel area. Finally, the video image is converted into a processed image sequence using a sliding window filter to facilitate feature extraction and lane change classification on each video segment.

[0038] The overall network structure of this application is as follows Figure 2 As shown, the visual feature encoding module performs feature extraction, and the frequency domain self-attention module is used for the final lane change result recognition output.

[0039] The visual feature encoding module is developed based on the popular mobile visual convolutional network MobileNetV3, which takes the image sequence as input and uses the calculation result of the second-to-last layer of the network as the visual feature output. Figure 3 As shown, unlike the existing MobileNetV3, the down-sampled inverse residual module (Inverted Residual module) in MobileNetV3 is replaced by an inverted pooling module (Inverted Pool module). The Inverted Pool module of this embodiment includes a first point-by-point convolution layer (1*1Pointwise conv), a maximum pooling layer (2*2MaxPool), a multi-scale CBAM module, and a second point-by-point convolution layer (1*1Pointwise conv).

[0040] Specifically, in the Inverted Pool module of this embodiment, the Depthwise convolution in the Inverted Residual module is replaced with the maximum pooling, which reduces the rotation invariance of the model and improves the model's attention to the rotation of the steering wheel. A multi-scale CBAM (Convolutional Block Attention Module) block is introduced between the maximum pooling layer and the second Pointwise convolution.

[0041] like Figure 4 As shown, the multi-scale CBAM module of the present application includes a multi-scale channel attention module and a multi-scale spatial attention module. The multi-scale CBAM module has two inputs, one of which is the input feature map of the maximum pooling layer (the first feature map), that is, the feature map before downsampling; the maximum pooling layer downsamples the first feature map, and the output of the maximum pooling layer is the second feature map, that is, the feature map after downsampling.

[0042] In this embodiment, the multi-scale CBAM module performs the following operations:

[0043] Input the first feature map before the maximum pooling layer and the second feature map after the maximum pooling layer into the multi-scale channel attention module to obtain a third feature map;

[0044] Multiplying the second feature map by the third feature map to obtain a fourth feature map;

[0045] Combine the fourth feature map with the first feature Figure 1 And send it to the multi-scale spatial attention module to obtain the fifth feature map;

[0046] Multiply the fourth feature map by the fifth feature map to obtain the output feature map.

[0047] This embodiment introduces the feature map before downsampling during the encoding process as an additional input to the channel attention and spatial attention modules to further improve the model's sensitivity to steering wheel rotation. Figure 5 and Figure 6As shown in the figure, the multi-scale channel attention module performs maximum and mean pooling on the first feature map and the second feature map respectively, connects them through a multi-layer perceptron (MLP) with shared weights, and finally obtains channel attention through Sigmoid activation. The multi-scale spatial attention module uses a convolution with a step size of 2 to convolve the link where the first feature map is located after performing channel-level mean and maximum pooling to adjust the size of the output feature map to match the size of the convolution result with a step size of 1 of the fourth feature map, and connects the convolution results of the two feature maps, and then performs another convolution and obtains spatial attention through Sigmoid activation. The convolution result with a step size of 1 of the fourth feature map is the convolution result obtained by convolving the fourth feature map with a step size of 1.

[0048] Step S2: Obtain GPS speed information of each frame in the image sequence, and fuse it with the visual features to obtain time domain fusion features.

[0049] This embodiment obtains GPS speed information of each frame in the image sequence and fuses it with the visual features. Assume that each image sequence contains L frames, and the visual feature encoding result of each frame is N-dimensional. Then the final time domain fusion feature is F=L×(N+1)-dimensional.

[0050] Step S3: Convert the time domain fusion features into frequency domain fusion features, perform lane change behavior detection through the frequency domain self-attention module, and output the lane change behavior recognition result.

[0051] In this embodiment, the time domain fusion feature F is converted into the frequency domain fusion feature F P , to better characterize the characteristic signal. Then, the lane changing behavior is detected through the frequency domain self-attention module.

[0052] like Figure 7 As shown in Figure 1, the lane-changing behavior is detected through the frequency-domain self-attention module, which includes the following processes:

[0053] (1) For each column of the frequency domain fusion feature, an MLP is used to extract the key frequency points, thereby obtaining a compressed fusion feature containing the key frequency points.

[0054] This step refines the frequency dimension features and P For each column of , an MLP is used to extract the key frequency points, thereby obtaining the compressed fusion features containing the key frequency points Its dimension is L′×(N+1), L′<L.

[0055] (2) For each row of the frequency domain fusion feature, an MLP is used to extract the key feature points, thereby obtaining a compressed fusion feature containing the key feature points.

[0056] This step refines the feature dimension and refines F P For each row of , an MLP is used to extract key feature points, thereby obtaining a compressed fusion feature containing key feature points Its dimension is L×N′, N′<N+1.

[0057] (3) The compressed fusion features containing key frequency points and the compressed fusion features containing key feature points are concatenated, and then three vectors Q, K, and V are obtained through three different MLPs. The three vectors are input into the self-attention module and finally passed through another layer of MLP to obtain the classification result of whether the current image sequence contains lane changes.

[0058] Will and The three vectors Q, K and V are concatenated and obtained through three different MLPs, which are input into the self-attention module and finally pass through another layer of MLP to obtain the classification result of whether the current image sequence contains lane changes.

[0059] It should be noted that the above-mentioned MLP layer and self-attention module are relatively mature technologies in this technical field and will not be described in detail here.

[0060] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A lane-changing behavior monitoring method integrating in-vehicle driving video and GPS speed information, characterized in that: The lane change behavior monitoring method integrating in-vehicle driving video and GPS speed information includes: Obtain a steering wheel area image sequence captured by the in-vehicle camera, input it into the visual feature encoding module for feature extraction, and obtain visual features; Obtain the GPS speed information of each frame in the image sequence, fuse it with the visual features, and obtain the time domain fusion features; The time domain fusion features are converted into frequency domain fusion features, and the lane change behavior is detected through the frequency domain self-attention module, and the lane change behavior recognition result is output; The visual feature encoding module adopts a mobile visual convolutional network MobileNetV3, and replaces the down-sampled inverse residual module Inverted Residual in the mobile visual convolutional network MobileNetV3 with an inverted pooling module Inverted Pool module, and the inverted pooling module Inverted Pool module includes: The first point-wise convolution layer, the maximum pooling layer, the multi-scale CBAM module and the second point-wise convolution layer; The multi-scale CBAM module includes a multi-scale channel attention module and a multi-scale spatial attention module. The multi-scale CBAM module performs the following operations: Input the first feature map before the maximum pooling layer and the second feature map after the maximum pooling layer into the multi-scale channel attention module to obtain a third feature map; Multiplying the second feature map by the third feature map to obtain a fourth feature map; Send the fourth feature map and the first feature map together to the multi-scale spatial attention module to obtain the fifth feature map; Multiply the fourth feature map by the fifth feature map to obtain the output feature map.

2. The lane change behavior monitoring method integrating in-vehicle driving video and GPS speed information according to claim 1 is characterized in that: The multi-scale channel attention module performs the following operations: The first feature map and the second feature map are subjected to maximum and mean pooling respectively, connected through a multi-layer perceptron with shared weights, and finally channel attention is obtained through Sigmoid activation.

3. The lane change behavior monitoring method integrating in-vehicle driving video and GPS speed information according to claim 1 is characterized in that: The multi-scale spatial attention module performs the following operations: The link where the first feature map is located is convolved with a convolution with a stride of 2 to adjust the size of the output feature map to match the size of the convolution result with a stride of 1 of the fourth feature map, and the convolution results of the two feature maps are connected, and then convolved once and activated by Sigmoid to obtain spatial attention.

4. The lane change behavior monitoring method integrating in-vehicle driving video and GPS speed information according to claim 1 is characterized in that: The lane-changing behavior detection is performed by the frequency-domain self-attention module, and the lane-changing behavior recognition result is output, including: A multi-layer perceptron (MLP) is used to extract key frequency points from each column of frequency domain fusion features, thereby obtaining compressed fusion features containing key frequency points; A multi-layer perceptron MLP is used to extract key feature points from each row of the frequency domain fusion feature, thereby obtaining a compressed fusion feature containing the key feature points; The compressed fusion features containing key frequency points and the compressed fusion features containing key feature points are concatenated, and then three vectors Q, K and V are obtained through three different multi-layer perceptrons MLP, which are input into the self-attention module and finally passed through another layer of multi-layer perceptron MLP to obtain the classification result of whether the current image sequence contains lane changes.

Citation Information

Patent Citations

  • Driving behavior analysis and processing method and device, equipment and storage medium

    CN110765807A

  • Converged network lane line detection method based on attention mechanism and terminal equipment

    CN111950467A